驱动包是否有问题
驱动包是否有问题
./metax-driver-3.8.0.10-rpm-aarch64.run -- --dkms -f
Verifying archive integrity... 100% MD5 checksums are OK. All good.
Uncompressing MetaX driver installer 100%
uninstalling metax-linux-dkms-3.8.23-1.aarch64
Uninstall of metax-linux module (version 3.8.23) beginning:
Module metax-linux-3.8.23 for kernel 6.6.0-25.02.2510.013.uos25.aarch64 (aarch64).
Before uninstall, this module version was ACTIVE on this kernel.
metax.ko.xz:
- Uninstallation
- Deleting from: /lib/modules/6.6.0-25.02.2510.013.uos25.aarch64/extra/
- Original module
- No original module was found for this module on this kernel.
- Use the dkms install command to reinstall any previous module version.
Running the post_remove script:
depmod....
Deleting module metax-linux-3.8.23 completely from the DKMS tree.
uninstalling mxsmt-3.7.0-1.ky10.aarch64
uninstalling mxfw-3.7.0-1.noarch
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_train.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_infer.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_train.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_infer.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_train.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_infer.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn.so.8 不是符号链接
上次元数据过期检查:2:42:44 前,执行于 2026年07月17日 星期五 07时19分35秒。
软件包 dkms-3.0.12-1.uos25.02.noarch 已安装。
软件包 autoconf-2.71-10.uos25.noarch 已安装。
软件包 automake-1.16.5-9.uos25.noarch 已安装。
软件包 elfutils-libelf-devel-0.190-11.uos25.aarch64 已安装。
依赖关系解决。
无需任何处理。
完毕!
rpm/mxfw_.rpm
rpm/mxsmt_.rpm
dkms_rpm/metax-linux-dkms[-_]*.rpm
installing packages under /opt/mxdriver
installing rpm/mxfw_3.8.0-1.noarch.rpm
Verifying... ################################# [100%]
准备中... ################################# [100%]
deepin_rpm_verify: HC disabled, allowing mxfw
正在升级/安装...
1:mxfw-3.8.0-1 ################################# [100%]
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_train.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_infer.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_train.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_infer.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_train.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_infer.so.8 不是符号链接
/sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn.so.8 不是符号链接
deepin_rpm_verify: HC disabled, allowing mxfw
installing rpm/mxsmt_3.8.0-1.aarch64.rpm
Verifying... ################################# [100%]
准备中... ################################# [100%]
deepin_rpm_verify: HC disabled, allowing mxsmt
正在升级/安装...
1:mxsmt-3.8.0-1.ky10 ################################# [100%]
ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_train.so.8 不是符号链接
ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_infer.so.8 不是符号链接
ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_train.so.8 不是符号链接
ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_infer.so.8 不是符号链接
ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_train.so.8 不是符号链接
ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_infer.so.8 不是符号链接
ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn.so.8 不是符号链接
deepin_rpm_verify: HC disabled, allowing mxsmt
installing dkms_rpm/metax-linux-dkms-3.9.10-1.aarch64.rpm
Verifying... ################################# [100%]
准备中... ################################# [100%]
deepin_rpm_verify: HC disabled, allowing metax-linux-dkms
remove old metax.ko from initramfs
正在升级/安装...
1:metax-linux-dkms-3.9.10-1 ################################# [100%]
Loading new metax-linux-3.9.10 DKMS files...
Building for 6.6.0-25.02.2510.013.uos25.aarch64
Building initial module for 6.6.0-25.02.2510.013.uos25.aarch64
autoreconf: export WARNINGS=
autoreconf: Entering directory '.'
autoreconf: configure.ac: not using Gettext
autoreconf: running: aclocal --force
autoreconf: configure.ac: tracing
autoreconf: configure.ac: not using Libtool
autoreconf: configure.ac: not using Intltool
autoreconf: configure.ac: not using Gtkdoc
autoreconf: running: /usr/bin/autoconf --force
configure.ac:23: warning: The macro `AC_HELP_STRING' is obsolete.
configure.ac:23: You should run autoupdate.
./lib/autoconf/general.m4:204: AC_HELP_STRING is expanded from...
./lib/autoconf/general.m4:1534: AC_ARG_ENABLE is expanded from...
m4/config.m4:14: AC_METAX_CONFIG is expanded from...
configure.ac:23: the top level
autoreconf: running: /usr/bin/autoheader --force
autoreconf: configure.ac: not using Automake
autoreconf: 'config/install-sh' is updated
autoreconf: Leaving directory '.'
fgrep: warning: fgrep is obsolescent; using grep -F
Error! Bad return status for module build on kernel: 6.6.0-25.02.2510.013.uos25.aarch64 (aarch64)
Consult /var/lib/dkms/metax-linux/3.9.10/build/make.log for more information.
警告:%post(metax-linux-dkms-3.9.10-1.aarch64) 脚本执行失败,退出状态码为 10
deepin_rpm_verify: HC disabled, allowing metax-linux-dkms
沐曦C500 卡使用最新版本的SDK3.7.1版本容器镜像无论是sglang还是vllm 引擎启动大模型都非常吃内存。只运行32B大模型宿主机128G的内存都会报内存溢出无法启动服务。尤其是SGLang 对内存的消耗更高。
之前发布的帖子让通过降级sdk到3.5版本才解决问题。
原贴地址:developer.metax-tech.com/forum/t/sglang-duo-qia-bing-xing-qi-dong-da-mo-xing-fu-wu-bu-guan-mo-xing-da-xiao-duo-xiao-du-bao-nei-cun-yi-chu/617/#post-2966
0.5.9 版本可以正常使用
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 16Gi 103Gi 25Mi 6.2Gi 108Gi
Swap: 4.0Gi 548Mi 3.5Gi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 19Gi 99Gi 25Mi 6.8Gi 104Gi
Swap: 4.0Gi 548Mi 3.5Gi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 21Gi 97Gi 25Mi 6.9Gi 103Gi
Swap: 4.0Gi 548Mi 3.5Gi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 22Gi 96Gi 25Mi 7.1Gi 102Gi
Swap: 4.0Gi 548Mi 3.5Gi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 23Gi 94Gi 25Mi 7.1Gi 100Gi
Swap: 4.0Gi 548Mi 3.5Gi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 24Gi 93Gi 25Mi 7.4Gi 100Gi
Swap: 4.0Gi 548Mi 3.5Gi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 106Gi 16Gi 20Mi 2.0Gi 17Gi
Swap: 4.0Gi 3.8Gi 197Mi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 112Gi 8.2Gi 27Mi 4.9Gi 12Gi
Swap: 4.0Gi 3.8Gi 200Mi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 111Gi 7.5Gi 27Mi 7.0Gi 13Gi
Swap: 4.0Gi 3.8Gi 247Mi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 35Gi 83Gi 23Mi 7.1Gi 89Gi
Swap: 4.0Gi 3.2Gi 847Mi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 37Gi 81Gi 23Mi 7.4Gi 87Gi
Swap: 4.0Gi 3.0Gi 998Mi
[root@muxi models]# free -h
total used free shared buff/cache available
Mem: 124Gi 4.2Gi 112Gi 23Mi 8.9Gi 120Gi
Swap: 4.0Gi 1.3Gi 2.7Gi 上述是在启动sglang服务过程中 内存的变化 一直在增长 知道内存溢出后然后全部释放
docker run -d \
--network host \
--ipc=host \
--name=sglang-mx \
--shm-size 100G \
--ulimit memlock=-1 \
--ulimit nofile=1024000 \
--device=/dev/dri \
--device=/dev/mxcd \
--device=/dev/mem \
--group-add video \
--privileged \
--security-opt seccomp=unconfined \
--security-opt apparmor=unconfined \
-v /data/models/:/models:ro \
cr.metax-tech.com/public-ai-release/maca/sglang:0.5.11-maca.ai3.7.1.107-torch2.8-py310-kylinv11-amd64 \
sleep infinity 添加了特权参数--privileged 依然是内存溢出失败
cd /opt/maca/samples/mccl_tests/perf
[root@muxi perf]# ls
cluster.sh dragonfly.sh function mccl_perf mccl.sh mxccl_perf nccl_perf README.md xccl.sh
[root@muxi perf]# bash mccl.sh 2
The test is all_reduce_perf, the maca version is /opt/maca-3.7.2
main_process = 211548
===============================
1024 512 bfloat16 sum -1 15.48 0.07 0.07 0 15.45 0.07 0.07 0
2048 1024 bfloat16 sum -1 16.25 0.13 0.13 0 16.14 0.13 0.13 0
4096 2048 bfloat16 sum -1 17.20 0.24 0.24 0 17.00 0.24 0.24 0
8192 4096 bfloat16 sum -1 17.35 0.47 0.47 0 17.04 0.48 0.48 0
16384 8192 bfloat16 sum -1 17.96 0.91 0.91 0 17.59 0.93 0.93 0
32768 16384 bfloat16 sum -1 17.98 1.82 1.82 0 17.95 1.83 1.83 0
65536 32768 bfloat16 sum -1 21.06 3.11 3.11 0 21.03 3.12 3.12 0
131072 65536 bfloat16 sum -1 27.76 4.72 4.72 0 27.74 4.72 4.72 0
262144 131072 bfloat16 sum -1 34.12 7.68 7.68 0 33.89 7.74 7.74 0
524288 262144 bfloat16 sum -1 46.10 11.37 11.37 0 46.17 11.35 11.35 0
1048576 524288 bfloat16 sum -1 70.71 14.83 14.83 0 70.77 14.82 14.82 0
2097152 1048576 bfloat16 sum -1 121.17 17.31 17.31 0 121.02 17.33 17.33 0
4194304 2097152 bfloat16 sum -1 225.56 18.59 18.59 0 225.20 18.62 18.62 0
8388608 4194304 bfloat16 sum -1 436.27 19.23 19.23 0 436.00 19.24 19.24 0
16777216 8388608 bfloat16 sum -1 856.35 19.59 19.59 0 856.41 19.59 19.59 0
33554432 16777216 bfloat16 sum -1 1697.19 19.77 19.77 0 1696.73 19.78 19.78 0
67108864 33554432 bfloat16 sum -1 3378.11 19.87 19.87 0 3377.85 19.87 19.87 0
134217728 67108864 bfloat16 sum -1 6611.09 20.30 20.30 0 6611.62 20.30 20.30 0
268435456 134217728 bfloat16 sum -1 13191.31 20.35 20.35 0 13191.46 20.35 20.35 0
536870912 268435456 bfloat16 sum -1 26340.36 20.38 20.38 0 26337.84 20.38 20.38 0
1073741824 536870912 bfloat16 sum -1 52597.00 20.41 20.41 0 52566.05 20.43 20.43 0
[root@muxi perf]#
现在TP=1 是可以的,但是我现在有是想在多卡情况下运行更大的模型的,但是总是报OOM 内存溢出才会选择用小模型验证的
单卡TP=1 可以
dmesg -T | grep -i err
[二 7月 7 18:39:16 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 102
[二 7月 7 18:39:45 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 103
[二 7月 7 18:40:26 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 104
[二 7月 7 18:40:42 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 105
[二 7月 7 18:40:46 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 106
[二 7月 7 18:40:53 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 107
[二 7月 7 18:41:11 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 108
[二 7月 7 18:41:30 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 109
[二 7月 7 18:41:40 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 110
[二 7月 7 18:42:58 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 111
[二 7月 7 19:02:49 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 112
[二 7月 7 19:03:33 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 113
[二 7月 7 19:03:39 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 114
[二 7月 7 19:04:04 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 115
[二 7月 7 19:04:12 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 116
[二 7月 7 19:04:19 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 117
[二 7月 7 19:04:54 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 118
[二 7月 7 19:05:50 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 119
[二 7月 7 19:06:01 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 120
[二 7月 7 19:06:05 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 121
[二 7月 7 19:06:11 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 122
[二 7月 7 19:06:23 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 123
[二 7月 7 19:06:25 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 124
[二 7月 7 19:06:29 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 125
[二 7月 7 19:06:58 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 126
[二 7月 7 19:07:08 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 127
[二 7月 7 19:08:19 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 128
[二 7月 7 19:16:17 2026] METAX.MC.ERROR failed to get user pages, -4
[二 7月 7 19:16:17 2026] METAX.B3800.D0.MC.ERROR init_user_pages failed, -4
[二 7月 7 19:16:17 2026] MXCD.IOCTL.ERROR alloc memory failed, -4
[二 7月 7 19:16:18 2026] METAX.MC.ERROR failed to get user pages, -4
[二 7月 7 19:16:18 2026] METAX.B3800.D0.MC.ERROR init_user_pages failed, -4
[二 7月 7 19:16:18 2026] MXCD.IOCTL.ERROR alloc memory failed, -4
[二 7月 7 19:48:11 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 129
[二 7月 7 19:48:19 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 130
[二 7月 7 19:50:38 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 131
[二 7月 7 19:50:50 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 132
[二 7月 7 19:51:05 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 133
[二 7月 7 19:51:19 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 134
[二 7月 7 19:54:55 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 135
[二 7月 7 19:55:01 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 136
[二 7月 7 19:55:29 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 137
[二 7月 7 19:56:11 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 138
[二 7月 7 19:56:13 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 139
[二 7月 7 19:56:34 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 140
[二 7月 7 19:57:58 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 141
[二 7月 7 19:58:15 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 142
[二 7月 7 19:59:22 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 143
[二 7月 7 20:00:24 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 144
[二 7月 7 20:00:28 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 145
[二 7月 7 20:00:51 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 146
[二 7月 7 20:01:28 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 147
[二 7月 7 20:02:09 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 148
[二 7月 7 20:02:23 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 149
[二 7月 7 20:03:10 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 150
[二 7月 7 20:04:34 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 151
[二 7月 7 20:05:35 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 152
[二 7月 7 20:05:41 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 153
[二 7月 7 20:06:06 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 154
[二 7月 7 20:06:08 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 155
[二 7月 7 20:06:22 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 156
[二 7月 7 20:06:59 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 157
[二 7月 7 20:07:28 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 158
[二 7月 7 20:08:09 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 159
[二 7月 7 20:08:29 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 160
[二 7月 7 20:08:31 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 161
[二 7月 7 20:09:45 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 162
[二 7月 7 20:09:49 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 163
[二 7月 7 20:11:05 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 164
[二 7月 7 20:12:10 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 165
[二 7月 7 20:12:39 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 166
[二 7月 7 20:14:00 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 167
[二 7月 7 20:14:08 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 168
[二 7月 7 20:15:12 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 169
[二 7月 7 20:16:13 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 170
[二 7月 7 20:16:46 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 171
[二 7月 7 20:18:16 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 172
[二 7月 7 20:18:20 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 173
[二 7月 7 20:18:37 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 174
[二 7月 7 20:19:04 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 175
[二 7月 7 20:19:55 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 176
[二 7月 7 20:20:36 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 177
[二 7月 7 20:20:42 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 178
[二 7月 7 20:20:44 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 179
[二 7月 7 20:21:04 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 180
[二 7月 7 20:21:16 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 181
[二 7月 7 20:22:03 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 182
[二 7月 7 20:22:18 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 183
[二 7月 7 20:23:34 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 184
[二 7月 7 20:23:42 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 185
[二 7月 7 20:24:48 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 186
[二 7月 7 20:25:23 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 187
[二 7月 7 20:26:02 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 188
[二 7月 7 20:26:14 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 189
[二 7月 7 20:27:01 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 190
[二 7月 7 20:27:09 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 191
[二 7月 7 20:27:28 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 192
[二 7月 7 20:27:36 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 193
[二 7月 7 20:27:58 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 194
[二 7月 7 20:28:21 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 195
[二 7月 7 20:28:33 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 196
[二 7月 7 20:28:35 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 197
[二 7月 7 20:28:47 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 198
[二 7月 7 20:29:53 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 199
[二 7月 7 20:30:30 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 200
[二 7月 7 20:31:33 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 201
[二 7月 7 20:31:47 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 202
[二 7月 7 20:32:14 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 203
[二 7月 7 20:33:28 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 204
[二 7月 7 20:35:01 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 205
[二 7月 7 20:35:18 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 206
[二 7月 7 20:37:31 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 207
[二 7月 7 20:37:43 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 208
[二 7月 7 20:39:09 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 209
[二 7月 7 20:39:59 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 210
[二 7月 7 20:40:01 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 211
[二 7月 7 20:41:21 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 212
[二 7月 7 20:43:38 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 213
[三 7月 8 09:00:56 2026] perf: interrupt took too long (2506 > 2500), lowering kernel.perf_event_max_sample_rate to 79800
[三 7月 8 09:21:56 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 214
[三 7月 8 09:22:08 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 215
[三 7月 8 09:33:16 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 216
[三 7月 8 09:46:05 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 217
[三 7月 8 09:46:21 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 218
[三 7月 8 09:47:21 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 219
[三 7月 8 10:10:36 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 220
[三 7月 8 10:13:10 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 221
[三 7月 8 10:20:35 2026] METAX.MC.ERROR failed to get user pages, -4
[三 7月 8 10:20:35 2026] METAX.B3800.D0.MC.ERROR init_user_pages failed, -4
[三 7月 8 10:20:35 2026] MXCD.IOCTL.ERROR alloc memory failed, -4
[三 7月 8 10:21:05 2026] METAX.B9800.D0.PCI.INFO correctable AER [ 0] Receiver Error, count 222
[三 7月 8 10:32:58 2026] ? common_interrupt+0xf/0xa0
[三 7月 8 10:32:58 2026] METAX.MC.ERROR failed to get user pages, -4
[三 7月 8 10:32:58 2026] METAX.B3800.D0.MC.ERROR init_user_pages failed, -4
[三 7月 8 10:32:58 2026] MXCD.IOCTL.ERROR alloc memory failed, -4
[三 7月 8 10:33:00 2026] METAX.MC.ERROR failed to get user pages, -4
[三 7月 8 10:33:00 2026] METAX.B3800.D0.MC.ERROR init_user_pages failed, -4
[三 7月 8 10:33:00 2026] MXCD.IOCTL.ERROR alloc memory failed, -4
[三 7月 8 10:35:14 2026] METAX.MC.ERROR failed to get user pages, -4
[三 7月 8 10:35:14 2026] METAX.B3800.D0.MC.ERROR init_user_pages failed, -4
[三 7月 8 10:35:14 2026] MXCD.IOCTL.ERROR alloc memory failed, -4
使用python -m sglang.launch_server --model-path /models/DeepSeek-R1-Distill-Qwen-7B \
--host 0.0.0.0 \
--port 9100 \
--served-model-name deepseek \
--tensor-parallel-size 2 \
--mem-fraction-static 0.7 \
--trust-remote-code \
--attention-backend flashinfer \
--enable-metrics \
--disable-radix-cache \
--disable-chunked-prefix-cache \
--disable-cuda-graph 启动命令 仍然报内存溢出错误日志如下:
python -m sglang.launch_server --model-path /models/DeepSeek-R1-Distill-Qwen-7B \
--host 0.0.0.0 \
--port 9100 \
--served-model-name deepseek \
--tensor-parallel-size 2 \
--mem-fraction-static 0.7 \
--trust-remote-code \
--attention-backend flashinfer \
--enable-metrics \
--disable-radix-cache \
--disable-chunked-prefix-cache \
--disable-cuda-graph
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.4'
Build_Maca_Version = '3.7.1.5'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '4d9c16e'
Vllm Op Version = 0.19.0
SGlang Op Version = 0.5.10
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.7.
/opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/sglang/launch_server.py:60: UserWarning: 'python -m sglang.launch_server' is still supported, but 'sglang serve' is the recommended entrypoint.
Example: sglang serve --model-path <model> [options]
warnings.warn(
[2026-07-08 10:34:21] Cuda graph max bs is adjusted to 128.
[2026-07-08 10:34:22] flashinfer.comm is not available, falling back to standard implementation
[2026-07-08 10:34:23] server_args=ServerArgs(model_path='/models/DeepSeek-R1-Distill-Qwen-7B', tokenizer_path='/models/DeepSeek-R1-Distill-Qwen-7B', tokenizer_mode='auto', tokenizer_backend='huggingface', tokenizer_worker_num=1, detokenizer_worker_num=1, skip_tokenizer_init=False, load_format='auto', model_loader_extra_config='{}', trust_remote_code=True, context_length=None, is_embedding=False, enable_multimodal=None, revision=None, model_impl='auto', host='0.0.0.0', port=9100, fastapi_root_path='', grpc_mode=False, skip_server_warmup=False, warmups=None, nccl_port=None, checkpoint_engine_wait_weights_before_ready=False, ssl_keyfile=None, ssl_certfile=None, ssl_ca_certs=None, ssl_keyfile_password=None, enable_ssl_refresh=False, enable_http2=False, dtype='auto', quantization=None, quantization_param_path=None, kv_cache_dtype='auto', enable_fp32_lm_head=False, modelopt_quant=None, modelopt_checkpoint_restore_path=None, modelopt_checkpoint_save_path=None, modelopt_export_path=None, quantize_and_serve=False, rl_quant_profile=None, mem_fraction_static=0.7, max_running_requests=None, max_queued_requests=None, max_total_tokens=None, chunked_prefill_size=8200, enable_dynamic_chunking=False, max_prefill_tokens=16384, prefill_max_requests=None, schedule_policy='fcfs', enable_priority_scheduling=False, disable_priority_preemption=False, default_priority_value=None, abort_on_priority_when_disabled=False, schedule_low_priority_values_first=False, priority_scheduling_preemption_threshold=10, schedule_conservativeness=1.0, page_size=1, swa_full_tokens_ratio=0.8, disable_hybrid_swa_memory=False, radix_eviction_policy='lru', enable_prefill_delayer=False, prefill_delayer_max_delay_passes=30, prefill_delayer_token_usage_low_watermark=None, prefill_delayer_forward_passes_buckets=None, prefill_delayer_wait_seconds_buckets=None, device='cuda', tp_size=2, pp_size=1, pp_max_micro_batch_size=None, pp_async_batch_depth=0, stream_interval=1, batch_notify_size=16, stream_response_default_include_usage=False, incremental_streaming_output=False, enable_streaming_session=False, random_seed=147146915, constrained_json_whitespace_pattern=None, constrained_json_disable_any_whitespace=False, watchdog_timeout=300, soft_watchdog_timeout=None, dist_timeout=None, download_dir=None, model_checksum=None, base_gpu_id=0, gpu_id_step=1, sleep_on_idle=False, use_ray=False, custom_sigquit_handler=None, log_level='info', log_level_http=None, log_requests=False, log_requests_level=2, log_requests_format='text', log_requests_target=None, uvicorn_access_log_exclude_prefixes=[], crash_dump_folder=None, show_time_cost=False, enable_metrics=True, grpc_http_sidecar_port=None, enable_mfu_metrics=False, enable_metrics_for_all_schedulers=False, tokenizer_metrics_custom_labels_header='x-custom-labels', tokenizer_metrics_allowed_custom_labels=None, extra_metric_labels=None, bucket_time_to_first_token=None, bucket_inter_token_latency=None, bucket_e2e_request_latency=None, prompt_tokens_buckets=None, generation_tokens_buckets=None, gc_warning_threshold_secs=0.0, decode_log_interval=40, enable_request_time_stats_logging=False, kv_events_config=None, enable_trace=False, otlp_traces_endpoint='localhost:4317', export_metrics_to_file=False, export_metrics_to_file_dir=None, api_key=None, admin_api_key=None, served_model_name='deepseek', weight_version='default', chat_template=None, hf_chat_template_name=None, completion_template=None, file_storage_path='sglang_storage', enable_cache_report=False, reasoning_parser=None, strip_thinking_cache=False, tool_call_parser=None, tool_server=None, sampling_defaults='model', dp_size=1, load_balance_method='round_robin', attn_cp_size=1, moe_dp_size=1, dist_init_addr=None, nnodes=1, node_rank=0, json_model_override_args='{}', preferred_sampling_params=None, enable_lora=None, enable_lora_overlap_loading=None, max_lora_rank=None, lora_target_modules=None, lora_paths=None, max_loaded_loras=None, max_loras_per_batch=8, lora_eviction_policy='lru', lora_backend='csgmv', max_lora_chunk_size=16, experts_shared_outer_loras=None, lora_use_virtual_experts=False, lora_strict_loading=False, attention_backend='flashinfer', decode_attention_backend=None, prefill_attention_backend=None, sampling_backend='flashinfer', grammar_backend='xgrammar', mm_attention_backend=None, fp8_gemm_runner_backend='auto', fp4_gemm_runner_backend='auto', nsa_prefill_backend='flashmla_sparse', nsa_decode_backend='flashmla_kv', disable_flashinfer_autotune=False, mamba_backend='triton', speculative_algorithm=None, speculative_draft_model_path=None, speculative_draft_model_revision=None, speculative_draft_load_format=None, speculative_num_steps=None, speculative_eagle_topk=None, speculative_num_draft_tokens=None, speculative_dflash_block_size=None, speculative_dflash_draft_window_size=None, speculative_accept_threshold_single=1.0, speculative_accept_threshold_acc=1.0, speculative_token_map=None, speculative_attention_mode='prefill', speculative_draft_attention_backend=None, speculative_moe_runner_backend='auto', speculative_moe_a2a_backend=None, speculative_draft_model_quantization=None, speculative_adaptive=False, speculative_adaptive_config=None, speculative_skip_dp_mlp_sync=False, speculative_ngram_min_bfs_breadth=1, speculative_ngram_max_bfs_breadth=10, speculative_ngram_match_type='BFS', speculative_ngram_max_trie_depth=18, speculative_ngram_capacity=10000000, speculative_ngram_external_corpus_path=None, speculative_ngram_external_sam_budget=0, speculative_ngram_external_corpus_max_tokens=10000000, enable_multi_layer_eagle=False, ep_size=1, moe_a2a_backend='none', moe_runner_backend='auto', record_nolora_graph=True, flashinfer_mxfp4_moe_precision='default', enable_flashinfer_allreduce_fusion=False, enforce_disable_flashinfer_allreduce_fusion=False, enable_aiter_allreduce_fusion=False, deepep_mode='auto', ep_num_redundant_experts=0, ep_dispatch_algorithm=None, init_expert_location='trivial', enable_eplb=False, eplb_algorithm='auto', eplb_rebalance_num_iterations=1000, eplb_rebalance_layers_per_chunk=None, eplb_min_rebalancing_utilization_threshold=1.0, expert_distribution_recorder_mode=None, expert_distribution_recorder_buffer_size=1000, enable_expert_distribution_metrics=False, deepep_config=None, moe_dense_tp_size=None, elastic_ep_backend=None, enable_elastic_expert_backup=False, mooncake_ib_device=None, elastic_ep_rejoin=False, max_mamba_cache_size=None, mamba_ssm_dtype=None, mamba_full_memory_ratio=0.9, mamba_scheduler_strategy='no_buffer', mamba_track_interval=256, linear_attn_backend='triton', linear_attn_decode_backend=None, linear_attn_prefill_backend=None, enable_hierarchical_cache=False, hicache_ratio=2.0, hicache_size=0, hicache_write_policy='write_through', hicache_io_backend='kernel', hicache_mem_layout='layer_first', hicache_storage_backend=None, hicache_storage_prefetch_policy='best_effort', hicache_storage_backend_extra_config=None, enable_hisparse=False, hisparse_config=None, enable_lmcache=False, kt_weight_path=None, kt_method='AMXINT4', kt_cpuinfer=None, kt_threadpool_count=2, kt_num_gpu_experts=None, kt_max_deferred_experts_per_token=None, dllm_algorithm=None, dllm_algorithm_config=None, cpu_offload_gb=0, offload_group_size=-1, offload_num_in_group=1, offload_prefetch_step=1, offload_mode='cpu', enable_mis=False, disable_radix_cache=True, cuda_graph_max_bs=128, cuda_graph_bs=None, disable_cuda_graph=True, disable_cuda_graph_padding=False, enable_breakable_cuda_graph=False, enable_profile_cuda_graph=False, enable_cudagraph_gc=False, debug_cuda_graph=False, enable_layerwise_nvtx_marker=False, enable_nccl_nvls=False, enable_symm_mem=False, disable_flashinfer_cutlass_moe_fp4_allgather=False, enable_tokenizer_batch_encode=False, disable_tokenizer_batch_decode=False, disable_outlines_disk_cache=False, disable_custom_all_reduce=False, enable_mscclpp=False, enable_torch_symm_mem=False, pre_warm_nccl=False, disable_overlap_schedule=False, enable_mixed_chunk=False, enable_dp_attention=False, enable_dp_attention_local_control_broadcast=False, enable_dp_lm_head=False, enable_two_batch_overlap=False, enable_single_batch_overlap=False, tbo_token_distribution_threshold=0.48, enable_torch_compile=False, disable_piecewise_cuda_graph=False, enforce_piecewise_cuda_graph=False, enable_torch_compile_debug_mode=False, torch_compile_max_bs=32, piecewise_cuda_graph_max_tokens=8200, piecewise_cuda_graph_tokens=[4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], piecewise_cuda_graph_compiler='eager', torchao_config='', enable_nan_detection=False, enable_p2p_check=False, triton_attention_reduce_in_fp32=False, triton_attention_num_kv_splits=8, triton_attention_split_tile_size=None, num_continuous_decode_steps=1, delete_ckpt_after_loading=False, enable_memory_saver=False, enable_weights_cpu_backup=False, enable_draft_weights_cpu_backup=False, allow_auto_truncate=False, enable_custom_logit_processor=False, flashinfer_mla_disable_ragged=False, disable_shared_experts_fusion=False, enforce_shared_experts_fusion=False, disable_chunked_prefix_cache=True, disable_fast_image_processor=False, keep_mm_feature_on_device=False, enable_return_hidden_states=False, enable_return_routed_experts=False, scheduler_recv_interval=1, numa_node=None, enable_deterministic_inference=False, rl_on_policy_target=None, enable_attn_tp_input_scattered=False, gc_threshold=None, enable_nsa_prefill_context_parallel=False, nsa_prefill_cp_mode='round-robin-split', enable_fused_qk_norm_rope=False, enable_precise_embedding_interpolation=False, enable_fused_moe_sum_all_reduce=False, enable_prefill_context_parallel=False, prefill_cp_mode='in-seq-split', enable_dynamic_batch_tokenizer=False, dynamic_batch_tokenizer_batch_size=32, dynamic_batch_tokenizer_batch_timeout=0.002, debug_tensor_dump_output_folder=None, debug_tensor_dump_layers=None, debug_tensor_dump_input_file=None, debug_tensor_dump_inject=False, disaggregation_mode='null', disaggregation_transfer_backend='mooncake', disaggregation_bootstrap_port=8998, disaggregation_ib_device=None, disaggregation_decode_enable_radix_cache=False, disaggregation_decode_enable_offload_kvcache=False, num_reserved_decode_tokens=512, disaggregation_decode_polling_interval=1, encoder_only=False, language_only=False, encoder_transfer_backend='zmq_to_scheduler', encoder_urls=[], enable_adaptive_dispatch_to_encoder=False, custom_weight_loader=[], weight_loader_disable_mmap=False, weight_loader_prefetch_checkpoints=False, weight_loader_prefetch_num_threads=4, remote_instance_weight_loader_seed_instance_ip=None, remote_instance_weight_loader_seed_instance_service_port=None, remote_instance_weight_loader_send_weights_group_ports=None, remote_instance_weight_loader_backend='nccl', remote_instance_weight_loader_start_seed_via_transfer_engine=False, engine_info_bootstrap_port=6789, modelexpress_config=None, enable_pdmux=False, pdmux_config_path=None, sm_group_num=8, enable_broadcast_mm_inputs_process=False, enable_prefix_mm_cache=False, mm_enable_dp_encoder=False, mm_process_config={}, limit_mm_data_per_request=None, enable_mm_global_cache=False, decrypted_config_file=None, decrypted_draft_config_file=None, forward_hooks=None, enable_quant_communications=False, msprobe_dump_config=None)
[2026-07-08 10:34:24] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
[2026-07-08 10:34:24] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
[2026-07-08 10:34:24] Using default HuggingFace chat template with detected content format: string
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.4'
Build_Maca_Version = '3.7.1.5'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '4d9c16e'
Vllm Op Version = 0.19.0
SGlang Op Version = 0.5.10
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.7.
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.4'
Build_Maca_Version = '3.7.1.5'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '4d9c16e'
Vllm Op Version = 0.19.0
SGlang Op Version = 0.5.10
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.7.
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.4'
Build_Maca_Version = '3.7.1.5'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '4d9c16e'
Vllm Op Version = 0.19.0
SGlang Op Version = 0.5.10
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.7.
/opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
flashinfer.comm is not available, falling back to standard implementation
flashinfer.comm is not available, falling back to standard implementation
flashinfer.comm is not available, falling back to standard implementation
[2026-07-08 10:34:53 TP1] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
[2026-07-08 10:34:53 TP1] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
[2026-07-08 10:34:53 TP1] Init torch distributed begin.
[2026-07-08 10:34:53 TP0] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
[2026-07-08 10:34:53 TP0] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
[2026-07-08 10:34:53 TP0] Init torch distributed begin.
[2026-07-08 10:34:53] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
[2026-07-08 10:34:53] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
[Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[2026-07-08 10:34:54 TP0] sglang is using nccl==2.16.5
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[2026-07-08 10:34:54 TP0] DCP disabled, dcp_size=1, tp_size=2
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[2026-07-08 10:34:54 TP0] Init torch distributed ends. elapsed=1.01 s, mem usage=0.08 GB
[2026-07-08 10:34:54 TP1] Init torch distributed ends. elapsed=1.10 s, mem usage=0.08 GB
[2026-07-08 10:34:55 TP1] Load weight begin. avail mem=62.03 GB
[2026-07-08 10:34:55 TP0] Load weight begin. avail mem=62.03 GB
Multi-thread loading shards: 0% Completed | 0/2 [00:00<?, ?it/s]
muxi:4519:5239 [0] /workspace/out/Release/build/linux/x86_64/mccl/macaify/src/misc/socket.cc:38 MCCL WARN socketProgressOpt: Call to recv from 127.0.0.1<55875> failed : Connection reset by peer
Traceback (most recent call last):
File "/opt/conda/lib/python3.10/runpy.py", line 196, in _run_module_as_main
return _run_code(code, main_globals, None,
File "/opt/conda/lib/python3.10/runpy.py", line 86, in _run_code
exec(code, run_globals)
File "/opt/conda/lib/python3.10/site-packages/sglang/launch_server.py", line 75, in <module>
run_server(server_args)
File "/opt/conda/lib/python3.10/site-packages/sglang/launch_server.py", line 56, in run_server
launch_server(server_args)
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/entrypoints/http_server.py", line 2345, in launch_server
) = Engine._launch_subprocesses(
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/entrypoints/engine.py", line 822, in _launch_subprocesses
scheduler_init_result.wait_for_ready()
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/entrypoints/engine.py", line 623, in wait_for_ready
infos = _wait_for_scheduler_ready(scheduler_pipe_readers, scheduler_procs)
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/entrypoints/engine.py", line 1322, in _wait_for_scheduler_ready
raise _scheduler_died_error(j, scheduler_procs[j])
RuntimeError: Rank 1 scheduler died during initialization (exit code: -9). If exit code is -9 (SIGKILL), a common cause is the OS OOM killer. Run dmesg -T | grep -i oom to check.
沐曦卡和驱动信息:
mx-smi
mx-smi version: 2.3.1
=================== MetaX System Management Interface Log ===================
Timestamp : Wed Jul 8 10:26:24 2026
Attached GPUs : 4
+---------------------------------------------------------------------------------+
| MX-SMI 2.3.1 Kernel Mode Driver Version: 3.8.30 |
| MACA Version: 3.7.1.5 BIOS Version: 1.33.5.0 |
|------------------+-----------------+---------------------+----------------------|
| Board Name | GPU Persist-M | Bus-id | GPU-Util sGPU-M |
| Pwr:Usage/Cap | Temp Perf | Memory-Usage | GPU-State |
|==================+=================+=====================+======================|
| 0 MetaX C500 | 0 Off | 0000:38:00.0 | 0% Disabled |
| 31W / 350W | 37C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 1 MetaX C500 | 1 Off | 0000:49:00.0 | 0% Disabled |
| 32W / 350W | 35C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 2 MetaX C500 | 2 Off | 0000:98:00.0 | 0% Disabled |
| 34W / 350W | 39C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 3 MetaX C500 | 3 Off | 0000:c8:00.0 | 0% Disabled |
| 31W / 350W | 38C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
+---------------------------------------------------------------------------------+
| Process: |
| GPU PID Process Name GPU Memory |
| Usage(MiB) |
|=================================================================================|
| no process found |
+---------------------------------------------------------------------------------+
docker容器启动命令:
docker run -d \
--network host \
--ipc=host \
--name=sglang-mx \
--shm-size 100G \
--ulimit memlock=-1 \
--ulimit nofile=1024000 \
--device=/dev/dri \
--device=/dev/mxcd \
--device=/dev/mem \
--group-add video \
--security-opt seccomp=unconfined \
--security-opt apparmor=unconfined \
-v /data/models/:/models:ro \
cr.metax-tech.com/public-ai-release/maca/sglang:0.5.11-maca.ai3.7.1.107-torch2.8-py310-kylinv11-amd64 \
sleep infinity
sglang启动命令:
python -m sglang.launch_server --model-path /models/DeepSeek-R1-Distill-Qwen-7B --host 0.0.0.0 --port 9100 --served-model-name deepseek --tensor-parallel-size 2 --mem-fraction-static 0.7 --nccl-port 29500 --trust-remote-code --enable-metrics --attention-backend triton --cpu-offload-gb 0 --disable-cuda-graph
错误日志:
python -m sglang.launch_server --model-path /models/DeepSeek-R1-Distill-Qwen-7B --host 0.0.0.0 --port 9100 --served-model-name deepseek --tensor-parallel-size 2 --mem-fraction-static 0.7 --nccl-port 29500 --trust-remote-code --enable-metrics --attention-backend triton --cpu-offload-gb 0 --disable-cuda-graph
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.4'
Build_Maca_Version = '3.7.1.5'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '4d9c16e'
Vllm Op Version = 0.19.0
SGlang Op Version = 0.5.10
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.7.
W0708 10:19:42.990000 1905 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
W0708 10:19:42.990000 1905 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
/opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/sglang/launch_server.py:60: UserWarning: 'python -m sglang.launch_server' is still supported, but 'sglang serve' is the recommended entrypoint.
Example: sglang serve --model-path <model> [options]
warnings.warn(
[2026-07-08 10:19:46] Cuda graph max bs is adjusted to 128.
[2026-07-08 10:19:47] flashinfer.comm is not available, falling back to standard implementation
[2026-07-08 10:19:48] server_args=ServerArgs(model_path='/models/DeepSeek-R1-Distill-Qwen-7B', tokenizer_path='/models/DeepSeek-R1-Distill-Qwen-7B', tokenizer_mode='auto', tokenizer_backend='huggingface', tokenizer_worker_num=1, detokenizer_worker_num=1, skip_tokenizer_init=False, load_format='auto', model_loader_extra_config='{}', trust_remote_code=True, context_length=None, is_embedding=False, enable_multimodal=None, revision=None, model_impl='auto', host='0.0.0.0', port=9100, fastapi_root_path='', grpc_mode=False, skip_server_warmup=False, warmups=None, nccl_port=29500, checkpoint_engine_wait_weights_before_ready=False, ssl_keyfile=None, ssl_certfile=None, ssl_ca_certs=None, ssl_keyfile_password=None, enable_ssl_refresh=False, enable_http2=False, dtype='auto', quantization=None, quantization_param_path=None, kv_cache_dtype='auto', enable_fp32_lm_head=False, modelopt_quant=None, modelopt_checkpoint_restore_path=None, modelopt_checkpoint_save_path=None, modelopt_export_path=None, quantize_and_serve=False, rl_quant_profile=None, mem_fraction_static=0.7, max_running_requests=None, max_queued_requests=None, max_total_tokens=None, chunked_prefill_size=8200, enable_dynamic_chunking=False, max_prefill_tokens=16384, prefill_max_requests=None, schedule_policy='fcfs', enable_priority_scheduling=False, disable_priority_preemption=False, default_priority_value=None, abort_on_priority_when_disabled=False, schedule_low_priority_values_first=False, priority_scheduling_preemption_threshold=10, schedule_conservativeness=1.0, page_size=1, swa_full_tokens_ratio=0.8, disable_hybrid_swa_memory=False, radix_eviction_policy='lru', enable_prefill_delayer=False, prefill_delayer_max_delay_passes=30, prefill_delayer_token_usage_low_watermark=None, prefill_delayer_forward_passes_buckets=None, prefill_delayer_wait_seconds_buckets=None, device='cuda', tp_size=2, pp_size=1, pp_max_micro_batch_size=None, pp_async_batch_depth=0, stream_interval=1, batch_notify_size=16, stream_response_default_include_usage=False, incremental_streaming_output=False, enable_streaming_session=False, random_seed=186191225, constrained_json_whitespace_pattern=None, constrained_json_disable_any_whitespace=False, watchdog_timeout=300, soft_watchdog_timeout=None, dist_timeout=None, download_dir=None, model_checksum=None, base_gpu_id=0, gpu_id_step=1, sleep_on_idle=False, use_ray=False, custom_sigquit_handler=None, log_level='info', log_level_http=None, log_requests=False, log_requests_level=2, log_requests_format='text', log_requests_target=None, uvicorn_access_log_exclude_prefixes=[], crash_dump_folder=None, show_time_cost=False, enable_metrics=True, grpc_http_sidecar_port=None, enable_mfu_metrics=False, enable_metrics_for_all_schedulers=False, tokenizer_metrics_custom_labels_header='x-custom-labels', tokenizer_metrics_allowed_custom_labels=None, extra_metric_labels=None, bucket_time_to_first_token=None, bucket_inter_token_latency=None, bucket_e2e_request_latency=None, prompt_tokens_buckets=None, generation_tokens_buckets=None, gc_warning_threshold_secs=0.0, decode_log_interval=40, enable_request_time_stats_logging=False, kv_events_config=None, enable_trace=False, otlp_traces_endpoint='localhost:4317', export_metrics_to_file=False, export_metrics_to_file_dir=None, api_key=None, admin_api_key=None, served_model_name='deepseek', weight_version='default', chat_template=None, hf_chat_template_name=None, completion_template=None, file_storage_path='sglang_storage', enable_cache_report=False, reasoning_parser=None, strip_thinking_cache=False, tool_call_parser=None, tool_server=None, sampling_defaults='model', dp_size=1, load_balance_method='round_robin', attn_cp_size=1, moe_dp_size=1, dist_init_addr=None, nnodes=1, node_rank=0, json_model_override_args='{}', preferred_sampling_params=None, enable_lora=None, enable_lora_overlap_loading=None, max_lora_rank=None, lora_target_modules=None, lora_paths=None, max_loaded_loras=None, max_loras_per_batch=8, lora_eviction_policy='lru', lora_backend='csgmv', max_lora_chunk_size=16, experts_shared_outer_loras=None, lora_use_virtual_experts=False, lora_strict_loading=False, attention_backend='triton', decode_attention_backend=None, prefill_attention_backend=None, sampling_backend='flashinfer', grammar_backend='xgrammar', mm_attention_backend=None, fp8_gemm_runner_backend='auto', fp4_gemm_runner_backend='auto', nsa_prefill_backend='flashmla_sparse', nsa_decode_backend='flashmla_kv', disable_flashinfer_autotune=False, mamba_backend='triton', speculative_algorithm=None, speculative_draft_model_path=None, speculative_draft_model_revision=None, speculative_draft_load_format=None, speculative_num_steps=None, speculative_eagle_topk=None, speculative_num_draft_tokens=None, speculative_dflash_block_size=None, speculative_dflash_draft_window_size=None, speculative_accept_threshold_single=1.0, speculative_accept_threshold_acc=1.0, speculative_token_map=None, speculative_attention_mode='prefill', speculative_draft_attention_backend=None, speculative_moe_runner_backend='auto', speculative_moe_a2a_backend=None, speculative_draft_model_quantization=None, speculative_adaptive=False, speculative_adaptive_config=None, speculative_skip_dp_mlp_sync=False, speculative_ngram_min_bfs_breadth=1, speculative_ngram_max_bfs_breadth=10, speculative_ngram_match_type='BFS', speculative_ngram_max_trie_depth=18, speculative_ngram_capacity=10000000, speculative_ngram_external_corpus_path=None, speculative_ngram_external_sam_budget=0, speculative_ngram_external_corpus_max_tokens=10000000, enable_multi_layer_eagle=False, ep_size=1, moe_a2a_backend='none', moe_runner_backend='auto', record_nolora_graph=True, flashinfer_mxfp4_moe_precision='default', enable_flashinfer_allreduce_fusion=False, enforce_disable_flashinfer_allreduce_fusion=False, enable_aiter_allreduce_fusion=False, deepep_mode='auto', ep_num_redundant_experts=0, ep_dispatch_algorithm=None, init_expert_location='trivial', enable_eplb=False, eplb_algorithm='auto', eplb_rebalance_num_iterations=1000, eplb_rebalance_layers_per_chunk=None, eplb_min_rebalancing_utilization_threshold=1.0, expert_distribution_recorder_mode=None, expert_distribution_recorder_buffer_size=1000, enable_expert_distribution_metrics=False, deepep_config=None, moe_dense_tp_size=None, elastic_ep_backend=None, enable_elastic_expert_backup=False, mooncake_ib_device=None, elastic_ep_rejoin=False, max_mamba_cache_size=None, mamba_ssm_dtype=None, mamba_full_memory_ratio=0.9, mamba_scheduler_strategy='no_buffer', mamba_track_interval=256, linear_attn_backend='triton', linear_attn_decode_backend=None, linear_attn_prefill_backend=None, enable_hierarchical_cache=False, hicache_ratio=2.0, hicache_size=0, hicache_write_policy='write_through', hicache_io_backend='kernel', hicache_mem_layout='layer_first', hicache_storage_backend=None, hicache_storage_prefetch_policy='best_effort', hicache_storage_backend_extra_config=None, enable_hisparse=False, hisparse_config=None, enable_lmcache=False, kt_weight_path=None, kt_method='AMXINT4', kt_cpuinfer=None, kt_threadpool_count=2, kt_num_gpu_experts=None, kt_max_deferred_experts_per_token=None, dllm_algorithm=None, dllm_algorithm_config=None, cpu_offload_gb=0, offload_group_size=-1, offload_num_in_group=1, offload_prefetch_step=1, offload_mode='cpu', enable_mis=False, disable_radix_cache=False, cuda_graph_max_bs=128, cuda_graph_bs=None, disable_cuda_graph=True, disable_cuda_graph_padding=False, enable_breakable_cuda_graph=False, enable_profile_cuda_graph=False, enable_cudagraph_gc=False, debug_cuda_graph=False, enable_layerwise_nvtx_marker=False, enable_nccl_nvls=False, enable_symm_mem=False, disable_flashinfer_cutlass_moe_fp4_allgather=False, enable_tokenizer_batch_encode=False, disable_tokenizer_batch_decode=False, disable_outlines_disk_cache=False, disable_custom_all_reduce=False, enable_mscclpp=False, enable_torch_symm_mem=False, pre_warm_nccl=False, disable_overlap_schedule=False, enable_mixed_chunk=False, enable_dp_attention=False, enable_dp_attention_local_control_broadcast=False, enable_dp_lm_head=False, enable_two_batch_overlap=False, enable_single_batch_overlap=False, tbo_token_distribution_threshold=0.48, enable_torch_compile=False, disable_piecewise_cuda_graph=False, enforce_piecewise_cuda_graph=False, enable_torch_compile_debug_mode=False, torch_compile_max_bs=32, piecewise_cuda_graph_max_tokens=8200, piecewise_cuda_graph_tokens=[4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], piecewise_cuda_graph_compiler='eager', torchao_config='', enable_nan_detection=False, enable_p2p_check=False, triton_attention_reduce_in_fp32=False, triton_attention_num_kv_splits=8, triton_attention_split_tile_size=None, num_continuous_decode_steps=1, delete_ckpt_after_loading=False, enable_memory_saver=False, enable_weights_cpu_backup=False, enable_draft_weights_cpu_backup=False, allow_auto_truncate=False, enable_custom_logit_processor=False, flashinfer_mla_disable_ragged=False, disable_shared_experts_fusion=False, enforce_shared_experts_fusion=False, disable_chunked_prefix_cache=False, disable_fast_image_processor=False, keep_mm_feature_on_device=False, enable_return_hidden_states=False, enable_return_routed_experts=False, scheduler_recv_interval=1, numa_node=None, enable_deterministic_inference=False, rl_on_policy_target=None, enable_attn_tp_input_scattered=False, gc_threshold=None, enable_nsa_prefill_context_parallel=False, nsa_prefill_cp_mode='round-robin-split', enable_fused_qk_norm_rope=False, enable_precise_embedding_interpolation=False, enable_fused_moe_sum_all_reduce=False, enable_prefill_context_parallel=False, prefill_cp_mode='in-seq-split', enable_dynamic_batch_tokenizer=False, dynamic_batch_tokenizer_batch_size=32, dynamic_batch_tokenizer_batch_timeout=0.002, debug_tensor_dump_output_folder=None, debug_tensor_dump_layers=None, debug_tensor_dump_input_file=None, debug_tensor_dump_inject=False, disaggregation_mode='null', disaggregation_transfer_backend='mooncake', disaggregation_bootstrap_port=8998, disaggregation_ib_device=None, disaggregation_decode_enable_radix_cache=False, disaggregation_decode_enable_offload_kvcache=False, num_reserved_decode_tokens=512, disaggregation_decode_polling_interval=1, encoder_only=False, language_only=False, encoder_transfer_backend='zmq_to_scheduler', encoder_urls=[], enable_adaptive_dispatch_to_encoder=False, custom_weight_loader=[], weight_loader_disable_mmap=False, weight_loader_prefetch_checkpoints=False, weight_loader_prefetch_num_threads=4, remote_instance_weight_loader_seed_instance_ip=None, remote_instance_weight_loader_seed_instance_service_port=None, remote_instance_weight_loader_send_weights_group_ports=None, remote_instance_weight_loader_backend='nccl', remote_instance_weight_loader_start_seed_via_transfer_engine=False, engine_info_bootstrap_port=6789, modelexpress_config=None, enable_pdmux=False, pdmux_config_path=None, sm_group_num=8, enable_broadcast_mm_inputs_process=False, enable_prefix_mm_cache=False, mm_enable_dp_encoder=False, mm_process_config={}, limit_mm_data_per_request=None, enable_mm_global_cache=False, decrypted_config_file=None, decrypted_draft_config_file=None, forward_hooks=None, enable_quant_communications=False, msprobe_dump_config=None)
[2026-07-08 10:19:49] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
[2026-07-08 10:19:49] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
[2026-07-08 10:19:49] Using default HuggingFace chat template with detected content format: string
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.4'
Build_Maca_Version = '3.7.1.5'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '4d9c16e'
Vllm Op Version = 0.19.0
SGlang Op Version = 0.5.10
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.7.
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.4'
Build_Maca_Version = '3.7.1.5'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '4d9c16e'
Vllm Op Version = 0.19.0
SGlang Op Version = 0.5.10
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.7.
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.4'
Build_Maca_Version = '3.7.1.5'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '4d9c16e'
Vllm Op Version = 0.19.0
SGlang Op Version = 0.5.10
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.7.
W0708 10:20:13.533000 2113 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
W0708 10:20:13.533000 2113 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
W0708 10:20:13.533000 2111 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
W0708 10:20:13.533000 2111 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
W0708 10:20:13.533000 2112 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
W0708 10:20:13.533000 2112 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
/opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
flashinfer.comm is not available, falling back to standard implementation
flashinfer.comm is not available, falling back to standard implementation
flashinfer.comm is not available, falling back to standard implementation
[2026-07-08 10:20:20 TP0] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
[2026-07-08 10:20:20 TP0] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
[2026-07-08 10:20:20] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
[2026-07-08 10:20:20 TP1] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
[2026-07-08 10:20:20] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
[2026-07-08 10:20:20 TP0] Init torch distributed begin.
[2026-07-08 10:20:20 TP1] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
[2026-07-08 10:20:20 TP1] Init torch distributed begin.
[Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[2026-07-08 10:20:20 TP0] sglang is using nccl==2.16.5
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[2026-07-08 10:20:20 TP0] DCP disabled, dcp_size=1, tp_size=2
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[2026-07-08 10:20:20 TP1] Init torch distributed ends. elapsed=0.61 s, mem usage=0.08 GB
[2026-07-08 10:20:20 TP0] Init torch distributed ends. elapsed=0.65 s, mem usage=0.08 GB
[2026-07-08 10:20:22 TP0] Load weight begin. avail mem=62.03 GB
[2026-07-08 10:20:22 TP1] Load weight begin. avail mem=62.03 GB
Multi-thread loading shards: 50% Completed | 1/2 [00:13<00:13, 13.38s/it]
muxi:2111:2831 [0] /workspace/out/Release/build/linux/x86_64/mccl/macaify/src/misc/socket.cc:38 MCCL WARN socketProgressOpt: Call to recv from 127.0.0.1<48913> failed : Connection reset by peer
muxi:2111:2831 [0] /workspace/out/Release/build/linux/x86_64/mccl/macaify/src/misc/socket.cc:567 MCCL WARN socketStartConnect: Connect to 127.0.0.1<48913> failed : Software caused connection abort
muxi:2111:2831 [0] /workspace/out/Release/build/linux/x86_64/mccl/macaify/src/misc/socket.cc:567 MCCL WARN socketStartConnect: Connect to 127.0.0.1<48913> failed : Software caused connection abort
Multi-thread loading shards: 100% Completed | 2/2 [00:16<00:00, 8.41s/it]
[2026-07-08 10:20:38 TP0] Load weight end. elapsed=16.99 s, type=Qwen2ForCausalLM, avail mem=54.71 GB, mem usage=7.32 GB.
muxi:2111:2831 [0] /workspace/out/Release/build/linux/x86_64/mccl/macaify/src/misc/socket.cc:567 MCCL WARN socketStartConnect: Connect to 127.0.0.1<48913> failed : Software caused connection abort
[2026-07-08 10:20:39 TP0] Scheduler hit an exception: Traceback (most recent call last):
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/scheduler.py", line 3943, in run_scheduler_process
scheduler = Scheduler(
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/scheduler.py", line 439, in init
self.init_model_worker()
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/scheduler.py", line 714, in init_model_worker
self.init_tp_model_worker()
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/scheduler.py", line 669, in init_tp_model_worker
self.tp_worker = TpModelWorker(**worker_kwargs)
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/tp_worker.py", line 260, in init
self._init_model_runner()
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/tp_worker.py", line 345, in _init_model_runner
self._model_runner = ModelRunner(
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/model_executor/model_runner.py", line 510, in init
self.initialize(pre_model_load_memory)
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/model_executor/model_runner.py", line 627, in initialize
self.load_model()
File "/opt/conda/lib/python3.10/site-packages/sglang/srt/model_executor/model_runner.py", line 1535, in load_model
raise ValueError(
ValueError: TP rank 0 could finish the model loading, but there are other ranks that didn't finish loading. It is likely due to unexpected failures (e.g., OOM) or a slow node.
[2026-07-08 10:20:39] Received sigquit from a child process. It usually means the child failed.