还是必须要2台8卡*64G 的服务器 才能部署起来?
还是必须要2台8卡*64G 的服务器 才能部署起来?
就是单台服务器8卡是能部署起来的是吗?
就是8卡 是能跑起来的是吧,显存上是满足的是吗?
一、软硬件信息
1.服务器厂家: H3C
2.沐曦GPU型号:C550
3.操作系统内核版本:
Linux muxi-gpu-server 6.14.0-27-generic #27~24.04.1-Ubuntu SMP PREEMPT_DYNAMIC Tue Jul 22 17:38:49 UTC 2 x86_64 x86_64 x86_64 GNU/Linux
4.是否开启CPU虚拟化:
app@muxi-gpu-server:~$ lscpu | grep -i virtualization
Virtualization: AMD-V
5、镜像为vllm-metax:0.22.0-maca.ai3.8.0.5-torch2.10-py312-kylinv11-amd64
6.mx-smi回显:
app@muxi-gpu-server:~$ mx-smi
mx-smi version: 2.3.4
=================== MetaX System Management Interface Log ===================
Timestamp : Mon Aug 3 01:52:32 2026
Attached GPUs : 8
+---------------------------------------------------------------------------------+
| MX-SMI 2.3.4 Kernel Mode Driver Version: 3.9.10 |
| MACA Version: 3.8.0.23 BIOS Version: 1.22.3.0 |
|------------------+-----------------+---------------------+----------------------|
| Board Name | GPU Persist-M | Bus-id | GPU-Util sGPU-M |
| Pwr:Usage/Cap | Temp Perf | Memory-Usage | GPU-State |
|==================+=================+=====================+======================|
| 0 MetaX C550 | 0 N/A | 0000:23:00.0 | 16% Disabled |
| 131W / 450W | 46C N/A | 34589/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 1 MetaX C550 | 1 N/A | 0000:26:00.0 | 13% Disabled |
| 131W / 450W | 41C N/A | 34589/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 2 MetaX C550 | 2 N/A | 0000:63:00.0 | 10% Disabled |
| 129W / 450W | 41C N/A | 34589/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 3 MetaX C550 | 3 N/A | 0000:66:00.0 | 9% Disabled |
| 135W / 450W | 47C N/A | 34589/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 4 MetaX C550 | 4 N/A | 0000:a3:00.0 | 9% Disabled |
| 132W / 450W | 46C N/A | 34589/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 5 MetaX C550 | 5 N/A | 0000:a4:00.0 | 13% Disabled |
| 130W / 450W | 41C N/A | 34589/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 6 MetaX C550 | 6 N/A | 0000:e3:00.0 | 13% Disabled |
| 127W / 450W | 41C N/A | 34589/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 7 MetaX C550 | 7 N/A | 0000:e4:00.0 | 8% Disabled |
| 135W / 450W | 46C N/A | 34589/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
+---------------------------------------------------------------------------------+
| Process: |
| GPU PID Process Name GPU Memory |
| Usage(MiB) |
|=================================================================================|
| 0 82079 VLLM::Worker_TP 33718 |
| 1 82080 VLLM::Worker_TP 33718 |
| 2 82081 VLLM::Worker_TP 33718 |
| 3 82082 VLLM::Worker_TP 33718 |
| 4 82083 VLLM::Worker_TP 33718 |
| 5 82084 VLLM::Worker_TP 33718 |
| 6 82085 VLLM::Worker_TP 33718 |
| 7 82086 VLLM::Worker_TP 33718 |
+---------------------------------------------------------------------------------+
End of Log
app@muxi-gpu-server:~$
根据以上信息,我的8卡能部署 Qwen3.5-397B-A17B吗?有什么参数推荐和要求
服务器未关闭ACS导致pcie逻辑隔离,卡间无法通信 。修改grub配置后系统 恢复正常通信
补充测试日志:app@muxi-gpu-server:~$ bash /opt/maca/samples/mccl_tests/perf/mccl.sh 1
The test is all_reduce_perf, the maca version is /opt/maca-3.8.0
main_process = 501528
===============================
1024 512 bfloat16 sum -1 5.33 0.19 0.00 0 0.28 3.70 0.00 0
2048 1024 bfloat16 sum -1 5.29 0.39 0.00 0 0.28 7.20 0.00 0
4096 2048 bfloat16 sum -1 5.34 0.77 0.00 0 0.27 15.20 0.00 0
8192 4096 bfloat16 sum -1 5.20 1.57 0.00 0 0.27 30.40 0.00 0
16384 8192 bfloat16 sum -1 5.32 3.08 0.00 0 0.27 60.79 0.00 0
32768 16384 bfloat16 sum -1 5.48 5.98 0.00 0 0.28 116.82 0.00 0
65536 32768 bfloat16 sum -1 6.09 10.76 0.00 0 0.27 244.08 0.00 0
131072 65536 bfloat16 sum -1 7.39 17.74 0.00 0 0.27 479.24 0.00 0
262144 131072 bfloat16 sum -1 9.21 28.45 0.00 0 0.28 951.52 0.00 0
524288 262144 bfloat16 sum -1 13.63 38.46 0.00 0 0.27 1923.99 0.00 0
1048576 524288 bfloat16 sum -1 22.20 47.23 0.00 0 0.27 3876.44 0.00 0
2097152 1048576 bfloat16 sum -1 39.87 52.61 0.00 0 0.28 7527.47 0.00 0
4194304 2097152 bfloat16 sum -1 29.72 141.15 0.00 0 0.27 15279.80 0.00 0
8388608 4194304 bfloat16 sum -1 34.49 243.21 0.00 0 0.27 30885.89 0.00 0
16777216 8388608 bfloat16 sum -1 46.04 364.44 0.00 0 0.27 61567.77 0.00 0
33554432 16777216 bfloat16 sum -1 68.91 486.96 0.00 0 0.27 123589.07 0.00 0
67108864 33554432 bfloat16 sum -1 115.29 582.11 0.00 0 0.27 249939.90 0.00 0
134217728 67108864 bfloat16 sum -1 208.51 643.71 0.00 0 0.21 640963.36 0.00 0
268435456 134217728 bfloat16 sum -1 397.83 674.74 0.00 0 0.19 1440104.38 0.00 0
536870912 268435456 bfloat16 sum -1 758.84 707.49 0.00 0 0.19 2849633.29 0.00 0
1073741824 536870912 bfloat16 sum -1 1499.05 716.28 0.00 0 0.29 3683505.40 0.00 0
app@muxi-gpu-server:~$ bash /opt/maca/samples/mccl_tests/perf/mccl.sh 2
The test is all_reduce_perf, the maca version is /opt/maca-3.8.0
main_process = 501623
===============================
^C^Capp@muxi-gpu-server:~$ bash /opt/maca/samples/mccl_tests/perf/mccl.sh 8
The test is all_reduce_perf, the maca version is /opt/maca-3.8.0
main_process = 503189
===============================
^C^Capp@muxi-gpu-server:~$ bash /opt/maca/samples/mccl_tests/perf/mccl.sh 4
The test is all_reduce_perf, the maca version is /opt/maca-3.8.0
main_process = 504027
===============================
就是这个性能测试脚本的作用是什么?bash /opt/maca/samples/mccl_tests/perf/mccl.sh 1 ,用一张卡能测试出数据,但是2、4、8 多张卡的时候都没有数据,这个能判断是多卡互联通信存在问题吗?
我在不启动docker的情况下执行bash /opt/maca/samples/mccl_tests/perf/mccl.sh 1、bash /opt/maca/samples/mccl_tests/perf/mccl.sh 4、bash /opt/maca/samples/mccl_tests/perf/mccl.sh 8,日志如下 ,这个能说明是八卡互联通信存在问题吗?
app@muxi-gpu-server:~$ bash /opt/maca/samples/mccl_tests/perf/mccl.sh 1
The test is all_reduce_perf, the maca version is /opt/maca-3.8.0
main_process = 457322
===============================
1024 512 bfloat16 sum -1 5.85 0.18 0.00 0 0.28 3.69 0.00 0
2048 1024 bfloat16 sum -1 5.17 0.40 0.00 0 0.30 6.88 0.00 0
4096 2048 bfloat16 sum -1 5.52 0.74 0.00 0 0.27 15.03 0.00 0
8192 4096 bfloat16 sum -1 5.17 1.58 0.00 0 0.27 30.51 0.00 0
16384 8192 bfloat16 sum -1 5.40 3.03 0.00 0 0.27 60.57 0.00 0
32768 16384 bfloat16 sum -1 6.11 5.36 0.00 0 0.27 120.65 0.00 0
65536 32768 bfloat16 sum -1 6.30 10.40 0.00 0 0.28 237.88 0.00 0
131072 65536 bfloat16 sum -1 7.32 17.89 0.00 0 0.29 454.32 0.00 0
262144 131072 bfloat16 sum -1 9.32 28.12 0.00 0 0.27 954.99 0.00 0
524288 262144 bfloat16 sum -1 13.45 38.99 0.00 0 0.27 1945.41 0.00 0
1048576 524288 bfloat16 sum -1 22.11 47.42 0.00 0 0.27 3862.16 0.00 0
2097152 1048576 bfloat16 sum -1 39.14 53.58 0.00 0 0.27 7667.83 0.00 0
4194304 2097152 bfloat16 sum -1 28.19 148.78 0.00 0 0.29 14533.28 0.00 0
8388608 4194304 bfloat16 sum -1 35.12 238.88 0.00 0 0.28 30338.55 0.00 0
16777216 8388608 bfloat16 sum -1 46.05 364.34 0.00 0 0.28 59388.38 0.00 0
33554432 16777216 bfloat16 sum -1 72.91 460.20 0.00 0 0.28 118734.72 0.00 0
67108864 33554432 bfloat16 sum -1 115.19 582.59 0.00 0 0.28 243589.34 0.00 0
134217728 67108864 bfloat16 sum -1 207.48 646.89 0.00 0 0.30 440780.72 0.00 0
268435456 134217728 bfloat16 sum -1 399.89 671.28 0.00 0 0.28 963861.60 0.00 0
536870912 268435456 bfloat16 sum -1 761.13 705.36 0.00 0 0.25 2187738.03 0.00 0
1073741824 536870912 bfloat16 sum -1 1495.97 717.76 0.00 0 0.21 5177154.41 0.00 0
app@muxi-gpu-server:~$
已按要求reboot后 还是一样的问题,无法成功
已按要求将驱动和SDK升级到3.8.0.23版本,但是还是报No available shared memory broadcast block found日志,无法正常部署成功。
Attached GPUs : 8
+---------------------------------------------------------------------------------+
| MX-SMI 2.3.4 Kernel Mode Driver Version: 3.9.10 |
| MACA Version: 3.8.0.23 BIOS Version: 1.22.3.0 |
|------------------+-----------------+---------------------+----------------------|
| Board Name | GPU Persist-M | Bus-id | GPU-Util sGPU-M |
| Pwr:Usage/Cap | Temp Perf | Memory-Usage | GPU-State |
|==================+=================+=====================+======================|
| 0 MetaX C550 | 0 N/A | 0000:23:00.0 | 29% Disabled |
| 128W / 450W | 47C N/A | 43534/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 1 MetaX C550 | 1 N/A | 0000:26:00.0 | 29% Disabled |
| 128W / 450W | 42C N/A | 43534/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 2 MetaX C550 | 2 N/A | 0000:63:00.0 | 29% Disabled |
| 128W / 450W | 42C N/A | 43534/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 3 MetaX C550 | 3 N/A | 0000:66:00.0 | 29% Disabled |
| 135W / 450W | 50C N/A | 43534/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 4 MetaX C550 | 4 N/A | 0000:a3:00.0 | 29% Disabled |
| 130W / 450W | 48C N/A | 43534/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 5 MetaX C550 | 5 N/A | 0000:a4:00.0 | 29% Disabled |
| 129W / 450W | 42C N/A | 43534/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 6 MetaX C550 | 6 N/A | 0000:e3:00.0 | 29% Disabled |
| 125W / 450W | 42C N/A | 43534/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 7 MetaX C550 | 7 N/A | 0000:e4:00.0 | 29% Disabled |
| 134W / 450W | 48C N/A | 43534/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
详细日志如下:
这个有升级的步骤吗?这个是升级 裸机的显卡驱动和sdk吗?具体是升级到 3.8.0.x 的小版本是多少呢?还是任意的版本都可以?
app@muxi-gpu-server:~$ bash /opt/maca/samples/mccl_tests/perf/mccl.sh 8
The test is all_reduce_perf, the maca version is /opt/maca-3.5.3
/opt/maca/samples/mccl_tests/perf/mccl_perf/all_reduce_perf: error while loading shared libraries: libToolsExt_cu.so: cannot open shared object file: No such file or directory
/opt/maca/samples/mccl_tests/perf/mccl_perf/all_reduce_perf: error while loading shared libraries: libToolsExt_cu.so: cannot open shared object file: No such file or directory
/opt/maca/samples/mccl_tests/perf/mccl_perf/all_reduce_perf: error while loading shared libraries: libToolsExt_cu.so: cannot open shared object file: No such file or directory
/opt/maca/samples/mccl_tests/perf/mccl_perf/all_reduce_perf: error while loading shared libraries: libToolsExt_cu.so: cannot open shared object file: No such file or directory
/opt/maca/samples/mccl_tests/perf/mccl_perf/all_reduce_perf: error while loading shared libraries: libToolsExt_cu.so: cannot open shared object file: No such file or directory
/opt/maca/samples/mccl_tests/perf/mccl_perf/all_reduce_perf: error while loading shared libraries: libToolsExt_cu.so: cannot open shared object file: No such file or directory
/opt/maca/samples/mccl_tests/perf/mccl_perf/all_reduce_perf: error while loading shared libraries: libToolsExt_cu.so: cannot open shared object file: No such file or directory
Primary job terminated normally, but 1 process returned
a non-zero exit code. Per user-direction, the job has been aborted.
/opt/maca/samples/mccl_tests/perf/mccl_perf/all_reduce_perf: error while loading shared libraries: libToolsExt_cu.so: cannot open shared object file: No such file or directory
app@muxi-gpu-server:~$
lspci | grep 9999
lsmod | grep metax
sudo grep -P "METAX.ERROR|MXCD.ERROR|MXGVM.*ERROR" /var/log/syslog
等日志补充
裸金属执行
dmesg -T | grep -i err 后日志补充如下
已添加参数。vllm serve /data/models/Qwen3.5-122B-A10B \
--pipeline-parallel-size 1 \
--tensor-parallel-size 8 \
--trust-remote-code \
--dtype bfloat16 \
--distributed-executor-backend mp \
--gpu-memory-utilization 0.85 \
--max-model-len 131072 \
--max-num-batched-tokens 131072 \
--no-async-scheduling \
--mm-encoder-tp-mode data \
--mm-processor-cache-type shm \
--limit-mm-per-prompt '{"image": 5, "video": 1}' \
--skip-mm-profiling \
--enable-prefix-caching \
--served-model-name Qwen3.5-122B-A10B \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--host 0.0.0.0 \
--attention-backend FLASH_ATTN \
--port 40022
还是不行,报了新的错误。[11:16:29.449][MXKW][E]queues.c :844 : [mxkwCreateQueueBlock][Hint]ioctl create queue block timeout, gpu_id:22863 type:21. Retrying.
[11:16:29.449][MXKW][E]queues.c :844 : [mxkwCreateQueueBlock][Hint]ioctl create queue block timeout, gpu_id:15197 type:21. Retrying.
[11:16:29.449][MXKW][E]queues.c :844 : [mxkwCreateQueueBlock][Hint]ioctl create queue block timeout, gpu_id:27405 type:21. Retrying.
[11:16:29.449][MXKW][E]queues.c :844 : [mxkwCreateQueueBlock][Hint]ioctl create queue block timeout, gpu_id:52076 type:21. Retrying.
[11:16:29.451][MXKW][E]queues.c :844 : [mxkwCreateQueueBlock][Hint]ioctl create queue block timeout, gpu_id:26782 type:21. Retrying.
[11:16:29.458][MXKW][E]queues.c :844 : [mxkwCreateQueueBlock][Hint]ioctl create queue block timeout, gpu_id:39740 type:21. Retrying.
一、软硬件信息
1.服务器厂家: H3C
2.沐曦GPU型号:C550
3.操作系统内核版本:
Linux muxi-gpu-server 6.14.0-27-generic #27~24.04.1-Ubuntu SMP PREEMPT_DYNAMIC Tue Jul 22 17:38:49 UTC 2 x86_64 x86_64 x86_64 GNU/Linux
4.是否开启CPU虚拟化:
app@muxi-gpu-server:~$ lscpu | grep -i virtualization
Virtualization: AMD-V
5.mx-smi回显:
app@muxi-gpu-server:~$ mx-smi
mx-smi version: 2.2.12
=================== MetaX System Management Interface Log ===================
Timestamp : Thu Jul 30 02:21:15 2026
Attached GPUs : 8
+---------------------------------------------------------------------------------+
| MX-SMI 2.2.12 Kernel Mode Driver Version: 3.6.11 |
| MACA Version: 3.5.3.18 BIOS Version: 1.22.3.0 |
|------------------+-----------------+---------------------+----------------------|
| Board Name | GPU Persist-M | Bus-id | GPU-Util sGPU-M |
| Pwr:Usage/Cap | Temp Perf | Memory-Usage | GPU-State |
|==================+=================+=====================+======================|
| 0 MetaX C550 | 0 N/A | 0000:23:00.0 | 0% Disabled |
| NA / NA | 44C N/A | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 1 MetaX C550 | 1 N/A | 0000:26:00.0 | 0% Disabled |
| NA / NA | 39C N/A | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 2 MetaX C550 | 2 N/A | 0000:63:00.0 | 0% Disabled |
| NA / NA | 39C N/A | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 3 MetaX C550 | 3 N/A | 0000:66:00.0 | 0% Disabled |
| NA / NA | 45C N/A | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 4 MetaX C550 | 4 N/A | 0000:a3:00.0 | 0% Disabled |
| NA / NA | 45C N/A | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 5 MetaX C550 | 5 N/A | 0000:a4:00.0 | 0% Disabled |
| NA / NA | 40C N/A | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 6 MetaX C550 | 6 N/A | 0000:e3:00.0 | 0% Disabled |
| NA / NA | 40C N/A | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 7 MetaX C550 | 7 N/A | 0000:e4:00.0 | 0% Disabled |
| NA / NA | 45C N/A | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
+---------------------------------------------------------------------------------+
| Process: |
| GPU PID Process Name GPU Memory |
| Usage(MiB) |
|=================================================================================|
| no process found |
+---------------------------------------------------------------------------------+
End of Log
6.docker info回显:
End of Log
app@muxi-gpu-server:~$ docker info
Client: Docker Engine - Community
Version: 29.4.0
Context: default
Debug Mode: false
Plugins:
buildx: Docker Buildx (Docker Inc.)
Version: v0.33.0
Path: /usr/libexec/docker/cli-plugins/docker-buildx
compose: Docker Compose (Docker Inc.)
Version: v5.1.3
Path: /usr/libexec/docker/cli-plugins/docker-compose
Server:
Containers: 24
Running: 18
Paused: 0
Stopped: 6
Images: 28
Server Version: 29.4.0
Storage Driver: overlayfs
driver-type: io.containerd.snapshotter.v1
Logging Driver: json-file
Cgroup Driver: systemd
Cgroup Version: 2
Plugins:
Volume: local
Network: bridge host ipvlan macvlan null overlay
Log: awslogs fluentd gcplogs gelf journald json-file local splunk syslog
CDI spec directories:
/etc/cdi
/var/run/cdi
Swarm: inactive
Runtimes: io.containerd.runc.v2 runc
Default Runtime: runc
Init Binary: docker-init
containerd version: 77c84241c7cbdd9b4eca2591793e3d4f4317c590
runc version: v1.3.5-0-g488fc13e
init version: de40ad0
Security Options:
apparmor
seccomp
Profile: builtin
cgroupns
Kernel Version: 6.14.0-27-generic
Operating System: Ubuntu 24.04.4 LTS
OSType: linux
Architecture: x86_64
CPUs: 192
Total Memory: 1008GiB
Name: muxi-gpu-server
ID: f402f581-1f0c-4c03-b987-89406ef8f43c
Docker Root Dir: /data/docker
Debug Mode: false
Experimental: false
Insecure Registries:
harbor.local:9090
::1/128
127.0.0.0/8
Registry Mirrors:
docker.1ms.run/
Live Restore Enabled: false
Firewall Backend: iptables
7.镜像版本:
cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.22.0-maca.ai3.8.0.5-torch2.10-py310-ubuntu22.04-amd64
8.启动容器命令:
docker run -it \
--name vllm-qwen \
--privileged \
--network=host \
-p 40022:40022 \
--ipc=host \
--security-opt seccomp=unconfined \
--security-opt apparmor=unconfined \
--ulimit memlock=-1 \
-v /data/gpustack/models/Qwen3.5-122B-A10B:/data/models/Qwen3.5-122B-A10B \
cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.22.0-maca.ai3.8.0.5-torch2.10-py310-ubuntu22.04-amd64 \
bash
9.容器内执行命令:
vllm serve /data/models/Qwen3.5-122B-A10B \
--pipeline-parallel-size 1 \
--tensor-parallel-size 8 \
--trust-remote-code \
--dtype bfloat16 \
--distributed-executor-backend mp \
--gpu-memory-utilization 0.85 \
--max-model-len 131072 \
--max-num-batched-tokens 131072 \
--no-async-scheduling \
--mm-encoder-tp-mode data \
--mm-processor-cache-type shm \
--limit-mm-per-prompt '{"image": 5, "video": 1}' \
--skip-mm-profiling \
--enable-prefix-caching \
--served-model-name Qwen3.5-122B-A10B \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--host 0.0.0.0 \
--port 40022
二、问题现象
一直无法启动,中间一直报(EngineCore pid=1283) INFO 07-30 10:55:49 [shm_broadcast.py:698] No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization).