你好 这就是崩溃日志
你好 这就是崩溃日志
感谢 目前多卡通讯有问题 只能单卡部署 真是伤脑筋
您好 无异常 现在单卡跑qwen3.6 9b是正常的 但是多卡跑就是一段时间会挂掉
你好,我想只测试gpu1~5 应该怎么操作呢 0 6 7现在部署了别的模型
请问用最新的镜像能解决这个问题么
这是相关的报告文档 我在九月1日更新了驱动
压力测试是正常的 请问怎么解决呢 现在二十小时左右就会崩溃一次
请问有参考命令行吗 mxvs --help无此命令
新华三的服务器 貌似是联想开天的贴牌?KA742H G3
★处理器:2海光7490
★内存:1TB DDR4
★硬盘1:23.84TB SSD
★硬盘2:2960 GB SSD
★网卡: 双口25GbE SFP28 光纤网卡(含模块)
★GPU: 8C500 共计512G显存
★其他:机架导轨
★服务:三年7*24原厂工程师上门售后服务
是不是双卡之间通信有问题
海光
MetaX C500 × 8(单卡 64GB / 65536 MiB)
Kylin-Server-V10-SP3-2403-Release-20240426-x86_64
Kernel: 4.19.90-89.11.v2401.ky10.x86_64
(GRUB 默认启动项已锁定该内核;同机另一内核 4.19.90-89.26 因驱动模块不兼容无法启动)
未开启
Client:
Version: 28.5.1
Context: default
Server:
Containers: 41 Running: 14 Stopped: 27
Images: 19
Server Version: 28.5.1
Storage Driver: overlay2
Cgroup Driver: cgroupfs
Cgroup Version: 1
Runtimes: io.containerd.runc.v2 runc
Default Runtime: runc
Kernel Version: 4.19.90-89.11.v2401.ky10.x86_64
Operating System: Kylin Linux Advanced Server V10 (Halberd)
OSType: linux
Architecture: x86_64
CPUs: 256
Total Memory: 1006 GiB
cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64
8. 启动容器命令
docker run -d \
--device=/dev/mxcd \
--device=/dev/dri/card2 --device=/dev/dri/renderD129 \
--device=/dev/dri/card3 --device=/dev/dri/renderD130 \
--device=/dev/dri/card4 --device=/dev/dri/renderD131 \
--device=/dev/dri/card5 --device=/dev/dri/renderD132 \
--device=/dev/dri/card1 --device=/dev/dri/renderD128 \
--device=/dev/dri/card6 --device=/dev/dri/renderD133 \
--group-add video \
--network=host --name vllm02-27b \
--restart unless-stopped \
--security-opt seccomp=unconfined --security-opt apparmor=unconfined \
--ulimit memlock=-1 \
-v /data:/data -v /data2:/data2 \
cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64 \
sleep infinity
(注:GPU 设备与卡序号映射非顺序:GPU0=card2/renderD129 ... GPU4=card1/renderD128 夹在中间 ... GPU5=card6/renderD133)
```bash
vllm serve /data2/Qwen3.6-27B \
--served-model-name qwen3.6-27b \
--tensor-parallel-size 2 \
--gpu-memory-utilization 0.85 \
--max-model-len 16384 \
--dtype auto \
--port 8016 \
--trust-remote-code \
-enforce-eager
现象二(2026-09-02 起持续至今,Qwen3.6-27B 官方 BF16,tp2,KMD 已升级 3.9.25)
你好 我用QWEN3.6按照手册部署tp2 报错mccl通讯错误
===== 27B 引擎崩溃日志(2026-09-02 17:15,运行约 7 小时后)=====
[APIServer pid=287] INFO: Application shutdown complete.
[APIServer pid=287] INFO: Finished server process [287]
--- 第 1 层:TCPStore 连接被对端重置 ---
Traceback (most recent call last):
File "/opt/conda/lib/python3.10/socket.py", line 1140, in recv_into
return self._sock.recv_into(b)
ConnectionResetError: [Errno 104] Connection reset by peer
The above exception was the direct cause of the following exception:
--- 第 2 层:c10d 分布式心跳监控线程异常 ---
Traceback (most recent call last):
File "/opt/conda/lib/python3.10/site-packages/torch/multiprocessing/.../distributed/c10d/Utils.hpp:653 (most recent call first):
frame #1: c10::Error::Error(...)
(0x7f7fd9a9990a2 in /opt/conda/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
frame #2: c10d::TCPStore::check(std::vector<std::string> const&) + 0x242
(0x7f7fd9a9990a2 in /opt/conda/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x3e8
(0x7f7fcd49cc08 in /opt/conda/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
frame #4: <unknown function> + 0xd3bb5
(0x7f7fe0a6f1b65 in /opt/conda/bin/../lib/libstdc++.so.6)
frame #5: <unknown function> + 0x9448
(0x7f7fe0bc6644 in /lib/x86_64-linux-gnu/libc.so.6)
frame #6: clone + 0x44
(0x7f7fe0bc6b64 in /lib/x86_64-linux-gnu/libc.so.6)
[rank3]:[W002 17:15:02.795520658 TCPStore.cpp:125] [c10d] recvValue failed on
SocketImpl(fd=14, addr=[::ffff:802e:4009:7c7f:0]) Connection closed by peer.
[rank3]:[W002 17:15:02.795520658 ProcessGroupNCCL.cpp:1808] [PG ID 0 PG GUID 0 Rank 3]
Failed to check the "should dump" flag on TCPStore, (maybe TCP Store server has shut down
too early), with error: Failed to recv, got 0 bytes. Connection was likely closed.
Did the remote server shutdown or crash
--- 第 3 层:MCCL(沐曦跨卡通信库)反复连接失败 ---
localhost:829:050 [0] /workspace/out/Release/build/linux/x86_64/mccl/macafify/src/misc/socket.cc:568
MCCL WARN socketStartConnect: Connect to 6.0.63.1:34343 failed: Software caused connection abort
(以上 MCCL WARN 重复出现多次)
[rank3]:[W002 17:15:02.795520658 TCPStore.cpp:106] [c10d] sendBytes failed on
SocketImpl(fd=14, addr=[::]:29057, remote=[ffff:ffff::802e:4009:7c7f:0])
[1137]: Broken pipe
[rank3]:[W002 17:15:03.807503454 ProcessGroupNCCL.cpp:1808] [PG ID 0 PG GUID 0 Rank 3]
Failed to check the "should dump" flag on TCPStore, ... with error: Broken pipe
--- 收尾:共享内存泄漏提示 ---
/opt/conda/lib/python3.10/multiprocessing/resource_tracker.py:224: UserWarning:
resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
===== 崩溃时环境状态(截屏底部)=====
服务器 uptime: 47 hours(即崩溃前服务器已稳定运行 47 小时)
磁盘: / 66%, /boot 27%, /data2 1%, /data3 2%
我看到了 准备用QWEN3.6 27b 请问这个适合tp2吗
你好 目前最新支持到哪个模型呢
1.服务器厂家:海光
2.沐曦GPU型号:C500*8
3.操作系统内核版本:Kylin-Server-V10-SP3-2403-Release-20240426-x86_64
4.是否开启CPU虚拟化:未开启
5.mx-smi回显:
mx-smi version: 2.2.8
=================== MetaX System Management Interface Log ===================
Timestamp : Tue Sep 1 14:55:07 2026
Attached GPUs : 8
+---------------------------------------------------------------------------------+
| MX-SMI 2.2.8 Kernel Mode Driver Version: 3.0.11 |
| MACA Version: 3.2.1.10 BIOS Version: 1.27.5.0 |
|------------------------------------+---------------------+----------------------+
| GPU NAME Persistence-M | Bus-id | GPU-Util sGPU-M |
| Temp Pwr:Usage/Cap Perf | Memory-Usage | GPU-State |
|====================================+=====================+======================|
| 0 MetaX C500 Off | 0000:05:00.0 | 0% Native |
| 44C 59W / 350W P0 | 858/65536 MiB | Available |
+------------------------------------+---------------------+----------------------+
| 1 MetaX C500 Off | 0000:0b:00.0 | 0% Native |
| 46C 62W / 350W P0 | 858/65536 MiB | Available |
+------------------------------------+---------------------+----------------------+
| 2 MetaX C500 Off | 0000:0e:00.0 | 0% Native |
| 44C 58W / 350W P0 | 858/65536 MiB | Available |
+------------------------------------+---------------------+----------------------+
| 3 MetaX C500 Off | 0000:0f:00.0 | 0% Native |
| 47C 63W / 350W P0 | 858/65536 MiB | Available |
+------------------------------------+---------------------+----------------------+
| 4 MetaX C500 Off | 0000:55:00.0 | 0% Native |
| 41C 44W / 350W P0 | 858/65536 MiB | Available |
+------------------------------------+---------------------+----------------------+
| 5 MetaX C500 Off | 0000:56:00.0 | 0% Native |
| 42C 42W / 350W P0 | 858/65536 MiB | Available |
+------------------------------------+---------------------+----------------------+
| 6 MetaX C500 Off | 0000:5b:00.0 | 0% Native |
| 42C 55W / 350W P9 | 56393/65536 MiB | Available |
+------------------------------------+---------------------+----------------------+
| 7 MetaX C500 Off | 0000:5e:00.0 | 0% Native |
| 44C 57W / 350W P9 | 56337/65536 MiB | Available |
+------------------------------------+---------------------+----------------------+
+---------------------------------------------------------------------------------+
| Process: |
| GPU PID Process Name GPU Memory |
| Usage(MiB) |
|=================================================================================|
| 6 4176164 VLLM::EngineCor 55534 |
| 7 4173582 VLLM::EngineCor 55478 |
+---------------------------------------------------------------------------------+
End of Log
6.docker info回显:
Client:
Version: 28.5.1
Context: default
Debug Mode: false
Server:
Containers: 41
Running: 14
Paused: 0
Stopped: 27
Images: 19
Server Version: 28.5.1
Storage Driver: overlay2
Backing Filesystem: xfs
Supports d_type: true
Using metacopy: false
Native Overlay Diff: true
userxattr: false
Logging Driver: json-file
Cgroup Driver: cgroupfs
Cgroup Version: 1
Plugins:
Volume: local
Network: bridge host ipvlan macvlan null overlay
Log: awslogs fluentd gcplogs gelf journald json-file local splunk syslog
CDI spec directories:
/etc/cdi
/var/run/cdi
Swarm: inactive
Runtimes: io.containerd.runc.v2 runc
Default Runtime: runc
Init Binary: docker-init
containerd version: b98a3aace656320842a23f4a392a33f46af97866
runc version: v1.3.0-0-g4ca628d
init version: de40ad0
Security Options:
seccomp
Profile: builtin
Kernel Version: 4.19.90-89.11.v2401.ky10.x86_64
Operating System: Kylin Linux Advanced Server V10 (Halberd)
OSType: linux
Architecture: x86_64
CPUs: 256
Total Memory: 1006GiB
Name: localhost.localdomain
ID: 4e75250f-0952-4b17-a2fe-a49d1b7cedc8
Docker Root Dir: /var/lib/docker
Debug Mode: false
Experimental: false
Insecure Registries:
::1/128
127.0.0.0/8
Registry Mirrors:
docker.m.daocloud.io/
Live Restore Enabled: false
Product License: Community Engine
7.镜像版本:
cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64
8.启动容器命令:
docker run -d \
--device=/dev/mxcd \
--device=/dev/dri/card2 --device=/dev/dri/renderD129 \
--device=/dev/dri/card3 --device=/dev/dri/renderD130 \
--device=/dev/dri/card4 --device=/dev/dri/renderD131 \
--device=/dev/dri/card5 --device=/dev/dri/renderD132 \
--device=/dev/dri/card1 --device=/dev/dri/renderD128 \
--device=/dev/dri/card6 --device=/dev/dri/renderD133 \
--group-add video \
--network=host --name vllm02-27b \
--restart unless-stopped \
--security-opt seccomp=unconfined --security-opt apparmor=unconfined \
--ulimit memlock=-1 \
-v /data:/data \
cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64 \
9.容器内执行命令:
vllm serve /data/Qwen3.5-27B-W8A8 \
--served-model-name qwen3.5-27b \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.85 \
--max-model-len 8192 \
--dtype half \
--port 8016 \
--trust-remote-code \
--reasoning-parser qwen3