你好 我用QWEN3.6按照手册部署tp2 报错mccl通讯错误
===== 27B 引擎崩溃日志(2026-09-02 17:15,运行约 7 小时后)=====
[APIServer pid=287] INFO: Application shutdown complete.
[APIServer pid=287] INFO: Finished server process [287]
--- 第 1 层:TCPStore 连接被对端重置 ---
Traceback (most recent call last):
File "/opt/conda/lib/python3.10/socket.py", line 1140, in recv_into
return self._sock.recv_into(b)
ConnectionResetError: [Errno 104] Connection reset by peer
The above exception was the direct cause of the following exception:
--- 第 2 层:c10d 分布式心跳监控线程异常 ---
Traceback (most recent call last):
File "/opt/conda/lib/python3.10/site-packages/torch/multiprocessing/.../distributed/c10d/Utils.hpp:653 (most recent call first):
frame #1: c10::Error::Error(...)
(0x7f7fd9a9990a2 in /opt/conda/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
frame #2: c10d::TCPStore::check(std::vector<std::string> const&) + 0x242
(0x7f7fd9a9990a2 in /opt/conda/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x3e8
(0x7f7fcd49cc08 in /opt/conda/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
frame #4: <unknown function> + 0xd3bb5
(0x7f7fe0a6f1b65 in /opt/conda/bin/../lib/libstdc++.so.6)
frame #5: <unknown function> + 0x9448
(0x7f7fe0bc6644 in /lib/x86_64-linux-gnu/libc.so.6)
frame #6: clone + 0x44
(0x7f7fe0bc6b64 in /lib/x86_64-linux-gnu/libc.so.6)
[rank3]:[W002 17:15:02.795520658 TCPStore.cpp:125] [c10d] recvValue failed on
SocketImpl(fd=14, addr=[::ffff:802e:4009:7c7f:0]) Connection closed by peer.
[rank3]:[W002 17:15:02.795520658 ProcessGroupNCCL.cpp:1808] [PG ID 0 PG GUID 0 Rank 3]
Failed to check the "should dump" flag on TCPStore, (maybe TCP Store server has shut down
too early), with error: Failed to recv, got 0 bytes. Connection was likely closed.
Did the remote server shutdown or crash
--- 第 3 层:MCCL(沐曦跨卡通信库)反复连接失败 ---
localhost:829:050 [0] /workspace/out/Release/build/linux/x86_64/mccl/macafify/src/misc/socket.cc:568
MCCL WARN socketStartConnect: Connect to 6.0.63.1:34343 failed: Software caused connection abort
(以上 MCCL WARN 重复出现多次)
[rank3]:[W002 17:15:02.795520658 TCPStore.cpp:106] [c10d] sendBytes failed on
SocketImpl(fd=14, addr=[::]:29057, remote=[ffff:ffff::802e:4009:7c7f:0])
[1137]: Broken pipe
[rank3]:[W002 17:15:03.807503454 ProcessGroupNCCL.cpp:1808] [PG ID 0 PG GUID 0 Rank 3]
Failed to check the "should dump" flag on TCPStore, ... with error: Broken pipe
--- 收尾:共享内存泄漏提示 ---
/opt/conda/lib/python3.10/multiprocessing/resource_tracker.py:224: UserWarning:
resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
===== 崩溃时环境状态(截屏底部)=====
服务器 uptime: 47 hours(即崩溃前服务器已稳定运行 47 小时)
磁盘: / 66%, /boot 27%, /data2 1%, /data3 2%