我看最新的镜像有llm-d的标签,llm-d部署相关的教程有吗
我想在单卡沐曦 N260 上部署 llama.cpp,加载 GGUF 模型,并使用 llama-server 提供 OpenAI 兼容 API,希望确认目前是否有经过验证的部署方案
镜像:
vllm-metax:0.25.0-maca.ai3.8.2.108-torch2.10-py310-kylinv11-amd64
模型:
DeepSeek-V4-Flash-0731-W8A8
配置:
TP=8
Expert Parallel enabled
max_num_seqs=8
max_num_batched_tokens=8192
现象:
默认 CUDA Graph 模式在 capture_model/_capture_cudagraphs 阶段,
MacaGemmMoe 内核触发 Xnack Error/ATU Fault 和 mcErrorIllegalAddress。
临时规避:
加入 --enforce-eager 后模型启动、普通对话及结构化工具调用均正常。
如果后续必须追求 CUDA Graph 性能,是不是要等新的镜像更新还是有别的方式可以解决?
疑问如题,希望可以解释得详细点
看到官网提供适配的arm推理镜像版本都比较老,如果arm机器要部署qwen3.6该如何部署呢
使用两台服务器、每台8张沐曦 C500,通过 vLLM + Ray 双机部署 DeepSeek-R1-0528。
模型可以正常启动并提供推理服务,但连续运行一段时间后,正在执行的请求停止生成 Token。等待约300秒后,vLLM报 RayChannelTimeoutError,随后 node2 的 Raylet 在 experimental_mutable_object_provider.cc:154 触发致命断言并退出,最终导致整个模型服务不可用。
历史上该问题曾在运行数小时后出现;本次完整采集的复现中,Ray Session 大约在15:27启动,16:32发生故障,约运行64分钟后复现。
二、软硬件环境
本次采集时,两台服务器均能正常识别8张GPU。GPU拓扑为每4张卡通过 MetaXLink 互联,两个四卡组之间跨NUMA/CPU互联。
MX-SMI:
2.2.12
Kernel Mode Driver:
3.6.17
MACA:
3.5.3.18
驱动和MACA版本来自 node1 的 mx-smi 输出。
Ray:2.53.0
vLLM:0.14.0
PyTorch:2.8.0+metax3.5.3.9
Python:3.10
两台节点容器内的软件版本一致。
运行时日志中确认的主要配置如下:
模型:
/mnt/data/models/DeepSeek-R1-0528
dtype:
torch.bfloat16
quantization:
compressed-tensors
max_seq_len:
32768
tensor_parallel_size:
8
pipeline_parallel_size:
2
data_parallel_size:
1
disable_custom_all_reduce:
True
enable_prefix_caching:
True
enable_chunked_prefill:
True
enforce_eager:
False
Compilation Mode:
VLLM_COMPILE
CUDAGraph Mode:
FULL_AND_PIECEWISE
即单机内部采用 TP=8,两台机器采用 PP=2,共使用16张GPU。运行过程中启用了 vLLM Compile、CUDAGraph 和 Ray Compiled DAG 相关执行路径。
Ray、Gloo 和 Socket 通信使用:
网卡:p5p1
node1:
192.168.1.204
node2:
192.168.1.205
当前速率:
1000Mb/s Full Duplex
MTU:
1500
故障后两台机器之间 Ping 正常,丢包率为0%,往返延迟约0.1ms。node1 到 node2 的路由明确经过 p5p1。
MCCL日志显示跨节点通信使用:
NET/IB
GDRDMA
故障清理阶段出现了到 192.168.101.24 的 MCCL Socket 连接异常。
故障前生成速度正常:
16:26:31 generation throughput:50.3 tokens/s
16:26:41 generation throughput:51.2 tokens/s
16:26:51 generation throughput:50.3 tokens/s
16:27:01 generation throughput:51.0 tokens/s
随后吞吐快速下降:
16:27:11
Avg generation throughput:10.5 tokens/s
Running:2
Waiting:0
16:27:21
Avg generation throughput:0.0 tokens/s
Running:2
Waiting:0
GPU KV cache usage:0.6%
当时仍显示有两个请求正在运行,没有等待请求,但已经不再生成Token。
停止生成约5分钟后,EngineCore报错:
ray.exceptions.RayChannelTimeoutError:
System error: Timed out waiting for object available to read.
ObjectID:
004fc8a14e43c330d8a5ea50e2389bc5b2cba1de0100000002e1f505
随后日志提示:
If the execution is expected to take a long time,
increase RAY_CGRAPH_get_timeout which is currently 300 seconds.
Otherwise, this may indicate that the execution is hanging.
调用栈位于:
ray/dag/compiled_dag_node.py
ray/experimental/channel/common.py
ray/experimental/channel/shared_memory_channel.py
ray._raylet.CoreWorker.get_objects
说明 EngineCore 在 Ray Compiled DAG/Shared Memory Channel 中等待执行结果超过300秒。
超时后立即出现:
Shutting down Ray distributed executor
Tearing down compiled DAG
Cancelling compiled worker
Waiting for worker tasks to exit
API请求开始返回:
HTTP/1.1 500 Internal Server Error
EngineDeadError
在清理 Compiled DAG 和取消 Worker 时,RayWorkerWrapper报错:
raylet_client.cc:202:
Error pushing mutable object:
RpcError: RPC error: Socket closed
rpc_code: 14
发生故障的 Raylet 位于:
ip=192.168.1.205
即 node2。
完整关键错误:
experimental_mutable_object_provider.cc:154:
An unexpected system state has occurred.
You have likely discovered a bug in Ray.
Check failed:
object_manager_->WriteAcquire(
info.local_object_id,
total_data_size,
nullptr,
total_metadata_size,
info.num_readers,
object_backing_store
)
Status not OK:
ChannelError: Channel closed.
调用栈关键位置:
ray::core::experimental::MutableObjectProvider::HandlePushMutableObject
ray::raylet::NodeManager::HandlePushMutableObject
ray::rpc::ServerCallImpl<>::HandleRequestImpl
该错误触发 Raylet异常退出。
Raylet异常后,node1上的MCCL日志出现:
MCCL WARN socketProgressOpt:
Call to recv from 192.168.101.24 failed:
Connection reset by peer
后续还有:
MCCL WARN socketStartConnect:
Connect to 192.168.101.24 failed:
Software caused connection abort
这些MCCL错误发生在 RayChannelTimeout 和 node2 Raylet异常之后,因此目前判断更像是进程退出和通信清理产生的后续报错。
node1的 GCS Server记录:
Node is dead because the health check failed.
death reason:
UNEXPECTED_TERMINATION
death message:
health check failed due to missing too many heartbeats
node_name:
192.168.1.205
随后大量Actor报:
Node with actor is already dead
Disconnected: GRPC client is shut down
两台服务器物理内存均为1TiB。
node2故障后采集结果:
物理内存总量:1.0TiB
可用内存:995GiB
Swap使用量:0
容器内存使用:
317.1GiB / 1006GiB
Docker Memory Limit:
0,未限制
Docker OOMKilled:
false
Docker ExitCode:
0
容器 ShmSize:
107374182400 bytes,约100GiB
node1容器状态同样为:
Container state:running
OOMKilled:false
ExitCode:0
Memory limit:0
ShmSize:107374182400
Ray Object Store配置约为:
object_store_memory:
约102GB
plasma_directory:
/dev/shm
故障后两台宿主机最近12小时均未发现明确的 Linux OOM Killer记录。
因此目前基本排除:
宿主机物理内存不足
Docker内存限制触发
容器被OOM Kill
/dev/shm容量不足
磁盘空间不足
虽然本次16:32故障附近没有新的Link Down记录,但 p5p1 确实存在历史链路不稳定问题。
node1在最近12小时记录到8次Link Down,其中几次恢复后只协商到100Mbps:
09:42:11 Link Down
09:42:14 Link Up 100 Mbps
09:43:03 Link Down
09:45:22 Link Up 100 Mbps
09:47:56 Link Down
09:47:59 Link Up 100 Mbps
之后恢复到1000Mbps。
模型本轮启动前最近一次为:
15:21:51 Link Down
15:21:54 Link Up 1000 Mbps
node2也记录到:
15:20:59 Link Down
15:21:03 Link Up 1000 Mbps
但本次模型停止生成发生在16:27,Ray报错发生在16:32,因此现有时间线不能证明本次故障由 p5p1 Link Down 直接触发。
目前可以确认的故障链路为:
模型正在运行
↓
某个Ray Worker、GPU Rank、MCCL执行或Compiled DAG执行静默卡住
↓
连续约300秒没有产生可读取的Object
↓
RayChannelTimeoutError
↓
vLLM关闭Ray Distributed Executor
↓
Compiled DAG开始teardown并关闭Channel
↓
Worker继续push mutable object
↓
Socket closed rpc_code=14
↓
node2 Raylet在WriteAcquire时遇到Channel closed
↓
experimental_mutable_object_provider.cc触发CHECK failed
↓
node2 Raylet异常退出
↓
GCS将node2标记为dead
↓
整个vLLM服务down
目前尚不能确定最初300秒卡住的底层原因是:
1. Ray 2.53.0 Compiled DAG / Mutable Object Channel自身问题;
2. vLLM 0.14.0与当前Ray版本的兼容性问题;
3. 某个MetaX GPU Worker或GPU Rank静默卡死;
4. MCCL/RDMA Collective静默阻塞;
烦请官方协助确认以下问题:
vLLM 0.14.0 + Ray 2.53.0 + torch 2.8.0+metax3.5.3.9 是否为当前推荐和验证过的版本组合?
当前版本中,Ray Compiled DAG、Shared Memory Channel和Mutable Object Channel是否存在已知稳定性问题?
以下Raylet错误是否属于已知问题?
experimental_mutable_object_provider.cc:154
WriteAcquire(...)
ChannelError: Channel closed
在vLLM已经开始 teardown 后,Raylet收到 PushMutableObject 并直接触发 CHECK failed,是否可能是 Ray Compiled DAG teardown过程中的竞态问题?
是否有官方推荐的 Ray、vLLM、MACA、驱动版本组合,可以规避该问题?
如何进一步定位是哪个GPU Rank、MCCL Collective或Ray Worker在16:27左右首先卡住?是否有推荐的MCCL超时、异步错误检测或Worker Watchdog环境变量?
RAY_CGRAPH_get_timeout=300 是否建议调整?本次只有两个运行请求、KV Cache使用率仅约0.6%,但执行完全停止,因此暂时认为单纯增大超时时间只能延迟失败,无法解决底层卡死。
log见附件
我在六机上进行glm5.1 PD分离操作,启动decode节点报错,报错内容如文件
如标题,glm5.1-w8a8做PD分离,decode节点最少需要几个?
metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$ bash mccl.sh 8
The test is all_reduce_perf, the maca version is /opt/maca-3.7.1
main_process = 7324
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7325: Test failure common.cu:1271
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7326: Test failure common.cu:1271
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7327: Test failure common.cu:1271
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7328: Test failure common.cu:1271
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7329: Test failure common.cu:1271
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7330: Test failure common.cu:1271
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7331: Test failure common.cu:1271
===============================
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7324: Test failure common.cu:1271
Primary job terminated normally, but 1 process returned
a non-zero exit code. Per user-direction, the job has been aborted.
mpirun detected that one or more processes exited with non-zero status, thus causing
the job to be terminated. The first process to do so was:
Process name: [[25858,1],4]
Exit code: 2
metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$ bash mccl.sh 2
The test is all_reduce_perf, the maca version is /opt/maca-3.7.1
main_process = 7378
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7379: Test failure common.cu:1271
===============================
metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
.. metax-host-104 pid 7378: Test failure common.cu:1271
Primary job terminated normally, but 1 process returned
a non-zero exit code. Per user-direction, the job has been aborted.
mpirun detected that one or more processes exited with non-zero status, thus causing
the job to be terminated. The first process to do so was:
Process name: [[25940,1],0]
Exit code: 2
metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$
单机mccl测试失败是怎么回事
mx-smi回显信息正常如下
metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$ mx-smi
mx-smi version: 2.3.1
=================== MetaX System Management Interface Log ===================
Timestamp : Wed Jun 17 10:13:47 2026
Attached GPUs : 8
+---------------------------------------------------------------------------------+
| MX-SMI 2.3.1 Kernel Mode Driver Version: 3.8.1 |
| MACA Version: 3.7.1.5 BIOS Version: 1.29.1.0 |
|------------------+-----------------+---------------------+----------------------|
| Board Name | GPU Persist-M | Bus-id | GPU-Util sGPU-M |
| Pwr:Usage/Cap | Temp Perf | Memory-Usage | GPU-State |
|==================+=================+=====================+======================|
| 0 MetaX C550 | 0 Off | 0000:2b:00.0 | 0% Disabled |
| 54W / 450W | 32C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 1 MetaX C550 | 1 Off | 0000:3a:00.0 | 0% Disabled |
| 56W / 450W | 33C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 2 MetaX C550 | 2 Off | 0000:4d:00.0 | 0% Disabled |
| 52W / 450W | 33C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 3 MetaX C550 | 3 Off | 0000:5c:00.0 | 0% Disabled |
| 56W / 450W | 33C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 4 MetaX C550 | 4 Off | 0000:aa:00.0 | 0% Disabled |
| 53W / 450W | 32C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 5 MetaX C550 | 5 Off | 0000:ba:00.0 | 0% Disabled |
| 52W / 450W | 33C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 6 MetaX C550 | 6 Off | 0000:ca:00.0 | 0% Disabled |
| 54W / 450W | 34C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 7 MetaX C550 | 7 Off | 0000:da:00.0 | 0% Disabled |
| 53W / 450W | 33C P0 | 858/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
+---------------------------------------------------------------------------------+
| Process: |
| GPU PID Process Name GPU Memory |
| Usage(MiB) |
|=================================================================================|
| no process found |
+---------------------------------------------------------------------------------+
End of Log
metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$
双卡16卡开启sgpu后,双机部署模型的性能会比不开启sgpu双击部署模型的性能低吗
如果要在裸金属宿主机环境下进行模型部署,安装maca-vllm-metax-0.19.0-py310-3.5.3.502-linux-x86_64.tar.xz过后,官方的vllm还需要安装吗?
缺陷报告:MCCL 跨机 RoCE 内存注册失败 (ibv_reg_mr Invalid argument)
【环境信息】
硬件: 双节点共 16 张沐曦 MetaX C500 (8卡/节点),机头是两台H3c uniserver R5330 G7(C500*16)
网络: 40Gbps RoCE 网卡 (设备名 rocep6s0, rocep95s0)---这是ibstat显示的信息,网卡用的是400G的迈洛斯 cx7
软件: MACA 3.5.3 / MCCL 2.16.5
【问题现象】
在使用 mpirun 执行跨机 all_reduce_perf 测试时,若开启 IB/RoCE 硬件加速,MCCL 在初始化阶段崩溃,报错:
MCCL WARN Call to ibv_reg_mr failed with error Invalid argument
【已完成的排查与隔离】
网络层正常: RoCE 高速网段(192.168.100.x)跨机 ping 测试延迟 < 0.1ms,双端 firewalld 与 SELinux 已关闭。
系统限制正常: 双端通过 MPI 验证 ulimit -l 均为 unlimited,排除 memlock 限制导致的问题。
iova2 降级测试: 增加 -x MCCL_IB_PCI_RELAXED_ORDERING=0 后,报错从 ibv_reg_mr_iova2 failed 退化为 ibv_reg_mr failed,依然返回 Invalid argument (retcode 2)。
TCP 降级对照组(核心证据): 增加 -x MCCL_IB_DISABLE=1 强制走普通 TCP/Socket 通信后,测试完美通过(#wrong 0)。
mccl测试脚本内容:
[root@localhost /opt/maca/samples/mccl_tests/perf]# cat cluster.sh
MACA_PATH="${MACA_PATH:-/opt/maca}"
HOST_IP=192.168.1.204:8,192.168.1.205:8
GPU_NUM=16
TEST_DIR=$MACA_PATH/samples/mccl_tests/perf/mccl_perf
BENCH_NAMES="all_reduce_perf"
if [[ -z "$1" || -z "$2" || -z "$3" ]]; then
echo "Use the default ip addr. Run with parameters for custom ip addr, for example: bash cluster.sh ip_1:proc_count,ip_2:proc_count gpu_num test_name"
else
HOST_IP=$1
GPU_NUM=$2
if [ "$3" = "all" ]; then
BENCH_NAMES="all_reduce_perf all_gather_perf reduce_scatter_perf sendrecv_perf alltoall_perf"
else
if [ -e "$TEST_DIR/$3" ]; then
BENCH_NAMES=$3
else
echo "$TEST_DIR/$3 dose not exist!"
exit 1
fi
fi
fi
IP_MASK="192.168.100.0/24"
IB_PORT=rocep6s0,rocep95s0
PERF_ENV="-x FORCE_ACTIVE_WAIT=2"
LIB_PATH_ENV="-x MACA_PATH=${MACA_PATH} -x LD_LIBRARY_PATH=${MACA_PATH}/lib:/${MACA_PATH}/ompi/lib:/${MACA_PATH}/ucx/lib"
ENV_VAR="-x MCCL_IB_HCA=rocep6s0,rocep95s0 -x MCCL_SOCKET_IFNAME=p50p1,p51p1 -x MCCL_CROSS_NIC=1 ${PERF_ENV} ${LIB_PATH_ENV} -x MCCL_IB_DISABLE=0"
MPI_PROCESS_NUM=${GPU_NUM}
MPI_RUN_OPT="--allow-run-as-root -mca btl_tcp_if_include ${IP_MASK} -mca oob_tcp_if_include ${IP_MASK} -mca pml ^ucx -mca osc ^ucx -mca btl ^openib"
for BENCH in ${BENCH_NAMES}; do
echo -n "The test is ${BENCH}, the maca version is " && realpath ${MACA_PATH}
${MACA_PATH}/ompi/bin/mpirun -np ${MPI_PROCESS_NUM} ${MPI_RUN_OPT} -host ${HOST_IP} ${ENV_VAR} ${TEST_DIR}/${BENCH} -b 1K -e 1G -d float -f 2 -g 1 -n 10
done
报错信息如附件内容
请问在进行 ray配置的时候,如果机器有ib卡,是不是必须要要映射计算网口?
比如以下模型那些支持,那些不支持:
Qwen3-235B
Qwen3-VL
Qwen3-embeding
Qwen3-rerank
Z-Image
GLM4.6/7
Deepseek-V3.2
Deepseek-V3
Deepseek-R1
安装信息:python -m pip install paddle-metax-gpu==3.3.0 -i www.paddlepaddle.org.cn/packages/stable/maca/
模型运行脚本如上传文件:
使用模型:
PP-OCRv5_server_det
PP-LCNet_x1_0_doc_ori
PP-LCNet_x1_0_textline_ori
PP-OCRv5_server_rec
UVDoc