MetaX-Tech Developer Forum 论坛首页
  • 沐曦开发者
search
Sign in

lishuai

  • Members
  • Joined 2025年12月19日
  • message 帖子
  • forum 主题
  • favorite 关注者
  • favorite_border Follows
  • person_outline 详细信息

lishuai has started 19 threads.

  • See post chevron_right
    lishuai
    Members
    llm-d相关 已解决 2026年9月24日 09:02

    我看最新的镜像有llm-d的标签,llm-d部署相关的教程有吗

  • See post chevron_right
    lishuai
    Members
    N260 是否支持 llama.cpp GPU 推理,有没有经过验证的版本和部署方式 已解决 2026年9月17日 12:47

    我想在单卡沐曦 N260 上部署 llama.cpp,加载 GGUF 模型,并使用 llama-server 提供 OpenAI 兼容 API,希望确认目前是否有经过验证的部署方案

  • See post chevron_right
    lishuai
    Members
    DeepSeek-V4-Flash-0731-W8A8问题 已解决 2026年9月14日 15:31

    镜像:
    vllm-metax:0.25.0-maca.ai3.8.2.108-torch2.10-py310-kylinv11-amd64

    模型:
    DeepSeek-V4-Flash-0731-W8A8

    配置:
    TP=8
    Expert Parallel enabled
    max_num_seqs=8
    max_num_batched_tokens=8192

    现象:
    默认 CUDA Graph 模式在 capture_model/_capture_cudagraphs 阶段,
    MacaGemmMoe 内核触发 Xnack Error/ATU Fault 和 mcErrorIllegalAddress。

    临时规避:
    加入 --enforce-eager 后模型启动、普通对话及结构化工具调用均正常。

    如果后续必须追求 CUDA Graph 性能,是不是要等新的镜像更新还是有别的方式可以解决?

  • See post chevron_right
    lishuai
    Members
    modelzoo的vllm镜像和vllm-metax的镜像有什么区别? 已解决 2026年9月9日 09:16

    疑问如题,希望可以解释得详细点

  • See post chevron_right
    lishuai
    Members
    SRIOV相关问题 已解决 2026年8月26日 17:20

    请问,SRIOV能实现的效果是类似与sgpu的效果吗?

  • See post chevron_right
    lishuai
    Members
    arm镜像问题 已解决 2026年8月14日 14:35

    看到官网提供适配的arm推理镜像版本都比较老,如果arm机器要部署qwen3.6该如何部署呢

  • See post chevron_right
    lishuai
    Members
    双机部署模型运行一段时间后ray报错 已解决 2026年7月23日 18:28

    使用两台服务器、每台8张沐曦 C500,通过 vLLM + Ray 双机部署 DeepSeek-R1-0528。

    模型可以正常启动并提供推理服务,但连续运行一段时间后,正在执行的请求停止生成 Token。等待约300秒后,vLLM报 RayChannelTimeoutError,随后 node2 的 Raylet 在 experimental_mutable_object_provider.cc:154 触发致命断言并退出,最终导致整个模型服务不可用。

    历史上该问题曾在运行数小时后出现;本次完整采集的复现中,Ray Session 大约在15:27启动,16:32发生故障,约运行64分钟后复现。

    二、软硬件环境

    1. 服务器与GPU
      服务器数量:2台
      每台GPU数量:8张
      GPU型号:MetaX C500
      单卡显存:65536 MiB
      总GPU数量:16张
      cpu:海光7390*2

    本次采集时,两台服务器均能正常识别8张GPU。GPU拓扑为每4张卡通过 MetaXLink 互联,两个四卡组之间跨NUMA/CPU互联。

    1. 系统和驱动
      Kernel:
      4.19.90-52.23.v2207.gfb08.ky10.x86_64

    MX-SMI:
    2.2.12

    Kernel Mode Driver:
    3.6.17

    MACA:
    3.5.3.18
    驱动和MACA版本来自 node1 的 mx-smi 输出。

    3. 推理软件版本

    Ray:2.53.0
    vLLM:0.14.0
    PyTorch:2.8.0+metax3.5.3.9
    Python:3.10
    

    两台节点容器内的软件版本一致。

    三、模型部署配置

    运行时日志中确认的主要配置如下:

    模型:
    /mnt/data/models/DeepSeek-R1-0528
    
    dtype:
    torch.bfloat16
    
    quantization:
    compressed-tensors
    
    max_seq_len:
    32768
    
    tensor_parallel_size:
    8
    
    pipeline_parallel_size:
    2
    
    data_parallel_size:
    1
    
    disable_custom_all_reduce:
    True
    
    enable_prefix_caching:
    True
    
    enable_chunked_prefill:
    True
    
    enforce_eager:
    False
    
    Compilation Mode:
    VLLM_COMPILE
    
    CUDAGraph Mode:
    FULL_AND_PIECEWISE
    

    即单机内部采用 TP=8,两台机器采用 PP=2,共使用16张GPU。运行过程中启用了 vLLM Compile、CUDAGraph 和 Ray Compiled DAG 相关执行路径。

    四、网络配置

    Ray、Gloo 和 Socket 通信使用:

    网卡:p5p1
    
    node1:
    192.168.1.204
    
    node2:
    192.168.1.205
    
    当前速率:
    1000Mb/s Full Duplex
    
    MTU:
    1500
    

    故障后两台机器之间 Ping 正常,丢包率为0%,往返延迟约0.1ms。node1 到 node2 的路由明确经过 p5p1。

    MCCL日志显示跨节点通信使用:

    NET/IB
    GDRDMA
    

    故障清理阶段出现了到 192.168.101.24 的 MCCL Socket 连接异常。

    五、问题现象及完整时间线

    1. 16:27左右模型停止产生Token

    故障前生成速度正常:

    16:26:31  generation throughput:50.3 tokens/s
    16:26:41  generation throughput:51.2 tokens/s
    16:26:51  generation throughput:50.3 tokens/s
    16:27:01  generation throughput:51.0 tokens/s
    

    随后吞吐快速下降:

    16:27:11
    Avg generation throughput:10.5 tokens/s
    Running:2
    Waiting:0
    
    16:27:21
    Avg generation throughput:0.0 tokens/s
    Running:2
    Waiting:0
    GPU KV cache usage:0.6%
    

    当时仍显示有两个请求正在运行,没有等待请求,但已经不再生成Token。

    2. 16:32:03触发 RayChannelTimeout

    停止生成约5分钟后,EngineCore报错:

    ray.exceptions.RayChannelTimeoutError:
    System error: Timed out waiting for object available to read.
    
    ObjectID:
    004fc8a14e43c330d8a5ea50e2389bc5b2cba1de0100000002e1f505
    

    随后日志提示:

    If the execution is expected to take a long time,
    increase RAY_CGRAPH_get_timeout which is currently 300 seconds.
    
    Otherwise, this may indicate that the execution is hanging.
    

    调用栈位于:

    ray/dag/compiled_dag_node.py
    ray/experimental/channel/common.py
    ray/experimental/channel/shared_memory_channel.py
    ray._raylet.CoreWorker.get_objects
    

    说明 EngineCore 在 Ray Compiled DAG/Shared Memory Channel 中等待执行结果超过300秒。

    3. vLLM开始关闭 Ray Distributed Executor

    超时后立即出现:

    Shutting down Ray distributed executor
    Tearing down compiled DAG
    Cancelling compiled worker
    Waiting for worker tasks to exit
    

    API请求开始返回:

    HTTP/1.1 500 Internal Server Error
    EngineDeadError
    

    4. 16:32:03.673出现 Socket closed

    在清理 Compiled DAG 和取消 Worker 时,RayWorkerWrapper报错:

    raylet_client.cc:202:
    Error pushing mutable object:
    RpcError: RPC error: Socket closed
    rpc_code: 14
    

    5. 16:32:04.443,node2 Raylet触发致命断言

    发生故障的 Raylet 位于:

    ip=192.168.1.205
    

    即 node2。

    完整关键错误:

    experimental_mutable_object_provider.cc:154:
    
    An unexpected system state has occurred.
    You have likely discovered a bug in Ray.
    
    Check failed:
    object_manager_->WriteAcquire(
        info.local_object_id,
        total_data_size,
        nullptr,
        total_metadata_size,
        info.num_readers,
        object_backing_store
    )
    
    Status not OK:
    ChannelError: Channel closed.
    

    调用栈关键位置:

    ray::core::experimental::MutableObjectProvider::HandlePushMutableObject
    ray::raylet::NodeManager::HandlePushMutableObject
    ray::rpc::ServerCallImpl<>::HandleRequestImpl
    

    该错误触发 Raylet异常退出。

    6. 16:32:05出现MCCL连接中断

    Raylet异常后,node1上的MCCL日志出现:

    MCCL WARN socketProgressOpt:
    Call to recv from 192.168.101.24 failed:
    Connection reset by peer
    

    后续还有:

    MCCL WARN socketStartConnect:
    Connect to 192.168.101.24 failed:
    Software caused connection abort
    

    这些MCCL错误发生在 RayChannelTimeout 和 node2 Raylet异常之后,因此目前判断更像是进程退出和通信清理产生的后续报错。

    7. 16:32:18,Head节点将node2标记为死亡

    node1的 GCS Server记录:

    Node is dead because the health check failed.
    
    death reason:
    UNEXPECTED_TERMINATION
    
    death message:
    health check failed due to missing too many heartbeats
    
    node_name:
    192.168.1.205
    

    随后大量Actor报:

    Node with actor is already dead
    Disconnected: GRPC client is shut down
    

    六、内存和共享内存情况

    两台服务器物理内存均为1TiB。

    node2故障后采集结果:

    物理内存总量:1.0TiB
    可用内存:995GiB
    Swap使用量:0
    
    容器内存使用:
    317.1GiB / 1006GiB
    
    Docker Memory Limit:
    0,未限制
    
    Docker OOMKilled:
    false
    
    Docker ExitCode:
    0
    
    容器 ShmSize:
    107374182400 bytes,约100GiB
    

    node1容器状态同样为:

    Container state:running
    OOMKilled:false
    ExitCode:0
    Memory limit:0
    ShmSize:107374182400
    

    Ray Object Store配置约为:

    object_store_memory:
    约102GB
    
    plasma_directory:
    /dev/shm
    

    故障后两台宿主机最近12小时均未发现明确的 Linux OOM Killer记录。

    因此目前基本排除:

    宿主机物理内存不足
    Docker内存限制触发
    容器被OOM Kill
    /dev/shm容量不足
    磁盘空间不足
    

    七、p5p1网卡异常情况

    虽然本次16:32故障附近没有新的Link Down记录,但 p5p1 确实存在历史链路不稳定问题。

    node1在最近12小时记录到8次Link Down,其中几次恢复后只协商到100Mbps:

    09:42:11 Link Down
    09:42:14 Link Up 100 Mbps
    
    09:43:03 Link Down
    09:45:22 Link Up 100 Mbps
    
    09:47:56 Link Down
    09:47:59 Link Up 100 Mbps
    

    之后恢复到1000Mbps。

    模型本轮启动前最近一次为:

    15:21:51 Link Down
    15:21:54 Link Up 1000 Mbps
    

    node2也记录到:

    15:20:59 Link Down
    15:21:03 Link Up 1000 Mbps
    

    但本次模型停止生成发生在16:27,Ray报错发生在16:32,因此现有时间线不能证明本次故障由 p5p1 Link Down 直接触发。

    八、当前初步判断

    目前可以确认的故障链路为:

    模型正在运行
        ↓
    某个Ray Worker、GPU Rank、MCCL执行或Compiled DAG执行静默卡住
        ↓
    连续约300秒没有产生可读取的Object
        ↓
    RayChannelTimeoutError
        ↓
    vLLM关闭Ray Distributed Executor
        ↓
    Compiled DAG开始teardown并关闭Channel
        ↓
    Worker继续push mutable object
        ↓
    Socket closed rpc_code=14
        ↓
    node2 Raylet在WriteAcquire时遇到Channel closed
        ↓
    experimental_mutable_object_provider.cc触发CHECK failed
        ↓
    node2 Raylet异常退出
        ↓
    GCS将node2标记为dead
        ↓
    整个vLLM服务down
    

    目前尚不能确定最初300秒卡住的底层原因是:
    1. Ray 2.53.0 Compiled DAG / Mutable Object Channel自身问题;
    2. vLLM 0.14.0与当前Ray版本的兼容性问题;
    3. 某个MetaX GPU Worker或GPU Rank静默卡死;
    4. MCCL/RDMA Collective静默阻塞;

    九、希望沐曦官方协助确认的问题

    烦请官方协助确认以下问题:

    1. vLLM 0.14.0 + Ray 2.53.0 + torch 2.8.0+metax3.5.3.9 是否为当前推荐和验证过的版本组合?

    2. 当前版本中,Ray Compiled DAG、Shared Memory Channel和Mutable Object Channel是否存在已知稳定性问题?

    3. 以下Raylet错误是否属于已知问题?
      experimental_mutable_object_provider.cc:154
      WriteAcquire(...)
      ChannelError: Channel closed

    4. 在vLLM已经开始 teardown 后,Raylet收到 PushMutableObject 并直接触发 CHECK failed,是否可能是 Ray Compiled DAG teardown过程中的竞态问题?

    5. 是否有官方推荐的 Ray、vLLM、MACA、驱动版本组合,可以规避该问题?

    6. 如何进一步定位是哪个GPU Rank、MCCL Collective或Ray Worker在16:27左右首先卡住?是否有推荐的MCCL超时、异步错误检测或Worker Watchdog环境变量?

    7. RAY_CGRAPH_get_timeout=300 是否建议调整?本次只有两个运行请求、KV Cache使用率仅约0.6%,但执行完全停止,因此暂时认为单纯增大超时时间只能延迟失败,无法解决底层卡死。

    log见附件

  • See post chevron_right
    lishuai
    Members
    glm5.1 PD分离问题 已解决 2026年7月3日 17:25

    我在六机上进行glm5.1 PD分离操作,启动decode节点报错,报错内容如文件

  • See post chevron_right
    lishuai
    Members
    C550做glm5.1的PD分离问题 已解决 2026年6月30日 08:56

    如标题,glm5.1-w8a8做PD分离,decode节点最少需要几个?

  • See post chevron_right
    lishuai
    Members
    mccl测试失败 已解决 2026年6月17日 18:14

    metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$ bash mccl.sh 8
    The test is all_reduce_perf, the maca version is /opt/maca-3.7.1
    main_process = 7324
    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7325: Test failure common.cu:1271
    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7326: Test failure common.cu:1271
    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7327: Test failure common.cu:1271
    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7328: Test failure common.cu:1271
    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7329: Test failure common.cu:1271
    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7330: Test failure common.cu:1271
    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7331: Test failure common.cu:1271
    ===============================

    nThread 1 nGpus 1 minBytes 1024 maxBytes 1073741824 step: 2(factor) warmup iters: 5 iters: 10 agg iters: 1 validation: 1 graph: 0

    Using devices

    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7324: Test failure common.cu:1271


    Primary job terminated normally, but 1 process returned
    a non-zero exit code. Per user-direction, the job has been aborted.



    mpirun detected that one or more processes exited with non-zero status, thus causing
    the job to be terminated. The first process to do so was:

    Process name: [[25858,1],4]
    Exit code: 2


    metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$ bash mccl.sh 2
    The test is all_reduce_perf, the maca version is /opt/maca-3.7.1
    main_process = 7378
    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7379: Test failure common.cu:1271
    ===============================

    nThread 1 nGpus 1 minBytes 1024 maxBytes 1073741824 step: 2(factor) warmup iters: 5 iters: 10 agg iters: 1 validation: 1 graph: 0

    Using devices

    metax-host-104: Test CUDA failure common.cu:1349 'initialization error'
    .. metax-host-104 pid 7378: Test failure common.cu:1271


    Primary job terminated normally, but 1 process returned
    a non-zero exit code. Per user-direction, the job has been aborted.



    mpirun detected that one or more processes exited with non-zero status, thus causing
    the job to be terminated. The first process to do so was:

    Process name: [[25940,1],0]
    Exit code: 2


    metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$

    单机mccl测试失败是怎么回事

    mx-smi回显信息正常如下
    metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$ mx-smi
    mx-smi version: 2.3.1

    =================== MetaX System Management Interface Log ===================
    Timestamp : Wed Jun 17 10:13:47 2026

    Attached GPUs : 8
    +---------------------------------------------------------------------------------+
    | MX-SMI 2.3.1 Kernel Mode Driver Version: 3.8.1 |
    | MACA Version: 3.7.1.5 BIOS Version: 1.29.1.0 |
    |------------------+-----------------+---------------------+----------------------|
    | Board Name | GPU Persist-M | Bus-id | GPU-Util sGPU-M |
    | Pwr:Usage/Cap | Temp Perf | Memory-Usage | GPU-State |
    |==================+=================+=====================+======================|
    | 0 MetaX C550 | 0 Off | 0000:2b:00.0 | 0% Disabled |
    | 54W / 450W | 32C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 1 MetaX C550 | 1 Off | 0000:3a:00.0 | 0% Disabled |
    | 56W / 450W | 33C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 2 MetaX C550 | 2 Off | 0000:4d:00.0 | 0% Disabled |
    | 52W / 450W | 33C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 3 MetaX C550 | 3 Off | 0000:5c:00.0 | 0% Disabled |
    | 56W / 450W | 33C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 4 MetaX C550 | 4 Off | 0000:aa:00.0 | 0% Disabled |
    | 53W / 450W | 32C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 5 MetaX C550 | 5 Off | 0000:ba:00.0 | 0% Disabled |
    | 52W / 450W | 33C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 6 MetaX C550 | 6 Off | 0000:ca:00.0 | 0% Disabled |
    | 54W / 450W | 34C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 7 MetaX C550 | 7 Off | 0000:da:00.0 | 0% Disabled |
    | 53W / 450W | 33C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+

    +---------------------------------------------------------------------------------+
    | Process: |
    | GPU PID Process Name GPU Memory |
    | Usage(MiB) |
    |=================================================================================|
    | no process found |
    +---------------------------------------------------------------------------------+

    End of Log
    metax@metax-host-104:/opt/maca/samples/mccl_tests/perf$

  • See post chevron_right
    lishuai
    Members
    vbios版本如何降级? 已解决 2026年6月17日 16:20

    问题如标题

  • See post chevron_right
    lishuai
    Members
    SGPU问题 已解决 2026年6月8日 16:24

    双卡16卡开启sgpu后,双机部署模型的性能会比不开启sgpu双击部署模型的性能低吗

  • See post chevron_right
    lishuai
    Members
    宿主机如何安装vllm-metax? 已解决 2026年5月26日 11:52

    如果要在裸金属宿主机环境下进行模型部署,安装maca-vllm-metax-0.19.0-py310-3.5.3.502-linux-x86_64.tar.xz过后,官方的vllm还需要安装吗?

  • See post chevron_right
    lishuai
    Members
    mccl测试问题 已解决 2026年3月23日 16:32

    缺陷报告:MCCL 跨机 RoCE 内存注册失败 (ibv_reg_mr Invalid argument)
    【环境信息】

    硬件: 双节点共 16 张沐曦 MetaX C500 (8卡/节点),机头是两台H3c uniserver R5330 G7(C500*16)
    网络: 40Gbps RoCE 网卡 (设备名 rocep6s0, rocep95s0)---这是ibstat显示的信息,网卡用的是400G的迈洛斯 cx7
    软件: MACA 3.5.3 / MCCL 2.16.5

    【问题现象】
    在使用 mpirun 执行跨机 all_reduce_perf 测试时,若开启 IB/RoCE 硬件加速,MCCL 在初始化阶段崩溃,报错:
    MCCL WARN Call to ibv_reg_mr failed with error Invalid argument
    【已完成的排查与隔离】
    网络层正常: RoCE 高速网段(192.168.100.x)跨机 ping 测试延迟 < 0.1ms,双端 firewalld 与 SELinux 已关闭。
    系统限制正常: 双端通过 MPI 验证 ulimit -l 均为 unlimited,排除 memlock 限制导致的问题。
    iova2 降级测试: 增加 -x MCCL_IB_PCI_RELAXED_ORDERING=0 后,报错从 ibv_reg_mr_iova2 failed 退化为 ibv_reg_mr failed,依然返回 Invalid argument (retcode 2)。
    TCP 降级对照组(核心证据): 增加 -x MCCL_IB_DISABLE=1 强制走普通 TCP/Socket 通信后,测试完美通过(#wrong 0)。


    mccl测试脚本内容:
    [root@localhost /opt/maca/samples/mccl_tests/perf]# cat cluster.sh

    !/bin/bash

    MACA_PATH="${MACA_PATH:-/opt/maca}"

    HOST_IP=192.168.1.204:8,192.168.1.205:8
    GPU_NUM=16

    TEST_DIR=$MACA_PATH/samples/mccl_tests/perf/mccl_perf

    BENCH_NAMES="all_reduce_perf all_gather_perf reduce_scatter_perf sendrecv_perf alltoall_perf"

    BENCH_NAMES="all_reduce_perf"

    if [[ -z "$1" || -z "$2" || -z "$3" ]]; then
    echo "Use the default ip addr. Run with parameters for custom ip addr, for example: bash cluster.sh ip_1:proc_count,ip_2:proc_count gpu_num test_name"
    else
    HOST_IP=$1
    GPU_NUM=$2

    if [ "$3" = "all" ]; then
    BENCH_NAMES="all_reduce_perf all_gather_perf reduce_scatter_perf sendrecv_perf alltoall_perf"
    else
    if [ -e "$TEST_DIR/$3" ]; then
    BENCH_NAMES=$3
    else
    echo "$TEST_DIR/$3 dose not exist!"
    exit 1
    fi
    fi
    fi

    IP_MASK="$(echo "$HOST_IP" | cut -d. -f1-3).0/24"

    IP_MASK="192.168.100.0/24"
    IB_PORT=rocep6s0,rocep95s0

    PERF_ENV="-x FORCE_ACTIVE_WAIT=2"
    LIB_PATH_ENV="-x MACA_PATH=${MACA_PATH} -x LD_LIBRARY_PATH=${MACA_PATH}/lib:/${MACA_PATH}/ompi/lib:/${MACA_PATH}/ucx/lib"

    ENV_VAR="-x MCCL_IB_HCA=${IB_PORT} -x MCCL_CROSS_NIC=1 ${PERF_ENV} ${LIB_PATH_ENV}"

    ENV_VAR="-x MCCL_IB_HCA=rocep6s0,rocep95s0 -x MCCL_SOCKET_IFNAME=p50p1,p51p1 -x MCCL_CROSS_NIC=1 ${PERF_ENV} ${LIB_PATH_ENV} -x MCCL_IB_DISABLE=0"
    MPI_PROCESS_NUM=${GPU_NUM}
    MPI_RUN_OPT="--allow-run-as-root -mca btl_tcp_if_include ${IP_MASK} -mca oob_tcp_if_include ${IP_MASK} -mca pml ^ucx -mca osc ^ucx -mca btl ^openib"

    for BENCH in ${BENCH_NAMES}; do
    echo -n "The test is ${BENCH}, the maca version is " && realpath ${MACA_PATH}
    ${MACA_PATH}/ompi/bin/mpirun -np ${MPI_PROCESS_NUM} ${MPI_RUN_OPT} -host ${HOST_IP} ${ENV_VAR} ${TEST_DIR}/${BENCH} -b 1K -e 1G -d float -f 2 -g 1 -n 10
    done

    报错信息如附件内容

  • See post chevron_right
    lishuai
    Members
    双机16卡C500部署metax-tech/DeepSeek-R1-0528-W8A8问题 已解决 2026年2月5日 13:40

    请问在进行 ray配置的时候,如果机器有ib卡,是不是必须要要映射计算网口?

  • See post chevron_right
    lishuai
    Members
    如何判断某些模型当前沐曦已经适配的vllm版本是支持的 已解决 2026年2月4日 09:57

    比如以下模型那些支持,那些不支持:
    Qwen3-235B
    Qwen3-VL
    Qwen3-embeding
    Qwen3-rerank
    Z-Image
    GLM4.6/7
    Deepseek-V3.2
    Deepseek-V3
    Deepseek-R1

  • See post chevron_right
    lishuai
    Members
    C500针对OCR识别的性能方面有没有什么优化方向 已解决 2026年2月3日 16:13

    安装信息:python -m pip install paddle-metax-gpu==3.3.0 -i www.paddlepaddle.org.cn/packages/stable/maca/
    模型运行脚本如上传文件:

    使用模型:
    PP-OCRv5_server_det
    PP-LCNet_x1_0_doc_ori
    PP-LCNet_x1_0_textline_ori
    PP-OCRv5_server_rec
    UVDoc

  • See post chevron_right
    lishuai
    Members
    英伟达的libdevice对应的是沐曦的哪个文件? 已解决 2026年2月2日 15:23

    问题如标题

  • 沐曦开发者论坛
powered by misago