• Members 20 posts
    2026年9月1日 11:52

    八张c6500 的kmd版本比较低是3.0.11 mxsmi是2.3.4 maca是3.8.2.6 bios是1.27.5.0 启动参数是nohup vllm serve /data/Qwen/Qwen3.8-27b --served-model-name qwen3.8-27b --tensor-parallel-size 4 --gpu-memory-utilization 0.82 --max-model-len 4000 --dtype half --port 8016 --host 0.0.0.0 --trust-remote-code --reasoning-parser qwen3 --max-num-seqs 5 --disable-log-requests > /data/vllm-27b.log 2>&1 &

    a9f1360366dd15c996d214d989a49380.jpg

    JPG, 485.0 KB, uploaded by czy7570 on 2026年9月1日.

  • Members 20 posts
    2026年9月1日 11:52

    麻烦帮忙看看 是用了四张卡

  • arrow_forward

    Thread has been moved from 公共.

  • Members 20 posts
    2026年9月1日 14:58

    1.服务器厂家:海光
    2.沐曦GPU型号:C500*8
    3.操作系统内核版本:Kylin-Server-V10-SP3-2403-Release-20240426-x86_64
    4.是否开启CPU虚拟化:未开启
    5.mx-smi回显:

    mx-smi version: 2.2.8

    =================== MetaX System Management Interface Log ===================
    Timestamp : Tue Sep 1 14:55:07 2026

    Attached GPUs : 8
    +---------------------------------------------------------------------------------+
    | MX-SMI 2.2.8 Kernel Mode Driver Version: 3.0.11 |
    | MACA Version: 3.2.1.10 BIOS Version: 1.27.5.0 |
    |------------------------------------+---------------------+----------------------+
    | GPU NAME Persistence-M | Bus-id | GPU-Util sGPU-M |
    | Temp Pwr:Usage/Cap Perf | Memory-Usage | GPU-State |
    |====================================+=====================+======================|
    | 0 MetaX C500 Off | 0000:05:00.0 | 0% Native |
    | 44C 59W / 350W P0 | 858/65536 MiB | Available |
    +------------------------------------+---------------------+----------------------+
    | 1 MetaX C500 Off | 0000:0b:00.0 | 0% Native |
    | 46C 62W / 350W P0 | 858/65536 MiB | Available |
    +------------------------------------+---------------------+----------------------+
    | 2 MetaX C500 Off | 0000:0e:00.0 | 0% Native |
    | 44C 58W / 350W P0 | 858/65536 MiB | Available |
    +------------------------------------+---------------------+----------------------+
    | 3 MetaX C500 Off | 0000:0f:00.0 | 0% Native |
    | 47C 63W / 350W P0 | 858/65536 MiB | Available |
    +------------------------------------+---------------------+----------------------+
    | 4 MetaX C500 Off | 0000:55:00.0 | 0% Native |
    | 41C 44W / 350W P0 | 858/65536 MiB | Available |
    +------------------------------------+---------------------+----------------------+
    | 5 MetaX C500 Off | 0000:56:00.0 | 0% Native |
    | 42C 42W / 350W P0 | 858/65536 MiB | Available |
    +------------------------------------+---------------------+----------------------+
    | 6 MetaX C500 Off | 0000:5b:00.0 | 0% Native |
    | 42C 55W / 350W P9 | 56393/65536 MiB | Available |
    +------------------------------------+---------------------+----------------------+
    | 7 MetaX C500 Off | 0000:5e:00.0 | 0% Native |
    | 44C 57W / 350W P9 | 56337/65536 MiB | Available |
    +------------------------------------+---------------------+----------------------+

    +---------------------------------------------------------------------------------+
    | Process: |
    | GPU PID Process Name GPU Memory |
    | Usage(MiB) |
    |=================================================================================|
    | 6 4176164 VLLM::EngineCor 55534 |
    | 7 4173582 VLLM::EngineCor 55478 |
    +---------------------------------------------------------------------------------+

    End of Log

    6.docker info回显:
    Client:
    Version: 28.5.1
    Context: default
    Debug Mode: false

    Server:
    Containers: 41
    Running: 14
    Paused: 0
    Stopped: 27
    Images: 19
    Server Version: 28.5.1
    Storage Driver: overlay2
    Backing Filesystem: xfs
    Supports d_type: true
    Using metacopy: false
    Native Overlay Diff: true
    userxattr: false
    Logging Driver: json-file
    Cgroup Driver: cgroupfs
    Cgroup Version: 1
    Plugins:
    Volume: local
    Network: bridge host ipvlan macvlan null overlay
    Log: awslogs fluentd gcplogs gelf journald json-file local splunk syslog
    CDI spec directories:
    /etc/cdi
    /var/run/cdi
    Swarm: inactive
    Runtimes: io.containerd.runc.v2 runc
    Default Runtime: runc
    Init Binary: docker-init
    containerd version: b98a3aace656320842a23f4a392a33f46af97866
    runc version: v1.3.0-0-g4ca628d
    init version: de40ad0
    Security Options:
    seccomp
    Profile: builtin
    Kernel Version: 4.19.90-89.11.v2401.ky10.x86_64
    Operating System: Kylin Linux Advanced Server V10 (Halberd)
    OSType: linux
    Architecture: x86_64
    CPUs: 256
    Total Memory: 1006GiB
    Name: localhost.localdomain
    ID: 4e75250f-0952-4b17-a2fe-a49d1b7cedc8
    Docker Root Dir: /var/lib/docker
    Debug Mode: false
    Experimental: false
    Insecure Registries:
    ::1/128
    127.0.0.0/8
    Registry Mirrors:
    docker.m.daocloud.io/
    Live Restore Enabled: false
    Product License: Community Engine

    7.镜像版本:
    cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64
    8.启动容器命令:
    docker run -d \
    --device=/dev/mxcd \
    --device=/dev/dri/card2 --device=/dev/dri/renderD129 \
    --device=/dev/dri/card3 --device=/dev/dri/renderD130 \
    --device=/dev/dri/card4 --device=/dev/dri/renderD131 \
    --device=/dev/dri/card5 --device=/dev/dri/renderD132 \
    --device=/dev/dri/card1 --device=/dev/dri/renderD128 \
    --device=/dev/dri/card6 --device=/dev/dri/renderD133 \
    --group-add video \
    --network=host --name vllm02-27b \
    --restart unless-stopped \
    --security-opt seccomp=unconfined --security-opt apparmor=unconfined \
    --ulimit memlock=-1 \
    -v /data:/data \
    cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64 \

    9.容器内执行命令:
    vllm serve /data/Qwen3.5-27B-W8A8 \
    --served-model-name qwen3.5-27b \
    --tensor-parallel-size 4 \
    --gpu-memory-utilization 0.85 \
    --max-model-len 8192 \
    --dtype half \
    --port 8016 \
    --trust-remote-code \
    --reasoning-parser qwen3

  • Members 20 posts
    2026年9月1日 15:01

    附上完整日志

    insert_drive_file
    vllm-27b.log

    Text, 81.6 KB, uploaded by czy7570 on 2026年9月1日.

  • Members 949 posts
    2026年9月1日 15:03

    尊敬的开发者您好,此镜像不支持qwen3.8-27B,等待后续镜像中心镜像更新

  • Members 20 posts
    2026年9月2日 10:56

    你好 目前最新支持到哪个模型呢

  • Members 949 posts
  • Members 20 posts
    2026年9月2日 11:32

    我看到了 准备用QWEN3.6 27b 请问这个适合tp2吗

  • Members 20 posts
    2026年9月3日 14:52

    你好 我用QWEN3.6按照手册部署tp2 报错mccl通讯错误
    ===== 27B 引擎崩溃日志(2026-09-02 17:15,运行约 7 小时后)=====

    [APIServer pid=287] INFO: Application shutdown complete.
    [APIServer pid=287] INFO: Finished server process [287]

    --- 第 1 层:TCPStore 连接被对端重置 ---
    Traceback (most recent call last):
    File "/opt/conda/lib/python3.10/socket.py", line 1140, in recv_into
    return self._sock.recv_into(b)
    ConnectionResetError: [Errno 104] Connection reset by peer

    The above exception was the direct cause of the following exception:

    --- 第 2 层:c10d 分布式心跳监控线程异常 ---
    Traceback (most recent call last):
    File "/opt/conda/lib/python3.10/site-packages/torch/multiprocessing/.../distributed/c10d/Utils.hpp:653 (most recent call first):
    frame #1: c10::Error::Error(...)
    (0x7f7fd9a9990a2 in /opt/conda/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
    frame #2: c10d::TCPStore::check(std::vector<std::string> const&) + 0x242
    (0x7f7fd9a9990a2 in /opt/conda/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
    frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x3e8
    (0x7f7fcd49cc08 in /opt/conda/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
    frame #4: <unknown function> + 0xd3bb5
    (0x7f7fe0a6f1b65 in /opt/conda/bin/../lib/libstdc++.so.6)
    frame #5: <unknown function> + 0x9448
    (0x7f7fe0bc6644 in /lib/x86_64-linux-gnu/libc.so.6)
    frame #6: clone + 0x44
    (0x7f7fe0bc6b64 in /lib/x86_64-linux-gnu/libc.so.6)

    [rank3]:[W002 17:15:02.795520658 TCPStore.cpp:125] [c10d] recvValue failed on
    SocketImpl(fd=14, addr=[::ffff:802e:4009:7c7f:0]) Connection closed by peer.
    [rank3]:[W002 17:15:02.795520658 ProcessGroupNCCL.cpp:1808] [PG ID 0 PG GUID 0 Rank 3]
    Failed to check the "should dump" flag on TCPStore, (maybe TCP Store server has shut down
    too early), with error: Failed to recv, got 0 bytes. Connection was likely closed.
    Did the remote server shutdown or crash

    --- 第 3 层:MCCL(沐曦跨卡通信库)反复连接失败 ---
    localhost:829:050 [0] /workspace/out/Release/build/linux/x86_64/mccl/macafify/src/misc/socket.cc:568
    MCCL WARN socketStartConnect: Connect to 6.0.63.1:34343 failed: Software caused connection abort
    (以上 MCCL WARN 重复出现多次)

    [rank3]:[W002 17:15:02.795520658 TCPStore.cpp:106] [c10d] sendBytes failed on
    SocketImpl(fd=14, addr=[::]:29057, remote=[ffff:ffff::802e:4009:7c7f:0])
    [1137]: Broken pipe
    [rank3]:[W002 17:15:03.807503454 ProcessGroupNCCL.cpp:1808] [PG ID 0 PG GUID 0 Rank 3]
    Failed to check the "should dump" flag on TCPStore, ... with error: Broken pipe

    --- 收尾:共享内存泄漏提示 ---
    /opt/conda/lib/python3.10/multiprocessing/resource_tracker.py:224: UserWarning:
    resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown

    ===== 崩溃时环境状态(截屏底部)=====
    服务器 uptime: 47 hours(即崩溃前服务器已稳定运行 47 小时)
    磁盘: / 66%, /boot 27%, /data2 1%, /data3 2%

  • Members 9 posts
    2026年9月4日 09:23

    大概多久就可以发布这个镜像呢 我们需要部署qwen3.8-27B 麒麟v10系统支持吗

  • Members 949 posts
    2026年9月4日 09:35

    尊敬的开发者您好,等待近期镜像中心更新。若需申请POC镜像,请通过GPU购买商务渠道获取。

  • Members 20 posts
    2026年9月7日 10:25
    1. 服务器厂家

    海光

    1. 沐曦 GPU 型号

    MetaX C500 × 8(单卡 64GB / 65536 MiB)

    1. 操作系统内核版本
    Kylin-Server-V10-SP3-2403-Release-20240426-x86_64
    Kernel: 4.19.90-89.11.v2401.ky10.x86_64
    (GRUB 默认启动项已锁定该内核;同机另一内核 4.19.90-89.26 因驱动模块不兼容无法启动)
    
    1. 是否开启 CPU 虚拟化

    未开启

    1. mx-smi 回显
      MX-SMI 2.3.4 Kernel Mode Driver Version: 3.9.25
      MACA Version: 3.8.2.6 BIOS Version: 1.27.5.0
      GPU0-7: MetaX C500, 65536 MiB

    6. docker info 回显

    Client:
     Version: 28.5.1
     Context: default
    Server:
     Containers: 41  Running: 14  Stopped: 27
     Images: 19
     Server Version: 28.5.1
     Storage Driver: overlay2
     Cgroup Driver: cgroupfs
     Cgroup Version: 1
     Runtimes: io.containerd.runc.v2 runc
     Default Runtime: runc
     Kernel Version: 4.19.90-89.11.v2401.ky10.x86_64
     Operating System: Kylin Linux Advanced Server V10 (Halberd)
     OSType: linux
     Architecture: x86_64
     CPUs: 256
     Total Memory: 1006 GiB
    
    1. 镜像版本

    cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64
    8. 启动容器命令

    docker run -d \
        --device=/dev/mxcd \
        --device=/dev/dri/card2 --device=/dev/dri/renderD129 \
        --device=/dev/dri/card3 --device=/dev/dri/renderD130 \
        --device=/dev/dri/card4 --device=/dev/dri/renderD131 \
        --device=/dev/dri/card5 --device=/dev/dri/renderD132 \
        --device=/dev/dri/card1 --device=/dev/dri/renderD128 \
        --device=/dev/dri/card6 --device=/dev/dri/renderD133 \
        --group-add video \
        --network=host --name vllm02-27b \
        --restart unless-stopped \
        --security-opt seccomp=unconfined --security-opt apparmor=unconfined \
        --ulimit memlock=-1 \
        -v /data:/data -v /data2:/data2 \
        cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64 \
        sleep infinity
    

    (注:GPU 设备与卡序号映射非顺序:GPU0=card2/renderD129 ... GPU4=card1/renderD128 夹在中间 ... GPU5=card6/renderD133)

    1. 容器内执行命令

    ```bash
    vllm serve /data2/Qwen3.6-27B \
    --served-model-name qwen3.6-27b \
    --tensor-parallel-size 2 \
    --gpu-memory-utilization 0.85 \
    --max-model-len 16384 \
    --dtype auto \
    --port 8016 \
    --trust-remote-code \
    -enforce-eager


    10. 问题现象(重点,请厂商优先阅读)

    现象二(2026-09-02 起持续至今,Qwen3.6-27B 官方 BF16,tp2,KMD 已升级 3.9.25)

    • 驱动升级后现象一未再出现,但新问题:持续服务约 7 小时后引擎崩溃,多次复现
    • 日志关键内容(两次崩溃完全一致):
      MCCL WARN socketStartConnect: Connect to 6.0.63.1:xxxxx failed: Software caused connection abort
      [rankN] TCPStore recvValue failed / sendBytes failed / Broken pipe
      [rankN] ProcessGroupNCCL.cpp:1808 Failed to check the "should dump" flag on TCPStore,
      (maybe TCP Store server has shut down too early)
      [rankN] Aborted at ProcessGroupNCCL::HeartbeatMonitor::runLoop()
    • 崩溃时间点无固定规律(空闲或请求中都可能),服务器 uptime 正常(非服务器重启导致)
    • 跨卡通信地址 6.0.63.1 为本机网卡 IP 段(用户 ssh 来源 6.0.63.x),疑似 MCCL 单机多卡通信走 TCP socket 绑定外部网卡,链路异常即崩
  • Members 20 posts
    2026年9月7日 10:43

    是不是双卡之间通信有问题

  • Members 949 posts
    2026年9月7日 10:48

    尊敬的开发者您好,请提供服务器型号

  • Members 20 posts
    2026年9月7日 11:10

    新华三的服务器 貌似是联想开天的贴牌?KA742H G3
    ★处理器:2海光7490
    ★内存:1TB DDR4
    ★硬盘1:2
    3.84TB SSD
    ★硬盘2:2960 GB SSD
    ★网卡: 双口25GbE SFP28 光纤网卡(含模块)
    ★GPU: 8
    C500 共计512G显存
    ★其他:机架导轨
    ★服务:三年7*24原厂工程师上门售后服务