MetaX-Tech Developer Forum 论坛首页
  • 沐曦开发者
search
Sign in

qdwzy0831

  • Members
  • Joined 2026年8月31日
  • message 帖子
  • forum 主题
  • favorite 关注者
  • favorite_border Follows
  • person_outline 详细信息

qdwzy0831 has started 3 threads.

  • See post chevron_right
    qdwzy0831
    Members
    metax-vllm镜像ubuntu24.04版本申请 已解决 2026年9月4日 17:21

    请问能否提供下基于ubuntu24.04版本的metax-vllm的最新镜像啊,宿主机是24.04的版本动不了,vllm运行的时候老是缺库,没法正常使用

  • See post chevron_right
    qdwzy0831
    Members
    c550单机8卡部署hy3w8a8时出现ReasoningParser问题 已解决 2026年9月4日 10:33

    一、软硬件信息
    1.服务器厂家:
    浪潮CS5820H3
    2.沐曦GPU型号:mxC550
    3.操作系统内核版本:Linux master 6.8.0-100-generic
    4.是否开启CPU虚拟化:是
    5.mx-smi回显:
    mx-smi version: 2.3.4

    =================== MetaX System Management Interface Log ===================
    Timestamp : Fri Sep 4 10:08:34 2026

    Attached GPUs : 8
    +---------------------------------------------------------------------------------+
    | MX-SMI 2.3.4 Kernel Mode Driver Version: 3.9.25 |
    | MACA Version: 3.8.2.6 BIOS Version: 1.35.4.0 |
    |------------------+-----------------+---------------------+----------------------|
    | Board Name | GPU Persist-M | Bus-id | GPU-Util sGPU-M |
    | Pwr:Usage/Cap | Temp Perf | Memory-Usage | GPU-State |
    |==================+=================+=====================+======================|
    | 0 MetaX C550 | 0 Off | 0000:04:00.0 | 0% Disabled |
    | 94W / 450W | 31C P0 | 860/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 1 MetaX C550 | 1 Off | 0000:05:00.0 | 0% Disabled |
    | 92W / 450W | 31C P0 | 860/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 2 MetaX C550 | 2 Off | 0000:24:00.0 | 0% Disabled |
    | 90W / 450W | 28C P0 | 860/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 3 MetaX C550 | 3 Off | 0000:25:00.0 | 0% Disabled |
    | 92W / 450W | 32C P0 | 860/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 4 MetaX C550 | 4 Off | 0000:83:00.0 | 0% Disabled |
    | 91W / 450W | 29C P0 | 860/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 5 MetaX C550 | 5 Off | 0000:86:00.0 | 0% Disabled |
    | 93W / 450W | 31C P0 | 860/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 6 MetaX C550 | 6 Off | 0000:a3:00.0 | 0% Disabled |
    | 92W / 450W | 31C P0 | 860/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 7 MetaX C550 | 7 Off | 0000:a6:00.0 | 0% Disabled |
    | 93W / 450W | 32C P0 | 860/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+

    +---------------------------------------------------------------------------------+
    | Process: |
    | GPU PID Process Name GPU Memory |
    | Usage(MiB) |
    |=================================================================================|
    | no process found |
    +---------------------------------------------------------------------------------+
    6.docker info回显:
    Client: Docker Engine - Community
    Version: 28.3.3
    Context: default
    Debug Mode: false
    Plugins:
    buildx: Docker Buildx (Docker Inc.)
    Version: v0.33.0
    Path: /usr/libexec/docker/cli-plugins/docker-buildx
    compose: Docker Compose (Docker Inc.)
    Version: v5.1.2
    Path: /usr/libexec/docker/cli-plugins/docker-compose

    Server:
    Containers: 27
    Running: 27
    Paused: 0
    Stopped: 0
    Images: 15
    Server Version: 28.3.3
    Storage Driver: overlay2
    Backing Filesystem: xfs
    Supports d_type: true
    Using metacopy: false
    Native Overlay Diff: true
    userxattr: false
    Logging Driver: json-file
    Cgroup Driver: systemd
    Cgroup Version: 2
    Plugins:
    Volume: local
    Network: bridge host ipvlan macvlan null overlay
    Log: awslogs fluentd gcplogs gelf journald json-file local splunk syslog
    CDI spec directories:
    /etc/cdi
    /var/run/cdi
    Swarm: inactive
    Runtimes: io.containerd.runc.v2 runc
    Default Runtime: runc
    Init Binary: docker-init
    containerd version: 301b2dac98f15c27117da5c8af12118a041a31d9
    runc version: v1.3.4-0-gd6d73eb8
    init version: de40ad0
    Security Options:
    apparmor
    seccomp
    Profile: builtin
    cgroupns
    Kernel Version: 6.8.0-53-generic
    Operating System: Ubuntu 24.04.3 LTS
    OSType: linux
    Architecture: x86_64
    CPUs: 255
    Total Memory: 1.472TiB
    Name: gpu02
    ID: 158650d3-e28b-4fd2-8cd1-1af270e4f8bd
    Docker Root Dir: /var/lib/data/docker
    Debug Mode: false
    Experimental: false
    Insecure Registries:
    0.0.0.0/0
    ::1/128
    127.0.0.0/8
    Live Restore Enabled: false
    7.镜像版本:vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64
    8.启动容器命令:
    9.容器内执行命令:

    !/bin/bash

    export LD_LIBRARY_PATH=/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64:/usr/local/cuda/compat/lib:/usr/local/nvidia/lib:/usr/local/nvidia/lib64:/usr/local/cuda/compat/lib:/usr/local/nvidia/lib:/usr/local/nvidia/lib64:/usr/local/lib
    export MACA_SMALL_PAGESIZE_ENABLE=1
    export MACA_GRAPH_LAUNCH_MODE=1
    export MACA_D2D_COPY_MODE=1
    export MACA_DIRECT_DISPATCH=1
    export MACA_MEMCPY_MODE=1
    export VLLM_DISABLE_SHARED_EXPERTS_STREAM=1

    vllm serve /model \
    --host 0.0.0.0 \
    --port 8080 \
    --trust-remote-code \
    --tensor-parallel-size 8 \
    --served-model-name HY3-W8A8 \
    --dtype auto \
    --gpu-memory-utilization 0.90 \
    --max-model-len 8192 \
    --max-num-batched-tokens 32768 \
    --max-num-seqs 16 \
    --speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
    --enable-auto-tool-choice \
    --tool-call-parser hy_v3 \
    --reasoning-parser hy_v3 \
    --no-async-scheduling \
    --enable-prefix-caching \
    --tokenizer-mode auto \
    --attention-backend FLASH_ATTN \
    --enforce-eager

    二、问题现象
    用metax-0.24.0镜像部署metax-hy3W8A8模型 ,报错 (APIServer pid=837) RuntimeError: HYV3ReasoningParser reasoning parser could not locate think start/end tokens in the tokenizer!

  • See post chevron_right
    qdwzy0831
    Members
    c550部署metax-deepseekv4flash0731w8a8出现问题 已解决 2026年8月31日 14:21

    使用的镜像为vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py312-ubuntu22.04-amd64
    无法使用mtp或dspark,且模型部署后,调用工具时出现upstream error: ("error":{"message":"cannot import name 'normalize_tool_choice' From 'xgrammar" (/opt/conda/lib/python3.12/site- packages/xgrammar/init.py)","bype":"InternalServerError","param"null,"code":500]}
    检查镜像后发现normalize_tool_choice未在镜像的init.py文件内出现。
    请问如何处理?

  • 沐曦开发者论坛
powered by misago