请问能否提供下基于ubuntu24.04版本的metax-vllm的最新镜像啊,宿主机是24.04的版本动不了,vllm运行的时候老是缺库,没法正常使用
请问能否提供下基于ubuntu24.04版本的metax-vllm的最新镜像啊,宿主机是24.04的版本动不了,vllm运行的时候老是缺库,没法正常使用
请问咱们现在有哪个镜像支持呢,我试过那个day0的vllm 0.21.0镜像也是一样的问题
一、软硬件信息
1.服务器厂家:
浪潮CS5820H3
2.沐曦GPU型号:mxC550
3.操作系统内核版本:Linux master 6.8.0-100-generic
4.是否开启CPU虚拟化:是
5.mx-smi回显:
mx-smi version: 2.3.4
=================== MetaX System Management Interface Log ===================
Timestamp : Fri Sep 4 10:08:34 2026
Attached GPUs : 8
+---------------------------------------------------------------------------------+
| MX-SMI 2.3.4 Kernel Mode Driver Version: 3.9.25 |
| MACA Version: 3.8.2.6 BIOS Version: 1.35.4.0 |
|------------------+-----------------+---------------------+----------------------|
| Board Name | GPU Persist-M | Bus-id | GPU-Util sGPU-M |
| Pwr:Usage/Cap | Temp Perf | Memory-Usage | GPU-State |
|==================+=================+=====================+======================|
| 0 MetaX C550 | 0 Off | 0000:04:00.0 | 0% Disabled |
| 94W / 450W | 31C P0 | 860/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 1 MetaX C550 | 1 Off | 0000:05:00.0 | 0% Disabled |
| 92W / 450W | 31C P0 | 860/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 2 MetaX C550 | 2 Off | 0000:24:00.0 | 0% Disabled |
| 90W / 450W | 28C P0 | 860/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 3 MetaX C550 | 3 Off | 0000:25:00.0 | 0% Disabled |
| 92W / 450W | 32C P0 | 860/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 4 MetaX C550 | 4 Off | 0000:83:00.0 | 0% Disabled |
| 91W / 450W | 29C P0 | 860/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 5 MetaX C550 | 5 Off | 0000:86:00.0 | 0% Disabled |
| 93W / 450W | 31C P0 | 860/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 6 MetaX C550 | 6 Off | 0000:a3:00.0 | 0% Disabled |
| 92W / 450W | 31C P0 | 860/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
| 7 MetaX C550 | 7 Off | 0000:a6:00.0 | 0% Disabled |
| 93W / 450W | 32C P0 | 860/65536 MiB | Available |
+------------------+-----------------+---------------------+----------------------+
+---------------------------------------------------------------------------------+
| Process: |
| GPU PID Process Name GPU Memory |
| Usage(MiB) |
|=================================================================================|
| no process found |
+---------------------------------------------------------------------------------+
6.docker info回显:
Client: Docker Engine - Community
Version: 28.3.3
Context: default
Debug Mode: false
Plugins:
buildx: Docker Buildx (Docker Inc.)
Version: v0.33.0
Path: /usr/libexec/docker/cli-plugins/docker-buildx
compose: Docker Compose (Docker Inc.)
Version: v5.1.2
Path: /usr/libexec/docker/cli-plugins/docker-compose
Server:
Containers: 27
Running: 27
Paused: 0
Stopped: 0
Images: 15
Server Version: 28.3.3
Storage Driver: overlay2
Backing Filesystem: xfs
Supports d_type: true
Using metacopy: false
Native Overlay Diff: true
userxattr: false
Logging Driver: json-file
Cgroup Driver: systemd
Cgroup Version: 2
Plugins:
Volume: local
Network: bridge host ipvlan macvlan null overlay
Log: awslogs fluentd gcplogs gelf journald json-file local splunk syslog
CDI spec directories:
/etc/cdi
/var/run/cdi
Swarm: inactive
Runtimes: io.containerd.runc.v2 runc
Default Runtime: runc
Init Binary: docker-init
containerd version: 301b2dac98f15c27117da5c8af12118a041a31d9
runc version: v1.3.4-0-gd6d73eb8
init version: de40ad0
Security Options:
apparmor
seccomp
Profile: builtin
cgroupns
Kernel Version: 6.8.0-53-generic
Operating System: Ubuntu 24.04.3 LTS
OSType: linux
Architecture: x86_64
CPUs: 255
Total Memory: 1.472TiB
Name: gpu02
ID: 158650d3-e28b-4fd2-8cd1-1af270e4f8bd
Docker Root Dir: /var/lib/data/docker
Debug Mode: false
Experimental: false
Insecure Registries:
0.0.0.0/0
::1/128
127.0.0.0/8
Live Restore Enabled: false
7.镜像版本:vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py310-ubuntu22.04-amd64
8.启动容器命令:
9.容器内执行命令:
export LD_LIBRARY_PATH=/usr/local/cuda/lib64:/usr/local/cuda/extras/CUPTI/lib64:/usr/local/cuda/compat/lib:/usr/local/nvidia/lib:/usr/local/nvidia/lib64:/usr/local/cuda/compat/lib:/usr/local/nvidia/lib:/usr/local/nvidia/lib64:/usr/local/lib
export MACA_SMALL_PAGESIZE_ENABLE=1
export MACA_GRAPH_LAUNCH_MODE=1
export MACA_D2D_COPY_MODE=1
export MACA_DIRECT_DISPATCH=1
export MACA_MEMCPY_MODE=1
export VLLM_DISABLE_SHARED_EXPERTS_STREAM=1
vllm serve /model \
--host 0.0.0.0 \
--port 8080 \
--trust-remote-code \
--tensor-parallel-size 8 \
--served-model-name HY3-W8A8 \
--dtype auto \
--gpu-memory-utilization 0.90 \
--max-model-len 8192 \
--max-num-batched-tokens 32768 \
--max-num-seqs 16 \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
--enable-auto-tool-choice \
--tool-call-parser hy_v3 \
--reasoning-parser hy_v3 \
--no-async-scheduling \
--enable-prefix-caching \
--tokenizer-mode auto \
--attention-backend FLASH_ATTN \
--enforce-eager
二、问题现象
用metax-0.24.0镜像部署metax-hy3W8A8模型 ,报错 (APIServer pid=837) RuntimeError: HYV3ReasoningParser reasoning parser could not locate think start/end tokens in the tokenizer!
请问调用工具的报错如何解决?
使用的镜像为vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py312-ubuntu22.04-amd64
无法使用mtp或dspark,且模型部署后,调用工具时出现upstream error: ("error":{"message":"cannot import name 'normalize_tool_choice' From 'xgrammar" (/opt/conda/lib/python3.12/site- packages/xgrammar/init.py)","bype":"InternalServerError","param"null,"code":500]}
检查镜像后发现normalize_tool_choice未在镜像的init.py文件内出现。
请问如何处理?