• Members 19 posts
    2026年7月1日 15:53

    硬件 C500 * 8

    我使用的命令是:

    #!/bin/bash
    export MACA_SMALL_PAGESIZE_ENABLE=1
    export MACA_VLLM_ENABLE_MCTLASS_FUSED_MOE=0
    export MACA_VLLM_ENABLE_MCTLASS_PYTHON_API=1
    
    vllm serve /llm_models/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8 --served-model-name DeepSeek-V4-Flash --trust-remote-code \
    --kv-cache-dtype bfloat16 --block-size 256 --port 8001 --max-model-len 1048576 --gpu-memory-utilization 0.9 \
    --tokenizer-mode deepseek_v4 --tool-call-parser deepseek_v4 \
    --enable-auto-tool-choice --reasoning-parser deepseek_v4 --default-chat-template-kwargs '{"enable_thinking": true}' \
    --compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE", "custom_ops":["all"]}' \
    --max-num-seq 32 -tp 8 --speculative_config '{"method": "mtp", "num_speculative_tokens": 1}'
    

    然后在加载完模型之后产生巨量的tilelang的编译报错,然后失败退出

  • arrow_forward

    Thread has been moved from 公共.

  • Members 19 posts
    2026年7月1日 16:02

    一、软硬件信息
    1.服务器厂家: 超云
    2.沐曦GPU型号: C500
    3.操作系统内核版本: 6.8.0-124-generic
    4.是否开启CPU虚拟化:是
    5.mx-smi回显:

    mx-smi  version: 2.3.4
    
    =================== MetaX System Management Interface Log ===================
    Timestamp                                         : Wed Jul  1 23:56:36 2026
    
    Attached GPUs                                     : 8
    +---------------------------------------------------------------------------------+
    | MX-SMI 2.3.4                       Kernel Mode Driver Version: 3.9.10           |
    | MACA Version: 3.7.0.38             BIOS Version: 1.35.2.0                       |
    |------------------+-----------------+---------------------+----------------------|
    | Board       Name | GPU   Persist-M | Bus-id              | GPU-Util      sGPU-M |
    | Pwr:Usage/Cap    | Temp       Perf | Memory-Usage        | GPU-State            |
    |==================+=================+=====================+======================|
    | 0     MetaX C500 | 0           Off | 0000:04:00.0        | 0%          Disabled |
    | 38W / 350W       | 40C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 1     MetaX C500 | 1           Off | 0000:05:00.0        | 0%          Disabled |
    | 39W / 350W       | 41C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 2     MetaX C500 | 2           Off | 0000:06:00.0        | 0%          Disabled |
    | 42W / 350W       | 43C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 3     MetaX C500 | 3           Off | 0000:07:00.0        | 0%          Disabled |
    | 38W / 350W       | 41C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 4     MetaX C500 | 4           Off | 0000:0b:00.0        | 0%          Disabled |
    | 40W / 350W       | 42C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 5     MetaX C500 | 5           Off | 0000:0c:00.0        | 0%          Disabled |
    | 41W / 350W       | 43C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 6     MetaX C500 | 6           Off | 0000:0d:00.0        | 0%          Disabled |
    | 42W / 350W       | 41C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 7     MetaX C500 | 7           Off | 0000:0e:00.0        | 0%          Disabled |
    | 39W / 350W       | 40C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    
    +---------------------------------------------------------------------------------+
    | Process:                                                                        |
    |  GPU                    PID         Process Name                 GPU Memory     |
    |                                                                  Usage(MiB)     |
    |=================================================================================|
    |  no process found                                                               |
    +---------------------------------------------------------------------------------+
    
    End of Log
    

    6.docker info回显:

    Client: Docker Engine - Community
     Version:    29.3.0
     Context:    default
     Debug Mode: false
     Plugins:
      buildx: Docker Buildx (Docker Inc.)
        Version:  v0.31.1
        Path:     /usr/libexec/docker/cli-plugins/docker-buildx
      compose: Docker Compose (Docker Inc.)
        Version:  v5.1.0
        Path:     /usr/libexec/docker/cli-plugins/docker-compose
      model: Docker Model Runner (Docker Inc.)
        Version:  v1.1.8
        Path:     /usr/libexec/docker/cli-plugins/docker-model
    
    Server:
     Containers: 12
      Running: 4
      Paused: 0
      Stopped: 8
     Images: 15
     Server Version: 29.3.0
     Storage Driver: overlayfs
      driver-type: io.containerd.snapshotter.v1
     Logging Driver: json-file
     Cgroup Driver: systemd
     Cgroup Version: 2
     Plugins:
      Volume: local
      Network: bridge host ipvlan macvlan null overlay
      Log: awslogs fluentd gcplogs gelf journald json-file local splunk syslog
     CDI spec directories:
      /etc/cdi
      /var/run/cdi
     Swarm: inactive
     Runtimes: io.containerd.runc.v2 runc
     Default Runtime: runc
     Init Binary: docker-init
     containerd version: dea7da592f5d1d2b7755e3a161be07f43fad8f75
     runc version: v1.3.4-0-gd6d73eb8
     init version: de40ad0
     Security Options:
      apparmor
      seccomp
       Profile: builtin
      cgroupns
     Kernel Version: 6.8.0-124-generic
     Operating System: Ubuntu-Server 24.04.4 LTS (Noble Numbat)
     OSType: linux
     Architecture: x86_64
     CPUs: 192
     Total Memory: 503.5GiB
     Name: server2
     ID: 404bb473-c286-41a5-a18e-b94481c6f73e
     Docker Root Dir: /var/lib/docker
     Debug Mode: false
     Experimental: false
     Insecure Registries:
      ::1/128
      127.0.0.0/8
     Live Restore Enabled: false
     Firewall Backend: iptables
    

    7.镜像版本:cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.21.0-maca.ai3.7.1.106-torch2.8-py312-ubuntu22.04-amd64
    8.启动容器命令:

    docker run -d \
        --device=/dev/dri \
        --device=/dev/mxcd \
        --device=/dev/infiniband \
        --group-add video \
        --security-opt seccomp=unconfined \
        --security-opt apparmor=unconfined \
        --shm-size 100gb \
        --ulimit memlock=-1 \
        --privileged=true \
        --network host \
        -v /mnt/modelscope:/llm_models \
        --name vllm-deepseek-v4-flash \
        cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.21.0-maca.ai3.7.1.106-torch2.8-py312-ubuntu22.04-amd64 bash -c "source /etc/profile.d/conda.sh && source /llm_models/run-deepseek-v4-entry.sh"
    

    9.容器内执行命令(/llm_models/run-deepseek-v4-entry.sh的内容):

    #!/bin/bash
    export MACA_SMALL_PAGESIZE_ENABLE=1
    export MACA_VLLM_ENABLE_MCTLASS_FUSED_MOE=0
    export MACA_VLLM_ENABLE_MCTLASS_PYTHON_API=1
    
    vllm serve /llm_models/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8 --served-model-name DeepSeek-V4-Flash --trust-remote-code \
    --kv-cache-dtype bfloat16 --block-size 256 --port 8001 --max-model-len 1048576 --gpu-memory-utilization 0.9 \
    --tokenizer-mode deepseek_v4 --tool-call-parser deepseek_v4 \
    --enable-auto-tool-choice --reasoning-parser deepseek_v4 --default-chat-template-kwargs '{"enable_thinking": true}' \
    --compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE", "custom_ops":["all"]}' \
    --max-num-seq 32 -tp 8 --speculative_config '{"method": "mtp", "num_speculative_tokens": 1}'
    

    二、问题现象
    请描述详细的问题现象日志。若日志过长,请上传附件(txt格式)。

    insert_drive_file
    dsv4-error.log

    Text, 876.9 KB, uploaded by e411 on 2026年7月1日.

  • Members 794 posts
    2026年7月1日 16:07

    尊敬的开发者您好,请在裸金属执行
    dmesg -T | grep -i err

  • Members 19 posts
    2026年7月1日 16:52
    [Wed Jul  1 22:58:53 2026] AMD-Vi: Disabling interrupt remapping
    [Wed Jul  1 22:58:54 2026] ACPI: Using IOAPIC for interrupt routing
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKA configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKB configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKC configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKD configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKE configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKF configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKG configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKH configured for IRQ 0
    [Wed Jul  1 22:58:57 2026] RAS: Correctable Errors collector initialized.
    [Wed Jul  1 22:59:12 2026] hrtimer: interrupt took 30958313 ns
    

    我认为这和卡没有任何关系,是vllm metax内部的一些软件错误

  • Members 794 posts
    2026年7月1日 17:05

    尊敬的开发者您好,请去除source /etc/profile.d/conda.sh尝试

  • Members 19 posts
    2026年7月1日 17:18

    source /etc/profile.d/conda.sh 是为了加载vllm所在的虚拟环境,因为docker run -d拉起来的bash是非交互的,因此不会自动加载conda环境,为了正确加载后续vllm命令,必须加载一下conda.sh。

  • Members 794 posts
    2026年7月1日 17:19

    尊敬的开发者您好,请使用docker -itd启动尝试

  • Members 19 posts
    2026年7月2日 10:35

    没有变化,日志仍然是之前一样。

  • Members 19 posts
    2026年7月2日 10:40

    我进一步排查后,发现其中的tilelang有问题,我对比了贵公司官网发布的tilelang安装包与vllm镜像内的tilelang,发现vllm镜像内tilelang是上游版本,而不是贵公司适配过的tilelang maca版本,其中不含任何maca代码,只有cuda,vllm调用到了cuda,导致maca工具链编译报错,请联系贵公司研发重新适配打包这个vllm的镜像,谢谢!

  • Members 19 posts
    2026年7月2日 10:42

    注意到vllm镜像里的tilelang没有maca适配代码

    image.png

    PNG, 102.4 KB, uploaded by e411 on 2026年7月2日.

    image.png

    PNG, 122.9 KB, uploaded by e411 on 2026年7月2日.

  • Members 794 posts
    2026年7月2日 10:47

    尊敬的开发者您好,经内部确认,此镜像不支持DS V4部署

  • Members 19 posts
    2026年7月2日 10:52

    你们现在有比较新的dsv4的验证镜像吗,我这边真的很想部署一个dsv4来用(

  • Members 794 posts
    2026年7月2日 11:03

    尊敬的开发者您好,请开启个人主题,左上角倒数第三个,收件人写shuai_chen,申请最新POC镜像

  • Members 7 posts
    2026年7月5日 13:52

    deepseek就是你们国产信创厂商的大腿,应该第一时间0day适配所有设备才有出路,否则为啥花大价钱买这种瘸腿产品

  • arrow_forward

    Thread has been moved from 解决中.

  • Members 19 posts
    2026年7月20日 08:07

    vllm-metax:0.22.0 7月16发布的这个新版本是否支持deepseek v4了?

  • Members 794 posts
    2026年7月20日 10:10

    尊敬的开发者您好,0.22.0版本支持DS V4部署

  • Members 19 posts
    2026年7月21日 11:30
    INFO 07-20 17:29:07 [__init__.py:44] Available plugins for group vllm.platform_plugins:
    INFO 07-20 17:29:07 [__init__.py:46] - metax -> vllm_metax:register
    INFO 07-20 17:29:07 [__init__.py:49] All plugins in this group will be loaded. Set `VLLM_PLUGINS` to control which plugins to load.
    INFO 07-20 17:29:07 [__init__.py:238] Platform plugin metax is activated
    INFO 07-20 17:29:07 [envs.py:123] Plugin sets VLLM_USE_FLASHINFER_SAMPLER to False. Reason: flashinfer sampler are not supported on maca
    INFO 07-20 17:29:07 [envs.py:123] Plugin sets VLLM_ENGINE_READY_TIMEOUT_S to 3600. Reason: set timeout to 3600s for model loading
    INFO 07-20 17:29:07 [envs.py:123] Plugin sets VLLM_FLOAT32_MATMUL_PRECISION to high. Reason: set float32 matmul precision to high for better performance on Maca platform
    INFO 07-20 17:29:07 [envs.py:123] Plugin sets VLLM_USE_V2_MODEL_RUNNER to False. Reason: v2 model runner is still under development and not fully tested on Maca platform, disable it by default
    WARNING 07-20 17:29:09 [config.py:70] Support for Transformers v4 is deprecated. The Transformers v4 codepath will become unmaintained in vLLM v0.22.0 and will be removed in vLLM v0.24.0. Please upgrade to Transformers v5: pip install --upgrade transformers
    INFO Print the version information of mcoplib during compilation.
    
    Version info:Mcoplib_Version = '0.4.7'
    Build_Maca_Version = '3.8.0.23'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '3c71b39'
    Vllm Op Version = 0.22.0
    SGlang Op Version  = 0.5.12
    
    INFO Staring Check the current MACA version of the operating environment.
    
    INFO: Release major.minor matching,  successful:3.8.
    
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'awq' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.awq.MacaAWQConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'awq_marlin' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.awq_marlin.MacaAWQMarlinConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'compressed-tensors' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.compressed_tensors.MacaCompressedTensorsConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'gptq' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.auto_gptq.MacaAutoGPTQConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'auto_gptq' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.auto_gptq.MacaAutoGPTQConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'moe_wna16' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.moe_wna16.MacaMoeWNA16Config'>.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepSeekMTPModel is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_mtp:DeepSeekMTP.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepseekV2ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v2:DeepseekV2ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepseekV3ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v2:DeepseekV3ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepseekV32ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v2:DeepseekV3ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture GlmMoeDsaForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v2:GlmMoeDsaForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepseekV4ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v4:DeepseekV4ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepSeekV4MTPModel is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v4_mtp:DeepSeekV4MTP.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture Step3p5MTP is already registered, and will be overwritten by the new model class vllm_metax.models.step3p5_mtp:Step3p5MTP.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture Qwen3OmniMoeForConditionalGeneration is already registered, and will be overwritten by the new model class vllm_metax.models.qwen3_omni_moe_thinker:Qwen3OmniMoeThinkerForConditionalGeneration.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture MiMoV2ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.mimo_v2:MiMoV2ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture MiMoV2FlashForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.mimo_v2:MiMoV2FlashForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture MiMoV2MTPModel is already registered, and will be overwritten by the new model class vllm_metax.models.mimo_v2_mtp:MiMoV2MTP.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture MiMoV2OmniMTPModel is already registered, and will be overwritten by the new model class vllm_metax.models.mimo_v2_mtp:MiMoV2OmniMTP.
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]        █     █     █▄   ▄█
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]  ▄▄ ▄█ █     █     █ ▀▄▀ █  version 0.22.0
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]   █▄█▀ █     █     █     █  model   /llm_models/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]    ▀▀  ▀▀▀▀▀ ▀▀▀▀▀ ▀     ▀
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:278] non-default args: {'model_tag': '/llm_models/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8', 'default_chat_template_kwargs': {'enable_thinking': True}, 'enable_auto_tool_choice': True, 'tool_call_parser': 'deepseek_v4', 'port': 8001, 'model': '/llm_models/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8', 'tokenizer_mode': 'deepseek_v4', 'trust_remote_code': True, 'max_model_len': 512000, 'served_model_name': ['deepseek-v4-flash'], 'reasoning_parser': 'deepseek_v4', 'tensor_parallel_size': 8, 'block_size': 256, 'gpu_memory_utilization': 0.8, 'kv_cache_dtype': 'bfloat16', 'max_num_seqs': 2, 'enable_chunked_prefill': True, 'async_scheduling': True, 'compilation_config': {'mode': None, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': None, 'splitting_ops': None, 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': None, 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.FULL_AND_PIECEWISE: (2, 1)>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': None, 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': None, 'pass_config': {}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': None, 'static_all_moe_layers': []}}
    (APIServer pid=13) INFO 07-20 17:29:28 [config.py:431] Replacing legacy 'type' key with 'rope_type'
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] Error in inspecting model architecture 'DeepseekV4ForCausalLM'
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] Traceback (most recent call last):
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 1385, in _run_in_subprocess
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     returned.check_returncode()
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/subprocess.py", line 502, in check_returncode
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     raise CalledProcessError(self.returncode, self.args, self.stdout,
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] subprocess.CalledProcessError: Command '['/opt/conda/bin/python3.12', '-m', 'vllm.model_executor.models.registry']' returned non-zero exit status 1.
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] The above exception was the direct cause of the following exception:
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] Traceback (most recent call last):
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 948, in _try_inspect_model_cls
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     return model.inspect_model_cls()
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]            ^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/logging_utils/log_time.py", line 21, in _wrapper
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     result = func(*args, **kwargs)
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]              ^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 909, in inspect_model_cls
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     mi = _run_in_subprocess(
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]          ^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 1388, in _run_in_subprocess
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     raise RuntimeError(
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] RuntimeError: Error raised in subprocess:
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] <frozen runpy>:128: RuntimeWarning: 'vllm.model_executor.models.registry' found in sys.modules after import of package 'vllm.model_executor.models', but prior to execution of 'vllm.model_executor.models.registry'; this may result in unpredictable behaviour
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] Traceback (most recent call last):
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen runpy>", line 198, in _run_module_as_main
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen runpy>", line 88, in _run_code
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 1411, in <module>
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     _run()
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 1404, in _run
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     result = fn()
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]              ^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 910, in <lambda>
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     lambda: _ModelInfo.from_model_cls(self.load_model_cls())
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]                                       ^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 923, in load_model_cls
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     mod = importlib.import_module(self.module_name)
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/importlib/__init__.py", line 90, in import_module
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     return _bootstrap._gcd_import(name[level:], package, level)
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 1387, in _gcd_import
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 1360, in _find_and_load
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 1331, in _find_and_load_unlocked
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 935, in _load_unlocked
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap_external>", line 999, in exec_module
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 488, in _call_with_frames_removed
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm_metax/models/deepseek_v4.py", line 23, in <module>
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     from vllm_metax.customized.pluggable_layer.deepseek_v4_attention import (
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm_metax/customized/pluggable_layer/deepseek_v4_attention.py", line 9, in <module>
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     from vllm.model_executor.layers.deepseek_v4_attention import (
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] ModuleNotFoundError: No module named 'vllm.model_executor.layers.deepseek_v4_attention'
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]
    (APIServer pid=13) Traceback (most recent call last):
    (APIServer pid=13)   File "/opt/conda/bin/vllm", line 8, in <module>
    (APIServer pid=13)     sys.exit(main())
    (APIServer pid=13)              ^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/cli/main.py", line 92, in main
    (APIServer pid=13)     args.dispatch_function(args)
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/cli/serve.py", line 148, in cmd
    (APIServer pid=13)     uvloop.run(run_server(args))
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/uvloop/__init__.py", line 96, in run
    (APIServer pid=13)     return __asyncio.run(
    (APIServer pid=13)            ^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/asyncio/runners.py", line 195, in run
    (APIServer pid=13)     return runner.run(main)
    (APIServer pid=13)            ^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/asyncio/runners.py", line 118, in run
    (APIServer pid=13)     return self._loop.run_until_complete(task)
    (APIServer pid=13)            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "uvloop/loop.pyx", line 1518, in uvloop.loop.Loop.run_until_complete
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/uvloop/__init__.py", line 48, in wrapper
    (APIServer pid=13)     return await main
    (APIServer pid=13)            ^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py", line 678, in run_server
    (APIServer pid=13)     await run_server_worker(listen_address, sock, args, **uvicorn_kwargs)
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py", line 692, in run_server_worker
    (APIServer pid=13)     async with build_async_engine_client(
    (APIServer pid=13)                ^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/contextlib.py", line 210, in __aenter__
    (APIServer pid=13)     return await anext(self.gen)
    (APIServer pid=13)            ^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py", line 100, in build_async_engine_client
    (APIServer pid=13)     async with build_async_engine_client_from_engine_args(
    (APIServer pid=13)                ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/contextlib.py", line 210, in __aenter__
    (APIServer pid=13)     return await anext(self.gen)
    (APIServer pid=13)            ^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py", line 124, in build_async_engine_client_from_engine_args
    (APIServer pid=13)     vllm_config = engine_args.create_engine_config(usage_context=usage_context)
    (APIServer pid=13)                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/engine/arg_utils.py", line 1730, in create_engine_config
    (APIServer pid=13)     model_config = self.create_model_config()
    (APIServer pid=13)                    ^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/engine/arg_utils.py", line 1560, in create_model_config
    (APIServer pid=13)     return ModelConfig(
    (APIServer pid=13)            ^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/pydantic/_internal/_dataclasses.py", line 121, in __init__
    (APIServer pid=13)     s.__pydantic_validator__.validate_python(ArgsKwargs(args, kwargs), self_instance=s)
    (APIServer pid=13) pydantic_core._pydantic_core.ValidationError: 1 validation error for ModelConfig
    (APIServer pid=13)   Value error, Model architectures ['DeepseekV4ForCausalLM'] failed to be inspected. Please check the logs for more details. [type=value_error, input_value=ArgsKwargs((), {'model': ...nderer_num_workers': 1}), input_type=ArgsKwargs]
    (APIServer pid=13)     For further information visit https://errors.pydantic.dev/2.13/v/value_error
    

    您好,部署按照之前命令无法成功,请你再确认一下

  • Members 19 posts
    2026年7月21日 16:44

    您好,采用的运行参数和本主题中前述参数完全一样,区别只有换用了7.16的发布版本vllm 0.22.0. 此命令在前面使用的vllm 0.21.0 poc中是可以正常运行的,所以希望您能确认下这个公开版是不是真支持dsv4了

  • Members 794 posts
    2026年7月21日 16:45

    尊敬的开发者您好,经内部确认0.22.0不支持DS V4,等待镜像中心后续发布。POC镜像没有更新。