MetaX-Tech Developer Forum 论坛首页
  • 沐曦开发者
search
Sign in

e411

  • Members
  • Joined 2026年5月8日
  • message 帖子
  • forum 主题
  • favorite 关注者
  • favorite_border Follows
  • person_outline 详细信息

e411 has posted 26 messages.

  • See post chevron_right
    e411
    Members
    deepseek v4 flash sglang启动参数? 已解决 2026年9月24日 14:59

    memfrac大概可以设置0.87,0.9会爆显存;max model len对应的大概能有450K;chunked_prefill_size 8192就行,SGLANG_DSPARK_KERNEL_COMMIT_KV_PROJ需要配置防止走fp8,别的不用全部回退,SGLANG_DSPARK_KERNEL_EXPAND_PREFILL这个必须配置,编译会报错。

    docker run -d \
      --device=/dev/dri --device=/dev/mxcd --device=/dev/infiniband \
      --group-add video --security-opt seccomp=unconfined \
      --security-opt apparmor=unconfined --shm-size 100gb --ulimit memlock=-1 \
      --privileged=true --network host \
      -v /mnt/nvme0n1p1/modelscope:/llm_models \
      -e TORCHINDUCTOR_COMPILE_THREADS=1 \
      -e PYTORCH_ALLOC_CONF=expandable_segments:True \
      -e SGLANG_DEFAULT_THINKING=1 \
      -e SGLANG_DSPARK_KERNEL_EXPAND_PREFILL=torch \
      -e SGLANG_DSPARK_KERNEL_COMMIT_KV_PROJ=torch \
      -e SGLANG_MOE_CONFIG_DIR=/llm_models/moe_configs \
      --name sglang-0514-dspark \
      cr.metax-tech.com/public-ai-release/maca/sglang:0.5.14-maca.ai3.8.2.122-torch2.10-py312-kylinv11-amd64 \
      bash -c 'python3 /llm_models/sglang-0514-thinkfix.py; echo ===PATCH-PREAMBLE-DONE===; grep -q SGLANG-0514-DEFTHINK /opt/conda/lib/python3.12/site-packages/sglang/srt/entrypoints/openai/serving_chat.py || { echo ===THINKFIX-NOT-APPLIED===; exit 9; }; python -m sglang.launch_server --host 0.0.0.0 --port 30000 --model-path /llm_models/metax-tech--DeepSeek-V4-Flash-0731-W8A8/snapshots/master --served-model-name deepseek-v4-flash --trust-remote-code --tp 8 --quantization w8a8_int8 --reasoning-parser deepseek-v4 --tool-call-parser deepseekv4 --mem-fraction-static 0.87 --speculative-algorithm DSPARK --speculative-dspark-block-size 5 --kv-cache-dtype bf16 --mamba-ssm-dtype bfloat16 --enable-cache-report --chunked-prefill-size 8192 --enable-hierarchical-cache --hicache-ratio 1 --hicache-io-backend direct --watchdog-timeout 1800 --context-length 460800'
    

    参考我的运行命令;THINKFIX是为了默认开启thinking,打了个猴子补丁你可以忽略

  • See post chevron_right
    e411
    Members
    deepseek v4 flash sglang启动参数? 已解决 2026年9月24日 08:56

    看到sglang 0.5.14镜像以及发布,这个版本是否支持v4 0731部署了?dspark是否可用?如可用命令行参数怎么写?谢谢!

  • See post chevron_right
    e411
    Members
    deepseek v4 flash sglang启动参数? 已解决 2026年9月21日 16:56

    实测sglang 0.5.12 + dsv4部署效果好于vllm;vllm-metax(即使是最新0.26.0)疑似精度或者推理逻辑有问题,会导致长上下文的性能严重退化,tool call和thinking也不太正常

  • See post chevron_right
    e411
    Members
    deepseek v4 flash sglang启动参数? 已解决 2026年9月21日 10:49

    deepseek v4 flash sglang启动参数?

  • See post chevron_right
    e411
    Members
    Qwen/Qwen3.8-27B-FP8这个模型可以在c500单卡上跑么 已解决 2026年9月18日 14:55

    不能,c500不支持fp8,所以所有fp8模型都不可以。如果是int8的话,应该可以跑int8尺寸大概30GB单卡塞得下

  • See post chevron_right
    e411
    Members
    GLM-5.3-Flash-W8A8 C500 8卡是否可部署,支持部署的vllm或者sglang镜像预计何时发布? 已解决 2026年9月18日 14:10

    www.modelscope.cn/models/metax-tech/GLM-5.3-Flash-W8A8

  • See post chevron_right
    e411
    Members
    8卡C500怎么部署DeepSeek-V4-Flash-0731-W8A8 已解决 2026年8月31日 10:44
    #!/bin/bash
    export MACA_SMALL_PAGESIZE_ENABLE=1
    export MACA_VLLM_ENABLE_MCTLASS_FUSED_MOE=0
    export MACA_VLLM_ENABLE_MCTLASS_PYTHON_API=1
    
    vllm serve /llm_models/metax-tech--DeepSeek-V4-Flash-0731-W8A8/snapshots/master --served-model-name deepseek-v4-flash --trust-remote-code \
    --kv-cache-dtype bfloat16 --block-size 256 --port 8001 --max-model-len 512000 --gpu-memory-utilization 0.8 \
    --tokenizer-mode deepseek_v4 --tool-call-parser deepseek_v4 \
    --enable-auto-tool-choice --reasoning-parser deepseek_v4 --default-chat-template-kwargs '{"enable_thinking": true}' \
    --compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE", "custom_ops":["all"]}' \
    --max-num-seq 4 -tp 8
    

    我的启动命令。其实vllm metax 0.24.0已经能跑,但是不能开mtp或者dspark,只开模型本体可以推理

  • See post chevron_right
    e411
    Members
    sglang 0.5.12重新发布的原因?和之前版本有何区别? 已解决 2026年8月23日 19:28

    我看到2026-08-21重新发布了一个使用新版本的maca的sglang 0.5.12镜像,这个镜像和之前07-22的sglang 0.5.12有何区别?和08-07发布的0.5.13版本,支持的模型有何区别?

    名称版本
    sglang:0.5.12-maca.ai3.8.2.7-torch2.10-py310-kylinv11-amd64
    推荐安装MACA Driver 3.8.2.xsglang 0.5.12Pytorch 2.10Python 3.10
    更新时间
    2026-08-21 14:31:17
    基本描述
    提供了面向曦云C500系列的 sglang 推理运行环境,适配多个 sglang 版本,支持 DeepSeek、Kimi-K2、Qwen 等多个系列模型,可用于提供 LLM 推理和性能验证服务。

    名称版本
    sglang:0.5.13-maca.ai3.8.1.3-torch2.10-py312-kylinv11-amd64
    推荐安装MACA Driver 3.8.1.xsglang 0.5.13Pytorch 2.10Python 3.12
    更新时间
    2026-08-07 19:05:52
    基本描述
    提供了面向曦云C500系列的 sglang 推理运行环境,适配多个 sglang 版本,支持 DeepSeek、Kimi-K2、Qwen 等多个系列模型,可用于提供 LLM 推理和性能验证服务。

    名称版本
    sglang:0.5.12-maca.ai3.8.0.3-torch2.8-py310-kylinv11-amd64
    推荐安装MACA Driver 3.8.0.xsglang 0.5.12Pytorch 2.8Python 3.10
    更新时间
    2026-07-22 17:51:33
    基本描述
    提供了面向曦云C500系列的 sglang 推理运行环境,适配多个 sglang 版本,支持 DeepSeek、Kimi-K2、Qwen 等多个系列模型,可用于提供 LLM 推理和性能验证服务。

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月21日 16:44

    您好,采用的运行参数和本主题中前述参数完全一样,区别只有换用了7.16的发布版本vllm 0.22.0. 此命令在前面使用的vllm 0.21.0 poc中是可以正常运行的,所以希望您能确认下这个公开版是不是真支持dsv4了

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月21日 11:30
    INFO 07-20 17:29:07 [__init__.py:44] Available plugins for group vllm.platform_plugins:
    INFO 07-20 17:29:07 [__init__.py:46] - metax -> vllm_metax:register
    INFO 07-20 17:29:07 [__init__.py:49] All plugins in this group will be loaded. Set `VLLM_PLUGINS` to control which plugins to load.
    INFO 07-20 17:29:07 [__init__.py:238] Platform plugin metax is activated
    INFO 07-20 17:29:07 [envs.py:123] Plugin sets VLLM_USE_FLASHINFER_SAMPLER to False. Reason: flashinfer sampler are not supported on maca
    INFO 07-20 17:29:07 [envs.py:123] Plugin sets VLLM_ENGINE_READY_TIMEOUT_S to 3600. Reason: set timeout to 3600s for model loading
    INFO 07-20 17:29:07 [envs.py:123] Plugin sets VLLM_FLOAT32_MATMUL_PRECISION to high. Reason: set float32 matmul precision to high for better performance on Maca platform
    INFO 07-20 17:29:07 [envs.py:123] Plugin sets VLLM_USE_V2_MODEL_RUNNER to False. Reason: v2 model runner is still under development and not fully tested on Maca platform, disable it by default
    WARNING 07-20 17:29:09 [config.py:70] Support for Transformers v4 is deprecated. The Transformers v4 codepath will become unmaintained in vLLM v0.22.0 and will be removed in vLLM v0.24.0. Please upgrade to Transformers v5: pip install --upgrade transformers
    INFO Print the version information of mcoplib during compilation.
    
    Version info:Mcoplib_Version = '0.4.7'
    Build_Maca_Version = '3.8.0.23'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '3c71b39'
    Vllm Op Version = 0.22.0
    SGlang Op Version  = 0.5.12
    
    INFO Staring Check the current MACA version of the operating environment.
    
    INFO: Release major.minor matching,  successful:3.8.
    
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'awq' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.awq.MacaAWQConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'awq_marlin' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.awq_marlin.MacaAWQMarlinConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'compressed-tensors' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.compressed_tensors.MacaCompressedTensorsConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'gptq' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.auto_gptq.MacaAutoGPTQConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'auto_gptq' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.auto_gptq.MacaAutoGPTQConfig'>.
    WARNING 07-20 17:29:28 [__init__.py:87] The quantization method 'moe_wna16' already exists and will be overwritten by the quantization config <class 'vllm_metax.quant_config.moe_wna16.MacaMoeWNA16Config'>.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepSeekMTPModel is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_mtp:DeepSeekMTP.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepseekV2ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v2:DeepseekV2ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepseekV3ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v2:DeepseekV3ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepseekV32ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v2:DeepseekV3ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture GlmMoeDsaForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v2:GlmMoeDsaForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepseekV4ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v4:DeepseekV4ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture DeepSeekV4MTPModel is already registered, and will be overwritten by the new model class vllm_metax.models.deepseek_v4_mtp:DeepSeekV4MTP.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture Step3p5MTP is already registered, and will be overwritten by the new model class vllm_metax.models.step3p5_mtp:Step3p5MTP.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture Qwen3OmniMoeForConditionalGeneration is already registered, and will be overwritten by the new model class vllm_metax.models.qwen3_omni_moe_thinker:Qwen3OmniMoeThinkerForConditionalGeneration.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture MiMoV2ForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.mimo_v2:MiMoV2ForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture MiMoV2FlashForCausalLM is already registered, and will be overwritten by the new model class vllm_metax.models.mimo_v2:MiMoV2FlashForCausalLM.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture MiMoV2MTPModel is already registered, and will be overwritten by the new model class vllm_metax.models.mimo_v2_mtp:MiMoV2MTP.
    WARNING 07-20 17:29:28 [registry.py:983] Model architecture MiMoV2OmniMTPModel is already registered, and will be overwritten by the new model class vllm_metax.models.mimo_v2_mtp:MiMoV2OmniMTP.
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]        █     █     █▄   ▄█
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]  ▄▄ ▄█ █     █     █ ▀▄▀ █  version 0.22.0
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]   █▄█▀ █     █     █     █  model   /llm_models/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]    ▀▀  ▀▀▀▀▀ ▀▀▀▀▀ ▀     ▀
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:344]
    (APIServer pid=13) INFO 07-20 17:29:28 [utils.py:278] non-default args: {'model_tag': '/llm_models/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8', 'default_chat_template_kwargs': {'enable_thinking': True}, 'enable_auto_tool_choice': True, 'tool_call_parser': 'deepseek_v4', 'port': 8001, 'model': '/llm_models/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8', 'tokenizer_mode': 'deepseek_v4', 'trust_remote_code': True, 'max_model_len': 512000, 'served_model_name': ['deepseek-v4-flash'], 'reasoning_parser': 'deepseek_v4', 'tensor_parallel_size': 8, 'block_size': 256, 'gpu_memory_utilization': 0.8, 'kv_cache_dtype': 'bfloat16', 'max_num_seqs': 2, 'enable_chunked_prefill': True, 'async_scheduling': True, 'compilation_config': {'mode': None, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': None, 'splitting_ops': None, 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': None, 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.FULL_AND_PIECEWISE: (2, 1)>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': None, 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': None, 'pass_config': {}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': None, 'static_all_moe_layers': []}}
    (APIServer pid=13) INFO 07-20 17:29:28 [config.py:431] Replacing legacy 'type' key with 'rope_type'
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] Error in inspecting model architecture 'DeepseekV4ForCausalLM'
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] Traceback (most recent call last):
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 1385, in _run_in_subprocess
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     returned.check_returncode()
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/subprocess.py", line 502, in check_returncode
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     raise CalledProcessError(self.returncode, self.args, self.stdout,
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] subprocess.CalledProcessError: Command '['/opt/conda/bin/python3.12', '-m', 'vllm.model_executor.models.registry']' returned non-zero exit status 1.
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] The above exception was the direct cause of the following exception:
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] Traceback (most recent call last):
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 948, in _try_inspect_model_cls
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     return model.inspect_model_cls()
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]            ^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/logging_utils/log_time.py", line 21, in _wrapper
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     result = func(*args, **kwargs)
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]              ^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 909, in inspect_model_cls
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     mi = _run_in_subprocess(
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]          ^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 1388, in _run_in_subprocess
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     raise RuntimeError(
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] RuntimeError: Error raised in subprocess:
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] <frozen runpy>:128: RuntimeWarning: 'vllm.model_executor.models.registry' found in sys.modules after import of package 'vllm.model_executor.models', but prior to execution of 'vllm.model_executor.models.registry'; this may result in unpredictable behaviour
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] Traceback (most recent call last):
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen runpy>", line 198, in _run_module_as_main
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen runpy>", line 88, in _run_code
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 1411, in <module>
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     _run()
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 1404, in _run
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     result = fn()
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]              ^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 910, in <lambda>
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     lambda: _ModelInfo.from_model_cls(self.load_model_cls())
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]                                       ^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm/model_executor/models/registry.py", line 923, in load_model_cls
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     mod = importlib.import_module(self.module_name)
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/importlib/__init__.py", line 90, in import_module
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     return _bootstrap._gcd_import(name[level:], package, level)
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 1387, in _gcd_import
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 1360, in _find_and_load
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 1331, in _find_and_load_unlocked
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 935, in _load_unlocked
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap_external>", line 999, in exec_module
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "<frozen importlib._bootstrap>", line 488, in _call_with_frames_removed
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm_metax/models/deepseek_v4.py", line 23, in <module>
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     from vllm_metax.customized.pluggable_layer.deepseek_v4_attention import (
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]   File "/opt/conda/lib/python3.12/site-packages/vllm_metax/customized/pluggable_layer/deepseek_v4_attention.py", line 9, in <module>
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]     from vllm.model_executor.layers.deepseek_v4_attention import (
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950] ModuleNotFoundError: No module named 'vllm.model_executor.layers.deepseek_v4_attention'
    (APIServer pid=13) ERROR 07-20 17:30:00 [registry.py:950]
    (APIServer pid=13) Traceback (most recent call last):
    (APIServer pid=13)   File "/opt/conda/bin/vllm", line 8, in <module>
    (APIServer pid=13)     sys.exit(main())
    (APIServer pid=13)              ^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/cli/main.py", line 92, in main
    (APIServer pid=13)     args.dispatch_function(args)
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/cli/serve.py", line 148, in cmd
    (APIServer pid=13)     uvloop.run(run_server(args))
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/uvloop/__init__.py", line 96, in run
    (APIServer pid=13)     return __asyncio.run(
    (APIServer pid=13)            ^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/asyncio/runners.py", line 195, in run
    (APIServer pid=13)     return runner.run(main)
    (APIServer pid=13)            ^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/asyncio/runners.py", line 118, in run
    (APIServer pid=13)     return self._loop.run_until_complete(task)
    (APIServer pid=13)            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "uvloop/loop.pyx", line 1518, in uvloop.loop.Loop.run_until_complete
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/uvloop/__init__.py", line 48, in wrapper
    (APIServer pid=13)     return await main
    (APIServer pid=13)            ^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py", line 678, in run_server
    (APIServer pid=13)     await run_server_worker(listen_address, sock, args, **uvicorn_kwargs)
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py", line 692, in run_server_worker
    (APIServer pid=13)     async with build_async_engine_client(
    (APIServer pid=13)                ^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/contextlib.py", line 210, in __aenter__
    (APIServer pid=13)     return await anext(self.gen)
    (APIServer pid=13)            ^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py", line 100, in build_async_engine_client
    (APIServer pid=13)     async with build_async_engine_client_from_engine_args(
    (APIServer pid=13)                ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/contextlib.py", line 210, in __aenter__
    (APIServer pid=13)     return await anext(self.gen)
    (APIServer pid=13)            ^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/api_server.py", line 124, in build_async_engine_client_from_engine_args
    (APIServer pid=13)     vllm_config = engine_args.create_engine_config(usage_context=usage_context)
    (APIServer pid=13)                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/engine/arg_utils.py", line 1730, in create_engine_config
    (APIServer pid=13)     model_config = self.create_model_config()
    (APIServer pid=13)                    ^^^^^^^^^^^^^^^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/vllm/engine/arg_utils.py", line 1560, in create_model_config
    (APIServer pid=13)     return ModelConfig(
    (APIServer pid=13)            ^^^^^^^^^^^^
    (APIServer pid=13)   File "/opt/conda/lib/python3.12/site-packages/pydantic/_internal/_dataclasses.py", line 121, in __init__
    (APIServer pid=13)     s.__pydantic_validator__.validate_python(ArgsKwargs(args, kwargs), self_instance=s)
    (APIServer pid=13) pydantic_core._pydantic_core.ValidationError: 1 validation error for ModelConfig
    (APIServer pid=13)   Value error, Model architectures ['DeepseekV4ForCausalLM'] failed to be inspected. Please check the logs for more details. [type=value_error, input_value=ArgsKwargs((), {'model': ...nderer_num_workers': 1}), input_type=ArgsKwargs]
    (APIServer pid=13)     For further information visit https://errors.pydantic.dev/2.13/v/value_error
    

    您好,部署按照之前命令无法成功,请你再确认一下

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月20日 08:07

    vllm-metax:0.22.0 7月16发布的这个新版本是否支持deepseek v4了?

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月2日 10:52

    你们现在有比较新的dsv4的验证镜像吗,我这边真的很想部署一个dsv4来用(

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月2日 10:42

    注意到vllm镜像里的tilelang没有maca适配代码

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月2日 10:40

    我进一步排查后,发现其中的tilelang有问题,我对比了贵公司官网发布的tilelang安装包与vllm镜像内的tilelang,发现vllm镜像内tilelang是上游版本,而不是贵公司适配过的tilelang maca版本,其中不含任何maca代码,只有cuda,vllm调用到了cuda,导致maca工具链编译报错,请联系贵公司研发重新适配打包这个vllm的镜像,谢谢!

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月2日 10:35

    没有变化,日志仍然是之前一样。

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月1日 17:18

    source /etc/profile.d/conda.sh 是为了加载vllm所在的虚拟环境,因为docker run -d拉起来的bash是非交互的,因此不会自动加载conda环境,为了正确加载后续vllm命令,必须加载一下conda.sh。

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月1日 16:52
    [Wed Jul  1 22:58:53 2026] AMD-Vi: Disabling interrupt remapping
    [Wed Jul  1 22:58:54 2026] ACPI: Using IOAPIC for interrupt routing
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKA configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKB configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKC configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKD configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKE configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKF configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKG configured for IRQ 0
    [Wed Jul  1 22:58:54 2026] ACPI: PCI: Interrupt link LNKH configured for IRQ 0
    [Wed Jul  1 22:58:57 2026] RAS: Correctable Errors collector initialized.
    [Wed Jul  1 22:59:12 2026] hrtimer: interrupt took 30958313 ns
    

    我认为这和卡没有任何关系,是vllm metax内部的一些软件错误

  • See post chevron_right
    e411
    Members
    6月30日发布的vllm 0.21.0是否支持deepseek v4部署?如果支持,正确的启动命令是? 已解决 2026年7月1日 16:02

    一、软硬件信息
    1.服务器厂家: 超云
    2.沐曦GPU型号: C500
    3.操作系统内核版本: 6.8.0-124-generic
    4.是否开启CPU虚拟化:是
    5.mx-smi回显:

    mx-smi  version: 2.3.4
    
    =================== MetaX System Management Interface Log ===================
    Timestamp                                         : Wed Jul  1 23:56:36 2026
    
    Attached GPUs                                     : 8
    +---------------------------------------------------------------------------------+
    | MX-SMI 2.3.4                       Kernel Mode Driver Version: 3.9.10           |
    | MACA Version: 3.7.0.38             BIOS Version: 1.35.2.0                       |
    |------------------+-----------------+---------------------+----------------------|
    | Board       Name | GPU   Persist-M | Bus-id              | GPU-Util      sGPU-M |
    | Pwr:Usage/Cap    | Temp       Perf | Memory-Usage        | GPU-State            |
    |==================+=================+=====================+======================|
    | 0     MetaX C500 | 0           Off | 0000:04:00.0        | 0%          Disabled |
    | 38W / 350W       | 40C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 1     MetaX C500 | 1           Off | 0000:05:00.0        | 0%          Disabled |
    | 39W / 350W       | 41C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 2     MetaX C500 | 2           Off | 0000:06:00.0        | 0%          Disabled |
    | 42W / 350W       | 43C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 3     MetaX C500 | 3           Off | 0000:07:00.0        | 0%          Disabled |
    | 38W / 350W       | 41C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 4     MetaX C500 | 4           Off | 0000:0b:00.0        | 0%          Disabled |
    | 40W / 350W       | 42C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 5     MetaX C500 | 5           Off | 0000:0c:00.0        | 0%          Disabled |
    | 41W / 350W       | 43C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 6     MetaX C500 | 6           Off | 0000:0d:00.0        | 0%          Disabled |
    | 42W / 350W       | 41C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    | 7     MetaX C500 | 7           Off | 0000:0e:00.0        | 0%          Disabled |
    | 39W / 350W       | 40C          P0 | 860/65536 MiB       | Available            |
    +------------------+-----------------+---------------------+----------------------+
    
    +---------------------------------------------------------------------------------+
    | Process:                                                                        |
    |  GPU                    PID         Process Name                 GPU Memory     |
    |                                                                  Usage(MiB)     |
    |=================================================================================|
    |  no process found                                                               |
    +---------------------------------------------------------------------------------+
    
    End of Log
    

    6.docker info回显:

    Client: Docker Engine - Community
     Version:    29.3.0
     Context:    default
     Debug Mode: false
     Plugins:
      buildx: Docker Buildx (Docker Inc.)
        Version:  v0.31.1
        Path:     /usr/libexec/docker/cli-plugins/docker-buildx
      compose: Docker Compose (Docker Inc.)
        Version:  v5.1.0
        Path:     /usr/libexec/docker/cli-plugins/docker-compose
      model: Docker Model Runner (Docker Inc.)
        Version:  v1.1.8
        Path:     /usr/libexec/docker/cli-plugins/docker-model
    
    Server:
     Containers: 12
      Running: 4
      Paused: 0
      Stopped: 8
     Images: 15
     Server Version: 29.3.0
     Storage Driver: overlayfs
      driver-type: io.containerd.snapshotter.v1
     Logging Driver: json-file
     Cgroup Driver: systemd
     Cgroup Version: 2
     Plugins:
      Volume: local
      Network: bridge host ipvlan macvlan null overlay
      Log: awslogs fluentd gcplogs gelf journald json-file local splunk syslog
     CDI spec directories:
      /etc/cdi
      /var/run/cdi
     Swarm: inactive
     Runtimes: io.containerd.runc.v2 runc
     Default Runtime: runc
     Init Binary: docker-init
     containerd version: dea7da592f5d1d2b7755e3a161be07f43fad8f75
     runc version: v1.3.4-0-gd6d73eb8
     init version: de40ad0
     Security Options:
      apparmor
      seccomp
       Profile: builtin
      cgroupns
     Kernel Version: 6.8.0-124-generic
     Operating System: Ubuntu-Server 24.04.4 LTS (Noble Numbat)
     OSType: linux
     Architecture: x86_64
     CPUs: 192
     Total Memory: 503.5GiB
     Name: server2
     ID: 404bb473-c286-41a5-a18e-b94481c6f73e
     Docker Root Dir: /var/lib/docker
     Debug Mode: false
     Experimental: false
     Insecure Registries:
      ::1/128
      127.0.0.0/8
     Live Restore Enabled: false
     Firewall Backend: iptables
    

    7.镜像版本:cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.21.0-maca.ai3.7.1.106-torch2.8-py312-ubuntu22.04-amd64
    8.启动容器命令:

    docker run -d \
        --device=/dev/dri \
        --device=/dev/mxcd \
        --device=/dev/infiniband \
        --group-add video \
        --security-opt seccomp=unconfined \
        --security-opt apparmor=unconfined \
        --shm-size 100gb \
        --ulimit memlock=-1 \
        --privileged=true \
        --network host \
        -v /mnt/modelscope:/llm_models \
        --name vllm-deepseek-v4-flash \
        cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.21.0-maca.ai3.7.1.106-torch2.8-py312-ubuntu22.04-amd64 bash -c "source /etc/profile.d/conda.sh && source /llm_models/run-deepseek-v4-entry.sh"
    

    9.容器内执行命令(/llm_models/run-deepseek-v4-entry.sh的内容):

    #!/bin/bash
    export MACA_SMALL_PAGESIZE_ENABLE=1
    export MACA_VLLM_ENABLE_MCTLASS_FUSED_MOE=0
    export MACA_VLLM_ENABLE_MCTLASS_PYTHON_API=1
    
    vllm serve /llm_models/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8 --served-model-name DeepSeek-V4-Flash --trust-remote-code \
    --kv-cache-dtype bfloat16 --block-size 256 --port 8001 --max-model-len 1048576 --gpu-memory-utilization 0.9 \
    --tokenizer-mode deepseek_v4 --tool-call-parser deepseek_v4 \
    --enable-auto-tool-choice --reasoning-parser deepseek_v4 --default-chat-template-kwargs '{"enable_thinking": true}' \
    --compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE", "custom_ops":["all"]}' \
    --max-num-seq 32 -tp 8 --speculative_config '{"method": "mtp", "num_speculative_tokens": 1}'
    

    二、问题现象
    请描述详细的问题现象日志。若日志过长,请上传附件(txt格式)。

  • 沐曦开发者论坛
powered by misago