MetaX-Tech Developer Forum 论坛首页
  • 沐曦开发者
search
Sign in
  • chevron_right Threads
  • label 产品&运维
  • label 已解决

8卡C500怎么部署DeepSeek-V4-Flash-0731-W8A8

wuying31231
2026年8月28日
chat_bubble_outline 2
  • link
    wuying31231
    Members 4 posts
    2026年8月28日 19:39 2026年8月28日 19:39
    link

    请问启动命令该怎么写

  • arrow_forward

    Thread has been moved from 产品&运维.

    • By shuai_chen on 2026年8月31日 10:18.
  • link
    shuai_chen
    Members 949 posts
    2026年8月31日 10:25 2026年8月31日 10:25
    link

    尊敬的开发者您好,等待近期镜像中心更新。若需申请POC镜像,请通过GPU购买商务渠道获取。

  • link
    e411
    Members 26 posts
    2026年8月31日 10:44 2026年8月31日 10:44
    link
    #!/bin/bash
    export MACA_SMALL_PAGESIZE_ENABLE=1
    export MACA_VLLM_ENABLE_MCTLASS_FUSED_MOE=0
    export MACA_VLLM_ENABLE_MCTLASS_PYTHON_API=1
    
    vllm serve /llm_models/metax-tech--DeepSeek-V4-Flash-0731-W8A8/snapshots/master --served-model-name deepseek-v4-flash --trust-remote-code \
    --kv-cache-dtype bfloat16 --block-size 256 --port 8001 --max-model-len 512000 --gpu-memory-utilization 0.8 \
    --tokenizer-mode deepseek_v4 --tool-call-parser deepseek_v4 \
    --enable-auto-tool-choice --reasoning-parser deepseek_v4 --default-chat-template-kwargs '{"enable_thinking": true}' \
    --compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE", "custom_ops":["all"]}' \
    --max-num-seq 4 -tp 8
    

    我的启动命令。其实vllm metax 0.24.0已经能跑,但是不能开mtp或者dspark,只开模型本体可以推理

  • arrow_forward

    Thread has been moved from 解决中.

    • By shuai_chen on 2026年9月7日 10:33.
arrow_upward Go to top
  • 沐曦开发者论坛
powered by misago