请问下 modelscope.cn/models/metax-tech/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8 这个量化模型在8卡沐曦C500上是否能够部署,如果要部署应使用哪个镜像,使用什么参数启动?
服务器:单机双路CPU,配备 8张沐曦(MetaX)C500-P GPU,单卡显存 64GB(总计512GB)。
驱动版本:Kernel Driver 3.3.12,MACA 3.2.1.10,BIOS 1.29.1.0。
请问下 modelscope.cn/models/metax-tech/DeepSeek-V4-Flash-FlexSMQ-AWQ-W8A8 这个量化模型在8卡沐曦C500上是否能够部署,如果要部署应使用哪个镜像,使用什么参数启动?
服务器:单机双路CPU,配备 8张沐曦(MetaX)C500-P GPU,单卡显存 64GB(总计512GB)。
驱动版本:Kernel Driver 3.3.12,MACA 3.2.1.10,BIOS 1.29.1.0。
尊敬的开发者您好,请使用vllm metax 0.24.0镜像,命令参考
export MACA_SMALL_PAGESIZE_ENABLE=1
export MACA_VLLM_ENABLE_MCTLASS_FUSED_MOE=1
export MACA_VLLM_ENABLE_MCTLASS_PYTHON_API=1
vllm serve /data/models/model_quant_opt/DeepSeek/DeepSeek-V4-flash_W8A8/ --trust-remote-code \
--kv-cache-dtype bfloat16 --block-size 256 --gpu-memory-utilization 0.85 \
--tokenizer-mode deepseek_v4 --tool-call-parser deepseek_v4 \
--enable-auto-tool-choice --reasoning-parser deepseek_v4 \
--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE", "custom_ops":["all"]}' \
--max-num-seq 32 -tp 8 --speculative_config '{"method": "mtp", "num_speculative_tokens": 1}'
按上面配置,模型从这里下载modelscope.cn/models/metax-tech/DeepSeek-V4-Flash-0731-W8A8/,报这个错误是什么问题
(APIServer pid=14574) File "/opt/conda/lib/python3.12/site-packages/transformers/configuration_utils.py", line 838, in _get_config_dict
(APIServer pid=14574) raise OSError(f"It looks like the config file at '{resolved_config_file}' is not a valid JSON file.")
(APIServer pid=14574) OSError: It looks like the config file at '/workspace/share/metax-tech/DeepSeek-V4-Flash-0731-W8A8/config.json' is not a valid JSON file.
尊敬的开发者您好,此镜像不支持DeepSeek-V4-Flash-0731-W8A8
应该用哪个镜像?我下载的是vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py312-ubuntu22.04-amd64
尊敬的开发者您好,请等待近期镜像中心更新。若需申请POC镜像,请通过GPU购买商务渠道获取。