镜像:cr.metax-tech.com/public-ai-release/maca/sglang:0.5.14-maca.ai3.8.2.122-torch2.10-py310-ubuntu22.04-amd64
模型:DeepSeek-V4-Flash-0731-W8A8
硬件:曦云 C500 ×16(单机)
【启动参数】
python3 -m sglang.launch_server \
--model-path "$MODEL" --served-model-name deepseek-v4-flash \
--host 0.0.0.0 --port "$PORT" \
--tp-size 8 --context-length 98304 --chunked-prefill-size 8192 \
--trust-remote-code --kv-cache-dtype bfloat16 \
--speculative-algorithm DSPARK \
--tool-call-parser deepseekv4 --reasoning-parser deepseek-v4
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
【问题 1】这个镜像目前是否只支持切到8张卡?是否支持16张?
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
【问题 2】加DSPARK后第一次前向就报错;按提示回退后性能反而退化4倍。
不加--speculative-algorithm时一切正常。加上DSPARK后第一次前向报:
File ".../sglang/srt/speculative/dspark_components/kernels/commit_kv_proj.py", line 138,
in _dequant_linear_weight
assert weight.dtype == torch.float8_e4m3fn, (
AssertionError: unsupported wkv weight dtype torch.int8 for the fused commit kv proj;
set SGLANG_DSPARK_KERNEL_COMMIT_KV_PROJ=torch
按报错提示把22个SGLANG_DSPARK_KERNEL_*=torch 全部回退后,服务能起来、请求也正常,
但同机同口径压测(1024 输入/强制输出256token/唯一前缀,8 并发):
(1)不投机:单人55.8 tok/s ,TPOT17.92 ms;
(2)DSPARK + 22 个回退:单人13.6 tok/s ,TPOT 73.47 ms。