memfrac大概可以设置0.87,0.9会爆显存;max model len对应的大概能有450K;chunked_prefill_size 8192就行,SGLANG_DSPARK_KERNEL_COMMIT_KV_PROJ需要配置防止走fp8,别的不用全部回退,SGLANG_DSPARK_KERNEL_EXPAND_PREFILL这个必须配置,编译会报错。
docker run -d \
--device=/dev/dri --device=/dev/mxcd --device=/dev/infiniband \
--group-add video --security-opt seccomp=unconfined \
--security-opt apparmor=unconfined --shm-size 100gb --ulimit memlock=-1 \
--privileged=true --network host \
-v /mnt/nvme0n1p1/modelscope:/llm_models \
-e TORCHINDUCTOR_COMPILE_THREADS=1 \
-e PYTORCH_ALLOC_CONF=expandable_segments:True \
-e SGLANG_DEFAULT_THINKING=1 \
-e SGLANG_DSPARK_KERNEL_EXPAND_PREFILL=torch \
-e SGLANG_DSPARK_KERNEL_COMMIT_KV_PROJ=torch \
-e SGLANG_MOE_CONFIG_DIR=/llm_models/moe_configs \
--name sglang-0514-dspark \
cr.metax-tech.com/public-ai-release/maca/sglang:0.5.14-maca.ai3.8.2.122-torch2.10-py312-kylinv11-amd64 \
bash -c 'python3 /llm_models/sglang-0514-thinkfix.py; echo ===PATCH-PREAMBLE-DONE===; grep -q SGLANG-0514-DEFTHINK /opt/conda/lib/python3.12/site-packages/sglang/srt/entrypoints/openai/serving_chat.py || { echo ===THINKFIX-NOT-APPLIED===; exit 9; }; python -m sglang.launch_server --host 0.0.0.0 --port 30000 --model-path /llm_models/metax-tech--DeepSeek-V4-Flash-0731-W8A8/snapshots/master --served-model-name deepseek-v4-flash --trust-remote-code --tp 8 --quantization w8a8_int8 --reasoning-parser deepseek-v4 --tool-call-parser deepseekv4 --mem-fraction-static 0.87 --speculative-algorithm DSPARK --speculative-dspark-block-size 5 --kv-cache-dtype bf16 --mamba-ssm-dtype bfloat16 --enable-cache-report --chunked-prefill-size 8192 --enable-hierarchical-cache --hicache-ratio 1 --hicache-io-backend direct --watchdog-timeout 1800 --context-length 460800'
参考我的运行命令;THINKFIX是为了默认开启thinking,打了个猴子补丁你可以忽略