vllm-metax:0.25.0-maca.ai3.8.2.108-torch2.10-py310-ubuntu22.04-amd64
请问这个镜像支持Qwen3.8-27B吗
vllm-metax:0.25.0-maca.ai3.8.2.108-torch2.10-py310-ubuntu22.04-amd64
请问这个镜像支持Qwen3.8-27B吗
使用的是cr.metax-tech.com/public-ai-release/maca/sglang:0.5.12-maca.ai3.8.2.7-torch2.10-py312-ubuntu22.04-amd64 镜像
导致服务崩溃的直接原因是 SGLang 的模型加载器无法识别 Draft 模型的架构名称 DFlash2DraftModel
官方镜像啥时候能支持?
错误日志:
Capturing batches (bs=1 avail_mem=11.82 GB): 100%|█| 3/3 [00:20<00:00, 6.76s/i
[2026-08-30 18:23:41] Capture cuda graph end. Time elapsed: 20.78 s. mem usage=0.26 GB. avail mem=11.80 GB.
[2026-08-30 18:23:41] Disable piecewise CUDA graph because --disable-piecewise-cuda-graph is set
[2026-08-30 18:23:42] Init torch distributed begin.
[2026-08-30 18:23:42] Init torch distributed ends. elapsed=0.00 s, mem usage=0.00 GB
[2026-08-30 18:23:42] Load weight begin. avail mem=11.80 GB
[2026-08-30 18:23:42] Scheduler hit an exception: Traceback (most recent call last):
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 4050, in run_scheduler_process
scheduler = Scheduler(
^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 444, in init
self.init_model_worker()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 740, in init_model_worker
self.maybe_init_draft_worker()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 724, in maybe_init_draft_worker
self.draft_worker = DraftWorkerClass(**draft_worker_kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/speculative/dflash_worker.py", line 143, in init
self.draft_worker = TpModelWorker(
^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 262, in init
self._init_model_runner()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 347, in _init_model_runner
self._model_runner = ModelRunner(
^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 539, in init
self.initialize(pre_model_load_memory)
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 659, in initialize
self.load_model()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 1491, in load_model
self.model = self.loader.load_model(
^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 715, in load_model
quant_config = _get_quantization_config(model_config, self.load_config)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 198, in _get_quantization_config
model_class, _ = get_model_architecture(model_config)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/utils.py", line 222, in get_model_architecture
architectures = resolve_transformers_arch(model_config, architectures)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/utils.py", line 152, in resolve_transformers_arch
raise ValueError(
ValueError: Cannot find model module. 'DFlash2DraftModel' is not a registered model in the Transformers library (only relevant if the model is meant to be in Transformers) and 'AutoModel' is not present in the model config's 'auto_map' (relevant if the model is custom).
[2026-08-30 18:23:42] Received sigquit from a child process. It usually means the child failed.
start_sglang_server.sh: line 38: 35 Killed /opt/conda/bin/python3 -m sglang.launch_server --model-path /models/Qwen3.8-27B-W8A8-INT8 --served-model-name "Qwen3.8-27B" --host 0.0.0.0 --port 23600 --tp 1 --trust-remote-code --attention-backend flashinfer --mem-fraction-static 0.80 --context-length 131072 --chunked-prefill-size 16384 --max-running-requests 3 --schedule-policy fcfs --schedule-conservativeness 1.0 --mamba-scheduler-strategy extra_buffer --speculative-algorithm DFLASH --speculative-draft-model-path /models/Qwen3.8-27B-DFlash2 --speculative-num-draft-tokens 8 --tool-call-parser qwen3_coder --reasoning-parser qwen3
服务器镜像版本:cr.metax-tech.com/public-ai-release/maca/vllm-metax:0.24.0-maca.ai3.8.2.5-torch2.10-py312-ubuntu22.04-amd64
(APIServer pid=35) INFO: 127.0.0.1:54738 - "POST /v1/chat/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=35) ERROR: Exception in ASGI application
(APIServer pid=35) Traceback (most recent call last):
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/uvicorn/protocols/http/httptools_impl.py", line 422, in run_asgi
(APIServer pid=35) result = await app( # type: ignore[func-returns-value]
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/uvicorn/middleware/proxy_headers.py", line 63, in call
(APIServer pid=35) return await self.app(scope, receive, send)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/applications.py", line 1159, in call
(APIServer pid=35) await super().call(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/applications.py", line 96, in call
(APIServer pid=35) await self.middleware_stack(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/middleware/errors.py", line 186, in call
(APIServer pid=35) raise exc
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/middleware/errors.py", line 164, in call
(APIServer pid=35) await self.app(scope, receive, _send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/middleware/cors.py", line 88, in call
(APIServer pid=35) await self.app(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/prometheus_fastapi_instrumentator/middleware.py", line 180, in call
(APIServer pid=35) raise exc
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/prometheus_fastapi_instrumentator/middleware.py", line 178, in call
(APIServer pid=35) await self.app(scope, receive, send_wrapper)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/middleware/exceptions.py", line 63, in call
(APIServer pid=35) await wrap_app_handling_exceptions(self.app, conn)(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/_exception_handler.py", line 53, in wrapped_app
(APIServer pid=35) raise exc
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/_exception_handler.py", line 42, in wrapped_app
(APIServer pid=35) await app(scope, receive, sender)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/middleware/asyncexitstack.py", line 18, in call
(APIServer pid=35) await self.app(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/routing.py", line 670, in call
(APIServer pid=35) await self.middleware_stack(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/routing.py", line 690, in app
(APIServer pid=35) await route.handle(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/routing.py", line 280, in handle
(APIServer pid=35) await self.app(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/routing.py", line 134, in app
(APIServer pid=35) await wrap_app_handling_exceptions(app, request)(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/_exception_handler.py", line 53, in wrapped_app
(APIServer pid=35) raise exc
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/_exception_handler.py", line 42, in wrapped_app
(APIServer pid=35) await app(scope, receive, sender)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/routing.py", line 120, in app
(APIServer pid=35) response = await f(request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/routing.py", line 674, in app
(APIServer pid=35) raw_response = await run_endpoint_function(
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/routing.py", line 328, in run_endpoint_function
(APIServer pid=35) return await dependant.call(values)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/serve/utils/api_utils.py", line 91, in wrapper
(APIServer pid=35) return handler_task.result()
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/serve/utils/api_utils.py", line 112, in wrapper
(APIServer pid=35) return await func(*args, kwargs)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/api_router.py", line 61, in create_chat_completion
(APIServer pid=35) generator = await handler.create_chat_completion(request, raw_request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 247, in create_chat_completion
(APIServer pid=35) return await self._with_kv_transfer_rejection_cleanup(
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 176, in _with_kv_transfer_rejection_cleanup
(APIServer pid=35) return await awaitable
(APIServer pid=35) ^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 267, in _create_chat_completion
(APIServer pid=35) result = await self.render_chat_request(request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 233, in render_chat_request
(APIServer pid=35) return await self.openai_serving_render.render_chat(request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/serve/render/serving.py", line 387, in render_chat
(APIServer pid=35) conversation, engine_inputs = await self.preprocess_chat(
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/serve/render/serving.py", line 964, in preprocess_chat
(APIServer pid=35) ).adjust_request(
(APIServer pid=35) ^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/parser/abstract_parser.py", line 461, in adjust_request
(APIServer pid=35) request = self._apply_structural_tag(request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/parser/abstract_parser.py", line 487, in _apply_structural_tag
(APIServer pid=35) structure_tag = self._tool_parser.get_structural_tag(
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/tool_parsers/abstract_tool_parser.py", line 178, in get_structural_tag
(APIServer pid=35) from vllm.tool_parsers.structural_tag_registry import get_model_structural_tag
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/tool_parsers/structural_tag_registry.py", line 12, in <module>
(APIServer pid=35) from xgrammar import StructuralTag, normalize_tool_choice
(APIServer pid=35) ImportError: cannot import name 'normalize_tool_choice' from 'xgrammar' (/opt/conda/lib/python3.12/site-packages/xgrammar/init.py)
(APIServer pid=35) INFO: 127.0.0.1:34610 - "POST /v1/chat/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=35) ERROR: Exception in ASGI application
(APIServer pid=35) Traceback (most recent call last):
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/uvicorn/protocols/http/httptools_impl.py", line 422, in run_asgi
(APIServer pid=35) result = await app( # type: ignore[func-returns-value]
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/uvicorn/middleware/proxy_headers.py", line 63, in call
(APIServer pid=35) return await self.app(scope, receive, send)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/applications.py", line 1159, in call
(APIServer pid=35) await super().call(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/applications.py", line 96, in call
(APIServer pid=35) await self.middleware_stack(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/middleware/errors.py", line 186, in call
(APIServer pid=35) raise exc
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/middleware/errors.py", line 164, in call
(APIServer pid=35) await self.app(scope, receive, _send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/middleware/cors.py", line 88, in call
(APIServer pid=35) await self.app(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/prometheus_fastapi_instrumentator/middleware.py", line 180, in call
(APIServer pid=35) raise exc
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/prometheus_fastapi_instrumentator/middleware.py", line 178, in call
(APIServer pid=35) await self.app(scope, receive, send_wrapper)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/middleware/exceptions.py", line 63, in call
(APIServer pid=35) await wrap_app_handling_exceptions(self.app, conn)(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/_exception_handler.py", line 53, in wrapped_app
(APIServer pid=35) raise exc
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/_exception_handler.py", line 42, in wrapped_app
(APIServer pid=35) await app(scope, receive, sender)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/middleware/asyncexitstack.py", line 18, in call
(APIServer pid=35) await self.app(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/routing.py", line 670, in call
(APIServer pid=35) await self.middleware_stack(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/routing.py", line 690, in app
(APIServer pid=35) await route.handle(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/routing.py", line 280, in handle
(APIServer pid=35) await self.app(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/routing.py", line 134, in app
(APIServer pid=35) await wrap_app_handling_exceptions(app, request)(scope, receive, send)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/_exception_handler.py", line 53, in wrapped_app
(APIServer pid=35) raise exc
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/starlette/_exception_handler.py", line 42, in wrapped_app
(APIServer pid=35) await app(scope, receive, sender)
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/routing.py", line 120, in app
(APIServer pid=35) response = await f(request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/routing.py", line 674, in app
(APIServer pid=35) raw_response = await run_endpoint_function(
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/fastapi/routing.py", line 328, in run_endpoint_function
(APIServer pid=35) return await dependant.call(values)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/serve/utils/api_utils.py", line 91, in wrapper
(APIServer pid=35) return handler_task.result()
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/serve/utils/api_utils.py", line 112, in wrapper
(APIServer pid=35) return await func(*args, kwargs)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/api_router.py", line 61, in create_chat_completion
(APIServer pid=35) generator = await handler.create_chat_completion(request, raw_request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 247, in create_chat_completion
(APIServer pid=35) return await self._with_kv_transfer_rejection_cleanup(
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 176, in _with_kv_transfer_rejection_cleanup
(APIServer pid=35) return await awaitable
(APIServer pid=35) ^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 267, in _create_chat_completion
(APIServer pid=35) result = await self.render_chat_request(request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 233, in render_chat_request
(APIServer pid=35) return await self.openai_serving_render.render_chat(request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/serve/render/serving.py", line 387, in render_chat
(APIServer pid=35) conversation, engine_inputs = await self.preprocess_chat(
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/entrypoints/serve/render/serving.py", line 964, in preprocess_chat
(APIServer pid=35) ).adjust_request(
(APIServer pid=35) ^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/parser/abstract_parser.py", line 461, in adjust_request
(APIServer pid=35) request = self._apply_structural_tag(request)
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/parser/abstract_parser.py", line 487, in _apply_structural_tag
(APIServer pid=35) structure_tag = self._tool_parser.get_structural_tag(
(APIServer pid=35) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/tool_parsers/abstract_tool_parser.py", line 178, in get_structural_tag
(APIServer pid=35) from vllm.tool_parsers.structural_tag_registry import get_model_structural_tag
(APIServer pid=35) File "/opt/conda/lib/python3.12/site-packages/vllm/tool_parsers/structural_tag_registry.py", line 12, in <module>
(APIServer pid=35) from xgrammar import StructuralTag, normalize_tool_choice
(APIServer pid=35) ImportError: cannot import name 'normalize_tool_choice' from 'xgrammar' (/opt/conda/lib/python3.12/site-packages/xgrammar/init.py)