在宿主机安装吗
操作系统: 银河麒麟 V10(ky10)
内核版本: 4.19.90-52.23
CPU 架构:aarch64(ARM64)
CPU: 飞腾腾云 S5000C,64核
GPU: 2× 沐曦曦云 C500(64GB/卡)
内存: 128GB DDR5
COU虚拟化 :CPU 虚拟化已开启,KVM 工作正常
镜像版本:cr.metax-tech.com/public-ai-release/maca/sglang:0.5.9-maca.ai3.5.3.208-torch2.8-py312-kylin2309a-arm64
docker info :
Containers: 17
Running: 12
Paused: 0
Stopped: 5
Images: 12
Server Version: 18.09.0
Storage Driver: overlay2
Backing Filesystem: xfs
Supports d_type: true
Native Overlay Diff: true
Logging Driver: json-file
Cgroup Driver: cgroupfs
Hugetlb Pagesize: 2MB, 512MB (default is 512MB)
Plugins:
Volume: local
Network: bridge host macvlan null overlay
Log: awslogs fluentd gcplogs gelf journald json-file local logentries splunk syslog
Swarm: inactive
Runtimes: runc
Default Runtime: runc
Init Binary: docker-init
containerd version:
runc version: N/A
init version: N/A (expected: )
Security Options:
seccomp
Profile: default
Kernel Version: 4.19.90-52.23.v2207.gfb08.ky10.aarch64
Operating System: Kylin Linux Advanced Server V10 (GFB)
OSType: linux
Architecture: aarch64
CPUs: 128
Total Memory: 126.4GiB
Name: localhost.localdomain
ID: OHYA:TRAQ:HPVT:2AEC:OA6P:JFT2:MGWD:YWNM:NJ2Y:K56I:SZFD:FOV7
Docker Root Dir: /var/lib/docker
Debug Mode (client): false
Debug Mode (server): false
Registry: index.docker.io/v1/
Labels:
Experimental: false
Insecure Registries:
127.0.0.0/8
Registry Mirrors:
docker.1ms.run/
docker.xuanyuan.me/
Live Restore Enabled: true
容器启动命令:
docker run -it -d --device=/dev/dri --device=/dev/mxcd --group-add video --name qwen38 --network=host --privileged=true --security-opt seccomp=unconfined --security-opt apparmor=unconfined --shm-size 100gb --ulimit memlock=-1 -v /dev/shm/:/dev/shm/ -v /root/models:/root/models 831b60593042 /opt/conda/bin/python3 -m sglang.launch_server --host 0.0.0.0 --model-path /root/models/metax-tech/Qwen3.8-27B-W8A8 --tp 2 --trust-remote-code --port 30000 --context-length 262144 --mem-fraction-static 0.85
日志:
就是使用的 最新的镜像 为什么还是不行 部署的是 Qwen3.8-27B-W8A8
Qwen3.6-27B-W8A8 能完美启动 但是 Qwen3.8-27B-W8A8 启动就显示
ValueError: Unsupported copy between dtypes: target.dtype=torch.int8, loaded_weight.dtype=torch.bfloat16
附上 启动命令 docker run -it -d --device=/dev/dri --device=/dev/mxcd --group-add video --name qwen38 --network=host --privileged=true --security-opt seccomp=unconfined --security-opt apparmor=unconfined --shm-size 100gb --ulimit memlock=-1 -v /dev/shm/:/dev/shm/ -v /root/models:/root/models 831b60593042 /opt/conda/bin/python3 -m sglang.launch_server --host 0.0.0.0 --model-path /root/models/metax-tech/Qwen3.8-27B-W8A8 --tp 2 --trust-remote-code --port 30000 --context-length 262144 --mem-fraction-static 0.85
是你们的镜像中心下吗 能否提供地址 谢谢
Qwen3.8 GPU C500X2 有没有专门的SGLang 镜像使用
docker run -it -d --device=/dev/dri --device=/dev/mxcd --group-add video --name qwen38 --network=host --privileged=true --security-opt seccomp=unconfined --security-opt apparmor=unconfined --shm-size 100gb --ulimit memlock=-1 -v /dev/shm/:/dev/shm/ -v /root/models:/root/models 831b60593042 /opt/conda/bin/python3 -m sglang.launch_server --host 0.0.0.0 --model-path /root/models/metax-tech/Qwen3.8-27B-W8A8 --tp 2 --trust-remote-code --port 30000 --context-length 262144 --mem-fraction-static 0.85
启动不起来
[root@localhost ~/models/metax-tech]# docker logs -f qwen38
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.1'
Build_Maca_Version = '3.5.3.21'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '8c775a3'
Vllm Op Version = 0.15.0
SGlang Op Version = 0.5.7 && 0.5.8
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.5.
/opt/conda/lib/python3.12/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.12/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
W0917 17:30:35.466000 1 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
W0917 17:30:35.466000 1 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
Disabling overlap schedule since mamba no_buffer is not compatible with overlap schedule, try to use --disable-radix-cache if overlap schedule is necessary
Cuda graph max bs is adjusted to 128.
Embedding tp size is adjusted to 2.
[2026-09-17 17:30:37] server_args=ServerArgs(model_path='/root/models/metax-tech/Qwen3.8-27B-W8A8', tokenizer_path='/root/models/metax-tech/Qwen3.8-27B-W8A8', tokenizer_mode='auto', tokenizer_worker_num=1, skip_tokenizer_init=False, load_format='auto', model_loader_extra_config='{}', trust_remote_code=True, context_length=262144, is_embedding=False, enable_multimodal=None, revision=None, model_impl='auto', host='0.0.0.0', port=30000, fastapi_root_path='', grpc_mode=False, skip_server_warmup=False, warmups=None, nccl_port=None, checkpoint_engine_wait_weights_before_ready=False, dtype='auto', quantization=None, quantization_param_path=None, kv_cache_dtype='auto', enable_fp32_lm_head=False, modelopt_quant=None, modelopt_checkpoint_restore_path=None, modelopt_checkpoint_save_path=None, modelopt_export_path=None, quantize_and_serve=False, rl_quant_profile=None, mem_fraction_static=0.85, max_running_requests=None, max_queued_requests=None, max_total_tokens=None, chunked_prefill_size=8200, enable_dynamic_chunking=False, max_prefill_tokens=16384, prefill_max_requests=None, schedule_policy='fcfs', enable_priority_scheduling=False, abort_on_priority_when_disabled=False, schedule_low_priority_values_first=False, priority_scheduling_preemption_threshold=10, schedule_conservativeness=1.0, page_size=1, swa_full_tokens_ratio=0.8, disable_hybrid_swa_memory=False, radix_eviction_policy='lru', enable_prefill_delayer=False, prefill_delayer_max_delay_passes=30, prefill_delayer_token_usage_low_watermark=None, prefill_delayer_forward_passes_buckets=None, prefill_delayer_wait_seconds_buckets=None, device='cuda', tp_size=2, pp_size=1, pp_max_micro_batch_size=None, pp_async_batch_depth=0, stream_interval=1, stream_output=False, random_seed=453567547, constrained_json_whitespace_pattern=None, constrained_json_disable_any_whitespace=False, watchdog_timeout=300, soft_watchdog_timeout=None, dist_timeout=None, download_dir=None, model_checksum=None, base_gpu_id=0, gpu_id_step=1, sleep_on_idle=False, custom_sigquit_handler=None, log_level='info', log_level_http=None, log_requests=False, log_requests_level=2, log_requests_format='text', log_requests_target=None, uvicorn_access_log_exclude_prefixes=[], crash_dump_folder=None, show_time_cost=False, enable_metrics=False, enable_metrics_for_all_schedulers=False, tokenizer_metrics_custom_labels_header='x-custom-labels', tokenizer_metrics_allowed_custom_labels=None, extra_metric_labels=None, bucket_time_to_first_token=None, bucket_inter_token_latency=None, bucket_e2e_request_latency=None, collect_tokens_histogram=False, prompt_tokens_buckets=None, generation_tokens_buckets=None, gc_warning_threshold_secs=0.0, decode_log_interval=40, enable_request_time_stats_logging=False, kv_events_config=None, enable_trace=False, otlp_traces_endpoint='localhost:4317', export_metrics_to_file=False, export_metrics_to_file_dir=None, api_key=None, admin_api_key=None, served_model_name='/root/models/metax-tech/Qwen3.8-27B-W8A8', weight_version='default', chat_template=None, hf_chat_template_name=None, completion_template=None, file_storage_path='sglang_storage', enable_cache_report=False, reasoning_parser=None, tool_call_parser=None, tool_server=None, sampling_defaults='model', dp_size=1, load_balance_method='round_robin', attn_cp_size=1, moe_dp_size=1, dist_init_addr=None, nnodes=1, node_rank=0, json_model_override_args='{}', preferred_sampling_params=None, enable_lora=None, enable_lora_overlap_loading=None, max_lora_rank=None, lora_target_modules=None, lora_paths=None, max_loaded_loras=None, max_loras_per_batch=8, lora_eviction_policy='lru', lora_backend='csgmv', max_lora_chunk_size=16, attention_backend='flashinfer', decode_attention_backend=None, prefill_attention_backend=None, sampling_backend='flashinfer', grammar_backend='xgrammar', mm_attention_backend=None, fp8_gemm_runner_backend='auto', fp4_gemm_runner_backend='flashinfer_cutlass', nsa_prefill_backend='flashmla_sparse', nsa_decode_backend='flashmla_kv', disable_flashinfer_autotune=False, mamba_backend='triton', speculative_algorithm=None, speculative_draft_model_path=None, speculative_draft_model_revision=None, speculative_draft_load_format=None, speculative_num_steps=None, speculative_eagle_topk=None, speculative_num_draft_tokens=None, speculative_accept_threshold_single=1.0, speculative_accept_threshold_acc=1.0, speculative_token_map=None, speculative_attention_mode='prefill', speculative_draft_attention_backend=None, speculative_moe_runner_backend='auto', speculative_moe_a2a_backend=None, speculative_draft_model_quantization=None, speculative_ngram_min_match_window_size=1, speculative_ngram_max_match_window_size=12, speculative_ngram_min_bfs_breadth=1, speculative_ngram_max_bfs_breadth=10, speculative_ngram_match_type='BFS', speculative_ngram_branch_length=18, speculative_ngram_capacity=10000000, enable_multi_layer_eagle=False, speculative_suffix_max_tree_depth=24, speculative_suffix_max_cached_requests=10000, speculative_suffix_max_spec_factor=1.0, speculative_suffix_min_token_prob=0.1, ep_size=1, moe_a2a_backend='none', moe_runner_backend='auto', flashinfer_mxfp4_moe_precision='default', enable_flashinfer_allreduce_fusion=False, deepep_mode='auto', ep_num_redundant_experts=0, ep_dispatch_algorithm=None, init_expert_location='trivial', enable_eplb=False, eplb_algorithm='auto', eplb_rebalance_num_iterations=1000, eplb_rebalance_layers_per_chunk=None, eplb_min_rebalancing_utilization_threshold=1.0, expert_distribution_recorder_mode=None, expert_distribution_recorder_buffer_size=1000, enable_expert_distribution_metrics=False, deepep_config=None, moe_dense_tp_size=None, moe_shared_expert_tp_size=None, elastic_ep_backend=None, mooncake_ib_device=None, embedding_tp_size=2, lmhead_tp_size=None, max_mamba_cache_size=None, mamba_ssm_dtype=None, mamba_full_memory_ratio=0.9, mamba_scheduler_strategy='no_buffer', mamba_track_interval=256, enable_hierarchical_cache=False, hicache_ratio=2.0, hicache_size=0, hicache_write_policy='write_through', hicache_io_backend='kernel', hicache_mem_layout='layer_first', disable_hicache_numa_detect=False, hicache_storage_backend=None, hicache_storage_prefetch_policy='best_effort', hicache_storage_backend_extra_config=None, hierarchical_sparse_attention_extra_config=None, enable_lmcache=False, kt_weight_path=None, kt_method='AMXINT4', kt_cpuinfer=None, kt_threadpool_count=2, kt_num_gpu_experts=None, kt_max_deferred_experts_per_token=None, dllm_algorithm=None, dllm_algorithm_config=None, enable_double_sparsity=False, ds_channel_config_path=None, ds_heavy_channel_num=32, ds_heavy_token_num=256, ds_heavy_channel_type='qk', ds_sparse_decode_threshold=4096, cpu_offload_gb=0, offload_group_size=-1, offload_num_in_group=1, offload_prefetch_step=1, offload_mode='cpu', multi_item_scoring_delimiter=None, disable_radix_cache=False, cuda_graph_max_bs=128, cuda_graph_bs=None, disable_cuda_graph=False, disable_cuda_graph_padding=False, enable_profile_cuda_graph=False, enable_cudagraph_gc=False, enable_layerwise_nvtx_marker=False, enable_nccl_nvls=False, enable_symm_mem=False, disable_flashinfer_cutlass_moe_fp4_allgather=False, enable_tokenizer_batch_encode=False, disable_tokenizer_batch_decode=False, disable_outlines_disk_cache=False, disable_custom_all_reduce=False, enable_mscclpp=False, enable_torch_symm_mem=False, disable_overlap_schedule=True, enable_mixed_chunk=False, enable_dp_attention=False, enable_dp_lm_head=False, enable_two_batch_overlap=False, enable_single_batch_overlap=False, tbo_token_distribution_threshold=0.48, enable_torch_compile=False, enable_piecewise_cuda_graph=False, enable_torch_compile_debug_mode=False, torch_compile_max_bs=32, piecewise_cuda_graph_max_tokens=8200, piecewise_cuda_graph_tokens=[4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], piecewise_cuda_graph_compiler='eager', torchao_config='', enable_nan_detection=False, enable_p2p_check=False, triton_attention_reduce_in_fp32=False, triton_attention_num_kv_splits=8, triton_attention_split_tile_size=None, num_continuous_decode_steps=1, delete_ckpt_after_loading=False, enable_memory_saver=False, enable_weights_cpu_backup=False, enable_draft_weights_cpu_backup=False, allow_auto_truncate=False, enable_custom_logit_processor=False, flashinfer_mla_disable_ragged=False, disable_shared_experts_fusion=False, disable_chunked_prefix_cache=False, disable_fast_image_processor=False, keep_mm_feature_on_device=False, enable_return_hidden_states=False, enable_return_routed_experts=False, scheduler_recv_interval=1, numa_node=None, enable_deterministic_inference=False, rl_on_policy_target=None, enable_attn_tp_input_scattered=False, enable_nsa_prefill_context_parallel=False, nsa_prefill_cp_mode='round-robin-split', enable_fused_qk_norm_rope=False, enable_precise_embedding_interpolation=False, enable_dynamic_batch_tokenizer=False, dynamic_batch_tokenizer_batch_size=32, dynamic_batch_tokenizer_batch_timeout=0.002, debug_tensor_dump_output_folder=None, debug_tensor_dump_layers=None, debug_tensor_dump_input_file=None, debug_tensor_dump_inject=False, disaggregation_mode='null', disaggregation_transfer_backend='mooncake', disaggregation_bootstrap_port=8998, disaggregation_decode_tp=None, disaggregation_decode_dp=None, disaggregation_prefill_pp=1, disaggregation_ib_device=None, disaggregation_decode_enable_offload_kvcache=False, num_reserved_decode_tokens=512, disaggregation_decode_polling_interval=1, encoder_only=False, language_only=False, encoder_transfer_backend='zmq_to_scheduler', encoder_urls=[], custom_weight_loader=[], weight_loader_disable_mmap=False, remote_instance_weight_loader_seed_instance_ip=None, remote_instance_weight_loader_seed_instance_service_port=None, remote_instance_weight_loader_send_weights_group_ports=None, remote_instance_weight_loader_backend='nccl', remote_instance_weight_loader_start_seed_via_transfer_engine=False, enable_pdmux=False, pdmux_config_path=None, sm_group_num=8, mm_max_concurrent_calls=32, mm_per_request_timeout=10.0, enable_broadcast_mm_inputs_process=False, enable_prefix_mm_cache=False, mm_enable_dp_encoder=False, mm_process_config={}, limit_mm_data_per_request=None, decrypted_config_file=None, decrypted_draft_config_file=None, forward_hooks=None)
[2026-09-17 17:30:38] Ignore import error when loading sglang.srt.multimodal.processors.glmasr: cannot import name 'GlmAsrConfig' from 'transformers' (/opt/conda/lib/python3.12/site-packages/transformers/init.py)
[2026-09-17 17:30:41] Using default HuggingFace chat template with detected content format: openai
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.1'
Build_Maca_Version = '3.5.3.21'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '8c775a3'
Vllm Op Version = 0.15.0
SGlang Op Version = 0.5.7 && 0.5.8
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.5.
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.1'
Build_Maca_Version = '3.5.3.21'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '8c775a3'
Vllm Op Version = 0.15.0
SGlang Op Version = 0.5.7 && 0.5.8
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.5.
INFO Print the version information of mcoplib during compilation.
Version info:Mcoplib_Version = '0.4.1'
Build_Maca_Version = '3.5.3.21'
GIT_BRANCH = 'HEAD'
GIT_COMMIT = '8c775a3'
Vllm Op Version = 0.15.0
SGlang Op Version = 0.5.7 && 0.5.8
INFO Staring Check the current MACA version of the operating environment.
INFO: Release major.minor matching, successful:3.5.
/opt/conda/lib/python3.12/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.12/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.12/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.12/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
W0917 17:30:59.509000 214 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
W0917 17:30:59.509000 214 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
/opt/conda/lib/python3.12/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
/opt/conda/lib/python3.12/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
warnings.warn(_BETA_TRANSFORMS_WARNING)
W0917 17:31:02.273000 213 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
W0917 17:31:02.273000 213 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
W0917 17:31:04.795000 212 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
W0917 17:31:04.795000 212 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
[2026-09-17 17:31:06 TP1] Mamba selective_state_update backend initialized: triton
[2026-09-17 17:31:07 TP1] Init torch distributed begin.
[2026-09-17 17:31:07 TP1] Failed to import pynvml with ModuleNotFoundError("No module named 'pynvml'")
[2026-09-17 17:31:10 TP0] Mamba selective_state_update backend initialized: triton
[2026-09-17 17:31:10 TP0] Init torch distributed begin.
[2026-09-17 17:31:10 TP0] Failed to import pynvml with ModuleNotFoundError("No module named 'pynvml'")
[Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
[2026-09-17 17:31:10 TP0] sglang is using nccl==2.16.5
[2026-09-17 17:31:11 TP0] DCP disabled, dcp_size=1, tp_size=2
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[2026-09-17 17:31:11 TP1] Init torch distributed ends. elapsed=4.53 s, mem usage=0.70 GB
[2026-09-17 17:31:11 TP0] Init torch distributed ends. elapsed=0.73 s, mem usage=0.70 GB
[2026-09-17 17:31:12 TP0] Ignore import error when loading sglang.srt.models.glm_ocr: No module named 'transformers.models.glm_ocr'
[2026-09-17 17:31:12 TP1] Ignore import error when loading sglang.srt.models.glm_ocr: No module named 'transformers.models.glm_ocr'
[2026-09-17 17:31:12 TP0] Ignore import error when loading sglang.srt.models.glm_ocr_nextn: No module named 'transformers.models.glm_ocr'
[2026-09-17 17:31:12 TP1] Ignore import error when loading sglang.srt.models.glm_ocr_nextn: No module named 'transformers.models.glm_ocr'
[2026-09-17 17:31:12 TP0] Ignore import error when loading sglang.srt.models.glmasr: cannot import name 'GlmAsrConfig' from 'transformers' (/opt/conda/lib/python3.12/site-packages/transformers/init.py)
[2026-09-17 17:31:12 TP1] Ignore import error when loading sglang.srt.models.glmasr: cannot import name 'GlmAsrConfig' from 'transformers' (/opt/conda/lib/python3.12/site-packages/transformers/init.py)
[2026-09-17 17:31:13 TP1] Load weight begin. avail mem=61.40 GB
[2026-09-17 17:31:13 TP0] Load weight begin. avail mem=61.40 GB
/opt/conda/lib/python3.12/site-packages/compressed_tensors/quantization/quant_args.py:368: UserWarning: No observer is used for dynamic quant., setting to None
warnings.warn(
/opt/conda/lib/python3.12/site-packages/compressed_tensors/quantization/quant_args.py:368: UserWarning: No observer is used for dynamic quant., setting to None
warnings.warn(
[2026-09-17 17:31:13 TP1] Multimodal attention backend not set. Use triton_attn.
[2026-09-17 17:31:13 TP1] Using triton_attn as multimodal attention backend.
[2026-09-17 17:31:13 TP0] Multimodal attention backend not set. Use triton_attn.
[2026-09-17 17:31:13 TP0] Using triton_attn as multimodal attention backend.
torch_dtype is deprecated! Use dtype instead!
torch_dtype is deprecated! Use dtype instead!
[2026-09-17 17:31:14 TP1] using attn output gate!
[2026-09-17 17:31:14 TP0] using attn output gate!
[2026-09-17 17:31:15 TP1] Scheduler hit an exception: Traceback (most recent call last):
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 3146, in run_scheduler_process
scheduler = Scheduler(
^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 368, in init
self.init_model_worker()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 566, in init_model_worker
self.init_tp_model_worker()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 525, in init_tp_model_worker
self.tp_worker = TpModelWorker(
^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 247, in init
self._init_model_runner()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 330, in _init_model_runner
self._model_runner = ModelRunner(
^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 422, in init
self.initialize(min_per_gpu_memory)
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 502, in initialize
self.load_model()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 1011, in load_model
self.model = self.loader.load_model(
^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 677, in load_model
self.load_weights_and_postprocess(
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 686, in load_weights_and_postprocess
model.load_weights(weights)
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/models/qwen3_5.py", line 1150, in load_weights
weight_loader(param, loaded_weight)
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/linear.py", line 428, in weight_loader_v2
param.load_column_parallel_weight(
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/parameter.py", line 169, in load_column_parallel_weight
copy_with_check(self.data, loaded_weight)
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/parameter.py", line 65, in copy_with_check
raise ValueError(
ValueError: Unsupported copy between dtypes: target.dtype=torch.int8, loaded_weight.dtype=torch.bfloat16
[2026-09-17 17:31:15] Received sigquit from a child process. It usually means the child failed.
Loading safetensors checkpoint shards: 0% 0/66 [00:00<?, ?it/s][2026-09-17 17:31:15 TP0] Scheduler hit an exception: Traceback (most recent call last):
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 3146, in run_scheduler_process
scheduler = Scheduler(
^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 368, in init
self.init_model_worker()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 566, in init_model_worker
self.init_tp_model_worker()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 525, in init_tp_model_worker
self.tp_worker = TpModelWorker(
^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 247, in init
self._init_model_runner()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 330, in _init_model_runner
self._model_runner = ModelRunner(
^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 422, in init
self.initialize(min_per_gpu_memory)
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 502, in initialize
self.load_model()
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 1011, in load_model
self.model = self.loader.load_model(
^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 677, in load_model
self.load_weights_and_postprocess(
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 686, in load_weights_and_postprocess
model.load_weights(weights)
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/models/qwen3_5.py", line 1150, in load_weights
weight_loader(param, loaded_weight)
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/linear.py", line 428, in weight_loader_v2
param.load_column_parallel_weight(
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/parameter.py", line 169, in load_column_parallel_weight
copy_with_check(self.data, loaded_weight)
File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/parameter.py", line 65, in copy_with_check
raise ValueError(
ValueError: Unsupported copy between dtypes: target.dtype=torch.int8, loaded_weight.dtype=torch.bfloat16
Loading safetensors checkpoint shards: 0% 0/66 [00:00<?, ?it/s]