MetaX-Tech Developer Forum 论坛首页
  • 沐曦开发者
search
Sign in

zhangpei

  • Members
  • Joined 2026年9月17日
  • message 帖子
  • forum 主题
  • favorite 关注者
  • favorite_border Follows
  • person_outline 详细信息

zhangpei has posted 5 messages.

  • See post chevron_right
    zhangpei
    Members
    Qwen3.8 GPU C500X2 有没有专门的SGLang 镜像使用 已解决 2026年9月22日 13:58

    在宿主机安装吗

  • See post chevron_right
    zhangpei
    Members
    Qwen3.8 GPU C500X2 有没有专门的SGLang 镜像使用 已解决 2026年9月22日 11:41

    操作系统: 银河麒麟 V10(ky10)
    内核版本: 4.19.90-52.23
    CPU 架构:aarch64(ARM64)
    CPU: 飞腾腾云 S5000C,64核
    GPU: 2× 沐曦曦云 C500(64GB/卡)
    内存: 128GB DDR5
    COU虚拟化 :CPU 虚拟化已开启,KVM 工作正常
    镜像版本:cr.metax-tech.com/public-ai-release/maca/sglang:0.5.9-maca.ai3.5.3.208-torch2.8-py312-kylin2309a-arm64
    docker info :
    Containers: 17
    Running: 12
    Paused: 0
    Stopped: 5
    Images: 12
    Server Version: 18.09.0
    Storage Driver: overlay2
    Backing Filesystem: xfs
    Supports d_type: true
    Native Overlay Diff: true
    Logging Driver: json-file
    Cgroup Driver: cgroupfs
    Hugetlb Pagesize: 2MB, 512MB (default is 512MB)
    Plugins:
    Volume: local
    Network: bridge host macvlan null overlay
    Log: awslogs fluentd gcplogs gelf journald json-file local logentries splunk syslog
    Swarm: inactive
    Runtimes: runc
    Default Runtime: runc
    Init Binary: docker-init
    containerd version:
    runc version: N/A
    init version: N/A (expected: )
    Security Options:
    seccomp
    Profile: default
    Kernel Version: 4.19.90-52.23.v2207.gfb08.ky10.aarch64
    Operating System: Kylin Linux Advanced Server V10 (GFB)
    OSType: linux
    Architecture: aarch64
    CPUs: 128
    Total Memory: 126.4GiB
    Name: localhost.localdomain
    ID: OHYA:TRAQ:HPVT:2AEC:OA6P:JFT2:MGWD:YWNM:NJ2Y:K56I:SZFD:FOV7
    Docker Root Dir: /var/lib/docker
    Debug Mode (client): false
    Debug Mode (server): false
    Registry: index.docker.io/v1/
    Labels:
    Experimental: false
    Insecure Registries:
    127.0.0.0/8
    Registry Mirrors:
    docker.1ms.run/
    docker.xuanyuan.me/
    Live Restore Enabled: true
    容器启动命令:
    docker run -it -d --device=/dev/dri --device=/dev/mxcd --group-add video --name qwen38 --network=host --privileged=true --security-opt seccomp=unconfined --security-opt apparmor=unconfined --shm-size 100gb --ulimit memlock=-1 -v /dev/shm/:/dev/shm/ -v /root/models:/root/models 831b60593042 /opt/conda/bin/python3 -m sglang.launch_server --host 0.0.0.0 --model-path /root/models/metax-tech/Qwen3.8-27B-W8A8 --tp 2 --trust-remote-code --port 30000 --context-length 262144 --mem-fraction-static 0.85
    日志:

  • See post chevron_right
    zhangpei
    Members
    Qwen3.8 GPU C500X2 有没有专门的SGLang 镜像使用 已解决 2026年9月22日 09:55

    就是使用的 最新的镜像 为什么还是不行 部署的是 Qwen3.8-27B-W8A8
    Qwen3.6-27B-W8A8 能完美启动 但是 Qwen3.8-27B-W8A8 启动就显示
    ValueError: Unsupported copy between dtypes: target.dtype=torch.int8, loaded_weight.dtype=torch.bfloat16
    附上 启动命令 docker run -it -d --device=/dev/dri --device=/dev/mxcd --group-add video --name qwen38 --network=host --privileged=true --security-opt seccomp=unconfined --security-opt apparmor=unconfined --shm-size 100gb --ulimit memlock=-1 -v /dev/shm/:/dev/shm/ -v /root/models:/root/models 831b60593042 /opt/conda/bin/python3 -m sglang.launch_server --host 0.0.0.0 --model-path /root/models/metax-tech/Qwen3.8-27B-W8A8 --tp 2 --trust-remote-code --port 30000 --context-length 262144 --mem-fraction-static 0.85

  • See post chevron_right
    zhangpei
    Members
    Qwen3.8 GPU C500X2 有没有专门的SGLang 镜像使用 已解决 2026年9月17日 18:08

    是你们的镜像中心下吗 能否提供地址 谢谢

  • See post chevron_right
    zhangpei
    Members
    Qwen3.8 GPU C500X2 有没有专门的SGLang 镜像使用 已解决 2026年9月17日 18:03

    Qwen3.8 GPU C500X2 有没有专门的SGLang 镜像使用

    docker run -it -d --device=/dev/dri --device=/dev/mxcd --group-add video --name qwen38 --network=host --privileged=true --security-opt seccomp=unconfined --security-opt apparmor=unconfined --shm-size 100gb --ulimit memlock=-1 -v /dev/shm/:/dev/shm/ -v /root/models:/root/models 831b60593042 /opt/conda/bin/python3 -m sglang.launch_server --host 0.0.0.0 --model-path /root/models/metax-tech/Qwen3.8-27B-W8A8 --tp 2 --trust-remote-code --port 30000 --context-length 262144 --mem-fraction-static 0.85

    启动不起来

    [root@localhost ~/models/metax-tech]# docker logs -f qwen38
    INFO Print the version information of mcoplib during compilation.

    Version info:Mcoplib_Version = '0.4.1'
    Build_Maca_Version = '3.5.3.21'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '8c775a3'
    Vllm Op Version = 0.15.0
    SGlang Op Version = 0.5.7 && 0.5.8

    INFO Staring Check the current MACA version of the operating environment.

    INFO: Release major.minor matching, successful:3.5.

    /opt/conda/lib/python3.12/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.12/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    W0917 17:30:35.466000 1 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
    W0917 17:30:35.466000 1 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
    Disabling overlap schedule since mamba no_buffer is not compatible with overlap schedule, try to use --disable-radix-cache if overlap schedule is necessary
    Cuda graph max bs is adjusted to 128.
    Embedding tp size is adjusted to 2.
    [2026-09-17 17:30:37] server_args=ServerArgs(model_path='/root/models/metax-tech/Qwen3.8-27B-W8A8', tokenizer_path='/root/models/metax-tech/Qwen3.8-27B-W8A8', tokenizer_mode='auto', tokenizer_worker_num=1, skip_tokenizer_init=False, load_format='auto', model_loader_extra_config='{}', trust_remote_code=True, context_length=262144, is_embedding=False, enable_multimodal=None, revision=None, model_impl='auto', host='0.0.0.0', port=30000, fastapi_root_path='', grpc_mode=False, skip_server_warmup=False, warmups=None, nccl_port=None, checkpoint_engine_wait_weights_before_ready=False, dtype='auto', quantization=None, quantization_param_path=None, kv_cache_dtype='auto', enable_fp32_lm_head=False, modelopt_quant=None, modelopt_checkpoint_restore_path=None, modelopt_checkpoint_save_path=None, modelopt_export_path=None, quantize_and_serve=False, rl_quant_profile=None, mem_fraction_static=0.85, max_running_requests=None, max_queued_requests=None, max_total_tokens=None, chunked_prefill_size=8200, enable_dynamic_chunking=False, max_prefill_tokens=16384, prefill_max_requests=None, schedule_policy='fcfs', enable_priority_scheduling=False, abort_on_priority_when_disabled=False, schedule_low_priority_values_first=False, priority_scheduling_preemption_threshold=10, schedule_conservativeness=1.0, page_size=1, swa_full_tokens_ratio=0.8, disable_hybrid_swa_memory=False, radix_eviction_policy='lru', enable_prefill_delayer=False, prefill_delayer_max_delay_passes=30, prefill_delayer_token_usage_low_watermark=None, prefill_delayer_forward_passes_buckets=None, prefill_delayer_wait_seconds_buckets=None, device='cuda', tp_size=2, pp_size=1, pp_max_micro_batch_size=None, pp_async_batch_depth=0, stream_interval=1, stream_output=False, random_seed=453567547, constrained_json_whitespace_pattern=None, constrained_json_disable_any_whitespace=False, watchdog_timeout=300, soft_watchdog_timeout=None, dist_timeout=None, download_dir=None, model_checksum=None, base_gpu_id=0, gpu_id_step=1, sleep_on_idle=False, custom_sigquit_handler=None, log_level='info', log_level_http=None, log_requests=False, log_requests_level=2, log_requests_format='text', log_requests_target=None, uvicorn_access_log_exclude_prefixes=[], crash_dump_folder=None, show_time_cost=False, enable_metrics=False, enable_metrics_for_all_schedulers=False, tokenizer_metrics_custom_labels_header='x-custom-labels', tokenizer_metrics_allowed_custom_labels=None, extra_metric_labels=None, bucket_time_to_first_token=None, bucket_inter_token_latency=None, bucket_e2e_request_latency=None, collect_tokens_histogram=False, prompt_tokens_buckets=None, generation_tokens_buckets=None, gc_warning_threshold_secs=0.0, decode_log_interval=40, enable_request_time_stats_logging=False, kv_events_config=None, enable_trace=False, otlp_traces_endpoint='localhost:4317', export_metrics_to_file=False, export_metrics_to_file_dir=None, api_key=None, admin_api_key=None, served_model_name='/root/models/metax-tech/Qwen3.8-27B-W8A8', weight_version='default', chat_template=None, hf_chat_template_name=None, completion_template=None, file_storage_path='sglang_storage', enable_cache_report=False, reasoning_parser=None, tool_call_parser=None, tool_server=None, sampling_defaults='model', dp_size=1, load_balance_method='round_robin', attn_cp_size=1, moe_dp_size=1, dist_init_addr=None, nnodes=1, node_rank=0, json_model_override_args='{}', preferred_sampling_params=None, enable_lora=None, enable_lora_overlap_loading=None, max_lora_rank=None, lora_target_modules=None, lora_paths=None, max_loaded_loras=None, max_loras_per_batch=8, lora_eviction_policy='lru', lora_backend='csgmv', max_lora_chunk_size=16, attention_backend='flashinfer', decode_attention_backend=None, prefill_attention_backend=None, sampling_backend='flashinfer', grammar_backend='xgrammar', mm_attention_backend=None, fp8_gemm_runner_backend='auto', fp4_gemm_runner_backend='flashinfer_cutlass', nsa_prefill_backend='flashmla_sparse', nsa_decode_backend='flashmla_kv', disable_flashinfer_autotune=False, mamba_backend='triton', speculative_algorithm=None, speculative_draft_model_path=None, speculative_draft_model_revision=None, speculative_draft_load_format=None, speculative_num_steps=None, speculative_eagle_topk=None, speculative_num_draft_tokens=None, speculative_accept_threshold_single=1.0, speculative_accept_threshold_acc=1.0, speculative_token_map=None, speculative_attention_mode='prefill', speculative_draft_attention_backend=None, speculative_moe_runner_backend='auto', speculative_moe_a2a_backend=None, speculative_draft_model_quantization=None, speculative_ngram_min_match_window_size=1, speculative_ngram_max_match_window_size=12, speculative_ngram_min_bfs_breadth=1, speculative_ngram_max_bfs_breadth=10, speculative_ngram_match_type='BFS', speculative_ngram_branch_length=18, speculative_ngram_capacity=10000000, enable_multi_layer_eagle=False, speculative_suffix_max_tree_depth=24, speculative_suffix_max_cached_requests=10000, speculative_suffix_max_spec_factor=1.0, speculative_suffix_min_token_prob=0.1, ep_size=1, moe_a2a_backend='none', moe_runner_backend='auto', flashinfer_mxfp4_moe_precision='default', enable_flashinfer_allreduce_fusion=False, deepep_mode='auto', ep_num_redundant_experts=0, ep_dispatch_algorithm=None, init_expert_location='trivial', enable_eplb=False, eplb_algorithm='auto', eplb_rebalance_num_iterations=1000, eplb_rebalance_layers_per_chunk=None, eplb_min_rebalancing_utilization_threshold=1.0, expert_distribution_recorder_mode=None, expert_distribution_recorder_buffer_size=1000, enable_expert_distribution_metrics=False, deepep_config=None, moe_dense_tp_size=None, moe_shared_expert_tp_size=None, elastic_ep_backend=None, mooncake_ib_device=None, embedding_tp_size=2, lmhead_tp_size=None, max_mamba_cache_size=None, mamba_ssm_dtype=None, mamba_full_memory_ratio=0.9, mamba_scheduler_strategy='no_buffer', mamba_track_interval=256, enable_hierarchical_cache=False, hicache_ratio=2.0, hicache_size=0, hicache_write_policy='write_through', hicache_io_backend='kernel', hicache_mem_layout='layer_first', disable_hicache_numa_detect=False, hicache_storage_backend=None, hicache_storage_prefetch_policy='best_effort', hicache_storage_backend_extra_config=None, hierarchical_sparse_attention_extra_config=None, enable_lmcache=False, kt_weight_path=None, kt_method='AMXINT4', kt_cpuinfer=None, kt_threadpool_count=2, kt_num_gpu_experts=None, kt_max_deferred_experts_per_token=None, dllm_algorithm=None, dllm_algorithm_config=None, enable_double_sparsity=False, ds_channel_config_path=None, ds_heavy_channel_num=32, ds_heavy_token_num=256, ds_heavy_channel_type='qk', ds_sparse_decode_threshold=4096, cpu_offload_gb=0, offload_group_size=-1, offload_num_in_group=1, offload_prefetch_step=1, offload_mode='cpu', multi_item_scoring_delimiter=None, disable_radix_cache=False, cuda_graph_max_bs=128, cuda_graph_bs=None, disable_cuda_graph=False, disable_cuda_graph_padding=False, enable_profile_cuda_graph=False, enable_cudagraph_gc=False, enable_layerwise_nvtx_marker=False, enable_nccl_nvls=False, enable_symm_mem=False, disable_flashinfer_cutlass_moe_fp4_allgather=False, enable_tokenizer_batch_encode=False, disable_tokenizer_batch_decode=False, disable_outlines_disk_cache=False, disable_custom_all_reduce=False, enable_mscclpp=False, enable_torch_symm_mem=False, disable_overlap_schedule=True, enable_mixed_chunk=False, enable_dp_attention=False, enable_dp_lm_head=False, enable_two_batch_overlap=False, enable_single_batch_overlap=False, tbo_token_distribution_threshold=0.48, enable_torch_compile=False, enable_piecewise_cuda_graph=False, enable_torch_compile_debug_mode=False, torch_compile_max_bs=32, piecewise_cuda_graph_max_tokens=8200, piecewise_cuda_graph_tokens=[4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], piecewise_cuda_graph_compiler='eager', torchao_config='', enable_nan_detection=False, enable_p2p_check=False, triton_attention_reduce_in_fp32=False, triton_attention_num_kv_splits=8, triton_attention_split_tile_size=None, num_continuous_decode_steps=1, delete_ckpt_after_loading=False, enable_memory_saver=False, enable_weights_cpu_backup=False, enable_draft_weights_cpu_backup=False, allow_auto_truncate=False, enable_custom_logit_processor=False, flashinfer_mla_disable_ragged=False, disable_shared_experts_fusion=False, disable_chunked_prefix_cache=False, disable_fast_image_processor=False, keep_mm_feature_on_device=False, enable_return_hidden_states=False, enable_return_routed_experts=False, scheduler_recv_interval=1, numa_node=None, enable_deterministic_inference=False, rl_on_policy_target=None, enable_attn_tp_input_scattered=False, enable_nsa_prefill_context_parallel=False, nsa_prefill_cp_mode='round-robin-split', enable_fused_qk_norm_rope=False, enable_precise_embedding_interpolation=False, enable_dynamic_batch_tokenizer=False, dynamic_batch_tokenizer_batch_size=32, dynamic_batch_tokenizer_batch_timeout=0.002, debug_tensor_dump_output_folder=None, debug_tensor_dump_layers=None, debug_tensor_dump_input_file=None, debug_tensor_dump_inject=False, disaggregation_mode='null', disaggregation_transfer_backend='mooncake', disaggregation_bootstrap_port=8998, disaggregation_decode_tp=None, disaggregation_decode_dp=None, disaggregation_prefill_pp=1, disaggregation_ib_device=None, disaggregation_decode_enable_offload_kvcache=False, num_reserved_decode_tokens=512, disaggregation_decode_polling_interval=1, encoder_only=False, language_only=False, encoder_transfer_backend='zmq_to_scheduler', encoder_urls=[], custom_weight_loader=[], weight_loader_disable_mmap=False, remote_instance_weight_loader_seed_instance_ip=None, remote_instance_weight_loader_seed_instance_service_port=None, remote_instance_weight_loader_send_weights_group_ports=None, remote_instance_weight_loader_backend='nccl', remote_instance_weight_loader_start_seed_via_transfer_engine=False, enable_pdmux=False, pdmux_config_path=None, sm_group_num=8, mm_max_concurrent_calls=32, mm_per_request_timeout=10.0, enable_broadcast_mm_inputs_process=False, enable_prefix_mm_cache=False, mm_enable_dp_encoder=False, mm_process_config={}, limit_mm_data_per_request=None, decrypted_config_file=None, decrypted_draft_config_file=None, forward_hooks=None)
    [2026-09-17 17:30:38] Ignore import error when loading sglang.srt.multimodal.processors.glmasr: cannot import name 'GlmAsrConfig' from 'transformers' (/opt/conda/lib/python3.12/site-packages/transformers/init.py)
    [2026-09-17 17:30:41] Using default HuggingFace chat template with detected content format: openai
    INFO Print the version information of mcoplib during compilation.

    Version info:Mcoplib_Version = '0.4.1'
    Build_Maca_Version = '3.5.3.21'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '8c775a3'
    Vllm Op Version = 0.15.0
    SGlang Op Version = 0.5.7 && 0.5.8

    INFO Staring Check the current MACA version of the operating environment.

    INFO: Release major.minor matching, successful:3.5.

    INFO Print the version information of mcoplib during compilation.

    Version info:Mcoplib_Version = '0.4.1'
    Build_Maca_Version = '3.5.3.21'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '8c775a3'
    Vllm Op Version = 0.15.0
    SGlang Op Version = 0.5.7 && 0.5.8

    INFO Staring Check the current MACA version of the operating environment.

    INFO: Release major.minor matching, successful:3.5.

    INFO Print the version information of mcoplib during compilation.

    Version info:Mcoplib_Version = '0.4.1'
    Build_Maca_Version = '3.5.3.21'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '8c775a3'
    Vllm Op Version = 0.15.0
    SGlang Op Version = 0.5.7 && 0.5.8

    INFO Staring Check the current MACA version of the operating environment.

    INFO: Release major.minor matching, successful:3.5.

    /opt/conda/lib/python3.12/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.12/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.12/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.12/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    W0917 17:30:59.509000 214 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
    W0917 17:30:59.509000 214 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
    /opt/conda/lib/python3.12/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.12/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    W0917 17:31:02.273000 213 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
    W0917 17:31:02.273000 213 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
    W0917 17:31:04.795000 212 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
    W0917 17:31:04.795000 212 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
    [2026-09-17 17:31:06 TP1] Mamba selective_state_update backend initialized: triton
    [2026-09-17 17:31:07 TP1] Init torch distributed begin.
    [2026-09-17 17:31:07 TP1] Failed to import pynvml with ModuleNotFoundError("No module named 'pynvml'")
    [2026-09-17 17:31:10 TP0] Mamba selective_state_update backend initialized: triton
    [2026-09-17 17:31:10 TP0] Init torch distributed begin.
    [2026-09-17 17:31:10 TP0] Failed to import pynvml with ModuleNotFoundError("No module named 'pynvml'")
    [Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
    [Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
    [Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
    [Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
    [2026-09-17 17:31:10 TP0] sglang is using nccl==2.16.5
    [2026-09-17 17:31:11 TP0] DCP disabled, dcp_size=1, tp_size=2
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [2026-09-17 17:31:11 TP1] Init torch distributed ends. elapsed=4.53 s, mem usage=0.70 GB
    [2026-09-17 17:31:11 TP0] Init torch distributed ends. elapsed=0.73 s, mem usage=0.70 GB
    [2026-09-17 17:31:12 TP0] Ignore import error when loading sglang.srt.models.glm_ocr: No module named 'transformers.models.glm_ocr'
    [2026-09-17 17:31:12 TP1] Ignore import error when loading sglang.srt.models.glm_ocr: No module named 'transformers.models.glm_ocr'
    [2026-09-17 17:31:12 TP0] Ignore import error when loading sglang.srt.models.glm_ocr_nextn: No module named 'transformers.models.glm_ocr'
    [2026-09-17 17:31:12 TP1] Ignore import error when loading sglang.srt.models.glm_ocr_nextn: No module named 'transformers.models.glm_ocr'
    [2026-09-17 17:31:12 TP0] Ignore import error when loading sglang.srt.models.glmasr: cannot import name 'GlmAsrConfig' from 'transformers' (/opt/conda/lib/python3.12/site-packages/transformers/init.py)
    [2026-09-17 17:31:12 TP1] Ignore import error when loading sglang.srt.models.glmasr: cannot import name 'GlmAsrConfig' from 'transformers' (/opt/conda/lib/python3.12/site-packages/transformers/init.py)
    [2026-09-17 17:31:13 TP1] Load weight begin. avail mem=61.40 GB
    [2026-09-17 17:31:13 TP0] Load weight begin. avail mem=61.40 GB
    /opt/conda/lib/python3.12/site-packages/compressed_tensors/quantization/quant_args.py:368: UserWarning: No observer is used for dynamic quant., setting to None
    warnings.warn(
    /opt/conda/lib/python3.12/site-packages/compressed_tensors/quantization/quant_args.py:368: UserWarning: No observer is used for dynamic quant., setting to None
    warnings.warn(
    [2026-09-17 17:31:13 TP1] Multimodal attention backend not set. Use triton_attn.
    [2026-09-17 17:31:13 TP1] Using triton_attn as multimodal attention backend.
    [2026-09-17 17:31:13 TP0] Multimodal attention backend not set. Use triton_attn.
    [2026-09-17 17:31:13 TP0] Using triton_attn as multimodal attention backend.
    torch_dtype is deprecated! Use dtype instead!
    torch_dtype is deprecated! Use dtype instead!
    [2026-09-17 17:31:14 TP1] using attn output gate!
    [2026-09-17 17:31:14 TP0] using attn output gate!
    [2026-09-17 17:31:15 TP1] Scheduler hit an exception: Traceback (most recent call last):
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 3146, in run_scheduler_process
    scheduler = Scheduler(
    ^^^^^^^^^^
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 368, in init
    self.init_model_worker()
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 566, in init_model_worker
    self.init_tp_model_worker()
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 525, in init_tp_model_worker
    self.tp_worker = TpModelWorker(
    ^^^^^^^^^^^^^^
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 247, in init
    self._init_model_runner()
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 330, in _init_model_runner
    self._model_runner = ModelRunner(
    ^^^^^^^^^^^^
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 422, in init
    self.initialize(min_per_gpu_memory)
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 502, in initialize
    self.load_model()
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 1011, in load_model
    self.model = self.loader.load_model(
    ^^^^^^^^^^^^^^^^^^^^^^^
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 677, in load_model
    self.load_weights_and_postprocess(
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 686, in load_weights_and_postprocess
    model.load_weights(weights)
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/models/qwen3_5.py", line 1150, in load_weights
    weight_loader(param, loaded_weight)
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/linear.py", line 428, in weight_loader_v2
    param.load_column_parallel_weight(
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/parameter.py", line 169, in load_column_parallel_weight
    copy_with_check(self.data, loaded_weight)
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/parameter.py", line 65, in copy_with_check
    raise ValueError(
    ValueError: Unsupported copy between dtypes: target.dtype=torch.int8, loaded_weight.dtype=torch.bfloat16

    [2026-09-17 17:31:15] Received sigquit from a child process. It usually means the child failed.
    Loading safetensors checkpoint shards: 0% 0/66 [00:00<?, ?it/s][2026-09-17 17:31:15 TP0] Scheduler hit an exception: Traceback (most recent call last):
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 3146, in run_scheduler_process
    scheduler = Scheduler(
    ^^^^^^^^^^
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 368, in init
    self.init_model_worker()
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 566, in init_model_worker
    self.init_tp_model_worker()
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/scheduler.py", line 525, in init_tp_model_worker
    self.tp_worker = TpModelWorker(
    ^^^^^^^^^^^^^^
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 247, in init
    self._init_model_runner()
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/managers/tp_worker.py", line 330, in _init_model_runner
    self._model_runner = ModelRunner(
    ^^^^^^^^^^^^
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 422, in init
    self.initialize(min_per_gpu_memory)
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 502, in initialize
    self.load_model()
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_executor/model_runner.py", line 1011, in load_model
    self.model = self.loader.load_model(
    ^^^^^^^^^^^^^^^^^^^^^^^
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 677, in load_model
    self.load_weights_and_postprocess(
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/model_loader/loader.py", line 686, in load_weights_and_postprocess
    model.load_weights(weights)
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/models/qwen3_5.py", line 1150, in load_weights
    weight_loader(param, loaded_weight)
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/linear.py", line 428, in weight_loader_v2
    param.load_column_parallel_weight(
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/parameter.py", line 169, in load_column_parallel_weight
    copy_with_check(self.data, loaded_weight)
    File "/opt/conda/lib/python3.12/site-packages/sglang/srt/layers/parameter.py", line 65, in copy_with_check
    raise ValueError(
    ValueError: Unsupported copy between dtypes: target.dtype=torch.int8, loaded_weight.dtype=torch.bfloat16

    Loading safetensors checkpoint shards: 0% 0/66 [00:00<?, ?it/s]

  • 沐曦开发者论坛
powered by misago