MetaX-Tech Developer Forum 论坛首页
  • 沐曦开发者
search
Sign in

yuanhong

  • Members
  • Joined 2026年7月8日
  • message 帖子
  • forum 主题
  • favorite 关注者
  • favorite_border Follows
  • person_outline 详细信息

yuanhong has started 3 threads.

  • See post chevron_right
    yuanhong
    Members
    3.8版本的驱动包依赖的3.9版本的rpm包造成驱动安装失败 已解决 2026年7月17日 10:06

    ./metax-driver-3.8.0.10-rpm-aarch64.run -- --dkms -f
    Verifying archive integrity... 100% MD5 checksums are OK. All good.
    Uncompressing MetaX driver installer 100%
    uninstalling metax-linux-dkms-3.8.23-1.aarch64

    Uninstall of metax-linux module (version 3.8.23) beginning:
    Module metax-linux-3.8.23 for kernel 6.6.0-25.02.2510.013.uos25.aarch64 (aarch64).
    Before uninstall, this module version was ACTIVE on this kernel.

    metax.ko.xz:
    - Uninstallation
    - Deleting from: /lib/modules/6.6.0-25.02.2510.013.uos25.aarch64/extra/
    - Original module
    - No original module was found for this module on this kernel.
    - Use the dkms install command to reinstall any previous module version.

    Running the post_remove script:
    depmod....
    Deleting module metax-linux-3.8.23 completely from the DKMS tree.
    uninstalling mxsmt-3.7.0-1.ky10.aarch64
    uninstalling mxfw-3.7.0-1.noarch
    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_train.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_infer.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_train.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_infer.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_train.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_infer.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn.so.8 不是符号链接

    上次元数据过期检查:2:42:44 前,执行于 2026年07月17日 星期五 07时19分35秒。
    软件包 dkms-3.0.12-1.uos25.02.noarch 已安装。
    软件包 autoconf-2.71-10.uos25.noarch 已安装。
    软件包 automake-1.16.5-9.uos25.noarch 已安装。
    软件包 elfutils-libelf-devel-0.190-11.uos25.aarch64 已安装。
    依赖关系解决。
    无需任何处理。
    完毕!

    rpm/mxfw_.rpm
    rpm/mxsmt_
    .rpm
    dkms_rpm/metax-linux-dkms[-_]*.rpm
    installing packages under /opt/mxdriver
    installing rpm/mxfw_3.8.0-1.noarch.rpm
    Verifying... ################################# [100%]
    准备中... ################################# [100%]
    deepin_rpm_verify: HC disabled, allowing mxfw
    正在升级/安装...
    1:mxfw-3.8.0-1 ################################# [100%]
    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_train.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_infer.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_train.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_infer.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_train.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_infer.so.8 不是符号链接

    /sbin/ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn.so.8 不是符号链接

    deepin_rpm_verify: HC disabled, allowing mxfw
    installing rpm/mxsmt_3.8.0-1.aarch64.rpm
    Verifying... ################################# [100%]
    准备中... ################################# [100%]
    deepin_rpm_verify: HC disabled, allowing mxsmt
    正在升级/安装...
    1:mxsmt-3.8.0-1.ky10 ################################# [100%]
    ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_train.so.8 不是符号链接

    ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_ops_infer.so.8 不是符号链接

    ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_train.so.8 不是符号链接

    ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_cnn_infer.so.8 不是符号链接

    ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_train.so.8 不是符号链接

    ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn_adv_infer.so.8 不是符号链接

    ldconfig: /usr/local/cuda-12.8/targets/sbsa-linux/lib/libcudnn.so.8 不是符号链接

    deepin_rpm_verify: HC disabled, allowing mxsmt
    installing dkms_rpm/metax-linux-dkms-3.9.10-1.aarch64.rpm
    Verifying... ################################# [100%]
    准备中... ################################# [100%]
    deepin_rpm_verify: HC disabled, allowing metax-linux-dkms
    remove old metax.ko from initramfs
    正在升级/安装...
    1:metax-linux-dkms-3.9.10-1 ################################# [100%]
    Loading new metax-linux-3.9.10 DKMS files...
    Building for 6.6.0-25.02.2510.013.uos25.aarch64
    Building initial module for 6.6.0-25.02.2510.013.uos25.aarch64
    autoreconf: export WARNINGS=
    autoreconf: Entering directory '.'
    autoreconf: configure.ac: not using Gettext
    autoreconf: running: aclocal --force
    autoreconf: configure.ac: tracing
    autoreconf: configure.ac: not using Libtool
    autoreconf: configure.ac: not using Intltool
    autoreconf: configure.ac: not using Gtkdoc
    autoreconf: running: /usr/bin/autoconf --force
    configure.ac:23: warning: The macro `AC_HELP_STRING' is obsolete.
    configure.ac:23: You should run autoupdate.
    ./lib/autoconf/general.m4:204: AC_HELP_STRING is expanded from...
    ./lib/autoconf/general.m4:1534: AC_ARG_ENABLE is expanded from...
    m4/config.m4:14: AC_METAX_CONFIG is expanded from...
    configure.ac:23: the top level
    autoreconf: running: /usr/bin/autoheader --force
    autoreconf: configure.ac: not using Automake
    autoreconf: 'config/install-sh' is updated
    autoreconf: Leaving directory '.'
    fgrep: warning: fgrep is obsolescent; using grep -F
    Error! Bad return status for module build on kernel: 6.6.0-25.02.2510.013.uos25.aarch64 (aarch64)
    Consult /var/lib/dkms/metax-linux/3.9.10/build/make.log for more information.
    警告:%post(metax-linux-dkms-3.9.10-1.aarch64) 脚本执行失败,退出状态码为 10
    deepin_rpm_verify: HC disabled, allowing metax-linux-dkms

  • See post chevron_right
    yuanhong
    Members
    沐曦C500 卡使用最新版本的SDK3.7.1版本容器镜像无论是sglang还是vllm 引擎启动大模型都非常吃内存 已解决 2026年7月10日 15:11

    沐曦C500 卡使用最新版本的SDK3.7.1版本容器镜像无论是sglang还是vllm 引擎启动大模型都非常吃内存。只运行32B大模型宿主机128G的内存都会报内存溢出无法启动服务。尤其是SGLang 对内存的消耗更高。
    之前发布的帖子让通过降级sdk到3.5版本才解决问题。
    原贴地址:developer.metax-tech.com/forum/t/sglang-duo-qia-bing-xing-qi-dong-da-mo-xing-fu-wu-bu-guan-mo-xing-da-xiao-duo-xiao-du-bao-nei-cun-yi-chu/617/#post-2966

  • See post chevron_right
    yuanhong
    Members
    sglang 多卡并行启动大模型服务不管模型大小多小都报内存溢出 已解决 2026年7月8日 10:28

    沐曦卡和驱动信息:
    mx-smi
    mx-smi version: 2.3.1

    =================== MetaX System Management Interface Log ===================
    Timestamp : Wed Jul 8 10:26:24 2026

    Attached GPUs : 4
    +---------------------------------------------------------------------------------+
    | MX-SMI 2.3.1 Kernel Mode Driver Version: 3.8.30 |
    | MACA Version: 3.7.1.5 BIOS Version: 1.33.5.0 |
    |------------------+-----------------+---------------------+----------------------|
    | Board Name | GPU Persist-M | Bus-id | GPU-Util sGPU-M |
    | Pwr:Usage/Cap | Temp Perf | Memory-Usage | GPU-State |
    |==================+=================+=====================+======================|
    | 0 MetaX C500 | 0 Off | 0000:38:00.0 | 0% Disabled |
    | 31W / 350W | 37C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 1 MetaX C500 | 1 Off | 0000:49:00.0 | 0% Disabled |
    | 32W / 350W | 35C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 2 MetaX C500 | 2 Off | 0000:98:00.0 | 0% Disabled |
    | 34W / 350W | 39C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+
    | 3 MetaX C500 | 3 Off | 0000:c8:00.0 | 0% Disabled |
    | 31W / 350W | 38C P0 | 858/65536 MiB | Available |
    +------------------+-----------------+---------------------+----------------------+

    +---------------------------------------------------------------------------------+
    | Process: |
    | GPU PID Process Name GPU Memory |
    | Usage(MiB) |
    |=================================================================================|
    | no process found |
    +---------------------------------------------------------------------------------+

    docker容器启动命令:
    docker run -d \
    --network host \
    --ipc=host \
    --name=sglang-mx \
    --shm-size 100G \
    --ulimit memlock=-1 \
    --ulimit nofile=1024000 \
    --device=/dev/dri \
    --device=/dev/mxcd \
    --device=/dev/mem \
    --group-add video \
    --security-opt seccomp=unconfined \
    --security-opt apparmor=unconfined \
    -v /data/models/:/models:ro \
    cr.metax-tech.com/public-ai-release/maca/sglang:0.5.11-maca.ai3.7.1.107-torch2.8-py310-kylinv11-amd64 \
    sleep infinity

    sglang启动命令:
    python -m sglang.launch_server --model-path /models/DeepSeek-R1-Distill-Qwen-7B --host 0.0.0.0 --port 9100 --served-model-name deepseek --tensor-parallel-size 2 --mem-fraction-static 0.7 --nccl-port 29500 --trust-remote-code --enable-metrics --attention-backend triton --cpu-offload-gb 0 --disable-cuda-graph

    错误日志:
    python -m sglang.launch_server --model-path /models/DeepSeek-R1-Distill-Qwen-7B --host 0.0.0.0 --port 9100 --served-model-name deepseek --tensor-parallel-size 2 --mem-fraction-static 0.7 --nccl-port 29500 --trust-remote-code --enable-metrics --attention-backend triton --cpu-offload-gb 0 --disable-cuda-graph
    INFO Print the version information of mcoplib during compilation.

    Version info:Mcoplib_Version = '0.4.4'
    Build_Maca_Version = '3.7.1.5'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '4d9c16e'
    Vllm Op Version = 0.19.0
    SGlang Op Version = 0.5.10

    INFO Staring Check the current MACA version of the operating environment.

    INFO: Release major.minor matching, successful:3.7.

    W0708 10:19:42.990000 1905 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
    W0708 10:19:42.990000 1905 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
    /opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.10/site-packages/sglang/launch_server.py:60: UserWarning: 'python -m sglang.launch_server' is still supported, but 'sglang serve' is the recommended entrypoint.
    Example: sglang serve --model-path <model> [options]
    warnings.warn(
    [2026-07-08 10:19:46] Cuda graph max bs is adjusted to 128.
    [2026-07-08 10:19:47] flashinfer.comm is not available, falling back to standard implementation
    [2026-07-08 10:19:48] server_args=ServerArgs(model_path='/models/DeepSeek-R1-Distill-Qwen-7B', tokenizer_path='/models/DeepSeek-R1-Distill-Qwen-7B', tokenizer_mode='auto', tokenizer_backend='huggingface', tokenizer_worker_num=1, detokenizer_worker_num=1, skip_tokenizer_init=False, load_format='auto', model_loader_extra_config='{}', trust_remote_code=True, context_length=None, is_embedding=False, enable_multimodal=None, revision=None, model_impl='auto', host='0.0.0.0', port=9100, fastapi_root_path='', grpc_mode=False, skip_server_warmup=False, warmups=None, nccl_port=29500, checkpoint_engine_wait_weights_before_ready=False, ssl_keyfile=None, ssl_certfile=None, ssl_ca_certs=None, ssl_keyfile_password=None, enable_ssl_refresh=False, enable_http2=False, dtype='auto', quantization=None, quantization_param_path=None, kv_cache_dtype='auto', enable_fp32_lm_head=False, modelopt_quant=None, modelopt_checkpoint_restore_path=None, modelopt_checkpoint_save_path=None, modelopt_export_path=None, quantize_and_serve=False, rl_quant_profile=None, mem_fraction_static=0.7, max_running_requests=None, max_queued_requests=None, max_total_tokens=None, chunked_prefill_size=8200, enable_dynamic_chunking=False, max_prefill_tokens=16384, prefill_max_requests=None, schedule_policy='fcfs', enable_priority_scheduling=False, disable_priority_preemption=False, default_priority_value=None, abort_on_priority_when_disabled=False, schedule_low_priority_values_first=False, priority_scheduling_preemption_threshold=10, schedule_conservativeness=1.0, page_size=1, swa_full_tokens_ratio=0.8, disable_hybrid_swa_memory=False, radix_eviction_policy='lru', enable_prefill_delayer=False, prefill_delayer_max_delay_passes=30, prefill_delayer_token_usage_low_watermark=None, prefill_delayer_forward_passes_buckets=None, prefill_delayer_wait_seconds_buckets=None, device='cuda', tp_size=2, pp_size=1, pp_max_micro_batch_size=None, pp_async_batch_depth=0, stream_interval=1, batch_notify_size=16, stream_response_default_include_usage=False, incremental_streaming_output=False, enable_streaming_session=False, random_seed=186191225, constrained_json_whitespace_pattern=None, constrained_json_disable_any_whitespace=False, watchdog_timeout=300, soft_watchdog_timeout=None, dist_timeout=None, download_dir=None, model_checksum=None, base_gpu_id=0, gpu_id_step=1, sleep_on_idle=False, use_ray=False, custom_sigquit_handler=None, log_level='info', log_level_http=None, log_requests=False, log_requests_level=2, log_requests_format='text', log_requests_target=None, uvicorn_access_log_exclude_prefixes=[], crash_dump_folder=None, show_time_cost=False, enable_metrics=True, grpc_http_sidecar_port=None, enable_mfu_metrics=False, enable_metrics_for_all_schedulers=False, tokenizer_metrics_custom_labels_header='x-custom-labels', tokenizer_metrics_allowed_custom_labels=None, extra_metric_labels=None, bucket_time_to_first_token=None, bucket_inter_token_latency=None, bucket_e2e_request_latency=None, prompt_tokens_buckets=None, generation_tokens_buckets=None, gc_warning_threshold_secs=0.0, decode_log_interval=40, enable_request_time_stats_logging=False, kv_events_config=None, enable_trace=False, otlp_traces_endpoint='localhost:4317', export_metrics_to_file=False, export_metrics_to_file_dir=None, api_key=None, admin_api_key=None, served_model_name='deepseek', weight_version='default', chat_template=None, hf_chat_template_name=None, completion_template=None, file_storage_path='sglang_storage', enable_cache_report=False, reasoning_parser=None, strip_thinking_cache=False, tool_call_parser=None, tool_server=None, sampling_defaults='model', dp_size=1, load_balance_method='round_robin', attn_cp_size=1, moe_dp_size=1, dist_init_addr=None, nnodes=1, node_rank=0, json_model_override_args='{}', preferred_sampling_params=None, enable_lora=None, enable_lora_overlap_loading=None, max_lora_rank=None, lora_target_modules=None, lora_paths=None, max_loaded_loras=None, max_loras_per_batch=8, lora_eviction_policy='lru', lora_backend='csgmv', max_lora_chunk_size=16, experts_shared_outer_loras=None, lora_use_virtual_experts=False, lora_strict_loading=False, attention_backend='triton', decode_attention_backend=None, prefill_attention_backend=None, sampling_backend='flashinfer', grammar_backend='xgrammar', mm_attention_backend=None, fp8_gemm_runner_backend='auto', fp4_gemm_runner_backend='auto', nsa_prefill_backend='flashmla_sparse', nsa_decode_backend='flashmla_kv', disable_flashinfer_autotune=False, mamba_backend='triton', speculative_algorithm=None, speculative_draft_model_path=None, speculative_draft_model_revision=None, speculative_draft_load_format=None, speculative_num_steps=None, speculative_eagle_topk=None, speculative_num_draft_tokens=None, speculative_dflash_block_size=None, speculative_dflash_draft_window_size=None, speculative_accept_threshold_single=1.0, speculative_accept_threshold_acc=1.0, speculative_token_map=None, speculative_attention_mode='prefill', speculative_draft_attention_backend=None, speculative_moe_runner_backend='auto', speculative_moe_a2a_backend=None, speculative_draft_model_quantization=None, speculative_adaptive=False, speculative_adaptive_config=None, speculative_skip_dp_mlp_sync=False, speculative_ngram_min_bfs_breadth=1, speculative_ngram_max_bfs_breadth=10, speculative_ngram_match_type='BFS', speculative_ngram_max_trie_depth=18, speculative_ngram_capacity=10000000, speculative_ngram_external_corpus_path=None, speculative_ngram_external_sam_budget=0, speculative_ngram_external_corpus_max_tokens=10000000, enable_multi_layer_eagle=False, ep_size=1, moe_a2a_backend='none', moe_runner_backend='auto', record_nolora_graph=True, flashinfer_mxfp4_moe_precision='default', enable_flashinfer_allreduce_fusion=False, enforce_disable_flashinfer_allreduce_fusion=False, enable_aiter_allreduce_fusion=False, deepep_mode='auto', ep_num_redundant_experts=0, ep_dispatch_algorithm=None, init_expert_location='trivial', enable_eplb=False, eplb_algorithm='auto', eplb_rebalance_num_iterations=1000, eplb_rebalance_layers_per_chunk=None, eplb_min_rebalancing_utilization_threshold=1.0, expert_distribution_recorder_mode=None, expert_distribution_recorder_buffer_size=1000, enable_expert_distribution_metrics=False, deepep_config=None, moe_dense_tp_size=None, elastic_ep_backend=None, enable_elastic_expert_backup=False, mooncake_ib_device=None, elastic_ep_rejoin=False, max_mamba_cache_size=None, mamba_ssm_dtype=None, mamba_full_memory_ratio=0.9, mamba_scheduler_strategy='no_buffer', mamba_track_interval=256, linear_attn_backend='triton', linear_attn_decode_backend=None, linear_attn_prefill_backend=None, enable_hierarchical_cache=False, hicache_ratio=2.0, hicache_size=0, hicache_write_policy='write_through', hicache_io_backend='kernel', hicache_mem_layout='layer_first', hicache_storage_backend=None, hicache_storage_prefetch_policy='best_effort', hicache_storage_backend_extra_config=None, enable_hisparse=False, hisparse_config=None, enable_lmcache=False, kt_weight_path=None, kt_method='AMXINT4', kt_cpuinfer=None, kt_threadpool_count=2, kt_num_gpu_experts=None, kt_max_deferred_experts_per_token=None, dllm_algorithm=None, dllm_algorithm_config=None, cpu_offload_gb=0, offload_group_size=-1, offload_num_in_group=1, offload_prefetch_step=1, offload_mode='cpu', enable_mis=False, disable_radix_cache=False, cuda_graph_max_bs=128, cuda_graph_bs=None, disable_cuda_graph=True, disable_cuda_graph_padding=False, enable_breakable_cuda_graph=False, enable_profile_cuda_graph=False, enable_cudagraph_gc=False, debug_cuda_graph=False, enable_layerwise_nvtx_marker=False, enable_nccl_nvls=False, enable_symm_mem=False, disable_flashinfer_cutlass_moe_fp4_allgather=False, enable_tokenizer_batch_encode=False, disable_tokenizer_batch_decode=False, disable_outlines_disk_cache=False, disable_custom_all_reduce=False, enable_mscclpp=False, enable_torch_symm_mem=False, pre_warm_nccl=False, disable_overlap_schedule=False, enable_mixed_chunk=False, enable_dp_attention=False, enable_dp_attention_local_control_broadcast=False, enable_dp_lm_head=False, enable_two_batch_overlap=False, enable_single_batch_overlap=False, tbo_token_distribution_threshold=0.48, enable_torch_compile=False, disable_piecewise_cuda_graph=False, enforce_piecewise_cuda_graph=False, enable_torch_compile_debug_mode=False, torch_compile_max_bs=32, piecewise_cuda_graph_max_tokens=8200, piecewise_cuda_graph_tokens=[4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], piecewise_cuda_graph_compiler='eager', torchao_config='', enable_nan_detection=False, enable_p2p_check=False, triton_attention_reduce_in_fp32=False, triton_attention_num_kv_splits=8, triton_attention_split_tile_size=None, num_continuous_decode_steps=1, delete_ckpt_after_loading=False, enable_memory_saver=False, enable_weights_cpu_backup=False, enable_draft_weights_cpu_backup=False, allow_auto_truncate=False, enable_custom_logit_processor=False, flashinfer_mla_disable_ragged=False, disable_shared_experts_fusion=False, enforce_shared_experts_fusion=False, disable_chunked_prefix_cache=False, disable_fast_image_processor=False, keep_mm_feature_on_device=False, enable_return_hidden_states=False, enable_return_routed_experts=False, scheduler_recv_interval=1, numa_node=None, enable_deterministic_inference=False, rl_on_policy_target=None, enable_attn_tp_input_scattered=False, gc_threshold=None, enable_nsa_prefill_context_parallel=False, nsa_prefill_cp_mode='round-robin-split', enable_fused_qk_norm_rope=False, enable_precise_embedding_interpolation=False, enable_fused_moe_sum_all_reduce=False, enable_prefill_context_parallel=False, prefill_cp_mode='in-seq-split', enable_dynamic_batch_tokenizer=False, dynamic_batch_tokenizer_batch_size=32, dynamic_batch_tokenizer_batch_timeout=0.002, debug_tensor_dump_output_folder=None, debug_tensor_dump_layers=None, debug_tensor_dump_input_file=None, debug_tensor_dump_inject=False, disaggregation_mode='null', disaggregation_transfer_backend='mooncake', disaggregation_bootstrap_port=8998, disaggregation_ib_device=None, disaggregation_decode_enable_radix_cache=False, disaggregation_decode_enable_offload_kvcache=False, num_reserved_decode_tokens=512, disaggregation_decode_polling_interval=1, encoder_only=False, language_only=False, encoder_transfer_backend='zmq_to_scheduler', encoder_urls=[], enable_adaptive_dispatch_to_encoder=False, custom_weight_loader=[], weight_loader_disable_mmap=False, weight_loader_prefetch_checkpoints=False, weight_loader_prefetch_num_threads=4, remote_instance_weight_loader_seed_instance_ip=None, remote_instance_weight_loader_seed_instance_service_port=None, remote_instance_weight_loader_send_weights_group_ports=None, remote_instance_weight_loader_backend='nccl', remote_instance_weight_loader_start_seed_via_transfer_engine=False, engine_info_bootstrap_port=6789, modelexpress_config=None, enable_pdmux=False, pdmux_config_path=None, sm_group_num=8, enable_broadcast_mm_inputs_process=False, enable_prefix_mm_cache=False, mm_enable_dp_encoder=False, mm_process_config={}, limit_mm_data_per_request=None, enable_mm_global_cache=False, decrypted_config_file=None, decrypted_draft_config_file=None, forward_hooks=None, enable_quant_communications=False, msprobe_dump_config=None)
    [2026-07-08 10:19:49] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
    [2026-07-08 10:19:49] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
    [2026-07-08 10:19:49] Using default HuggingFace chat template with detected content format: string
    INFO Print the version information of mcoplib during compilation.

    Version info:Mcoplib_Version = '0.4.4'
    Build_Maca_Version = '3.7.1.5'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '4d9c16e'
    Vllm Op Version = 0.19.0
    SGlang Op Version = 0.5.10

    INFO Staring Check the current MACA version of the operating environment.

    INFO: Release major.minor matching, successful:3.7.

    INFO Print the version information of mcoplib during compilation.

    Version info:Mcoplib_Version = '0.4.4'
    Build_Maca_Version = '3.7.1.5'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '4d9c16e'
    Vllm Op Version = 0.19.0
    SGlang Op Version = 0.5.10

    INFO Staring Check the current MACA version of the operating environment.

    INFO: Release major.minor matching, successful:3.7.

    INFO Print the version information of mcoplib during compilation.

    Version info:Mcoplib_Version = '0.4.4'
    Build_Maca_Version = '3.7.1.5'
    GIT_BRANCH = 'HEAD'
    GIT_COMMIT = '4d9c16e'
    Vllm Op Version = 0.19.0
    SGlang Op Version = 0.5.10

    INFO Staring Check the current MACA version of the operating environment.

    INFO: Release major.minor matching, successful:3.7.

    W0708 10:20:13.533000 2113 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
    W0708 10:20:13.533000 2113 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
    W0708 10:20:13.533000 2111 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
    W0708 10:20:13.533000 2111 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
    W0708 10:20:13.533000 2112 site-packages/torch/utils/cpp_extension.py:2527] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation.
    W0708 10:20:13.533000 2112 site-packages/torch/utils/cpp_extension.py:2527] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.
    /opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.10/site-packages/torchvision/datapoints/init.py:12: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    /opt/conda/lib/python3.10/site-packages/torchvision/transforms/v2/init.py:54: UserWarning: The torchvision.datapoints and torchvision.transforms.v2 namespaces are still Beta. While we do not expect major breaking changes, some APIs may still change according to user feedback. Please submit any feedback you may have in this issue: github.com/pytorch/vision/issues/6753, and you can also check out github.com/pytorch/vision/issues/7319 to learn more about the APIs that we suspect might involve future changes. You can silence this warning by calling torchvision.disable_beta_transforms_warning().
    warnings.warn(_BETA_TRANSFORMS_WARNING)
    flashinfer.comm is not available, falling back to standard implementation
    flashinfer.comm is not available, falling back to standard implementation
    flashinfer.comm is not available, falling back to standard implementation
    [2026-07-08 10:20:20 TP0] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
    [2026-07-08 10:20:20 TP0] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
    [2026-07-08 10:20:20] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
    [2026-07-08 10:20:20 TP1] Fixing v5 tokenizer component mismatch for /models/DeepSeek-R1-Distill-Qwen-7B: pre_tokenizer Metaspace -> Sequence, decoder Sequence -> ByteLevel
    [2026-07-08 10:20:20] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
    [2026-07-08 10:20:20 TP0] Init torch distributed begin.
    [2026-07-08 10:20:20 TP1] Restoring add_bos_token=True for /models/DeepSeek-R1-Distill-Qwen-7B (was False after v5 loading)
    [2026-07-08 10:20:20 TP1] Init torch distributed begin.
    [Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
    [Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
    [Gloo] Rank 1 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
    [Gloo] Rank 0 is connected to 1 peer ranks. Expected number of connected peer ranks is : 1
    [2026-07-08 10:20:20 TP0] sglang is using nccl==2.16.5
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [2026-07-08 10:20:20 TP0] DCP disabled, dcp_size=1, tp_size=2
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
    [2026-07-08 10:20:20 TP1] Init torch distributed ends. elapsed=0.61 s, mem usage=0.08 GB
    [2026-07-08 10:20:20 TP0] Init torch distributed ends. elapsed=0.65 s, mem usage=0.08 GB
    [2026-07-08 10:20:22 TP0] Load weight begin. avail mem=62.03 GB
    [2026-07-08 10:20:22 TP1] Load weight begin. avail mem=62.03 GB
    Multi-thread loading shards: 50% Completed | 1/2 [00:13<00:13, 13.38s/it]
    muxi:2111:2831 [0] /workspace/out/Release/build/linux/x86_64/mccl/macaify/src/misc/socket.cc:38 MCCL WARN socketProgressOpt: Call to recv from 127.0.0.1<48913> failed : Connection reset by peer

    muxi:2111:2831 [0] /workspace/out/Release/build/linux/x86_64/mccl/macaify/src/misc/socket.cc:567 MCCL WARN socketStartConnect: Connect to 127.0.0.1<48913> failed : Software caused connection abort

    muxi:2111:2831 [0] /workspace/out/Release/build/linux/x86_64/mccl/macaify/src/misc/socket.cc:567 MCCL WARN socketStartConnect: Connect to 127.0.0.1<48913> failed : Software caused connection abort
    Multi-thread loading shards: 100% Completed | 2/2 [00:16<00:00, 8.41s/it]
    [2026-07-08 10:20:38 TP0] Load weight end. elapsed=16.99 s, type=Qwen2ForCausalLM, avail mem=54.71 GB, mem usage=7.32 GB.

    muxi:2111:2831 [0] /workspace/out/Release/build/linux/x86_64/mccl/macaify/src/misc/socket.cc:567 MCCL WARN socketStartConnect: Connect to 127.0.0.1<48913> failed : Software caused connection abort
    [2026-07-08 10:20:39 TP0] Scheduler hit an exception: Traceback (most recent call last):
    File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/scheduler.py", line 3943, in run_scheduler_process
    scheduler = Scheduler(
    File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/scheduler.py", line 439, in init
    self.init_model_worker()
    File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/scheduler.py", line 714, in init_model_worker
    self.init_tp_model_worker()
    File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/scheduler.py", line 669, in init_tp_model_worker
    self.tp_worker = TpModelWorker(**worker_kwargs)
    File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/tp_worker.py", line 260, in init
    self._init_model_runner()
    File "/opt/conda/lib/python3.10/site-packages/sglang/srt/managers/tp_worker.py", line 345, in _init_model_runner
    self._model_runner = ModelRunner(
    File "/opt/conda/lib/python3.10/site-packages/sglang/srt/model_executor/model_runner.py", line 510, in init
    self.initialize(pre_model_load_memory)
    File "/opt/conda/lib/python3.10/site-packages/sglang/srt/model_executor/model_runner.py", line 627, in initialize
    self.load_model()
    File "/opt/conda/lib/python3.10/site-packages/sglang/srt/model_executor/model_runner.py", line 1535, in load_model
    raise ValueError(
    ValueError: TP rank 0 could finish the model loading, but there are other ranks that didn't finish loading. It is likely due to unexpected failures (e.g., OOM) or a slow node.

    [2026-07-08 10:20:39] Received sigquit from a child process. It usually means the child failed.

  • 沐曦开发者论坛
powered by misago