KT Qwen3.5-35B-A3B 记录

(ktlab) root@DESKTOP-9TRG62N:~/ktransformers# /root/miniconda3/envs/ktlab/bin/python3.11 -m sglang.launch_server --host 0.0.0.0 --port 30000 --model "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B" --kt-weight-path "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B-Q4_K_M" --kt-cpuinfer 9 --kt-threadpool-count 1 --kt-num-gpu-experts 4 --kt-method LLAMAFILE --attention-backend triton --trust-remote-code --mem-fraction-static 0.6 --max-total-tokens 8192 --enable-mixed-chunk --disable-shared-experts-fusion --disable-cuda-graph --skip-server-warmup --disable-flashinfer-autotune

(ktlab) root@DESKTOP-9TRG62N:~/ktransformers# /root/miniconda3/envs/ktlab/bin/python3.11 -m sglang.launch_server --host 0.0.0.0 --port 30000 --model "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B" --kt-weight-path "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B-Q4_K_M" --kt-cpuinfer 9 --kt-threadpool-count 1 --kt-num-gpu-experts 8 --kt-method LLAMAFILE --attention-backend triton --trust-remote-code --mem-fraction-static 0.6 --max-total-tokens 8192 --enable-mixed-chunk --disable-shared-experts-fusion --disable-cuda-graph --skip-server-warmup --disable-flashinfer-autotune

这个显存占用不超过10G

(ktlab) root@DESKTOP-9TRG62N:~/ktransformers# /root/miniconda3/envs/ktlab/bin/python3.11 -m sglang.launch_server --host 0.0.0.0 --port 30000 --model "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B" --kt-weight-path "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B-Q4_K_M" --kt-cpuinfer 11 --kt-threadpool-count 1 --kt-num-gpu-experts 16 --kt-method LLAMAFILE --attention-backend triton --trust-remote-code --mem-fraction-static 0.88 --max-total-tokens 8192 --enable-mixed-chunk --disable-shared-experts-fusion --disable-cuda-graph --skip-server-warmup --disable-flashinfer-autotune

CPU利用率98%、内存使用20G

显存利用12.3G、利用率61%

2026-03-31 18:47:20 Prefill batch, #new-seq: 1, #new-token: 74, #cached-token: 0, full token usage: 0.01, mamba usage: 0.07, #running-req: 0, #queue-req: 1, input throughput (token/s): 0.00, cuda graph: False

2026-03-31 18:47:34 Prefill batch, #new-seq: 1, #new-token: 95, #cached-token: 0, full token usage: 0.02, mamba usage: 0.14, #running-req: 1, #queue-req: 0, input throughput (token/s): 5.28, cuda graph: False

2026-03-31 18:48:23 Decode batch, #running-req: 2, #full token: 249, full token usage: 0.03, mamba num: 4, mamba usage: 0.14, cuda graph: False, gen throughput (token/s): 0.20, #queue-req: 0

2026-03-31 18:49:06 Decode batch, #running-req: 2, #full token: 329, full token usage: 0.04, mamba num: 4, mamba usage: 0.14, cuda graph: False, gen throughput (token/s): 1.84, #queue-req: 0

2026-03-31 18:49:44 Decode batch, #running-req: 2, #full token: 409, full token usage: 0.05, mamba num: 4, mamba usage: 0.14, cuda graph: False, gen throughput (token/s): 2.15, #queue-req: 0

(ktlab) root@DESKTOP-9TRG62N:~/ktransformers# /root/miniconda3/envs/ktlab/bin/python3.11 -m sglang.launch_server --host 0.0.0.0 --port 30000 --model "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B" --kt-weight-path "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B-Q4_K_M" --kt-cpuinfer 6 --kt-threadpool-count 1 --kt-num-gpu-experts 24 --kt-method LLAMAFILE --attention-backend triton --trust-remote-code --mem-fraction-static 0.85 --max-total-tokens 4096 --chunked-prefill-size 1024 --enable-mixed-chunk --disable-shared-experts-fusion --disable-cuda-graph --skip-server-warmup --disable-flashinfer-autotune

CPU利用率75%、内存使用20.2G

显存利用12.9G、利用率42%

2026-03-31 18:58:51 Load weight end. elapsed=255.03 s, type=Qwen3_5MoeForConditionalGeneration, dtype=torch.bfloat16, avail mem=3.40 GB, mem usage=11.35 GB.

2026-03-31 18:58:51 Using KV cache dtype: torch.bfloat16

2026-03-31 18:58:51 Mamba Cache is allocated. max_mamba_cache_size: 9, conv_state size: 0.01GB, ssm_state size: 0.59GB

2026-03-31 18:58:51 KV Cache is allocated. #tokens: 4096, K size: 0.04 GB, V size: 0.04 GB

2026-03-31 18:58:51 Memory pool end. avail mem=2.66 GB

2026-03-31 18:58:53 Using hybrid linear attention backend for hybrid GDN models.

2026-03-31 18:58:53 CuTe DSL GDN decode enabled: False

2026-03-31 18:58:54 max_total_num_tokens=4096, chunked_prefill_size=1024, max_prefill_tokens=16384, max_running_requests=3, context_len=262144, available_gpu_mem=2.63 GB

2026-03-31 18:59:00\] INFO: Started server process \[7941

2026-03-31 18:59:00 INFO: Waiting for application startup.

2026-03-31 18:59:00 Using default chat sampling params from model generation config: {'repetition_penalty': 1.0, 'temperature': 1.0, 'top_k': 20, 'top_p': 0.95}

2026-03-31 18:59:01 INFO: Application startup complete.

2026-03-31 18:59:01 The server is fired up and ready to roll!

2026-03-31 18:59:01 INFO: Uvicorn running on http://0.0.0.0:30000 (Press CTRL+C to quit)

...

2026-03-31 19:01:36 Decode batch, #running-req: 1, #full token: 516, full token usage: 0.13, mamba num: 2, mamba usage: 0.22, cuda graph: False, gen throughput (token/s): 8.79, #queue-req: 1

2026-03-31 19:01:41 Decode batch, #running-req: 1, #full token: 556, full token usage: 0.14, mamba num: 2, mamba usage: 0.22, cuda graph: False, gen throughput (token/s): 8.82, #queue-req: 1

2026-03-31 19:01:46 Decode batch, #running-req: 1, #full token: 596, full token usage: 0.15, mamba num: 2, mamba usage: 0.22, cuda graph: False, gen throughput (token/s): 8.23, #queue-req: 1

2026-03-31 19:01:50 Decode batch, #running-req: 1, #full token: 636, full token usage: 0.16, mamba num: 2, mamba usage: 0.22, cuda graph: False, gen throughput (token/s): 8.82, #queue-req: 1

2026-03-31 19:01:55 Decode batch, #running-req: 1, #full token: 676, full token usage: 0.17, mamba num: 2, mamba usage: 0.22, cuda graph: False, gen throughput (token/s): 9.27, #queue-req: 1

2026-03-31 19:01:59 Decode batch, #running-req: 1, #full token: 716, full token usage: 0.17, mamba num: 2, mamba usage: 0.22, cuda graph: False, gen throughput (token/s): 8.92, #queue-req: 1

2026-03-31 19:02:03 Decode batch, #running-req: 1, #full token: 756, full token usage: 0.18, mamba num: 2, mamba usage: 0.22, cuda graph: False, gen throughput (token/s): 9.22, #queue-req: 1

2026-03-31 19:02:08 Decode batch, #running-req: 1, #full token: 796, full token usage: 0.19, mamba num: 2, mamba usage: 0.22, cuda graph: False, gen throughput (token/s): 8.81, #queue-req: 1

(ktlab) root@DESKTOP-9TRG62N:~/ktransformers# /root/miniconda3/envs/ktlab/bin/python3.11 -m sglang.launch_server --host 0.0.0.0 --port 30000 --model "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B" --kt-weight-path "/mnt/d/Caches/LLM Models/Qwen3.5/Qwen3.5-35B-A3B-Q4_K_M" --kt-cpuinfer 6 --kt-threadpool-count 1 --kt-num-gpu-experts 32 --kt-method LLAMAFILE --attention-backend triton --trust-remote-code --mem-fraction-static 0.96 --max-total-tokens 4096 --chunked-prefill-size 1024 --enable-mixed-chunk --disable-shared-experts-fusion --disable-cuda-graph --skip-server-warmup --disable-flashinfer-autotune

CPU利用率70%、内存使用20.6G

显存利用14.6G、利用率47%

2026-03-31 19:18:13 Load weight end. elapsed=280.22 s, type=Qwen3_5MoeForConditionalGeneration, dtype=torch.bfloat16, avail mem=1.50 GB, mem usage=13.24 GB.

2026-03-31 19:18:13 Using KV cache dtype: torch.bfloat16

2026-03-31 19:18:13 Mamba Cache is allocated. max_mamba_cache_size: 7, conv_state size: 0.01GB, ssm_state size: 0.47GB

2026-03-31 19:18:13 KV Cache is allocated. #tokens: 4096, K size: 0.04 GB, V size: 0.04 GB

2026-03-31 19:18:14 Memory pool end. avail mem=1.00 GB

2026-03-31 19:18:15 Using hybrid linear attention backend for hybrid GDN models.

2026-03-31 19:18:15 CuTe DSL GDN decode enabled: False

2026-03-31 19:18:17 max_total_num_tokens=4096, chunked_prefill_size=1024, max_prefill_tokens=16384, max_running_requests=2, context_len=262144, available_gpu_mem=1.01 GB

2026-03-31 19:18:22\] INFO: Started server process \[8142

2026-03-31 19:18:22 INFO: Waiting for application startup.

2026-03-31 19:18:22 Using default chat sampling params from model generation config: {'repetition_penalty': 1.0, 'temperature': 1.0, 'top_k': 20, 'top_p': 0.95}

2026-03-31 19:18:24 INFO: Application startup complete.

2026-03-31 19:18:24 The server is fired up and ready to roll!

2026-03-31 19:18:24 INFO: Uvicorn running on http://0.0.0.0:30000 (Press CTRL+C to quit)

nvcc warning : incompatible redefinition for option 'std', the last value of this option was used

nvcc warning : incompatible redefinition for option 'optimize', the last value of this option was used

nvcc fatal : Unsupported gpu architecture 'compute_120'

ninja: build stopped: subcommand failed.

2026-03-31 19:21:34 Decode batch, #running-req: 1, #full token: 241, full token usage: 0.06, mamba num: 2, mamba usage: 0.29, cuda graph: False, gen throughput (token/s): 9.26, #queue-req: 0

2026-03-31 19:21:39 Decode batch, #running-req: 1, #full token: 281, full token usage: 0.07, mamba num: 2, mamba usage: 0.29, cuda graph: False, gen throughput (token/s): 9.17, #queue-req: 0

2026-03-31 19:21:43 Decode batch, #running-req: 1, #full token: 321, full token usage: 0.08, mamba num: 2, mamba usage: 0.29, cuda graph: False, gen throughput (token/s): 8.87, #queue-req: 0

2026-03-31 19:21:48 Decode batch, #running-req: 1, #full token: 361, full token usage: 0.09, mamba num: 2, mamba usage: 0.29, cuda graph: False, gen throughput (token/s): 8.17, #queue-req: 0

2026-03-31 19:21:53 Decode batch, #running-req: 1, #full token: 401, full token usage: 0.10, mamba num: 2, mamba usage: 0.29, cuda graph: False, gen throughput (token/s): 8.47, #queue-req: 0

2026-03-31 19:21:58 Decode batch, #running-req: 1, #full token: 441, full token usage: 0.11, mamba num: 2, mamba usage: 0.29, cuda graph: False, gen throughput (token/s): 8.10, #queue-req: 0

相关推荐
Python图像识别3 小时前
40-【2027毕设】YOLO11面部口罩检测识别系统 - Python完整源码+PyQt5界面+训练模型+数据集
python·qt·课程设计
谢亮_vipxieliang3 小时前
Spring Boot 3.x 从零开始——环境搭建与第一个 REST 项目
java·spring boot·后端
赵大仁3 小时前
开源清单怎么维护才不烂掉?GitHub Actions 每周查死链 + 软 404 检测实录
python·ci/cd·自动化·开源项目·踩坑·github actions·json schema
玩AI的奶茶3 小时前
24GB 显存能跑多大的模型?参数量、精度与显存占用对照表
人工智能·python·算法·ai·aigc·gpu算力·算力租赁
kimnoic3 小时前
Python操作Git命令的详细指南
开发语言·python
Python图像识别3 小时前
39-【2027毕设】YOLO11行人摔倒检测系统 - Python完整源码+PyQt5界面+训练模型+数据集
python·qt·课程设计
Wang's Blog3 小时前
Java 项目部署之 Docker工具快速入门: Docker 架构拆解:镜像、容器、守护进程与 Registry
java·docker·架构
栗子~~3 小时前
SpringCloud Gateway 基于 Nacos 实现动态路由
java·spring cloud·gateway
计算机毕业编程指导师3 小时前
【计算机毕设】基于Hadoop的人口统计特征与肥胖风险关联分析的数据分析系统源码 毕业设计 选题推荐 毕设选题 数据分析 机器学习
大数据·hadoop·python·计算机·数据分析·课程设计·肥胖风险
量化吞吐机3 小时前
涨跌停价与最小变动价位,为什么要分开看?
python