MI50运行GLM-4.7-Flash的速度测试

模型版本:https://huggingface.co/unsloth/GLM-4.7-Flash-GGUF GLM-4.7-Flash-UD-Q4_K_XL.gguf

llama.cpp版本:b7933

复制代码
root@dev:~# llama-bench -m glm-4.7-flash.gguf 
ggml_cuda_init: found 1 ROCm devices:
  Device 0: AMD Radeon Graphics, gfx906:sramecc-:xnack- (0x906), VMM: no, Wave Size: 64
| model                          |       size |     params | backend    | ngl |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: |
| deepseek2 30B.A3B Q4_K - Medium |  16.31 GiB |    29.94 B | ROCm       |  99 |           pp512 |        918.25 ± 1.37 |
| deepseek2 30B.A3B Q4_K - Medium |  16.31 GiB |    29.94 B | ROCm       |  99 |           tg128 |         54.62 ± 0.12 |

实际使用体验:

复制代码
srv  params_from_: Chat format: GLM 4.5
slot get_availabl: id  1 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 1.000
slot launch_slot_: id  1 | task -1 | sampler chain: logits -> ?penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> top-p -> min-p -> ?xtc -> ?temp-ext -> dist 
slot launch_slot_: id  1 | task 12639 | processing task, is_child = 0
slot update_slots: id  1 | task 12639 | new prompt, n_ctx_slot = 32768, n_keep = 0, task.n_tokens = 32375
slot update_slots: id  1 | task 12639 | n_tokens = 32366, memory_seq_rm [32366, end)
slot update_slots: id  1 | task 12639 | prompt processing progress, n_tokens = 32375, batch.n_tokens = 9, progress = 1.000000
slot update_slots: id  1 | task 12639 | prompt done, n_tokens = 32375, batch.n_tokens = 9
slot init_sampler: id  1 | task 12639 | init sampler, took 4.34 ms, tokens: text = 32375, total = 32375
slot print_timing: id  1 | task 12639 | 
prompt eval time =     176.24 ms /     9 tokens (   19.58 ms per token,    51.07 tokens per second)
       eval time =    1334.39 ms /    46 tokens (   29.01 ms per token,    34.47 tokens per second)
      total time =    1510.63 ms /    55 tokens

上下文32k时,提示词解码速度:51.07 tokens per second, 生成速度:34.47 tokens per second

相关推荐
whyfail7 天前
开发平替实录:ZCode + 火山 Coding Plan + GLM-5.3-Flash,打不过 GPT-5.6-Sol,但只差一口气的成本
gpt·codex·glm·智谱·zcode
阆遤9 天前
llama.cpp 0.4.0-dev 本地编译指南
ai·编译·llama.cpp
JiMoKuangXiangQu13 天前
在 AllWinner T507 上部署 Qwen2.5-0.5B 大语言模型
人工智能·llama.cpp·推理引擎
智码看视界16 天前
.NET 10 推理大模型TensorSharp 3.3.0 部署实测:纯.NET推理引擎反超llama.cpp 1.5倍,DFlash2提速62%
c#·.net·llama.cpp·.net 10·tensorsharp·本地大模型推理·开源推理引擎
梦想的颜色17 天前
2026 年 9 月国内外主流大模型横向硬核评测:价格、算力、业务场景全维度对比,GPT‑6 Astra横空出世
gpt·claude·glm·minimax·kimi·deepseek·大模型测评
云卷云舒___________18 天前
Kimi K3新检查点疑现Arena,Omen Alpha匿名突袭,OpenAI GPT-6 Astra免除加价,AI圈谍影重重 | 9月5日 AI日报
glm·kimi·智谱·ai日报·k3·arena·omenalpha
云卷云舒___________25 天前
谷歌内测Gemini 3.8 Flash,智谱开源GLM-5.3权重,腾讯亮剑Hy4,开源大模型神仙打架 | 8月29日 AI日报
谷歌·腾讯·glm·gemini·混元·智谱·ai日报
deepseek231 个月前
智谱开源 GLM-5.3-Flash 拆解:320B 只激活 18B 的混合注意力架构与十万张国产卡
人工智能·大模型·glm
zhangfeng11331 个月前
AMD Instinct MI50(gfx906)上为 Qwen 系列模型优化并可用的 vLLM 相关仓库、Docker 镜像与实践指南。
人工智能·docker·ai编程·qwen·算子开发·vllm·mi50