单卡5090跑125B 大模型:从安装、验证到基准测试的完整操作教程

单卡5090跑125B 大模型:从安装、验证到基准测试的完整操作教程

作者:吴佳浩

撰稿时间:2026-10-09

测试环境:

  • GPU:RTX 5090(32GB VRAM,驱动 617.14,PCIe Gen5 ×8)
  • CPU:AMD Ryzen 9 9950X
  • Memory:94GB RAM
  • Storage:NVMe SSD
  • OS:Windows 11(10.0.26200)
  • Engine:Strata v0.1.40.2(CUDA 13.0)
  • Model:Qwen3.8-Flash-Next IQ3_S(125B 级 MoE,24,576 experts)

引言

面向动手读者的操作手册。 不止给出命令,更说明「看到什么才算对、报错先查哪里」。 上篇讲清了它为什么能跑,这篇把同一套流程落到你的机器上:判断硬件 → 备环境 → 安装 → 下载除错 → 启动 → 冒烟验证 → 基准测试 → 读结果 → 调优。 命令与脚本都在本机真实跑过,脚本原文随文附上,输出为真实回填。 注:全文命令与脚本均在本机真实执行,未经验证的推测会明确标注为推测。因为阿拉好奇所以我们一道来白相相,看看照着做完是不是真能把它跑起来。


目录

  1. 开工前:先判断你的机器能不能跑
  2. 环境准备:三个会静默失败的坑
  3. 安装:选型、下载源与断点续传
  4. 下载环节的三个真实故障与修法
  5. 启动与就绪判定
  6. 冒烟验证:四个接口跑通才算装好
  7. 基准测试:方法学先于脚本
  8. 四类测试的实操
  9. 五个常见误读
  10. 调优清单与故障速查
  11. 附录:脚本清单与基线数字

一、开工前:先判断你的机器能不能跑

不要先下载 80 GB 再发现显存不够。引擎自带自检,它只装 Python 和 venv,不下载模型:

bash 复制代码
cd /g/strata
env -u PYTHONPATH -u PYTHONHOME ./.venv/Scripts/python.exe setup.py --check

这台机器的真实输出:

text 复制代码
Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)

=== Step 1: checking your PC ===
  [ok] GPU: NVIDIA GeForce RTX 5090, 31.8 GB VRAM, compute capability 12.0, driver 617.14
  [ok] RAM: 94 GB
  [ok] CPU: AMD Ryzen 9 9950X 16-Core Processor (AVX-512)
  [ok] PCIe: 5.0 x8
  [!]  the GPU's PCIe link reads 8 of its 16 lanes right now (some cards narrow it when idle); ...
  Q2_0     needs ~48 GB RAM: fits
  IQ2_XS   needs ~48 GB RAM: fits
  IQ3_XXS  needs ~60 GB RAM: fits
  IQ3_S    needs ~62 GB RAM: fits
  IQ1_M    needs ~32 GB RAM: fits
  UD-Q4_K_XL needs ~48 GB RAM: EXPERIMENTAL, fits with 70 GiB of its experts in RAM, ...
  UD-IQ4_XS needs ~48 GB RAM: fits with 55 GiB of its experts in RAM, ...

This PC can run Strata. Run it again without --check to install.

看三个地方,缺一不可:

看什么 门槛 为什么
GPU 与 compute capability 12 GB VRAM 起,sm_75 以上 引擎按算力分发内核;太老只能走实验性路径
RAM 32 GB 起,决定能选哪一档 专家绝大部分驻留在内存,不是显存
PCIe 宽度 越宽越好 专家要靠它搬运,见第十节的第 1 条

输出末尾的物件表是选型的唯一依据 :fits 一列表示这档在你机器上装得下。把这几行的 RAM 需求记下来,第三节直接用。

二、环境准备:三个会静默失败的坑

这一步不产生任何性能数字,但它是全套流程里最容易浪费一天的地方。三个坑的共同特征是不报错。

2.1 解释器可能指向坏文件

先确认默认解释器是好的:

bash 复制代码
py -3 --version          # 本机指向 F:\python,直接抛 Fatal Python error: init_fs_encoding
python --version         # 3.11.15,但这个可能是别的工具链的私有环境

判断标准不是「能不能打印版本」,而是能不能真的初始化 。本机的 py -3 连 encodings 模块都找不到,属于坏安装。用已知可用的独立解释器建 venv:

bash 复制代码
env -u PYTHONPATH -u PYTHONHOME /f/mambaforge/python -m venv "G:/strata/.venv"

2.2 环境变量会污染子进程

宿主工具链常把 PYTHONPATH 注入子进程。后果不是报错,而是venv 里"看得见"外面的包:

bash 复制代码
# 错误方式:继承了宿主 PYTHONPATH,输出里会混进宿主目录
./.venv/Scripts/python.exe -c "import sys; [print(p) for p in sys.path]"

# 正确方式:清空后只剩自己的 site-packages
env -u PYTHONPATH -u PYTHONHOME ./.venv/Scripts/python.exe \
  -c "import sys; [print(p) for p in sys.path if 'site' in p]"
# → G:\strata\.venv\Lib\site-packages

之后每一步都要带 env -u PYTHONPATH -u PYTHONHOME。 本机的 pip list 在污染状态下会列出上百个本不属于这个 venv 的包,据此判断依赖是否装好必然出错。

2.3 MSYS 风格路径会让 venv 静默不建

Windows 原生 Python 不认 /g/... 这种路径:

bash 复制代码
python -m venv /g/strata/.venv        # ← exit 0,但目录根本没建
python -m venv "G:/strata/.venv"      # ← 正确

exit 0 加空目录,是最贵的失败形态。建完立刻验:

bash 复制代码
ls ./.venv/Scripts/python.exe && ./.venv/Scripts/python.exe --version

2.4 把这一步固化成一个脚本

bash 复制代码
#!/usr/bin/env bash
cd /g/strata
unset PYTHONPATH PYTHONHOME
export HF_ENDPOINT=https://hf-mirror.com
export STRATA_SOURCE=modelscope
exec ./.venv/Scripts/python.exe setup.py --yes --family qwen --model IQ3_S --source modelscope --no-start

用法就是 bash 该脚本,它会清空环境变量、指定 ModelScope 为下载源,并执行无交互安装。脚本里那条 setup.py 调用在第三节解释。

三、安装:选型、下载源与断点续传

3.1 选型:用第一节的 RAM 需求表做映射

你的 RAM 选哪档 取舍
32 GB Coder(IQ1_M) 只留 256/512 个专家,代码强、通用弱
48 GB IQ2_XS 或 Q2_0 大档装不下
64 GB IQ2_XS 起,IQ3_XXS / IQ3_S 也可 IQ3_S 需要少开其他程序
96 GB 以上 IQ3_S,或 Unsloth UD-IQ4_XS 余量充足

本机 94 GB RAM 属于最高那一档,选 IQ3_S(3.5-bit,官方称与 BF16 全模型在公开基准上持平)。下载量约 83.6 GB。

别为了省时间换小档,算一下就明白: 分片 2(PLE 表,28.8 GB)是各档共享的同一个文件,换档只省分片 1 的差额,而 Q2_0 还要多做一次约 40 GB 的 AVX-512 重打包。净收益很小,质量却要降一档。

3.2 下载源:先测再选,不要凭感觉

安装时用 --source modelscope。这不是随便选的,是实测对比的结果:

源 单连接实测
mirrors.aliyun.com 0.45 MB/s
mirrors.tuna.tsinghua.edu.cn 1.55 MB/s
registry.npmmirror.com 1.75 MB/s
ModelScope CDN 3.29 MB/s(最快)

想用 HuggingFace 镜像时,设 HF_ENDPOINT=https://hf-mirror.com 即可,引擎会走镜像而跳过 HEAD 探测。

关于并行加速: 对 ModelScope CDN 做过 1 / 4 / 8 / 16 连接的并行测速,结果是 1.73 → 4.46 → 4.97 → 5.26 MiB/s ,而单连接速度从 1.73 掉到 0.22 MiB/s 。每连接变慢说明瓶颈是每 IP 封顶,不是服务端限速------并行下载不会更快,别在这上面花时间。

3.3 安装命令

bash 复制代码
cd /g/strata
env -u PYTHONPATH -u PYTHONHOME \
  HF_ENDPOINT=https://hf-mirror.com STRATA_SOURCE=modelscope \
  ./.venv/Scripts/python.exe setup.py --yes --family qwen --model IQ3_S \
     --source modelscope --no-start

--no-start 是给自动化用的:服务在前台运行直到窗口关闭,会占住 shell。装完再单独启动。

安装阶段会依次做四件事:装依赖 → 下载引擎(约 0.14 GB)→ 下载模型(83.6 GB)→ 打包。

text 复制代码
=== Step 3: Python packages ===
  [ok] numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil installed
=== Step 4: the Strata engine ===
  [ok] llama.cpp 3cf0325 (gguf-py, ggml, mtmd)
  [ok] Strata engine downloaded
  [ok] engine: G:\strata\engine\strata.exe
=== Step 6: preparing the model for Strata ===
  index.txt: 1079 tensors, 302 served natively, 0 in extra.bin, arena 1.43 GiB
  tokenizer/: vocab 248320, merges 247587, pre qwen35, ...
  [ok] model prepared: G:\Strata-data\packs\iq3_s
  [ok] MTP draft layer: G:\Strata-data\mtp\rt
=== Step 7: writing the start script ===
  [ok] KV streaming on: the context's KV cache lives in RAM (1.8 GB), more experts fit in VRAM
  [ok] start script: run-iq3_s.bat

这一步如果中断,直接重跑同一条命令 ------下载是断点续传的,已完成的文件会被 already downloaded 跳过。但下一节的坑正是从「跳过」开始的。

四、下载环节的三个真实故障与修法

这一节是全文最该细看的部分。三个问题都不是「参数写错了」,而是上游代码的隐含假设 与静默的数据损坏。

4.1 现象:19.3 GB 就「下载完成」了

第一次装到分片 2(应为 28.8 GB)时,脚本在 19.3 GB 处判定完成并退出,报错:

text 复制代码
Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf is short: 19,310,770,859 of 28,800,138,432 bytes

根因 :setup.py 用 HEAD 的 Content-Length 来确定目标大小,而 ModelScope 的 resolve 端点不返回它。

python 复制代码
# setup.py:1348
if not total or have >= total:      # total 来自 HEAD,此时是 0
    break                           # → 第一次连接断开处就 break
...
if total and part.stat().st_size != total:   # total 为 0,这条校验整条跳过
    fail(f"could not finish downloading ...")
part.replace(dst)                   # 半成品被改名成正式文件

更危险的是它留下的标记 :因为没有哈希可校验,代码走了 mark(dst) 分支,写下一个只含时间戳的 .done 文件。而 .done 的语义是「这一步已完成」,于是下一次运行会直接跳过下载,然后在同一处再次失败------形成死循环。

4.2 现象:长度对、哈希错的静默损坏

按 4.1 的修法续传补齐长度后,哈希校验仍然失败:

text 复制代码
got      8ae37d5797657c00e47cc61eefc88c5b574e7798807d7920a443115b46b2da39
expected 316b46f3a2dbd68c900f43136ab9449f9dcc3725dfd8c794847c204bc161e113

长度完全正确、内容错误。 这类损坏能骗过所有基于长度的检查。

根因:下载器只校验了状态码,没校验区间起点。

python 复制代码
# 有缺陷的版本
if have and r.status != 206:
    f.seek(0); f.truncate(); have = 0     # 只判断"服务端是否忽略了 Range"
f.write(b)                                 # 然后无条件追加

一旦 CDN 返回 206 但区间起点不是请求的偏移 (缓存代理错位是常见成因),错位的字节会被当成「文件的这一段」直接追加。修法是断言 Content-Range 的起点:

python 复制代码
cr = r.headers.get("Content-Range") or ""
cr_start = int(cr.split()[1].split("-")[0]) if cr.startswith("bytes ") else None
if have and (status != 206 or (cr_start is not None and cr_start != have)):
    # 服务端忽略了 Range,或返回区间起点不是我要的偏移 → 丢弃重来
    f.seek(0); f.truncate(); have = 0

4.3 一个反例:稀疏采样证明不了大文件的完整性

我试过更省事的做法------从 CDN 抽 6 个偏移各取 64 KB 与本地比对,结果 6/6 全部 MATCH,据此判断文件「可以续传」。这个结论是错的。

text 复制代码
64 KB 探针,打在每个 64 MiB 的网格上
  → 覆盖率 = 64 KB / 64 MiB = 0.1%
  → 几十 MB 的坏区,只要探针没落进去,就整段隐藏
  → 28.8 GB 的文件,6 个探针能证明什么?什么也证明不了

补充证据:整份文件按 64 MiB 网格扫描,前 200/430 个块一个不匹配都没有,而全量哈希明确判定它是坏的。

结论:大文件完整性只认全量哈希;任何采样率低于「可能坏区大小 / 文件大小」的探测,都只是心理安慰。 最终处置是删档重下。

4.4 修复工具:一个带区间断言的下载器

python 复制代码
#!/usr/bin/env python3
"""Robust resumable fetch for a Strata model shard with a KNOWN size and SHA-256.

Why this exists: setup.py's downloader learns the expected size from a HEAD Content-Length. ModelScope's
resolve endpoint returns none for this file, so `total` became 0 and the loop bailed out at the first
connection drop (19.3 of 28.8 GB), then marked the file done. Here the total and the hash come from
ModelScope's API (its own published metadata), so progress cannot be mistaken for completion.

Resumes from whatever bytes already exist, retries on connection drops, verifies SHA-256 at the end and
only then moves the file into place and writes the `<file>.done` marker setup looks for.
"""
from __future__ import annotations

import hashlib, sys, time, urllib.error, urllib.request
from pathlib import Path

BASE = "https://www.modelscope.cn/models/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/resolve/master/"
ITEMS = {
    "IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf": (
        28800138432, "316b46f3a2dbd68c900f43136ab9449f9dcc3725dfd8c794847c204bc161e113"),
}
MODELS = Path(r"G:\Strata-data\models\IQ3_S")
CHUNK = 8 << 20
UA = {"User-Agent": "strata-fetch"}


def sample_ok(part: Path, url_rel: str, size: int) -> bool:
    """Confirm the existing prefix is still a real prefix of the remote file (guards against interleaved writers)."""
    off = max(0, size - CHUNK) if size < 3 * CHUNK else int(size * 0.9)
    off = (off // 4096) * 4096
    local = b""
    with open(part, "rb") as fh:
        fh.seek(off)
        local = fh.read(CHUNK)
    try:
        req = urllib.request.Request(BASE + url_rel, headers={**UA, "Range": f"bytes={off}-{off+CHUNK-1}"})
        with urllib.request.urlopen(req, timeout=120) as r:
            rem = r.read(CHUNK)
    except Exception as e:  # noqa: BLE001
        print(f"  prefix check at {off:,} could not be done ({e}); continuing anyway", flush=True)
        return True
    ok = local == rem
    print(f"  prefix check at {off:,}: {'MATCH' if ok else 'MISMATCH'}", flush=True)
    return ok


def fetch(rel: str, total: int, sha: str) -> int:
    dest = MODELS / Path(rel).name
    part = dest.with_name(dest.name + ".part")
    if not part.exists() and dest.exists():
        if dest.stat().st_size < total:
            print(f"moving the {dest.stat().st_size:,}-byte partial file to {part.name}", flush=True)
            dest.replace(part)
        else:
            part = dest                      # already full length; just verify below

    if part.exists():
        have = part.stat().st_size
        print(f"resuming from {have:,} / {total:,} bytes ({have/total:.1%})", flush=True)
        if have > 0 and not sample_ok(part, rel, have):
            print("  existing prefix is NOT coherent -> restarting from 0", flush=True)
            part.unlink()
            have = 0
    else:
        have = 0

    stall, last_report, started = 0, 0.0, time.time()
    for attempt in range(1, 501):
        if have >= total:
            break
        try:
            req = urllib.request.Request(BASE + rel, headers={**UA, "Range": f"bytes={have}-"})
            with urllib.request.urlopen(req, timeout=90) as r, open(part, "ab" if have else "wb") as f:
                status = getattr(r, "status", 200)
                # Every appended byte must provably be the file's byte at that offset. A 200 (Range ignored)
                # or a Content-Range whose start is not our offset means the stream is misaligned: redo it.
                cr = r.headers.get("Content-Range") or ""
                cr_start = None
                if cr.startswith("bytes "):
                    try:
                        cr_start = int(cr.split()[1].split("-")[0])
                    except (IndexError, ValueError):
                        cr_start = None
                if have and (status != 206 or (cr_start is not None and cr_start != have)):
                    print(f"  misaligned stream (status={status}, Content-Range={cr!r} expected start "
                          f"{have:,}): restarting from 0", flush=True)
                    f.seek(0)
                    f.truncate()
                    have = 0
                if not cr_start and cr == "":
                    print(f"  note: no Content-Range header (status={status})", flush=True)
                before = have
                while True:
                    b = r.read(CHUNK)
                    if not b:
                        break
                    f.write(b)
                    have += len(b)
                    now = time.time()
                    if now - last_report > 3:
                        last_report = now
                        rate = (have - 0) / max(1e-6, now - started)
                        print(f"  {have/1e9:7.2f} / {total/1e9:.2f} GB ({have/total:5.1%})  "
                              f"avg {rate/1e6:5.2f} MB/s  attempt {attempt}", flush=True)
                if have == before:
                    stall += 1
                else:
                    stall = 0
        except Exception as e:  # noqa: BLE001
            print(f"  attempt {attempt}: {type(e).__name__}: {e}", flush=True)
            stall += 1
        if have < total:
            time.sleep(min(30, 2 + 2 * stall))
    print(f"downloaded {have:,} / {total:,} bytes", flush=True)
    if have != total:
        print("INCOMPLETE", flush=True)
        return 1

    print(f"verifying SHA-256 of {part.name} ...", flush=True)
    h = hashlib.sha256()
    with open(part, "rb") as fh:
        while True:
            b = fh.read(16 << 20)
            if not b:
                break
            h.update(b)
    got = h.hexdigest()
    if got != sha:
        print(f"SHA-256 MISMATCH: {got} != {sha}", flush=True)
        return 2
    if part != dest:
        dest.unlink(missing_ok=True)
        part.replace(dest)
    dest.with_name(dest.name + ".done").write_text(f"sha256 {sha}", encoding="utf-8")
    print(f"OK {dest.name} verified; .done marker written", flush=True)
    return 0


if __name__ == "__main__":
    rc = 0
    for rel, (total, sha) in ITEMS.items():
        rc |= fetch(rel, total, sha)
    sys.exit(rc)

这个脚本的四个关键设计,正好对应上面三个故障:

设计 解决什么
尺寸与哈希取自 API 公布值,不读 HEAD 4.1 的 total = 0
每次追加前断言 Content-Range 起点 == 当前偏移 4.2 的静默错位
断线自动重试(最多 500 次),按实际文件大小续传 长下载必然遇到的断流
全量 SHA-256 通过后才改名并写 .done 杜绝「先标记、后校验」留下的毒标记

用法 :把 ITEMS 里的仓库路径、字节数、哈希换成你要下的那一片即可;已存在的 .part 会被自动续传,已存在且完整的文件会被跳过。

五、启动与就绪判定

bash 复制代码
cd /g/strata
env -u PYTHONPATH -u PYTHONHOME ./.venv/Scripts/python.exe serve/server.py \
  --engine strata --config "G:/strata/strata-iq3_s.json" --port 8080

--config 指向安装时生成的配置,内容就是引擎参数(专家缓存、投机解码、上下文长度、KV 策略)。先看一遍它再启动,后面调优全靠改它。

冷启动很慢,日志会明确提示:

text 复制代码
loading the model (the first start takes a minute or two) ...
[strata] loading the experts into RAM (about 55 GB) and locking part of them for the GPU.
         YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES NOW - this is normal.
[strata] still starting (201 s) - please wait ...
[strata] experts loaded: 46.84 GiB at 0.24 GiB/s (228 s so far)
[strata] filling the GPU's expert cache (11784 experts, 22.36 GiB of VRAM) ...
ready: http://127.0.0.1:8080/v1

判定就绪的唯一标准是 /health 返回 loaded: true,不是「窗口里出现 ready」。 用轮询而不是死等:

bash 复制代码
for i in $(seq 1 30); do
  curl -s -m 5 http://127.0.0.1:8080/health && break
  sleep 10
done
json 复制代码
{"status": "ok", "max_context": 131072, "model": "qwen3.8-flash-next-iq3_s",
 "images": false, "api_key": false, "loaded": true, "service": "strata"}

重要:loaded: true 不等于可以对外服务。 头几次请求仍是冷的,实测第 1 次 29.9 tok/s,要到第 5 次才稳定在 64--68 tok/s。服务上线前先打几次丢弃式请求。

六、冒烟验证:四个接口跑通才算装好

装好先别跑基准,先用一个脚本把「接口活着、输出是真的、坏输入不会崩」三件事一次确认。

python 复制代码
#!/usr/bin/env python3
"""Strata acceptance check: health, models, one real completion with engine timings, and an SSE TTFT probe."""
from __future__ import annotations

import json, sys, time, urllib.request

URL = "http://127.0.0.1:8080"


def get(path, timeout=60):
    with urllib.request.urlopen(URL + path, timeout=timeout) as r:
        return json.load(r)


def post(path, body, timeout=900, stream=False):
    req = urllib.request.Request(URL + path, data=json.dumps(body).encode(),
                                headers={"Content-Type": "application/json"})
    return urllib.request.urlopen(req, timeout=timeout)


def main():
    out = {"ts": time.strftime("%Y-%m-%d %H:%M:%S")}
    out["health"] = get("/health")
    out["models"] = get("/v1/models")
    try:
        out["metrics"] = get("/metrics")
    except Exception as e:  # noqa: BLE001
        out["metrics"] = {"error": str(e)}
    print("health:", json.dumps(out["health"]))
    print("models:", json.dumps(out["models"])[:300])

    # --- one real, blocking completion ---
    body = {"model": "strata", "max_tokens": 200, "temperature": 0, "seed": 42,
            "reasoning_effort": "none", "cache_prompt": False,
            "messages": [{"role": "user", "content":
                          "Explain in three sentences why a mixture-of-experts model can run on a consumer GPU."}]}
    t0 = time.perf_counter()
    with post("/v1/chat/completions", body) as r:
        d = json.load(r)
    wall = time.perf_counter() - t0
    txt = d["choices"][0]["message"].get("content") or ""
    out["completion"] = {"wall_s": round(wall, 3), "http": 200, "text": txt,
                         "finish_reason": d["choices"][0].get("finish_reason"),
                         "usage": d.get("usage"), "timings": d.get("timings")}
    t = d.get("timings") or {}
    print(f"\n[blocking] wall={wall:.2f}s prompt_n={t.get('prompt_n')} "
          f"prefill={t.get('prompt_per_second')} tok/s | generated={t.get('predicted_n')} "
          f"decode={t.get('predicted_per_second')} tok/s")
    print("[answer]", txt[:400].replace("\n", " "))
    print("[usage]", d.get("usage"))

    # --- SSE: client TTFT on a fresh prompt ---
    sbody = dict(body, stream=True, max_tokens=128,
                 messages=[{"role": "user", "content": f"[ttft {time.time_ns()}] Name four noble gases."}])
    t0 = time.perf_counter()
    ttft, chunks = None, []
    with post("/v1/chat/completions", sbody, stream=True) as r:
        buf = b""
        while True:
            c = r.read(1)
            if not c:
                break
            buf += c
            if not buf.endswith(b"\n"):
                continue
            line = buf.decode("utf-8", "replace").strip()
            buf = b""
            if not line.startswith("data:"):
                continue
            p = line[5:].strip()
            if p == "[DONE]":
                break
            try:
                j = json.loads(p)
            except ValueError:
                continue
            if j.get("timings"):
                out["sse_timings"] = j["timings"]
            for ch in j.get("choices", []):
                piece = (ch.get("delta") or {}).get("content")
                if piece:
                    if ttft is None:
                        ttft = time.perf_counter() - t0
                    chunks.append(piece)
    out["sse"] = {"ttft_s": round(ttft, 4) if ttft else None,
                  "wall_s": round(time.perf_counter() - t0, 3),
                  "text": "".join(chunks)[:300]}
    print(f"\n[SSE] client TTFT={out['sse']['ttft_s']}s wall={out['sse']['wall_s']}s")
    print("[sse answer]", out["sse"]["text"][:200].replace("\n", " "))

    with open(r"G:\tmp\strata-bench\verify.json", "w", encoding="utf-8") as fh:
        json.dump(out, fh, indent=1, ensure_ascii=False)
    print("\nwrote G:\\tmp\\strata-bench\\verify.json")
    return 0


if __name__ == "__main__":
    sys.exit(main())
bash 复制代码
python verify.py

真实输出:

text 复制代码
health: {"status": "ok", "max_context": 131072, "model": "qwen3.8-flash-next-iq3_s", "images": false, "api_key": false, "loaded": true, "service": "strata"}
models: {"object": "list", "data": [{"id": "qwen3.8-flash-next-iq3_s", "object": "model", "status": {"value": "loaded"}, "meta": {"n_ctx": 131072}, ...}]}

[blocking] wall=3.86s prompt_n=31 prefill=56.7 tok/s | generated=92 decode=27.9 tok/s
[answer] A mixture-of-experts model can run on a consumer GPU because it employs sparse activation, ...

[SSE] client TTFT=1.1104s wall=2.715s
[sse answer] Four noble gases are:  1. Helium (He) 2. Neon (Ne) 3. Argon (Ar) 4. Krypton (Kr)

注意 27.9 tok/s 这个数。 它是冷启动后的第一个请求,比稳态低一半以上。把它当成「模型性能」是这类测试最常见的误读------但它作为「服务能被真实调起来」的证据是合格的。

另外两个必须验的点:

检查 期望 本机实测
空 messages 数组 HTTP 400 且给出错误类型 {"error":{"type":"invalid_request_error","message":"No messages provided."}}
/v1/models 的 n_ctx 与配置的 --max-context 一致 131072

七、基准测试:方法学先于脚本

数字只有在方法公开时才有意义。先固定五条,再动手写脚本。

必须固定的条件 原因
每次请求带唯一首行 + cache_prompt=false 否则引擎会复用前缀,测到的不是全量读取
预热不计入统计 冷启动会系统性拉低首批数字
吞吐取自引擎自报 timings 与客户端墙钟混用会产生两套口径
TTFT 在客户端独立测量 服务端自报的 prompt_ms 不含排队与网络
思考模式固定(reasoning_effort) 思考 token 长度不同会污染跨档比较

验证方法有效性的标志:每次请求的 cache_n 应恒为 0(表示没有吃到前缀缓存)。如果你看到 cache_n 不为 0,说明测量无效,先修方法再看数字。

八、四类测试的实操

8.1 预填 / Decode / 首 Token 扫描

python 复制代码
#!/usr/bin/env python3
"""Strata benchmark harness: prefill (TTFT proxy + tok/s), decode tok/s, streaming client TTFT.

Method (follows docs/COMMUNITY_BENCHMARKS.md):
  * one warm-up + N measured runs per prompt size, model loading not timed
  * every request starts with a unique first line and cache_prompt=false, so each run reads the
    whole prompt fresh (engine reports cache_n ~= 0)
  * throughput is the ENGINE's own timings (prompt_per_second / predicted_per_second)
  * client TTFT is measured independently over the streaming SSE path
  * thinking is pinned per run via reasoning_effort so output composition is controlled
"""
from __future__ import annotations

import argparse, hashlib, json, statistics, sys, time, urllib.error, urllib.request
from pathlib import Path

ROOT = Path("G:/strata")
SOURCES = [("docs", "*.md"), ("src", "*.cpp"), ("src", "*.cu"), ("include", "*.hpp"),
           ("serve", "*.py"), ("tools", "*.py")]
CHARS_PER_TOKEN = 3.2


def haystack(n_chars: int) -> str:
    files = []
    for d, pat in SOURCES:
        files += sorted((ROOT / d).rglob(pat))
    parts, total = [], 0
    while total < n_chars:
        for f in files:
            try:
                t = f.read_text(encoding="utf-8", errors="replace")
            except OSError:
                continue
            parts.append(f"\n\n=== {f.relative_to(ROOT).as_posix()} ===\n{t}")
            total += len(parts[-1])
            if total >= n_chars:
                break
        if total == 0:
            raise SystemExit("no source text found to build prompts from")
    return "".join(parts)[:n_chars]


def post(url: str, body: dict, key: str, timeout: float, stream: bool = False):
    headers = {"Content-Type": "application/json"}
    if key:
        headers["Authorization"] = "Bearer " + key
    if stream:
        headers["Accept"] = "text/event-stream"
    req = urllib.request.Request(url.rstrip("/") + "/v1/chat/completions",
                                 data=json.dumps(body).encode(), headers=headers)
    return urllib.request.urlopen(req, timeout=timeout)


def run_stream(url, prompt, max_tokens, key, effort, timeout):
    """Streaming: returns client TTFT (s), wall (s), aggregated text, and the final payload."""
    body = {"model": "strata", "messages": [{"role": "user", "content": prompt}],
            "temperature": 0, "seed": 42, "max_tokens": max_tokens, "cache_prompt": False,
            "stream": True, "reasoning_effort": effort,
            "stream_options": {"include_usage": True}}
    t0 = time.perf_counter()
    ttft = None
    text, final = [], {}
    with post(url, body, key, timeout, stream=True) as r:
        buf = b""
        while True:
            chunk = r.read(1)
            if not chunk:
                break
            buf += chunk
            if not buf.endswith(b"\n"):
                continue
            line = buf.decode("utf-8", "replace").strip()
            buf = b""
            if not line.startswith("data:"):
                continue
            payload = line[5:].strip()
            if payload == "[DONE]":
                break
            try:
                d = json.loads(payload)
            except ValueError:
                continue
            if d.get("timings") or d.get("usage"):
                final.update({k: v for k, v in d.items() if k != "choices"})
            for ch in d.get("choices", []):
                piece = (ch.get("delta") or {}).get("content") or ""
                if piece:
                    if ttft is None:
                        ttft = time.perf_counter() - t0
                    text.append(piece)
    wall = time.perf_counter() - t0
    out = "".join(text)
    return {"ttft_s": round(ttft, 4) if ttft else None, "wall_s": round(wall, 3),
            "chars": len(out), "sha256": hashlib.sha256(out.encode()).hexdigest()[:16],
            "timings": final.get("timings"), "usage": final.get("usage")}


def run_blocking(url, prompt, max_tokens, key, effort, timeout):
    body = {"model": "strata", "messages": [{"role": "user", "content": prompt}],
            "temperature": 0, "seed": 42, "max_tokens": max_tokens, "cache_prompt": False,
            "reasoning_effort": effort}
    t0 = time.perf_counter()
    with post(url, body, key, timeout) as r:
        d = json.load(r)
    wall = time.perf_counter() - t0
    text = d["choices"][0]["message"].get("content") or ""
    return {"wall_s": round(wall, 3), "timings": d.get("timings"), "usage": d.get("usage"),
            "finish_reason": d["choices"][0].get("finish_reason"),
            "sha256": hashlib.sha256(text.encode()).hexdigest()[:16], "chars": len(text)}


def summ(runs, key, sub=None):
    v = []
    for r in runs:
        t = r.get("timings") or {}
        x = t.get(key) if sub is None else (t.get(sub) or {}).get(key)
        if x is not None:
            v.append(x)
    if not v:
        return None
    return {"median": round(statistics.median(v), 3), "min": min(v), "max": max(v), "n": len(v)}


def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--url", default="http://127.0.0.1:8080")
    ap.add_argument("--api-key", default="")
    ap.add_argument("--label", required=True)
    ap.add_argument("--sizes", default="1k,4k,16k,32k,64k,128k")
    ap.add_argument("--runs", type=int, default=3)
    ap.add_argument("--warmup", type=int, default=1)
    ap.add_argument("--max-tokens", type=int, default=256)
    ap.add_argument("--effort", default="none", choices=["none", "low", "medium", "high"])
    ap.add_argument("--timeout", type=float, default=3600)
    ap.add_argument("--stream-ttft", action="store_true")
    ap.add_argument("--out", required=True)
    a = ap.parse_args()

    props = {}
    for ep in ("/props", "/health", "/v1/models"):
        try:
            with urllib.request.urlopen(a.url.rstrip("/") + ep, timeout=30) as r:
                props[ep] = json.load(r)
        except Exception as e:  # noqa: BLE001
            props[ep] = {"error": str(e)}

    res = {"label": a.label, "endpoint": a.url, "sizes": {}, "params": vars(a),
           "engine": props, "started": time.strftime("%Y-%m-%d %H:%M:%S")}
    for size in a.sizes.split(","):
        size = size.strip()
        toks = int(float(size.lower().rstrip("k")) * 1024) if size.lower().endswith("k") else int(size)
        base = haystack(int(toks * CHARS_PER_TOKEN))
        entry = {"target_tokens": toks, "warmup": [], "runs": []}
        tag = lambda kind, i: f"[{a.label} {size} {kind} {i} {time.time_ns()}]\n"  # noqa: E731
        for i in range(a.warmup):
            try:
                entry["warmup"].append(run_blocking(a.url, tag("warmup", i) + base,
                                                    a.max_tokens, a.api_key, a.effort, a.timeout))
            except Exception as e:  # noqa: BLE001
                entry["warmup"].append({"error": str(e)})
        for i in range(a.runs):
            r = {"run": i + 1}
            try:
                r.update(run_blocking(a.url, tag("run", i + 1) + base,
                                      a.max_tokens, a.api_key, a.effort, a.timeout))
            except Exception as e:  # noqa: BLE001
                r["error"] = str(e)
            if a.stream_ttft and "error" not in r:
                try:
                    r["stream"] = run_stream(a.url, tag("srun", i + 1) + base,
                                             a.max_tokens, a.api_key, a.effort, a.timeout)
                except Exception as e:  # noqa: BLE001
                    r["stream_error"] = str(e)
            entry["runs"].append(r)
            t = r.get("timings") or {}
            st = (r.get("stream") or {}).get("ttft_s")
            print(f"{a.label} {size} run {i+1}: prompt_n={t.get('prompt_n')} cache_n={t.get('cache_n')} "
                  f"prefill={t.get('prompt_per_second')} tok/s decode={t.get('predicted_per_second')} tok/s "
                  f"generated={t.get('predicted_n')} wall={r.get('wall_s')}s ttft={st}s"
                  + (f" ERROR={r['error']}" if "error" in r else ""), flush=True)
        entry["prefill_tps"] = summ(entry["runs"], "prompt_per_second")
        entry["decode_tps"] = summ(entry["runs"], "predicted_per_second")
        entry["prompt_ms"] = summ(entry["runs"], "prompt_ms")
        sess = [r["stream"]["ttft_s"] for r in entry["runs"]
                if isinstance(r.get("stream"), dict) and r["stream"].get("ttft_s")]
        entry["ttft_s"] = ({"median": statistics.median(sess), "min": min(sess), "max": max(sess),
                            "n": len(sess)} if sess else None)
        res["sizes"][size] = entry
    res["finished"] = time.strftime("%Y-%m-%d %H:%M:%S")
    Path(a.out).write_text(json.dumps(res, indent=2), encoding="utf-8")
    print("written", a.out)
    return 0


if __name__ == "__main__":
    sys.exit(main())
bash 复制代码
python bench.py --url http://127.0.0.1:8080 --label iq3_s-131k \
  --sizes 1k,4k,16k,32k,64k,128k --runs 3 --warmup 1 \
  --max-tokens 256 --effort none --stream-ttft \
  --out bench-iq3s.json

真实输出(节选):

text 复制代码
iq3_s-131k 1k run 1: prompt_n=1035 cache_n=0 prefill=611.3 tok/s decode=53.4 tok/s generated=130 wall=4.14s ttft=1.9411s
iq3_s-131k 16k run 2: prompt_n=17662 cache_n=0 prefill=1214.7 tok/s decode=101.4 tok/s generated=224 wall=16.817s ttft=14.953s
iq3_s-131k 128k run 1: prompt_n=128846 cache_n=0 prefill=1032.6 tok/s decode=65.4 tok/s generated=143 wall=127.379s ttft=144.1314s

本机六档结果(3 次取中位):

提示长度 Prefill tok/s Decode 中位 tok/s 客户端 TTFT
1K 593.1 55.4 1.92 s
4K 685.8 37.8 5.98 s
16K 1,214.7 91.9 15.58 s
32K 1,232.7 70.2 28.54 s
64K 1,065.6 87.3 62.57 s
128K 1,032.6 65.4 131.41 s

8.2 长上下文召回

不要自己写 needle 测试,仓库自带一个:

bash 复制代码
cd /g/strata
env -u PYTHONPATH -u PYTHONHOME ./.venv/Scripts/python.exe tools/needle_bench.py \
  --lengths 8k,32k,128k --depths 10,50,90 \
  --url http://127.0.0.1:8080 --out needle.json
text 复制代码
   8k depth  10%: FOUND  (8,064 prompt tokens, 10 s)
  32k depth  50%: FOUND  (32,894 prompt tokens, 41 s)
 128k depth  10%: FOUND  (125,833 prompt tokens, 148 s)
 128k depth  90%: FOUND  (125,834 prompt tokens, 162 s)
9 of 9 found

只测最大长度是不够的,必须测三个深度。理由:

text 复制代码
只测最大长度
  → 只能证明「塞得进去」,不能证明「读得出来」

只测开头
  → 无法暴露 lost in the middle

三个深度全测
  → 10% / 50% / 90% 全部命中 = 中间位置没有系统性衰减

8.3 稳定性与并发

python 复制代码
#!/usr/bin/env python3
"""Strata stability / soak test.

Checks, against a live server:
  1. N sequential requests on a fixed prompt: error count, decode tok/s distribution (jitter),
     output determinism (temperature 0 -> same sha256) and finish reasons.
  2. A concurrency sweep (1,2,4 clients in parallel) with per-client latency and aggregate tok/s,
     to find the throughput/latency knee and confirm the server's one-request-at-a-time policy.
  3. Repeated long-prompt requests: detects degradation / expert-cache thrash across repeats.
  4. Health/robustness probes: malformed body, oversized max_tokens, /health + /metrics liveness.

Writes a JSON report.
"""
from __future__ import annotations

import argparse, concurrent.futures as cf, hashlib, json, statistics, sys, time, urllib.error, urllib.request

PROMPT = ("Write a Python function that merges two sorted linked lists. "
          "Then list three edge cases, each with one line of explanation.")


def post(url, body, key, timeout):
    headers = {"Content-Type": "application/json"}
    if key:
        headers["Authorization"] = "Bearer " + key
    req = urllib.request.Request(url.rstrip("/") + "/v1/chat/completions",
                                 data=json.dumps(body).encode(), headers=headers)
    t0 = time.perf_counter()
    with urllib.request.urlopen(req, timeout=timeout) as r:
        d = json.load(r)
    return time.perf_counter() - t0, d


def one(url, key, effort, max_tokens, timeout, prompt=PROMPT, seed=42):
    body = {"model": "strata", "messages": [{"role": "user", "content": prompt}],
            "temperature": 0, "seed": seed, "max_tokens": max_tokens,
            "cache_prompt": False, "reasoning_effort": effort}
    wall, d = post(url, body, key, timeout)
    t = d.get("timings") or {}
    txt = d["choices"][0]["message"].get("content") or ""
    return {"ok": True, "wall_s": round(wall, 3), "decode_tps": t.get("predicted_per_second"),
            "prefill_tps": t.get("prompt_per_second"), "prompt_n": t.get("prompt_n"),
            "generated": t.get("predicted_n"), "cache_n": t.get("cache_n"),
            "sha256": hashlib.sha256(txt.encode()).hexdigest()[:16],
            "finish_reason": d["choices"][0].get("finish_reason"), "chars": len(txt)}


def probe(url, key, path, body=None):
    try:
        req = urllib.request.Request(url.rstrip("/") + path,
                                     data=(json.dumps(body).encode() if body is not None else None),
                                     headers={"Content-Type": "application/json",
                                              **({"Authorization": "Bearer " + key} if key else {})})
        t0 = time.perf_counter()
        with urllib.request.urlopen(req, timeout=60) as r:
            raw = r.read().decode("utf-8", "replace")
            return {"path": path, "status": r.status, "ms": round((time.perf_counter() - t0) * 1000, 1),
                    "body": raw[:400]}
    except urllib.error.HTTPError as e:
        return {"path": path, "status": e.code, "ms": None, "body": e.read().decode("utf-8", "replace")[:400]}
    except Exception as e:  # noqa: BLE001
        return {"path": path, "status": None, "error": str(e)}


def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--url", default="http://127.0.0.1:8080")
    ap.add_argument("--api-key", default="")
    ap.add_argument("--label", required=True)
    ap.add_argument("--seq", type=int, default=12)
    ap.add_argument("--concurrency", default="1,2,4")
    ap.add_argument("--per-client", type=int, default=3)
    ap.add_argument("--effort", default="none")
    ap.add_argument("--max-tokens", type=int, default=256)
    ap.add_argument("--timeout", type=float, default=1800)
    ap.add_argument("--out", required=True)
    a = ap.parse_args()

    rep = {"label": a.label, "url": a.url, "params": vars(a), "started": time.strftime("%Y-%m-%d %H:%M:%S"),
           "probes": [probe(a.url, a.api_key, "/health"), probe(a.url, a.api_key, "/metrics"),
                      probe(a.url, a.api_key, "/v1/models"),
                      probe(a.url, a.api_key, "/v1/chat/completions", {"messages": []}),
                      probe(a.url, a.api_key, "/v1/chat/completions",
                            {"model": "strata", "messages": [{"role": "user", "content": "hi"}],
                             "max_tokens": 5, "reasoning_effort": "none"})]}

    # 1. sequential
    seq, errs = [], 0
    for i in range(a.seq):
        try:
            r = one(a.url, a.api_key, a.effort, a.max_tokens, a.timeout, seed=100 + i)
            seq.append(r)
            print(f"seq {i+1}/{a.seq}: decode={r['decode_tps']} tok/s wall={r['wall_s']}s "
                  f"gen={r['generated']} finish={r['finish_reason']} sha={r['sha256']}", flush=True)
        except Exception as e:  # noqa: BLE001
            errs += 1
            seq.append({"ok": False, "error": str(e)})
            print(f"seq {i+1}/{a.seq}: ERROR {e}", flush=True)
    d = [r["decode_tps"] for r in seq if r.get("ok") and r.get("decode_tps")]
    shas = {r["sha256"] for r in seq if r.get("ok")}
    rep["sequential"] = {
        "n": a.seq, "errors": errs, "ok": a.seq - errs,
        "decode_tps": ({"median": round(statistics.median(d), 3), "min": min(d), "max": max(d),
                        "stdev": round(statistics.pstdev(d), 3),
                        "cv_pct": round(100 * statistics.pstdev(d) / statistics.fmean(d), 2)} if d else None),
        "distinct_outputs": len(shas), "deterministic": len(shas) == 1 and a.effort == "none",
        "runs": seq}

    # 2. concurrency sweep
    sweep = {}
    for c in (int(x) for x in a.concurrency.split(",")):
        def worker(wid):
            out = []
            for j in range(a.per_client):
                try:
                    out.append(one(a.url, a.api_key, a.effort, a.max_tokens, a.timeout,
                                   seed=1000 + wid * 100 + j))
                except Exception as e:  # noqa: BLE001
                    out.append({"ok": False, "error": str(e)})
            return out
        t0 = time.perf_counter()
        with cf.ThreadPoolExecutor(max_workers=c) as ex:
            futs = [ex.submit(worker, w) for w in range(c)]
            allr = [r for f_ in futs for r in f_.result()]
        span = time.perf_counter() - t0
        ok = [r for r in allr if r.get("ok")]
        toks = sum(r.get("generated") or 0 for r in ok)
        lat = sorted(r["wall_s"] for r in ok)
        sweep[str(c)] = {
            "workers": c, "requests": len(allr), "errors": len(allr) - len(ok),
            "tokens": toks, "span_s": round(span, 2),
            "aggregate_tok_s": round(toks / span, 2) if span else None,
            "latency_s": ({"median": round(statistics.median(lat), 2), "min": lat[0],
                           "p90": lat[int(0.9 * (len(lat) - 1))], "max": lat[-1]} if lat else None),
            "per_request_decode_tps": ({"median": round(statistics.median(
                [r["decode_tps"] for r in ok if r.get("decode_tps")]), 2)} if ok else None)}
        print(f"concurrency {c}: aggregate={sweep[str(c)]['aggregate_tok_s']} tok/s "
              f"median_latency={sweep[str(c)]['latency_s']} errors={sweep[str(c)]['errors']}", flush=True)
    rep["concurrency"] = sweep
    rep["finished"] = time.strftime("%Y-%m-%d %H:%M:%S")
    from pathlib import Path
    Path(a.out).write_text(json.dumps(rep, indent=2), encoding="utf-8")
    print("written", a.out)
    return 0


if __name__ == "__main__":
    sys.exit(main())
bash 复制代码
python stability.py --url http://127.0.0.1:8080 --label iq3_s-131k \
  --seq 12 --concurrency 1,2,4 --per-client 3 \
  --max-tokens 256 --effort none --out stability-iq3s.json
text 复制代码
seq 1/12: decode=29.9 tok/s wall=9.519s gen=256 finish=length sha=3d068a586493adb3
seq 5/12: decode=64.3 tok/s wall=4.113s gen=256 finish=length sha=0a2bcb27d08ab6c8
seq 12/12: decode=64.4 tok/s wall=4.072s gen=256 finish=length sha=3d068a586493adb3
concurrency 1: aggregate=64.04 tok/s errors=0
concurrency 2: aggregate=63.72 tok/s errors=0
concurrency 4: aggregate=60.47 tok/s errors=0

8.4 资源占用采样

python 复制代码
#!/usr/bin/env python3
"""Sample GPU (nvidia-smi) and system RAM while a benchmark runs; write CSV + a summary JSON.

The summary is rewritten every few samples, so a hard kill still leaves a usable file.
Stop with --duration N, SIGINT/SIGTERM, or by killing it.

Usage: python monitor.py --out gpu.csv --summary gpu-summary.json --interval 0.5 --duration 600
"""
from __future__ import annotations

import argparse, csv, json, statistics, subprocess, sys, time
from pathlib import Path

Q = ("index,name,utilization.gpu,utilization.memory,memory.used,memory.total,power.draw,"
     "temperature.gpu,clocks.sm,clocks.mem")
KEYS = ["t_s", "gpu", "util_gpu", "util_mem", "vram_used_mib", "vram_total_mib", "power_w",
        "temp_c", "sm_mhz", "mem_mhz", "ram_used_gib", "ram_total_gib", "ram_pct"]


def sample():
    try:
        out = subprocess.run(["nvidia-smi", f"--query-gpu={Q}", "--format=csv,noheader,nounits"],
                             capture_output=True, text=True, timeout=10)
        gpus = [dict(zip(Q.split(","), [c.strip() for c in ln.split(",")])) for ln in
                out.stdout.strip().splitlines() if ln.strip()]
    except Exception:  # noqa: BLE001
        gpus = []
    ram = {}
    try:
        import psutil
        vm = psutil.virtual_memory()
        ram = {"ram_used_gib": round((vm.total - vm.available) / 2**30, 2),
               "ram_total_gib": round(vm.total / 2**30, 2), "ram_pct": vm.percent}
    except Exception:  # noqa: BLE001
        pass
    return gpus, ram


def f(x):
    try:
        return float(x)
    except (TypeError, ValueError):
        return None


def write_summary(path, rows, n, t0, interval):
    def st(k):
        v = [r[k] for r in rows if r.get(k) is not None]
        if not v:
            return None
        return {"min": round(min(v), 2), "median": round(statistics.median(v), 2),
                "max": round(max(v), 2), "mean": round(statistics.fmean(v), 2), "n": len(v)}
    summary = {"samples": n, "rows": len(rows), "duration_s": round(time.time() - t0, 1),
               "interval_s": interval, "complete": False,
               "vram_used_mib": st("vram_used_mib"), "vram_total_mib": st("vram_total_mib"),
               "util_gpu_pct": st("util_gpu"), "util_mem_pct": st("util_mem"),
               "power_w": st("power_w"), "temp_c": st("temp_c"),
               "ram_used_gib": st("ram_used_gib"), "ram_pct": st("ram_pct"), "sm_mhz": st("sm_mhz")}
    Path(path).write_text(json.dumps(summary, indent=2), encoding="utf-8")
    return summary


def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--out", required=True)
    ap.add_argument("--summary", required=True)
    ap.add_argument("--interval", type=float, default=0.5)
    ap.add_argument("--duration", type=float, default=0.0,
                    help="stop by itself after this many seconds (0 = until a signal)")
    a = ap.parse_args()
    rows, t0 = [], time.time()
    stop = {"v": False}
    for sig in ("SIGINT", "SIGTERM"):
        try:
            import signal
            signal.signal(getattr(signal, sig), lambda *_: stop.update(v=True))
        except Exception:  # noqa: BLE001, PERF203
            pass
    with open(a.out, "w", newline="", encoding="utf-8") as fh:
        w = csv.writer(fh)
        w.writerow(KEYS)
        n = 0
        while not stop["v"]:
            gpus, ram = sample()
            for g in gpus:
                row = {"t_s": round(time.time() - t0, 2), "gpu": g["index"],
                       "util_gpu": f(g["utilization.gpu"]), "util_mem": f(g["utilization.memory"]),
                       "vram_used_mib": f(g["memory.used"]), "vram_total_mib": f(g["memory.total"]),
                       "power_w": f(g["power.draw"]), "temp_c": f(g["temperature.gpu"]),
                       "sm_mhz": f(g["clocks.sm"]), "mem_mhz": f(g["clocks.mem"]),
                       "ram_used_gib": ram.get("ram_used_gib"), "ram_total_gib": ram.get("ram_total_gib"),
                       "ram_pct": ram.get("ram_pct")}
                rows.append(row)
                w.writerow([row[k] for k in KEYS])
            n += 1
            fh.flush()
            if n % 10 == 0:
                write_summary(a.summary, rows, n, t0, a.interval)
            if a.duration and (time.time() - t0) >= a.duration:
                break
            time.sleep(a.interval)
    summary = write_summary(a.summary, rows, n, t0, a.interval)
    if summary is not None:
        summary["complete"] = True
        Path(a.summary).write_text(json.dumps(summary, indent=2), encoding="utf-8")
        print(json.dumps(summary, indent=2))
    return 0


if __name__ == "__main__":
    sys.exit(main())

用法是让它与基准同时跑:

bash 复制代码
python monitor.py --out gpu.csv --summary gpu-summary.json --interval 1.0 --duration 2400 &
python bench.py --url http://127.0.0.1:8080 --label iq3_s-131k --sizes 1k,4k,16k,32k,64k,128k \
  --runs 3 --warmup 1 --max-tokens 256 --effort none --stream-ttft --out bench-iq3s.json
text 复制代码
{ "samples": 2280, "duration_s": 2400.7 }
指标 min 中位 max
VRAM 占用 (MiB) 31,556 31,747 31,938
GPU 利用率 (%) 1 99 100
功耗 (W) 75.1 192.6 265.1
温度 (°C) 39 49 55
系统 RAM (GiB) 75.8 78.5 81.2

九、五个常见误读

误读 事实 正确读法
「decode 随上下文变长而下降」 不单调:4K 的 37.8 低于 16K 的 91.9 投机接受率随生成内容波动,跨档只能作参考,看同档 min--max
「首批请求快 = 性能好」 冷启动首请求比稳态低一半以上 按 run 顺序看趋势,忽略前 4--5 次
「并发能提高吞吐」 64.04 → 63.72 → 60.47,只拉长延迟 引擎默认一次一个请求,多客户端只是排队
「功耗温度高说明跑满了」 功耗仅 192.6 W / 600 W,49 °C 算力不是瓶颈,去看搬运路径
「哈希对不上是小概率」 长度对、内容错是真实发生过的事故 全量哈希是唯一判据,别用采样

再补一条只有翻原始 run 才看得见的:

长档位的 finish_reason 是 tool_calls,短档位是 stop。 32K / 64K / 128K 全部以工具调用收尾,说明长档统计的「生成 token 数」里包含工具调用的 JSON 结构,与其他档位不是同一种输出形态------跨档比较时必须知道这一点。

十、调优清单与故障速查

10.1 按预期收益排序的调优点

# 做什么 依据
1 PCIe 从 ×8 修到 ×16 引擎自检与 nvidia-smi 双向确认 width.current=8 / max=16;专家缓存命中率 91--95%,2.4--5.3% 要走 PCIe
2 加 --vram-reserve-mib 1116 引擎自身告警:全部加载后只剩 96 MiB,低于建议的 700 MiB 保留
3 试 --prefill auto:32768 同型号参考实测在 14.7K--28.9K 提示上比 auto 快 14--25%
4 启动后先预热再对外服务 前 4--5 次请求慢 30--55%
5 多客户端时配 --conversation-cache-mib 8192 不加会重读约 90% 的提示
6 长提示为主时放大 --kv-resident,或改 --kv q4_0 64K 后预填回落的成因是 32K 以上 KV 走 RAM
7 单用户保持 parallel=1 实测并发不涨吞吐、只涨延迟
8 不要做功耗墙或超频调优 192.6 W / 600 W、49 °C,离限值很远
9 可选:跑一次 --calibrate 会实测几个引擎参数并保留最快组合

10.2 故障速查

现象 先查这里
is short: N of M bytes 4.1 的 Content-Length 问题;删掉对应的 .done 再重下
哈希对不上 4.2 的区间错位;用 4.4 的下载器重下,别用采样判断
重跑一直跳过下载但一直失败 有毒的 .done 标记;删掉它
pip list 结果明显不对 PYTHONPATH 污染;加 env -u PYTHONPATH -u PYTHONHOME
python -m venv 后目录是空的 路径写法;改用 G:/... 原生路径
装了但一跑就报内存不足 档位选大了;回第一节的 RAM 需求表
/health 迟迟不返回 冷启动 3 分钟正常;用轮询,不要杀进程
端口 8080 被占 已在运行,先 curl /health 确认

十一、附录:脚本清单与基线数字

本机基线数字(供对照)

项 数值
冷启动到 /health 228 s 装载 + 缓存填充,约 3 分钟
Prefill 峰值 1,232.7 tok/s(32K)
Decode 稳态 64--68 tok/s
TTFT 1.92 s(1K)/ 131.41 s(128K)
长上下文召回 9/9(含 128K 三深度)
VRAM 占用 31,747 / 32,607 MiB(97.4%)
功耗 中位 192.6 W(上限 600 W)
磁盘占用 代码 850 MB + 模型数据 86 GB

口径声明:以上均为本机(RTX 5090 32GB / Ryzen 9 9950X / 94 GB RAM / NVMe)实测,随机器浮动。凡本文未实际运行的步骤,均已明确标注为需自行验证;未做过多卡与 vision 的测试。


邦友们到这里,Strata 从环境准备、安装部署、启动验证,到基准测试与故障排查,就全部走完了。

如果你一路跟着操作下来,得到的不只是一个能够运行的 125B 级 MoE 模型,更重要的是得到了一套可以重复执行、可以重复验证、可以长期复用的测试流程。以后无论是升级 Strata、切换模型,还是更换硬件,都可以沿着同样的方法重新验证,而不是依赖别人截图里的几个数字。

上一篇,我们讨论的是 Strata 为什么能让服务器级 MoE 跑进一台普通 PC;这一篇,我们把整个过程真正落到了机器上。

接下来,我也会继续关注 Strata 后续版本,以及更多本地大模型部署框架和推理引擎的发展,继续把真实测试结果、踩过的坑和工程经验整理出来。

如果本文对你有所帮助,欢迎点赞、收藏、转发,也欢迎关注 「全栈架构师笔记」。下一次,我们继续一起把这些新东西跑明白、测明白,也讲明白。

相关推荐
数智工坊1 小时前
视觉SLAM第13讲|工程落地:双目视觉里程计系统架构设计与性能优化
人工智能·深度学习·线性代数·性能优化·矩阵·系统架构·机器人
newsxun1 小时前
校园餐全链条管理有了“实操手册” ——《学校食品安全与营养健康管理操作指南》在成都发布
大数据·人工智能
空堂与归1 小时前
GPT-6 智能界面上线,对话式 Agent 真没戏了?
人工智能·gpt·ai
码流子1 小时前
05-AI数字人
人工智能
moonsims1 小时前
AiBrainBox “高校算法创新”与“真实无人系统”的国产化智能端侧科研平台-高校 × 企业:构建面向低空与具身智能的产学研协同创新平台
人工智能
深蓝AI1 小时前
Redis 文档 MCP 揭示 Agent 检索真相:3 个工具、双读路径,引用为什么不能交给模型编
人工智能·redis
FPGA信号处理3 小时前
【模式识别】第三节课:分类误差的来源与线性分类器
人工智能·分类·数据挖掘
hrrrrxeeeee3 小时前
电商从业者AI技能提升:证书选择与商品、客服、运营场景落地
人工智能
小和尚同志8 小时前
1.8k star 的开源 token 使用量监控神器— TokenTracker
人工智能·ai编程