原文:New in llama.cpp: Decision Models
发布日期:2026 年 10 月 2 日
作者:Xuan-Son Nguyen (ngxson)、Victor Mustar、ggml-org
llama.cpp 服务器现在通过 /v1/systemone 端点支持决策模型。你发送一个状态(文本、JSON、截图)和类型化问题,模型在单次前向传播中返回每个选项的概率。
API 遵循 TypeSafe 的 Jev 模型引入的 System One 格式,因此现有客户端只需更换 base URL 即可。实现细节见 PR #29818。
什么是决策模型? 决策模型通过为你给定的选项打分来回答问题,而不是生成文本。聊天模型每个输出 token 需要一次前向传播,且输出仍需解析。决策模型只读取一次输入,其答案始终是你给出的选项之一,并附带概率。典型用途包括:路由请求、内容审核、检查 Agent 步骤是否成功、或选择下一步动作。
支持的模型
| 模型 | 大小 | 基于 | 语言 | 图片 | 许可证 | 速度* |
|---|---|---|---|---|---|---|
| Julia-1 | 144M | mmBERT-small | 50+ | 否 | Apache 2.0 | 3 ms |
| Laya | 421M | ModernBERT-large | 英语 | 否 | Apache 2.0 | 5 ms |
| Kev-4B | 4B | Qwen3.5-4B-Base | 英语 | 否 | Apache 2.0 | 12 ms |
| lev | 4B | Qwen3.5-4B | 英语 | 否 | Apache 2.0 | 36 ms |
| OpenJev | 27B | Qwen3.8-27B | en, de, fr, hi, zh, ja | 是 | CC BY-NC 4.0 | 43 ms |
| Clef | 27B | Qwen3.8-27B | 英语 | 是 | Apache 2.0 | --- |
* 在单块 NVIDIA RTX PRO 6000 上回答一个问题的中位时间。
在 Decision models 合集 中查找这些模型,更多模型即将推出。社区 Decision Index 展示了它们的对比。
快速开始
从 llama.app 获取最新版 llama.cpp(或运行 llama update),然后启动模型:
bash
llama serve -hf ggml-org/Kev-4B-GGUF
一个请求包含一个状态和一个或多个问题。问题有三种类型:
| 类型 | 你发送 | 你得到 |
|---|---|---|
choice |
选项(含可选描述) | 最佳选项 + 每个选项的概率 |
score |
2 到 10 个等级(从低到高) | 期望等级(可以落在两个等级之间) |
noul |
是/否问题 | 是(yes)的概率 |
发送包含状态和问题的请求:
bash
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "Customer message: I was charged twice for my order last week and nobody has replied.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "payments, charges, refunds, invoices",
"shipping": "delivery, tracking, lost or late parcels",
"technical": "bugs, errors, login problems"
}
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["can wait", "this week", "today", "right now"]
}
}
}'
响应(数值已四舍五入):
json
{
"model": "ggml-org/Kev-4B-GGUF",
"answers": {
"route": {
"type": "choice",
"choice": "billing",
"probabilities": {"billing": 0.9049, "shipping": 0.0275, "technical": 0.0676},
"confidence": 0.8574
},
"angry": {
"type": "noul",
"noul": 0.8208
},
"urgency": {
"type": "score",
"score": 2.2821,
"legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
"probabilities": {"0": 0.036, "1": 0.1937, "2": 0.2225, "3": 0.5478},
"confidence": 0.2821
}
},
"usage": {"input_tokens": 130, "output_tokens": 0}
}
完整参考见 服务器文档。
图片
某些模型(目前为 OpenJev)还可以读取图片,如文档或截图。视觉投影器会自动下载:
bash
llama serve -hf ggml-org/OpenJev-GGUF
例如,对上传的文档进行分类:
python
import base64
import requests
with open("document.png", "rb") as f:
image = "data:image/png;base64," + base64.b64encode(f.read()).decode()
response = requests.post("http://localhost:8080/v1/systemone", json={
"state": "A file uploaded by a customer.",
"images": [image],
"questions": {
"kind": {
"type": "choice",
"instructions": "What kind of document is this?",
"criteria": {"invoice": None, "receipt": None, "contract": None, "other": None},
},
},
})
print(response.json()["answers"]["kind"]["choice"]) # invoice
state 也可以是聊天消息列表。任何 image_url 部分(data URL)都会被当作图片读取,与聊天补全相同。
多模型,单服务器
在路由模式下,模型按需加载,你可以在每个请求中选择一个:
bash
llama serve
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'
/v1/models 列出所有 id。当只加载了一个模型时,model 字段会被忽略。
使用技巧
- 尝试不同大小的模型。 小模型更快,大模型知识更丰富。Decision Index 对它们进行了比较。
- 描述你的选项。 Julia-1 在仅有标签时将"我被扣了两次钱"路由到
shipping,而在每个选项都有描述时路由到billing(0.99)。 - 为每个模型选择置信度阈值。 常见的模式是对有信心的答案采取行动,其余的发送给人工。正确的阈值取决于模型:一个模糊的工单("Hi, quick question about my account")在 Julia-1 上得分 0.25,而在 Kev-4B 上得分 0.80。在选择阈值之前,请在你自己的示例上进行测试。
- 批量处理问题。 问题是独立回答的,Kev-4B、lev 和 OpenJev 只处理一次状态。
- 尝试不同的量化。 与任何 GGUF 一样,这些模型有多种精度可选,例如
llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0。