llama.cpp 新特性:决策模型

原文:New in llama.cpp: Decision Models

发布日期:2026 年 10 月 2 日

作者:Xuan-Son Nguyen (ngxson)、Victor Mustar、ggml-org

llama.cpp 服务器现在通过 /v1/systemone 端点支持决策模型。你发送一个状态(文本、JSON、截图)和类型化问题,模型在单次前向传播中返回每个选项的概率。

API 遵循 TypeSafe 的 Jev 模型引入的 System One 格式,因此现有客户端只需更换 base URL 即可。实现细节见 PR #29818。

什么是决策模型? 决策模型通过为你给定的选项打分来回答问题,而不是生成文本。聊天模型每个输出 token 需要一次前向传播,且输出仍需解析。决策模型只读取一次输入,其答案始终是你给出的选项之一,并附带概率。典型用途包括:路由请求、内容审核、检查 Agent 步骤是否成功、或选择下一步动作。

支持的模型

模型 大小 基于 语言 图片 许可证 速度*
Julia-1 144M mmBERT-small 50+ 否 Apache 2.0 3 ms
Laya 421M ModernBERT-large 英语 否 Apache 2.0 5 ms
Kev-4B 4B Qwen3.5-4B-Base 英语 否 Apache 2.0 12 ms
lev 4B Qwen3.5-4B 英语 否 Apache 2.0 36 ms
OpenJev 27B Qwen3.8-27B en, de, fr, hi, zh, ja 是 CC BY-NC 4.0 43 ms
Clef 27B Qwen3.8-27B 英语 是 Apache 2.0 ---

* 在单块 NVIDIA RTX PRO 6000 上回答一个问题的中位时间。

在 Decision models 合集 中查找这些模型,更多模型即将推出。社区 Decision Index 展示了它们的对比。

快速开始

从 llama.app 获取最新版 llama.cpp(或运行 llama update),然后启动模型:

bash 复制代码
llama serve -hf ggml-org/Kev-4B-GGUF

一个请求包含一个状态和一个或多个问题。问题有三种类型:

类型 你发送 你得到
choice 选项(含可选描述) 最佳选项 + 每个选项的概率
score 2 到 10 个等级(从低到高) 期望等级(可以落在两个等级之间)
noul 是/否问题 是(yes)的概率

发送包含状态和问题的请求:

bash 复制代码
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "payments, charges, refunds, invoices",
          "shipping": "delivery, tracking, lost or late parcels",
          "technical": "bugs, errors, login problems"
        }
      },
      "angry": {
        "type": "noul",
        "instructions": "Is the customer angry?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["can wait", "this week", "today", "right now"]
      }
    }
  }'

响应(数值已四舍五入):

json 复制代码
{
  "model": "ggml-org/Kev-4B-GGUF",
  "answers": {
    "route": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {"billing": 0.9049, "shipping": 0.0275, "technical": 0.0676},
      "confidence": 0.8574
    },
    "angry": {
      "type": "noul",
      "noul": 0.8208
    },
    "urgency": {
      "type": "score",
      "score": 2.2821,
      "legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
      "probabilities": {"0": 0.036, "1": 0.1937, "2": 0.2225, "3": 0.5478},
      "confidence": 0.2821
    }
  },
  "usage": {"input_tokens": 130, "output_tokens": 0}
}

完整参考见 服务器文档。

图片

某些模型(目前为 OpenJev)还可以读取图片,如文档或截图。视觉投影器会自动下载:

bash 复制代码
llama serve -hf ggml-org/OpenJev-GGUF

例如,对上传的文档进行分类:

python 复制代码
import base64
import requests

with open("document.png", "rb") as f:
    image = "data:image/png;base64," + base64.b64encode(f.read()).decode()

response = requests.post("http://localhost:8080/v1/systemone", json={
    "state": "A file uploaded by a customer.",
    "images": [image],
    "questions": {
        "kind": {
            "type": "choice",
            "instructions": "What kind of document is this?",
            "criteria": {"invoice": None, "receipt": None, "contract": None, "other": None},
        },
    },
})

print(response.json()["answers"]["kind"]["choice"])  # invoice

state 也可以是聊天消息列表。任何 image_url 部分(data URL)都会被当作图片读取,与聊天补全相同。

多模型,单服务器

在路由模式下,模型按需加载,你可以在每个请求中选择一个:

bash 复制代码
llama serve
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'

/v1/models 列出所有 id。当只加载了一个模型时,model 字段会被忽略。

使用技巧

  • 尝试不同大小的模型。 小模型更快,大模型知识更丰富。Decision Index 对它们进行了比较。
  • 描述你的选项。 Julia-1 在仅有标签时将"我被扣了两次钱"路由到 shipping,而在每个选项都有描述时路由到 billing(0.99)。
  • 为每个模型选择置信度阈值。 常见的模式是对有信心的答案采取行动,其余的发送给人工。正确的阈值取决于模型:一个模糊的工单("Hi, quick question about my account")在 Julia-1 上得分 0.25,而在 Kev-4B 上得分 0.80。在选择阈值之前,请在你自己的示例上进行测试。
  • 批量处理问题。 问题是独立回答的,Kev-4B、lev 和 OpenJev 只处理一次状态。
  • 尝试不同的量化。 与任何 GGUF 一样,这些模型有多种精度可选,例如 llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0。
相关推荐
鲲穹AI种草1 小时前
AI 生成 PPT 工具怎么选?鲲穹 PPT 功能实测与横向对比
人工智能·powerpoint
AI技术新视界1 小时前
Strata 的工作原理
人工智能·推理引擎·本地ai
microrain2 小时前
先应答,再入库:SagooIoT 接入 GB/T 32960 的四层宿主改造
物联网·golang·开源·sagooiot
2501_926978332 小时前
给自指系统接一个外部锚 —— 一个关于「用 AI 观察自己」的方法
人工智能·经验分享·笔记·机器学习·ai写作
不悔哥2 小时前
小智AI聊天机器人:ESP32上的开源语音助手拆解
人工智能·机器人·开源
枯木◊靠推文躺平版2 小时前
2026 企业级 AI 办公平台选型指南:如何评估可落地的办公 AI 能力
大数据·人工智能
空奈qwq2 小时前
ANN 全连接神经网络进阶:从反向传播到完整实战
人工智能·深度学习·神经网络
Dawson Zhu2 小时前
《Agentic Design Patterns》第 4 章导读:反思(Reflection)
人工智能·语言模型·架构·aigc·agi
llqbzllll3 小时前
Spring AI 工具调用不是反射一下就结束:用 2.0.1 跑通失败恢复与调用上限
人工智能·后端