1.2-DeepSeek 客户端:同步、SSE 流式与 tool_calls 解析

DeepSeek 客户端:同步、SSE 流式与 tool_calls 解析

源码仓库:

上一篇看到 main.cpp 里 DeepSeekClient 被 std::make_shared 出来,以 IModelClient 的抽象注入给 AgentRuntime。这篇打开这个客户端。它只做两件事:把一轮对话发给 DeepSeek Chat Completions,再把结果解析回来;区别只在同步和流式两条路。真正容易出错的不是 HTTP,而是 SSE 分帧、tool_calls 参数的增量拼接、以及流式失败时到底能不能重试。

只读一个文件的话,读 src/llm/deepseek_client.cpp(454 行);类型定义在 include/aiagent/llm/deepseek_types.hpp。


1. 问题背景:为什么把模型调用单独封一层

AgentRuntime 关心的是「消息历史 + 工具循环 + 压缩」,它不应该知道 HTTP、JSON、SSE 分帧和重试。所以这里抽象出一个只有两个方法的接口(include/aiagent/llm/deepseek_client.hpp):

cpp 复制代码
class IModelClient {
public:
    virtual ~IModelClient() = default;
    virtual ChatResult Chat(const std::vector<ChatMessage>& messages,
                            const std::vector<ToolSpec>& tools = {},
                            const std::optional<std::string>& tool_choice = std::nullopt) = 0;
    virtual void ChatStream(const std::vector<ChatMessage>& messages,
                            const std::vector<ToolSpec>& tools,
                            const std::function<void(const StreamEvent&)>& on_event) = 0;
};

两个方法对应两种调用方式:Chat 一次性拿回完整结果,ChatStream 通过回调逐块吐事件。AgentRuntime 依赖这个接口,生产环境注入 DeepSeekClient,测试里注入一个假模型(tests/agent_runtime_test.cpp、tests/web_gateway_test.cpp 里都这么干)。这就是为什么整套测试不需要真实 API Key。

封在这一层里的具体职责有四块:

  • 构造请求体(消息序列化、工具描述、流式开关);
  • 发请求(同步 POST 或流式 POST)并处理状态码;
  • 解析响应(非流式 JSON / 流式 SSE / tool_calls);
  • 重试、退避与日志。

它并不自己实现 HTTP 和 SSE,而是复用 A2A 协议库已经写好的两个组件:a2a::core::HttpClient 和 a2a::core::SseParser(vendor/a2a-protocol/)。这也是上一篇说的「协议库是编译期依赖」的一个直接体现。


2. 整体设计:请求构造 → 传输 → 解析

两条路径共用前面一半,分叉在后面:

flowchart LR M[&#34;messages + tools&#34;] --> B[&#34;BuildChatRequestBody<br/>stream 决定是否带 stream_options&#34;] B --> T{&#34;调用方式&#34;} T -->|Chat| P[&#34;HttpClient.Post&#34;] T -->|ChatStream| S[&#34;HttpClient.PostStream<br/>on_chunk 回调&#34;] S --> SSE[&#34;SseParser.Feed<br/>按空行切事件&#34;] SSE --> PK[&#34;ParseStreamChunk<br/>逐帧解析&#34;] P --> PR[&#34;ParseChatResponse&#34;] PR --> R[&#34;ChatResult&#34;] PK --> E[&#34;StreamEvent 回调给上层&#34;]

三类数据结构把职责分开:

结构 出现位置 作用
ChatMessage / ToolSpec 请求 上层给的消息与工具描述
ChatResult / StreamChunk 响应解析 原始响应解出来的中间结果
StreamEvent 回调 流式场景给上层的高层事件

这种分层的好处是:ParseChatResponse 和 ParseStreamChunk 都是纯函数,输入字符串输出结构体,可以脱离网络单测(tests/llm_client_test.cpp 里 TestResponseParsing 就是喂 JSON 字符串)。网络层只负责把字节喂给它们。


3. 源码精读

3.1 类型设计里藏着的两个决定

先看请求侧的两个类型(include/aiagent/llm/deepseek_types.hpp):

cpp 复制代码
struct ToolCall {
    std::string id;
    std::string name;
    std::string arguments;  // 原始 JSON 字符串,方便增量拼接与回传
    int index = 0;          // 流式工具调用 delta 中的序号
    bool has_index = false; // 非流式响应中不存在该字段
};

struct ChatMessage {
    std::string role;                 // system / user / assistant / tool
    std::string content;
    std::string reasoning_content;    // DeepSeek 思考模型的推理文本
    std::string tool_call_id;         // role == tool 时使用
    std::string name;                 // 可选
    std::vector<ToolCall> tool_calls; // assistant 携带工具调用时使用
};

两个决定值得说:

  • arguments 保持原始 JSON 字符串,不解析成 nlohmann::json。 因为流式下模型是分片吐参数的({"city":" 然后 北京"}),必须能原样首尾相接;解析成对象反而会把不完整的 JSON 变成错误。等拼完整了,上层 CompositeToolProvider 再按工具自己的 schema 解析。
  • index 与 has_index 分开。 流式的 tool_calls delta 带 index,用来把分片归并到同一个调用;非流式响应的 tool_calls 没有这个字段。用一个 has_index 布尔区分「下标是 0」和「根本没有下标」,避免把多个调用错误地都归到 0 号。

具体看一个并行工具调用的例子。流式响应会分几帧下发,靠 index 认领(以 OpenAI 兼容格式为例,下面是三帧的原文):

text 复制代码
// 第 1 帧:两个调用同时出现,带 id / name / index
{"choices":[{"delta":{"tool_calls":[
  {"index":0,"id":"call_a","function":{"name":"read","arguments":"{\"path\":"}},
  {"index":1,"id":"call_b","function":{"name":"bash","arguments":"{\"cmd\":"}}
]},"finish_reason":null}]}

// 第 2 帧:只有 arguments 片段,靠 index 归属
{"choices":[{"delta":{"tool_calls":[
  {"index":0,"function":{"arguments":"\"a.txt\"}"}},
  {"index":1,"function":{"arguments":"\"ls\"}"}}
]},"finish_reason":null}]}

// 第 3 帧:结束标记
{"choices":[{"delta":{},"finish_reason":"tool_calls"}]}

非流式则是一次性给全,且不带 index:

json 复制代码
{"choices":[{"message":{"tool_calls":[
  {"id":"call_a","function":{"name":"read","arguments":"{\"path\":\"a.txt\"}"}},
  {"id":"call_b","function":{"name":"bash","arguments":"{\"cmd\":\"ls\"}"}}
]},"finish_reason":"tool_calls"}]}

如果没有 index,第 2 帧里两个 arguments 片段就无法判断该拼给谁;若简单按数组下标 0/1 认领,遇到「只回了 index=1 一个调用」的帧又会把片段错拼到 0 号。pending_tool_calls 用 std::map<size_t, ToolCall> 以 index 为键,就是为解决这个问题。ToolCallsFromJson() 同时被非流式与流式解析复用,流式正常会带 index,但万一某个兼容实现不带,就用 pending_tool_calls.size() 顺序兜底------这也是要单独用一个 has_index 区分「index 就是 0」的原因。

响应侧则把「一个 SSE 帧」和「给上层的事件」分开:

cpp 复制代码
// 单个 SSE data 帧解析后的增量。
struct StreamChunk {
    std::string content;
    std::string reasoning_content;
    std::vector<ToolCall> tool_calls;
    std::string finish_reason;
    ChatUsage usage;  // 仅在 include_usage 时,最后一帧携带
};

// ChatStream 回调中的高层事件。
struct StreamEvent {
    std::string content;
    std::string reasoning_content;
    std::string finish_reason;
    std::vector<ToolCall> tool_calls;
    ChatUsage usage;
};

StreamChunk 是「这一帧有什么」,StreamEvent 是「要告诉上层什么」。两者的字段现在接近,但语义不同:SseParser 已经滤掉了 : 开头的注释/心跳行,但一个 data 帧解析出来仍可能各字段都为空(例如空 delta),所以不能假设每帧都有内容;而 StreamEvent 只有在至少一个字段非空时才会派发。这个区别在 3.4 节会看到具体后果。

3.2 请求体构造:两条路共用同一个函数

BuildChatRequestBody() 同时服务同步和流式,靠 stream 参数分叉:

cpp 复制代码
std::string BuildChatRequestBody(const std::string& model,
                                 const std::vector<ChatMessage>& messages,
                                 const std::vector<ToolSpec>& tools,
                                 bool stream,
                                 const std::optional<std::string>& tool_choice) {
    nlohmann::json body{
        {"model", model},
        {"stream", stream},
    };
    if (stream) {
        // 让流式响应的最后一帧带上 usage,供 Agent 做上下文计量。
        body["stream_options"] = {{"include_usage", true}};
    }

    nlohmann::json message_array = nlohmann::json::array();
    for (const ChatMessage& message : messages) {
        message_array.push_back(MessageToJson(message));
    }
    body["messages"] = std::move(message_array);
    // ... tools / tool_choice
    return body.dump();
}

单条消息的序列化在 MessageToJson() 里,有几个容易踩的兼容性点:

cpp 复制代码
nlohmann::json MessageToJson(const ChatMessage& message) {
    nlohmann::json object;
    object["role"] = message.role;
    if (message.role == "tool") {
        object["tool_call_id"] = message.tool_call_id;
    }
    if (!message.name.empty()) {
        object["name"] = message.name;
    }
    if (!message.reasoning_content.empty()) {
        object["reasoning_content"] = message.reasoning_content;
    }
    if (!message.tool_calls.empty()) {
        object["tool_calls"] = ToolCallsToJson(message.tool_calls);
    }
    // OpenAI/DeepSeek 兼容格式要求 content 至少是字符串;空内容按空串发送。
    object["content"] = message.content;
    return object;
}
  • tool_call_id 只在 role == "tool" 时出现。工具结果必须带上它对应的调用 id,否则模型无法把结果与请求配对。
  • content 无条件写出 ,即使为空字符串。当 assistant 只发起 tool_calls 而没有文本时,content 是被省略还是空串,不同兼容实现处理不一致;这里统一发 ""。
  • reasoning_content 只在非空时写。它是 DeepSeek 思考模型的推理文本,回传与否由调用方决定:AgentRuntime 的同步路径会把 assistant 的 reasoning_content 一起塞回下一轮请求(src/agent/agent_runtime.cpp),流式路径目前只回填 content 与 tool_calls,不回传 reasoning。

工具描述的序列化绕了一个弯。请求里有两处和 tool 有关:tools(我们提供给模型的工具描述)直接把 ToolSpec::parameters 塞进去;assistant 消息里的 tool_calls(模型返回的调用)在 ToolCallsToJson() 里给 arguments 兜了个默认值:

cpp 复制代码
array.push_back({{"id", call.id},
                 {"type", "function"},
                 {"function",
                  {{"name", call.name},
                   {"arguments", call.arguments.empty() ? "{}" : call.arguments}}}});

空参数的工具调用如果发 arguments: "",服务端可能判为非法 JSON,所以兜成 "{}"。这类兼容性细节没什么技术含量,但漏了就是 400。

3.3 同步调用:Chat 只管重试,DoChat 只管一次请求

同步路径拆成两层:Chat() 负责重试循环与日志,DoChat() 负责发一次请求。这样重试逻辑不污染请求逻辑:

cpp 复制代码
ChatResult DeepSeekClient::DoChat(const std::vector<ChatMessage>& messages,
                                  const std::vector<ToolSpec>& tools,
                                  const std::optional<std::string>& tool_choice) {
    if (config_.api_key.empty()) {
        throw DeepSeekException(0, "DeepSeek api_key is not configured");
    }

    a2a::core::HttpClient http;
    http.SetTimeoutSeconds(config_.timeout_seconds);
    http.AddHeader("Authorization", "Bearer " + config_.api_key);
    http.AddHeader("Content-Type", "application/json");

    const auto response = http.Post(DeepSeekChatEndpoint(config_),
                                    RequestBody(messages, tools, false, tool_choice),
                                    "application/json");
    if (!response.IsSuccess()) {
        throw DeepSeekException(response.status_code, DeepSeekErrorMessage(response.body));
    }
    return ParseChatResponse(response.body);
}

这里有几个约定值得记住:

  • 每次请求新建一个 HttpClient,不共享。 HttpClient 内部每个请求独立创建并释放 curl easy handle,所以多线程下各用各的最省心。header 也随对象走,避免并发改同一个 header map。
  • 状态码非 2xx 一律抛 DeepSeekException,把状态码带上。 上层(Chat)靠状态码决定要不要重试;AgentRuntime 靠 IsContextOverflowMessage() 判断是不是上下文超长。
  • 错误信息来自 DeepSeekErrorMessage(body) :优先取响应体里 error.message,解析失败就截断原 body 前 500 字节。不把原始响应整个塞进异常,避免日志里出现超长内容。
  • URL 由 DeepSeekChatEndpoint() 推导 :base_url 已经以 /chat/completions 结尾就原样用,否则去掉尾部 / 再拼。这样 https://api.deepseek.com 和 https://api.deepseek.com/ 都能用。

解析本身是一个纯函数:

cpp 复制代码
ChatResult ParseChatResponse(const std::string& body) {
    const nlohmann::json root = nlohmann::json::parse(body);
    if (root.contains("error")) {
        throw DeepSeekException(0, DeepSeekErrorMessage(body));
    }
    ChatResult result;
    result.id = JsonString(root, "id");
    if (root.contains("choices") && root["choices"].is_array()) {
        for (const auto& choice_value : root["choices"]) {
            ChatChoice choice;
            const auto& message =
                    choice_value.contains("message") ? choice_value["message"] : choice_value;
            choice.content = JsonString(message, "content");
            choice.tool_calls = ToolCallsFromJson(message.value("tool_calls", nlohmann::json::array()));
            choice.finish_reason = JsonString(choice_value, "finish_reason");
            // ... role / reasoning_content
            result.choices.push_back(std::move(choice));
        }
    }
    // ... usage
    return result;
}

choice_value.contains("message") ? message : choice_value 这个三元表达式是为了兼容两种返回形态:Chat Completions 把内容放在 choices[].message,但有些实现(测试里的假响应也可以)直接放在 choices[] 上。多几行代码换来解析器对两种格式都成立。

3.4 流式调用:SSE 分帧与 tool_calls 增量拼接

先看一次带工具调用的流式请求,从 AgentRuntime 发起、到回调结束的完整过程:

sequenceDiagram autonumber participant AR as AgentRuntime participant DC as DeepSeekClient participant HC as a2a HttpClient participant SP as a2a SseParser participant DS as DeepSeek API AR->>DC: ChatStream(messages, tools, on_event) DC->>DC: BuildChatRequestBody(stream=true, include_usage) DC->>HC: PostStream(url, body, on_chunk) HC->>DS: POST /chat/completions (stream) DS-->>HC: 200 + 分块 SSE body loop 每个网络分片 HC->>SP: Feed(chunk) SP->>DC: 一个 data 帧(空行为界) DC->>DC: ParseStreamChunk(data) alt data == [DONE] DC->>DC: 忽略,流结束 else 文本增量帧 DC-->>AR: StreamEvent{content / reasoning_content} else tool_calls 分片帧 DC->>DC: pending_tool_calls[index].arguments += 片段 else finish_reason == tool_calls DC-->>AR: StreamEvent{tool_calls=[完整调用]} else usage 帧(choices 为空) DC-->>AR: StreamEvent{usage} end end HC-->>DC: PostStream 返回 alt 本回合抛异常 DC->>DC: !delivered_any && 429/5xx/网络异常 → 退避后整轮重试 DC->>DC: 否则把 DeepSeekException 抛给 AgentRuntime end

图里有四处是本节重点:网络分片如何喂给 SseParser、ParseStreamChunk 怎么给帧分类、tool_calls 怎么跨帧累加、异常时重试的边界在哪。下面逐一展开。

流式请求用 HttpClient::PostStream() 拿到「原始网络分片」,再交给 SseParser 按事件边界切分。注意回调拿到的是任意长度的网络块,不是完整的一行:

cpp 复制代码
std::map<std::size_t, ToolCall> pending_tool_calls;
a2a::core::SseParser parser([&](const std::string& data) {
    if (parse_failed || data == "[DONE]") {
        return;
    }
    const StreamChunk chunk = ParseStreamChunk(data);
    // ... 把 role/content/usage 转成 StreamEvent 回调
    for (const ToolCall& call : chunk.tool_calls) {
        const std::size_t call_index =
                call.has_index ? static_cast<std::size_t>(call.index)
                               : pending_tool_calls.size();
        ToolCall& pending = pending_tool_calls[call_index];
        if (!call.id.empty()) {
            pending.id = call.id;
        }
        if (!call.name.empty()) {
            pending.name = call.name;
        }
        pending.arguments += call.arguments;   // 关键:累加,而不是覆盖
    }
    if (chunk.finish_reason == "tool_calls" && !pending_tool_calls.empty()) {
        StreamEvent event;
        event.finish_reason = chunk.finish_reason;
        for (auto& [index, call] : pending_tool_calls) {
            event.tool_calls.push_back(std::move(call));
        }
        pending_tool_calls.clear();
        if (on_event) {
            on_event(event);
        }
    }
});

const auto response = http.PostStream(DeepSeekChatEndpoint(config_),
                                      RequestBody(messages, tools, true, std::nullopt),
                                      "application/json",
                                      [&](std::string_view chunk) { parser.Feed(chunk); });

关键过程用一张放大的时序图看更清楚(只看 tool_calls 的拼接):

sequenceDiagram participant DS as DeepSeek SSE participant SP as SseParser participant CS as ChatStream participant AR as AgentRuntime DS-->>SP: data: delta.tool_calls[0] {index:0, id, name, arguments:&#34;{\&#34;city\&#34;:\&#34;&#34;} SP->>CS: ParseStreamChunk CS->>CS: pending_tool_calls[0].arguments += &#34;{\&#34;city\&#34;:\&#34;&#34; DS-->>SP: data: delta.tool_calls[0] {arguments:&#34;北京\&#34;}&#34;} SP->>CS: ParseStreamChunk CS->>CS: pending_tool_calls[0].arguments += &#34;北京\&#34;}&#34; DS-->>SP: data: {finish_reason:&#34;tool_calls&#34;} SP->>CS: ParseStreamChunk CS-->>AR: StreamEvent{tool_calls:[{id,name,arguments:&#34;{\&#34;city\&#34;:\&#34;北京\&#34;}&#34;}]}

这里有三个必须理解的点:

  1. arguments 是跨帧累加的。 模型可能一个字符一个字符地吐 JSON,每一帧只带片段。代码用 pending[call_index] 累积,id 和 name 只在非空时覆盖(它们通常只在首帧出现一次)。
  2. 上层只在最后拿到一次完整的 tool_calls。 中间帧的 tool_calls 被消费进 pending_tool_calls,不通过事件下发;等 finish_reason == "tool_calls" 时把攒好的调用一次性组装成 StreamEvent。这样 AgentRuntime 不需要自己拼分片,拿到就是完整可执行的调用。
  3. [DONE] 是协议结束标记,不解析。 SSE 流最后会有 data: [DONE],直接 return。

SseParser 的职责比看上去细。它按「空行 = 事件边界」切分,只认 data: 字段,多行 data 用换行拼接,忽略 : 开头的注释/心跳行,兼容 LF 与 CRLF(vendor/a2a-protocol/src/core/sse_parser.cpp)。所以 DeepSeekClient 这一层不用自己处理分帧,只处理语义。

3.5 重试与退避:同步能重试,流式要看有没有派发过事件

Chat() 和 ChatStream() 各有一个重试循环,规则不同。同步路径:

cpp 复制代码
bool IsRetryableStatusCode(int status_code) {
    return status_code == 429 || status_code >= 500;  // 限流与 5xx 值得重试
}

void SleepBackoff(int attempt) {
    constexpr long kBaseDelayMs = 500;
    constexpr long kMaxDelayMs = 5000;
    long delay = kBaseDelayMs * (1L << attempt);  // 500ms * 2^attempt
    if (delay > kMaxDelayMs) {
        delay = kMaxDelayMs;
    }
    std::this_thread::sleep_for(std::chrono::milliseconds(delay));
}

同步 Chat 的重试条件是:IsRetryableStatusCode(status) 且 attempt < max_retries。429 和 5xx 可重试;401/400 这类客户端错误直接抛出;网络层异常(a2a::core::JsonRpcException,对应连接失败/超时)也重试。等待时间是指数退避 500ms * 2^attempt,封顶 5 秒。测试里 TestRetryOnTransientServerErrors 让服务端前两次返回 500、第三次成功,断言客户端命中了 3 次。

流式路径多一个变量 delivered_any:

cpp 复制代码
// 一旦已向调用方派发过事件,中途失败不再重试------否则会产生重复的
// token/工具调用片段,破坏 Agent 侧的消息拼接。
bool delivered_any = false;
// ...
} catch (const DeepSeekException& e) {
    const bool retry = !delivered_any &&
                       IsRetryableStatusCode(e.status_code()) &&
                       attempt < config_.max_retries;
    // ...
}

这就是流式重试的核心权衡:重试的前提是「还没吐过任何东西」 。如果已经推了几个 token 出去,再重试会从头再推一遍,前端会看到重复文本,AgentRuntime 的工具调用拼接也可能被打乱。所以流式只在连接阶段(尚未派发事件)失败时重试,一旦开始产出就宁可报错。测试里 TestStreamUsageEvent 只覆盖成功路径,重试边界靠这段注释和 delivered_any 保证。

3.6 usage:为上下文计量留的接口

流式请求默认带上 stream_options.include_usage = true,DeepSeek 会在流的最后额外发一帧 usage。这一帧的特点是 choices 为空数组,usage 在顶层:

cpp 复制代码
StreamChunk ParseStreamChunk(const std::string& data) {
    const nlohmann::json root = nlohmann::json::parse(data);
    // ...
    StreamChunk chunk;
    // include_usage 的 usage 帧通常 choices 为空,需独立于 choices 解析。
    if (root.contains("usage") && root["usage"].is_object()) {
        chunk.usage.prompt_tokens = JsonInt(root["usage"], "prompt_tokens");
        chunk.usage.completion_tokens = JsonInt(root["usage"], "completion_tokens");
        chunk.usage.total_tokens = JsonInt(root["usage"], "total_tokens");
    }
    if (root.contains("choices") && root["choices"].is_array() &&
        !root["choices"].empty()) {
        // ... 解析 delta
    }
    return chunk;
}

解析 usage 时必须放在 choices 判断之外 ,否则这一帧会被整个跳过------这是最容易漏的一处。这个 usage 最终通过 StreamEvent.usage 传给 AgentRuntime,成为 2.4/2.5 里「上下文计量」的真实 token 来源:有 usage 就用真实值,其后新增的尾部消息再估算。非流式响应里 usage 在顶层,ParseChatResponse 直接解析。


4. 那些值得说的细节

工具调用参数为什么要保留原始字符串。 前面提过一次,这里展开:流式 arguments 分片下,片段之间可能正好切在字符串转义、UTF-8 多字节字符的中间。任何「收到一帧就 json::parse」的写法都会间歇性失败。正确做法是字节级拼接,拼完再解析。代价是 ToolCall.arguments 在到达工具执行前一直是字符串,调用方要自己记住「它可能不是合法 JSON」。

非流式与流式的 tool_calls 结构不同。 流式 delta 带 index,非流式不带。代码用 has_index 区分,并且流式归并时用 pending_tool_calls[call_index] 这个 std::map 按键累积。如果不区分,多个并行工具调用(index 0、1、2)会被错误地合成一个。tests/llm_client_test.cpp 的 TestResponseParsing 专门断言了 has_index 和 index。

空事件过滤与工具片段的错位。 在为 role/content/reasoning/finish_reason/usage 构造 StreamEvent 时,代码有一个「至少有一个字段非空」的过滤条件;而 tool_calls 的累积在这个过滤之外。这样纯工具调用的 delta(没有可见文本)不会被丢掉,但也不会产生一堆空事件。这是「帧」与「事件」分开建模的直接好处。

流式重试的边界是有意保守的。 有人会想「既然 5xx 可重试,流式为什么不行」。答案在 3.5:已经派发的 token 没法撤回。真正完整的做法是记录已发送偏移、重连时要求服务端续传,但 Chat Completions 没有这个语义,所以这里选择「宁可失败也不重复」。这是个明确的限制,不是遗漏。


5. 测试与验证

tests/llm_client_test.cpp 用进程内 httplib::Server 冒充 DeepSeek 端点,覆盖了纯解析、请求体、重试三块:

用例 覆盖点
TestConfigParsing [deepseek] 解析、缺 api_key 报错、endpoint 推导
TestRequestBody 消息/tool_calls/reasoning/tool 结果/tools 序列化、include_usage 只在流式出现
TestResponseParsing 非流式 JSON、流式帧、tool_calls 分片、usage 帧、错误信息提取
TestRetryOnTransientServerErrors 5xx 两次后成功,命中 3 次
TestNoRetryOnClientError 401 不重试,只命中一次
TestNoRetryWhenDisabled max_retries=0 关闭重试
TestStreamUsageEvent 流式 usage 帧透传到 StreamEvent

跑起来(不需要 API Key):

bash 复制代码
cmake -S . -B build -DCMAKE_BUILD_TYPE=Debug
cmake --build build -j
ctest --test-dir build -R aiagent_llm_client -V

想对真实模型说话,用命令行示例(需要 conf/conf.ini 或 DEEPSEEK_API_KEY):

bash 复制代码
./build/aiagent_deepseek_cli --config conf/conf.ini "你好"
./build/aiagent_deepseek_cli --config conf/conf.ini --stream "介绍一下自己"

examples/deepseek_chat.cpp 把 ChatStream 的 event.content 直接 std::cout,是最短的流式演示;它没有接工具执行,所以工具调用只在 AgentRuntime 那层才会被真正执行(下一篇)。


6. 小结

  1. DeepSeekClient 把 HTTP/JSON/SSE/重试封在 IModelClient 后面,只暴露 Chat 与 ChatStream;AgentRuntime 和测试都不依赖真实网络。
  2. 同步与流式共用 BuildChatRequestBody,区别只在 stream_options.include_usage 和响应解析方式;tool_calls.arguments 全程保持原始字符串。
  3. 最容易出错的两处是流式 tool_calls 的跨帧累加、以及 usage 帧 choices 为空要独立解析;流式重试的边界是「未派发过任何事件」,这是为了避免重复 token。
  4. 可靠性、日志、上下文计量是分三次加到客户端上的,delivered_any、include_usage、指数退避都能在改动史里找到对应记录。

下一篇进入 AgentRuntime,看它怎么拿着这个客户端跑工具调用循环、把 tool_calls 翻译成真正的工具执行,并把结果回填成下一轮消息。

相关推荐
Zelman1 小时前
第06章-超节点
人工智能·后端
用户79457223954131 小时前
当卖 token 的遇上修高速公路的:AI 周期的真空期在哪
人工智能·agent
Raas1001 小时前
AI网关能做语义缓存吗?MAI Gateway(魔芋企业级AI网关)实战能力深度解读
java·后端·spring
2601_968900771 小时前
意图召回与相关性过滤实战:把出价放在排序之后
人工智能
知几蜗牛1 小时前
Java 17标准HTTP调用语音转写API的超时与错误处理
人工智能
静水深码1 小时前
AI 能翻好谐音梗吗
人工智能·后端
9i编程1 小时前
8. 教 AI 上班:带出我的数字同事 —— 亲测第一轮:7 个问题逐条改,打包运行再回看
人工智能·openai·ai编程
柏修1 小时前
Java工程师转型Agent Engineer - 24周详细学习计划
后端
牛气冲天的星球1 小时前
服务器资源监控之grafana+Prometheus配置
后端
用户360055579001 小时前
【FHE 同态加密】我们如何实现同态加密推理(九):逐层尺度自适应——同态加密里为什么每层都要单独定标(scale 管理)
人工智能