Agent 自主决策机制:什么时候该检索、什么时候该直接回答

摘要:自主决策是 AI Agent 区别于传统对话系统的核心特征。本文系统剖析 Agent 在运行时面临的四类关键决策------检索决策、工具选择决策、停止决策和降级决策,探讨每种决策的触发条件、评估维度和实现策略。通过一个完整的决策引擎实战代码,展示如何将决策逻辑从硬编码升级为可配置、可解释的规则系统。文章涵盖置信度评估、多工具路由、循环终止条件判断、错误降级策略等核心主题,并提供可直接复用的工程实现方案。
版本声明:本文适用于 2024-2025 年主流 LLM Agent 框架(LangChain/LangGraph、LlamaIndex、AutoGen 等),基于 OpenAI GPT-4o / Claude 3.5 Sonnet 等模型能力编写。不同模型的决策能力差异较大,实际落地时需根据具体模型调整阈值参数。


文章目录

    • [一、自主决策是 Agent 的核心特征:不是"能不能"而是"该不该"](#一、自主决策是 Agent 的核心特征:不是"能不能"而是"该不该")
      • [1.1 从被动响应到主动决策](#1.1 从被动响应到主动决策)
      • [1.2 决策的四种类型](#1.2 决策的四种类型)
      • [1.3 为什么决策比执行更难](#1.3 为什么决策比执行更难)
      • [1.4 决策的评估维度](#1.4 决策的评估维度)
    • 二、检索决策:何时检索、检索什么、检索多少
      • [2.1 检索的三个层次](#2.1 检索的三个层次)
      • [2.2 检索触发的判断框架](#2.2 检索触发的判断框架)
      • [2.3 置信度评估:检索决策的核心指标](#2.3 置信度评估:检索决策的核心指标)
      • [2.4 检索量的动态控制](#2.4 检索量的动态控制)
      • [2.5 检索决策的流程图](#2.5 检索决策的流程图)
    • 三、工具选择决策:多个工具时如何选择
      • [3.1 工具选择的复杂性](#3.1 工具选择的复杂性)
      • [3.2 工具选择策略](#3.2 工具选择策略)
      • [3.3 工具选择的决策矩阵](#3.3 工具选择的决策矩阵)
      • [3.4 工具链编排:多工具组合决策](#3.4 工具链编排:多工具组合决策)
    • 四、停止决策:何时停止循环、何时输出最终结果
      • [4.1 停止决策的重要性](#4.1 停止决策的重要性)
      • [4.2 停止条件的设计](#4.2 停止条件的设计)
      • [4.3 停止决策的状态机](#4.3 停止决策的状态机)
    • 五、降级决策:出错时重试还是换方案
      • [5.1 错误处理的工程现实](#5.1 错误处理的工程现实)
      • [5.2 降级策略矩阵](#5.2 降级策略矩阵)
      • [5.3 降级决策引擎实现](#5.3 降级决策引擎实现)
      • [5.4 降级决策的流程](#5.4 降级决策的流程)
    • 六、实战:一个完整的决策引擎实现
      • [6.1 决策引擎架构](#6.1 决策引擎架构)
      • [6.2 完整实现](#6.2 完整实现)
      • [6.3 决策引擎的调用流程](#6.3 决策引擎的调用流程)
      • [6.4 决策日志与可观测性](#6.4 决策日志与可观测性)
    • 七、适用边界与风险提示
      • [7.1 适用场景](#7.1 适用场景)
      • [7.2 不适用场景](#7.2 不适用场景)
      • [7.3 风险提示](#7.3 风险提示)
      • [7.4 成本模型](#7.4 成本模型)
    • 八、总结
    • 参考资料

一、自主决策是 Agent 的核心特征:不是"能不能"而是"该不该"

1.1 从被动响应到主动决策

传统对话系统的工作模式是"输入-输出"的单次映射:用户问一个问题,系统给一个回答。系统本身不做任何决策------它不需要决定是否需要更多信息,不需要判断是否该调用外部工具,不需要评估回答是否充分。

Agent 的本质区别在于:它在回答用户问题之前,需要先回答一系列"元问题":

  • 我现在掌握的信息足够回答这个问题吗?
  • 如果不够,我应该去哪里获取?
  • 获取到信息后,这些信息是否足以支撑一个高质量回答?
  • 如果获取过程中出了错,我应该重试还是换一种方式?
  • 我是否已经可以输出最终答案了,还是需要再做一轮处理?

这就是自主决策------Agent 在运行时对自身行为路径的动态选择能力。

1.2 决策的四种类型

在 Agent 的运行生命周期中,决策行为可以归纳为四个核心类型:

决策类型 核心问题 典型场景 复杂度
检索决策 是否需要获取外部信息 知识不足、信息过时、需要事实依据 中
工具选择决策 使用哪个工具完成任务 多工具可用、功能重叠、成本差异 高
停止决策 是否可以输出最终结果 循环终止、答案完整性判断 中
降级决策 出错后如何恢复 工具失败、模型超时、结果质量低 高

这四类决策相互交织,构成了 Agent 的完整决策链路。下面我们逐一深入剖析。

1.3 为什么决策比执行更难

执行一个工具调用是确定性的------参数对了,API 就会返回结果。但决策本身是概率性的------"该不该检索"没有一个绝对的正确答案,它取决于对当前上下文、知识边界、用户意图的综合判断。

python 复制代码
# 一个简单的决策示例:是否需要检索外部知识
def should_retrieve(query: str, context: list, confidence: float) -> bool:
    """判断是否需要检索外部知识"""
    # 规则1:置信度低于阈值,必须检索
    if confidence < 0.6:
        return True
    
    # 规则2:查询包含时间敏感词,需要检索最新信息
    time_sensitive_keywords = ["今天", "最新", "当前", "现在", "最近", "2024", "2025"]
    if any(kw in query for kw in time_sensitive_keywords):
        return True
    
    # 规则3:上下文已有相关信息且置信度足够,跳过检索
    if context and confidence >= 0.8:
        return False
    
    # 默认:倾向于检索
    return True

代码解释 :上面的 should_retrieve 函数展示了一个基于规则的检索决策器。它综合了三个维度的判断:模型置信度(反映模型对自身回答的确定程度)、时间敏感性(查询是否涉及需要实时信息的词汇)、上下文充分性(当前对话历史中是否已有足够信息)。这种规则驱动的方式虽然简单,但可解释性强、易于调参,是工程化落地的首选方案。

图:决策引擎四层架构:输入层(查询/历史/状态)→决策层(检索决策/工具选择/停止决策/降级决策并行)→行动层(执行/等待/停止/重试)→输出层(最终答案/中间结果/错误报告)

1.4 决策的评估维度

任何决策都不是非黑即白的。一个好的决策框架需要考虑多个维度:
#mermaid-svg-j9sDpXZqqJVd6Who{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;fill:#333;}@keyframes edge-animation-frame{from{stroke-dashoffset:0;}}@keyframes dash{to{stroke-dashoffset:0;}}#mermaid-svg-j9sDpXZqqJVd6Who .edge-animation-slow{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 50s linear infinite;stroke-linecap:round;}#mermaid-svg-j9sDpXZqqJVd6Who .edge-animation-fast{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 20s linear infinite;stroke-linecap:round;}#mermaid-svg-j9sDpXZqqJVd6Who .error-icon{fill:#552222;}#mermaid-svg-j9sDpXZqqJVd6Who .error-text{fill:#552222;stroke:#552222;}#mermaid-svg-j9sDpXZqqJVd6Who .edge-thickness-normal{stroke-width:1px;}#mermaid-svg-j9sDpXZqqJVd6Who .edge-thickness-thick{stroke-width:3.5px;}#mermaid-svg-j9sDpXZqqJVd6Who .edge-pattern-solid{stroke-dasharray:0;}#mermaid-svg-j9sDpXZqqJVd6Who .edge-thickness-invisible{stroke-width:0;fill:none;}#mermaid-svg-j9sDpXZqqJVd6Who .edge-pattern-dashed{stroke-dasharray:3;}#mermaid-svg-j9sDpXZqqJVd6Who .edge-pattern-dotted{stroke-dasharray:2;}#mermaid-svg-j9sDpXZqqJVd6Who .marker{fill:#333333;stroke:#333333;}#mermaid-svg-j9sDpXZqqJVd6Who .marker.cross{stroke:#333333;}#mermaid-svg-j9sDpXZqqJVd6Who svg{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;}#mermaid-svg-j9sDpXZqqJVd6Who p{margin:0;}#mermaid-svg-j9sDpXZqqJVd6Who .label{font-family:"trebuchet ms",verdana,arial,sans-serif;color:#333;}#mermaid-svg-j9sDpXZqqJVd6Who .cluster-label text{fill:#333;}#mermaid-svg-j9sDpXZqqJVd6Who .cluster-label span{color:#333;}#mermaid-svg-j9sDpXZqqJVd6Who .cluster-label span p{background-color:transparent;}#mermaid-svg-j9sDpXZqqJVd6Who .label text,#mermaid-svg-j9sDpXZqqJVd6Who span{fill:#333;color:#333;}#mermaid-svg-j9sDpXZqqJVd6Who .node rect,#mermaid-svg-j9sDpXZqqJVd6Who .node circle,#mermaid-svg-j9sDpXZqqJVd6Who .node ellipse,#mermaid-svg-j9sDpXZqqJVd6Who .node polygon,#mermaid-svg-j9sDpXZqqJVd6Who .node path{fill:#ECECFF;stroke:#9370DB;stroke-width:1px;}#mermaid-svg-j9sDpXZqqJVd6Who .rough-node .label text,#mermaid-svg-j9sDpXZqqJVd6Who .node .label text,#mermaid-svg-j9sDpXZqqJVd6Who .image-shape .label,#mermaid-svg-j9sDpXZqqJVd6Who .icon-shape .label{text-anchor:middle;}#mermaid-svg-j9sDpXZqqJVd6Who .node .katex path{fill:#000;stroke:#000;stroke-width:1px;}#mermaid-svg-j9sDpXZqqJVd6Who .rough-node .label,#mermaid-svg-j9sDpXZqqJVd6Who .node .label,#mermaid-svg-j9sDpXZqqJVd6Who .image-shape .label,#mermaid-svg-j9sDpXZqqJVd6Who .icon-shape .label{text-align:center;}#mermaid-svg-j9sDpXZqqJVd6Who .node.clickable{cursor:pointer;}#mermaid-svg-j9sDpXZqqJVd6Who .root .anchor path{fill:#333333!important;stroke-width:0;stroke:#333333;}#mermaid-svg-j9sDpXZqqJVd6Who .arrowheadPath{fill:#333333;}#mermaid-svg-j9sDpXZqqJVd6Who .edgePath .path{stroke:#333333;stroke-width:2.0px;}#mermaid-svg-j9sDpXZqqJVd6Who .flowchart-link{stroke:#333333;fill:none;}#mermaid-svg-j9sDpXZqqJVd6Who .edgeLabel{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-j9sDpXZqqJVd6Who .edgeLabel p{background-color:rgba(232,232,232, 0.8);}#mermaid-svg-j9sDpXZqqJVd6Who .edgeLabel rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-j9sDpXZqqJVd6Who .labelBkg{background-color:rgba(232, 232, 232, 0.5);}#mermaid-svg-j9sDpXZqqJVd6Who .cluster rect{fill:#ffffde;stroke:#aaaa33;stroke-width:1px;}#mermaid-svg-j9sDpXZqqJVd6Who .cluster text{fill:#333;}#mermaid-svg-j9sDpXZqqJVd6Who .cluster span{color:#333;}#mermaid-svg-j9sDpXZqqJVd6Who div.mermaidTooltip{position:absolute;text-align:center;max-width:200px;padding:2px;font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:12px;background:hsl(80, 100%, 96.2745098039%);border:1px solid #aaaa33;border-radius:2px;pointer-events:none;z-index:100;}#mermaid-svg-j9sDpXZqqJVd6Who .flowchartTitleText{text-anchor:middle;font-size:18px;fill:#333;}#mermaid-svg-j9sDpXZqqJVd6Who rect.text{fill:none;stroke-width:0;}#mermaid-svg-j9sDpXZqqJVd6Who .icon-shape,#mermaid-svg-j9sDpXZqqJVd6Who .image-shape{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-j9sDpXZqqJVd6Who .icon-shape p,#mermaid-svg-j9sDpXZqqJVd6Who .image-shape p{background-color:rgba(232,232,232, 0.8);padding:2px;}#mermaid-svg-j9sDpXZqqJVd6Who .icon-shape .label rect,#mermaid-svg-j9sDpXZqqJVd6Who .image-shape .label rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-j9sDpXZqqJVd6Who .label-icon{display:inline-block;height:1em;overflow:visible;vertical-align:-0.125em;}#mermaid-svg-j9sDpXZqqJVd6Who .node .label-icon path{fill:currentColor;stroke:revert;stroke-width:revert;}#mermaid-svg-j9sDpXZqqJVd6Who :root{--mermaid-font-family:"trebuchet ms",verdana,arial,sans-serif;} 高
中
低
低成本
高成本
需要准确
允许近似
高质量
低质量
决策输入
置信度评估
直接回答
成本评估
必须检索
执行检索
用户期望评估
检索结果
结果质量评估
降级策略
换工具/换查询/重试

这个决策流程图展示了一个完整的决策链路:从输入到最终输出,涉及置信度评估、成本评估、用户期望评估、结果质量评估等多个判断节点。每个节点都是一个决策点,每个决策点都可能有不同的分支。


二、检索决策:何时检索、检索什么、检索多少

2.1 检索的三个层次

检索决策不是简单的"检索或不检索"的二选一,而是包含三个层次的渐进决策:

层次一:是否需要检索?

这是最基本的决策。如果模型自身的参数化知识已经能够回答用户的问题,检索就是多余的开销。但"能回答"和"能准确回答"是两回事------模型可能对自己的知识过于自信,产生"知识幻觉"。

层次二:检索什么?

确定需要检索后,下一步是决定检索什么类型的信息。是检索内部知识库?是搜索互联网?是查询数据库?不同类型的检索适用于不同场景。

层次三:检索多少?

检索结果的数量直接影响回答质量和响应延迟。检索太少可能遗漏关键信息,检索太多则引入噪声、增加处理时间和 token 消耗。

2.2 检索触发的判断框架

触发条件 检索类型 检索范围 示例
知识边界外 知识库/互联网 宽范围检索 "公司2024年Q3财报数据"
时间敏感 实时搜索 窄范围检索 "今天的天气"
事实核查 知识库验证 精确检索 "Python 3.12的match语句语法"
多角度分析 多源检索 宽范围+多源 "比较React和Vue的性能差异"
上下文不足 对话历史检索 窄范围 用户引用了10轮对话前提到的内容

2.3 置信度评估:检索决策的核心指标

置信度是检索决策最重要的信号。但 LLM 的置信度评估是一个开放问题------模型没有像传统分类器那样的概率输出。实践中,有几种常见的置信度评估方法:

python 复制代码
import json
from typing import Optional

class ConfidenceEstimator:
    """LLM 置信度评估器"""
    
    def __init__(self, llm_client):
        self.llm = llm_client
    
    def estimate_via_self_assessment(self, query: str, draft_answer: str) -> float:
        """方法一:自我评估法 - 让模型评估自己的答案"""
        prompt = f"""请评估你对以下问题的回答的置信度(0.0-1.0)。
        
        问题:{query}
        回答:{draft_answer}
        
        评估维度:
        1. 事实准确性:你确定这些事实是正确的吗?
        2. 完整性:回答是否涵盖了问题的所有方面?
        3. 时效性:信息是否可能过时?
        
        返回JSON格式:{{"confidence": 0.85, "reasoning": "简要说明"}}
        """
        response = self.llm.generate(prompt)
        result = json.loads(response)
        return result.get("confidence", 0.5)
    
    def estimate_via_uncertainty_markers(self, answer: str) -> float:
        """方法二:不确定性标记法 - 通过语言特征推断"""
        uncertainty_markers = [
            "可能", "也许", "大概", "似乎", "据我所知",
            "我不确定", "可能不准确", "建议查证",
            "approximately", "roughly", "i think", "maybe"
        ]
        marker_count = sum(1 for marker in uncertainty_markers if marker in answer.lower())
        
        # 标记越多,置信度越低
        base_confidence = 0.9
        penalty = marker_count * 0.15
        return max(0.2, base_confidence - penalty)
    
    def estimate_via_consistency(self, query: str, n: int = 3) -> float:
        """方法三:一致性法 - 多次采样计算一致性"""
        answers = []
        for _ in range(n):
            answer = self.llm.generate(query, temperature=0.7)
            answers.append(answer)
        
        # 计算答案间的语义相似度
        # 简化实现:使用关键词重叠度
        keyword_sets = [set(a.split()) for a in answers]
        overlaps = []
        for i in range(len(keyword_sets)):
            for j in range(i+1, len(keyword_sets)):
                overlap = len(keyword_sets[i] & keyword_sets[j]) / len(keyword_sets[i] | keyword_sets[j])
                overlaps.append(overlap)
        
        avg_overlap = sum(overlaps) / len(overlaps) if overlaps else 0
        return avg_overlap

代码解释 :这段代码实现了三种置信度评估方法。estimate_via_self_assessment 通过让模型自我评估来获取置信度,这是最直观但也是最容易过度自信的方法。estimate_via_uncertainty_markers 通过检测答案中的不确定性语言标记来推断置信度,简单但粗糙。estimate_via_consistency 通过多次采样计算答案一致性,理论上最可靠但成本最高(需要多次 LLM 调用)。工程实践中,通常组合使用方法一和方法二:先用不确定性标记做初步筛选,再对边缘案例使用自我评估。

2.4 检索量的动态控制

python 复制代码
from dataclasses import dataclass
from typing import List, Dict

@dataclass
class RetrievalConfig:
    """检索配置:根据查询复杂度动态调整"""
    
    # 基础参数
    top_k: int = 5                    # 检索结果数量
    min_score: float = 0.7            # 最低相关度阈值
    max_tokens: int = 4000            # 检索内容最大token数
    
    # 自适应参数
    query_complexity: str = "medium"  # simple/medium/complex
    enable_rerank: bool = True        # 是否启用重排序
    enable_multi_query: bool = False # 是否启用多查询扩展


def determine_retrieval_config(query: str, history_depth: int) -> RetrievalConfig:
    """根据查询特征确定检索配置"""
    
    config = RetrievalConfig()
    
    # 评估查询复杂度
    query_length = len(query)
    has_comparison = any(kw in query for kw in ["比较", "对比", "区别", "vs", "相比"])
    has_multi_questions = query.count("?") > 1 or query.count("?") > 1
    has_analysis = any(kw in query for kw in ["分析", "为什么", "如何", "原理"])
    
    if has_comparison or has_multi_questions:
        config.query_complexity = "complex"
        config.top_k = 8
        config.max_tokens = 6000
        config.enable_multi_query = True
    elif has_analysis or query_length > 100:
        config.query_complexity = "medium"
        config.top_k = 5
        config.max_tokens = 4000
    else:
        config.query_complexity = "simple"
        config.top_k = 3
        config.max_tokens = 2000
        config.enable_rerank = False
    
    # 历史上下文丰富时减少检索量
    if history_depth > 10:
        config.top_k = max(2, config.top_k - 2)
        config.max_tokens = int(config.max_tokens * 0.8)
    
    return config

代码解释 :RetrievalConfig 定义了检索的配置参数,determine_retrieval_config 函数根据查询的复杂度动态调整检索策略。对于比较类、多问题类查询(复杂),增加检索量和多查询扩展;对于简单事实查询,减少检索量以降低延迟。同时考虑了对话历史深度------如果上下文已经很丰富,适当减少检索量以避免信息冗余。这种动态调整策略在实践中能有效平衡回答质量和响应速度。

图:检索决策三层判断框架:是否需要检索→查询改写策略选择→检索方法选择,含置信度分支和上下文充分性检查

2.5 检索决策的流程图

#mermaid-svg-8LIst42wcc1C3oQj{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;fill:#333;}@keyframes edge-animation-frame{from{stroke-dashoffset:0;}}@keyframes dash{to{stroke-dashoffset:0;}}#mermaid-svg-8LIst42wcc1C3oQj .edge-animation-slow{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 50s linear infinite;stroke-linecap:round;}#mermaid-svg-8LIst42wcc1C3oQj .edge-animation-fast{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 20s linear infinite;stroke-linecap:round;}#mermaid-svg-8LIst42wcc1C3oQj .error-icon{fill:#552222;}#mermaid-svg-8LIst42wcc1C3oQj .error-text{fill:#552222;stroke:#552222;}#mermaid-svg-8LIst42wcc1C3oQj .edge-thickness-normal{stroke-width:1px;}#mermaid-svg-8LIst42wcc1C3oQj .edge-thickness-thick{stroke-width:3.5px;}#mermaid-svg-8LIst42wcc1C3oQj .edge-pattern-solid{stroke-dasharray:0;}#mermaid-svg-8LIst42wcc1C3oQj .edge-thickness-invisible{stroke-width:0;fill:none;}#mermaid-svg-8LIst42wcc1C3oQj .edge-pattern-dashed{stroke-dasharray:3;}#mermaid-svg-8LIst42wcc1C3oQj .edge-pattern-dotted{stroke-dasharray:2;}#mermaid-svg-8LIst42wcc1C3oQj .marker{fill:#333333;stroke:#333333;}#mermaid-svg-8LIst42wcc1C3oQj .marker.cross{stroke:#333333;}#mermaid-svg-8LIst42wcc1C3oQj svg{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;}#mermaid-svg-8LIst42wcc1C3oQj p{margin:0;}#mermaid-svg-8LIst42wcc1C3oQj .label{font-family:"trebuchet ms",verdana,arial,sans-serif;color:#333;}#mermaid-svg-8LIst42wcc1C3oQj .cluster-label text{fill:#333;}#mermaid-svg-8LIst42wcc1C3oQj .cluster-label span{color:#333;}#mermaid-svg-8LIst42wcc1C3oQj .cluster-label span p{background-color:transparent;}#mermaid-svg-8LIst42wcc1C3oQj .label text,#mermaid-svg-8LIst42wcc1C3oQj span{fill:#333;color:#333;}#mermaid-svg-8LIst42wcc1C3oQj .node rect,#mermaid-svg-8LIst42wcc1C3oQj .node circle,#mermaid-svg-8LIst42wcc1C3oQj .node ellipse,#mermaid-svg-8LIst42wcc1C3oQj .node polygon,#mermaid-svg-8LIst42wcc1C3oQj .node path{fill:#ECECFF;stroke:#9370DB;stroke-width:1px;}#mermaid-svg-8LIst42wcc1C3oQj .rough-node .label text,#mermaid-svg-8LIst42wcc1C3oQj .node .label text,#mermaid-svg-8LIst42wcc1C3oQj .image-shape .label,#mermaid-svg-8LIst42wcc1C3oQj .icon-shape .label{text-anchor:middle;}#mermaid-svg-8LIst42wcc1C3oQj .node .katex path{fill:#000;stroke:#000;stroke-width:1px;}#mermaid-svg-8LIst42wcc1C3oQj .rough-node .label,#mermaid-svg-8LIst42wcc1C3oQj .node .label,#mermaid-svg-8LIst42wcc1C3oQj .image-shape .label,#mermaid-svg-8LIst42wcc1C3oQj .icon-shape .label{text-align:center;}#mermaid-svg-8LIst42wcc1C3oQj .node.clickable{cursor:pointer;}#mermaid-svg-8LIst42wcc1C3oQj .root .anchor path{fill:#333333!important;stroke-width:0;stroke:#333333;}#mermaid-svg-8LIst42wcc1C3oQj .arrowheadPath{fill:#333333;}#mermaid-svg-8LIst42wcc1C3oQj .edgePath .path{stroke:#333333;stroke-width:2.0px;}#mermaid-svg-8LIst42wcc1C3oQj .flowchart-link{stroke:#333333;fill:none;}#mermaid-svg-8LIst42wcc1C3oQj .edgeLabel{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-8LIst42wcc1C3oQj .edgeLabel p{background-color:rgba(232,232,232, 0.8);}#mermaid-svg-8LIst42wcc1C3oQj .edgeLabel rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-8LIst42wcc1C3oQj .labelBkg{background-color:rgba(232, 232, 232, 0.5);}#mermaid-svg-8LIst42wcc1C3oQj .cluster rect{fill:#ffffde;stroke:#aaaa33;stroke-width:1px;}#mermaid-svg-8LIst42wcc1C3oQj .cluster text{fill:#333;}#mermaid-svg-8LIst42wcc1C3oQj .cluster span{color:#333;}#mermaid-svg-8LIst42wcc1C3oQj div.mermaidTooltip{position:absolute;text-align:center;max-width:200px;padding:2px;font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:12px;background:hsl(80, 100%, 96.2745098039%);border:1px solid #aaaa33;border-radius:2px;pointer-events:none;z-index:100;}#mermaid-svg-8LIst42wcc1C3oQj .flowchartTitleText{text-anchor:middle;font-size:18px;fill:#333;}#mermaid-svg-8LIst42wcc1C3oQj rect.text{fill:none;stroke-width:0;}#mermaid-svg-8LIst42wcc1C3oQj .icon-shape,#mermaid-svg-8LIst42wcc1C3oQj .image-shape{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-8LIst42wcc1C3oQj .icon-shape p,#mermaid-svg-8LIst42wcc1C3oQj .image-shape p{background-color:rgba(232,232,232, 0.8);padding:2px;}#mermaid-svg-8LIst42wcc1C3oQj .icon-shape .label rect,#mermaid-svg-8LIst42wcc1C3oQj .image-shape .label rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-8LIst42wcc1C3oQj .label-icon{display:inline-block;height:1em;overflow:visible;vertical-align:-0.125em;}#mermaid-svg-8LIst42wcc1C3oQj .node .label-icon path{fill:currentColor;stroke:revert;stroke-width:revert;}#mermaid-svg-8LIst42wcc1C3oQj :root{--mermaid-font-family:"trebuchet ms",verdana,arial,sans-serif;} >0.85 高置信
0.6-0.85 中等
<0.6 低置信
是
否
是
否
相关度>0.7
相关度<0.7
用户查询
置信度评估
直接回答
时间敏感?
执行检索
上下文充分?
检索配置
检索执行
结果质量评估
使用检索结果
扩大检索范围
生成回答


三、工具选择决策:多个工具时如何选择

3.1 工具选择的复杂性

当 Agent 只有一个工具时,决策很简单------用或不用。但实际系统中,Agent 往往配置了多个工具:搜索引擎、代码执行器、计算器、知识库检索、API 调用等。选择哪个工具,或以什么顺序组合使用多个工具,是一个复杂的决策问题。

工具选择的核心挑战包括:

  1. 功能重叠:搜索引擎和知识库都能回答"什么是RAG",选哪个?
  2. 成本差异:搜索引擎调用免费,专业API每次$0.05,如何权衡?
  3. 延迟差异:本地计算器秒回,搜索引擎3秒,知识库检索500ms
  4. 质量差异:不同工具对不同类型问题的回答质量不同
  5. 组合使用:有些问题需要先搜索再计算,有些需要先检索再总结

3.2 工具选择策略

python 复制代码
from typing import List, Dict, Optional
from dataclasses import dataclass, field
from enum import Enum

class ToolCategory(Enum):
    SEARCH = "search"           # 搜索类工具
    KNOWLEDGE = "knowledge"     # 知识库类工具
    COMPUTE = "compute"        # 计算类工具
    CODE = "code"             # 代码执行类工具
    API = "api"               # 外部API类工具

@dataclass
class ToolProfile:
    """工具能力画像"""
    name: str
    category: ToolCategory
    description: str
    latency_ms: int              # 平均延迟(毫秒)
    cost_per_call: float         # 每次调用成本(美元)
    accuracy_score: float        # 准确度评分(0-1)
    best_for: List[str]          # 最适合的场景
    limitations: List[str]        # 已知限制

class ToolSelector:
    """多工具智能选择器"""
    
    def __init__(self, tools: List[ToolProfile]):
        self.tools = {t.name: t for t in tools}
        self.tools_by_category = {}
        for t in tools:
            self.tools_by_category.setdefault(t.category, []).append(t)
    
    def select(self, query: str, intent: str, budget: float = 0.5) -> List[str]:
        """选择最适合的工具组合"""
        
        # Step 1: 根据意图过滤工具类别
        candidate_categories = self._intent_to_categories(intent)
        
        # Step 2: 在候选类别中选择最优工具
        candidates = []
        for cat in candidate_categories:
            tools_in_cat = self.tools_by_category.get(cat, [])
            if not tools_in_cat:
                continue
            
            # 按性价比排序:准确度/延迟/成本的综合评分
            best = max(tools_in_cat, key=lambda t: self._score_tool(t, query))
            candidates.append(best)
        
        # Step 3: 预算过滤
        selected = []
        remaining_budget = budget
        for tool in candidates:
            if tool.cost_per_call <= remaining_budget:
                selected.append(tool.name)
                remaining_budget -= tool.cost_per_call
        
        return selected if selected else ["direct_answer"]
    
    def _intent_to_categories(self, intent: str) -> List[ToolCategory]:
        """将查询意图映射到工具类别"""
        intent_map = {
            "factual_lookup": [ToolCategory.KNOWLEDGE, ToolCategory.SEARCH],
            "realtime_info": [ToolCategory.SEARCH],
            "calculation": [ToolCategory.COMPUTE, ToolCategory.CODE],
            "data_analysis": [ToolCategory.CODE, ToolCategory.API],
            "knowledge_synthesis": [ToolCategory.KNOWLEDGE, ToolCategory.SEARCH],
            "code_execution": [ToolCategory.CODE],
        }
        return intent_map.get(intent, [ToolCategory.SEARCH])
    
    def _score_tool(self, tool: ToolProfile, query: str) -> float:
        """综合评分工具适配度"""
        accuracy = tool.accuracy_score * 0.5
        latency_score = (1 - min(tool.latency_ms / 5000, 1)) * 0.3
        cost_score = (1 - min(tool.cost_per_call / 0.1, 1)) * 0.2
        return accuracy + latency_score + cost_score

代码解释 :ToolSelector 实现了一个多维度工具选择器。首先通过 _intent_to_categories 将用户意图映射到候选工具类别(如"事实查询"映射到知识库和搜索),然后在每个类别中通过 _score_tool 选择综合评分最高的工具(评分权重:准确度50%、延迟30%、成本20%),最后通过预算过滤确保总成本不超过预设上限。这种分层选择策略既保证了工具选择的准确性,又控制了调用成本。

3.3 工具选择的决策矩阵

查询意图 首选工具 备选工具 选择依据
实时信息查询 搜索引擎 API调用 时效性优先
事实核查 知识库 搜索引擎 准确性优先
数学计算 计算器 代码执行 延迟优先
数据分析 代码执行 API调用 功能优先
知识综合 知识库+搜索 --- 多源互补
代码运行 代码执行 --- 唯一选择

3.4 工具链编排:多工具组合决策

有些问题不是单个工具能解决的,需要多个工具按特定顺序协作:
代码执行器 知识库 搜索引擎 Agent 用户 代码执行器 知识库 搜索引擎 Agent 用户 #mermaid-svg-XeA1VcYbc6nwljdN{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;fill:#333;}@keyframes edge-animation-frame{from{stroke-dashoffset:0;}}@keyframes dash{to{stroke-dashoffset:0;}}#mermaid-svg-XeA1VcYbc6nwljdN .edge-animation-slow{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 50s linear infinite;stroke-linecap:round;}#mermaid-svg-XeA1VcYbc6nwljdN .edge-animation-fast{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 20s linear infinite;stroke-linecap:round;}#mermaid-svg-XeA1VcYbc6nwljdN .error-icon{fill:#552222;}#mermaid-svg-XeA1VcYbc6nwljdN .error-text{fill:#552222;stroke:#552222;}#mermaid-svg-XeA1VcYbc6nwljdN .edge-thickness-normal{stroke-width:1px;}#mermaid-svg-XeA1VcYbc6nwljdN .edge-thickness-thick{stroke-width:3.5px;}#mermaid-svg-XeA1VcYbc6nwljdN .edge-pattern-solid{stroke-dasharray:0;}#mermaid-svg-XeA1VcYbc6nwljdN .edge-thickness-invisible{stroke-width:0;fill:none;}#mermaid-svg-XeA1VcYbc6nwljdN .edge-pattern-dashed{stroke-dasharray:3;}#mermaid-svg-XeA1VcYbc6nwljdN .edge-pattern-dotted{stroke-dasharray:2;}#mermaid-svg-XeA1VcYbc6nwljdN .marker{fill:#333333;stroke:#333333;}#mermaid-svg-XeA1VcYbc6nwljdN .marker.cross{stroke:#333333;}#mermaid-svg-XeA1VcYbc6nwljdN svg{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;}#mermaid-svg-XeA1VcYbc6nwljdN p{margin:0;}#mermaid-svg-XeA1VcYbc6nwljdN .actor{stroke:hsl(259.6261682243, 59.7765363128%, 87.9019607843%);fill:#ECECFF;}#mermaid-svg-XeA1VcYbc6nwljdN text.actor>tspan{fill:black;stroke:none;}#mermaid-svg-XeA1VcYbc6nwljdN .actor-line{stroke:hsl(259.6261682243, 59.7765363128%, 87.9019607843%);}#mermaid-svg-XeA1VcYbc6nwljdN .innerArc{stroke-width:1.5;stroke-dasharray:none;}#mermaid-svg-XeA1VcYbc6nwljdN .messageLine0{stroke-width:1.5;stroke-dasharray:none;stroke:#333;}#mermaid-svg-XeA1VcYbc6nwljdN .messageLine1{stroke-width:1.5;stroke-dasharray:2,2;stroke:#333;}#mermaid-svg-XeA1VcYbc6nwljdN #arrowhead path{fill:#333;stroke:#333;}#mermaid-svg-XeA1VcYbc6nwljdN .sequenceNumber{fill:white;}#mermaid-svg-XeA1VcYbc6nwljdN #sequencenumber{fill:#333;}#mermaid-svg-XeA1VcYbc6nwljdN #crosshead path{fill:#333;stroke:#333;}#mermaid-svg-XeA1VcYbc6nwljdN .messageText{fill:#333;stroke:none;}#mermaid-svg-XeA1VcYbc6nwljdN .labelBox{stroke:hsl(259.6261682243, 59.7765363128%, 87.9019607843%);fill:#ECECFF;}#mermaid-svg-XeA1VcYbc6nwljdN .labelText,#mermaid-svg-XeA1VcYbc6nwljdN .labelText>tspan{fill:black;stroke:none;}#mermaid-svg-XeA1VcYbc6nwljdN .loopText,#mermaid-svg-XeA1VcYbc6nwljdN .loopText>tspan{fill:black;stroke:none;}#mermaid-svg-XeA1VcYbc6nwljdN .loopLine{stroke-width:2px;stroke-dasharray:2,2;stroke:hsl(259.6261682243, 59.7765363128%, 87.9019607843%);fill:hsl(259.6261682243, 59.7765363128%, 87.9019607843%);}#mermaid-svg-XeA1VcYbc6nwljdN .note{stroke:#aaaa33;fill:#fff5ad;}#mermaid-svg-XeA1VcYbc6nwljdN .noteText,#mermaid-svg-XeA1VcYbc6nwljdN .noteText>tspan{fill:black;stroke:none;}#mermaid-svg-XeA1VcYbc6nwljdN .activation0{fill:#f4f4f4;stroke:#666;}#mermaid-svg-XeA1VcYbc6nwljdN .activation1{fill:#f4f4f4;stroke:#666;}#mermaid-svg-XeA1VcYbc6nwljdN .activation2{fill:#f4f4f4;stroke:#666;}#mermaid-svg-XeA1VcYbc6nwljdN .actorPopupMenu{position:absolute;}#mermaid-svg-XeA1VcYbc6nwljdN .actorPopupMenuPanel{position:absolute;fill:#ECECFF;box-shadow:0px 8px 16px 0px rgba(0,0,0,0.2);filter:drop-shadow(3px 5px 2px rgb(0 0 0 / 0.4));}#mermaid-svg-XeA1VcYbc6nwljdN .actor-man line{stroke:hsl(259.6261682243, 59.7765363128%, 87.9019607843%);fill:#ECECFF;}#mermaid-svg-XeA1VcYbc6nwljdN .actor-man circle,#mermaid-svg-XeA1VcYbc6nwljdN line{stroke:hsl(259.6261682243, 59.7765363128%, 87.9019607843%);fill:#ECECFF;stroke-width:2px;}#mermaid-svg-XeA1VcYbc6nwljdN :root{--mermaid-font-family:"trebuchet ms",verdana,arial,sans-serif;} "分析公司2024年营收增长率" 决策:需要搜索+计算 搜索"公司2024年营收" 返回营收数据 检索"公司2023年营收"(历史数据) 返回历史数据 决策:数据完整,可以计算 执行增长率计算代码 返回计算结果 "2024年营收增长率为15.3%..."

这个时序图展示了一个典型的多工具组合场景:Agent 首先通过搜索引擎获取最新数据,再从知识库获取历史数据,最后用代码执行器进行计算。每一步之间的决策------"数据是否完整"、"是否需要更多信息"------都是动态的。


四、停止决策:何时停止循环、何时输出最终结果

4.1 停止决策的重要性

Agent 系统通常采用循环(ReAct、Plan-Execute 等)架构:观察-思考-行动-观察-思考-行动......直到得出最终答案。这个循环何时终止,直接影响系统的效率和质量:

  • 过早停止:信息不充分,答案质量差
  • 过晚停止:浪费资源,增加延迟,可能陷入死循环

停止决策就是在"够了"和"还不够"之间找到正确的平衡点。

4.2 停止条件的设计

python 复制代码
from dataclasses import dataclass
from typing import Optional, List

@dataclass
class StopDecision:
    """停止决策结果"""
    should_stop: bool
    reason: str
    confidence: float
    next_action: Optional[str] = None  # 如果不停止,下一步做什么

class StopDecisionMaker:
    """停止决策器"""
    
    MAX_ITERATIONS = 10              # 硬性上限
    MAX_TOKENS = 50000               # Token预算上限
    MAX_LATENCY_MS = 30000           # 延迟上限
    QUALITY_THRESHOLD = 0.85         # 质量阈值
    
    def evaluate(self, state: dict) -> StopDecision:
        """评估是否应该停止"""
        
        # 硬性停止条件
        if state["iteration"] >= self.MAX_ITERATIONS:
            return StopDecision(True, "达到最大迭代次数", 1.0, "force_stop")
        
        if state["total_tokens"] >= self.MAX_TOKENS:
            return StopDecision(True, "达到Token预算上限", 1.0, "force_stop")
        
        if state["elapsed_ms"] >= self.MAX_LATENCY_MS:
            return StopDecision(True, "达到延迟上限", 1.0, "force_stop")
        
        # 质量停止条件
        answer_completeness = self._assess_completeness(state)
        if answer_completeness >= self.QUALITY_THRESHOLD:
            return StopDecision(True, "答案质量达标", answer_completeness, "final_answer")
        
        # 循环检测:如果最近3次行动相似,可能陷入循环
        if self._detect_loop(state["action_history"][-3:]):
            return StopDecision(True, "检测到行动循环", 0.7, "force_stop")
        
        # 信息充分性检查
        if self._info_sufficient(state):
            return StopDecision(True, "信息已充分", 0.9, "final_answer")
        
        # 继续执行
        return StopDecision(False, "需要更多信息", 0.3, state.get("next_hint", "continue"))
    
    def _assess_completeness(self, state: dict) -> float:
        """评估答案完整性"""
        prompt = f"""评估当前信息是否足以完整回答用户的问题。
        
        用户问题:{state['query']}
        当前收集的信息:{state['collected_info']}
        
        评分维度(0-1):
        1. 事实覆盖度:是否涵盖了问题的关键事实?
        2. 逻辑完整性:论证链是否完整?
        3. 细节充分性:是否有足够的细节支撑?
        
        返回平均分。"""
        # 简化:实际中调用LLM评估
        return 0.87  # 示例值
    
    def _detect_loop(self, recent_actions: List[str]) -> bool:
        """检测行动循环"""
        if len(recent_actions) < 3:
            return False
        # 如果最近3次行动完全相同,判定为循环
        if len(set(recent_actions)) == 1:
            return True
        # 如果行动模式重复(A-B-A-B)
        if recent_actions[0] == recent_actions[2] and recent_actions[1] != recent_actions[0]:
            return True
        return False
    
    def _info_sufficient(self, state: dict) -> bool:
        """信息充分性检查"""
        # 检查是否已有足够的信息片段
        info_pieces = state.get("collected_info", [])
        if len(info_pieces) >= 3:  # 至少3个信息片段
            # 检查信息是否覆盖了问题的关键方面
            query_aspects = state.get("query_aspects", [])
            covered = sum(1 for aspect in query_aspects 
                         if any(aspect in info for info in info_pieces))
            coverage = covered / len(query_aspects) if query_aspects else 0
            return coverage >= 0.8
        return False

代码解释 :StopDecisionMaker 实现了一个多维度的停止决策器。硬性条件(迭代次数、Token预算、延迟上限)是不可逾越的安全网,防止 Agent 无限运行。质量条件通过 _assess_completeness 评估答案的完整性,当完整性达到阈值时主动停止。_detect_loop 通过分析最近的行动历史检测循环模式,防止 Agent 在相同操作间反复跳转。_info_sufficient 通过信息覆盖率判断是否已经收集了足够的信息来回答用户问题的各个方面。这种多层次的停止条件设计确保了 Agent 既不会过早终止,也不会无限运行。

4.3 停止决策的状态机

#mermaid-svg-ebDWo8jf3tl5Gd5i{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;fill:#333;}@keyframes edge-animation-frame{from{stroke-dashoffset:0;}}@keyframes dash{to{stroke-dashoffset:0;}}#mermaid-svg-ebDWo8jf3tl5Gd5i .edge-animation-slow{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 50s linear infinite;stroke-linecap:round;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edge-animation-fast{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 20s linear infinite;stroke-linecap:round;}#mermaid-svg-ebDWo8jf3tl5Gd5i .error-icon{fill:#552222;}#mermaid-svg-ebDWo8jf3tl5Gd5i .error-text{fill:#552222;stroke:#552222;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edge-thickness-normal{stroke-width:1px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edge-thickness-thick{stroke-width:3.5px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edge-pattern-solid{stroke-dasharray:0;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edge-thickness-invisible{stroke-width:0;fill:none;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edge-pattern-dashed{stroke-dasharray:3;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edge-pattern-dotted{stroke-dasharray:2;}#mermaid-svg-ebDWo8jf3tl5Gd5i .marker{fill:#333333;stroke:#333333;}#mermaid-svg-ebDWo8jf3tl5Gd5i .marker.cross{stroke:#333333;}#mermaid-svg-ebDWo8jf3tl5Gd5i svg{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;}#mermaid-svg-ebDWo8jf3tl5Gd5i p{margin:0;}#mermaid-svg-ebDWo8jf3tl5Gd5i defs #statediagram-barbEnd{fill:#333333;stroke:#333333;}#mermaid-svg-ebDWo8jf3tl5Gd5i g.stateGroup text{fill:#9370DB;stroke:none;font-size:10px;}#mermaid-svg-ebDWo8jf3tl5Gd5i g.stateGroup text{fill:#333;stroke:none;font-size:10px;}#mermaid-svg-ebDWo8jf3tl5Gd5i g.stateGroup .state-title{font-weight:bolder;fill:#131300;}#mermaid-svg-ebDWo8jf3tl5Gd5i g.stateGroup rect{fill:#ECECFF;stroke:#9370DB;}#mermaid-svg-ebDWo8jf3tl5Gd5i g.stateGroup line{stroke:#333333;stroke-width:1;}#mermaid-svg-ebDWo8jf3tl5Gd5i .transition{stroke:#333333;stroke-width:1;fill:none;}#mermaid-svg-ebDWo8jf3tl5Gd5i .stateGroup .composit{fill:white;border-bottom:1px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .stateGroup .alt-composit{fill:#e0e0e0;border-bottom:1px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .state-note{stroke:#aaaa33;fill:#fff5ad;}#mermaid-svg-ebDWo8jf3tl5Gd5i .state-note text{fill:black;stroke:none;font-size:10px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .stateLabel .box{stroke:none;stroke-width:0;fill:#ECECFF;opacity:0.5;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edgeLabel .label rect{fill:#ECECFF;opacity:0.5;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edgeLabel{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-ebDWo8jf3tl5Gd5i .edgeLabel p{background-color:rgba(232,232,232, 0.8);}#mermaid-svg-ebDWo8jf3tl5Gd5i .edgeLabel rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-ebDWo8jf3tl5Gd5i .edgeLabel .label text{fill:#333;}#mermaid-svg-ebDWo8jf3tl5Gd5i .label div .edgeLabel{color:#333;}#mermaid-svg-ebDWo8jf3tl5Gd5i .stateLabel text{fill:#131300;font-size:10px;font-weight:bold;}#mermaid-svg-ebDWo8jf3tl5Gd5i .node circle.state-start{fill:#333333;stroke:#333333;}#mermaid-svg-ebDWo8jf3tl5Gd5i .node .fork-join{fill:#333333;stroke:#333333;}#mermaid-svg-ebDWo8jf3tl5Gd5i .node circle.state-end{fill:#9370DB;stroke:white;stroke-width:1.5;}#mermaid-svg-ebDWo8jf3tl5Gd5i .end-state-inner{fill:white;stroke-width:1.5;}#mermaid-svg-ebDWo8jf3tl5Gd5i .node rect{fill:#ECECFF;stroke:#9370DB;stroke-width:1px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .node polygon{fill:#ECECFF;stroke:#9370DB;stroke-width:1px;}#mermaid-svg-ebDWo8jf3tl5Gd5i #statediagram-barbEnd{fill:#333333;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-cluster rect{fill:#ECECFF;stroke:#9370DB;stroke-width:1px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .cluster-label,#mermaid-svg-ebDWo8jf3tl5Gd5i .nodeLabel{color:#131300;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-cluster rect.outer{rx:5px;ry:5px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-state .divider{stroke:#9370DB;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-state .title-state{rx:5px;ry:5px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-cluster.statediagram-cluster .inner{fill:white;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-cluster.statediagram-cluster-alt .inner{fill:#f0f0f0;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-cluster .inner{rx:0;ry:0;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-state rect.basic{rx:5px;ry:5px;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-state rect.divider{stroke-dasharray:10,10;fill:#f0f0f0;}#mermaid-svg-ebDWo8jf3tl5Gd5i .note-edge{stroke-dasharray:5;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-note rect{fill:#fff5ad;stroke:#aaaa33;stroke-width:1px;rx:0;ry:0;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-note rect{fill:#fff5ad;stroke:#aaaa33;stroke-width:1px;rx:0;ry:0;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-note text{fill:black;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram-note .nodeLabel{color:black;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagram .edgeLabel{color:red;}#mermaid-svg-ebDWo8jf3tl5Gd5i #dependencyStart,#mermaid-svg-ebDWo8jf3tl5Gd5i #dependencyEnd{fill:#333333;stroke:#333333;stroke-width:1;}#mermaid-svg-ebDWo8jf3tl5Gd5i .statediagramTitleText{text-anchor:middle;font-size:18px;fill:#333;}#mermaid-svg-ebDWo8jf3tl5Gd5i :root{--mermaid-font-family:"trebuchet ms",verdana,arial,sans-serif;} 有可行计划
无法制定计划
工具执行完成
执行出错
信息充分
需要更多操作
信息部分可用
优化后达标
需要补充信息
降级成功
无法恢复
输出最终答案
输出降级答案
Planning
Executing
Stopped
Evaluating
ErrorHandling
AnswerReady
Refining

这个状态机展示了 Agent 从规划到最终输出的完整生命周期。每个状态转换都是一个决策点:从 Planning 到 Executing 需要确认计划可行,从 Evaluating 到 AnswerReady 需要确认信息充分,从 ErrorHandling 到 Executing 需要确认降级方案有效。


五、降级决策:出错时重试还是换方案

5.1 错误处理的工程现实

在理想的 Agent 系统中,每一步都能顺利执行。但在真实环境中,错误无处不在:API 超时、模型幻觉、检索结果不相关、代码执行报错......降级决策就是在这些错误发生时,决定"再试一次"还是"换个方案"。

好的降级策略需要回答三个问题:

  1. 这个错误是暂时的还是永久的?
  2. 重试有没有可能成功?
  3. 如果重试不行,有没有替代方案?

5.2 降级策略矩阵

错误类型 临时性 重试价值 降级策略 典型示例
网络超时 是 高 指数退避重试 API调用超时
限流(429) 是 中 延迟后重试/换备用Key 搜索API限流
结果不相关 否 低 换查询/换工具 检索结果偏离主题
格式错误 否 中 修正提示词后重试 JSON解析失败
工具不可用 否 低 切换备用工具 API服务下线
模型幻觉 否 低 增加约束/换模型 生成不存在的事实

5.3 降级决策引擎实现

python 复制代码
import time
from typing import Any, Optional, Callable
from dataclasses import dataclass
from enum import Enum

class ErrorType(Enum):
    TIMEOUT = "timeout"
    RATE_LIMIT = "rate_limit"
    IRRELEVANT_RESULT = "irrelevant_result"
    FORMAT_ERROR = "format_error"
    TOOL_UNAVAILABLE = "tool_unavailable"
    HALLUCINATION = "hallucination"
    UNKNOWN = "unknown"

@dataclass
class FallbackAction:
    """降级行动"""
    action: str          # retry / switch_tool / simplify_query / give_up
    delay_ms: int        # 延迟时间
    params: dict         # 行动参数
    
class FallbackDecider:
    """降级决策器"""
    
    MAX_RETRIES = 3
    BASE_DELAY_MS = 1000
    
    def decide(self, error: Exception, error_type: ErrorType, 
               retry_count: int, context: dict) -> FallbackAction:
        """根据错误类型决定降级行动"""
        
        # 已经达到最大重试次数
        if retry_count >= self.MAX_RETRIES:
            return self._try_alternative(context)
        
        if error_type == ErrorType.TIMEOUT:
            # 超时:指数退避重试
            delay = self.BASE_DELAY_MS * (2 ** retry_count)
            return FallbackAction("retry", delay, {"timeout": context.get("timeout", 30) * 1.5})
        
        elif error_type == ErrorType.RATE_LIMIT:
            # 限流:等待后重试或切换到备用服务
            if retry_count < 2:
                retry_after = context.get("retry_after", 5)
                return FallbackAction("retry", retry_after * 1000, {})
            else:
                return FallbackAction("switch_tool", 0, {"preferred": "backup_search"})
        
        elif error_type == ErrorType.IRRELEVANT_RESULT:
            # 结果不相关:换查询方式
            return FallbackAction("simplify_query", 0, {
                "strategy": "keyword_focused",
                "remove_context": True
            })
        
        elif error_type == ErrorType.FORMAT_ERROR:
            # 格式错误:增强提示词
            return FallbackAction("retry", 0, {
                "enhanced_prompt": True,
                "format_hint": context.get("expected_format", "json")
            })
        
        elif error_type == ErrorType.TOOL_UNAVAILABLE:
            # 工具不可用:直接切换
            return FallbackAction("switch_tool", 0, context.get("alternatives", {}))
        
        elif error_type == ErrorType.HALLUCINATION:
            # 幻觉:增加约束重试
            return FallbackAction("retry", 0, {
                "add_constraints": True,
                "require_citation": True
            })
        
        else:
            # 未知错误:尝试一次后降级
            if retry_count == 0:
                return FallbackAction("retry", 2000, {})
            return self._try_alternative(context)
    
    def _try_alternative(self, context: dict) -> FallbackAction:
        """尝试替代方案"""
        alternatives = context.get("alternative_tools", [])
        if alternatives:
            return FallbackAction("switch_tool", 0, {"tool": alternatives[0]})
        
        # 最终降级:使用模型自身知识
        return FallbackAction("give_up", 0, {
            "use_internal_knowledge": True,
            "add_disclaimer": "此回答基于模型内部知识,未经外部验证"
        })

代码解释 :FallbackDecider 实现了一个基于错误类型的降级决策器。核心逻辑是对不同类型的错误采取不同的恢复策略:超时错误使用指数退避重试(delay 按 2 的幂次增长);限流错误先等待后重试,多次限流则切换备用服务;结果不相关时改变查询策略(如从语义检索切换到关键词检索);格式错误时增强提示词重新尝试;工具不可用直接切换到替代工具。当所有重试都失败后,_try_alternative 会尝试替代工具或最终降级到使用模型内部知识(并附加免责声明)。这种分类型的降级策略能最大化恢复成功率,同时避免无意义的重试。

5.4 降级决策的流程

#mermaid-svg-OxC2IMQvGphcozFZ{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;fill:#333;}@keyframes edge-animation-frame{from{stroke-dashoffset:0;}}@keyframes dash{to{stroke-dashoffset:0;}}#mermaid-svg-OxC2IMQvGphcozFZ .edge-animation-slow{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 50s linear infinite;stroke-linecap:round;}#mermaid-svg-OxC2IMQvGphcozFZ .edge-animation-fast{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 20s linear infinite;stroke-linecap:round;}#mermaid-svg-OxC2IMQvGphcozFZ .error-icon{fill:#552222;}#mermaid-svg-OxC2IMQvGphcozFZ .error-text{fill:#552222;stroke:#552222;}#mermaid-svg-OxC2IMQvGphcozFZ .edge-thickness-normal{stroke-width:1px;}#mermaid-svg-OxC2IMQvGphcozFZ .edge-thickness-thick{stroke-width:3.5px;}#mermaid-svg-OxC2IMQvGphcozFZ .edge-pattern-solid{stroke-dasharray:0;}#mermaid-svg-OxC2IMQvGphcozFZ .edge-thickness-invisible{stroke-width:0;fill:none;}#mermaid-svg-OxC2IMQvGphcozFZ .edge-pattern-dashed{stroke-dasharray:3;}#mermaid-svg-OxC2IMQvGphcozFZ .edge-pattern-dotted{stroke-dasharray:2;}#mermaid-svg-OxC2IMQvGphcozFZ .marker{fill:#333333;stroke:#333333;}#mermaid-svg-OxC2IMQvGphcozFZ .marker.cross{stroke:#333333;}#mermaid-svg-OxC2IMQvGphcozFZ svg{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;}#mermaid-svg-OxC2IMQvGphcozFZ p{margin:0;}#mermaid-svg-OxC2IMQvGphcozFZ .label{font-family:"trebuchet ms",verdana,arial,sans-serif;color:#333;}#mermaid-svg-OxC2IMQvGphcozFZ .cluster-label text{fill:#333;}#mermaid-svg-OxC2IMQvGphcozFZ .cluster-label span{color:#333;}#mermaid-svg-OxC2IMQvGphcozFZ .cluster-label span p{background-color:transparent;}#mermaid-svg-OxC2IMQvGphcozFZ .label text,#mermaid-svg-OxC2IMQvGphcozFZ span{fill:#333;color:#333;}#mermaid-svg-OxC2IMQvGphcozFZ .node rect,#mermaid-svg-OxC2IMQvGphcozFZ .node circle,#mermaid-svg-OxC2IMQvGphcozFZ .node ellipse,#mermaid-svg-OxC2IMQvGphcozFZ .node polygon,#mermaid-svg-OxC2IMQvGphcozFZ .node path{fill:#ECECFF;stroke:#9370DB;stroke-width:1px;}#mermaid-svg-OxC2IMQvGphcozFZ .rough-node .label text,#mermaid-svg-OxC2IMQvGphcozFZ .node .label text,#mermaid-svg-OxC2IMQvGphcozFZ .image-shape .label,#mermaid-svg-OxC2IMQvGphcozFZ .icon-shape .label{text-anchor:middle;}#mermaid-svg-OxC2IMQvGphcozFZ .node .katex path{fill:#000;stroke:#000;stroke-width:1px;}#mermaid-svg-OxC2IMQvGphcozFZ .rough-node .label,#mermaid-svg-OxC2IMQvGphcozFZ .node .label,#mermaid-svg-OxC2IMQvGphcozFZ .image-shape .label,#mermaid-svg-OxC2IMQvGphcozFZ .icon-shape .label{text-align:center;}#mermaid-svg-OxC2IMQvGphcozFZ .node.clickable{cursor:pointer;}#mermaid-svg-OxC2IMQvGphcozFZ .root .anchor path{fill:#333333!important;stroke-width:0;stroke:#333333;}#mermaid-svg-OxC2IMQvGphcozFZ .arrowheadPath{fill:#333333;}#mermaid-svg-OxC2IMQvGphcozFZ .edgePath .path{stroke:#333333;stroke-width:2.0px;}#mermaid-svg-OxC2IMQvGphcozFZ .flowchart-link{stroke:#333333;fill:none;}#mermaid-svg-OxC2IMQvGphcozFZ .edgeLabel{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-OxC2IMQvGphcozFZ .edgeLabel p{background-color:rgba(232,232,232, 0.8);}#mermaid-svg-OxC2IMQvGphcozFZ .edgeLabel rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-OxC2IMQvGphcozFZ .labelBkg{background-color:rgba(232, 232, 232, 0.5);}#mermaid-svg-OxC2IMQvGphcozFZ .cluster rect{fill:#ffffde;stroke:#aaaa33;stroke-width:1px;}#mermaid-svg-OxC2IMQvGphcozFZ .cluster text{fill:#333;}#mermaid-svg-OxC2IMQvGphcozFZ .cluster span{color:#333;}#mermaid-svg-OxC2IMQvGphcozFZ div.mermaidTooltip{position:absolute;text-align:center;max-width:200px;padding:2px;font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:12px;background:hsl(80, 100%, 96.2745098039%);border:1px solid #aaaa33;border-radius:2px;pointer-events:none;z-index:100;}#mermaid-svg-OxC2IMQvGphcozFZ .flowchartTitleText{text-anchor:middle;font-size:18px;fill:#333;}#mermaid-svg-OxC2IMQvGphcozFZ rect.text{fill:none;stroke-width:0;}#mermaid-svg-OxC2IMQvGphcozFZ .icon-shape,#mermaid-svg-OxC2IMQvGphcozFZ .image-shape{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-OxC2IMQvGphcozFZ .icon-shape p,#mermaid-svg-OxC2IMQvGphcozFZ .image-shape p{background-color:rgba(232,232,232, 0.8);padding:2px;}#mermaid-svg-OxC2IMQvGphcozFZ .icon-shape .label rect,#mermaid-svg-OxC2IMQvGphcozFZ .image-shape .label rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-OxC2IMQvGphcozFZ .label-icon{display:inline-block;height:1em;overflow:visible;vertical-align:-0.125em;}#mermaid-svg-OxC2IMQvGphcozFZ .node .label-icon path{fill:currentColor;stroke:revert;stroke-width:revert;}#mermaid-svg-OxC2IMQvGphcozFZ :root{--mermaid-font-family:"trebuchet ms",verdana,arial,sans-serif;} 是
否
是
是
否
否
是
否
是
否
执行操作
成功?
继续流程
错误类型分析
临时性错误?
重试次数<上限?
指数退避重试
切换方案
有替代方案?
最终降级
切换后成功?
使用内部知识+免责声明


图:降级决策完整流程:错误类型分析(临时/永久)→重试判断→方案切换→最终降级(内部知识+免责声明)

六、实战:一个完整的决策引擎实现

6.1 决策引擎架构

将前面讨论的所有决策类型整合到一个统一的框架中,我们得到一个完整的 Agent 决策引擎。这个引擎在每个 Agent 循环中被调用,决定下一步行动。

6.2 完整实现

python 复制代码
"""
Agent 决策引擎 - 完整实现
整合检索决策、工具选择、停止决策和降级决策
"""
import time
import json
from typing import Dict, List, Optional, Any, Tuple
from dataclasses import dataclass, field
from enum import Enum


# ============ 状态定义 ============

class DecisionType(Enum):
    RETRIEVE = "retrieve"
    USE_TOOL = "use_tool"
    ANSWER = "answer"
    RETRY = "retry"
    STOP = "stop"

@dataclass
class Decision:
    """决策结果"""
    type: DecisionType
    confidence: float
    reasoning: str
    action_params: Dict[str, Any] = field(default_factory=dict)
    fallback_plan: Optional[str] = None

@dataclass  
class AgentState:
    """Agent 运行状态"""
    query: str
    iteration: int = 0
    total_tokens: int = 0
    elapsed_ms: int = 0
    collected_info: List[str] = field(default_factory=list)
    action_history: List[str] = field(default_factory=list)
    error_count: int = 0
    last_error: Optional[str] = None
    confidence: float = 0.0
    intent: str = "unknown"


# ============ 决策引擎 ============

class DecisionEngine:
    """Agent 决策引擎"""
    
    # 配置参数
    CONFIDENCE_THRESHOLD = 0.75
    MAX_ITERATIONS = 8
    MAX_TOKENS = 40000
    MAX_LATENCY_MS = 25000
    MAX_ERRORS = 5
    QUALITY_THRESHOLD = 0.85
    
    def __init__(self, llm_client, tools: List[ToolProfile]):
        self.llm = llm_client
        self.confidence_estimator = ConfidenceEstimator(llm_client)
        self.tool_selector = ToolSelector(tools)
        self.stop_maker = StopDecisionMaker()
        self.fallback = FallbackDecider()
    
    def decide(self, state: AgentState) -> Decision:
        """主决策入口"""
        state.iteration += 1
        start_time = time.time()
        
        # 第一优先级:检查停止条件
        stop_check = self._check_stop(state)
        if stop_check:
            return stop_check
        
        # 第二优先级:检查错误恢复
        if state.error_count > 0 and state.last_error:
            fallback = self._check_fallback(state)
            if fallback:
                return fallback
        
        # 第三优先级:评估当前信息充分性
        confidence = self.confidence_estimator.estimate_via_uncertainty_markers(
            " ".join(state.collected_info)
        )
        state.confidence = confidence
        
        if confidence >= self.CONFIDENCE_THRESHOLD:
            # 信息充分,可以回答
            return Decision(
                type=DecisionType.ANSWER,
                confidence=confidence,
                reasoning=f"置信度{confidence:.2f}超过阈值{self.CONFIDENCE_THRESHOLD}",
                action_params={"use_info": state.collected_info}
            )
        
        # 第四优先级:需要获取更多信息
        # 决定是检索还是使用工具
        intent = self._classify_intent(state.query)
        state.intent = intent
        
        if intent in ["factual_lookup", "realtime_info", "knowledge_synthesis"]:
            # 需要检索
            retrieval_config = determine_retrieval_config(state.query, state.iteration)
            return Decision(
                type=DecisionType.RETRIEVE,
                confidence=confidence,
                reasoning=f"置信度{confidence:.2f}不足,需要检索补充信息",
                action_params={
                    "config": retrieval_config.__dict__,
                    "query": state.query
                }
            )
        else:
            # 需要使用工具
            tools = self.tool_selector.select(state.query, intent)
            return Decision(
                type=DecisionType.USE_TOOL,
                confidence=confidence,
                reasoning=f"意图为{intent},选择工具:{tools}",
                action_params={"tools": tools}
            )
    
    def _check_stop(self, state: AgentState) -> Optional[Decision]:
        """检查停止条件"""
        # 硬性停止
        if state.iteration >= self.MAX_ITERATIONS:
            return Decision(
                type=DecisionType.STOP,
                confidence=1.0,
                reasoning="达到最大迭代次数",
                action_params={"force": True}
            )
        
        if state.total_tokens >= self.MAX_TOKENS:
            return Decision(
                type=DecisionType.STOP,
                confidence=1.0,
                reasoning="达到Token预算上限",
                action_params={"force": True}
            )
        
        if state.elapsed_ms >= self.MAX_LATENCY_MS:
            return Decision(
                type=DecisionType.STOP,
                confidence=1.0,
                reasoning="达到延迟上限",
                action_params={"force": True}
            )
        
        if state.error_count >= self.MAX_ERRORS:
            return Decision(
                type=DecisionType.STOP,
                confidence=0.9,
                reasoning="错误次数过多",
                action_params={"force": True}
            )
        
        # 循环检测
        if len(state.action_history) >= 3:
            recent = state.action_history[-3:]
            if len(set(recent)) == 1:
                return Decision(
                    type=DecisionType.STOP,
                    confidence=0.8,
                    reasoning="检测到行动循环",
                    action_params={"force": True}
                )
        
        return None
    
    def _check_fallback(self, state: AgentState) -> Optional[Decision]:
        """检查降级策略"""
        error_type = self._classify_error(state.last_error)
        fallback_action = self.fallback.decide(
            error=Exception(state.last_error),
            error_type=error_type,
            retry_count=state.error_count,
            context={"query": state.query}
        )
        
        if fallback_action.action == "retry":
            return Decision(
                type=DecisionType.RETRY,
                confidence=0.5,
                reasoning=f"错误恢复:{fallback_action.action}",
                action_params={"delay": fallback_action.delay_ms, **fallback_action.params}
            )
        elif fallback_action.action == "switch_tool":
            return Decision(
                type=DecisionType.USE_TOOL,
                confidence=0.5,
                reasoning=f"工具切换:{fallback_action.params}",
                action_params=fallback_action.params
            )
        elif fallback_action.action == "give_up":
            return Decision(
                type=DecisionType.ANSWER,
                confidence=0.4,
                reasoning="降级为使用内部知识",
                action_params={
                    "use_internal_knowledge": True,
                    "disclaimer": "此回答基于模型内部知识,未经外部验证"
                }
            )
        
        return None
    
    def _classify_intent(self, query: str) -> str:
        """分类查询意图"""
        intent_rules = {
            "realtime_info": ["今天", "最新", "当前", "现在", "最近", "实时"],
            "calculation": ["计算", "求", "等于", "多少", "百分比", "增长率"],
            "data_analysis": ["分析", "统计", "趋势", "对比", "报表"],
            "factual_lookup": ["什么是", "是什么", "定义", "解释", "原理"],
            "knowledge_synthesis": ["比较", "总结", "概述", "区别", "优缺点"],
            "code_execution": ["运行代码", "执行", "写脚本", "编程"],
        }
        
        for intent, keywords in intent_rules.items():
            if any(kw in query for kw in keywords):
                return intent
        
        return "factual_lookup"  # 默认
    
    def _classify_error(self, error_msg: str) -> ErrorType:
        """分类错误类型"""
        error_lower = error_msg.lower()
        if "timeout" in error_lower or "timed out" in error_lower:
            return ErrorType.TIMEOUT
        if "rate limit" in error_lower or "429" in error_lower:
            return ErrorType.RATE_LIMIT
        if "not relevant" in error_lower or "irrelevant" in error_lower:
            return ErrorType.IRRELEVANT_RESULT
        if "format" in error_lower or "parse" in error_lower or "json" in error_lower:
            return ErrorType.FORMAT_ERROR
        if "unavailable" in error_lower or "connection" in error_lower:
            return ErrorType.TOOL_UNAVAILABLE
        return ErrorType.UNKNOWN

代码解释 :这是决策引擎的完整核心实现。DecisionEngine 类将四类决策(检索、工具选择、停止、降级)整合到一个统一的 decide 方法中。决策按优先级执行:首先检查硬性停止条件(防止无限循环),然后检查错误恢复(如果有错误需要先处理),接着评估信息充分性(决定是否可以输出答案),最后如果信息不足则决定是检索还是使用工具。_classify_intent 方法通过关键词匹配将查询分类到不同意图,_classify_error 将错误消息映射到错误类型枚举。这个设计的优势在于:所有决策逻辑集中管理、易于调试和修改,且每个决策都有明确的优先级和触发条件。

6.3 决策引擎的调用流程

python 复制代码
def run_agent_with_decision_engine(
    query: str, 
    engine: DecisionEngine,
    tool_executor: Callable,
    retrieval_executor: Callable
) -> str:
    """使用决策引擎运行Agent"""
    
    state = AgentState(query=query)
    start_time = time.time()
    
    while True:
        # 记录耗时
        state.elapsed_ms = int((time.time() - start_time) * 1000)
        
        # 获取决策
        decision = engine.decide(state)
        print(f"[迭代 {state.iteration}] 决策: {decision.type.value} "
              f"(置信度: {decision.confidence:.2f}, 原因: {decision.reasoning})")
        
        # 执行决策
        if decision.type == DecisionType.STOP:
            if state.collected_info:
                return f"基于已收集信息的回答:\n{''.join(state.collected_info)}"
            else:
                return "抱歉,无法处理您的请求。"
        
        elif decision.type == DecisionType.ANSWER:
            if decision.action_params.get("use_internal_knowledge"):
                disclaimer = decision.action_params.get("disclaimer", "")
                answer = engine.llm.generate(query)
                return f"{answer}\n\n⚠️ {disclaimer}"
            else:
                info = decision.action_params.get("use_info", [])
                return engine.llm.generate(f"{query}\n参考信息:{info}")
        
        elif decision.type == DecisionType.RETRIEVE:
            config = decision.action_params["config"]
            results = retrieval_executor(
                decision.action_params["query"],
                top_k=config["top_k"],
                min_score=config["min_score"]
            )
            state.collected_info.extend(results)
            state.action_history.append("retrieve")
        
        elif decision.type == DecisionType.USE_TOOL:
            tools = decision.action_params["tools"]
            if tools == ["direct_answer"]:
                answer = engine.llm.generate(query)
                return answer
            
            for tool_name in tools:
                try:
                    result = tool_executor(tool_name, query)
                    state.collected_info.append(result)
                    state.action_history.append(f"tool:{tool_name}")
                except Exception as e:
                    state.error_count += 1
                    state.last_error = str(e)
                    state.action_history.append(f"error:{tool_name}")
                    break  # 出错后让决策引擎决定下一步
        
        elif decision.type == DecisionType.RETRY:
            delay = decision.action_params.get("delay", 0)
            if delay > 0:
                time.sleep(delay / 1000)
            state.error_count = 0  # 重置错误计数
            state.last_error = None
    
    # 理论上不会到达这里
    return "处理超时"

代码解释 :run_agent_with_decision_engine 是决策引擎的运行循环。每一轮迭代中,先记录当前耗时,然后调用 engine.decide 获取决策。根据决策类型执行不同操作:STOP 时返回已收集的信息或错误提示;ANSWER 时调用 LLM 生成最终答案(如果是从降级路径来的,会附加免责声明);RETRIEVE 时调用检索执行器获取外部信息;USE_TOOL 时遍历选中的工具执行操作;RETRY 时等待指定时间后重置错误状态。这个设计的关键在于:决策和执行完全分离------决策引擎只负责"做什么",执行函数只负责"怎么做",两者通过 AgentState 共享状态。

6.4 决策日志与可观测性

工程实践中,决策的可观测性至关重要。每一步决策都应该记录:决策类型、置信度、推理过程、执行时间、结果状态。这些日志不仅用于调试,还用于后续优化决策参数。


七、适用边界与风险提示

7.1 适用场景

本文描述的决策框架适用于以下场景:

  • 知识密集型问答系统:需要从外部知识库获取信息并综合回答的场景
  • 多工具编排系统:配置了多个工具,需要根据查询动态选择的场景
  • 高可靠性要求场景:对错误恢复和降级有明确要求的场景
  • 长对话系统:需要管理多轮对话上下文的场景

7.2 不适用场景

  • 简单FAQ系统:固定问答映射,无需动态决策
  • 单工具系统:只有一个工具可用,无需工具选择
  • 实时性要求极高的系统:决策本身引入额外延迟
  • 预算极度受限的系统:多轮决策增加LLM调用成本

7.3 风险提示

过度自信风险:模型的自我评估置信度可能偏高,导致跳过必要的检索步骤。建议定期校准置信度阈值。

无限循环风险:即使有最大迭代次数保护,Agent 仍可能在"检索-评估-检索"的循环中浪费大量资源。建议设置 Token 预算和延迟上限作为双重保险。

降级传播风险:降级答案可能比正常答案质量低很多,如果用户没有意识到答案是降级的,可能造成误导。建议始终标注降级答案。

决策一致性风险:同样的输入可能因为模型采样的随机性导致不同的决策路径。对一致性要求高的场景,建议使用 temperature=0 或降低决策步骤的随机性。

参数敏感性风险:本文中的各种阈值(0.75、0.85 等)需要根据实际模型和场景调优,直接套用可能导致决策效果不佳。

7.4 成本模型

决策类型 额外LLM调用 平均Token消耗 平均延迟增加
置信度评估 1次/轮 ~200 tokens ~500ms
工具选择 0-1次/轮 ~100 tokens ~200ms
停止决策 0-1次/轮 ~150 tokens ~300ms
降级决策 0次(规则驱动) ~0 tokens <10ms
完整决策引擎 1-2次/轮 ~500 tokens ~1000ms

决策本身是有成本的。一个复杂的决策引擎每轮可能消耗 500+ tokens 和增加 1+ 秒延迟。对于简单查询,这可能使总延迟翻倍。因此,需要根据查询复杂度动态启用/禁用决策模块。


八、总结

Agent 的自主决策能力是其区别于传统对话系统的根本特征。一个优秀的决策系统需要在"足够好"和"最好"之间取得平衡------不是每次都追求最优解,而是在合理的成本内做出可靠的判断。

本文讨论的四类决策构成了 Agent 运行时的核心决策链路:

检索决策解决了"何时需要外部信息"的问题。通过置信度评估、时间敏感性检测和上下文充分性分析,Agent 能在"直接回答"和"先检索再回答"之间做出合理选择。关键在于建立可靠的置信度评估机制------无论是自我评估法、不确定性标记法还是一致性法,都需要结合具体场景校准。

工具选择决策解决了"用哪个工具"的问题。在多工具环境中,通过意图分类、能力画像和综合评分,Agent 能为不同类型的查询匹配合适的工具。核心权衡是准确度、延迟和成本的三维平衡。

停止决策解决了"何时该输出"的问题。硬性条件(迭代上限、Token预算、延迟上限)是安全网,质量条件(答案完整性、信息覆盖率)是主动停止机制。循环检测是防止 Agent 陷入无效循环的关键防线。

降级决策解决了"出错怎么办"的问题。分类型的错误处理策略------临时错误指数退避重试、永久错误切换替代方案------能最大化恢复成功率。最终降级到模型内部知识时,必须附加免责声明。

工程实践中最重要的三条经验:

第一,决策和执行必须分离。决策引擎只负责"做什么",执行器只负责"怎么做",通过共享状态协作。这种分离使决策逻辑可测试、可调优、可替换。

第二,所有决策必须可观测。每一步决策都应记录决策类型、置信度、推理过程和执行结果。没有可观测性就没有优化------你无法改进你看不见的东西。

第三,决策参数必须持续校准。置信度阈值、最大迭代次数、质量阈值等参数不是一劳永逸的。随着模型升级、用户需求变化、知识库更新,这些参数需要定期重新校准。

决策不是一次性的设计,而是持续优化的过程。好的决策系统会随着使用数据的积累越来越好------这就是工程化的价值所在。


参考资料

  1. Yao, S. et al. (2023). "ReAct: Synergizing Reasoning and Acting in Language Models." ICLR 2023.
  2. Shinn, N. et al. (2023). "Reflexion: Language Agents with Verbal Reinforcement Learning." NeurIPS 2023.
  3. LangChain Documentation - Agent Architectures: https://python.langchain.com/docs/modules/agents/
  4. LangGraph Documentation - Stateful Multi-Actor Applications: https://langchain-ai.github.io/langgraph/
  5. Wang, L. et al. (2024). "A Survey on Large Language Model based Autonomous Agents." Frontiers of Computer Science.
  6. AutoGen Documentation - Multi-Agent Conversation Framework: https://microsoft.github.io/autogen/
  7. LlamaIndex Documentation - Agent Patterns: https://docs.llamaindex.ai/en/stable/module_guides/deploying/agents/
  8. Schick, T. et al. (2023). "Toolformer: Language Models Can Teach Themselves to Use Tools." NeurIPS 2023.
  9. OpenAI Function Calling Guide: https://platform.openai.com/docs/guides/function-calling
  10. Anthropic Claude Tool Use Documentation: https://docs.anthropic.com/en/docs/build-with-claude/tool-use
相关推荐
沉默王二1 小时前
再见了 WebUI,DeepSeek 桌面版真不错。
openai·agent·deepseek
打工仔折腾 AI1 小时前
蓝耘元生代实测:用WorkBuddy与TextIn xParse拆解43页建模论文
人工智能·后端·python·数学建模·langchain·ai agent 实战
2601_962202981 小时前
万象生鲜系统:首衡集配移动端审批高效
大数据·人工智能
兔兔爱学习兔兔爱学习1 小时前
AI Infra知识点
人工智能
程序员清风1 小时前
基于 Spring Boot 构建生产级 AI 应用平台
大数据·人工智能·spring boot
星空inn1 小时前
OpenRouter 匿名模型 Space Bunny:上线当天宕机,社区三条指纹指向 MiniMax
ai·minimax
效率工作实验室1 小时前
用 AI Agent 做游戏数据分析,哪些场景效果最好?
ai·agent·智能体·游戏数据·游戏数据分析
程序员清风1 小时前
Portkey AI Gateway 实战:构建可治理的多模型网关
人工智能·gateway
小小王app小程序开发1 小时前
AI 智能体小程序开发实战:架构设计、核心难点与落地避坑
人工智能