AI Agent 第七周完整学习笔记
第七周主题:LangGraph Tool Agent 高级能力
学习目标:从"有状态 Workflow"升级到真正可调用 Tool、可循环、可持久化、可人工审批的 Agent Runtime。
统一环境:
Google Colab + Python + LangChain + LangGraph + langchain-openai + ChatOpenAI + 阿里云百炼 + qwen-plus
第七周总体路线
text
Day 31
MessagesState 深入
HumanMessage / AIMessage / ToolMessage
多轮消息
State ≠ Memory
State ≠ Context
↓
Day 32
@tool
bind_tools()
AIMessage.tool_calls
ToolNode
ToolMessage
tools_condition
↓
Day 33
完整 Tool Calling Agent
agent → tools → agent
Agent Loop
create_agent 对比
↓
Day 34
Checkpointer
Checkpoint
thread_id
Persistence
StateSnapshot
get_state()
get_state_history()
↓
Day 35
Human-in-the-loop
interrupt()
Command(resume=...)
人工审批
高风险 Tool
第七周最终目标架构
text
thread_id
│
↓
Checkpointer
│
↓
MessagesState
│
↓
Agent
↓
LLM
↓
Tool Calls
↓
Runtime Policy
↙ ↘
Low Risk High Risk
↓ ↓
ToolNode interrupt
↓ ↓
ToolMessage Human
↓ ↙ ↘
Agent Approve Reject
↑ ↓ ↓
└──────── ToolNode END
↓
Agent
Day 31|MessagesState 深入 + Qwen 多轮对话
一、为什么 Agent 需要 messages
真正 Agent 的 State 中,经常会出现:
python
{
"messages": [
HumanMessage(...),
AIMessage(...),
ToolMessage(...),
AIMessage(...)
]
}
因为 Agent 的运行过程本质是:
text
用户输入
↓
HumanMessage
模型输出
↓
AIMessage
模型请求 Tool
↓
AIMessage(tool_calls)
Tool 执行结果
↓
ToolMessage
模型继续处理
↓
AIMessage
所以:
text
state["messages"]
是 Agent 最核心的状态之一。
二、四种主要 Message
| Message | 谁产生 | 作用 |
|---|---|---|
SystemMessage |
程序 | 角色、规则、约束 |
HumanMessage |
用户 | 用户输入 |
AIMessage |
模型 | 模型回复 / Tool Call |
ToolMessage |
Tool Runtime | Tool 执行结果 |
三、MessagesState 是什么
python
from langgraph.graph import MessagesState
可以近似理解:
python
class MessagesState(TypedDict):
messages: Annotated[list, add_messages]
也就是:
text
messages字段
+
add_messages Reducer
四、为什么不能普通 list
如果:
python
class State(TypedDict):
messages: list
默认更新行为:
text
新值覆盖旧值
例如:
text
旧:HumanMessage("你好")
新:AIMessage("你好")
最后可能只剩:
text
AIMessage
所以 messages 通常需要 add_messages。
五、实验:MessagesState + Qwen
python
from langgraph.graph import (
MessagesState,
StateGraph,
START,
END
)
from langchain_core.messages import (
HumanMessage,
AIMessage
)
def chat_node(state: MessagesState):
print(
"模型本次看到的消息数量:",
len(state["messages"])
)
response = model.invoke(
state["messages"]
)
print(
"模型返回类型:",
type(response).__name__
)
print(
"模型回答:",
response.content
)
return {
"messages": [
response
]
}
builder = StateGraph(MessagesState)
builder.add_node(
"chat",
chat_node
)
builder.add_edge(
START,
"chat"
)
builder.add_edge(
"chat",
END
)
chat_graph = builder.compile()
预期输出
text
没有输出
运行:
python
result = chat_graph.invoke({
"messages": [
HumanMessage(
content="你好,请简单介绍一下LangGraph。"
)
]
})
预期输出
text
模型本次看到的消息数量: 1
模型返回类型: AIMessage
模型回答: LangGraph 是......
六、为什么 state["messages"][-1] 经常出现
python
last_message = state["messages"][-1]
表示:
text
获取当前 messages 的最后一条消息
注意:最后一条不一定是 HumanMessage,也可能是 AIMessage 或 ToolMessage。
七、为什么模型调用通常传完整 messages
python
model.invoke(
state["messages"]
)
因为模型理解上下文通常需要相关历史。
而 Router 常常只需要:
python
state["messages"][-1]
来判断最新一轮状态。
八、MessagesState ≠ Memory
Day31 的 MessagesState 解决的是:
text
一次 Graph State 中
messages 如何保存和更新
它不会自动解决:
text
多次独立 invoke 之间
历史 State 保存在哪里
跨 invoke 保存要等 Day34 的:
text
Checkpointer
+
thread_id
九、State ≠ Context
假设 State 中有:
text
messages
customer_id
retry_count
status
但如果:
python
model.invoke(
state["messages"]
)
模型实际看到的只有 messages。
所以:
text
State
=
Graph当前拥有的工作数据
Context
=
本次Model Call真正看到的数据
Day 31 最终记忆
text
MessagesState
=
messages + add_messages
HumanMessage
=
用户输入
AIMessage
=
模型输出 / Tool Call
ToolMessage
=
Tool Observation
MessagesState
≠
跨invoke持久化Memory
State
≠
Context
Day 32|ToolNode + Tool Calling
一、完整核心链
text
@tool
↓
Tool Schema
↓
bind_tools()
↓
AIMessage.tool_calls
↓
ToolNode
↓
Python Tool
↓
ToolMessage
↓
messages
二、几个角色
| 组件 | 作用 |
|---|---|
@tool |
定义 Tool |
bind_tools() |
把 Tool Schema 告诉模型 |
AIMessage.tool_calls |
模型提出 Tool 请求 |
ToolNode |
真正执行 Tool |
ToolMessage |
Tool 执行结果 |
最重要:
text
bind_tools()
≠
执行Tool
AIMessage.tool_calls
≠
Tool已经执行
三、三个 Tool
python
from langchain_core.tools import tool
from typing import Literal
@tool
def query_customer(
customer_id: str
) -> dict:
"""
根据CRM客户编号查询客户基本信息。
"""
customers = {
"10001": {
"name": "张三",
"level": "VIP",
"assets": 2000000
},
"10002": {
"name": "李四",
"level": "NORMAL",
"assets": 600000
}
}
if customer_id not in customers:
return {
"success": False,
"error": f"客户不存在:{customer_id}"
}
customer = customers[customer_id]
return {
"success": True,
"customer_id": customer_id,
"name": customer["name"],
"level": customer["level"],
"assets": customer["assets"]
}
@tool
def calculator(
a: float,
b: float,
operation: Literal[
"add",
"subtract",
"multiply",
"divide"
]
) -> dict:
"""
执行两个数字之间的数学运算。
"""
if operation == "add":
result = a + b
elif operation == "subtract":
result = a - b
elif operation == "multiply":
result = a * b
elif operation == "divide":
if b == 0:
return {
"success": False,
"error": "除数不能为0"
}
result = a / b
else:
return {
"success": False,
"error": f"不支持的操作:{operation}"
}
return {
"success": True,
"result": result
}
@tool
def get_weather(
city: str
) -> dict:
"""
查询指定城市天气。
"""
weather_data = {
"北京": {
"temperature": 30,
"weather": "晴"
},
"上海": {
"temperature": 26,
"weather": "多云"
},
"广州": {
"temperature": 32,
"weather": "阵雨"
}
}
if city not in weather_data:
return {
"success": False,
"error": f"暂时没有{city}的天气数据"
}
data = weather_data[city]
return {
"success": True,
"city": city,
"temperature": data["temperature"],
"weather": data["weather"]
}
预期输出
text
没有输出
四、bind_tools
python
tools = [
query_customer,
calculator,
get_weather
]
model_with_tools = (
model.bind_tools(
tools
)
)
预期输出
text
没有输出
五、AIMessage.tool_calls
python
from langchain_core.messages import (
SystemMessage,
HumanMessage,
AIMessage,
ToolMessage
)
messages = [
SystemMessage(
content=(
"你是CRM智能助手。"
"客户资料必须调用query_customer工具。"
)
),
HumanMessage(
content="客户10001是谁?"
)
]
response = (
model_with_tools.invoke(
messages
)
)
print(
response.tool_calls
)
预期输出
类似:
text
[
{
'name': 'query_customer',
'args': {
'customer_id': '10001'
},
'id': 'call_xxx'
}
]
六、Tool Call 三个关键字段
text
name
=
调用哪个Tool
args
=
传什么参数
id
=
这次Tool Call唯一标识
Tool Call 是:
text
Action请求
不是执行结果。
七、ToolNode
python
from langgraph.prebuilt import ToolNode
tool_node = ToolNode(
tools
)
预期输出
text
没有输出
职责:
text
读取 AIMessage.tool_calls
↓
找到 Tool
↓
执行 Tool
↓
生成 ToolMessage
八、执行 ToolNode
python
tool_result = (
tool_node.invoke({
"messages": [
response
]
})
)
print(
tool_result
)
预期输出核心
text
ToolMessage
name:
query_customer
content:
{
success: true,
customer_id: 10001,
name: 张三,
level: VIP,
assets: 2000000
}
tool_call_id:
call_xxx
九、tools_condition
python
from langgraph.prebuilt import tools_condition
它本质:
text
最后一个AIMessage有tool_calls
→ tools
没有tool_calls
→ END
Day 32 最终记忆
text
@tool
=
定义Tool
bind_tools()
=
让模型知道Tool
AIMessage.tool_calls
=
模型提出Action
ToolNode
=
真正执行Action
ToolMessage
=
Observation
tools_condition
=
判断是否进入ToolNode
Day 33|完整 Tool Calling Agent Graph
一、完整 Agent Loop
Day32:
text
agent
↓
tools
↓
END
Day33:
text
agent
↓
tools
↓
agent
关键代码:
python
builder.add_edge(
"tools",
"agent"
)
二、Agent Node
python
def agent_node(
state: MessagesState
):
response = (
model_with_tools.invoke(
state["messages"]
)
)
return {
"messages": [
response
]
}
预期输出
text
没有输出
三、完整 Graph
python
from langgraph.graph import (
StateGraph,
START,
END
)
from langgraph.prebuilt import (
ToolNode,
tools_condition
)
builder = StateGraph(
MessagesState
)
builder.add_node(
"agent",
agent_node
)
builder.add_node(
"tools",
ToolNode(tools)
)
builder.add_edge(
START,
"agent"
)
builder.add_conditional_edges(
"agent",
tools_condition
)
builder.add_edge(
"tools",
"agent"
)
agent_graph = (
builder.compile()
)
预期输出
text
没有输出
四、执行客户查询
text
START
↓
agent 第1次
↓
AIMessage(tool_call)
↓
tools_condition
↓
tools
↓
ToolMessage
↓
agent 第2次
↓
AIMessage(final)
↓
tools_condition
↓
END
最终 messages:
text
SystemMessage
HumanMessage
AIMessage(tool_call)
ToolMessage
AIMessage(final)
五、多 Tool 任务
例如:
text
查询客户10001,
然后按照折扣率计算1000元客户价
可能:
text
agent
↓
query_customer
↓
agent
↓
calculator
↓
agent
↓
final answer
Agent Node 执行次数不是固定的。
六、谁决定继续循环
真正决定有没有 Tool Call 的是:
text
LLM
tools_condition 只是读取模型已经做出的结构化决定。
所以:
text
模型
=
提出Action
Runtime
=
验证、路由、执行Action
七、create_agent 对比
python
from langchain.agents import create_agent
ready_agent = create_agent(
model=model,
tools=tools,
system_prompt=(
"你是CRM智能助手。"
"客户数据必须通过工具查询。"
)
)
预期输出
text
没有输出
create_agent() 帮你封装标准:
text
Messages State
Model Node
Tool Calling
ToolNode
Tool Loop
Conditional Routing
Stop Condition
自己手写 Graph 的价值:
text
理解底层Agent Loop
Day 33 最终记忆
text
agent
=
Model Node
tools_condition
=
判断有没有Tool Call
ToolNode
=
执行Tool
ToolMessage
=
Observation
tools → agent
=
Agent Loop
AIMessage没有tool_calls
=
Stop Condition
Day 34|Checkpointer + thread_id + Persistence
一、为什么需要 Persistence
没有 Checkpointer:
text
第一次invoke:
我叫张三
第二次invoke:
我叫什么?
→ 不知道
Day34:
text
State
↓
Checkpoint
↓
Checkpointer
↓
thread_id
↓
下一次invoke
↓
恢复State
二、核心概念
| 概念 | 含义 |
|---|---|
| State | Graph 当前工作数据 |
| Checkpoint | State 某一时刻快照 |
| Checkpointer | 保存 / 恢复 Checkpoint |
| Thread | 一条状态历史 |
| thread_id | Thread 唯一标识 |
| Persistence | State 跨 invoke 保存 |
| Context | 本次真正送给模型的数据 |
三、InMemorySaver
python
from langgraph.checkpoint.memory import InMemorySaver
checkpointer = (
InMemorySaver()
)
预期输出
text
没有输出
四、compile 配置 Checkpointer
python
chat_graph = (
builder.compile(
checkpointer=
checkpointer
)
)
预期输出
text
没有输出
五、thread_id
python
config = {
"configurable": {
"thread_id":
"chat-001"
}
}
预期输出
text
没有输出
同一个 thread_id:
text
继续同一条会话 / 状态线程
六、第一轮
python
result1 = chat_graph.invoke(
{
"messages": [
HumanMessage(
content=(
"我叫张三,"
"请记住这个名字。"
)
)
]
},
config=config
)
预期输出
类似:
text
好的,我知道你叫张三。
七、第二轮只传新消息
python
result2 = chat_graph.invoke(
{
"messages": [
HumanMessage(
content=
"我刚才说我叫什么?"
)
]
},
config=config
)
print(
result2[
"messages"
][-1].content
)
预期输出
类似:
text
你刚才说你叫张三。
八、为什么第二轮知道张三
text
thread_id=chat-001
↓
Checkpointer找到旧Checkpoint
↓
恢复MessagesState
↓
加入新的HumanMessage
↓
add_messages合并
↓
模型看到完整相关历史
九、换 thread_id
text
chat-001
和
chat-002
是不同 Thread。
所以新的 Thread 不会自动拥有旧 Thread 的状态。
十、thread_id 不要简单等于 user_id
一个用户可能同时有:
text
会话A
会话B
会话C
更合理:
text
user_id
↓
多个conversation_id
↓
多个thread_id
实际项目:
text
thread_id
≈
conversation_id
十一、get_state
python
snapshot = (
chat_graph.get_state(
config
)
)
print(
type(snapshot).__name__
)
print(
snapshot.values
)
print(
snapshot.next
)
预期输出
类似:
text
StateSnapshot
{
"messages": [...]
}
()
十二、get_state_history
python
history = list(
chat_graph.get_state_history(
config
)
)
print(
"Checkpoint数量:",
len(history)
)
预期输出
text
多个Checkpoint
用于:
text
调试Agent
分析执行轨迹
查看Thread历史状态
十三、Checkpointer ≠ Long-term Memory
Checkpointer 更像:
text
Thread级短期状态
跨 Thread 长期用户信息:
text
Long-term Store
是另一层能力。
十四、InMemorySaver 局限
如果:
text
Colab Runtime重启
Python进程退出
内存状态就没了。
生产通常考虑数据库型 Checkpointer,例如 PostgresSaver。
Day 34 最终记忆
text
State
=
当前Graph工作数据
Checkpoint
=
某时刻State快照
Checkpointer
=
保存和恢复Checkpoint
Thread
=
一串Checkpoint
thread_id
=
指定哪一条Thread
Persistence
=
State跨invoke继续
Checkpointer
≠
Long-term User Memory
Day 35|Human-in-the-loop + interrupt + resume
一、为什么需要人工审批
低风险:
text
查客户
查天气
读取信息
通常可自动执行。
高风险:
text
删除
正式发送
转账
修改生产数据
覆盖文件
应该:
text
模型提出动作
↓
Runtime判断风险
↓
人工审批
↓
再执行
二、核心 API
暂停:
python
interrupt(...)
恢复:
python
Command(
resume=...
)
三、简单审批 State
python
from typing_extensions import TypedDict
class ApprovalState(TypedDict):
customer_id: str
operation: str
approval: str
result: str
预期输出
text
没有输出
四、human_review Node
python
from langgraph.types import (
interrupt,
Command
)
def human_review(
state: ApprovalState
):
decision = interrupt({
"type":
"approval",
"message":
"这是一个高风险操作,需要人工确认。",
"operation":
state["operation"],
"options": [
"approve",
"reject"
]
})
return {
"approval":
decision
}
预期输出
text
没有输出
第一次运行到 interrupt():
text
Graph暂停
不是失败。
五、人工批准
python
Command(
resume="approve"
)
恢复后:
text
interrupt()返回"approve"
↓
Node继续
↓
Router
↓
执行操作
六、人工拒绝
python
Command(
resume="reject"
)
路线:
text
human_review
↓
reject
↓
cancel / END
七、非常重要:Resume 后 Node 会重新从头执行
例如:
python
def demo_interrupt_node(
state
):
print(
"Node从头开始执行"
)
answer = interrupt(
"请输入YES继续"
)
print(
"interrupt之后:",
answer
)
第一次:
text
Node从头开始执行
↓
interrupt
↓
暂停
Resume:
text
Node从头开始执行
↓
interrupt返回resume值
↓
interrupt之后:YES
八、为什么副作用危险
错误:
python
def dangerous_node(
state
):
send_email()
approval = interrupt(
"是否继续?"
)
Resume 会重跑 Node,可能再次 send_email()。
所以:
text
interrupt前
尽量不要做不可幂等副作用
九、正确结构
text
human_review
↓
interrupt
↓
人工确认
↓
execute_action
也就是:
text
审批Node
和
真正副作用Node
分开
Day 35 综合:高风险 Tool Agent
一、目标架构
text
START
↓
agent
↓
tool_router
↙ ↓ ↘
END safe_tool risky_tool
↓ ↓
tools human_review
↓ ↙ ↘
agent approve reject
↓ ↓
tools END
↓
agent
二、两个 Tool
python
@tool
def query_customer(
customer_id: str
) -> dict:
"""
查询客户信息。
"""
return {
"customer_id": customer_id,
"name": "张三",
"level": "VIP"
}
@tool
def delete_customer(
customer_id: str
) -> dict:
"""
删除指定客户。
这是高风险操作。
"""
return {
"success": True,
"message": f"客户{customer_id}已删除"
}
预期输出
text
没有输出
三、Agent State
python
class AgentApprovalState(
MessagesState
):
approval: str
预期输出
text
没有输出
四、Agent Node
python
def agent_node(
state: AgentApprovalState
):
response = (
model_with_tools.invoke(
state["messages"]
)
)
return {
"messages": [
response
]
}
预期输出
text
没有输出
五、自定义 Tool Router
python
def tool_router(
state: AgentApprovalState
):
last_message = (
state["messages"][-1]
)
if not getattr(
last_message,
"tool_calls",
None
):
return "end"
tool_names = [
call["name"]
for call
in last_message.tool_calls
]
if "delete_customer" in tool_names:
return "review"
return "tools"
预期输出
text
没有输出
六、Human Review Node
python
def agent_human_review(
state: AgentApprovalState
):
last_message = (
state["messages"][-1]
)
tool_calls = (
last_message.tool_calls
)
decision = interrupt({
"type":
"tool_approval",
"message":
"Agent准备执行高风险Tool,请人工确认。",
"tool_calls":
tool_calls
})
return {
"approval":
decision
}
预期输出
text
没有输出
七、审批 Router
python
def review_router(
state: AgentApprovalState
):
if state["approval"] == "approve":
return "tools"
return "end"
预期输出
text
没有输出
八、完整 HITL Graph
python
from langgraph.prebuilt import ToolNode
builder = StateGraph(
AgentApprovalState
)
builder.add_node(
"agent",
agent_node
)
builder.add_node(
"tools",
ToolNode(tools)
)
builder.add_node(
"human_review",
agent_human_review
)
builder.add_edge(
START,
"agent"
)
builder.add_conditional_edges(
"agent",
tool_router,
{
"end": END,
"tools": "tools",
"review": "human_review"
}
)
builder.add_conditional_edges(
"human_review",
review_router,
{
"tools": "tools",
"end": END
}
)
builder.add_edge(
"tools",
"agent"
)
agent_checkpointer = (
InMemorySaver()
)
hitl_agent = (
builder.compile(
checkpointer=
agent_checkpointer
)
)
预期输出
text
没有输出
九、低风险 Tool 路线
text
用户:查询客户10001
↓
agent
↓
query_customer Tool Call
↓
tool_router
↓
tools
↓
ToolNode
↓
ToolMessage
↓
agent
↓
final answer
↓
END
不会 interrupt。
十、高风险 Tool 路线
text
用户:删除客户10001
↓
agent
↓
delete_customer Tool Call
↓
tool_router
↓
human_review
↓
interrupt
↓
暂停
此时高风险 Tool 还没有执行。
人工批准:
python
Command(
resume="approve"
)
继续:
text
human_review
↓
tools
↓
delete_customer
↓
ToolMessage
↓
agent
↓
final answer
↓
END
人工拒绝:
python
Command(
resume="reject"
)
路线:
text
human_review
↓
END
高风险 Tool 不执行。
第七周知识串联
text
Day 31
MessagesState
↓
消息状态
Day 32
ToolNode
↓
Tool执行
Day 33
agent → tools → agent
↓
Agent Loop
Day 34
Checkpointer + thread_id
↓
Persistence
Day 35
interrupt + resume
↓
Human-in-the-loop
第七周最重要的工程原则
- LLM 负责提出 Tool Call,不负责执行 Tool。
- ToolNode / Runtime 才真正执行 Tool。
- Tool Result 必须以 ToolMessage 回到 messages。
tools → agent才形成完整 Agent Loop。- AIMessage 没有 tool_calls 时通常结束当前 Agent Loop。
- MessagesState 不是跨 invoke 的持久化 Memory。
- 跨 invoke 状态需要 Checkpointer + thread_id。
- thread_id 更适合映射 conversation_id,而不是简单 user_id。
- Checkpointer 是 Thread 状态持久化,不等于长期用户记忆。
- 高风险 Tool 应由 Runtime Policy 和人工审批控制,不要只靠 Prompt。
- interrupt 前避免不可幂等副作用,因为 Resume 会重新执行 Node。
- 审批 Node 和真正执行副作用的 Node 最好分开。
第七周最终验收题
- MessagesState 是什么?
- add_messages 为什么重要?
state["messages"][-1]是什么?- State 和 Context 有什么区别?
- bind_tools() 做什么?
- AIMessage.tool_calls 是 Tool 执行结果吗?
- ToolNode 是什么?
- ToolMessage 是什么?
- tools_condition 做什么?
- 为什么需要 tools → agent 回边?
- Agent Loop 什么时候结束?
- create_agent() 帮我们封装了什么?
- Checkpoint 是什么?
- Checkpointer 是什么?
- thread_id 是什么?
- 一个 Thread 是否只有一个 Checkpoint?
- 为什么换 thread_id 就没有之前状态?
- InMemorySaver 适合生产吗?
- get_state() 返回什么?
- get_state_history() 有什么用?
- Checkpointer 和 Long-term Store 有什么区别?
- interrupt() 做什么?
- Command(resume=...) 做什么?
- 为什么 interrupt 依赖 Checkpointer?
- Resume 后 Node 是从 interrupt 下一行直接继续吗?
- 为什么 interrupt 前不要放不可重复副作用?
- 为什么高风险 Tool 应该先审批再执行?
- Day31~Day35 最终组合成了什么?
第七周最终答案压缩版
text
MessagesState
=
Agent消息状态
add_messages
=
消息合并Reducer
AIMessage.tool_calls
=
模型提出Action
ToolNode
=
执行Action
ToolMessage
=
Observation
tools_condition
=
判断是否进入Tool
tools → agent
=
Agent Loop
Checkpoint
=
State快照
Checkpointer
=
保存 / 恢复State快照
thread_id
=
哪一条状态线程
Persistence
=
State跨invoke继续
interrupt
=
暂停Graph等待外部输入
Command(resume=...)
=
恢复Graph
Human-in-the-loop
=
关键步骤由人参与最终决定
第七周一句话总结
一个成熟的 Agent Runtime,不只是"LLM 会调用 Tool",还要能够维护消息状态、形成受控循环、跨执行保存 State、识别高风险动作,并在需要时暂停让人审核后再安全恢复执行。
下一周预告|RAG 基础
text
Document
↓
Chunk
↓
Embedding
↓
Vector Store
↓
Retrieval
↓
TopK
↓
Context
↓
LLM
第八周会从:
text
Agent会调用Tool
继续升级到:
text
Agent能够从知识库中检索真实业务知识