LangGraph 并行执行与异步优化

上一篇我们聊了状态管理、条件路由和人工介入,这篇解决一个更实际的问题:怎么让 LangGraph 跑得快一点。

先说结论:LangGraph 的并行不是"魔法",它依赖两个前提------节点之间无数据依赖 + 调用异步接口(ainvoke/astream) 。满足这两点,框架会自动把无依赖的节点放进同一个 Superstep 并发执行,总耗时从"相加"变成"取最大值"。


一、先搞懂 Superstep:LangGraph 的"节拍器"

LangGraph 底层用的是 Pregel 模型,把执行过程切成一轮一轮的 Superstep(超步) :

less 复制代码
Superstep 1: 节点 A、B、C 并发执行
       ↓ 全部完成,全局同步
Superstep 2: 节点 D、E 并发执行(依赖 A/B/C 的输出)
       ↓ 全部完成,全局同步
Superstep 3: 节点 F 执行
       ↓
END

每个 Superstep 内部:

  1. 并发计算:所有被激活的节点同时跑,彼此读的是同一份状态快照
  2. 消息投递:节点输出写入 Channel 缓冲区
  3. 全局同步栅栏:等所有节点跑完 → 执行 Reducer 合并状态 → 存 Checkpoint → 进入下一轮

关键推论:

  • 同一 Superstep 内的节点天然并行,总耗时 = 最慢那个节点的耗时
  • 有依赖关系的节点一定在不同 Superstep,无法并行
  • 并行节点同时写同一个 State 字段时,必须声明 Reducer ,否则报 InvalidUpdateError

二、并行执行的两种姿势

2.1 静态扇出(Fan-out):分支数固定

最简单的方式------从一个节点引出多条边,目标节点自动并发。

python 复制代码
from typing import TypedDict, Annotated
import operator
from langgraph.graph import StateGraph, START, END

class State(TypedDict):
    query: str
    scores: Annotated[list[float], operator.add]  # 必须声明 Reducer!

def evaluate_a(state: State) -> dict:
    return {"scores": [0.9]}

def evaluate_b(state: State) -> dict:
    return {"scores": [0.8]}

def evaluate_c(state: State) -> dict:
    return {"scores": [0.7]}

def aggregate(state: State) -> dict:
    avg = sum(state["scores"]) / len(state["scores"])
    return {"final_score": avg}

graph = StateGraph(State)
graph.add_node("eval_a", evaluate_a)
graph.add_node("eval_b", evaluate_b)
graph.add_node("eval_c", evaluate_c)
graph.add_node("aggregate", aggregate)

# 三条边从 START 出发 → 三个节点进入同一 Superstep
graph.add_edge(START, "eval_a")
graph.add_edge(START, "eval_b")
graph.add_edge(START, "eval_c")

# 三个节点都指向 aggregate → 等全部完成后才执行(扇入同步)
graph.add_edge("eval_a", "aggregate")
graph.add_edge("eval_b", "aggregate")
graph.add_edge("eval_c", "aggregate")
graph.add_edge("aggregate", END)

app = graph.compile()

执行流程:

sql 复制代码
START → 并发执行 eval_a、eval_b、eval_c
         ↓ 全部完成
      aggregate(汇总三个分数)
         ↓
        END

适用场景:固定数量的并行任务,比如多维度评估、多数据源查询。

2.2 Send API:分支数运行时动态确定

当并发数量在编译期不确定时(比如"检索返回 N 条文档,每条都要处理"),用 Send。

python 复制代码
from langgraph.types import Send

class State(TypedDict, total=False):
    query: str
    docs: list[str]
    results: Annotated[list, operator.add]

def dispatcher(state: State) -> dict:
    # 模拟检索,返回动态数量的文档
    docs = retrieve_docs(state["query"])  # 可能返回 3 条,也可能返回 20 条
    return {"docs": docs}

def route_to_workers(state: State) -> list:
    # 为每条文档启动一个并发 worker
    return [Send("worker", {"doc": doc, "query": state["query"]})
            for doc in state["docs"]]

def worker(state: State) -> dict:
    doc = state["doc"]
    result = process_doc(doc)  # 耗时操作
    return {"results": [result]}

def aggregator(state: State) -> dict:
    return {"final": state["results"]}

graph = StateGraph(State)
graph.add_node("dispatcher", dispatcher)
graph.add_node("worker", worker)
graph.add_node("aggregator", aggregator)

graph.add_edge(START, "dispatcher")
# 条件边返回 Send 列表 → 框架为每个 Send 启动一个并发实例
graph.add_conditional_edges("dispatcher", route_to_workers, {"worker": "worker"})
graph.add_edge("worker", "aggregator")
graph.add_edge("aggregator", END)

app = graph.compile()

两种方式的对比:

维度 静态扇出 Send API
并发数量 编译期固定 运行时动态确定
实现方式 多条 add_edge 条件边返回 Send 列表
典型场景 固定多维度评估 文档分块处理、动态任务分发
复杂度 低 稍高

三、异步化:让并行真正跑起来

3.1 同步 vs 异步:差的不是一点点

上面的例子如果用同步节点,LangGraph 在同步模式(invoke)下会通过线程池并行;但如果节点内部是阻塞 I/O(如同步 HTTP 请求),线程会被占住,并发能力大打折扣。

正确做法是全异步化:

python 复制代码
import asyncio
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o-mini")

# 异步节点:用 async def
async def evaluate_a(state: State) -> dict:
    # 异步调用 LLM,不阻塞事件循环
    response = await llm.ainvoke(f"评估维度A: {state['query']}")
    return {"scores": [0.9]}

async def evaluate_b(state: State) -> dict:
    response = await llm.ainvoke(f"评估维度B: {state['query']}")
    return {"scores": [0.8]}

# 异步调用入口:用 ainvoke / astream
app = graph.compile()
result = await app.ainvoke({"query": "LangGraph 并行优化"})

实测效果 :4 核 8G 服务器上,同步 stream() + FastAPI 同步路由,50 并发就明显延迟;切换到 astream() + 异步路由后,同样硬件轻松支撑 500+ 并发。

3.2 异步节点的定义规范

python 复制代码
# 正确:async def + await
async def my_node(state: State) -> dict:
    result = await some_async_api(state["query"])
    return {"data": result}

# 错误:同步节点里调异步 API(会阻塞)
def my_node(state: State) -> dict:
    result = asyncio.run(some_async_api(state["query"]))  # 每个节点都起新事件循环,极慢
    return {"data": result}

关键原则:

  • 节点函数用 async def
  • 内部 I/O 操作用 await(llm.ainvoke()、httpx.AsyncClient 等)
  • 调用入口用 ainvoke() / astream() / astream_events()

3.3 限制并发数:防止 API 限流

并发不是越多越好。LLM API 通常有速率限制,并发过高会触发 429。

python 复制代码
import asyncio

# 用 Semaphore 限制最大并发数
semaphore = asyncio.Semaphore(5)  # 最多 5 个并发

async def worker_with_limit(state: State) -> dict:
    async with semaphore:
        result = await llm.ainvoke(state["query"])
        return {"results": [result.content]}

四、流式输出:边跑边看

并行优化了服务端吞吐量 ,流式输出优化了用户感知延迟。两者配合,体验最好。

4.1 三种流式模式

LangGraph 的 stream() / astream() 支持多种 stream_mode:

ini 复制代码
# 模式1:values ------ 每步返回完整 State(数据量大,适合调试)
async for chunk in app.astream(initial_state, stream_mode="values"):
    print(chunk)

# 模式2:updates ------ 只返回本次节点的增量更新(轻量,推荐)
async for mode, updates in app.astream(initial_state, stream_mode="updates"):
    for node_name, update in updates.items():
        print(f"[{node_name}] {update}")

# 模式3:messages ------ 逐 token 输出 LLM 回复(聊天场景必备)
async for msg, metadata in app.astream(initial_state, stream_mode="messages"):
    print(msg.content, end="", flush=True)

怎么选:

模式 返回内容 适用场景
values 完整 State 快照 调试、状态可视化
updates 节点增量更新 生产环境默认选择
messages LLM token 流 聊天界面打字机效果

可以组合使用:stream_mode=["updates", "messages"]

4.2 astream_events:细粒度事件订阅

比 astream 更底层,可以订阅到 LLM token、工具调用、节点开始/结束等事件:

ini 复制代码
async for event in app.astream_events(initial_state, version="v2"):
    kind = event["event"]
    if kind == "on_chat_model_stream":
        # LLM 逐字输出
        print(event["data"]["chunk"].content, end="", flush=True)
    elif kind == "on_tool_start":
        # 工具开始执行
        print(f"\n 调用工具: {event['name']}")
    elif kind == "on_node_end":
        # 节点完成
        print(f"\n 节点完成: {event['name']}")

建议 :代码里显式写 version="v2",避免 LangGraph 升级后事件结构变化。

4.3 StreamWriter:自定义进度事件

工具节点执行时不产生 token,前端会"静默"。可以用 StreamWriter 推送自定义进度:

python 复制代码
from langgraph.types import StreamWriter

async def search_node(state: State, writer: StreamWriter) -> dict:
    writer({"type": "status", "msg": "正在搜索..."})
    results = await search_api(state["query"])
    writer({"type": "status", "msg": f"找到 {len(results)} 条结果"})
    return {"search_results": results}

前端通过 stream_mode="custom" 接收这些事件,渲染进度条或状态提示。


五、完整案例:多源资讯并行检索 + 流式汇总

把上面所有技巧串起来,做一个"并行检索多源资讯 → 汇总生成报告"的 Agent。

python 复制代码
from typing import TypedDict, Annotated, List
import operator
import asyncio
from langgraph.graph import StateGraph, START, END
from langgraph.types import Send, StreamWriter
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o-mini")

# ---- 状态定义 ----
class ResearchState(TypedDict, total=False):
    query: str
    sources: List[str]          # 数据源列表
    raw_results: Annotated[List[dict], operator.add]  # Reducer 合并
    summary: str
    status: str

# ---- 异步节点 ----
async def fetch_source(state: ResearchState, writer: StreamWriter) -> dict:
    """为每个数据源启动一个并发检索任务"""
    source = state["source"]
    writer({"type": "fetch_start", "source": source})
    try:
        # 模拟异步 API 调用
        result = await asyncio.sleep(1) or {"source": source, "content": f"{source} 的内容"}
        writer({"type": "fetch_done", "source": source})
        return {"raw_results": [result]}
    except Exception as e:
        writer({"type": "fetch_error", "source": source, "error": str(e)})
        return {"raw_results": [{"source": source, "error": str(e)}]}

def dispatch_sources(state: ResearchState) -> list:
    """动态决定并发几个数据源"""
    sources = state.get("sources", ["news", "blog", "paper"])
    return [Send("fetch_source", {**state, "source": s}) for s in sources]

async def summarize(state: ResearchState, writer: StreamWriter) -> dict:
    writer({"type": "summarize_start"})
    prompt = f"根据以下检索结果,总结关于 '{state['query']}' 的报告:\n{state['raw_results']}"
    response = await llm.ainvoke(prompt)
    writer({"type": "summarize_done"})
    return {"summary": response.content, "status": "done"}

# ---- 建图 ----
graph = StateGraph(ResearchState)
graph.add_node("fetch_source", fetch_source)
graph.add_node("summarize", summarize)

graph.add_edge(START, "dispatch_placeholder")  # 占位,实际用条件边
graph.add_conditional_edges(START, lambda s: "dispatch")  # 简化示意
# 实际写法:
graph.add_node("dispatch", lambda s: {"sources": s.get("sources", ["news", "blog", "paper"])})
graph.add_edge(START, "dispatch")
graph.add_conditional_edges("dispatch", dispatch_sources, {"fetch_source": "fetch_source"})
graph.add_edge("fetch_source", "summarize")
graph.add_edge("summarize", END)

app = graph.compile()

# ---- 异步流式执行 ----
async def run():
    initial = {"query": "LangGraph 并行优化", "sources": ["news", "blog", "paper", "twitter"]}
    async for mode, chunk in app.astream(initial, stream_mode=["updates", "custom"]):
        if mode == "custom":
            print(f"  {chunk}")
        elif mode == "updates":
            for node, update in chunk.items():
                print(f"[{node}] {update}")

asyncio.run(run())

执行效果:

css 复制代码
  {"type": "fetch_start", "source": "news"}
  {"type": "fetch_start", "source": "blog"}
  {"type": "fetch_start", "source": "paper"}
  {"type": "fetch_start", "source": "twitter"}
  {"type": "fetch_done", "source": "news"}
  {"type": "fetch_done", "source": "blog"}
  ...(4 个源并发,总耗时 ≈ 最慢那个,而非 4 倍)
  {"type": "summarize_start"}
  {"type": "summarize_done"}
[summarize] {'summary': '...'}

六、工程实践:五个容易踩的坑

6.1 忘记声明 Reducer

并发节点同时写同一个 State 字段,不声明 Reducer 直接报错 InvalidUpdateError。

ruby 复制代码
# 错误
class State(TypedDict):
    results: list  # 并发写入会冲突

# 正确
class State(TypedDict):
    results: Annotated[list, operator.add]  # 自动合并

6.2 同步节点假装异步

python 复制代码
# 错误:节点是同步的,ainvoke 也救不了
def slow_node(state):
    time.sleep(3)  # 阻塞线程
    return {"data": "ok"}

# 正确
async def slow_node(state):
    await asyncio.sleep(3)  # 不阻塞事件循环
    return {"data": "ok"}

6.3 并发数不设限

检索 20 个数据源就启动 20 个并发,大概率触发 API 限流。用 asyncio.Semaphore 兜底。

6.4 流式静默期过长

工具节点执行期间没有 token 输出,用户以为卡死了。用 StreamWriter 推进度,或者设置前端超时降级。

6.5 State 膨胀

stream_mode="values" 每次返回完整 State,长对话中 messages 列表越来越大,带宽和内存都吃紧。生产环境优先用 updates 或 messages,定期清理历史消息。


七、性能优化速查表

优化方向 具体做法 预期收益
并行执行 无依赖节点放同一 Superstep 耗时从相加变取最大值
全异步化 节点用 async def,调用 ainvoke 并发能力提升 10x+
限制并发 asyncio.Semaphore(N) 避免 API 限流
流式输出 astream + stream_mode="updates" 用户感知延迟降低
进度反馈 StreamWriter 推自定义事件 消除静默期焦虑
状态精简 State 只存引用,不存大对象 减少序列化开销
分布式 Checkpoint Redis/PostgreSQL 替代 SQLite 支持水平扩展

总结

并行和异步这两件事,本质上是把等待时间利用起来:

  • 并行:让多个独立任务同时等,而不是排队等
  • 异步:等的时候不占着线程,把资源让给别人

LangGraph 在这两层都做了封装,你只需要记住三条规则:

  1. 无依赖 → 加边 → 自动并行
  2. 要并发 → 用 async def + ainvoke
  3. 要体验 → 用 astream + StreamWriter
相关推荐
TMT星球2 小时前
旗舰手机的中场战事:AI为核,影像为王
人工智能·智能手机
霸道流氓气质2 小时前
AI应用成本优化完全指南:Token压缩、语义缓存与模型路由的Java生产级实战
java·人工智能·缓存
Gu0Qiang2 小时前
从 0 到 1 打造 AI 提示流编排器:别把大模型当机械拼图!Case #7 字段冻结陷阱与代码回滚复盘(开源系列 15)
人工智能·github
骇客野人2 小时前
SDD AI-Native(Spec-Driven Development,规约驱动开发,AI 原生研发范式)
人工智能·ai-native
用户547455508122 小时前
用开源的Toonflow和MiniMax H3一步步复刻万妖
人工智能
RobinDevNotes2 小时前
JAX 分布式训练,和 PyTorch 有什么不一样
人工智能·深度学习·ajax
奕鼎竜瑆2 小时前
[新手小白也能学会] 01-PyTorch框架使用(上)
人工智能·pytorch·python
用户837133200762 小时前
接口返回文章 ID 后,怎样确认发布真的完成了?
人工智能
Daorigin_com2 小时前
道本科技携手DeepSeek:以AI重塑合同全生命周期管理
前端·人工智能·科技·网络安全·数据挖掘·前端框架·传媒
guslegend2 小时前
AutoDebug Agent:用真实反馈做出会修缺陷的 Agent
人工智能