上一章我们介绍了:
Enterprise AI Gateway
它位于:
Context Engine
↓
Agent Runtime
↓
AI Gateway
↓
Model / Agent / Tool / MCP
↓
Enterprise Systems
但真正进入生产环境之后,一个新的问题会出现:
Gateway到底应该负责什么?
如果只是简单地:
Request
↓
LLM API
那么它实际上只是一个API Proxy。
真正的Enterprise AI Gateway需要解决:
身份
路由
模型选择
负载均衡
限流
降级
重试
缓存
成本
安全
审计
监控
甚至还需要进一步处理:
Agent Routing
Tool Routing
MCP Routing
Tenant Routing
Policy Routing
因此:
Enterprise AI Gateway不是简单的网络代理,而是Enterprise AI运行时的重要控制入口。
一、为什么企业需要AI Gateway?
假设一家企业同时使用:
GPT
Claude
Gemini
Qwen
DeepSeek
企业私有模型
如果每个Agent直接连接模型:
Agent A → Model A
Agent B → Model B
Agent C → Model C
Agent D → Model D
很快就会出现:
API Key分散
模型配置分散
成本无法统一统计
限流难以管理
故障难以切换
审计困难
更复杂的是:
Agent
↓
Tool
↓
MCP
↓
ERP / WMS / MES
每个Agent自己维护连接。
最终形成:
Agent A
/ | \
/ | \
Model Tool MCP
Agent B
/ | \
/ | \
Model Tool MCP
Agent C
/ | \
/ | \
Model Tool MCP
这会造成大量重复基础设施。
所以需要:
Agents
│
↓
AI Gateway
│
┌───────────────┼───────────────┐
↓ ↓ ↓
Models Tools MCP
│ │ │
└───────────────┼───────────────┘
↓
Enterprise Systems
二、AI Gateway到底是什么?
可以把AI Gateway理解成:
Enterprise AI的统一流量入口。
传统互联网系统:
User
↓
API Gateway
↓
Microservices
Enterprise AI:
User
↓
Agent
↓
AI Gateway
↓
Models / Tools / MCP
↓
Enterprise Systems
因此AI Gateway实际上承担了类似:
API Gateway
+
Model Gateway
+
Agent Gateway
+
Tool Gateway
+
MCP Gateway
的职责。
三、AI Gateway整体架构
可以设计为:
Enterprise AI Gateway
│
┌──────────────────────────┼──────────────────────────┐
↓ ↓ ↓
Identity Routing Policy
│ │ │
Authentication Model Route Security
Authorization Agent Route Permission
Tenant Tool Route Compliance
MCP Route
└──────────────────────────┼──────────────────────────┘
↓
Traffic Management
│
┌─────────────────────────┼─────────────────────────┐
↓ ↓ ↓
Rate Limit Load Balance Retry
↓ ↓ ↓
Quota Routing Timeout
│
↓
Provider Adapter
│
┌────────────────────┼────────────────────┐
↓ ↓ ↓
Model A Model B Model C
│ │ │
└────────────────────┼────────────────────┘
↓
Observability
│
Logs / Metrics / Trace
↓
Billing
四、第一层:Identity
所有请求首先需要知道:
谁发起?
属于哪个Tenant?
是什么角色?
使用什么Agent?
例如:
{
"tenant_id": "tenant-a",
"user_id": "u-1001",
"agent_id": "procurement-agent",
"role": "procurement_manager"
}
Gateway根据这些信息决定:
允许访问什么?
五、Tenant隔离
在AI SaaS环境中:
Tenant A
Tenant B
Tenant C
可能共享:
Gateway
Model
GPU
Tool Runtime
但必须隔离:
Data
Permission
Quota
Cost
Logs
Knowledge
Agent
Configuration
例如:
Tenant A
↓
AI Gateway
↓
A的Quota
↓
A的Model Policy
Tenant B:
Tenant B
↓
AI Gateway
↓
B的Quota
↓
B的Model Policy
因此Gateway需要携带:
tenant_id
贯穿整个请求链路。
六、第二层:Routing
这是AI Gateway最重要的能力之一。
例如:
用户问题
↓
Gateway
↓
选择模型
但是模型选择不应该简单写死:
所有请求 → Model A
而应该根据:
任务
模型能力
成本
延迟
数据敏感度
Tenant Policy
进行路由。
七、Model Routing
例如:
简单分类
↓
Small Model
复杂推理:
Complex Reasoning
↓
Large Model
企业内部敏感任务:
Sensitive Data
↓
Private Model
通用任务:
General Task
↓
Cloud Model
于是:
AI Gateway
│
┌─────────────┼─────────────┐
↓ ↓ ↓
Small Model Large Model Private Model
│ │ │
Low Cost Reasoning Internal
八、基于任务的模型路由
例如:
Task Type
可以定义:
classification
summarization
translation
chat
reasoning
coding
vision
embedding
路由:
classification
→ Small Model
summarization
→ Mid Model
reasoning
→ Large Model
sensitive
→ Private Model
这样可以避免:
所有任务都使用最大模型。
九、基于成本的路由
例如:
Model A
输入成本:低
输出成本:低
Model B
输入成本:中
输出成本:中
Model C
输入成本:高
输出成本:高
Gateway可以根据:
Task
+
Budget
+
Tenant Quota
选择模型。
例如:
月度预算剩余:
¥2,000
普通任务:
→ Model A
高价值任务:
→ Model B
超高成本任务:
→ 需要审批
这就是:
Cost-aware Routing
十、基于数据敏感度的路由
这是企业场景特别重要的一种方式。
例如:
普通文本
↓
Cloud Model
但是:
员工薪资
客户隐私
生产配方
财务数据
内部合同
可能需要:
Private Model
于是:
Request
↓
Data Classification
│
┌────────┴────────┐
↓ ↓
General Sensitive
↓ ↓
Cloud Model Private Model
这实际上是:
Data-aware Model Routing
十一、第三层:Load Balancing
企业可能部署多个模型实例:
Model A
├── Instance 1
├── Instance 2
└── Instance 3
Gateway负责:
Request
↓
Load Balancer
↓
Instance
常见策略:
Round Robin
Weighted
Least Connections
Latency Based
Health Based
十二、模型健康检查
假设:
Model A
Instance 1 ✓
Instance 2 ✓
Instance 3 ✗
Gateway应该自动:
Instance 3
↓
Remove
而不是继续发送请求。
形成:
Health Check
↓
Healthy?
├── Yes → Available
└── No → Remove
十三、第四层:Rate Limit
企业AI非常容易出现:
Agent Loop
例如Agent发生错误:
Tool
↓
Error
↓
Retry
↓
Error
↓
Retry
↓
Retry
↓
Retry
如果没有限流:
Token
↓
疯狂消耗
甚至可能:
成本暴涨
因此必须:
Rate Limit
十四、Rate Limit的维度
可以按照:
Tenant
User
Agent
API
Model
Tool
分别限制。
例如:
Tenant A
1000 req/min
User A
100 req/min
Agent A
300 req/min
Tool X
50 req/min
十五、Token Quota
除了请求次数:
1000 requests
还需要控制:
Token
例如:
Tenant A
Daily Token:
10M
达到:
80%
触发:
Warning
达到:
100%
执行:
Limit
十六、第五层:Timeout
AI请求最大的特点之一:
延迟不稳定。
例如:
Model A
1.2s
Model B
3.5s
Model C
8.7s
因此Gateway需要:
Timeout
例如:
Connect Timeout
5s
Read Timeout
30s
Total Timeout
60s
十七、第六层:Retry
模型请求可能失败:
Timeout
429
5xx
Network Error
Gateway可以执行:
Retry
但不是所有错误都应该重试。
例如:
429
→ Retry
503
→ Retry
400
→ Don't Retry
Permission Denied
→ Don't Retry
因此:
Retry必须结合Error Classification。
十八、Retry Backoff
不能:
失败
↓
立即Retry
↓
失败
↓
立即Retry
应该:
第一次:
1s
第二次:
2s
第三次:
4s
也就是:
Exponential Backoff
例如:
Retry 1 → 1s
Retry 2 → 2s
Retry 3 → 4s
同时设置:
Max Retry
避免无限重试。
十九、第七层:Fallback
这是企业AI非常关键的能力。
假设:
Primary Model
↓
Unavailable
Gateway可以:
Fallback Model
形成:
Request
↓
Model A
↓
Failure
↓
Model B
↓
Success
例如:
GPT
↓
Unavailable
↓
Claude
↓
Success
或者:
Cloud Model
↓
Network Failure
↓
Private Model
二十、Fallback需要考虑模型能力
不能简单:
A失败
↓
随便找B
因为:
模型能力
Context Window
Tool Calling
Structured Output
Vision
Reasoning
可能不同。
所以Gateway需要维护:
Model Capability Registry
例如:
Model A
✓ Tool Calling
✓ JSON
✓ Vision
Model B
✓ Tool Calling
✗ Vision
Model C
✓ JSON
✓ Reasoning
然后选择:
Compatible Model
二十一、Circuit Breaker
如果某模型连续失败:
Model A
↓
Failure
↓
Failure
↓
Failure
Gateway可以暂时:
Circuit Open
停止继续请求。
形成:
Closed
↓
Failure
↓
Open
↓
Wait
↓
Half Open
↓
Success
↓
Closed
这样可以避免:
一个故障模型拖垮整个AI平台。
二十二、第八层:Caching
有些AI请求高度重复。
例如:
"公司的报销制度是什么?"
大量用户可能重复询问。
Gateway可以:
Query
↓
Cache
↓
Hit
↓
Return
没有命中:
Cache Miss
↓
Model
↓
Response
↓
Cache
二十三、什么内容适合Cache?
适合:
FAQ
固定知识
稳定摘要
Embedding
公共配置
模型Metadata
不适合:
实时库存
实时价格
实时生产状态
高敏感个性化数据
所以Cache必须考虑:
Freshness
Security
Tenant
二十四、Tenant-aware Cache
这是Multi-Tenant系统特别容易出现的问题。
错误:
Cache Key:
"库存是多少?"
Tenant A查询:
100
Tenant B查询:
Cache Hit
→ 100
这就是严重的数据隔离问题。
正确:
tenant-a:inventory:SKU001
tenant-b:inventory:SKU001
因此:
Tenant必须成为Cache Key的一部分。
二十五、第九层:Cost Control
Enterprise AI真正进入生产后:
Cost
会变成核心问题。
例如一天:
Requests:
125,680
Tokens:
86M
Model Cost:
¥8,620
如果没有Cost Tracking:
谁花的?
哪个Agent花的?
哪个Tenant花的?
哪个Tool花的?
全部不知道。
二十六、Cost Attribution
Gateway应该记录:
Tenant
User
Agent
Task
Model
Tokens
Latency
Cost
例如:
{
"tenant": "A",
"agent": "procurement-agent",
"model": "model-x",
"input_tokens": 1200,
"output_tokens": 800,
"cost": 0.08
}
最终可以得到:
Tenant A
├── WMS Agent ¥120
├── ERP Agent ¥180
├── CRM Agent ¥90
└── Finance Agent ¥320
二十七、Cost Budget
可以设置:
Tenant Budget
Agent Budget
Task Budget
User Budget
例如:
Tenant A
Monthly Budget = ¥10,000
达到:
80%
通知。
达到:
100%
执行:
Block
或者:
Fallback to Low-Cost Model
二十八、第十层:Observability
Gateway天然是整个AI系统的观测点。
一条请求:
User
↓
Agent
↓
Gateway
↓
Model
↓
Tool
↓
ERP
Gateway可以记录:
Request ID
Trace ID
Tenant
Agent
Model
Tool
Latency
Tokens
Cost
Status
Error
这样可以形成完整Trace。
二十九、Trace
例如:
Trace ID:
trace-001
下面包含:
Span 1
Agent Request
↓
Span 2
Context Retrieval
↓
Span 3
LLM Call
↓
Span 4
Tool Call
↓
Span 5
ERP API
最终:
Total Latency = 4.8s
其中:
RAG = 0.7s
LLM = 2.1s
Tool = 1.2s
ERP = 0.8s
FDE就可以快速发现:
到底慢在哪里。
三十、Gateway Error Classification
企业AI错误可以分成:
Authentication Error
Authorization Error
Rate Limit
Timeout
Model Error
Tool Error
MCP Error
Enterprise API Error
Context Error
Policy Error
例如:
401
→ Authentication
403
→ Authorization
429
→ Rate Limit
500
→ Provider / Internal Error
不同错误进入不同处理流程。
三十一、AI Gateway完整请求链路
现在把前面的能力全部组合:
User
↓
Authentication
↓
Tenant Resolution
↓
Policy Check
↓
Request Classification
↓
Routing
↓
Quota Check
↓
Rate Limit
↓
Cache Check
↓
Load Balance
↓
Model / Agent / Tool
↓
Timeout
↓
Retry
↓
Fallback
↓
Response
↓
Cost Tracking
↓
Observability
↓
Audit
这已经不是一个简单的Proxy。
而是:
AI Traffic Control Plane。
三十二、AI Gateway与API Gateway的区别
两者并不完全相同。
| 能力 | API Gateway | AI Gateway |
|---|---|---|
| Authentication | ✓ | ✓ |
| Authorization | ✓ | ✓ |
| Rate Limit | ✓ | ✓ |
| Routing | ✓ | ✓ |
| Load Balance | ✓ | ✓ |
| Token Control | - | ✓ |
| Model Routing | - | ✓ |
| Prompt Policy | - | ✓ |
| Model Fallback | - | ✓ |
| Tool Routing | - | ✓ |
| Agent Routing | - | ✓ |
| MCP Routing | - | ✓ |
| AI Cost | - | ✓ |
| AI Evaluation | - | ✓ |
因此:
API Gateway
=
API流量管理
AI Gateway
=
AI流量 + 模型 + Agent + Tool + MCP管理
三十三、AI Gateway与MCP Gateway
两者也不是同一个概念。
AI Gateway
主要解决:
AI请求统一入口
模型
Agent
Tool
策略
成本
观测
而:
MCP Gateway
重点解决:
MCP Server
MCP Tool
Resource
Prompt
Protocol
Connection
可以理解为:
AI Gateway
↓
MCP Gateway
↓
MCP Servers
↓
Enterprise Tools
三十四、Enterprise AI Gateway完整架构
最终可以形成:
Enterprise AI
│
↓
AI Gateway
│
┌─────────────────────┼─────────────────────┐
↓ ↓ ↓
Identity Routing Policy
│ │ │
Tenant Model Route Security
User Agent Route Permission
Role Tool Route Compliance
MCP Route
└─────────────────────┼─────────────────────┘
↓
Traffic Control
│
┌──────────────────────┼──────────────────────┐
↓ ↓ ↓
Rate Limit Load Balance Quota
↓ ↓ ↓
Timeout Retry Cost
│
↓
Provider Adapter
│
┌───────────────────┼───────────────────┐
↓ ↓ ↓
Model Agent Tool
↓ ↓ ↓
LLM A/B/C Agent Mesh MCP
└───────────────────┼───────────────────┘
↓
Enterprise Systems
│
ERP / WMS / MES / CRM / OA / BI
↓
Observability
↓
Evaluation
↓
Audit
三十五、FDE如何设计AI Gateway?
当客户提出:
"我们需要接入多个模型。"
FDE不应该直接开始写:
if model == xxx
而应该首先建立:
Model Registry
Model
├── Name
├── Provider
├── Endpoint
├── Capability
├── Context Window
├── Pricing
├── Region
├── Security Level
└── Status
三十六、Agent Registry
同时建立:
Agent
├── Agent ID
├── Version
├── Tenant
├── Model Policy
├── Tool Policy
├── Knowledge
├── Cost Limit
└── Status
例如:
procurement-agent:v3
使用:
Model:
Reasoning Model
Tools:
WMS
ERP
Approval:
Required
Budget:
¥100/day
三十七、Tool Registry
Tool也应该统一管理:
Tool
├── Tool ID
├── Schema
├── Endpoint
├── Permission
├── Risk
├── Timeout
├── Retry
└── Audit
例如:
create_purchase_order
配置:
Risk:
High
Permission:
Procurement Manager
Approval:
Required
Timeout:
10s
三十八、Policy Engine
最终形成:
Request
↓
Policy Engine
↓
Allow / Deny / Review
例如:
金额 < ¥10,000
→ Auto
¥10,000 ~ ¥100,000
→ Manager Approval
> ¥100,000
→ Director Approval
这与前面的:
Human-in-the-loop
直接连接起来。
三十九、AI Gateway与Human-in-the-loop
完整流程:
Agent
↓
AI Gateway
↓
Policy Engine
↓
Risk Check
↓
需要审批?
├── No
│ ↓
│ Tool
│
└── Yes
↓
Human Approval
↓
AI Gateway
↓
Enterprise API
所以Gateway实际上成为:
AI行为与企业系统之间的重要安全边界。
四十、AI Gateway不是越复杂越好
FDE需要特别注意:
不要一开始就建设:
几十个Gateway
几百条Policy
复杂Service Mesh
复杂Multi-Agent
正确方法应该是:
Business Scenario
↓
最小可用Gateway
↓
验证
↓
Observability
↓
Cost
↓
Security
↓
Scale
先解决:
客户真正需要的问题。
四十一、从PoC到Production
PoC:
Agent
↓
LLM
Production:
Agent
↓
Context Engine
↓
AI Gateway
↓
Policy
↓
Routing
↓
Rate Limit
↓
Model
↓
Tool
↓
Enterprise System
↓
Audit
这就是FDE需要完成的:
从AI Demo到Enterprise Production。
四十二、本章核心总结
Enterprise AI Gateway主要解决:
谁可以访问?
↓
访问什么?
↓
路由到哪里?
↓
使用哪个模型?
↓
允许多少请求?
↓
允许消耗多少Token?
↓
失败怎么办?
↓
成本是多少?
↓
出了问题如何追踪?
可以浓缩成:
Identity
+
Routing
+
Policy
+
Traffic
+
Fallback
+
Cost
+
Observability
+
Audit
最终:
AI Gateway把分散的AI能力,变成统一、可控制、可观测、可运营的企业AI基础设施。
四十三、FDE能力再次升级
到第27章,FDE的能力模型已经逐渐变成:
Business
+
Software Engineering
+
AI Engineering
+
Data Engineering
+
Context Engineering
+
Runtime Engineering
+
Gateway Engineering
+
Enterprise Architecture
这时候FDE已经不只是:
AI应用开发者
而开始接近:
Enterprise AI Infrastructure Engineer。
四十四、下一章预告
《FDE前沿部署工程师实战教程》28
Enterprise AI Reliability:让Agent真正稳定运行
当我们拥有:
Context Engine
Agent Runtime
Workflow Runtime
AI Gateway
Model Gateway
Tool Gateway
MCP
Enterprise Systems
新的问题又出现:
如果Agent每天运行10万次,怎么保证它稳定?
下一章将进入:
Reliability
重点讨论:
Availability
Latency
Timeout
Retry
Circuit Breaker
Fallback
Idempotency
Queue
Backpressure
Rate Limit
Distributed Lock
State Recovery
Checkpoint
Disaster Recovery
并最终建立:
Enterprise AI
│
AI Gateway
│
Agent Runtime
│
┌───────────┼───────────┐
↓ ↓ ↓
Agent A Agent B Agent C
│ │ │
└───────────┼───────────┘
↓
Reliability Layer
│
┌─────────────────┼─────────────────┐
↓ ↓ ↓
Retry Fallback Recovery
↓ ↓ ↓
Circuit Breaker Checkpoint State Store
↓
Enterprise Systems
下一章将回答一个真正的生产环境问题:
"Agent能不能稳定运行一年,而不仅仅是Demo运行十分钟?"