LLM 应用限流与熔断机制完全指南:从多层防护架构到Java生产级弹性实战
本文深入解析 2026 年 LLM 应用限流与熔断的最新范式演进,涵盖网关层限流、Resilience4j 熔断、Bulkhead 舱壁隔离、自适应限流、成本熔断、多租户配额管理及生产级监控告警方案,帮助 Java 团队构建高可用的 LLM 弹性架构。
项目概述与应用场景
1.1 业务痛点
LLM API 调用相比普通 HTTP 请求有更严格的 QPS 限制和更高的延迟不确定性。生产环境必须在外层做多层限流/熔断/降级,避免级联故障导致整个服务不可用。
text
典型故障链路:
大流量 → LLM API 超限(429) → 重试风暴 → 全部失败 → 上游雪崩
2026 年新增痛点:
- Token 成本失控:一次意外的大流量可能消耗数万元 Token 费用
- 多模型路由复杂:OpenAI、Claude、通义千问等多提供商需要统一的配额管理
- Agent 循环风险:Agent 自主决策可能导致无限循环调用,消耗大量 Token
1.2 应用场景
| 场景 | 描述 | 核心需求 |
|---|---|---|
| 多租户 LLM 服务 | 共享 LLM 实例给多个租户 | 按租户隔离配额 |
| 高并发推理 | 突发流量冲击 | 限流 + 排队 |
| 长文本处理 | 单次调用消耗大量 Token | 成本熔断 |
| 混合模型路由 | 多 LLM 提供商 | 统一配额管理 |
| 成本管控 | Token 预算控制 | 超预算熔断 |
| Agent 循环保护(2026新增) | Agent 自主决策可能无限循环 | 最大迭代次数+Token预算 |
| MCP 工具调用限流(2026新增) | MCP 工具调用需要独立配额 | 工具级限流 |
1.3 多层防护架构
text
用户请求 → [API 网关限流] → [应用层令牌桶] → [LLM API 熔断器] → LLM
↓ ↓ ↓
每秒N请求 每分钟M请求 连续失败K次熔断
2026 年新增第四层:Token 预算熔断层------当 Token 消耗超过预算时,自动降级到小模型或返回缓存。
前置知识
- Spring Cloud Gateway / Resilience4j
- 限流算法基础(令牌桶/滑动窗口)
- 熔断器模式概念
注:
博客:
https://blog.csdn.net/badao_liumang_qizhi
技术选型与架构设计
2.1 多层防护(2026 年增强)
text
┌─────────────────────────────────────────────────────────────────────┐
│ 多层限流熔断架构 │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ ┌────────────────────────────────────────────────────────┐ │
│ │ 第1层: API 网关 (Spring Cloud Gateway) │ │
│ │ ┌──────────┐ ┌──────────┐ ┌──────────────────────┐ │ │
│ │ │Redis │ │IP/用户 │ │API Key │ │ │
│ │ │Rate │ │限流 │ │认证+鉴权 │ │ │
│ │ │Limiter │ │(全局QPS) │ │(每租户配额) │ │ │
│ │ └──────────┘ └──────────┘ └──────────────────────┘ │ │
│ └────────────────────┬───────────────────────────────────┘ │
│ │ │
│ ┌────────────────────▼───────────────────────────────────┐ │
│ │ 第2层: 应用层 (Resilience4j) │ │
│ │ ┌──────────┐ ┌──────────┐ ┌──────────────────────┐ │ │
│ │ │Circuit │ │Bulkhead │ │Rate │ │ │
│ │ │Breaker │ │(舱壁) │ │Limiter │ │ │
│ │ │(熔断器) │ │(并发隔离) │ │(令牌桶) │ │ │
│ │ └──────────┘ └──────────┘ └──────────────────────┘ │ │
│ └────────────────────┬───────────────────────────────────┘ │
│ │ │
│ ┌────────────────────▼───────────────────────────────────┐ │
│ │ 第3层: 调用层 (Http Client + Retry) │ │
│ │ ┌──────────┐ ┌──────────┐ ┌──────────────────────┐ │ │
│ │ │Retry │ │Timeout │ │Fallback │ │ │
│ │ │(重试) │ │(超时) │ │(降级) │ │ │
│ │ └──────────┘ └──────────┘ └──────────────────────┘ │ │
│ └────────────────────┬───────────────────────────────────┘ │
│ │ │
│ ┌────────────────────▼───────────────────────────────────┐ │
│ │ 第4层: Token 预算熔断层(2026新增) │ │
│ │ ┌──────────┐ ┌──────────┐ ┌──────────────────────┐ │ │
│ │ │Token │ │成本 │ │模型降级 │ │ │
│ │ │预算监控 │ │熔断 │ │(大模型→小模型) │ │ │
│ │ └──────────┘ └──────────┘ └──────────────────────┘ │ │
│ └────────────────────┬───────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────┐ │
│ │ LLM API │ │
│ └─────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────┘
2.2 限流算法对比
| 算法 | 优点 | 缺点 | 适用场景 |
|---|---|---|---|
| 令牌桶 | 允许突发流量 | 参数调优复杂 | 通用 API 限流 |
| 滑动窗口 | 精确无突发 | 内存消耗大 | 精确保留 |
| 漏桶 | 匀速处理 | 不能应对突发 | 下游保护 |
| 计数器 | 简单 | 临界时间窗口问题 | 简单场景 |
| 自适应限流(2026) | 基于 P99 延迟动态调整 | 实现复杂 | 生产级 LLM 服务 |
开发环境搭建
pom.xml
xml
<properties>
<java.version>21</java.version>
<resilience4j.version>2.3.0</resilience4j.version>
<spring-boot.version>3.5.0</spring-boot.version>
</properties>
<dependencies>
<!-- Web -->
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-web</artifactId>
</dependency>
<!-- WebFlux -->
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-webflux</artifactId>
</dependency>
<!-- Gateway -->
<dependency>
<groupId>org.springframework.cloud</groupId>
<artifactId>spring-cloud-starter-gateway</artifactId>
</dependency>
<!-- Resilience4j -->
<dependency>
<groupId>io.github.resilience4j</groupId>
<artifactId>resilience4j-spring-boot3</artifactId>
<version>${resilience4j.version}</version>
</dependency>
<dependency>
<groupId>io.github.resilience4j</groupId>
<artifactId>resilience4j-all</artifactId>
<version>${resilience4j.version}</version>
</dependency>
<dependency>
<groupId>io.github.resilience4j</groupId>
<artifactId>resilience4j-reactor</artifactId>
<version>${resilience4j.version}</version>
</dependency>
<!-- Redis -->
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-data-redis</artifactId>
</dependency>
<!-- Caffeine Cache -->
<dependency>
<groupId>com.github.ben-manes.caffeine</groupId>
<artifactId>caffeine</artifactId>
</dependency>
<!-- Micrometer -->
<dependency>
<groupId>io.micrometer</groupId>
<artifactId>micrometer-registry-prometheus</artifactId>
</dependency>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-actuator</artifactId>
</dependency>
</dependencies>
application.yml
yaml
server:
port: 8095
spring:
application:
name: llm-resilience-service
cloud:
gateway:
routes:
- id: chat-service
uri: http://localhost:8080
predicates:
- Path=/api/v1/chat/**
filters:
- name: RequestRateLimiter
args:
redis-rate-limiter.replenishRate: 100
redis-rate-limiter.burstCapacity: 200
key-resolver: "#{@userKeyResolver}"
resilience4j:
circuitbreaker:
instances:
llmApi:
registerHealthIndicator: true
slidingWindowType: COUNT_BASED
slidingWindowSize: 10
minimumNumberOfCalls: 5
permittedNumberOfCallsInHalfOpenState: 5
automaticTransitionFromOpenToHalfOpenEnabled: true
waitDurationInOpenState: 30s
failureRateThreshold: 50
slowCallRateThreshold: 80
slowCallDurationThreshold: 5s
recordExceptions:
- java.util.concurrent.TimeoutException
- org.springframework.web.reactive.function.client.WebClientResponseException$ServiceUnavailable
ignoreExceptions:
- com.example.llm.exception.AuthenticationException
- com.example.llm.exception.ContentViolationException
ratelimiter:
instances:
llmApi:
registerHealthIndicator: true
limitForPeriod: 60
limitRefreshPeriod: 1s
timeoutDuration: 3s
bulkhead:
instances:
llmApi:
maxConcurrentCalls: 50
maxWaitDuration: 500ms
retry:
instances:
llmApi:
maxAttempts: 3
waitDuration: 1s
enableExponentialBackoff: true
exponentialBackoffMultiplier: 2
timelimiter:
instances:
llmApi:
timeoutDuration: 30s
cancelRunningFuture: true
完整代码实现
2.3 API 网关层限流
java
@Configuration
public class GatewayRateLimitConfig {
@Bean
public RedisRateLimiter redisRateLimiter() {
return new RedisRateLimiter(100, 200);
}
@Bean
public KeyResolver userKeyResolver() {
return exchange -> {
String userId = exchange.getRequest().getHeaders().getFirst("X-User-Id");
if (userId != null) return Mono.just(userId);
String apiKey = exchange.getRequest().getHeaders().getFirst("X-Api-Key");
if (apiKey != null) return Mono.just("apikey:" + apiKey);
return Mono.just("ip:" + exchange.getRequest()
.getRemoteAddress().getAddress().getHostAddress());
};
}
}
2.4 Resilience4j 熔断器配置
java
@Configuration
public class LlmCircuitBreakerConfig {
@Bean
public CircuitBreaker llmCircuitBreaker() {
CircuitBreakerConfig config = CircuitBreakerConfig.custom()
.failureRateThreshold(50)
.slowCallRateThreshold(80)
.slowCallDurationThreshold(Duration.ofSeconds(5))
.waitDurationInOpenState(Duration.ofSeconds(30))
.permittedNumberOfCallsInHalfOpenState(5)
.slidingWindowType(SlidingWindowType.COUNT_BASED)
.slidingWindowSize(10)
.minimumNumberOfCalls(5)
.automaticTransitionFromOpenToHalfOpenEnabled(true)
.recordExceptions(
TimeoutException.class,
IOException.class,
ResourceAccessException.class)
.ignoreExceptions(
AuthenticationException.class,
ContentViolationException.class,
RateLimitExceededException.class)
.build();
CircuitBreakerRegistry registry = CircuitBreakerRegistry.ofDefault();
registry.circuitBreaker("llmApi", config);
return registry.circuitBreaker("llmApi");
}
}
2.5 综合弹性保护 LLM 服务
java
@Service
@Slf4j
public class ResilientLlmService {
private final LlmClientService llClient;
private final Cache<String, String> fallbackCache;
private final MeterRegistry meterRegistry;
/**
* 完整的弹性调用:限流→熔断→舱壁→调用→降级
*/
public String chat(String prompt, String userId) {
long startTime = System.currentTimeMillis();
String result = "";
try {
// 1. 应用层限流检查 (RateLimiter)
RateLimiter rateLimiter = RateLimiterRegistry.ofDefault()
.rateLimiter("llmApi");
boolean permissionGranted = rateLimiter.acquirePermission();
if (!permissionGranted) {
meterRegistry.counter("llm.rate_limited", "user", userId).increment();
return handleRateLimited(prompt);
}
// 2. 舱壁 + 熔断 + 重试 + 超时 → 调用
CircuitBreaker circuitBreaker = CircuitBreakerRegistry.ofDefault()
.circuitBreaker("llmApi");
Bulkhead bulkhead = BulkheadRegistry.ofDefault().bulkhead("llmApi");
Retry retry = RetryRegistry.ofDefault().retry("llmApi");
Callable<String> decorated = Decorators
.ofCallable(() -> doCall(prompt))
.withCircuitBreaker(circuitBreaker)
.withBulkhead(bulkhead)
.withRateLimiter(rateLimiter)
.withRetry(retry)
.decorate();
result = decorated.call();
meterRegistry.counter("llm.call.success", "user", userId).increment();
} catch (RequestNotPermitted e) {
log.warn("熔断器打开 [{}/{}], 执行降级: user={}", userId, prompt.hashCode());
result = handleCircuitBroken(prompt);
meterRegistry.counter("llm.circuit_open", "user", userId).increment();
} catch (BulkheadFullException e) {
log.warn("并发调用过多,执行降级: user={}", userId);
result = handleBulkheadFull(prompt);
meterRegistry.counter("llm.bulkhead_full", "user", userId).increment();
} catch (Exception e) {
log.error("LLM 调用异常: {}", e.getMessage());
result = handleGenericError(prompt, e);
meterRegistry.counter("llm.call.error",
"type", e.getClass().getSimpleName()).increment();
} finally {
long elapsed = System.currentTimeMillis() - startTime;
meterRegistry.timer("llm.call.duration", "user", userId)
.record(elapsed, TimeUnit.MILLISECONDS);
}
return result;
}
private String doCall(String prompt) {
return llClient.generate(prompt);
}
private String handleRateLimited(String prompt) {
String key = DigestUtils.md5Hex(prompt);
String cached = fallbackCache.getIfPresent(key);
if (cached != null) return cached;
throw new RateLimitExceededException("请求过于频繁,请稍后重试");
}
private String handleCircuitBroken(String prompt) {
return "AI 服务暂时繁忙,我们正在恢复中。请稍后重试。";
}
private String handleBulkheadFull(String prompt) {
return "系统当前请求过多,建议您稍后重试或排队等待。";
}
private String handleGenericError(String prompt, Exception e) {
return "请求处理中发生错误,稍后会自动恢复。如持续异常请联系管理员。";
}
}
2.6 自适应限流
java
/**
* 基于 P99 延迟的自适应限流 --- 自动调整限流阈值
*/
@Service
@Slf4j
public class AdaptiveRateLimiterService {
private final AtomicInteger currentLimit = new AtomicInteger(100);
private final EvictingQueue<Long> latencyWindow = EvictingQueue.create(100);
private final MeterRegistry meterRegistry;
public void recordLatency(long latencyMs) {
latencyWindow.add(latencyMs);
}
@Scheduled(fixedDelay = 5000)
public void adjustLimit() {
List<Long> latencies = new ArrayList<>(latencyWindow);
if (latencies.size() < 10) return;
List<Long> sorted = latencies.stream().sorted().toList();
long p99 = sorted.get((int)(sorted.size() * 0.99));
int oldLimit = currentLimit.get();
int newLimit;
String action;
if (p99 > 10000) {
newLimit = Math.max(10, oldLimit / 4);
action = "SHRINK_AGGRESSIVE";
} else if (p99 > 5000) {
newLimit = Math.max(20, oldLimit / 2);
action = "SHRINK";
} else if (p99 > 3000) {
newLimit = (int)(oldLimit * 0.8);
action = "SHRINK_SLIGHT";
} else if (p99 < 1000 && oldLimit < 500) {
newLimit = oldLimit + 10;
action = "EXPAND";
} else {
return;
}
currentLimit.set(newLimit);
log.info("自适应限流调整: p99={}ms, limit: {} → {}, action={}",
p99, oldLimit, newLimit, action);
meterRegistry.gauge("adaptive_rate_limit.value", newLimit);
meterRegistry.counter("adaptive_rate_limit.adjustments",
"action", action).increment();
}
public int getCurrentLimit() { return currentLimit.get(); }
}
2.7 Token 预算熔断
java
/**
* Token 预算熔断器 --- 当 Token 消耗超过预算时自动降级
*/
@Service
@Slf4j
public class TokenBudgetCircuitBreaker {
private final RedisTemplate<String, Long> redisTemplate;
private final Map<String, Long> dailyBudget = Map.of(
"gpt-4o", 1_000_000L, // 大模型日预算
"qwen-turbo", 10_000_000L // 小模型日预算
);
/**
* 检查并扣减 Token 预算
*/
public BudgetCheckResult checkAndConsume(String model, long estimatedTokens) {
String today = LocalDate.now().toString();
String key = "token:budget:" + model + ":" + today;
Long used = redisTemplate.opsForValue().get(key);
used = used != null ? used : 0L;
long budget = dailyBudget.getOrDefault(model, 1_000_000L);
if (used + estimatedTokens > budget) {
log.warn("Token 预算超支: model={}, used={}, budget={}",
model, used, budget);
return BudgetCheckResult.exceeded(
suggestFallbackModel(model), budget - used);
}
redisTemplate.opsForValue().increment(key, estimatedTokens);
redisTemplate.expire(key, Duration.ofDays(2));
return BudgetCheckResult.allowed(budget - used - estimatedTokens);
}
/**
* 降级建议:大模型超预算时建议切换到小模型
*/
private String suggestFallbackModel(String model) {
return switch (model) {
case "gpt-4o", "claude-3-opus" -> "qwen-turbo";
case "qwen-max" -> "qwen-plus";
default -> "qwen-turbo";
};
}
public record BudgetCheckResult(boolean allowed, String fallbackModel,
long remaining) {
static BudgetCheckResult allowed(long remaining) {
return new BudgetCheckResult(true, null, remaining);
}
static BudgetCheckResult exceeded(String fallback, long remaining) {
return new BudgetCheckResult(false, fallback, remaining);
}
}
}
2.8 Agent 循环保护
java
/**
* Agent 循环保护 --- 防止 Agent 自主决策导致无限循环
*/
@Service
public class AgentLoopGuard {
private static final int MAX_ITERATIONS = 10;
private static final long MAX_TOKENS_PER_SESSION = 50_000;
private final RedisTemplate<String, Integer> redisTemplate;
/**
* 检查 Agent 是否应该继续循环
*/
public LoopGuardResult checkLoop(String sessionId, int currentIteration,
long tokensUsed) {
// 1. 检查迭代次数
if (currentIteration >= MAX_ITERATIONS) {
log.warn("Agent 循环次数超限: session={}, iterations={}",
sessionId, currentIteration);
return LoopGuardResult.stop("MAX_ITERATIONS_REACHED");
}
// 2. 检查 Token 预算
if (tokensUsed >= MAX_TOKENS_PER_SESSION) {
log.warn("Agent Token 预算耗尽: session={}, tokens={}",
sessionId, tokensUsed);
return LoopGuardResult.stop("TOKEN_BUDGET_EXHAUSTED");
}
// 3. 检查循环检测(相同工具调用重复)
String key = "agent:loop:" + sessionId;
Integer repeatCount = redisTemplate.opsForValue().get(key);
if (repeatCount != null && repeatCount >= 3) {
log.warn("Agent 重复调用检测: session={}", sessionId);
return LoopGuardResult.stop("REPEATED_CALL_DETECTED");
}
return LoopGuardResult.continueLoop();
}
public record LoopGuardResult(boolean shouldContinue, String reason) {
static LoopGuardResult continueLoop() {
return new LoopGuardResult(true, null);
}
static LoopGuardResult stop(String reason) {
return new LoopGuardResult(false, reason);
}
}
}
2.9 滑动窗口限流(Redis Lua 脚本)
java
@Service
@Slf4j
public class SlidingWindowRateLimiter {
private final StringRedisTemplate redisTemplate;
private static final String LUA_SCRIPT = """
local key = KEYS[1]
local now = tonumber(ARGV[1])
local window = tonumber(ARGV[2])
local limit = tonumber(ARGV[3])
redis.call('ZREMRANGEBYSCORE', key, 0, now - window)
local current = redis.call('ZCARD', key)
if current < limit then
redis.call('ZADD', key, now, now .. ':' .. math.random())
redis.call('EXPIRE', key, window / 1000)
return 1
else
return 0
end
""";
private final RedisScript<Long> redisScript =
new DefaultRedisScript<>(LUA_SCRIPT, Long.class);
public boolean tryAcquire(String key, int limit, long windowMs) {
long now = System.currentTimeMillis();
Long result = redisTemplate.execute(redisScript, List.of(key),
String.valueOf(now), String.valueOf(windowMs), String.valueOf(limit));
return result != null && result == 1;
}
public boolean tryAcquireByTenant(String tenantId) {
return tryAcquire("ratelimit:tenant:" + tenantId, 60, 60_000);
}
public boolean tryAcquireByUser(String userId) {
return tryAcquire("ratelimit:user:" + userId, 10, 10_000);
}
public boolean tryAcquireByIp(String ip) {
return tryAcquire("ratelimit:ip:" + ip, 30, 60_000);
}
}
2.10 注解驱动限流(AOP 切面)
java
@Target(ElementType.METHOD)
@Retention(RetentionPolicy.RUNTIME)
public @interface RateLimited {
String key() default "";
int limit() default 100;
long window() default 60;
LimitType type() default LimitType.TOTAL;
enum LimitType { TOTAL, PER_USER, PER_IP, PER_TENANT }
}
@Aspect
@Component
@Slf4j
public class RateLimitAspect {
private final SlidingWindowRateLimiter rateLimiter;
private final ExpressionParser parser = new SpelExpressionParser();
@Around("@annotation(rateLimited)")
public Object around(ProceedingJoinPoint pjp, RateLimited rateLimited)
throws Throwable {
String key = resolveKey(pjp, rateLimited.key());
String fullKey = "limit:" + key + ":" + pjp.getSignature().getName();
boolean acquired = rateLimiter.tryAcquire(fullKey,
rateLimited.limit(), rateLimited.window() * 1000L);
if (!acquired) {
log.warn("方法 {} 触发限流, key={}",
pjp.getSignature().getName(), key);
throw new RateLimitExceededException(
String.format("调用过于频繁, %d 秒内最多 %d 次",
rateLimited.window(), rateLimited.limit()));
}
return pjp.proceed();
}
private String resolveKey(ProceedingJoinPoint pjp, String spel) {
if (spel.isEmpty()) return "default";
MethodSignature methodSig = (MethodSignature) pjp.getSignature();
StandardEvaluationContext context = new StandardEvaluationContext();
Object[] args = pjp.getArgs();
String[] paramNames = methodSig.getParameterNames();
for (int i = 0; i < paramNames.length; i++) {
context.setVariable(paramNames[i], args[i]);
}
return parser.parseExpression(spel).getValue(context, String.class);
}
}
运行与测试
bash
# 启动
mvn spring-boot:run
# 测试限流
for i in {1..10}; do curl -X POST http://localhost:8095/api/v1/chat/generate \
-d '{"prompt":"hello"}' & done
# 模拟 40% 失败触发熔断
# 观察熔断器打开 → 半开 → 关闭
生产运维与案例分析
5.1 监控指标
| 指标 | 含义 | 告警阈值 |
|---|---|---|
llm.call.success |
调用成功次数 | --- |
llm.call.error |
调用失败次数 | 错误率 > 5% |
llm.rate_limited |
限流触发次数 | 触发率 > 10% |
llm.circuit_open |
熔断触发次数 | > 0(需立即告警) |
llm.bulkhead_full |
舱壁满触发次数 | 触发率 > 5% |
llm.call.duration |
调用延迟分布 | P99 > 30s |
adaptive_rate_limit.value |
当前自适应限流值 | 持续下降 |
token.budget.remaining |
Token 预算剩余 | < 20% |
5.2 案例:某金融平台 LLM 限流实践
背景:金融平台使用 LLM 提供智能客服,日均调用 50 万次,高峰期 QPS 达 2000。
挑战:
- LLM API 限流 429 频繁触发
- 重试风暴导致级联故障
- Token 成本月度超支 40%
解决方案:
| 层次 | 措施 | 效果 |
|---|---|---|
| 网关层 | Redis 令牌桶限流,按租户隔离 | 429 错误率下降 85% |
| 应用层 | Resilience4j 熔断 + 舱壁 | 级联故障消除 |
| 成本层 | Token 预算熔断 + 模型降级 | 成本降低 35% |
| 自适应 | P99 延迟动态调整限流阈值 | 高峰期稳定性提升 |
效果:
- 系统可用性从 99.5% 提升到 99.95%
- Token 成本从超支 40% 变为节省 15%
- 平均响应延迟降低 30%
5.3 常见故障排查
| 现象 | 原因 | 解法 |
|---|---|---|
| 熔断器频繁打开 | 失败率阈值设置过低 | 调整 failureRateThreshold |
| 限流过于严格 | limitForPeriod 设置过低 |
根据实际 QPS 调整 |
| 降级返回过多 | 舱壁容量不足 | 增加 maxConcurrentCalls |
| Token 预算频繁超支 | 大模型调用过多 | 配置模型路由规则 |
| 自适应限流震荡 | 调整步长过大 | 使用更平滑的调整策略 |
总结
LLM 应用限流与熔断机制的关键技术点:
| 维度 | 2026 年实践 | 核心价值 |
|---|---|---|
| 网关层限流 | Redis 令牌桶 + 多维度 Key Resolver | 按用户/租户/IP 隔离 |
| 熔断器配置 | 失败率 + 慢调用双重熔断条件 | 快速隔离故障 |
| 舱壁隔离 | Bulkhead 限制并发调用数 | 防止雪崩 |
| 自适应限流 | 基于 P99 延迟动态调整阈值 | 高峰期稳定性 |
| Token 预算熔断 | 日预算 + 模型降级 | 成本控制 |
| Agent 循环保护 | 最大迭代 + Token 预算 + 重复检测 | 防止无限循环 |
| 多级降级 | 快速失败 + 旧缓存 + 兜底话术 | 用户体验保障 |
| 注解驱动 | AOP 声明式限流 | 业务代码零侵入 |
参考资源: