一句话摘要:LLM API 调用不是普通 HTTP 请求,它的长超时、流式响应和高并发特性会让默认连接池配置在生产中悄悄制造 P99 延迟劣化。本文从真实事故出发,系统讲解连接池核心参数、Keep-Alive 陷阱、连接泄漏检测与高并发调优实践。
背景:一次"莫名其妙"的 P99 劣化
某团队的 AI 写作助手在日活 5 万后开始出现奇怪的现象:平均延迟 1.2s,P99 却飙到 18s,而且这个问题只在早晚高峰出现,低峰期完全正常。
他们排查了一圈:模型 API 端的延迟监控显示正常,自己服务的 CPU/Memory 没异常,日志里没有明显报错。最后用 netstat 一看:
bash
$ netstat -an | grep 443 | grep -E 'ESTABLISHED|CLOSE_WAIT|TIME_WAIT' | wc -l
847
CLOSE_WAIT 状态的连接积压了 600+ 个。问题找到了------连接池配置完全没针对 LLM 场景调整过,用的是框架默认值。
这类问题的典型特征:
- 表现为 P99/P999 劣化,平均延迟看起来正常
- 高并发时才出现,压测低并发不复现
- 监控里看不到明显错误,只是"慢"
为什么 LLM 调用对连接池特别敏感
LLM API 调用和普通 REST 请求有三个核心差异,每一个都会影响连接池行为:
1. 超时时间极长
普通 API:超时 1-5 秒
LLM API(非流式):超时 30-120 秒
LLM API(流式,等第一个 token):超时 30-60 秒
LLM API(流式,总持续时间):可能 3-10 分钟
超时越长,连接被"占用"的时间越长。如果连接池上限是 10,10 个并发请求就能把连接池打满,后续请求开始排队等待。
2. 流式响应占用连接更久
普通请求:连接占用时间 ≈ 服务端处理时间(100ms~2s)
流式请求:连接占用时间 ≈ 服务端处理时间 + 全部 token 传输时间(可能 20-60s)
一个流式请求,从发出到读完最后一个 token,连接始终被占用。如果你的应用有 20% 的用户在用流式模式,这 20% 的请求会消耗不成比例的连接资源。
3. 并发峰值集中
AI 写作、AI 搜索这类应用有明显的早晚高峰。用户集中在同一时间触发请求,同时打满连接池的概率远高于分散均匀的微服务场景。
连接池的核心参数
以 Python httpx 为例(Node.js undici、Go net/http 的参数名不同,但概念一致):
python
import httpx
# 默认配置(危险!)
client = httpx.AsyncClient() # 等价于下方注释的默认值
# 显式配置(推荐)
client = httpx.AsyncClient(
limits=httpx.Limits(
max_connections=100, # 总连接上限(默认 100,但要结合实际调整)
max_keepalive_connections=20, # Keep-Alive 连接池大小(默认 20)
keepalive_expiry=30, # Keep-Alive 连接的最大空闲时间(秒)
),
timeout=httpx.Timeout(
connect=5.0, # TCP 握手超时
read=120.0, # 等待响应数据超时(非流式要设长)
write=10.0, # 发送请求体超时
pool=10.0, # 等待从连接池获取连接的超时(极重要!)
),
)
最容易被忽略的是 pool 超时 :它控制"排队等待连接池空出一个连接"的最大时间。如果不设,默认可能是 None(无限等待),高并发时请求会无限堆积,表现为 P99 劣化但没有报错。
Node.js undici 配置
javascript
import { Pool } from 'undici';
import OpenAI from 'openai'; // 示例:替换为你实际使用的 SDK
// undici Pool 是大模型 SDK Node.js 版本的底层实现
const pool = new Pool('https://api.your-llm-provider.com', {
connections: 50, // 最大连接数
pipelining: 1, // LLM 场景建议为 1(不用 pipeline)
keepAliveTimeout: 30_000, // Keep-Alive 超时(ms)
keepAliveMaxTimeout: 600_000, // Keep-Alive 最长存活(ms,流式场景设长)
headersTimeout: 10_000, // 等待响应头超时(ms)
bodyTimeout: 120_000, // 等待响应体超时(非流式,ms)
connectTimeout: 5_000, // TCP 连接超时(ms)
});
// 将 Pool 作为 fetch 的底层传给 SDK
const client = new OpenAI({
baseURL: 'https://api.your-llm-provider.com/v1',
fetch: (url, options) => pool.fetch(url, options),
});
Go net/http 配置
go
import (
"net/http"
"time"
"github.com/anthropics/anthropic-sdk-go"
)
transport := &http.Transport{
MaxIdleConns: 100, // 全局最大 Keep-Alive 连接
MaxIdleConnsPerHost: 50, // 每个 Host 的最大 Keep-Alive 连接
MaxConnsPerHost: 100, // 每个 Host 的最大总连接(含活跃)
IdleConnTimeout: 90 * time.Second, // Keep-Alive 空闲超时
TLSHandshakeTimeout: 5 * time.Second,
ResponseHeaderTimeout: 30 * time.Second, // 等响应头超时
// DisableKeepAlives: false, // 默认 false,不要改成 true
}
httpClient := &http.Client{
Transport: transport,
Timeout: 0, // 流式场景设 0 = 无总超时,靠上层 context 控制
}
client := anthropic.NewClient(
option.WithHTTPClient(httpClient),
)
Keep-Alive 的三个常见陷阱
陷阱 1:服务端先关闭连接,客户端不知道
这是 CLOSE_WAIT 积压的直接原因。
LLM API 服务端通常有自己的 Keep-Alive 超时(比如 60 秒不活动就关连接)。客户端的连接池以为连接还活着,把它放在池里复用,但实际上服务端已经关闭了。下一次用这个连接发请求时:
客户端发 TCP 段 → 服务端返回 RST(连接已关)→ 客户端收到 ConnectionResetError
解决方案 :客户端的 keepalive_expiry 要比服务端的超时短 10-20 秒。
如果你不知道服务端的 Keep-Alive 超时是多少,用保守值:
python
# httpx:设置 20 秒(通常比 API 服务端的 30-60s 超时短)
limits=httpx.Limits(keepalive_expiry=20)
javascript
// undici:设置 20 秒
keepAliveTimeout: 20_000
陷阱 2:流式响应期间连接被误判为空闲
某些连接池实现会把"正在等待下一个 SSE chunk"的连接误判为"空闲超时",提前关闭它,导致流式响应中断。
复现方式:
python
async with client.stream("POST", url, json=payload) as response:
async for chunk in response.aiter_bytes():
# 模拟慢消费(客户端处理 chunk 耗时)
await asyncio.sleep(2) # 如果 keepalive_expiry < 2,连接可能被杀掉
process(chunk)
解决方案 :流式请求的 keepalive_expiry 要比预期的 chunk 间隔长,或者为流式请求单独维护一个 client 实例:
python
# 非流式 client:激进配置,快速回收
sync_client = httpx.AsyncClient(
limits=httpx.Limits(
max_keepalive_connections=20,
keepalive_expiry=20,
),
timeout=httpx.Timeout(read=60.0, pool=5.0),
)
# 流式 client:保守配置,允许长连接
stream_client = httpx.AsyncClient(
limits=httpx.Limits(
max_keepalive_connections=10,
keepalive_expiry=300, # 5 分钟,覆盖长流
),
timeout=httpx.Timeout(read=None, pool=10.0), # read=None 表示不超时
)
陷阱 3:连接泄漏------用了但没还
流式响应最容易发生连接泄漏。原因是读取流时抛了异常,但没有正确关闭连接:
python
# ❌ 危险写法:异常时连接可能泄漏
async with client.stream("POST", url, json=payload) as response:
async for chunk in response.aiter_bytes():
result = json.loads(chunk) # 如果这里抛 JSONDecodeError,连接泄漏!
process(result)
# ✅ 正确写法:确保异常时也关闭连接
try:
async with client.stream("POST", url, json=payload) as response:
async for chunk in response.aiter_bytes():
try:
result = json.loads(chunk)
process(result)
except json.JSONDecodeError:
logger.warning("Invalid JSON chunk, skipping")
continue
except httpx.ReadTimeout:
logger.error("Stream read timeout")
raise
httpx 的 async with client.stream(...) 实际上会在上下文管理器退出时调用 response.aclose(),但如果你在 async for 外层 break 了,要手动关:
python
response = await client.send(request, stream=True)
try:
async for chunk in response.aiter_bytes():
if should_stop:
break # break 不会自动关闭!
process(chunk)
finally:
await response.aclose() # 必须显式关闭
连接泄漏的检测方法
方法 1:暴露连接池状态指标
python
# httpx 提供了连接池状态查询
import httpx
import asyncio
client = httpx.AsyncClient(
limits=httpx.Limits(max_connections=50, max_keepalive_connections=20)
)
async def get_pool_metrics():
pool = client._transport._pool
return {
"active": len(pool._requests), # 正在使用的连接数
"keepalive": len(pool._keepalive_connections), # Keep-Alive 池里的连接数
"connecting": len([c for c in pool._connections if c._connect_failed is False]),
}
# 定期采集,写入 Prometheus metrics
async def collect_pool_metrics():
while True:
metrics = await get_pool_metrics()
gauge_active_connections.set(metrics["active"])
gauge_keepalive_connections.set(metrics["keepalive"])
await asyncio.sleep(10)
方法 2:操作系统级监控
bash
# 查看进程的连接状态分布
PID=$(pgrep -f "python app.py")
ss -p -n | grep "pid=$PID" | awk '{print $2}' | sort | uniq -c | sort -rn
# 输出示例:
# 847 CLOSE_WAIT ← 这个数量在增长就是泄漏
# 23 ESTABLISHED
# 12 TIME_WAIT
# 实时监控
watch -n 5 "ss -p -n | grep 'pid=$PID' | awk '{print \$2}' | sort | uniq -c"
方法 3:周期性连接数告警
python
import asyncio
import logging
async def connection_leak_detector(client: httpx.AsyncClient, threshold: int = 40):
"""当 CLOSE_WAIT 连接超过阈值时告警"""
while True:
await asyncio.sleep(30)
try:
pool = client._transport._pool
# httpx 内部 API,版本升级可能变化,做好异常处理
active = len(getattr(pool, '_requests', []))
if active > threshold:
logging.warning(
f"Connection pool pressure: {active} active connections "
f"(threshold={threshold}). Possible leak or overload."
)
except Exception:
pass # 内部 API 访问失败不影响主逻辑
高并发场景的连接池调优
基准:按 QPS 和平均持续时间计算所需连接数
最小连接数 = QPS × 平均响应时间(秒)
安全系数 × 1.5(应对峰值)
示例:
- QPS = 50
- 平均响应时间 = 3s(非流式)
- 最小连接数 = 50 × 3 = 150
- 加安全系数:150 × 1.5 = 225
如果你的连接池 max_connections=100,50 QPS 下就会出现排队等待,P99 飙升。
流式场景调整:
- 流式平均持续时间 = 20s(输出 1000 tokens @ 50 token/s)
- QPS = 20(流式并发通常低于非流式)
- 最小连接数 = 20 × 20 = 400
这意味着流式场景需要显著更大的连接池,或者对流式请求的并发数做独立限流:
python
# 对流式请求独立限流,避免它们耗尽连接池
stream_semaphore = asyncio.Semaphore(30) # 最多 30 个并发流式请求
async def streaming_llm_call(prompt: str):
async with stream_semaphore:
async with stream_client.stream("POST", url, json={...}) as response:
async for chunk in response.aiter_bytes():
yield chunk
实际调优的参数清单
| 参数 | 默认值 | 非流式推荐 | 流式推荐 | 说明 |
|---|---|---|---|---|
| max_connections | 100 | QPS×响应时间×1.5 | QPS×响应时间×2 | 按计算公式设 |
| max_keepalive_connections | 20 | 50~100 | 10~30 | 流式长连接占资源,少设 |
| keepalive_expiry | 5s | 20~30s | 120~300s | 短于服务端超时 10s |
| pool 超时 | None | 5~10s | 10~20s | 必须设,防止无限排队 |
| read 超时 | 5s | 60~120s | None | 流式设 None,靠 context 控 |
连接池监控 Dashboard 关键指标
python
# 建议暴露的 Prometheus metrics
from prometheus_client import Gauge, Histogram
llm_pool_active = Gauge('llm_pool_active_connections', 'Active LLM connections')
llm_pool_keepalive = Gauge('llm_pool_keepalive_connections', 'Keepalive LLM connections')
llm_pool_wait_time = Histogram(
'llm_pool_wait_seconds',
'Time waiting for a connection from the pool',
buckets=[0.01, 0.05, 0.1, 0.5, 1.0, 5.0, 10.0]
)
# 在请求包装器里采集
async def llm_request_with_metrics(client, *args, **kwargs):
wait_start = time.monotonic()
# pool 超时会在这里抛 PoolTimeout,记录为等待时间异常
response = await client.request(*args, **kwargs)
llm_pool_wait_time.observe(time.monotonic() - wait_start)
return response
一个完整的生产级 LLM Client 封装
将上述实践整合成可直接使用的封装:
python
import asyncio
import httpx
import logging
import time
from contextlib import asynccontextmanager
from typing import AsyncIterator
from prometheus_client import Gauge, Histogram, Counter
logger = logging.getLogger(__name__)
# Prometheus metrics
pool_active = Gauge('llm_http_pool_active', 'Active connections')
pool_wait = Histogram('llm_http_pool_wait_seconds', 'Pool wait time',
buckets=[.01, .05, .1, .5, 1., 5., 10.])
pool_timeout_total = Counter('llm_http_pool_timeout_total', 'Pool timeout count')
stream_leak_total = Counter('llm_http_stream_leak_total', 'Stream connection leak events')
class LLMHttpClient:
"""生产级 LLM HTTP 客户端,内置连接池管理、泄漏检测与 metrics"""
def __init__(
self,
base_url: str,
api_key: str,
max_connections: int = 100,
max_keepalive: int = 30,
keepalive_expiry: float = 25.0,
pool_timeout: float = 8.0,
read_timeout: float = 90.0,
stream_max_connections: int = 40,
stream_keepalive_expiry: float = 300.0,
):
self.base_url = base_url
self.headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
}
# 非流式 client:快速回收,严格超时
self._sync_client = httpx.AsyncClient(
base_url=base_url,
headers=self.headers,
limits=httpx.Limits(
max_connections=max_connections,
max_keepalive_connections=max_keepalive,
keepalive_expiry=keepalive_expiry,
),
timeout=httpx.Timeout(
connect=5.0,
read=read_timeout,
write=10.0,
pool=pool_timeout,
),
)
# 流式 client:宽松 keepalive,read 超时由上层 context 控制
self._stream_client = httpx.AsyncClient(
base_url=base_url,
headers=self.headers,
limits=httpx.Limits(
max_connections=stream_max_connections,
max_keepalive_connections=10,
keepalive_expiry=stream_keepalive_expiry,
),
timeout=httpx.Timeout(
connect=5.0,
read=None, # 流式不设 read 超时,靠 asyncio.timeout 控制
write=10.0,
pool=pool_timeout + 5.0,
),
)
# 流式并发限制
self._stream_sem = asyncio.Semaphore(stream_max_connections)
async def post(self, path: str, json: dict) -> dict:
"""非流式请求"""
wait_start = time.monotonic()
try:
resp = await self._sync_client.post(path, json=json)
pool_wait.observe(time.monotonic() - wait_start)
resp.raise_for_status()
return resp.json()
except httpx.PoolTimeout:
pool_timeout_total.inc()
logger.error("LLM pool timeout on non-stream request")
raise
@asynccontextmanager
async def stream(
self, path: str, json: dict, total_timeout: float = 120.0
) -> AsyncIterator[httpx.Response]:
"""流式请求,内置并发限制和泄漏防护"""
async with self._stream_sem:
try:
async with asyncio.timeout(total_timeout):
async with self._stream_client.stream(
"POST", path, json=json
) as response:
response.raise_for_status()
yield response
except httpx.PoolTimeout:
pool_timeout_total.inc()
logger.error("LLM pool timeout on stream request")
raise
except Exception:
stream_leak_total.inc() # 非正常退出计数
raise
async def aclose(self):
await self._sync_client.aclose()
await self._stream_client.aclose()
async def health(self) -> dict:
"""连接池健康状态,用于 /healthz 接口"""
def _pool_info(client):
try:
pool = client._transport._pool
return {
"active": len(getattr(pool, '_requests', [])),
"keepalive": len(getattr(pool, '_keepalive_connections', [])),
}
except Exception:
return {"active": -1, "keepalive": -1}
return {
"sync_pool": _pool_info(self._sync_client),
"stream_pool": _pool_info(self._stream_client),
}
# 单例,作为应用级共享 client
_llm_client: LLMHttpClient | None = None
def get_llm_client() -> LLMHttpClient:
global _llm_client
if _llm_client is None:
raise RuntimeError("LLMHttpClient not initialized. Call init_llm_client() first.")
return _llm_client
def init_llm_client(api_key: str, **kwargs) -> LLMHttpClient:
global _llm_client
_llm_client = LLMHttpClient(
base_url="https://api.your-llm-provider.com",
api_key=api_key,
**kwargs,
)
return _llm_client
事故复盘:参数调整前后对比
回到文章开头的案例,这是他们调整前后的连接池配置和效果:
调整前(httpx 默认值):
python
client = httpx.AsyncClient()
# max_connections=100, max_keepalive_connections=20
# keepalive_expiry=5s, pool_timeout=None(无限等待)
# read_timeout=5s(LLM 经常超过 5s!)
问题:
read_timeout=5s导致大量超时报错(但被业务层重试掩盖了)pool_timeout=None导致高并发时请求无限排队,P99 飙升keepalive_expiry=5s比服务端短太多,大量连接在池里就已经失效
调整后:
python
client = httpx.AsyncClient(
limits=httpx.Limits(
max_connections=200,
max_keepalive_connections=50,
keepalive_expiry=25, # 短于服务端 ~30s
),
timeout=httpx.Timeout(
connect=5.0,
read=90.0, # 覆盖非流式最长响应时间
write=10.0,
pool=8.0, # 超过 8s 等不到连接就报错,不再无限排队
),
)
效果:
| 指标 | 调整前 | 调整后 |
|---|---|---|
| P99 延迟(高峰) | 18s | 2.1s |
| CLOSE_WAIT 连接数 | 600+ | <10 |
| 连接池等待超时报错 | 0(无限等待变成堆积) | 偶发,有告警 |
| 平均延迟 | 1.2s | 1.1s |
P99 从 18s 降到 2.1s,平均延迟基本不变------这是连接池问题的典型特征:平均值正常,尾部极差。
总结
LLM HTTP 连接池和普通服务的连接池有几个关键差异需要特别对待:
- 超时全家桶都要设 :connect、read、write、pool 四个维度,pool 超时最容易被漏掉
- 流式和非流式分开配置:它们对 keepalive_expiry 和 read_timeout 的要求完全相反
- keepalive_expiry 要短于服务端超时:否则复用失效连接,引发大量 ConnectionReset
- 按 QPS×响应时间 计算连接数上限:LLM 响应时间长,需要的连接数远超普通 API
- 流式请求必须显式关闭连接:break、异常都可能造成泄漏,用 finally + aclose() 保护
连接池配置错误不会立刻崩溃,它会以 P99 劣化的形式慢慢侵蚀用户体验,直到某次流量高峰才彻底暴露。提前建立监控、按场景调优,比事后排查要划算得多。
本文配套代码已整合到 [LLMHttpClient 封装](#本文配套代码已整合到 LLMHttpClient 封装,可直接用于生产。 "#%E5%AE%8C%E6%95%B4%E5%B0%81%E8%A3%85"),可直接用于生产。