一、十讲内容全景回顾
┌─────────────────────────────────────────────────────────────────────────┐
│ AI 应用可观测性:从埋点到诊断 │
├─────────────────────────────────────────────────────────────────────────┤
│ │
│ 第1讲 观测维度与指标体系 │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ RED 方法论: Rate / Error / Duration │ │
│ │ USE 方法论: Utilization / Saturation / Errors │ │
│ │ AI 专属指标: Token 消耗 / 模型延迟 / 幻觉率 / 决策置信度 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ ↓ │
│ 第2讲 Span 设计与 Context Propagation │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ W3C Trace Context 标准 │ │
│ │ Span 生命周期: Start → Tag → Log → End │ │
│ │ 跨服务传播: HTTP Header / gRPC Metadata / MQ Header │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ ↓ │
│ 第3讲 LLM 调用监控 │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ Token 计量与计费 │ │
│ │ 模型定价表与成本分摊 │ │
│ │ 告警引擎: 阈值 / 复合 / 静默 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ ↓ │
│ 第4讲 决策质量监控与 PSI 漂移检测 │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ 决策质量: 准确率 / 召回率 / 用户满意度 │ │
│ │ PSI 漂移检测: 特征分布变化监控 │ │
│ │ 自动回滚触发: 质量下降 → 回滚到上一版本 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ ↓ │
│ 第5讲 Prompt 管理与 AB Test │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ Prompt 版本管理与灰度发布 │ │
│ │ ABTestEngine: 分流 / 指标对比 / 显著性检验 │ │
│ │ InjectionDetector: Prompt 注入攻击检测 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ ↓ │
│ 第6讲 日志结构化与存储 │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ Event Schema: 统一日志格式 │ │
│ │ 高性能写入: 无锁 Ring Buffer / 批量压缩 │ │
│ │ 分层存储: ClickHouse(7d) → ES(30d) → S3(90d) │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ ↓ │
│ 第7讲 实时告警与自动化响应 │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ AlertLevel: P0 ~ P3 分级 │ │
│ │ RuleEngine: 阈值 / 复合 / 静默期 │ │
│ │ AutoResponder: 熔断 / 扩容 / 回滚 / 限流 │ │
│ │ StormSuppressor: 风暴抑制 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ ↓ │
│ 第8讲 分布式链路追踪与根因分析 │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ Trace / Span / SpanContext 数据结构 │ │
│ │ 采样策略: HeadBased / TailBased / Priority │ │
│ │ 根因分析: 首个错误 Span / 瓶颈检测 / 异常检测 │ │
│ │ 服务拓扑图: 依赖关系可视化 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ ↓ │
│ 第9讲 混沌工程与容灾演练 │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ 故障注入: 网络 / 服务 / 资源 / LLM │ │
│ │ 实验引擎: 稳态检查 → 注入 → 观察 → 回滚 │ │
│ │ 预置场景: LLM 降级 / 服务崩溃 / 资源耗尽 / 依赖故障 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ ↓ │
│ 第10讲 全景总结与生产实战(本讲) │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ 生产环境完整部署方案 │ │
│ │ Dashboard 设计: 从宏观到微观 │ │
│ │ SLA / SLO / SLI 体系 │ │
│ │ 运维 SOP: On-Call 手册 │ │
│ │ 未来演进方向 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────────┘
二、生产环境完整部署方案
2.1 整体架构
┌─────────────────────────────────────────────────────────────────────┐
│ 用户请求 │
└──────────────────────────┬──────────────────────────────────────────┘
│
┌──────────────────────────▼──────────────────────────────────────────┐
│ API Gateway (Envoy / Kong) │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Rate Limiter │ │ Auth Filter │ │ Trace ID Gen │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└──────────┬───────────────────────────────────────────────────────────┘
│
┌──────────▼───────────────────────────────────────────────────────────┐
│ Service Mesh (Istio) │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Sidecar Proxy: 自动注入 Trace / 采集 Metrics / 流量管控 │ │
│ └──────────────────────────────────────────────────────────────┘ │
└──────────┬───────────────────────────────────────────────────────────┘
│
┌──────────▼───────────────────────────────────────────────────────────┐
│ AI 应用服务集群 │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ NLP │ │ LLM │ │ Vector │ │ Reranker │ │
│ │ Service │ │ Proxy │ │ DB │ │ │ │
│ └──────────┘ └──────────┘ └──────────┘ └──────────┘ │
│ │
│ 每个 Pod 内置: │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ ① OpenTelemetry SDK (Go / Python) │ │
│ │ ② Prometheus Metrics Exporter (端口 :9090) │ │
│ │ ③ Structured Logger (JSON 格式 → stdout) │ │
│ │ ④ Health Check Handler (/healthz / /readyz) │ │
│ └──────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
│
┌──────────────────────────▼──────────────────────────────────────────┐
│ 可观测性基础设施 │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ OpenTelemetry│ │ Kafka │ │ Fluentd │ │
│ │ Collector │ │ (缓冲层) │ │ (日志采集) │ │
│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │ │
│ ┌──────▼─────────────────▼─────────────────▼───────┐ │
│ │ 数据处理层 │ │
│ │ ┌──────────┐ ┌──────────┐ ┌──────────────┐ │ │
│ │ │ Metrics │ │ Traces │ │ Logs │ │ │
│ │ │ Processor│ │ Processor│ │ Processor │ │ │
│ │ └────┬─────┘ └────┬─────┘ └──────┬───────┘ │ │
│ └───────┼─────────────┼───────────────┼───────────┘ │
│ │ │ │ │
│ ┌───────▼─────────────▼───────────────▼───────────┐ │
│ │ 存储层 │ │
│ │ ┌──────────┐ ┌──────────┐ ┌──────────────┐ │ │
│ │ │Victoria │ │Jaeger │ │ Loki / ES │ │ │
│ │ │Metrics │ │(Traces) │ │ (Logs) │ │ │
│ │ └──────────┘ └──────────┘ └──────────────┘ │ │
│ └──────────────────────────────────────────────────┘ │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Grafana │ │ AlertManager│ │ Incident │ │
│ │ (Dashboards)│ │ (告警管理) │ │ Management │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
2.2 Kubernetes 部署清单
# observability-stack.yaml
---
apiVersion: v1
kind: Namespace
metadata:
name: observability
---
# OpenTelemetry Collector
apiVersion: apps/v1
kind: Deployment
metadata:
name: otel-collector
namespace: observability
spec:
replicas: 3
selector:
matchLabels:
app: otel-collector
template:
metadata:
labels:
app: otel-collector
spec:
containers:
- name: otel-collector
image: otel/opentelemetry-collector-contrib:0.120.0
args:
- "--config=/etc/otel/config.yaml"
ports:
- containerPort: 4317 # gRPC
- containerPort: 4318 # HTTP
volumeMounts:
- name: config
mountPath: /etc/otel
resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: 2
memory: 2Gi
volumes:
- name: config
configMap:
name: otel-config
---
apiVersion: v1
kind: ConfigMap
metadata:
name: otel-config
namespace: observability
data:
config.yaml: |
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 5s
send_batch_size: 8192
memory_limiter:
check_interval: 1s
limit_mib: 1536
spike_limit_mib: 256
attributes:
actions:
- key: environment
value: production
action: upsert
filter:
error_mode: ignore
traces:
span:
- 'attributes["http.target"] == "/healthz"'
exporters:
prometheus:
endpoint: 0.0.0.0:8889
namespace: ai_app
otlp:
endpoint: jaeger:4317
tls:
insecure: true
loki:
endpoint: http://loki:3100/loki/api/v1/push
tenant_id: ai-app
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlp]
metrics:
receivers: [otlp]
processors: [memory_limiter, filter, batch]
exporters: [prometheus]
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [loki]
---
# VictoriaMetrics
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: victoria-metrics
namespace: observability
spec:
replicas: 2
selector:
matchLabels:
app: victoria-metrics
serviceName: victoria-metrics
template:
metadata:
labels:
app: victoria-metrics
spec:
containers:
- name: victoria-metrics
image: victoriametrics/victoria-metrics:v1.108.0
args:
- "-storageDataPath=/data"
- "-retentionPeriod=30d"
- "-search.maxUniqueTimeseries=1000000"
ports:
- containerPort: 8428
volumeMounts:
- name: data
mountPath: /data
resources:
requests:
cpu: 1
memory: 2Gi
limits:
cpu: 4
memory: 8Gi
volumeClaimTemplates:
- metadata:
name: data
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: ssd
resources:
requests:
storage: 500Gi
---
# Jaeger
apiVersion: apps/v1
kind: Deployment
metadata:
name: jaeger
namespace: observability
spec:
replicas: 2
selector:
matchLabels:
app: jaeger
template:
metadata:
labels:
app: jaeger
spec:
containers:
- name: jaeger
image: jaegertracing/all-in-one:1.63.0
env:
- name: COLLECTOR_OTLP_ENABLED
value: "true"
- name: SPAN_STORAGE_TYPE
value: "elasticsearch"
- name: ES_SERVER_URLS
value: "http://elasticsearch:9200"
ports:
- containerPort: 16686 # UI
- containerPort: 4317 # OTLP gRPC
resources:
requests:
cpu: 500m
memory: 1Gi
limits:
cpu: 2
memory: 4Gi
---
# Grafana
apiVersion: apps/v1
kind: Deployment
metadata:
name: grafana
namespace: observability
spec:
replicas: 2
selector:
matchLabels:
app: grafana
template:
metadata:
labels:
app: grafana
spec:
containers:
- name: grafana
image: grafana/grafana:11.4.0
env:
- name: GF_SECURITY_ADMIN_PASSWORD
valueFrom:
secretKeyRef:
name: grafana-secret
key: admin-password
- name: GF_INSTALL_PLUGINS
value: "grafana-piechart-panel,grafana-worldmap-panel"
ports:
- containerPort: 3000
volumeMounts:
- name: dashboards
mountPath: /etc/grafana/provisioning/dashboards
- name: datasources
mountPath: /etc/grafana/provisioning/datasources
resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: 1
memory: 1Gi
volumes:
- name: dashboards
configMap:
name: grafana-dashboards
- name: datasources
configMap:
name: grafana-datasources
---
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-datasources
namespace: observability
data:
datasources.yaml: |
apiVersion: 1
datasources:
- name: VictoriaMetrics
type: prometheus
url: http://victoria-metrics:8428
access: proxy
isDefault: true
- name: Jaeger
type: jaeger
url: http://jaeger:16686
access: proxy
- name: Loki
type: loki
url: http://loki:3100
access: proxy
---
# AI 应用服务示例
apiVersion: apps/v1
kind: Deployment
metadata:
name: llm-proxy
namespace: ai-app
spec:
replicas: 5
selector:
matchLabels:
app: llm-proxy
template:
metadata:
labels:
app: llm-proxy
annotations:
sidecar.istio.io/inject: "true"
prometheus.io/scrape: "true"
prometheus.io/port: "9090"
spec:
containers:
- name: llm-proxy
image: registry.ai-app/llm-proxy:v2.3.1
ports:
- containerPort: 8080
name: http
- containerPort: 9090
name: metrics
env:
- name: OTEL_EXPORTER_OTLP_ENDPOINT
value: "http://otel-collector.observability:4317"
- name: OTEL_SERVICE_NAME
value: "llm-proxy"
- name: OTEL_TRACES_SAMPLER
value: "parentbased_traceidratio"
- name: OTEL_TRACES_SAMPLER_ARG
value: "0.1"
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 10
periodSeconds: 15
readinessProbe:
httpGet:
path: /readyz
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: 2
memory: 2Gi
volumeMounts:
- name: config
mountPath: /etc/app
volumes:
- name: config
configMap:
name: llm-proxy-config
---
# HPA 自动扩缩容
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: llm-proxy-hpa
namespace: ai-app
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: llm-proxy
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
- type: Pods
pods:
metric:
name: llm_proxy_queue_depth
target:
type: AverageValue
averageValue: 100
behavior:
scaleUp:
stabilizationWindowSeconds: 60
policies:
- type: Percent
value: 100
periodSeconds: 15
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 10
periodSeconds: 60
三、Dashboard 设计
3.1 宏观大盘(CEO 视角)
┌─────────────────────────────────────────────────────────────────────────┐
│ AI 应用可观测性 · 宏观大盘 [Last 24h] [Auto-refresh] │
├─────────────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐│
│ │ 总请求量 │ │ 成功率 │ │ P99 延迟 │ │ 日活用户 ││
│ │ 1,234,567 │ │ 99.87% │ │ 892ms │ │ 89,123 ││
│ │ ▲ 12% vs 昨日 │ │ ▼ 0.05% │ │ ▲ 45ms │ │ ▲ 5% ││
│ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘│
│ │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ 服务健康状态 │ │
│ │ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ │ │
│ │ │API │ │NLP │ │LLM │ │VecDB│ │Rer │ │Cache│ │ │
│ │ │GW │ │Svc │ │Proxy│ │ │ │anker│ │ │ │ │
│ │ │ 🟢 │ │ 🟢 │ │ 🟡 │ │ 🟢 │ │ 🟢 │ │ 🟢 │ │ │
│ │ │200ms│ │50ms │ │1.2s │ │35ms │ │15ms │ │2ms │ │ │
│ │ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘ │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ 过去 24 小时 P99 延迟趋势 │ │
│ │ ▁▂▃▄▅▆▇█▇▆▅▄▃▂▁▂▃▄▅▆▇█▇▆▅▄▃▂▁▂▃▄▅▆▇█▇▆▅▄▃▂▁▂▃▄▅▆▇█▇▆▅▄▃▂▁ │ │
│ │ 00:00 04:00 08:00 12:00 16:00 20:00 现在 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ │
│ ┌─────────────────────────────┐ ┌─────────────────────────────────┐ │
│ │ 错误分布 Top 5 │ │ Token 消耗趋势 │ │
│ │ ┌─────────────────────┐ │ │ ▁▂▃▄▅▆▇█▇▆▅▄▃▂▁ │ │
│ │ │ LLM Timeout 45% │ │ │ 今日: 12.3M tokens │ │
│ │ │ Rate Limit 22% │ │ │ 费用: $246.78 │ │
│ │ │ Vector DB 15% │ │ └─────────────────────────────────┘ │
│ │ │ Auth Fail 10% │ │ │
│ │ │ Other 8% │ │ │
│ │ └─────────────────────┘ │ │
│ └─────────────────────────────┘ └─────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────────┘
3.2 服务详细面板(SRE 视角)
{
"title": "LLM Proxy 详细面板",
"panels": [
{
"title": "QPS & 延迟",
"type": "timeseries",
"queries": [
"rate(llm_proxy_requests_total[1m])",
"histogram_quantile(0.99, rate(llm_proxy_request_duration_seconds_bucket[5m]))",
"histogram_quantile(0.95, rate(llm_proxy_request_duration_seconds_bucket[5m]))",
"histogram_quantile(0.50, rate(llm_proxy_request_duration_seconds_bucket[5m]))"
]
},
{
"title": "Token 消耗明细",
"type": "stat",
"queries": [
"sum(rate(llm_proxy_prompt_tokens_total[5m]))",
"sum(rate(llm_proxy_completion_tokens_total[5m]))",
"sum(rate(llm_proxy_cost_usd_total[5m]))"
]
},
{
"title": "模型调用分布",
"type": "piechart",
"queries": [
"count by(model) (llm_proxy_requests_total)"
]
},
{
"title": "错误原因分布",
"type": "barchart",
"queries": [
"count by(error_type) (llm_proxy_errors_total)"
]
},
{
"title": "熔断器状态",
"type": "stat",
"queries": [
"llm_proxy_circuit_breaker_state{state=\"open\"}",
"llm_proxy_circuit_breaker_state{state=\"half-open\"}",
"llm_proxy_circuit_breaker_state{state=\"closed\"}"
]
},
{
"title": "最近 Trace 列表",
"type": "table",
"datasource": "Jaeger",
"queries": [
"service=llm-proxy & limit=20 & lookback=1h"
],
"columns": ["TraceID", "Duration", "Spans", "Services", "Errors"]
}
]
}
3.3 业务指标面板(产品经理视角)
{
"title": "AI 应用业务指标",
"panels": [
{
"title": "日活跃用户 (DAU)",
"type": "timeseries",
"query": "sum(increase(ai_app_active_users_total[24h]))"
},
{
"title": "用户满意度评分",
"type": "gauge",
"query": "avg(ai_app_user_satisfaction_score)",
"thresholds": {
"green": 4.0,
"yellow": 3.0,
"red": 2.0
}
},
{
"title": "问答采纳率",
"type": "timeseries",
"query": "sum(rate(ai_app_answer_accepted_total[1h])) / sum(rate(ai_app_answers_total[1h]))"
},
{
"title": "模型版本分布",
"type": "piechart",
"query": "count by(model_version) (ai_app_model_deployments)"
},
{
"title": "AB 实验效果对比",
"type": "stat",
"queries": [
"avg(ai_app_conversion_rate{experiment_group='control'})",
"avg(ai_app_conversion_rate{experiment_group='treatment'})"
]
},
{
"title": "Token 成本按部门",
"type": "barchart",
"query": "sum by(department) (ai_app_token_cost_total)"
}
]
}
四、SLA / SLO / SLI 体系
4.1 定义
SLA (Service Level Agreement) ------ 对外承诺
└── SLO (Service Level Objective) ------ 内部目标
└── SLI (Service Level Indicator) ------ 可测量指标
4.2 AI 应用 SLI 指标
slis:
# 可用性
availability:
definition: "成功响应的请求占比"
measurement: "successful_requests / total_requests * 100"
exclusion: "排除计划内维护窗口"
# 延迟
latency_p99:
definition: "最慢 1% 请求的响应时间"
measurement: "histogram_quantile(0.99, request_duration_seconds)"
exclusion: "排除超时已熔断的请求"
# 吞吐量
throughput:
definition: "每秒处理的请求数"
measurement: "rate(requests_total[1m])"
# 新鲜度
freshness:
definition: "数据从产生到可查询的时间"
measurement: "max(event_timestamp - ingestion_timestamp)"
# 正确性
correctness:
definition: "AI 回答被用户采纳的比例"
measurement: "accepted_answers / total_answers * 100"
# 成本效率
cost_efficiency:
definition: "每美元 token 产出的有效回答数"
measurement: "accepted_answers / total_cost_usd"
4.3 SLO 目标设定
slo_targets:
# Tier 1: 核心对话服务
tier_1:
description: "用户直接感知的核心链路"
targets:
availability: 99.99% # 全年宕机 < 52min
latency_p99: 2s # 99% 请求在 2s 内
correctness: 85% # 回答采纳率 > 85%
burn_rate:
warning: 10% / 7d # 7 天内消耗 10% 预算
critical: 20% / 1d # 1 天内消耗 20% 预算
consequences:
- "触发 P0 告警"
- "立即回滚最近变更"
- "全员 On-Call"
# Tier 2: 辅助功能
tier_2:
description: "增强体验但不阻塞核心功能"
targets:
availability: 99.9%
latency_p99: 5s
freshness: 30s
burn_rate:
warning: 20% / 7d
critical: 50% / 1d
consequences:
- "触发 P1 告警"
- "下个工作日修复"
# Tier 3: 后台批处理
tier_3:
description: "离线分析和模型训练"
targets:
availability: 99.0%
throughput: "每日处理 > 1M 条"
freshness: 1h
consequences:
- "触发 P2 告警"
- "本周内修复"
4.4 错误预算管理
error_budget:
# 以 Tier 1 为例:99.99% 可用性 = 每年 52min 错误预算
calculation:
annual_uptime: 99.99%
annual_downtime_budget: "52 minutes"
monthly_budget: "4.3 minutes"
weekly_budget: "1 minute"
consumption_tracking:
- date: "2026-09-20"
consumed: 12s
remaining_weekly: 48s
status: "🟢 正常"
- date: "2026-09-21"
consumed: 45s
remaining_weekly: 15s
status: "🟡 警告"
- date: "2026-09-22"
consumed: 72s
remaining_weekly: -12s
status: "🔴 超额"
actions_on_exhaustion:
yellow:
- "暂停所有非紧急变更"
- "启动根因分析"
red:
- "冻结所有变更"
- "全员介入修复"
- "通知 VP 级别"
五、运维 SOP:On-Call 手册
5.1 On-Call 流程
┌─────────────────────────────────────────────────────────────────┐
│ On-Call 响应流程 │
├─────────────────────────────────────────────────────────────────┤
│ │
│ 收到告警 │
│ │ │
│ ▼ │
│ ┌──────────────────────┐ │
│ │ 1. ACKNOWLEDGE │ ← 5分钟内必须确认 │
│ │ 确认告警 │ │
│ └──────────┬───────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────┐ │
│ │ 2. TRIAGE │ ← 判断级别 │
│ │ 分类定级 │ │
│ └──────────┬───────────┘ │
│ │ │
│ ┌──────┴──────┐ │
│ ▼ ▼ │
│ ┌────────┐ ┌────────────┐ │
│ │ P0/P1 │ │ P2/P3 │ │
│ │ 立即响应│ │ 工作时间处理│ │
│ └───┬────┘ └────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────┐ │
│ │ 3. DIAGNOSE │ ← 查看 Dashboard / Trace / Log │
│ │ 诊断根因 │ │
│ └──────────┬───────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────┐ │
│ │ 4. MITIGATE │ ← 回滚 / 重启 / 扩容 / 降级 │
│ │ 止血恢复 │ │
│ └──────────┬───────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────┐ │
│ │ 5. RESOLVE │ ← 确认指标恢复正常 │
│ │ 确认解决 │ │
│ └──────────┬───────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────┐ │
│ │ 6. POSTMORTEM │ ← 24h 内提交事故事后分析 │
│ │ 事后复盘 │ │
│ └──────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
5.2 常见故障处理手册
incident_handbook:
# 场景 1: LLM 响应超时飙升
scenario_llm_timeout:
symptoms:
- "P99 延迟 > 5s"
- "llm_proxy_timeout_errors 飙升"
- "用户反馈回复很慢"
diagnosis:
step_1: "检查 LLM Provider 状态页"
step_2: "查看 Jaeger 中 llm-proxy 的 Trace"
step_3: "检查 Token 消耗是否有异常突增"
step_4: "检查模型是否被限流"
mitigation:
option_a: "切换到备用模型 (gpt-3.5-turbo → claude-haiku)"
option_b: "降低 max_tokens 限制"
option_c: "启用本地缓存减少重复调用"
option_d: "扩容 llm-proxy 实例"
escalation:
if_not_resolved_in: "15分钟"
escalate_to: "AI Platform Team Lead"
# 场景 2: Vector DB 不可用
scenario_vector_db_down:
symptoms:
- "vector_db_errors 100%"
- "RAG 功能完全不可用"
- "告警: VectorDBConnectionFailed"
diagnosis:
step_1: "检查 Vector DB Pod 状态"
step_2: "查看 PVC 磁盘使用率"
step_3: "检查网络策略是否误拦截"
step_4: "查看 Vector DB 日志"
mitigation:
option_a: "重启 Vector DB Pod"
option_b: "扩容副本数"
option_c: "降级为关键词搜索 (BM25)"
option_d: "切换只读副本"
escalation:
if_not_resolved_in: "10分钟"
escalate_to: "DBA Team"
# 场景 3: 模型漂移导致回答质量下降
scenario_model_drift:
symptoms:
- "用户满意度评分下降 > 10%"
- "PSI 指标超出阈值"
- "负面反馈增多"
diagnosis:
step_1: "查看 PSI Dashboard 确认漂移维度"
step_2: "对比新旧模型版本的回答样本"
step_3: "检查 Prompt 是否被意外修改"
step_4: "检查上游数据分布是否变化"
mitigation:
option_a: "回滚到上一个稳定模型版本"
option_b: "切换 AB 实验组到对照组"
option_c: "临时增加后处理校验规则"
escalation:
if_not_resolved_in: "30分钟"
escalate_to: "ML Team Lead"
5.3 事后复盘模板
# 事故事后复盘报告
## 基本信息
- **事故编号**: INC-2026-0928-001
- **标题**: LLM 代理服务 P99 延迟飙升至 8s
- **严重级别**: P0
- **日期**: 2026-09-28
- **持续时间**: 23 分钟
- **影响范围**: 所有依赖 LLM 的用户请求
## 时间线
| 时间 | 事件 |
|------|------|
| 14:23 | 告警触发: P99 延迟 > 5s |
| 14:24 | On-Call 工程师确认告警 |
| 14:26 | 诊断: OpenAI API 响应变慢 |
| 14:28 | 执行降级: 切换到备用模型 |
| 14:35 | 指标恢复正常 |
| 14:46 | OpenAI 发布状态更新确认故障 |
| 15:00 | 切换回主模型 |
## 根因分析
- **直接原因**: OpenAI API 出现区域性延迟
- **根本原因**: 未配置多区域故障转移
- **促成因素**: 备用模型预热不足,首次切换延迟较高
## 改进措施
| 项目 | 负责人 | 截止日期 |
|------|--------|----------|
| 配置多区域 LLM Provider 故障转移 | @infra-team | 2026-10-05 |
| 备用模型保持 Warm Pool | @ml-team | 2026-10-03 |
| 添加 Provider 健康探测 | @sre-team | 2026-10-01 |
| 更新 On-Call 手册 LLM 降级章节 | @docs-team | 2026-09-30 |
## 附件
- [Grafana Dashboard 截图]
- [Jaeger Trace 链接]
- [告警记录]
六、未来演进方向
6.1 AI for Observability (AIOps)
aiops_roadmap:
phase_1: 异常检测自动化
capabilities:
- "基于历史数据的动态阈值"
- "多维度的异常关联分析"
- "季节性模式自动识别"
example: "自动区分工作日/周末流量差异,避免误告警"
phase_2: 根因分析智能化
capabilities:
- "因果推断: 从相关性到因果性"
- "知识图谱: 服务依赖关系推理"
- "自然语言根因描述"
example: "LLM 分析 Trace 后直接输出: 'llm-proxy 的 call_llm span 耗时 8.2s,占整条 Trace 的 98%,根因为 OpenAI API 区域性延迟'"
phase_3: 自愈系统
capabilities:
- "故障自动诊断 + 自动修复"
- "渐进式回滚: 1% → 5% → 20% → 100%"
- "混沌工程与自愈形成闭环"
example: "检测到模型漂移 → 自动触发 AB 实验回滚 → 验证指标恢复 → 通知团队"
phase_4: 预测性可观测性
capabilities:
- "容量预测: 提前 24h 预测资源需求"
- "故障预测: 基于模式识别的故障预警"
- "成本预测: Token 消耗和费用预估"
example: "预测今晚 20:00 会有流量高峰,提前扩容 3 个副本"
6.2 技术演进趋势
technology_trends:
# eBPF 深度观测
ebpf:
description: "无需修改代码即可观测内核态行为"
use_cases:
- "网络延迟的精确测量"
- "文件 I/O 性能分析"
- "系统调用追踪"
tools:
- "Pixie"
- "Cilium Tetragon"
- "Parca (连续分析)"
# OpenTelemetry 标准化
opentelemetry:
description: "可观测性的行业标准,统一 Metrics/Traces/Logs"
trends:
- "Profiling Signal 加入 OTel"
- "eBPF 与 OTel 融合"
- "OTel Operator 简化部署"
# 低成本存储
cost_effective_storage:
description: "海量观测数据的存储成本优化"
strategies:
- "列式存储 + 高压缩比"
- "对象存储作为冷备"
- "采样 + 聚合降低数据量"
tools:
- "Grafana Mimir"
- "Thanos"
- "ClickHouse"
# 持续验证
continuous_validation:
description: "可观测性本身也需要被观测"
practices:
- "SLI 覆盖率检查"
- "告警质量评分"
- "Dashboard 使用率统计"
七、总结:从埋点到诊断的完整链路
┌─────────────────────────────────────────────────────────────────────────┐
│ AI 应用可观测性:从埋点到诊断 │
│ │
│ 埋点层 (Instrumentation) │
│ ├── 代码埋点: OpenTelemetry SDK │
│ ├── 自动埋点: Istio Sidecar / eBPF │
│ └── 业务埋点: 自定义 Metrics / Events │
│ │ │
│ ▼ │
│ 采集层 (Collection) │
│ ├── Metrics: Prometheus Exporter (端口 :9090) │
│ ├── Traces: OTLP Exporter → OpenTelemetry Collector │
│ └── Logs
│ └── Logs: Structured JSON → stdout → Fluentd
│ │
│ ▼
│ 传输层 (Transport)
│ ├── 缓冲: Kafka (削峰填谷)
│ ├── 采样: Tail-Based Sampling (错误全采 / 慢请求半采 / 正常 1%)
│ └── 压缩: gzip / snappy (减少带宽)
│ │
│ ▼
│ 存储层 (Storage)
│ ├── 热存储 (7天): ClickHouse / VictoriaMetrics
│ ├── 温存储 (30天): Elasticsearch / Jaeger
│ └── 冷存储 (90天): S3 / GCS (Parquet 格式)
│ │
│ ▼
│ 分析层 (Analysis)
│ ├── 指标分析: PromQL 查询 / 聚合 / 预测
│ ├── 链路分析: Trace 查询 / 服务拓扑 / 根因定位
│ └── 日志分析: 全文检索 / 模式识别 / 聚类
│ │
│ ▼
│ 展示层 (Visualization)
│ ├── Grafana Dashboards: 宏观大盘 / 服务面板 / 业务面板
│ ├── Jaeger UI: Trace 详情 / 火焰图 / 比较视图
│ └── 自定义 Portal: 业务指标 / 成本报表 / SLA 看板
│ │
│ ▼
│ 行动层 (Action)
│ ├── 告警: AlertManager → 电话 / Slack / 邮件
│ ├── 自动化: AutoResponder → 熔断 / 扩容 / 回滚 / 限流
│ └── 混沌: Chaos Engine → 故障注入 / 韧性验证 / 改进闭环
│
└─────────────────────────────────────────────────────────────────────────┘
八、附录:常用命令速查
8.1 排查问题三板斧
# 1. 看指标 ------ 哪里慢了?
# P99 延迟最高的服务 Top 5
topk(5, histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (service, le)))
# 2. 看链路 ------ 为什么慢?
# 查询最近 1 小时最慢的 10 条 Trace
# 在 Jaeger UI 中执行:
# Service: llm-proxy
# Operation: call_llm
# Min Duration: 1s
# Lookback: 1h
# Limit: 10
# 3. 看日志 ------ 报了什么错?
# 查询特定 TraceID 的所有日志
{app="llm-proxy"} |= "trace_id=a1b2c3d4e5f6"
8.2 常用 PromQL 查询
# 服务可用性 (排除 503 熔断)
sum(rate(http_requests_total{status!~"5.."}[5m]))
/
sum(rate(http_requests_total[5m])) * 100
# P99 延迟趋势
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le))
# 错误率 (5xx / total)
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m])) * 100
# Token 消耗速率
sum(rate(llm_token_total[5m])) by (model)
# 每秒成本
sum(rate(llm_cost_usd_total[5m]))
# 熔断器打开次数
increase(circuit_breaker_open_total[1h])
# 最慢的 5 个端点
topk(5, avg by(endpoint) (http_request_duration_seconds_sum / http_request_duration_seconds_count))
# 各服务 Span 数量
count by(service.name) (spans{})
8.3 常用 kubectl 命令
# 查看所有 Pod 状态
kubectl get pods -n ai-app -o wide
# 查看 Pod 日志
kubectl logs -n ai-app deployment/llm-proxy --tail=100 -f
# 查看 Pod 资源使用
kubectl top pod -n ai-app
# 查看 HPA 状态
kubectl get hpa -n ai-app
# 查看事件
kubectl get events -n ai-app --sort-by='.lastTimestamp'
# 端口转发 (本地访问 Jaeger)
kubectl port-forward -n observability svc/jaeger 16686:16686
# 进入 Pod 调试
kubectl exec -it -n ai-app pod/llm-proxy-xxx -- sh
# 查看 ConfigMap
kubectl get configmap -n ai-app llm-proxy-config -o yaml
# 滚动重启
kubectl rollout restart -n ai-app deployment/llm-proxy
# 查看滚动状态
kubectl rollout status -n ai-app deployment/llm-proxy
九、推荐阅读与工具
9.1 必读书籍
| 书名 | 作者 | 推荐理由 |
|---|---|---|
| 《Site Reliability Engineering》 | Google SRE Team | SRE 圣经,定义了现代可观测性理念 |
| 《The Art of Monitoring》 | James Turnbull | 从零搭建监控体系的实操指南 |
| 《Distributed Tracing in Practice》 | Austin Parker 等 | 链路追踪的权威实践手册 |
| 《Chaos Engineering》 | Casey Rosenthal 等 | 混沌工程的奠基之作 |
| 《Observability Engineering》 | Charity Majors 等 | 可观测性三大支柱的系统讲解 |
9.2 开源工具推荐
| 类别 | 工具 | 说明 |
|---|---|---|
| Metrics | VictoriaMetrics | 高性能时序数据库,兼容 PromQL |
| Traces | Jaeger / Tempo | 分布式链路追踪 |
| Logs | Loki / Quickwit | 低成本日志存储 |
| Profiling | Parca / Pyroscope | 持续性能分析 |
| eBPF | Pixie / Cilium | 零侵入内核观测 |
| 混沌 | Chaos Mesh / Litmus | K8s 原生混沌工程 |
| 告警 | AlertManager / Keep | 告警管理和抑制 |
| Dashboard | Grafana / Perses | 可视化面板 |
9.3 在线工具
| 工具 | 用途 | 地址 |
|---|---|---|
| 数字转大写 | 金额转换 | zz365.top/daxie |
| JSON 格式化 | Trace 数据美化 | zz365.top/json |
| 时间戳转换 | Unix 时间 ↔ 可读时间 | zz365.top/timestamp |
| Cron 表达式 | 告警周期配置 | zz365.top/cron |
| Base64 编解码 | Token 解码调试 | zz365.top/base64 |
十、结语
至此,《AI 应用可观测性:从埋点到诊断》十讲全部结束。我们从最基础的观测维度出发,一步步构建了完整的可观测性体系:
第1讲 👉 知道要看什么(指标体系)
第2讲 👉 知道怎么串起来(Trace)
第3-5讲 👉 知道 AI 特有的坑(LLM / 漂移 / Prompt)
第6讲 👉 知道怎么存(日志结构)
第7讲 👉 知道怎么反应(告警响应)
第8讲 👉 知道怎么查根因(链路分析)
第9讲 👉 知道怎么验证韧性(混沌工程)
第10讲 👉 知道怎么落地(生产实战)
记住三条铁律:
- 没有度量就没有改进 ------ 先埋点,再优化
- 没有 Trace 的日志就是孤岛 ------ 永远带着 TraceID
- 没有自动化的告警就是噪音 ------ 告警必须能触发行动
祝你的 AI 应用永远稳定、高效、可观测!
🧰 开发之余的小工具推荐
设计混沌实验场景时,经常需要计算各种时间窗口和延迟参数。zz365.top 的在线计算器可以快速进行毫秒/秒/分钟的单位换算和百分比计算,帮助你在配置故障参数时更精确。所有计算纯前端完成,不需要联网。