第10讲:AI 应用可观测性全景总结与生产实战

一、十讲内容全景回顾

复制代码
┌─────────────────────────────────────────────────────────────────────────┐
│                     AI 应用可观测性:从埋点到诊断                          │
├─────────────────────────────────────────────────────────────────────────┤
│                                                                         │
│  第1讲 观测维度与指标体系                                               │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  RED 方法论: Rate / Error / Duration                           │   │
│  │  USE 方法论: Utilization / Saturation / Errors                 │   │
│  │  AI 专属指标: Token 消耗 / 模型延迟 / 幻觉率 / 决策置信度       │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                              ↓                                         │
│  第2讲 Span 设计与 Context Propagation                                 │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  W3C Trace Context 标准                                        │   │
│  │  Span 生命周期: Start → Tag → Log → End                        │   │
│  │  跨服务传播: HTTP Header / gRPC Metadata / MQ Header           │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                              ↓                                         │
│  第3讲 LLM 调用监控                                                   │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  Token 计量与计费                                                │   │
│  │  模型定价表与成本分摊                                             │   │
│  │  告警引擎: 阈值 / 复合 / 静默                                     │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                              ↓                                         │
│  第4讲 决策质量监控与 PSI 漂移检测                                      │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  决策质量: 准确率 / 召回率 / 用户满意度                           │   │
│  │  PSI 漂移检测: 特征分布变化监控                                  │   │
│  │  自动回滚触发: 质量下降 → 回滚到上一版本                          │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                              ↓                                         │
│  第5讲 Prompt 管理与 AB Test                                          │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  Prompt 版本管理与灰度发布                                       │   │
│  │  ABTestEngine: 分流 / 指标对比 / 显著性检验                      │   │
│  │  InjectionDetector: Prompt 注入攻击检测                          │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                              ↓                                         │
│  第6讲 日志结构化与存储                                                 │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  Event Schema: 统一日志格式                                     │   │
│  │  高性能写入: 无锁 Ring Buffer / 批量压缩                         │   │
│  │  分层存储: ClickHouse(7d) → ES(30d) → S3(90d)                  │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                              ↓                                         │
│  第7讲 实时告警与自动化响应                                              │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  AlertLevel: P0 ~ P3 分级                                      │   │
│  │  RuleEngine: 阈值 / 复合 / 静默期                               │   │
│  │  AutoResponder: 熔断 / 扩容 / 回滚 / 限流                       │   │
│  │  StormSuppressor: 风暴抑制                                      │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                              ↓                                         │
│  第8讲 分布式链路追踪与根因分析                                           │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  Trace / Span / SpanContext 数据结构                            │   │
│  │  采样策略: HeadBased / TailBased / Priority                    │   │
│  │  根因分析: 首个错误 Span / 瓶颈检测 / 异常检测                   │   │
│  │  服务拓扑图: 依赖关系可视化                                      │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                              ↓                                         │
│  第9讲 混沌工程与容灾演练                                                │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  故障注入: 网络 / 服务 / 资源 / LLM                             │   │
│  │  实验引擎: 稳态检查 → 注入 → 观察 → 回滚                        │   │
│  │  预置场景: LLM 降级 / 服务崩溃 / 资源耗尽 / 依赖故障             │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                              ↓                                         │
│  第10讲 全景总结与生产实战(本讲)                                      │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │  生产环境完整部署方案                                            │   │
│  │  Dashboard 设计: 从宏观到微观                                   │   │
│  │  SLA / SLO / SLI 体系                                          │   │
│  │  运维 SOP: On-Call 手册                                        │   │
│  │  未来演进方向                                                    │   │
│  └─────────────────────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────────────────────┘

二、生产环境完整部署方案

2.1 整体架构

复制代码
┌─────────────────────────────────────────────────────────────────────┐
│                        用户请求                                       │
└──────────────────────────┬──────────────────────────────────────────┘
                           │
┌──────────────────────────▼──────────────────────────────────────────┐
│                     API Gateway (Envoy / Kong)                       │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐               │
│  │ Rate Limiter │  │ Auth Filter  │  │ Trace ID Gen │               │
│  └──────────────┘  └──────────────┘  └──────────────┘               │
└──────────┬───────────────────────────────────────────────────────────┘
           │
┌──────────▼───────────────────────────────────────────────────────────┐
│                        Service Mesh (Istio)                          │
│  ┌──────────────────────────────────────────────────────────────┐   │
│  │ Sidecar Proxy: 自动注入 Trace / 采集 Metrics / 流量管控       │   │
│  └──────────────────────────────────────────────────────────────┘   │
└──────────┬───────────────────────────────────────────────────────────┘
           │
┌──────────▼───────────────────────────────────────────────────────────┐
│                    AI 应用服务集群                                    │
│                                                                     │
│  ┌──────────┐  ┌──────────┐  ┌──────────┐  ┌──────────┐            │
│  │ NLP      │  │ LLM      │  │ Vector   │  │ Reranker │            │
│  │ Service  │  │ Proxy    │  │ DB       │  │          │            │
│  └──────────┘  └──────────┘  └──────────┘  └──────────┘            │
│                                                                     │
│  每个 Pod 内置:                                                      │
│  ┌──────────────────────────────────────────────────────────────┐   │
│  │ ① OpenTelemetry SDK (Go / Python)                           │   │
│  │ ② Prometheus Metrics Exporter (端口 :9090)                  │   │
│  │ ③ Structured Logger (JSON 格式 → stdout)                    │   │
│  │ ④ Health Check Handler (/healthz / /readyz)                 │   │
│  └──────────────────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────────────────┘
                           │
┌──────────────────────────▼──────────────────────────────────────────┐
│                    可观测性基础设施                                   │
│                                                                     │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐               │
│  │ OpenTelemetry│  │   Kafka      │  │  Fluentd     │               │
│  │  Collector   │  │ (缓冲层)     │  │  (日志采集)   │               │
│  └──────┬───────┘  └──────┬───────┘  └──────┬───────┘               │
│         │                 │                 │                        │
│  ┌──────▼─────────────────▼─────────────────▼───────┐               │
│  │              数据处理层                            │               │
│  │  ┌──────────┐  ┌──────────┐  ┌──────────────┐   │               │
│  │  │ Metrics  │  │  Traces  │  │    Logs      │   │               │
│  │  │ Processor│  │ Processor│  │  Processor   │   │               │
│  │  └────┬─────┘  └────┬─────┘  └──────┬───────┘   │               │
│  └───────┼─────────────┼───────────────┼───────────┘               │
│          │             │               │                            │
│  ┌───────▼─────────────▼───────────────▼───────────┐               │
│  │              存储层                              │               │
│  │  ┌──────────┐  ┌──────────┐  ┌──────────────┐   │               │
│  │  │Victoria  │  │Jaeger    │  │  Loki / ES   │   │               │
│  │  │Metrics   │  │(Traces)  │  │  (Logs)      │   │               │
│  │  └──────────┘  └──────────┘  └──────────────┘   │               │
│  └──────────────────────────────────────────────────┘               │
│                                                                     │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐               │
│  │  Grafana     │  │  AlertManager│  │  Incident    │               │
│  │  (Dashboards)│  │  (告警管理)  │  │  Management  │               │
│  └──────────────┘  └──────────────┘  └──────────────┘               │
└─────────────────────────────────────────────────────────────────────┘

2.2 Kubernetes 部署清单

复制代码
# observability-stack.yaml
---
apiVersion: v1
kind: Namespace
metadata:
  name: observability
---
# OpenTelemetry Collector
apiVersion: apps/v1
kind: Deployment
metadata:
  name: otel-collector
  namespace: observability
spec:
  replicas: 3
  selector:
    matchLabels:
      app: otel-collector
  template:
    metadata:
      labels:
        app: otel-collector
    spec:
      containers:
      - name: otel-collector
        image: otel/opentelemetry-collector-contrib:0.120.0
        args:
          - "--config=/etc/otel/config.yaml"
        ports:
          - containerPort: 4317  # gRPC
          - containerPort: 4318  # HTTP
        volumeMounts:
          - name: config
            mountPath: /etc/otel
        resources:
          requests:
            cpu: 500m
            memory: 512Mi
          limits:
            cpu: 2
            memory: 2Gi
      volumes:
        - name: config
          configMap:
            name: otel-config
---
apiVersion: v1
kind: ConfigMap
metadata:
  name: otel-config
  namespace: observability
data:
  config.yaml: |
    receivers:
      otlp:
        protocols:
          grpc:
            endpoint: 0.0.0.0:4317
          http:
            endpoint: 0.0.0.0:4318
    
    processors:
      batch:
        timeout: 5s
        send_batch_size: 8192
      memory_limiter:
        check_interval: 1s
        limit_mib: 1536
        spike_limit_mib: 256
      attributes:
        actions:
          - key: environment
            value: production
            action: upsert
      filter:
        error_mode: ignore
        traces:
          span:
            - 'attributes["http.target"] == "/healthz"'
    
    exporters:
      prometheus:
        endpoint: 0.0.0.0:8889
        namespace: ai_app
      otlp:
        endpoint: jaeger:4317
        tls:
          insecure: true
      loki:
        endpoint: http://loki:3100/loki/api/v1/push
        tenant_id: ai-app
    
    service:
      pipelines:
        traces:
          receivers: [otlp]
          processors: [memory_limiter, batch]
          exporters: [otlp]
        metrics:
          receivers: [otlp]
          processors: [memory_limiter, filter, batch]
          exporters: [prometheus]
        logs:
          receivers: [otlp]
          processors: [memory_limiter, batch]
          exporters: [loki]
---
# VictoriaMetrics
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: victoria-metrics
  namespace: observability
spec:
  replicas: 2
  selector:
    matchLabels:
      app: victoria-metrics
  serviceName: victoria-metrics
  template:
    metadata:
      labels:
        app: victoria-metrics
    spec:
      containers:
      - name: victoria-metrics
        image: victoriametrics/victoria-metrics:v1.108.0
        args:
          - "-storageDataPath=/data"
          - "-retentionPeriod=30d"
          - "-search.maxUniqueTimeseries=1000000"
        ports:
          - containerPort: 8428
        volumeMounts:
          - name: data
            mountPath: /data
        resources:
          requests:
            cpu: 1
            memory: 2Gi
          limits:
            cpu: 4
            memory: 8Gi
  volumeClaimTemplates:
  - metadata:
      name: data
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: ssd
      resources:
        requests:
          storage: 500Gi
---
# Jaeger
apiVersion: apps/v1
kind: Deployment
metadata:
  name: jaeger
  namespace: observability
spec:
  replicas: 2
  selector:
    matchLabels:
      app: jaeger
  template:
    metadata:
      labels:
        app: jaeger
    spec:
      containers:
      - name: jaeger
        image: jaegertracing/all-in-one:1.63.0
        env:
          - name: COLLECTOR_OTLP_ENABLED
            value: "true"
          - name: SPAN_STORAGE_TYPE
            value: "elasticsearch"
          - name: ES_SERVER_URLS
            value: "http://elasticsearch:9200"
        ports:
          - containerPort: 16686  # UI
          - containerPort: 4317   # OTLP gRPC
        resources:
          requests:
            cpu: 500m
            memory: 1Gi
          limits:
            cpu: 2
            memory: 4Gi
---
# Grafana
apiVersion: apps/v1
kind: Deployment
metadata:
  name: grafana
  namespace: observability
spec:
  replicas: 2
  selector:
    matchLabels:
      app: grafana
  template:
    metadata:
      labels:
        app: grafana
    spec:
      containers:
      - name: grafana
        image: grafana/grafana:11.4.0
        env:
          - name: GF_SECURITY_ADMIN_PASSWORD
            valueFrom:
              secretKeyRef:
                name: grafana-secret
                key: admin-password
          - name: GF_INSTALL_PLUGINS
            value: "grafana-piechart-panel,grafana-worldmap-panel"
        ports:
          - containerPort: 3000
        volumeMounts:
          - name: dashboards
            mountPath: /etc/grafana/provisioning/dashboards
          - name: datasources
            mountPath: /etc/grafana/provisioning/datasources
        resources:
          requests:
            cpu: 500m
            memory: 512Mi
          limits:
            cpu: 1
            memory: 1Gi
      volumes:
        - name: dashboards
          configMap:
            name: grafana-dashboards
        - name: datasources
          configMap:
            name: grafana-datasources
---
apiVersion: v1
kind: ConfigMap
metadata:
  name: grafana-datasources
  namespace: observability
data:
  datasources.yaml: |
    apiVersion: 1
    datasources:
      - name: VictoriaMetrics
        type: prometheus
        url: http://victoria-metrics:8428
        access: proxy
        isDefault: true
      - name: Jaeger
        type: jaeger
        url: http://jaeger:16686
        access: proxy
      - name: Loki
        type: loki
        url: http://loki:3100
        access: proxy
---
# AI 应用服务示例
apiVersion: apps/v1
kind: Deployment
metadata:
  name: llm-proxy
  namespace: ai-app
spec:
  replicas: 5
  selector:
    matchLabels:
      app: llm-proxy
  template:
    metadata:
      labels:
        app: llm-proxy
      annotations:
        sidecar.istio.io/inject: "true"
        prometheus.io/scrape: "true"
        prometheus.io/port: "9090"
    spec:
      containers:
      - name: llm-proxy
        image: registry.ai-app/llm-proxy:v2.3.1
        ports:
          - containerPort: 8080
            name: http
          - containerPort: 9090
            name: metrics
        env:
          - name: OTEL_EXPORTER_OTLP_ENDPOINT
            value: "http://otel-collector.observability:4317"
          - name: OTEL_SERVICE_NAME
            value: "llm-proxy"
          - name: OTEL_TRACES_SAMPLER
            value: "parentbased_traceidratio"
          - name: OTEL_TRACES_SAMPLER_ARG
            value: "0.1"
        livenessProbe:
          httpGet:
            path: /healthz
            port: 8080
          initialDelaySeconds: 10
          periodSeconds: 15
        readinessProbe:
          httpGet:
            path: /readyz
            port: 8080
          initialDelaySeconds: 5
          periodSeconds: 10
        resources:
          requests:
            cpu: 500m
            memory: 512Mi
          limits:
            cpu: 2
            memory: 2Gi
        volumeMounts:
          - name: config
            mountPath: /etc/app
      volumes:
        - name: config
          configMap:
            name: llm-proxy-config
---
# HPA 自动扩缩容
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: llm-proxy-hpa
  namespace: ai-app
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: llm-proxy
  minReplicas: 3
  maxReplicas: 20
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 70
  - type: Pods
    pods:
      metric:
        name: llm_proxy_queue_depth
      target:
        type: AverageValue
        averageValue: 100
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 60
      policies:
      - type: Percent
        value: 100
        periodSeconds: 15
    scaleDown:
      stabilizationWindowSeconds: 300
      policies:
      - type: Percent
        value: 10
        periodSeconds: 60

三、Dashboard 设计

3.1 宏观大盘(CEO 视角)

复制代码
┌─────────────────────────────────────────────────────────────────────────┐
│  AI 应用可观测性 · 宏观大盘                    [Last 24h] [Auto-refresh] │
├─────────────────────────────────────────────────────────────────────────┤
│                                                                         │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐│
│  │ 总请求量      │  │ 成功率       │  │ P99 延迟     │  │ 日活用户     ││
│  │ 1,234,567    │  │ 99.87%       │  │ 892ms        │  │ 89,123      ││
│  │ ▲ 12% vs 昨日 │  │ ▼ 0.05%     │  │ ▲ 45ms       │  │ ▲ 5%        ││
│  └──────────────┘  └──────────────┘  └──────────────┘  └──────────────┘│
│                                                                         │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │ 服务健康状态                                                     │   │
│  │ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐               │   │
│  │ │API  │ │NLP  │ │LLM  │ │VecDB│ │Rer  │ │Cache│               │   │
│  │ │GW   │ │Svc  │ │Proxy│ │     │ │anker│ │     │               │   │
│  │ │ 🟢  │ │ 🟢  │ │ 🟡  │ │ 🟢  │ │ 🟢  │ │ 🟢  │               │   │
│  │ │200ms│ │50ms │ │1.2s │ │35ms │ │15ms │ │2ms  │               │   │
│  │ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘               │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                                                                         │
│  ┌─────────────────────────────────────────────────────────────────┐   │
│  │ 过去 24 小时 P99 延迟趋势                                        │   │
│  │ ▁▂▃▄▅▆▇█▇▆▅▄▃▂▁▂▃▄▅▆▇█▇▆▅▄▃▂▁▂▃▄▅▆▇█▇▆▅▄▃▂▁▂▃▄▅▆▇█▇▆▅▄▃▂▁   │   │
│  │ 00:00    04:00    08:00    12:00    16:00    20:00    现在     │   │
│  └─────────────────────────────────────────────────────────────────┘   │
│                                                                         │
│  ┌─────────────────────────────┐  ┌─────────────────────────────────┐  │
│  │ 错误分布 Top 5              │  │ Token 消耗趋势                  │  │
│  │ ┌─────────────────────┐    │  │ ▁▂▃▄▅▆▇█▇▆▅▄▃▂▁              │  │
│  │ │ LLM Timeout   45%   │    │  │ 今日: 12.3M tokens              │  │
│  │ │ Rate Limit    22%   │    │  │ 费用: $246.78                   │  │
│  │ │ Vector DB     15%   │    │  └─────────────────────────────────┘  │
│  │ │ Auth Fail     10%   │    │                                      │
│  │ │ Other          8%   │    │                                      │
│  │ └─────────────────────┘    │                                      │
│  └─────────────────────────────┘  └─────────────────────────────────┘  │
└─────────────────────────────────────────────────────────────────────────┘

3.2 服务详细面板(SRE 视角)

复制代码
{
  "title": "LLM Proxy 详细面板",
  "panels": [
    {
      "title": "QPS & 延迟",
      "type": "timeseries",
      "queries": [
        "rate(llm_proxy_requests_total[1m])",
        "histogram_quantile(0.99, rate(llm_proxy_request_duration_seconds_bucket[5m]))",
        "histogram_quantile(0.95, rate(llm_proxy_request_duration_seconds_bucket[5m]))",
        "histogram_quantile(0.50, rate(llm_proxy_request_duration_seconds_bucket[5m]))"
      ]
    },
    {
      "title": "Token 消耗明细",
      "type": "stat",
      "queries": [
        "sum(rate(llm_proxy_prompt_tokens_total[5m]))",
        "sum(rate(llm_proxy_completion_tokens_total[5m]))",
        "sum(rate(llm_proxy_cost_usd_total[5m]))"
      ]
    },
    {
      "title": "模型调用分布",
      "type": "piechart",
      "queries": [
        "count by(model) (llm_proxy_requests_total)"
      ]
    },
    {
      "title": "错误原因分布",
      "type": "barchart",
      "queries": [
        "count by(error_type) (llm_proxy_errors_total)"
      ]
    },
    {
      "title": "熔断器状态",
      "type": "stat",
      "queries": [
        "llm_proxy_circuit_breaker_state{state=\"open\"}",
        "llm_proxy_circuit_breaker_state{state=\"half-open\"}",
        "llm_proxy_circuit_breaker_state{state=\"closed\"}"
      ]
    },
    {
      "title": "最近 Trace 列表",
      "type": "table",
      "datasource": "Jaeger",
      "queries": [
        "service=llm-proxy & limit=20 & lookback=1h"
      ],
      "columns": ["TraceID", "Duration", "Spans", "Services", "Errors"]
    }
  ]
}

3.3 业务指标面板(产品经理视角)

复制代码
{
  "title": "AI 应用业务指标",
  "panels": [
    {
      "title": "日活跃用户 (DAU)",
      "type": "timeseries",
      "query": "sum(increase(ai_app_active_users_total[24h]))"
    },
    {
      "title": "用户满意度评分",
      "type": "gauge",
      "query": "avg(ai_app_user_satisfaction_score)",
      "thresholds": {
        "green": 4.0,
        "yellow": 3.0,
        "red": 2.0
      }
    },
    {
      "title": "问答采纳率",
      "type": "timeseries",
      "query": "sum(rate(ai_app_answer_accepted_total[1h])) / sum(rate(ai_app_answers_total[1h]))"
    },
    {
      "title": "模型版本分布",
      "type": "piechart",
      "query": "count by(model_version) (ai_app_model_deployments)"
    },
    {
      "title": "AB 实验效果对比",
      "type": "stat",
      "queries": [
        "avg(ai_app_conversion_rate{experiment_group='control'})",
        "avg(ai_app_conversion_rate{experiment_group='treatment'})"
      ]
    },
    {
      "title": "Token 成本按部门",
      "type": "barchart",
      "query": "sum by(department) (ai_app_token_cost_total)"
    }
  ]
}

四、SLA / SLO / SLI 体系

4.1 定义

复制代码
SLA (Service Level Agreement) ------ 对外承诺
  └── SLO (Service Level Objective) ------ 内部目标
        └── SLI (Service Level Indicator) ------ 可测量指标

4.2 AI 应用 SLI 指标

复制代码
slis:
  # 可用性
  availability:
    definition: "成功响应的请求占比"
    measurement: "successful_requests / total_requests * 100"
    exclusion: "排除计划内维护窗口"
    
  # 延迟
  latency_p99:
    definition: "最慢 1% 请求的响应时间"
    measurement: "histogram_quantile(0.99, request_duration_seconds)"
    exclusion: "排除超时已熔断的请求"
    
  # 吞吐量
  throughput:
    definition: "每秒处理的请求数"
    measurement: "rate(requests_total[1m])"
    
  # 新鲜度
  freshness:
    definition: "数据从产生到可查询的时间"
    measurement: "max(event_timestamp - ingestion_timestamp)"
    
  # 正确性
  correctness:
    definition: "AI 回答被用户采纳的比例"
    measurement: "accepted_answers / total_answers * 100"
    
  # 成本效率
  cost_efficiency:
    definition: "每美元 token 产出的有效回答数"
    measurement: "accepted_answers / total_cost_usd"

4.3 SLO 目标设定

复制代码
slo_targets:
  # Tier 1: 核心对话服务
  tier_1:
    description: "用户直接感知的核心链路"
    targets:
      availability: 99.99%    # 全年宕机 < 52min
      latency_p99: 2s         # 99% 请求在 2s 内
      correctness: 85%        # 回答采纳率 > 85%
    burn_rate:
      warning: 10% / 7d       # 7 天内消耗 10% 预算
      critical: 20% / 1d      # 1 天内消耗 20% 预算
    consequences:
      - "触发 P0 告警"
      - "立即回滚最近变更"
      - "全员 On-Call"
  
  # Tier 2: 辅助功能
  tier_2:
    description: "增强体验但不阻塞核心功能"
    targets:
      availability: 99.9%
      latency_p99: 5s
      freshness: 30s
    burn_rate:
      warning: 20% / 7d
      critical: 50% / 1d
    consequences:
      - "触发 P1 告警"
      - "下个工作日修复"
  
  # Tier 3: 后台批处理
  tier_3:
    description: "离线分析和模型训练"
    targets:
      availability: 99.0%
      throughput: "每日处理 > 1M 条"
      freshness: 1h
    consequences:
      - "触发 P2 告警"
      - "本周内修复"

4.4 错误预算管理

复制代码
error_budget:
  # 以 Tier 1 为例:99.99% 可用性 = 每年 52min 错误预算
  calculation:
    annual_uptime: 99.99%
    annual_downtime_budget: "52 minutes"
    monthly_budget: "4.3 minutes"
    weekly_budget: "1 minute"
  
  consumption_tracking:
    - date: "2026-09-20"
      consumed: 12s
      remaining_weekly: 48s
      status: "🟢 正常"
    - date: "2026-09-21"
      consumed: 45s
      remaining_weekly: 15s
      status: "🟡 警告"
    - date: "2026-09-22"
      consumed: 72s
      remaining_weekly: -12s
      status: "🔴 超额"
  
  actions_on_exhaustion:
    yellow:
      - "暂停所有非紧急变更"
      - "启动根因分析"
    red:
      - "冻结所有变更"
      - "全员介入修复"
      - "通知 VP 级别"

五、运维 SOP:On-Call 手册

5.1 On-Call 流程

复制代码
┌─────────────────────────────────────────────────────────────────┐
│                      On-Call 响应流程                            │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  收到告警                                                        │
│     │                                                            │
│     ▼                                                            │
│  ┌──────────────────────┐                                        │
│  │ 1. ACKNOWLEDGE       │ ← 5分钟内必须确认                      │
│  │    确认告警           │                                        │
│  └──────────┬───────────┘                                        │
│             │                                                    │
│             ▼                                                    │
│  ┌──────────────────────┐                                        │
│  │ 2. TRIAGE            │ ← 判断级别                            │
│  │    分类定级           │                                        │
│  └──────────┬───────────┘                                        │
│             │                                                    │
│      ┌──────┴──────┐                                            │
│      ▼             ▼                                             │
│  ┌────────┐  ┌────────────┐                                      │
│  │ P0/P1  │  │  P2/P3     │                                      │
│  │ 立即响应│  │ 工作时间处理│                                     │
│  └───┬────┘  └────────────┘                                      │
│      │                                                            │
│      ▼                                                            │
│  ┌──────────────────────┐                                        │
│  │ 3. DIAGNOSE          │ ← 查看 Dashboard / Trace / Log        │
│  │    诊断根因           │                                        │
│  └──────────┬───────────┘                                        │
│             │                                                    │
│             ▼                                                    │
│  ┌──────────────────────┐                                        │
│  │ 4. MITIGATE          │ ← 回滚 / 重启 / 扩容 / 降级           │
│  │    止血恢复           │                                        │
│  └──────────┬───────────┘                                        │
│             │                                                    │
│             ▼                                                    │
│  ┌──────────────────────┐                                        │
│  │ 5. RESOLVE           │ ← 确认指标恢复正常                     │
│  │    确认解决           │                                        │
│  └──────────┬───────────┘                                        │
│             │                                                    │
│             ▼                                                    │
│  ┌──────────────────────┐                                        │
│  │ 6. POSTMORTEM        │ ← 24h 内提交事故事后分析              │
│  │    事后复盘           │                                        │
│  └──────────────────────┘                                        │
└─────────────────────────────────────────────────────────────────┘

5.2 常见故障处理手册

复制代码
incident_handbook:
  
  # 场景 1: LLM 响应超时飙升
  scenario_llm_timeout:
    symptoms:
      - "P99 延迟 > 5s"
      - "llm_proxy_timeout_errors 飙升"
      - "用户反馈回复很慢"
    diagnosis:
      step_1: "检查 LLM Provider 状态页"
      step_2: "查看 Jaeger 中 llm-proxy 的 Trace"
      step_3: "检查 Token 消耗是否有异常突增"
      step_4: "检查模型是否被限流"
    mitigation:
      option_a: "切换到备用模型 (gpt-3.5-turbo → claude-haiku)"
      option_b: "降低 max_tokens 限制"
      option_c: "启用本地缓存减少重复调用"
      option_d: "扩容 llm-proxy 实例"
    escalation:
      if_not_resolved_in: "15分钟"
      escalate_to: "AI Platform Team Lead"
  
  # 场景 2: Vector DB 不可用
  scenario_vector_db_down:
    symptoms:
      - "vector_db_errors 100%"
      - "RAG 功能完全不可用"
      - "告警: VectorDBConnectionFailed"
    diagnosis:
      step_1: "检查 Vector DB Pod 状态"
      step_2: "查看 PVC 磁盘使用率"
      step_3: "检查网络策略是否误拦截"
      step_4: "查看 Vector DB 日志"
    mitigation:
      option_a: "重启 Vector DB Pod"
      option_b: "扩容副本数"
      option_c: "降级为关键词搜索 (BM25)"
      option_d: "切换只读副本"
    escalation:
      if_not_resolved_in: "10分钟"
      escalate_to: "DBA Team"
  
  # 场景 3: 模型漂移导致回答质量下降
  scenario_model_drift:
    symptoms:
      - "用户满意度评分下降 > 10%"
      - "PSI 指标超出阈值"
      - "负面反馈增多"
    diagnosis:
      step_1: "查看 PSI Dashboard 确认漂移维度"
      step_2: "对比新旧模型版本的回答样本"
      step_3: "检查 Prompt 是否被意外修改"
      step_4: "检查上游数据分布是否变化"
    mitigation:
      option_a: "回滚到上一个稳定模型版本"
      option_b: "切换 AB 实验组到对照组"
      option_c: "临时增加后处理校验规则"
    escalation:
      if_not_resolved_in: "30分钟"
      escalate_to: "ML Team Lead"

5.3 事后复盘模板

复制代码
# 事故事后复盘报告

## 基本信息
- **事故编号**: INC-2026-0928-001
- **标题**: LLM 代理服务 P99 延迟飙升至 8s
- **严重级别**: P0
- **日期**: 2026-09-28
- **持续时间**: 23 分钟
- **影响范围**: 所有依赖 LLM 的用户请求

## 时间线
| 时间 | 事件 |
|------|------|
| 14:23 | 告警触发: P99 延迟 > 5s |
| 14:24 | On-Call 工程师确认告警 |
| 14:26 | 诊断: OpenAI API 响应变慢 |
| 14:28 | 执行降级: 切换到备用模型 |
| 14:35 | 指标恢复正常 |
| 14:46 | OpenAI 发布状态更新确认故障 |
| 15:00 | 切换回主模型 |

## 根因分析
- **直接原因**: OpenAI API 出现区域性延迟
- **根本原因**: 未配置多区域故障转移
- **促成因素**: 备用模型预热不足,首次切换延迟较高

## 改进措施
| 项目 | 负责人 | 截止日期 |
|------|--------|----------|
| 配置多区域 LLM Provider 故障转移 | @infra-team | 2026-10-05 |
| 备用模型保持 Warm Pool | @ml-team | 2026-10-03 |
| 添加 Provider 健康探测 | @sre-team | 2026-10-01 |
| 更新 On-Call 手册 LLM 降级章节 | @docs-team | 2026-09-30 |

## 附件
- [Grafana Dashboard 截图]
- [Jaeger Trace 链接]
- [告警记录]

六、未来演进方向

6.1 AI for Observability (AIOps)

复制代码
aiops_roadmap:
  
  phase_1: 异常检测自动化
    capabilities:
      - "基于历史数据的动态阈值"
      - "多维度的异常关联分析"
      - "季节性模式自动识别"
    example: "自动区分工作日/周末流量差异,避免误告警"
  
  phase_2: 根因分析智能化
    capabilities:
      - "因果推断: 从相关性到因果性"
      - "知识图谱: 服务依赖关系推理"
      - "自然语言根因描述"
    example: "LLM 分析 Trace 后直接输出: 'llm-proxy 的 call_llm span 耗时 8.2s,占整条 Trace 的 98%,根因为 OpenAI API 区域性延迟'"
  
  phase_3: 自愈系统
    capabilities:
      - "故障自动诊断 + 自动修复"
      - "渐进式回滚: 1% → 5% → 20% → 100%"
      - "混沌工程与自愈形成闭环"
    example: "检测到模型漂移 → 自动触发 AB 实验回滚 → 验证指标恢复 → 通知团队"
  
  phase_4: 预测性可观测性
    capabilities:
      - "容量预测: 提前 24h 预测资源需求"
      - "故障预测: 基于模式识别的故障预警"
      - "成本预测: Token 消耗和费用预估"
    example: "预测今晚 20:00 会有流量高峰,提前扩容 3 个副本"

6.2 技术演进趋势

复制代码
technology_trends:
  
  # eBPF 深度观测
  ebpf:
    description: "无需修改代码即可观测内核态行为"
    use_cases:
      - "网络延迟的精确测量"
      - "文件 I/O 性能分析"
      - "系统调用追踪"
    tools:
      - "Pixie"
      - "Cilium Tetragon"
      - "Parca (连续分析)"
  
  # OpenTelemetry 标准化
  opentelemetry:
    description: "可观测性的行业标准,统一 Metrics/Traces/Logs"
    trends:
      - "Profiling Signal 加入 OTel"
      - "eBPF 与 OTel 融合"
      - "OTel Operator 简化部署"
  
  # 低成本存储
  cost_effective_storage:
    description: "海量观测数据的存储成本优化"
    strategies:
      - "列式存储 + 高压缩比"
      - "对象存储作为冷备"
      - "采样 + 聚合降低数据量"
    tools:
      - "Grafana Mimir"
      - "Thanos"
      - "ClickHouse"
  
  # 持续验证
  continuous_validation:
    description: "可观测性本身也需要被观测"
    practices:
      - "SLI 覆盖率检查"
      - "告警质量评分"
      - "Dashboard 使用率统计"

七、总结:从埋点到诊断的完整链路

复制代码
┌─────────────────────────────────────────────────────────────────────────┐
│                     AI 应用可观测性:从埋点到诊断                          │
│                                                                         │
│  埋点层 (Instrumentation)                                               │
│  ├── 代码埋点: OpenTelemetry SDK                                       │
│  ├── 自动埋点: Istio Sidecar / eBPF                                   │
│  └── 业务埋点: 自定义 Metrics / Events                                │
│         │                                                              │
│         ▼                                                              │
│  采集层 (Collection)                                                    │
│  ├── Metrics: Prometheus Exporter (端口 :9090)                        │
│  ├── Traces: OTLP Exporter → OpenTelemetry Collector                  │
│  └── Logs

│  └── Logs: Structured JSON → stdout → Fluentd
│         │
│         ▼
│  传输层 (Transport)
│  ├── 缓冲: Kafka (削峰填谷)
│  ├── 采样: Tail-Based Sampling (错误全采 / 慢请求半采 / 正常 1%)
│  └── 压缩: gzip / snappy (减少带宽)
│         │
│         ▼
│  存储层 (Storage)
│  ├── 热存储 (7天): ClickHouse / VictoriaMetrics
│  ├── 温存储 (30天): Elasticsearch / Jaeger
│  └── 冷存储 (90天): S3 / GCS (Parquet 格式)
│         │
│         ▼
│  分析层 (Analysis)
│  ├── 指标分析: PromQL 查询 / 聚合 / 预测
│  ├── 链路分析: Trace 查询 / 服务拓扑 / 根因定位
│  └── 日志分析: 全文检索 / 模式识别 / 聚类
│         │
│         ▼
│  展示层 (Visualization)
│  ├── Grafana Dashboards: 宏观大盘 / 服务面板 / 业务面板
│  ├── Jaeger UI: Trace 详情 / 火焰图 / 比较视图
│  └── 自定义 Portal: 业务指标 / 成本报表 / SLA 看板
│         │
│         ▼
│  行动层 (Action)
│  ├── 告警: AlertManager → 电话 / Slack / 邮件
│  ├── 自动化: AutoResponder → 熔断 / 扩容 / 回滚 / 限流
│  └── 混沌: Chaos Engine → 故障注入 / 韧性验证 / 改进闭环
│
└─────────────────────────────────────────────────────────────────────────┘

八、附录:常用命令速查

8.1 排查问题三板斧

复制代码
# 1. 看指标 ------ 哪里慢了?
# P99 延迟最高的服务 Top 5
topk(5, histogram_quantile(0.99, 
  sum(rate(http_request_duration_seconds_bucket[5m])) by (service, le)))

# 2. 看链路 ------ 为什么慢?
# 查询最近 1 小时最慢的 10 条 Trace
# 在 Jaeger UI 中执行:
#   Service: llm-proxy
#   Operation: call_llm
#   Min Duration: 1s
#   Lookback: 1h
#   Limit: 10

# 3. 看日志 ------ 报了什么错?
# 查询特定 TraceID 的所有日志
{app="llm-proxy"} |= "trace_id=a1b2c3d4e5f6"

8.2 常用 PromQL 查询

复制代码
# 服务可用性 (排除 503 熔断)
sum(rate(http_requests_total{status!~"5.."}[5m])) 
/ 
sum(rate(http_requests_total[5m])) * 100

# P99 延迟趋势
histogram_quantile(0.99, 
  sum(rate(http_request_duration_seconds_bucket[5m])) by (le))

# 错误率 (5xx / total)
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m])) * 100

# Token 消耗速率
sum(rate(llm_token_total[5m])) by (model)

# 每秒成本
sum(rate(llm_cost_usd_total[5m]))

# 熔断器打开次数
increase(circuit_breaker_open_total[1h])

# 最慢的 5 个端点
topk(5, avg by(endpoint) (http_request_duration_seconds_sum / http_request_duration_seconds_count))

# 各服务 Span 数量
count by(service.name) (spans{})

8.3 常用 kubectl 命令

复制代码
# 查看所有 Pod 状态
kubectl get pods -n ai-app -o wide

# 查看 Pod 日志
kubectl logs -n ai-app deployment/llm-proxy --tail=100 -f

# 查看 Pod 资源使用
kubectl top pod -n ai-app

# 查看 HPA 状态
kubectl get hpa -n ai-app

# 查看事件
kubectl get events -n ai-app --sort-by='.lastTimestamp'

# 端口转发 (本地访问 Jaeger)
kubectl port-forward -n observability svc/jaeger 16686:16686

# 进入 Pod 调试
kubectl exec -it -n ai-app pod/llm-proxy-xxx -- sh

# 查看 ConfigMap
kubectl get configmap -n ai-app llm-proxy-config -o yaml

# 滚动重启
kubectl rollout restart -n ai-app deployment/llm-proxy

# 查看滚动状态
kubectl rollout status -n ai-app deployment/llm-proxy

九、推荐阅读与工具

9.1 必读书籍

书名 作者 推荐理由
《Site Reliability Engineering》 Google SRE Team SRE 圣经,定义了现代可观测性理念
《The Art of Monitoring》 James Turnbull 从零搭建监控体系的实操指南
《Distributed Tracing in Practice》 Austin Parker 等 链路追踪的权威实践手册
《Chaos Engineering》 Casey Rosenthal 等 混沌工程的奠基之作
《Observability Engineering》 Charity Majors 等 可观测性三大支柱的系统讲解

9.2 开源工具推荐

类别 工具 说明
Metrics​ VictoriaMetrics 高性能时序数据库,兼容 PromQL
Traces​ Jaeger / Tempo 分布式链路追踪
Logs​ Loki / Quickwit 低成本日志存储
Profiling​ Parca / Pyroscope 持续性能分析
eBPF​ Pixie / Cilium 零侵入内核观测
混沌​ Chaos Mesh / Litmus K8s 原生混沌工程
告警​ AlertManager / Keep 告警管理和抑制
Dashboard​ Grafana / Perses 可视化面板

9.3 在线工具

工具 用途 地址
数字转大写 金额转换 zz365.top/daxie
JSON 格式化 Trace 数据美化 zz365.top/json
时间戳转换 Unix 时间 ↔ 可读时间 zz365.top/timestamp
Cron 表达式 告警周期配置 zz365.top/cron
Base64 编解码 Token 解码调试 zz365.top/base64

十、结语

至此,《AI 应用可观测性:从埋点到诊断》十讲全部结束。我们从最基础的观测维度出发,一步步构建了完整的可观测性体系:

复制代码
第1讲   👉 知道要看什么(指标体系)
第2讲   👉 知道怎么串起来(Trace)
第3-5讲 👉 知道 AI 特有的坑(LLM / 漂移 / Prompt)
第6讲   👉 知道怎么存(日志结构)
第7讲   👉 知道怎么反应(告警响应)
第8讲   👉 知道怎么查根因(链路分析)
第9讲   👉 知道怎么验证韧性(混沌工程)
第10讲  👉 知道怎么落地(生产实战)

记住三条铁律:

  1. 没有度量就没有改进 ------ 先埋点,再优化
  2. 没有 Trace 的日志就是孤岛 ------ 永远带着 TraceID
  3. 没有自动化的告警就是噪音 ------ 告警必须能触发行动

祝你的 AI 应用永远稳定、高效、可观测!


🧰 开发之余的小工具推荐

设计混沌实验场景时,经常需要计算各种时间窗口和延迟参数。zz365.top 的在线计算器可以快速进行毫秒/秒/分钟的单位换算和百分比计算,帮助你在配置故障参数时更精确。所有计算纯前端完成,不需要联网。


相关推荐
青柠之夏cc1 小时前
AI与教育:个性化学习助手让“因材施教“成为现实,但代价是隐私?
人工智能·学习
雷焰财经1 小时前
AI开始学会“越界”:当智能体拥有行动能力,人工智能的竞争已经进入安全深水区
人工智能·安全
可乐ea1 小时前
Agent 模型分级升级路由:用置信度阈值把贵模型省下来
人工智能·算法·机器学习·结构化输出·置信度路由·模型升级路由·大模型成本优化
hsfxuebao1 小时前
Loop Engineering 保姆级教程 + 项目实战
人工智能·后端
晚安code1 小时前
大模型是什么?讲透 Token、上下文窗口、Temperature 和 MoE
人工智能·深度学习·机器学习
xianghongtao01161 小时前
麦肯锡2026技术趋势02_智能体AI_研究解读
大数据·人工智能
AI日报派送佬1 小时前
2026年10月2日AI行业日报|AI监管全面加码,算力基建革新,视觉推理技术突破
人工智能·ai智能体·大模型技术·ai日报·算力基建·ai科研·ai产业落地
数字化顾问2 小时前
(138页PPT)四大咨询矿业集团流程梳理与优化报告(附下载方式)
大数据·人工智能
liuchangng2 小时前
Agent的state注入
人工智能·agent·jev·决策模型