Linux 主机监控(node-exporter)
原理
kube-prometheus 用 DaemonSet 在每个节点跑一个 node-exporter Pod,监听 9100 端口,暴露宿主机指标。
text
每个节点 → node-exporter Pod → :9100 → Prometheus 抓 → Grafana 展示
验证
# 1. 看 Pod(每个节点一个)
kubectl -n monitoring get pods -l app.kubernetes.io/name=node-exporter -o wide
# 2. 看端口监听
ss -tunlp | grep 9100
# tcp LISTEN ... 192.168.36.152:9100
# 3. 直接抓指标
curl -s http://192.168.36.152:9100/metrics | head
查看 Service
kubectl -n monitoring get svc node-exporter -o yaml
关键:Headless Service (clusterIP: None),端口名 https-metrics。
查看 ServiceMonitor
kubectl -n monitoring get servicemonitor node-exporter -o yaml
作用:告诉 Prometheus 去 monitoring 命名空间抓 node-exporter Service 的 https-metrics 端口。
Prometheus 验证
浏览器 http://<prometheus>:9090
- Status → Targets → 搜
node-exporter,应每节点一条且 UP
Grafana web界面配置监控模板: (模版ID:13105)
非云原生应用 Redis 集群监控
原理
Redis chart 自带 redis-exporter (监听 9121),通过 ServiceMonitor 让 Prometheus 抓取。
text
Redis Pod ──┬── redis 容器(6379)
└── redis-exporter 容器(9121)
↑
ServiceMonitor 抓取
↑
Prometheus
部署 Redis 集群
redis-values-metrics.yml:
global:
imageRegistry: "reg.westos.org"
security:
allowInsecureImages: true
redis:
password: "westos"
metrics:
enabled: true # 开启 exporter
serviceMonitor:
enabled: true
namespace: "monitoring" # ServiceMonitor 放到 monitoring
helm install redis-cluster \
-f /resources/charts/values/redis-values-metrics.yml \
/resources/charts/redis-23.2.6.tgz \
-n redis
验证部署
# Pod:每个 Redis 实例 2/2(redis + exporter)
kubectl -n redis get pods
# redis-cluster-master-0 2/2 Running
# redis-cluster-replicas-0 2/2 Running
# redis-cluster-replicas-1 2/2 Running
# redis-cluster-replicas-2 2/2 Running
# metrics Service
kubectl -n redis get svc redis-cluster-metrics
# ClusterIP ... 9121/TCP
验证 metrics 端点
# 直接 curl(从集群内 Pod)
curl 10.109.211.101:9121/metrics | head
# 或 port-forward 出来
kubectl -n redis port-forward svc/redis-cluster-metrics 9121:9121
curl http://127.0.0.1:9121/metrics | head
关键指标:
redis_up 1
redis_connected_clients
redis_memory_used_bytes
redis_commands_processed_total
查看 ServiceMonitor
kubectl -n monitoring describe servicemonitors redis-cluster
关键:port: metrics + selector 匹配 redis-cluster-metrics 的标签。
RBAC 授权
Prometheus 的 ServiceAccount 默认没有权限读取 redis 命名空间的 endpoints,需要手动授权:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: prometheus-endpoints-view
namespace: redis
rules:
- apiGroups: [""]
resources: ["endpoints", "pods", "services"]
verbs: ["get", "list", "watch"]
- apiGroups: ["discovery.k8s.io"]
resources: ["endpointslices"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: prometheus-k8s-endpoints
namespace: redis
subjects:
- kind: ServiceAccount
name: prometheus-k8s
namespace: monitoring
roleRef:
kind: Role
name: prometheus-endpoints-view
apiGroup: rbac.authorization.k8s.io
kubectl apply -f rbac.yaml
不加 RBAC → Prometheus 能发现 target,但状态是
DOWN,报forbidden。
Prometheus 验证
浏览器 http://<prometheus>:9090
- Targets → 搜
redis,应出现serviceMonitor/monitoring/redis-cluster/0,状态 UP
Grafana 看
导入官方 dashboard(ID: 763 Redis Dashboard)或使用内置面板。
非云原生 NGINX 监控
原理
Nginx 业务容器(暴露 /status)
↑
nginx-exporter(转成 Prometheus 格式,:9113)
↑
ServiceMonitor → Prometheus 抓取
准备 ConfigMap(提供 /status)
default.conf:
server {
listen 80;
server_name localhost;
location / {
root /usr/share/nginx/html;
index index.html index.htm;
}
location /status {
stub_status on;
}
error_page 500 502 503 504 /50x.html;
location = /50x.html {
root /usr/share/nginx/html;
}
}
kubectl create cm nginx-conf --from-file=./default.conf
部署 Nginx + Exporter
deploy-nginx.yml 关键部分:
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: nginx-server
spec:
selector:
matchLabels:
app: nginx
replicas: 3
template:
metadata:
labels:
app: nginx
annotations:
prometheus.io/scrape: 'true'
prometheus.io/port: '9113'
spec:
containers:
- name: nginx
image: reg.westos.org/nginx/nginx:1.29
ports:
- containerPort: 80
volumeMounts:
- name: nginxconf
mountPath: /etc/nginx/conf.d
- name: nginx-exporter
image: reg.westos.org/nginx/nginx-prometheus-exporter:1.5.1
args:
- '-nginx.scrape-uri=http://localhost/status'
resources:
limits:
memory: 128Mi
cpu: 500m
ports:
- containerPort: 9113
volumes:
- name: nginxconf
configMap:
name: nginx-conf
kubectl apply -f deploy-nginx.yml
kubectl get pod | grep nginx
创建 Service
service-nginx.yml:
---
apiVersion: v1
kind: Service
metadata:
name: webapp
labels:
app: nginx
spec:
selector:
app: nginx
ports:
- protocol: TCP
port: 80
targetPort: 80
name: web
- protocol: TCP
port: 9113
targetPort: 9113
name: metrics
kubectl apply -f service-nginx.yml
kubectl get svc webapp
测试指标
curl 10.99.152.210:9113/metrics | head
关键指标:nginx_up、nginx_connections_active、nginx_http_requests_total
创建 ServiceMonitor
nginx-servicemonitor.yml:
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: nginx-exporter
namespace: monitoring
labels:
k8sapp: nginx-exporter
namespace: monitoring
spec:
jobLabel: k8s-app
endpoints:
- port: metrics
interval: 30s
scheme: http
path: /metrics
selector:
matchLabels:
app: nginx
namespaceSelector:
matchNames:
- default
kubectl apply -f nginx-servicemonitor.yml
压测 + 验证
# 制造流量
while true; do curl -s http://10.99.152.210 > /dev/null; done
Prometheus
UI → Targets 搜 nginx-exporter,状态应为 UP。
Grafana web界面配置监控模板
https://github.com/nginx/nginx-prometheus-exporter/blob/main/grafana/dashboard.json
黑盒监控(Blackbox Exporter)
原理
从外部探测服务可用性,不关心内部指标。模拟用户请求(HTTP/HTTPS/TCP/ICMP),检查响应是否符合预期。
text
Blackbox Exporter → 探测目标(如 https://www.baidu.com)
↑
Prometheus 通过 Probe 资源抓取
| 黑盒监控 | 白盒监控 | |
|---|---|---|
| 视角 | 外部观察 | 内部指标 |
| 典型 | Blackbox Exporter | node-exporter、redis-exporter |
| 关注 | 服务是否可达、响应是否正常 | CPU、内存、QPS 等内部状态 |
确认 Blackbox Exporter 已部署
# 查看 Pod(3/3 表示 exporter + config-reloader + ?)
kubectl -n monitoring get pods -l app.kubernetes.io/name=blackbox-exporter
# 查看 Service
kubectl -n monitoring get svc -l app.kubernetes.io/name=blackbox-exporter
# 端口:9115(对外探测)、19115(Probe 用)
配置探针(Probe CRD)
blackbox-config.yml:
yaml
apiVersion: monitoring.coreos.com/v1
kind: Probe
metadata:
name: blackbox
namespace: monitoring
spec:
interval: 30s
jobName: blackbox
module: http_2xx # 探测模块:期望 HTTP 2xx
prober:
url: blackbox-exporter.monitoring.svc:19115
scheme: http
path: /probe
targets:
staticConfig:
static:
- https://www.baidu.com # 要探测的目标
kubectl apply -f blackbox-config.yml
说明 :Prometheus 通过 Probe 资源自动发现探测目标,调用 Blackbox Exporter 的
/probe接口。
Prometheus 查看指标
- Status → Targets → 搜
blackbox,应出现探测目标,状态 UP
Grafana 展示
导入官方模板:
- Dashboard ID:13659