《OpenShift / RHEL / DevSecOps 汇总目录》
本文已在 OpenShift 4.22 上验证。
文章目录
- 系统负载、业务负载
- 系统负载和业务负载隔离的技术手段
-
- [Performance Profile](#Performance Profile)
- [Workload Partitioning](#Workload Partitioning)
- 使用工作负载分区特性
-
- [只用 PerformanceProfile 特性](#只用 PerformanceProfile 特性)
-
- [创建 PerformanceProfile 前](#创建 PerformanceProfile 前)
- [创建 PerformanceProfile 后](#创建 PerformanceProfile 后)
- 用户负载验证
- [使用 PerformanceProfile + Workload Partitioning 特性](#使用 PerformanceProfile + Workload Partitioning 特性)
- 参考
系统负载、业务负载
一个 OpenShift 集群的管理平面用来运行容器平台的系统负载,数据平面运行用户的业务负载。
- 系统负载:或者称为平台负载、系统 pod、平台 pod,是用来跑系统组件(kubelet、CRI、OS 内核线程、kube-proxy 等)。
- 业务负载:或者称为用户负载、用户 pod。
当使用多个节点分别运行管理平面和数据平面的时候,系统负载和业务负载不会出现争抢节点资源的情况。但如果管理平面和数据平面在同一节点运行的时候(例如单节点或 3 节点 OpenShift 集群),系统负载和业务负载可能出现争抢节点 CPU 资源的情况。
为了解决上述问题,我们可以使用 Performance Profile + Workload Partitioning 将系统负载和业务负载隔离在不同的 CPU 内核上运行,这样即便是在高负载的时候,系统负载和业务负载之间才不会相互影响。
系统负载和业务负载隔离的技术手段
要实现系统负载和业务负载隔离,就是无论系统负载和业务负载哪一方的负载增加,都不应影响另一方正常运行和响应速度。在 OpenShift 中Performance Profile 和 Workload Partitioning 是实现系统负载和业务负载隔离的主要技术手段。
Performance Profile
Performance Profile 是 Node Tuning Operator 的一个 CRD,通过它可以设置 CPU、内存、NUMA 等的适用策略,以适配诸如低延迟、real-time、低功耗等不同的性能场景的需要。由于它是在 OpenShift 安装完后根据需要设置的,因此功能可以随时启用和关闭。
在 Performance Profile 中可以为 CPU 设置 reserved 和 isolated 的 cpuset,其中 reserved 是保留给平台负载用的 CPU,isolated 是保留给用户负载用的 CPU,并且通常 all-cpus = isolated + reserved。
yaml
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
name: openshift-node-performance-profile
spec:
cpu:
# Set core for OpenShift system components and the OS
reserved: "0-7"
# Set core for user workloads
isolated: "8-15"
。。。
需要注意的是 PerformanceProfile 虽能保证"业务 pod 只能跑在 isolated 的 cpuset",但不能保证"只有业务 pod 能跑在 isolated 的 cpuset",系统 Pod / 基础设施 Pod 也有可能跑到 isolated cpuset 上。这是因为 PerformanceProfile 使用的 CPU Manager 的 static 策略只对 QoS 为 Guaranteed 且请求整数核的 Pod 生效。而非 Guaranteed(Burstable/BestEffort)的 Pod 则不受 CPU Manager 管控,它们可以在所有 CPU(reserved+isolated)之间随意调度。
Workload Partitioning
PerformanceProfile 的 reserved 可以保护宿主机系统进程,但不保证 isolated 不被基础设施 Pod 占用,而 Workload Partitioning 可以解决 PerformanceProfile 隔离不彻底的问题。Workload Partitioning 不是简单 cpuset,它可使用 cgroup v2 + cpuset + IRQ partitioning 等多种隔离手段让系统负载和业务负载实现更充分的隔离,能够将基础设施 Pod 也锁死在 reserved。
无法单独使用 Workload Partitioning,它必须配合 PerformanceProfile 一起使用才能实现系统负载和业务负载完全隔离。而 PerformanceProfile 是可以不依赖 Workload Partitioning 而独立使用,但系统负载和业务负载无法实现完全隔离。
虽然 Workload Partitioning 有更强的隔离性来隔离系统负载和业务负载,但使用它也是有一定的条件和代价的。因为当系统负载和业务负载只能运行在各自的 cpuset 中,那么就需要各自部分有更精确的 CPU 数量划分,以避免即便 CPU 整体还有较多剩余也无法调度 Pod 的情况。因此 Workload Partitioning 通常主要用在以下典型场景:SNO - 单节点 OpenShift / 3-node OpenShift 集群、Telcom/NFV 强 SLA 和审计要求。而在以下场景中无法使用 Workload Partitioning:
- master / infra / worker 分开部署在不同的节点:平台与业务本来就不在同一台机器上,因此彼此已不会相互干扰。
- Workload Partitioning 只能在安装时启用且不可逆,因此无法对已安装完的 OpenShift 集群启用该功能。
- 配置不同的节点:PerformanceProfile 可按 MCP 分别定义 isolated/reserved;但 Workload Partitioning 的配置 cpuPartitioningMode: AllNodes 只能作用于所有节点,因此需要集群中的所有节点具有相同的配置。
使用工作负载分区特性
只用 PerformanceProfile 特性
如果只用 PerformanceProfile 特性,只能影响到 CPU Request 为整数且 QoS 为 Guaranteed 的 Pod。这种 Pod 被分配的 CPU 是下图 non system reserved CPUs 中的 exclusive CPUs,而其他类型的 Pod 可以使用下图中所有 shared pool 的 CPU。
bash
+----------------------+------------------------------+
| system reserved CPUs | non system reserved CPUs |
+----------------------+------------------------------+
| shared pool | exclusive CPUs | shared pool |
+----------------------+------------------------------+
------------->
创建 PerformanceProfile 前
- 首先查看节点 CPU 的核数,本环境的 CPU 共有 32 个内核。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- lscpu --extend
。。。
CPU NODE SOCKET CORE L1d:L1i:L2:L3 ONLINE
0 0 0 0 0:0:0:0 yes
1 0 0 1 1:1:1:0 yes
2 0 0 2 2:2:2:0 yes
3 0 0 3 3:3:3:0 yes
4 0 0 4 4:4:4:0 yes
5 0 0 5 5:5:5:0 yes
6 0 0 6 6:6:6:0 yes
7 0 0 7 7:7:7:0 yes
8 0 0 8 8:8:8:0 yes
9 0 0 9 9:9:9:0 yes
10 0 0 10 10:10:10:0 yes
11 0 0 11 11:11:11:0 yes
12 0 0 12 12:12:12:0 yes
13 0 0 13 13:13:13:0 yes
14 0 0 14 14:14:14:0 yes
15 0 0 15 15:15:15:0 yes
16 0 0 16 16:16:16:0 yes
17 0 0 17 17:17:17:0 yes
18 0 0 18 18:18:18:0 yes
19 0 0 19 19:19:19:0 yes
20 0 0 20 20:20:20:0 yes
21 0 0 21 21:21:21:0 yes
22 0 0 22 22:22:22:0 yes
23 0 0 23 23:23:23:0 yes
24 0 0 24 24:24:24:0 yes
25 0 0 25 25:25:25:0 yes
26 0 0 26 26:26:26:0 yes
27 0 0 27 27:27:27:0 yes
28 0 0 28 28:28:28:0 yes
29 0 0 29 29:29:29:0 yes
30 0 0 30 30:30:30:0 yes
31 0 0 31 31:31:31:0 yes
- 查看节点的 CPU 总容量和可分配量。
bash
$ oc describe node $(oc get nodes -o jsonpath='{.items[0].metadata.name}') | grep -E 'Capacity:|Allocatable:' -A 1
Capacity:
cpu: 32
--
Allocatable:
cpu: 31500m
- 检查系统的 kubelet 进程的 CPU 负载亲和性。确认在未使用 PerformanceProfile 的时候 kubelet 进程可以在 0-31 任意一个内核上运行。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c 'taskset -c -p $(pidof /usr/bin/kubelet)'
。。。
pid 3392's current affinity list: 0-31
创建 PerformanceProfile 后
- 创建 PerformanceProfile,它将 CPU 内核分为 2 部分,0-7 为 system reserved CPUs,8-31 为 non system reserved CPUs。
yaml
$ oc apply -f - <<EOF
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
name: openshift-node-performance-profile
spec:
cpu:
reserved: 0-7
isolated: 8-31
machineConfigPoolSelector:
pools.operator.machineconfiguration.openshift.io/master: ''
nodeSelector:
node-role.kubernetes.io/master: ''
numa:
topologyPolicy: "restricted"
workloadHints:
realTime: false
highPowerConsumption: false
perPodPowerManagement: false
EOF
- 查看基于 PerformanceProfile 生成的 kubeletconfig 和 Tuned 对象。
bash
$ oc get kubeletconfig performance-openshift-node-performance-profile -o jsonpath='reservedSystemCPUs={.spec.kubeletConfig.cpuManagerPolicy}{"\n"}cpuManagerPolicy={.spec.kubeletConfig.reservedSystemCPUs}{"\n"}'
cpuManagerPolicy=static
reservedSystemCPUs=0-7
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- chroot /host grep -E 'cpuManagerPolicy|reservedSystemCPUs' /etc/kubernetes/kubelet.conf
。。。
cpuManagerPolicy=static
reservedSystemCPUs=0-7
$ oc get Tuned openshift-node-performance-openshift-node-performance-profile -n openshift-cluster-node-tuning-operator
NAME VALID AGE
openshift-node-performance-openshift-node-performance-profile True 3h58m
- 查看 kubelet 进程可使用节点 CPU 的范围。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c 'taskset -c -p $(pidof /usr/bin/kubelet)'
。。。
pid 4426's current affinity list: 0-7
- 查看当前内核启动命令,注意
systemd.cpu_affinity和isolcpus=<cpulist>。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /proc/cmdline | tr ' ' '\n' | grep -E 'cpu'
。。。
tuned.non_isolcpus=000000ff
systemd.cpu_affinity=0,1,2,3,4,5,6,7
isolcpus=managed_irq,8-31
- 查看当前内核启动参数,注意 Allocatable 已经变化为 24。
bash
$ oc describe node $(oc get nodes -o jsonpath='{.items[0].metadata.name}') | grep -E 'Capacity|Allocatable' -A 1
Capacity:
cpu: 32
--
Allocatable:
cpu: 24
- 执行命令,查看 /host/etc/crio/crio.conf.d/ 中的 99-runtimes.conf 文件内容。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- ls -al /host/etc/crio/crio.conf.d/
。。。
total 16
drwxr-xr-x. 2 root root 120 Oct 7 06:12 .
drwxr-xr-x. 4 root root 78 Oct 7 06:12 ..
-rw-r--r--. 1 root root 3254 Oct 7 06:09 00-default
-rw-r--r--. 1 root root 625 Oct 7 06:09 99-runtimes.conf
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /host/etc/crio/crio.conf.d/99-runtimes.conf
。。。
[crio.runtime]
infra_ctr_cpuset = "0-7"
。。。
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- chroot /host cat /var/lib/kubelet/cpu_manager_state | jq
。。。
{
"policyName": "static",
"defaultCpuSet": "0-7,9-31",
"entries": {
"d0ccbdd0-95e6-4928-b0ea-551c7cd4a8dd": {
"fedora-minimal-exclusive": "8"
}
},
"checksum": 3215225415
}
用户负载验证
- 创建测试 Pod。其中名为 fedora-minimal-exclusive 的 Pod 的 QoS 为 Guaranteed,且申请的 CPU 为 1;而名为 fedora-minimal-shared 的 Pod 的 QoS 虽也为 Guaranteed,但申请的 CPU 不是整数。
yaml
$ oc new-project pp-demo
$ oc apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata:
labels:
app: fedora-minimal
name: fedora-minimal
spec:
replicas: 1
selector:
matchLabels:
app: fedora-minimal
template:
metadata:
labels:
app: fedora-minimal
spec:
containers:
- image: registry.fedoraproject.org/fedora-minimal:latest
name: fedora-minimal-exclusive
command:
- /bin/sleep
- "12345"
resources:
requests:
memory: "128Mi"
cpu: "1"
limits:
memory: "128Mi"
cpu: "1"
- image: registry.fedoraproject.org/fedora-minimal:latest
name: fedora-minimal-shared
command:
- /bin/sleep
- "54321"
resources:
requests:
memory: "128Mi"
cpu: "200m"
limits:
memory: "128Mi"
cpu: "200m"
EOF
- 执行以下命令,分别查看 2 个容器可以使用的 CPU。确认 fedora-minimal-exclusive 可用的 CPU 为 8 号,而 fedora-minimal-shared 可用的 CPU 为 0-7,9-31。
bash
$ oc -n pp-demo exec deploy/fedora-minimal -c fedora-minimal-exclusive -- grep -i Cpus_allowed_list /proc/self/status
Cpus_allowed_list: 8
$ oc -n pp-demo exec deploy/fedora-minimal -c fedora-minimal-shared -- grep -i Cpus_allowed_list /proc/self/status
Cpus_allowed_list: 0-7,9-31
使用 PerformanceProfile + Workload Partitioning 特性
为了解决 PerformanceProfile 分离工作负载的功能较弱的问题,我们可以结合 OpenShift 独有的 Workload Partitioning 特性实现功能更强的工作负载隔离。
启用 Workload Partitioning 特性
OpenShift 工作负载分区特性只能在安装集群时就启用,无法在安装完成后启用。以下是一个 3 节点 OpenShift 的 install-config.yaml 示例,其中 cpuPartitioningMode: AllNodes 的声明启用了 Workload Partitioning 功能。
yaml
apiVersion: v1
baseDomain: devcluster.openshift.com
cpuPartitioningMode: AllNodes
compute:
- architecture: amd64
name: worker
replicas: 0
controlPlane:
architecture: amd64
name: master
replicas: 3
。。。
先查看在启用 Workload Partitioning 的 OpenShift 中会有 2 个和 partitioning 相关的 machineconfig,然后再查看其中 spec.config.storage.files.path 位置的配置。
json
$ oc get machineconfig | grep partition
01-master-cpu-partitioning 3.2.0 11h
01-worker-cpu-partitioning 3.2.0 11h
$ oc get machineconfig 01-master-cpu-partitioning -o jsonpath={.spec.config.storage.files} | jq
[
{
"contents": {
"source": "data:text/plain;charset=utf-8;base64,CnsKICAibWFuYWdlbWVudCI6IHsKICAgICJjcHVzZXQiOiAiIgogIH0KfQo=",
"verification": {}
},
"group": {},
"mode": 420,
"path": "/etc/kubernetes/openshift-workload-pinning",
"user": {}
},
{
"contents": {
"source": "data:text/plain;charset=utf-8;base64,CltjcmlvLnJ1bnRpbWUud29ya2xvYWRzLm1hbmFnZW1lbnRdCmFjdGl2YXRpb25fYW5ub3RhdGlvbiA9ICJ0YXJnZXQud29ya2xvYWQub3BlbnNoaWZ0LmlvL21hbmFnZW1lbnQiCmFubm90YXRpb25fcHJlZml4ID0gInJlc291cmNlcy53b3JrbG9hZC5vcGVuc2hpZnQuaW8iCnJlc291cmNlcyA9IHsgImNwdXNoYXJlcyIgPSAwLCAiY3B1c2V0IiA9ICIiIH0K",
"verification": {}
},
"group": {},
"mode": 420,
"path": "/etc/crio/crio.conf.d/01-workload-pinning-default.conf",
"user": {}
}
]
可以根据上一步的结果,查看节点 /etc/kubernetes 和 /host/etc/crio/crio.conf.d/ 目录中包含哪些文件,其中如果未启用 Workload Partitioning,则不会有 openshift-workload-pinning 和 01-workload-pinning-default.conf 文件。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- chroot /host ls -al /etc/kubernetes/
。。。
drwxr-xr-x. 7 root root 4096 Sep 28 11:19 .
drwxr-xr-x. 100 root root 8192 Sep 28 10:14 ..
-rw-r--r--. 1 root root 102 Sep 28 10:08 apiserver-url.env
drwxr-xr-x. 2 root root 6 Sep 28 10:13 cloud
-rw-r--r--. 1 root root 0 Sep 28 10:08 cloud.conf
drwxr-xr-x. 3 root root 19 Sep 28 10:10 cni
-rw-r--r--. 1 root root 174 Sep 28 10:08 crio-metrics-proxy.cfg
-rw-------. 1 root root 10984 Sep 27 23:29 kubeconfig
-rw-r--r--. 1 root root 8332 Sep 28 11:19 kubelet-ca.crt
drwxr-xr-x. 3 root root 20 Sep 28 10:13 kubelet-plugins
-rw-r--r--. 1 root root 4119 Sep 28 10:08 kubelet.conf
drwxr-xr-x. 2 root root 158 Sep 28 10:13 manifests
-rw-r--r--. 1 root root 44 Sep 28 10:08 openshift-workload-pinning
drwxr-xr-x. 23 root root 4096 Sep 28 10:13 static-pod-resources
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- ls -al /host/etc/crio/crio.conf.d/
。。。
total 8
drwxr-xr-x. 2 root root 64 Sep 28 10:13 .
drwxr-xr-x. 4 root root 78 Sep 28 10:13 ..
-rw-r--r--. 1 root root 3254 Sep 28 10:08 00-default
-rw-r--r--. 1 root root 204 Sep 28 10:08 01-workload-pinning-default.conf
最后再查看 openshift-workload-pinning 和 01-workload-pinning-default.conf 文件的内容,确认其中的 "cpuset": "" 和 resources = { "cpushares" = 0, "cpuset" = "" }。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- chroot /host cat /etc/kubernetes/openshift-workload-pinning
。。。
{
"management": {
"cpuset": ""
}
}
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /host/etc/crio/crio.conf.d/01-workload-pinning-default.conf
。。。
[crio.runtime.workloads.management]
activation_annotation = "target.workload.openshift.io/management"
annotation_prefix = "resources.workload.openshift.io"
resources = { "cpushares" = 0, "cpuset" = "" }
用 PerformanceProfile 为系统负载和业务负载分配 CPU
Workload Partitioning 负责启用负载强制隔离功能,而将负载分配到 CPU 则需用 PerformanceProfile 实现。以下用一个 SNO 单节点 OpenShift 为例说明,该 OpenShift 已经启用了 Workload Partitioning,但还未创建 PerformanceProfile。
创建 PerformanceProfile 前
- 首先查看节点 CPU 的核数,本环境的 CPU 共有 32 个内核。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- lscpu --extend
。。。
CPU NODE SOCKET CORE L1d:L1i:L2:L3 ONLINE
0 0 0 0 0:0:0:0 yes
1 0 0 1 1:1:1:0 yes
2 0 0 2 2:2:2:0 yes
3 0 0 3 3:3:3:0 yes
4 0 0 4 4:4:4:0 yes
5 0 0 5 5:5:5:0 yes
6 0 0 6 6:6:6:0 yes
7 0 0 7 7:7:7:0 yes
8 0 0 8 8:8:8:0 yes
9 0 0 9 9:9:9:0 yes
10 0 0 10 10:10:10:0 yes
11 0 0 11 11:11:11:0 yes
12 0 0 12 12:12:12:0 yes
13 0 0 13 13:13:13:0 yes
14 0 0 14 14:14:14:0 yes
15 0 0 15 15:15:15:0 yes
16 0 0 16 16:16:16:0 yes
17 0 0 17 17:17:17:0 yes
18 0 0 18 18:18:18:0 yes
19 0 0 19 19:19:19:0 yes
20 0 0 20 20:20:20:0 yes
21 0 0 21 21:21:21:0 yes
22 0 0 22 22:22:22:0 yes
23 0 0 23 23:23:23:0 yes
24 0 0 24 24:24:24:0 yes
25 0 0 25 25:25:25:0 yes
26 0 0 26 26:26:26:0 yes
27 0 0 27 27:27:27:0 yes
28 0 0 28 28:28:28:0 yes
29 0 0 29 29:29:29:0 yes
30 0 0 30 30:30:30:0 yes
31 0 0 31 31:31:31:0 yes
- 查看节点的 CPU 总容量和可分配量。
bash
$ oc describe node $(oc get nodes -o jsonpath='{.items[0].metadata.name}') | grep -E 'Capacity:|Allocatable:' -A 1
Capacity:
cpu: 32
--
Allocatable:
cpu: 31500m
- 检查系统的 kubelet 进程的 CPU 负载亲和性。确认在未使用 PerformanceProfile 的时候 kubelet 进程可以在 0-31 任意一个内核上运行。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c 'taskset -c -p $(pidof /usr/bin/kubelet)'
。。。
pid 3392's current affinity list: 0-31
- 检查系统的 etcd 进程的 CPU 负载亲和性。确认在未使用 PerformanceProfile 的时候系统 etcd 进程可以在 0-31 任意一个内核上运行。
json
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- <<'EOF'
for pid in $(pidof etcd); do
echo -n "CPU affinity (Cpuset): "
taskset -c -p "$pid"
done
EOF
。。。
CPU affinity (Cpuset): pid 3987's current affinity list: 0-31
CPU affinity (Cpuset): pid 4000's current affinity list: 0-31
- 检查运行 etcd 的 pod 中包含的所有容器可以运行的 CPU 内核,确认此时这些容器可以运行在所有内核上。
bash
$ etcd_pod=$(oc get pod -n openshift-etcd -l app=etcd -o jsonpath='{range .items[*]}{.metadata.name}')
$ for con in $(oc get pod $etcd_pod -n openshift-etcd -o jsonpath='{.spec.containers[*].name}'); do
echo -n -e "$con -->\t"
oc -n openshift-etcd exec $etcd_pod -c $con -- grep -i Cpus_allowed_list /proc/self/status
done
etcdctl --> Cpus_allowed_list: 0-31
etcd --> Cpus_allowed_list: 0-31
etcd-metrics --> Cpus_allowed_list: 0-31
etcd-readyz --> Cpus_allowed_list: 0-31
etcd-rev --> Cpus_allowed_list: 0-31
- 查看和系统级 cluster-etcd-operator 相关的进程运行在哪些内核上。注意:psr 为运行内核编号,其中有一个进程运行在 19 号内核上。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c ' ps -e -o pid,psr,pcpu,cmd | grep cluster-etcd-operator | grep -v grep'
。。。
4164 3 0.3 cluster-etcd-operator readyz --target=https://localhost:2379 --listen-port=9980 --serving-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-serving-control-plane-cluster-wpt96-1.crt --serving-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-serving-control-plane-cluster-wpt96-1.key --client-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.crt --client-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.key --client-cacert-file=/etc/kubernetes/static-pod-certs/configmaps/etcd-all-bundles/server-ca-bundle.crt --listen-cipher-suites=TLS_AES_128_GCM_SHA256,TLS_AES_256_GCM_SHA384,TLS_CHACHA20_POLY1305_SHA256,TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_ECDSA_WITH_CHACHA20_POLY1305_SHA256,TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305_SHA256
4172 2 0.0 cluster-etcd-operator rev --endpoints=https://10.10.10.10:2379 --client-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.crt --client-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.key --client-cacert-file=/etc/kubernetes/static-pod-certs/configmaps/etcd-all-bundles/server-ca-bundle.crt
17060 19 0.6 cluster-etcd-operator operator --config=/var/run/configmaps/config/config.yaml --terminate-on-files=/var/run/secrets/serving-cert/tls.crt --terminate-on-files=/var/run/secrets/serving-cert/tls.key --terminate-on-files=/var/run/secrets/etcd-client/tls.crt --terminate-on-files=/var/run/secrets/etcd-client/tls.key --terminate-on-files=/var/run/configmaps/etcd-ca/ca-bundle.crt --terminate-on-files=/var/run/configmaps/etcd-service-ca/service-ca.crt
创建 PerformanceProfile 后
- 创建 PerformanceProfile,将包含 etcd 进程的系统负载限制在 0-15 的内核上运行,用户的业务负载限制在16-31 的内核上运行。
yaml
$ oc apply -f - <<EOF
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
name: openshift-node-performance-profile
spec:
cpu:
# Set core 0-15 for OpenShift system components and the OS
reserved: 0-15
# Set core 16-31 for user workloads
isolated: 16-31
machineConfigPoolSelector:
pools.operator.machineconfiguration.openshift.io/master: ''
nodeSelector:
node-role.kubernetes.io/master: ''
numa:
topologyPolicy: "restricted"
workloadHints:
realTime: false
highPowerConsumption: false
perPodPowerManagement: false
EOF
- 查看基于 PerformanceProfile 生成的 kubeletconfig 和 Tuned 对象。
bash
$ oc get kubeletconfig performance-openshift-node-performance-profile -o jsonpath='reservedSystemCPUs={.spec.kubeletConfig.cpuManagerPolicy}{"\n"}cpuManagerPolicy={.spec.kubeletConfig.reservedSystemCPUs}{"\n"}'
cpuManagerPolicy=static
reservedSystemCPUs=0-15
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- chroot /host grep -E 'cpuManagerPolicy|reservedSystemCPUs' /etc/kubernetes/kubelet.conf
。。。
cpuManagerPolicy=static
reservedSystemCPUs=0-15
$ oc get Tuned -n openshift-cluster-node-tuning-operator
NAME VALID AGE
default True 36h
openshift-node-performance-openshift-node-performance-profile True 3h58m
- 查看当前内核启动参数。当 Workload Partitioning 生效后 isolcpus 是设为
isolcpus=managed_irq,这和只用 PerformanceProfile 的isolcpus=<cpulist>格式不一样。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /proc/cmdline | tr ' ' '\n' | grep -E 'cpu'
。。。
tuned.non_isolcpus=0000ffff
systemd.cpu_affinity=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15
isolcpus=managed_irq,16-31
- 查看当前内核启动参数,注意 Allocatable 已经变化为 16。
bash
$ oc describe node $(oc get nodes -o jsonpath='{.items[0].metadata.name}') | grep -E 'Capacity|Allocatable' -A 1
Capacity:
cpu: 32
--
Allocatable:
cpu: 16
- 再次执行以下命令,查看 /host/etc/crio/crio.conf.d/ 中的 99-runtimes.conf 和 99-workload-pinning.conf 文件内容。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- ls -al /host/etc/crio/crio.conf.d/
。。。
total 16
drwxr-xr-x. 2 root root 120 Oct 7 06:12 .
drwxr-xr-x. 4 root root 78 Oct 7 06:12 ..
-rw-r--r--. 1 root root 3254 Oct 7 06:09 00-default
-rw-r--r--. 1 root root 204 Oct 7 06:09 01-workload-pinning-default.conf
-rw-r--r--. 1 root root 625 Oct 7 06:09 99-runtimes.conf
-rw-r--r--. 1 root root 208 Oct 7 06:09 99-workload-pinning.conf
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /host/etc/crio/crio.conf.d/99-runtimes.conf
。。。
[crio.runtime]
infra_ctr_cpuset = "0-15"
。。。
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /host/etc/crio/crio.conf.d/99-workload-pinning.conf
。。。
[crio.runtime.workloads.management]
activation_annotation = "target.workload.openshift.io/management"
annotation_prefix = "resources.workload.openshift.io"
resources = { "cpushares" = 0, "cpuset" = "0-15"
负载验证
系统负载验证
- 检查当前 kubelet 进程的 CPU 负载亲和性。确认在使用 PerformanceProfile 的时候系统 kubelet 进程限制在 0-15 内核上运行。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c 'taskset -c -p $(pidof /usr/bin/kubelet)'
。。。
pid 3392's current affinity list: 0-15
- 检查当前 etcd 进程的 CPU 负载亲和性。确认在使用 PerformanceProfile 后,系统 etcd pod 可以限制在 0-15 内核上运行。注意:如果为启用 Workload Partitioning,则和 etcd 相关的 pod 还是可以运行在 0-31 内核上。
json
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- <<'EOF'
for pid in $(pidof etcd); do
echo -n "CPU affinity (Cpuset): "
taskset -c -p "$pid"
done
EOF
。。。
CPU affinity (Cpuset): pid 3987's current affinity list: 0-15
CPU affinity (Cpuset): pid 4000's current affinity list: 0-15
- 再次查看节点中运行的所有和系统级 cluster-etcd-operator 相关的进程运行在哪个 CPU 内核上。在返回结果中可以看到这些系统负载都运行在前面由 PerformanceProfile 指定的 CPU reserved 区域,即以下的 PSR 为 14/0 ,都是在 0-15 号以内的 CPU 内核。注意:如果未启用 Workload Partitioning,则和 cluster-etcd-operator 相关的 pod 还是可以运行在 0-31 内核上。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c ' ps -e -o pid,psr,pcpu,cmd | grep cluster-etcd-operator | grep -v grep'
。。。
4312 14 0.2 cluster-etcd-operator readyz --target=https://localhost:2379 --listen-port=9980 --serving-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-serving-control-plane-cluster-wpt96-1.crt --serving-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-serving-control-plane-cluster-wpt96-1.key --client-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.crt --client-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.key --client-cacert-file=/etc/kubernetes/static-pod-certs/configmaps/etcd-all-bundles/server-ca-bundle.crt --listen-cipher-suites=TLS_AES_128_GCM_SHA256,TLS_AES_256_GCM_SHA384,TLS_CHACHA20_POLY1305_SHA256,TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_ECDSA_WITH_CHACHA20_POLY1305_SHA256,TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305_SHA256
4324 0 0.0 cluster-etcd-operator rev --endpoints=https://10.10.10.10:2379 --client-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.crt --client-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.key --client-cacert-file=/etc/kubernetes/static-pod-certs/configmaps/etcd-all-bundles/server-ca-bundle.crt
16655 2 0.6 cluster-etcd-operator operator --config=/var/run/configmaps/config/config.yaml --terminate-on-files=/var/run/secrets/serving-cert/tls.crt --terminate-on-files=/var/run/secrets/serving-cert/tls.key --terminate-on-files=/var/run/secrets/etcd-client/tls.crt --terminate-on-files=/var/run/secrets/etcd-client/tls.key --terminate-on-files=/var/run/configmaps/etcd-ca/ca-bundle.crt --terminate-on-files=/var/run/configmaps/etcd-service-ca/service-ca.crt
用户负载验证
- 执行命令,部署用户负载。
yaml
$ oc new-project wp-demo
$ oc apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata:
name: cpu-stress
spec:
replicas: 1
selector:
matchLabels:
app: cpu-stress
template:
metadata:
labels:
app: cpu-stress
spec:
containers:
- name: stress
image: polinux/stress-ng:latest
command: ["stress-ng"]
args:
- --cpu
- "4"
- --timeout
- "0"
EOF
- 在 Pod 运行后查看节点消耗 CPU 最多的进程,确认是前 4 个进程是运行 stress-ng 的应用负载,而且每个进程消耗的 CPU 都接近 100%。注意:返回结果中的 PSR 字段就是运行该进程使用的 CPU 内核编号,即 24/26/27/29。
bash
$ oc get deploy -n wp-demo
NAME READY UP-TO-DATE AVAILABLE AGE
cpu-stress 1/1 1 1 40s
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- ps -e -o pid,psr,pcpu,cmd --sort=-pcpu | grep stress-ng
。。。
PID PSR %CPU CMD
11037 26 99.4 stress-ng --cpu 4 --timeout 0
11039 24 99.4 stress-ng --cpu 4 --timeout 0
11042 29 99.4 stress-ng --cpu 4 --timeout 0
11041 27 99.3 stress-ng --cpu 4 --timeout 0
11030 16 0.0 stress-ng --cpu 4 --timeout 0
- 还可进入 node 内部,然后再通过
top 1命令查看所有 CPU 内核当前消耗情况。确认在 16-31 号 CPU 内核(以下的 24/26/27/29)使用率接近 100%。
bash
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}')
sh-5.1# top 1
Tasks: 1149 total, 5 running, 1144 sleeping, 0 stopped, 0 zombie
%Cpu0 : 2.3 us, 1.0 sy, 0.0 ni, 93.3 id, 0.0 wa, 0.7 hi, 2.7 si, 0.0 st
%Cpu1 : 2.3 us, 1.7 sy, 0.0 ni, 95.0 id, 0.0 wa, 0.7 hi, 0.3 si, 0.0 st
%Cpu2 : 1.7 us, 1.3 sy, 0.0 ni, 96.0 id, 0.3 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu3 : 2.3 us, 1.7 sy, 0.0 ni, 95.0 id, 0.0 wa, 0.7 hi, 0.3 si, 0.0 st
%Cpu4 : 2.3 us, 1.0 sy, 0.0 ni, 96.0 id, 0.0 wa, 0.7 hi, 0.0 si, 0.0 st
%Cpu5 : 2.6 us, 2.0 sy, 0.0 ni, 94.7 id, 0.0 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu6 : 2.3 us, 2.0 sy, 0.0 ni, 95.0 id, 0.0 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu7 : 3.6 us, 2.3 sy, 0.0 ni, 93.0 id, 0.3 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu8 : 11.3 us, 1.3 sy, 0.0 ni, 86.4 id, 0.0 wa, 0.7 hi, 0.3 si, 0.0 st
%Cpu9 : 2.7 us, 1.7 sy, 0.0 ni, 94.7 id, 0.0 wa, 0.7 hi, 0.3 si, 0.0 st
%Cpu10 : 3.0 us, 1.3 sy, 0.0 ni, 94.7 id, 0.0 wa, 0.7 hi, 0.3 si, 0.0 st
%Cpu11 : 2.3 us, 2.0 sy, 0.0 ni, 95.0 id, 0.0 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu12 : 3.3 us, 2.0 sy, 0.0 ni, 94.0 id, 0.0 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu13 : 3.3 us, 2.0 sy, 0.0 ni, 93.7 id, 0.3 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu14 : 2.3 us, 1.7 sy, 0.0 ni, 95.0 id, 0.0 wa, 0.7 hi, 0.3 si, 0.0 st
%Cpu15 : 3.0 us, 1.7 sy, 0.0 ni, 94.7 id, 0.0 wa, 0.7 hi, 0.0 si, 0.0 st
%Cpu16 : 0.7 us, 0.7 sy, 0.0 ni, 98.7 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu17 : 1.0 us, 0.7 sy, 0.0 ni, 98.0 id, 0.0 wa, 0.3 hi, 0.0 si, 0.0 st
%Cpu18 : 1.3 us, 1.6 sy, 0.0 ni, 96.4 id, 0.0 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu19 : 1.3 us, 0.7 sy, 0.0 ni, 98.0 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu20 : 3.3 us, 2.6 sy, 0.0 ni, 93.7 id, 0.0 wa, 0.3 hi, 0.0 si, 0.0 st
%Cpu21 : 1.0 us, 0.7 sy, 0.0 ni, 98.0 id, 0.0 wa, 0.3 hi, 0.0 si, 0.0 st
%Cpu22 : 3.0 us, 1.7 sy, 0.0 ni, 94.7 id, 0.0 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu23 : 0.7 us, 0.7 sy, 0.0 ni, 98.3 id, 0.0 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu24 : 99.3 us, 0.0 sy, 0.0 ni, 0.0 id, 0.0 wa, 0.7 hi, 0.0 si, 0.0 st
%Cpu25 : 2.0 us, 1.0 sy, 0.0 ni, 96.4 id, 0.0 wa, 0.3 hi, 0.3 si, 0.0 st
%Cpu26 : 99.0 us, 0.0 sy, 0.0 ni, 0.0 id, 0.0 wa, 0.7 hi, 0.3 si, 0.0 st
%Cpu27 : 99.3 us, 0.0 sy, 0.0 ni, 0.0 id, 0.0 wa, 0.7 hi, 0.0 si, 0.0 st
%Cpu28 : 1.0 us, 1.0 sy, 0.0 ni, 97.7 id, 0.0 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu29 : 99.0 us, 0.0 sy, 0.0 ni, 0.0 id, 0.0 wa, 0.7 hi, 0.3 si, 0.0 st
%Cpu30 : 3.6 us, 2.6 sy, 0.0 ni, 93.0 id, 0.3 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu31 : 0.7 us, 0.7 sy, 0.0 ni, 98.0 id, 0.0 wa, 0.3 hi, 0.3 si, 0.0 st
MiB Mem : 64271.7 total, 27358.1 free, 15650.5 used, 22074.4 buff/cache
MiB Swap: 0.0 total, 0.0 free, 0.0 used. 48621.2 avail Mem
PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
11037 1000850+ 20 0 53096 5892 3072 R 99.7 0.0 94:23.23 stress-ng-cpu
11041 1000850+ 20 0 53096 6148 3328 R 99.7 0.0 94:21.41 stress-ng-cpu
11039 1000850+ 20 0 53096 5636 3328 R 99.3 0.0 94:21.78 stress-ng-cpu
11042 1000850+ 20 0 53096 5892 3072 R 99.3 0.0 94:22.34 stress-ng-cpu
4110 root 20 0 5208828 2.0g 90112 S 18.5 3.3 29:14.28 kube-apiserver
3860 root 20 0 5038952 403808 61440 S 15.5 0.6 18:26.41 kubelet
18050 nobody 20 0 7987252 2.2g 153048 S 10.9 3.6 10:15.70 prometheus
4284 root 1 -19 12.0g 335020 174848 S 10.6 0.5 9:26.67 etcd
参考
https://kubernetes.io/docs/tasks/administer-cluster/cpu-management-policies/
https://andreaskaris.github.io/blog/openshift/cpu-isolation-in-openshift/
https://redhatquickcourses.github.io/ocp4-workload-partition/modules/index.html