OpenShift - 实现系统负载和业务负载分区隔离

《OpenShift / RHEL / DevSecOps 汇总目录》

本文已在 OpenShift 4.22 上验证。

文章目录

  • 系统负载、业务负载
  • 系统负载和业务负载隔离的技术手段
    • [Performance Profile](#Performance Profile)
    • [Workload Partitioning](#Workload Partitioning)
  • 使用工作负载分区特性
    • [只用 PerformanceProfile 特性](#只用 PerformanceProfile 特性)
      • [创建 PerformanceProfile 前](#创建 PerformanceProfile 前)
      • [创建 PerformanceProfile 后](#创建 PerformanceProfile 后)
      • 用户负载验证
    • [使用 PerformanceProfile + Workload Partitioning 特性](#使用 PerformanceProfile + Workload Partitioning 特性)
      • [启用 Workload Partitioning 特性](#启用 Workload Partitioning 特性)
      • [用 PerformanceProfile 为系统负载和业务负载分配 CPU](#用 PerformanceProfile 为系统负载和业务负载分配 CPU)
        • [创建 PerformanceProfile 前](#创建 PerformanceProfile 前)
        • [创建 PerformanceProfile 后](#创建 PerformanceProfile 后)
      • 负载验证
  • 参考

系统负载、业务负载

一个 OpenShift 集群的管理平面用来运行容器平台的系统负载,数据平面运行用户的业务负载。

  • 系统负载:或者称为平台负载、系统 pod、平台 pod,是用来跑系统组件(kubelet、CRI、OS 内核线程、kube-proxy 等)。
  • 业务负载:或者称为用户负载、用户 pod。

当使用多个节点分别运行管理平面和数据平面的时候,系统负载和业务负载不会出现争抢节点资源的情况。但如果管理平面和数据平面在同一节点运行的时候(例如单节点或 3 节点 OpenShift 集群),系统负载和业务负载可能出现争抢节点 CPU 资源的情况。

为了解决上述问题,我们可以使用 Performance Profile + Workload Partitioning 将系统负载和业务负载隔离在不同的 CPU 内核上运行,这样即便是在高负载的时候,系统负载和业务负载之间才不会相互影响。

系统负载和业务负载隔离的技术手段

要实现系统负载和业务负载隔离,就是无论系统负载和业务负载哪一方的负载增加,都不应影响另一方正常运行和响应速度。在 OpenShift 中Performance Profile 和 Workload Partitioning 是实现系统负载和业务负载隔离的主要技术手段。

Performance Profile

Performance Profile 是 Node Tuning Operator 的一个 CRD,通过它可以设置 CPU、内存、NUMA 等的适用策略,以适配诸如低延迟、real-time、低功耗等不同的性能场景的需要。由于它是在 OpenShift 安装完后根据需要设置的,因此功能可以随时启用和关闭。

在 Performance Profile 中可以为 CPU 设置 reserved 和 isolated 的 cpuset,其中 reserved 是保留给平台负载用的 CPU,isolated 是保留给用户负载用的 CPU,并且通常 all-cpus = isolated + reserved。

yaml 复制代码
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
  name: openshift-node-performance-profile
spec:
  cpu:
    # Set core for OpenShift system components and the OS
    reserved: "0-7"
    # Set core for user workloads
    isolated: "8-15"
。。。

需要注意的是 PerformanceProfile 虽能保证"业务 pod 只能跑在 isolated 的 cpuset",但不能保证"只有业务 pod 能跑在 isolated 的 cpuset",系统 Pod / 基础设施 Pod 也有可能跑到 isolated cpuset 上。这是因为 PerformanceProfile 使用的 CPU Manager 的 static 策略只对 QoS 为 Guaranteed 且请求整数核的 Pod 生效。而非 Guaranteed(Burstable/BestEffort)的 Pod 则不受 CPU Manager 管控,它们可以在所有 CPU(reserved+isolated)之间随意调度。

Workload Partitioning

PerformanceProfile 的 reserved 可以保护宿主机系统进程,但不保证 isolated 不被基础设施 Pod 占用,而 Workload Partitioning 可以解决 PerformanceProfile 隔离不彻底的问题。Workload Partitioning 不是简单 cpuset,它可使用 cgroup v2 + cpuset + IRQ partitioning 等多种隔离手段让系统负载和业务负载实现更充分的隔离,能够将基础设施 Pod 也锁死在 reserved。

无法单独使用 Workload Partitioning,它必须配合 PerformanceProfile 一起使用才能实现系统负载和业务负载完全隔离。而 PerformanceProfile 是可以不依赖 Workload Partitioning 而独立使用,但系统负载和业务负载无法实现完全隔离。

虽然 Workload Partitioning 有更强的隔离性来隔离系统负载和业务负载,但使用它也是有一定的条件和代价的。因为当系统负载和业务负载只能运行在各自的 cpuset 中,那么就需要各自部分有更精确的 CPU 数量划分,以避免即便 CPU 整体还有较多剩余也无法调度 Pod 的情况。因此 Workload Partitioning 通常主要用在以下典型场景:SNO - 单节点 OpenShift / 3-node OpenShift 集群、Telcom/NFV 强 SLA 和审计要求。而在以下场景中无法使用 Workload Partitioning:

  • master / infra / worker 分开部署在不同的节点:平台与业务本来就不在同一台机器上,因此彼此已不会相互干扰。
  • Workload Partitioning 只能在安装时启用且不可逆,因此无法对已安装完的 OpenShift 集群启用该功能。
  • 配置不同的节点:PerformanceProfile 可按 MCP 分别定义 isolated/reserved;但 Workload Partitioning 的配置 cpuPartitioningMode: AllNodes 只能作用于所有节点,因此需要集群中的所有节点具有相同的配置。

使用工作负载分区特性

只用 PerformanceProfile 特性

如果只用 PerformanceProfile 特性,只能影响到 CPU Request 为整数且 QoS 为 Guaranteed 的 Pod。这种 Pod 被分配的 CPU 是下图 non system reserved CPUs 中的 exclusive CPUs,而其他类型的 Pod 可以使用下图中所有 shared pool 的 CPU。

bash 复制代码
+----------------------+------------------------------+
| system reserved CPUs | non system reserved CPUs     |
+----------------------+------------------------------+
| shared pool          | exclusive CPUs | shared pool |
+----------------------+------------------------------+
                         ------------->

创建 PerformanceProfile 前

  1. 首先查看节点 CPU 的核数,本环境的 CPU 共有 32 个内核。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- lscpu --extend
。。。
CPU NODE SOCKET CORE L1d:L1i:L2:L3 ONLINE
  0    0      0    0 0:0:0:0          yes
  1    0      0    1 1:1:1:0          yes
  2    0      0    2 2:2:2:0          yes
  3    0      0    3 3:3:3:0          yes
  4    0      0    4 4:4:4:0          yes
  5    0      0    5 5:5:5:0          yes
  6    0      0    6 6:6:6:0          yes
  7    0      0    7 7:7:7:0          yes
  8    0      0    8 8:8:8:0          yes
  9    0      0    9 9:9:9:0          yes
 10    0      0   10 10:10:10:0       yes
 11    0      0   11 11:11:11:0       yes
 12    0      0   12 12:12:12:0       yes
 13    0      0   13 13:13:13:0       yes
 14    0      0   14 14:14:14:0       yes
 15    0      0   15 15:15:15:0       yes
 16    0      0   16 16:16:16:0       yes
 17    0      0   17 17:17:17:0       yes
 18    0      0   18 18:18:18:0       yes
 19    0      0   19 19:19:19:0       yes
 20    0      0   20 20:20:20:0       yes
 21    0      0   21 21:21:21:0       yes
 22    0      0   22 22:22:22:0       yes
 23    0      0   23 23:23:23:0       yes
 24    0      0   24 24:24:24:0       yes
 25    0      0   25 25:25:25:0       yes
 26    0      0   26 26:26:26:0       yes
 27    0      0   27 27:27:27:0       yes
 28    0      0   28 28:28:28:0       yes
 29    0      0   29 29:29:29:0       yes
 30    0      0   30 30:30:30:0       yes
 31    0      0   31 31:31:31:0       yes
  1. 查看节点的 CPU 总容量和可分配量。
bash 复制代码
$ oc describe node $(oc get nodes -o jsonpath='{.items[0].metadata.name}')  | grep -E 'Capacity:|Allocatable:' -A 1
Capacity:
  cpu:                            32
--
Allocatable:
  cpu:                            31500m
  1. 检查系统的 kubelet 进程的 CPU 负载亲和性。确认在未使用 PerformanceProfile 的时候 kubelet 进程可以在 0-31 任意一个内核上运行。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c 'taskset -c -p $(pidof /usr/bin/kubelet)'
。。。
pid 3392's current affinity list: 0-31

创建 PerformanceProfile 后

  1. 创建 PerformanceProfile,它将 CPU 内核分为 2 部分,0-7 为 system reserved CPUs,8-31 为 non system reserved CPUs。
yaml 复制代码
$ oc apply -f - <<EOF
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
  name: openshift-node-performance-profile
spec:
  cpu:
    reserved: 0-7
    isolated: 8-31
  machineConfigPoolSelector:
    pools.operator.machineconfiguration.openshift.io/master: ''
  nodeSelector:
    node-role.kubernetes.io/master: ''
  numa:
    topologyPolicy: "restricted"
  workloadHints:
    realTime: false
    highPowerConsumption: false
    perPodPowerManagement: false
EOF
  1. 查看基于 PerformanceProfile 生成的 kubeletconfig 和 Tuned 对象。
bash 复制代码
$ oc get kubeletconfig performance-openshift-node-performance-profile -o jsonpath='reservedSystemCPUs={.spec.kubeletConfig.cpuManagerPolicy}{"\n"}cpuManagerPolicy={.spec.kubeletConfig.reservedSystemCPUs}{"\n"}'
cpuManagerPolicy=static
reservedSystemCPUs=0-7
 
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- chroot /host grep -E 'cpuManagerPolicy|reservedSystemCPUs' /etc/kubernetes/kubelet.conf
。。。
cpuManagerPolicy=static
reservedSystemCPUs=0-7
 
$ oc get Tuned openshift-node-performance-openshift-node-performance-profile -n openshift-cluster-node-tuning-operator 
NAME                                                            VALID   AGE
openshift-node-performance-openshift-node-performance-profile   True    3h58m
  1. 查看 kubelet 进程可使用节点 CPU 的范围。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c 'taskset -c -p $(pidof /usr/bin/kubelet)'
。。。
pid 4426's current affinity list: 0-7
  1. 查看当前内核启动命令,注意 systemd.cpu_affinity 和isolcpus=<cpulist>。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /proc/cmdline | tr ' ' '\n' | grep -E 'cpu'
。。。
tuned.non_isolcpus=000000ff
systemd.cpu_affinity=0,1,2,3,4,5,6,7
isolcpus=managed_irq,8-31
  1. 查看当前内核启动参数,注意 Allocatable 已经变化为 24。
bash 复制代码
$ oc describe node $(oc get nodes -o jsonpath='{.items[0].metadata.name}')  | grep -E 'Capacity|Allocatable' -A 1
Capacity:
  cpu:                                     32
--
Allocatable:
  cpu:                                     24
  1. 执行命令,查看 /host/etc/crio/crio.conf.d/ 中的 99-runtimes.conf 文件内容。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- ls -al /host/etc/crio/crio.conf.d/
。。。
total 16
drwxr-xr-x. 2 root root  120 Oct  7 06:12 .
drwxr-xr-x. 4 root root   78 Oct  7 06:12 ..
-rw-r--r--. 1 root root 3254 Oct  7 06:09 00-default
-rw-r--r--. 1 root root  625 Oct  7 06:09 99-runtimes.conf

$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /host/etc/crio/crio.conf.d/99-runtimes.conf
。。。
[crio.runtime]
infra_ctr_cpuset = "0-7"
。。。

$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}')  -- chroot /host cat /var/lib/kubelet/cpu_manager_state | jq
。。。
{
  "policyName": "static",
  "defaultCpuSet": "0-7,9-31",
  "entries": {
    "d0ccbdd0-95e6-4928-b0ea-551c7cd4a8dd": {
      "fedora-minimal-exclusive": "8"
    }
  },
  "checksum": 3215225415
}

用户负载验证

  1. 创建测试 Pod。其中名为 fedora-minimal-exclusive 的 Pod 的 QoS 为 Guaranteed,且申请的 CPU 为 1;而名为 fedora-minimal-shared 的 Pod 的 QoS 虽也为 Guaranteed,但申请的 CPU 不是整数。
yaml 复制代码
$ oc new-project pp-demo
$ oc apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata:
  labels:
    app: fedora-minimal
  name: fedora-minimal
spec:
  replicas: 1
  selector:
    matchLabels:
      app: fedora-minimal
  template:
    metadata:
      labels:
        app: fedora-minimal
    spec:
      containers:
      - image: registry.fedoraproject.org/fedora-minimal:latest
        name: fedora-minimal-exclusive
        command:
          - /bin/sleep
          - "12345"
        resources:
          requests:
            memory: "128Mi"
            cpu: "1"
          limits:
            memory: "128Mi"
            cpu: "1"
      - image: registry.fedoraproject.org/fedora-minimal:latest
        name: fedora-minimal-shared
        command:
          - /bin/sleep
          - "54321"
        resources:
          requests:
            memory: "128Mi"
            cpu: "200m"
          limits:
            memory: "128Mi"
            cpu: "200m"
EOF
  1. 执行以下命令,分别查看 2 个容器可以使用的 CPU。确认 fedora-minimal-exclusive 可用的 CPU 为 8 号,而 fedora-minimal-shared 可用的 CPU 为 0-7,9-31。
bash 复制代码
$ oc -n pp-demo exec deploy/fedora-minimal -c fedora-minimal-exclusive -- grep -i Cpus_allowed_list /proc/self/status
Cpus_allowed_list:      8

$ oc -n pp-demo exec deploy/fedora-minimal -c fedora-minimal-shared -- grep -i Cpus_allowed_list /proc/self/status
Cpus_allowed_list:      0-7,9-31

使用 PerformanceProfile + Workload Partitioning 特性

为了解决 PerformanceProfile 分离工作负载的功能较弱的问题,我们可以结合 OpenShift 独有的 Workload Partitioning 特性实现功能更强的工作负载隔离。

启用 Workload Partitioning 特性

OpenShift 工作负载分区特性只能在安装集群时就启用,无法在安装完成后启用。以下是一个 3 节点 OpenShift 的 install-config.yaml 示例,其中 cpuPartitioningMode: AllNodes 的声明启用了 Workload Partitioning 功能。

yaml 复制代码
apiVersion: v1
baseDomain: devcluster.openshift.com
cpuPartitioningMode: AllNodes
compute:
  - architecture: amd64
    name: worker
    replicas: 0
controlPlane:
  architecture: amd64
  name: master
  replicas: 3
。。。

先查看在启用 Workload Partitioning 的 OpenShift 中会有 2 个和 partitioning 相关的 machineconfig,然后再查看其中 spec.config.storage.files.path 位置的配置。

json 复制代码
$ oc get machineconfig | grep partition
01-master-cpu-partitioning                                                                    3.2.0             11h
01-worker-cpu-partitioning                                                                    3.2.0             11h

$ oc get machineconfig 01-master-cpu-partitioning -o jsonpath={.spec.config.storage.files} | jq
[
  {
    "contents": {
      "source": "data:text/plain;charset=utf-8;base64,CnsKICAibWFuYWdlbWVudCI6IHsKICAgICJjcHVzZXQiOiAiIgogIH0KfQo=",
      "verification": {}
    },
    "group": {},
    "mode": 420,
    "path": "/etc/kubernetes/openshift-workload-pinning",
    "user": {}
  },
  {
    "contents": {
      "source": "data:text/plain;charset=utf-8;base64,CltjcmlvLnJ1bnRpbWUud29ya2xvYWRzLm1hbmFnZW1lbnRdCmFjdGl2YXRpb25fYW5ub3RhdGlvbiA9ICJ0YXJnZXQud29ya2xvYWQub3BlbnNoaWZ0LmlvL21hbmFnZW1lbnQiCmFubm90YXRpb25fcHJlZml4ID0gInJlc291cmNlcy53b3JrbG9hZC5vcGVuc2hpZnQuaW8iCnJlc291cmNlcyA9IHsgImNwdXNoYXJlcyIgPSAwLCAiY3B1c2V0IiA9ICIiIH0K",
      "verification": {}
    },
    "group": {},
    "mode": 420,
    "path": "/etc/crio/crio.conf.d/01-workload-pinning-default.conf",
    "user": {}
  }
]

可以根据上一步的结果,查看节点 /etc/kubernetes 和 /host/etc/crio/crio.conf.d/ 目录中包含哪些文件,其中如果未启用 Workload Partitioning,则不会有 openshift-workload-pinning 和 01-workload-pinning-default.conf 文件。

bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- chroot /host ls -al /etc/kubernetes/
。。。
drwxr-xr-x.   7 root root  4096 Sep 28 11:19 .
drwxr-xr-x. 100 root root  8192 Sep 28 10:14 ..
-rw-r--r--.   1 root root   102 Sep 28 10:08 apiserver-url.env
drwxr-xr-x.   2 root root     6 Sep 28 10:13 cloud
-rw-r--r--.   1 root root     0 Sep 28 10:08 cloud.conf
drwxr-xr-x.   3 root root    19 Sep 28 10:10 cni
-rw-r--r--.   1 root root   174 Sep 28 10:08 crio-metrics-proxy.cfg
-rw-------.   1 root root 10984 Sep 27 23:29 kubeconfig
-rw-r--r--.   1 root root  8332 Sep 28 11:19 kubelet-ca.crt
drwxr-xr-x.   3 root root    20 Sep 28 10:13 kubelet-plugins
-rw-r--r--.   1 root root  4119 Sep 28 10:08 kubelet.conf
drwxr-xr-x.   2 root root   158 Sep 28 10:13 manifests
-rw-r--r--.   1 root root    44 Sep 28 10:08 openshift-workload-pinning
drwxr-xr-x.  23 root root  4096 Sep 28 10:13 static-pod-resources

$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- ls -al /host/etc/crio/crio.conf.d/
。。。
total 8
drwxr-xr-x. 2 root root   64 Sep 28 10:13 .
drwxr-xr-x. 4 root root   78 Sep 28 10:13 ..
-rw-r--r--. 1 root root 3254 Sep 28 10:08 00-default
-rw-r--r--. 1 root root  204 Sep 28 10:08 01-workload-pinning-default.conf

最后再查看 openshift-workload-pinning 和 01-workload-pinning-default.conf 文件的内容,确认其中的 "cpuset": "" 和 resources = { "cpushares" = 0, "cpuset" = "" }。

bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- chroot /host cat /etc/kubernetes/openshift-workload-pinning
。。。
{
  "management": {
    "cpuset": ""
  }
}

$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /host/etc/crio/crio.conf.d/01-workload-pinning-default.conf
。。。
[crio.runtime.workloads.management]
activation_annotation = "target.workload.openshift.io/management"
annotation_prefix = "resources.workload.openshift.io"
resources = { "cpushares" = 0, "cpuset" = "" }

用 PerformanceProfile 为系统负载和业务负载分配 CPU

Workload Partitioning 负责启用负载强制隔离功能,而将负载分配到 CPU 则需用 PerformanceProfile 实现。以下用一个 SNO 单节点 OpenShift 为例说明,该 OpenShift 已经启用了 Workload Partitioning,但还未创建 PerformanceProfile。

创建 PerformanceProfile 前
  1. 首先查看节点 CPU 的核数,本环境的 CPU 共有 32 个内核。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- lscpu --extend
。。。
CPU NODE SOCKET CORE L1d:L1i:L2:L3 ONLINE
  0    0      0    0 0:0:0:0          yes
  1    0      0    1 1:1:1:0          yes
  2    0      0    2 2:2:2:0          yes
  3    0      0    3 3:3:3:0          yes
  4    0      0    4 4:4:4:0          yes
  5    0      0    5 5:5:5:0          yes
  6    0      0    6 6:6:6:0          yes
  7    0      0    7 7:7:7:0          yes
  8    0      0    8 8:8:8:0          yes
  9    0      0    9 9:9:9:0          yes
 10    0      0   10 10:10:10:0       yes
 11    0      0   11 11:11:11:0       yes
 12    0      0   12 12:12:12:0       yes
 13    0      0   13 13:13:13:0       yes
 14    0      0   14 14:14:14:0       yes
 15    0      0   15 15:15:15:0       yes
 16    0      0   16 16:16:16:0       yes
 17    0      0   17 17:17:17:0       yes
 18    0      0   18 18:18:18:0       yes
 19    0      0   19 19:19:19:0       yes
 20    0      0   20 20:20:20:0       yes
 21    0      0   21 21:21:21:0       yes
 22    0      0   22 22:22:22:0       yes
 23    0      0   23 23:23:23:0       yes
 24    0      0   24 24:24:24:0       yes
 25    0      0   25 25:25:25:0       yes
 26    0      0   26 26:26:26:0       yes
 27    0      0   27 27:27:27:0       yes
 28    0      0   28 28:28:28:0       yes
 29    0      0   29 29:29:29:0       yes
 30    0      0   30 30:30:30:0       yes
 31    0      0   31 31:31:31:0       yes
  1. 查看节点的 CPU 总容量和可分配量。
bash 复制代码
$ oc describe node $(oc get nodes -o jsonpath='{.items[0].metadata.name}')  | grep -E 'Capacity:|Allocatable:' -A 1
Capacity:
  cpu:                            32
--
Allocatable:
  cpu:                            31500m
  1. 检查系统的 kubelet 进程的 CPU 负载亲和性。确认在未使用 PerformanceProfile 的时候 kubelet 进程可以在 0-31 任意一个内核上运行。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c 'taskset -c -p $(pidof /usr/bin/kubelet)'
。。。
pid 3392's current affinity list: 0-31
  1. 检查系统的 etcd 进程的 CPU 负载亲和性。确认在未使用 PerformanceProfile 的时候系统 etcd 进程可以在 0-31 任意一个内核上运行。
json 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- <<'EOF'
	for pid in $(pidof etcd); do
	    echo -n "CPU affinity (Cpuset): "
	    taskset -c -p "$pid"
	done
EOF
。。。
CPU affinity (Cpuset): pid 3987's current affinity list: 0-31
CPU affinity (Cpuset): pid 4000's current affinity list: 0-31
  1. 检查运行 etcd 的 pod 中包含的所有容器可以运行的 CPU 内核,确认此时这些容器可以运行在所有内核上。
bash 复制代码
$ etcd_pod=$(oc get pod -n openshift-etcd -l app=etcd -o jsonpath='{range .items[*]}{.metadata.name}')
$ for con in $(oc get pod $etcd_pod -n openshift-etcd -o jsonpath='{.spec.containers[*].name}'); do
  echo -n -e "$con -->\t" 
  oc -n openshift-etcd exec $etcd_pod -c $con -- grep -i Cpus_allowed_list /proc/self/status
done
etcdctl -->     Cpus_allowed_list:      0-31
etcd -->        Cpus_allowed_list:      0-31
etcd-metrics -->        Cpus_allowed_list:      0-31
etcd-readyz --> Cpus_allowed_list:      0-31
etcd-rev -->    Cpus_allowed_list:      0-31
  1. 查看和系统级 cluster-etcd-operator 相关的进程运行在哪些内核上。注意:psr 为运行内核编号,其中有一个进程运行在 19 号内核上。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c ' ps -e -o pid,psr,pcpu,cmd | grep cluster-etcd-operator | grep -v grep'
。。。
   4164   3  0.3 cluster-etcd-operator readyz --target=https://localhost:2379 --listen-port=9980 --serving-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-serving-control-plane-cluster-wpt96-1.crt --serving-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-serving-control-plane-cluster-wpt96-1.key --client-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.crt --client-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.key --client-cacert-file=/etc/kubernetes/static-pod-certs/configmaps/etcd-all-bundles/server-ca-bundle.crt --listen-cipher-suites=TLS_AES_128_GCM_SHA256,TLS_AES_256_GCM_SHA384,TLS_CHACHA20_POLY1305_SHA256,TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_ECDSA_WITH_CHACHA20_POLY1305_SHA256,TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305_SHA256
   4172   2  0.0 cluster-etcd-operator rev --endpoints=https://10.10.10.10:2379 --client-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.crt --client-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.key --client-cacert-file=/etc/kubernetes/static-pod-certs/configmaps/etcd-all-bundles/server-ca-bundle.crt
  17060  19  0.6 cluster-etcd-operator operator --config=/var/run/configmaps/config/config.yaml --terminate-on-files=/var/run/secrets/serving-cert/tls.crt --terminate-on-files=/var/run/secrets/serving-cert/tls.key --terminate-on-files=/var/run/secrets/etcd-client/tls.crt --terminate-on-files=/var/run/secrets/etcd-client/tls.key --terminate-on-files=/var/run/configmaps/etcd-ca/ca-bundle.crt --terminate-on-files=/var/run/configmaps/etcd-service-ca/service-ca.crt
创建 PerformanceProfile 后
  1. 创建 PerformanceProfile,将包含 etcd 进程的系统负载限制在 0-15 的内核上运行,用户的业务负载限制在16-31 的内核上运行。
yaml 复制代码
$ oc apply -f - <<EOF
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
  name: openshift-node-performance-profile
spec:
  cpu:
    # Set core 0-15 for OpenShift system components and the OS
    reserved: 0-15
    # Set core 16-31 for user workloads
    isolated: 16-31
  machineConfigPoolSelector:
    pools.operator.machineconfiguration.openshift.io/master: ''
  nodeSelector:
    node-role.kubernetes.io/master: ''
  numa:
    topologyPolicy: "restricted"
  workloadHints:
    realTime: false
    highPowerConsumption: false
    perPodPowerManagement: false
EOF
  1. 查看基于 PerformanceProfile 生成的 kubeletconfig 和 Tuned 对象。
bash 复制代码
$ oc get kubeletconfig performance-openshift-node-performance-profile -o jsonpath='reservedSystemCPUs={.spec.kubeletConfig.cpuManagerPolicy}{"\n"}cpuManagerPolicy={.spec.kubeletConfig.reservedSystemCPUs}{"\n"}'
cpuManagerPolicy=static
reservedSystemCPUs=0-15

$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- chroot /host grep -E 'cpuManagerPolicy|reservedSystemCPUs' /etc/kubernetes/kubelet.conf
。。。
cpuManagerPolicy=static
reservedSystemCPUs=0-15
 
$ oc get Tuned -n openshift-cluster-node-tuning-operator 
NAME                                                            VALID   AGE
default                                                         True    36h
openshift-node-performance-openshift-node-performance-profile   True    3h58m
  1. 查看当前内核启动参数。当 Workload Partitioning 生效后 isolcpus 是设为isolcpus=managed_irq,这和只用 PerformanceProfile 的isolcpus=<cpulist>格式不一样。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /proc/cmdline | tr ' ' '\n' | grep -E 'cpu'
。。。
tuned.non_isolcpus=0000ffff
systemd.cpu_affinity=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15
isolcpus=managed_irq,16-31
  1. 查看当前内核启动参数,注意 Allocatable 已经变化为 16。
bash 复制代码
$ oc describe node $(oc get nodes -o jsonpath='{.items[0].metadata.name}')  | grep -E 'Capacity|Allocatable' -A 1
Capacity:
  cpu:                                     32
--
Allocatable:
  cpu:                                     16
  1. 再次执行以下命令,查看 /host/etc/crio/crio.conf.d/ 中的 99-runtimes.conf 和 99-workload-pinning.conf 文件内容。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- ls -al /host/etc/crio/crio.conf.d/
。。。
total 16
drwxr-xr-x. 2 root root  120 Oct  7 06:12 .
drwxr-xr-x. 4 root root   78 Oct  7 06:12 ..
-rw-r--r--. 1 root root 3254 Oct  7 06:09 00-default
-rw-r--r--. 1 root root  204 Oct  7 06:09 01-workload-pinning-default.conf
-rw-r--r--. 1 root root  625 Oct  7 06:09 99-runtimes.conf
-rw-r--r--. 1 root root  208 Oct  7 06:09 99-workload-pinning.conf

$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /host/etc/crio/crio.conf.d/99-runtimes.conf
。。。
[crio.runtime]
infra_ctr_cpuset = "0-15"
。。。

$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- cat /host/etc/crio/crio.conf.d/99-workload-pinning.conf
。。。
[crio.runtime.workloads.management]
activation_annotation = "target.workload.openshift.io/management"
annotation_prefix = "resources.workload.openshift.io"
resources = { "cpushares" = 0, "cpuset" = "0-15" 

负载验证

系统负载验证
  1. 检查当前 kubelet 进程的 CPU 负载亲和性。确认在使用 PerformanceProfile 的时候系统 kubelet 进程限制在 0-15 内核上运行。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c 'taskset -c -p $(pidof /usr/bin/kubelet)'
。。。
pid 3392's current affinity list: 0-15
  1. 检查当前 etcd 进程的 CPU 负载亲和性。确认在使用 PerformanceProfile 后,系统 etcd pod 可以限制在 0-15 内核上运行。注意:如果为启用 Workload Partitioning,则和 etcd 相关的 pod 还是可以运行在 0-31 内核上。
json 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- <<'EOF'
	for pid in $(pidof etcd); do
	    echo -n "CPU affinity (Cpuset): "
	    taskset -c -p "$pid"
	done
EOF
。。。
CPU affinity (Cpuset): pid 3987's current affinity list: 0-15
CPU affinity (Cpuset): pid 4000's current affinity list: 0-15
  1. 再次查看节点中运行的所有和系统级 cluster-etcd-operator 相关的进程运行在哪个 CPU 内核上。在返回结果中可以看到这些系统负载都运行在前面由 PerformanceProfile 指定的 CPU reserved 区域,即以下的 PSR 为 14/0 ,都是在 0-15 号以内的 CPU 内核。注意:如果未启用 Workload Partitioning,则和 cluster-etcd-operator 相关的 pod 还是可以运行在 0-31 内核上。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- bash -c ' ps -e -o pid,psr,pcpu,cmd | grep cluster-etcd-operator | grep -v grep'
。。。
   4312  14  0.2 cluster-etcd-operator readyz --target=https://localhost:2379 --listen-port=9980 --serving-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-serving-control-plane-cluster-wpt96-1.crt --serving-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-serving-control-plane-cluster-wpt96-1.key --client-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.crt --client-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.key --client-cacert-file=/etc/kubernetes/static-pod-certs/configmaps/etcd-all-bundles/server-ca-bundle.crt --listen-cipher-suites=TLS_AES_128_GCM_SHA256,TLS_AES_256_GCM_SHA384,TLS_CHACHA20_POLY1305_SHA256,TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_ECDSA_WITH_CHACHA20_POLY1305_SHA256,TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305_SHA256
   4324   0  0.0 cluster-etcd-operator rev --endpoints=https://10.10.10.10:2379 --client-cert-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.crt --client-key-file=/etc/kubernetes/static-pod-certs/secrets/etcd-all-certs/etcd-peer-control-plane-cluster-wpt96-1.key --client-cacert-file=/etc/kubernetes/static-pod-certs/configmaps/etcd-all-bundles/server-ca-bundle.crt
  16655   2  0.6 cluster-etcd-operator operator --config=/var/run/configmaps/config/config.yaml --terminate-on-files=/var/run/secrets/serving-cert/tls.crt --terminate-on-files=/var/run/secrets/serving-cert/tls.key --terminate-on-files=/var/run/secrets/etcd-client/tls.crt --terminate-on-files=/var/run/secrets/etcd-client/tls.key --terminate-on-files=/var/run/configmaps/etcd-ca/ca-bundle.crt --terminate-on-files=/var/run/configmaps/etcd-service-ca/service-ca.crt
用户负载验证
  1. 执行命令,部署用户负载。
yaml 复制代码
$ oc new-project wp-demo
$ oc apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata:
  name: cpu-stress
spec:
  replicas: 1
  selector:
    matchLabels:
      app: cpu-stress
  template:
    metadata:
      labels:
        app: cpu-stress
    spec:
      containers:
      - name: stress
        image: polinux/stress-ng:latest
        command: ["stress-ng"]
        args:
        - --cpu
        - "4"
        - --timeout
        - "0"
EOF
  1. 在 Pod 运行后查看节点消耗 CPU 最多的进程,确认是前 4 个进程是运行 stress-ng 的应用负载,而且每个进程消耗的 CPU 都接近 100%。注意:返回结果中的 PSR 字段就是运行该进程使用的 CPU 内核编号,即 24/26/27/29。
bash 复制代码
$ oc get deploy -n wp-demo
NAME         READY   UP-TO-DATE   AVAILABLE   AGE
cpu-stress   1/1     1            1           40s

$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}') -- ps -e -o pid,psr,pcpu,cmd --sort=-pcpu | grep stress-ng
。。。
    PID PSR %CPU CMD
  11037  26 99.4 stress-ng --cpu 4 --timeout 0
  11039  24 99.4 stress-ng --cpu 4 --timeout 0
  11042  29 99.4 stress-ng --cpu 4 --timeout 0
  11041  27 99.3 stress-ng --cpu 4 --timeout 0
  11030  16  0.0 stress-ng --cpu 4 --timeout 0
  
  1. 还可进入 node 内部,然后再通过 top 1 命令查看所有 CPU 内核当前消耗情况。确认在 16-31 号 CPU 内核(以下的 24/26/27/29)使用率接近 100%。
bash 复制代码
$ oc debug node/$(oc get nodes -o jsonpath='{.items[0].metadata.name}')
sh-5.1# top 1
Tasks: 1149 total,   5 running, 1144 sleeping,   0 stopped,   0 zombie
%Cpu0  :  2.3 us,  1.0 sy,  0.0 ni, 93.3 id,  0.0 wa,  0.7 hi,  2.7 si,  0.0 st
%Cpu1  :  2.3 us,  1.7 sy,  0.0 ni, 95.0 id,  0.0 wa,  0.7 hi,  0.3 si,  0.0 st
%Cpu2  :  1.7 us,  1.3 sy,  0.0 ni, 96.0 id,  0.3 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu3  :  2.3 us,  1.7 sy,  0.0 ni, 95.0 id,  0.0 wa,  0.7 hi,  0.3 si,  0.0 st
%Cpu4  :  2.3 us,  1.0 sy,  0.0 ni, 96.0 id,  0.0 wa,  0.7 hi,  0.0 si,  0.0 st
%Cpu5  :  2.6 us,  2.0 sy,  0.0 ni, 94.7 id,  0.0 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu6  :  2.3 us,  2.0 sy,  0.0 ni, 95.0 id,  0.0 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu7  :  3.6 us,  2.3 sy,  0.0 ni, 93.0 id,  0.3 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu8  : 11.3 us,  1.3 sy,  0.0 ni, 86.4 id,  0.0 wa,  0.7 hi,  0.3 si,  0.0 st
%Cpu9  :  2.7 us,  1.7 sy,  0.0 ni, 94.7 id,  0.0 wa,  0.7 hi,  0.3 si,  0.0 st
%Cpu10 :  3.0 us,  1.3 sy,  0.0 ni, 94.7 id,  0.0 wa,  0.7 hi,  0.3 si,  0.0 st
%Cpu11 :  2.3 us,  2.0 sy,  0.0 ni, 95.0 id,  0.0 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu12 :  3.3 us,  2.0 sy,  0.0 ni, 94.0 id,  0.0 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu13 :  3.3 us,  2.0 sy,  0.0 ni, 93.7 id,  0.3 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu14 :  2.3 us,  1.7 sy,  0.0 ni, 95.0 id,  0.0 wa,  0.7 hi,  0.3 si,  0.0 st
%Cpu15 :  3.0 us,  1.7 sy,  0.0 ni, 94.7 id,  0.0 wa,  0.7 hi,  0.0 si,  0.0 st
%Cpu16 :  0.7 us,  0.7 sy,  0.0 ni, 98.7 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
%Cpu17 :  1.0 us,  0.7 sy,  0.0 ni, 98.0 id,  0.0 wa,  0.3 hi,  0.0 si,  0.0 st
%Cpu18 :  1.3 us,  1.6 sy,  0.0 ni, 96.4 id,  0.0 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu19 :  1.3 us,  0.7 sy,  0.0 ni, 98.0 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
%Cpu20 :  3.3 us,  2.6 sy,  0.0 ni, 93.7 id,  0.0 wa,  0.3 hi,  0.0 si,  0.0 st
%Cpu21 :  1.0 us,  0.7 sy,  0.0 ni, 98.0 id,  0.0 wa,  0.3 hi,  0.0 si,  0.0 st
%Cpu22 :  3.0 us,  1.7 sy,  0.0 ni, 94.7 id,  0.0 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu23 :  0.7 us,  0.7 sy,  0.0 ni, 98.3 id,  0.0 wa,  0.0 hi,  0.3 si,  0.0 st
%Cpu24 : 99.3 us,  0.0 sy,  0.0 ni,  0.0 id,  0.0 wa,  0.7 hi,  0.0 si,  0.0 st
%Cpu25 :  2.0 us,  1.0 sy,  0.0 ni, 96.4 id,  0.0 wa,  0.3 hi,  0.3 si,  0.0 st
%Cpu26 : 99.0 us,  0.0 sy,  0.0 ni,  0.0 id,  0.0 wa,  0.7 hi,  0.3 si,  0.0 st
%Cpu27 : 99.3 us,  0.0 sy,  0.0 ni,  0.0 id,  0.0 wa,  0.7 hi,  0.0 si,  0.0 st
%Cpu28 :  1.0 us,  1.0 sy,  0.0 ni, 97.7 id,  0.0 wa,  0.0 hi,  0.3 si,  0.0 st
%Cpu29 : 99.0 us,  0.0 sy,  0.0 ni,  0.0 id,  0.0 wa,  0.7 hi,  0.3 si,  0.0 st
%Cpu30 :  3.6 us,  2.6 sy,  0.0 ni, 93.0 id,  0.3 wa,  0.0 hi,  0.3 si,  0.0 st
%Cpu31 :  0.7 us,  0.7 sy,  0.0 ni, 98.0 id,  0.0 wa,  0.3 hi,  0.3 si,  0.0 st
MiB Mem :  64271.7 total,  27358.1 free,  15650.5 used,  22074.4 buff/cache
MiB Swap:      0.0 total,      0.0 free,      0.0 used.  48621.2 avail Mem 

    PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND              
  11037 1000850+  20   0   53096   5892   3072 R  99.7   0.0  94:23.23 stress-ng-cpu        
  11041 1000850+  20   0   53096   6148   3328 R  99.7   0.0  94:21.41 stress-ng-cpu        
  11039 1000850+  20   0   53096   5636   3328 R  99.3   0.0  94:21.78 stress-ng-cpu        
  11042 1000850+  20   0   53096   5892   3072 R  99.3   0.0  94:22.34 stress-ng-cpu        
   4110 root      20   0 5208828   2.0g  90112 S  18.5   3.3  29:14.28 kube-apiserver       
   3860 root      20   0 5038952 403808  61440 S  15.5   0.6  18:26.41 kubelet              
  18050 nobody    20   0 7987252   2.2g 153048 S  10.9   3.6  10:15.70 prometheus           
   4284 root       1 -19   12.0g 335020 174848 S  10.6   0.5   9:26.67 etcd   

参考

https://github.com/openshift/cluster-node-tuning-operator/blob/main/docs/performanceprofile/performance_profile.md

https://kubernetes.io/docs/tasks/administer-cluster/cpu-management-policies/

https://andreaskaris.github.io/blog/openshift/cpu-isolation-in-openshift/

https://www.redhat.com/en/blog/your-guide-to-workload-partitioning-for-multi-node-clusters-in-openshift-4.13

https://redhatquickcourses.github.io/ocp4-workload-partition/modules/index.html

https://docs.redhat.com/pt-br/learn/learning-paths/how-use-workload-partitioning-red-hat-openshift-container-platform/enabling-workload-partitioning-red-hat-openshift-container-platform

https://docs.redhat.com/en/documentation/openshift_container_platform/4.20/html/scalability_and_performance/enabling-workload-partitioning

https://danielchg.github.io/posts/troubleshooting-isolcpus/

相关推荐
杀不死的坏蛋c4 小时前
k8s-原理-网络-安装
kubernetes
安易算力8 小时前
GPU集群调度实践:Slurm/K8s混合部署与GPU共享优化 —— 从批处理到在线推理的统一调度架构
容器·架构·kubernetes
MrSYJ1 天前
Docker端口映射咋做的,模拟下。
docker·云原生·kubernetes
张小凡vip1 天前
Kubernetes--k8s---了解和使用configmap挂载配置文件
云原生·容器·kubernetes
小小龙学IT2 天前
Kubernetes 深度实践:架构、API 与避坑指南
容器·架构·kubernetes
MrSYJ2 天前
别人再问你路由表是啥,这篇文章摔它脸上
docker·云原生·kubernetes
吴声子夜歌2 天前
Nginx应用与运维——Nginx在Kubernetes中的应用(一)
运维·nginx·kubernetes
吴声子夜歌2 天前
Nginx应用与运维——Nginx在Kubernetes中的应用(三)
运维·nginx·kubernetes
sbjdhjd2 天前
云安全 | Docker 容器逃逸复盘(一):从隔离边界到运行时链路,如何确认自己身处容器
网络安全·docker·云原生·容器·kubernetes·云计算·云安全