Linux 6.6内核 CPU 深度解析(五):CPU 空闲状态机 cpuidle

〇、全景:CPU 没活干时,不是"傻转",而是钻进一层层更深的省电状态

上一系列讲的是"怎么把 CPU 拉起来"(BSP/AP/hotplug)和"CPU 之间怎么喊话"(IPI)。但 CPU 大部分时间其实没活干 ------等待磁盘、等待网络、等待用户敲键盘。这时候如果让 CPU 空转(while(1)),功耗白白烧掉。

所以 x86 设计了一套 C-state(CPU 空闲状态) :CPU 空闲时进入越来越深的省电状态,深度越深越省电,但"醒过来"也越慢。而 Linux 用 cpuidle 框架 来管理这套状态------核心问题是:这一轮空闲会持续多久?该钻进多深的 C-state?
#mermaid-svg-6hQEO452j9Hw4gs2{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;fill:#333;}@keyframes edge-animation-frame{from{stroke-dashoffset:0;}}@keyframes dash{to{stroke-dashoffset:0;}}#mermaid-svg-6hQEO452j9Hw4gs2 .edge-animation-slow{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 50s linear infinite;stroke-linecap:round;}#mermaid-svg-6hQEO452j9Hw4gs2 .edge-animation-fast{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 20s linear infinite;stroke-linecap:round;}#mermaid-svg-6hQEO452j9Hw4gs2 .error-icon{fill:#552222;}#mermaid-svg-6hQEO452j9Hw4gs2 .error-text{fill:#552222;stroke:#552222;}#mermaid-svg-6hQEO452j9Hw4gs2 .edge-thickness-normal{stroke-width:1px;}#mermaid-svg-6hQEO452j9Hw4gs2 .edge-thickness-thick{stroke-width:3.5px;}#mermaid-svg-6hQEO452j9Hw4gs2 .edge-pattern-solid{stroke-dasharray:0;}#mermaid-svg-6hQEO452j9Hw4gs2 .edge-thickness-invisible{stroke-width:0;fill:none;}#mermaid-svg-6hQEO452j9Hw4gs2 .edge-pattern-dashed{stroke-dasharray:3;}#mermaid-svg-6hQEO452j9Hw4gs2 .edge-pattern-dotted{stroke-dasharray:2;}#mermaid-svg-6hQEO452j9Hw4gs2 .marker{fill:#333333;stroke:#333333;}#mermaid-svg-6hQEO452j9Hw4gs2 .marker.cross{stroke:#333333;}#mermaid-svg-6hQEO452j9Hw4gs2 svg{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;}#mermaid-svg-6hQEO452j9Hw4gs2 p{margin:0;}#mermaid-svg-6hQEO452j9Hw4gs2 .label{font-family:"trebuchet ms",verdana,arial,sans-serif;color:#333;}#mermaid-svg-6hQEO452j9Hw4gs2 .cluster-label text{fill:#333;}#mermaid-svg-6hQEO452j9Hw4gs2 .cluster-label span{color:#333;}#mermaid-svg-6hQEO452j9Hw4gs2 .cluster-label span p{background-color:transparent;}#mermaid-svg-6hQEO452j9Hw4gs2 .label text,#mermaid-svg-6hQEO452j9Hw4gs2 span{fill:#333;color:#333;}#mermaid-svg-6hQEO452j9Hw4gs2 .node rect,#mermaid-svg-6hQEO452j9Hw4gs2 .node circle,#mermaid-svg-6hQEO452j9Hw4gs2 .node ellipse,#mermaid-svg-6hQEO452j9Hw4gs2 .node polygon,#mermaid-svg-6hQEO452j9Hw4gs2 .node path{fill:#ECECFF;stroke:#9370DB;stroke-width:1px;}#mermaid-svg-6hQEO452j9Hw4gs2 .rough-node .label text,#mermaid-svg-6hQEO452j9Hw4gs2 .node .label text,#mermaid-svg-6hQEO452j9Hw4gs2 .image-shape .label,#mermaid-svg-6hQEO452j9Hw4gs2 .icon-shape .label{text-anchor:middle;}#mermaid-svg-6hQEO452j9Hw4gs2 .node .katex path{fill:#000;stroke:#000;stroke-width:1px;}#mermaid-svg-6hQEO452j9Hw4gs2 .rough-node .label,#mermaid-svg-6hQEO452j9Hw4gs2 .node .label,#mermaid-svg-6hQEO452j9Hw4gs2 .image-shape .label,#mermaid-svg-6hQEO452j9Hw4gs2 .icon-shape .label{text-align:center;}#mermaid-svg-6hQEO452j9Hw4gs2 .node.clickable{cursor:pointer;}#mermaid-svg-6hQEO452j9Hw4gs2 .root .anchor path{fill:#333333!important;stroke-width:0;stroke:#333333;}#mermaid-svg-6hQEO452j9Hw4gs2 .arrowheadPath{fill:#333333;}#mermaid-svg-6hQEO452j9Hw4gs2 .edgePath .path{stroke:#333333;stroke-width:2.0px;}#mermaid-svg-6hQEO452j9Hw4gs2 .flowchart-link{stroke:#333333;fill:none;}#mermaid-svg-6hQEO452j9Hw4gs2 .edgeLabel{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-6hQEO452j9Hw4gs2 .edgeLabel p{background-color:rgba(232,232,232, 0.8);}#mermaid-svg-6hQEO452j9Hw4gs2 .edgeLabel rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-6hQEO452j9Hw4gs2 .labelBkg{background-color:rgba(232, 232, 232, 0.5);}#mermaid-svg-6hQEO452j9Hw4gs2 .cluster rect{fill:#ffffde;stroke:#aaaa33;stroke-width:1px;}#mermaid-svg-6hQEO452j9Hw4gs2 .cluster text{fill:#333;}#mermaid-svg-6hQEO452j9Hw4gs2 .cluster span{color:#333;}#mermaid-svg-6hQEO452j9Hw4gs2 div.mermaidTooltip{position:absolute;text-align:center;max-width:200px;padding:2px;font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:12px;background:hsl(80, 100%, 96.2745098039%);border:1px solid #aaaa33;border-radius:2px;pointer-events:none;z-index:100;}#mermaid-svg-6hQEO452j9Hw4gs2 .flowchartTitleText{text-anchor:middle;font-size:18px;fill:#333;}#mermaid-svg-6hQEO452j9Hw4gs2 rect.text{fill:none;stroke-width:0;}#mermaid-svg-6hQEO452j9Hw4gs2 .icon-shape,#mermaid-svg-6hQEO452j9Hw4gs2 .image-shape{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-6hQEO452j9Hw4gs2 .icon-shape p,#mermaid-svg-6hQEO452j9Hw4gs2 .image-shape p{background-color:rgba(232,232,232, 0.8);padding:2px;}#mermaid-svg-6hQEO452j9Hw4gs2 .icon-shape .label rect,#mermaid-svg-6hQEO452j9Hw4gs2 .image-shape .label rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-6hQEO452j9Hw4gs2 .label-icon{display:inline-block;height:1em;overflow:visible;vertical-align:-0.125em;}#mermaid-svg-6hQEO452j9Hw4gs2 .node .label-icon path{fill:currentColor;stroke:revert;stroke-width:revert;}#mermaid-svg-6hQEO452j9Hw4gs2 :root{--mermaid-font-family:"trebuchet ms",verdana,arial,sans-serif;} do_idle 循环

idle 线程没活干
cpuidle_select

(governor 选一个 C-state)
menu governor 预测

这次空闲会持续多久
cpuidle_enter

进入选中的 C-state
state->enter

(intel_idle → mwait)
中断到来 → 唤醒

回到运行态

一句话主线:cpuidle 是"CPU 怎么省电"的状态机 ------idle 线程没活干时,先由 governor 预测这轮空闲会持续多久 ,据此在 driver 提供的 C-state 列表里选一个"足够省电、又不会醒得太慢"的深度,然后通过 mwait(或老的 hlt)指令真正钻进去,等中断来唤醒。


一、预备概念:C-state 与"三个状态机"

1.1 C-state:越深越省电,醒得越慢

x86 的 C-state 是一组功耗逐级降低、唤醒延迟逐级升高的空闲状态:

状态 名称 功耗 唤醒延迟 说明
C0 运行态 最高 0 正常执行指令
C1 Halt 低 极短 停时钟(hlt),几乎无副作用
C1E Enhanced Halt 更低 短 额外降频降电压
C2 Stop Grant 更低 较长 停部分时钟
C6 Deep Power Down 很低 长 关核心电源,flush 缓存

关键规律:深度越深,省电越多,但"进入 + 退出"的代价越大。C6 能把核心电源都关掉,但退出要重新上电、恢复缓存,延迟可达几十微秒。所以"进多深"是个权衡------这正是 cpuidle 框架要解决的核心问题。

1.2 进入 C-state 的硬件手段:hlt vs mwait

x86 提供两条指令让 CPU 进 C-state:

  • hlt:最老的空闲指令,让 CPU 停在 C1,等中断唤醒。简单但只能进 C1。
  • monitor + mwait :更先进。monitor 先设置一个监视地址 ,mwait 让 CPU 进入 C-state(可带一个"hint"指定深度,如 C1/C1E/C6),当地址被写、或中断到来时自动唤醒。能进更深的 C-state,是现代 CPU 的首选。

1.3 三个"状态机"别搞混

写到这里必须钉死三个粒度的区别(这也是这个系列一直强调的):

状态机 管什么 核心问题 对应
CPU hotplug (cpuhp_state) 怎么把 CPU 拉起来/拆掉 offline → online 的几十步 CPU #3 hotplug
cpuidle(C-state) CPU 空闲时怎么省电 进多深的 C-state 本篇
cpufreq(P-state) CPU 跑多快(调频) 多高的频率 下一篇

三者都是"状态",但 hotplug 是"存在与否"、cpuidle 是"睡多深"、cpufreq 是"跑多快"。别把 cpuhp_state 和 cpuidle 的 C-state 混为一谈。


二、cpuidle 框架三层:driver / device / state

cpuidle 框架把"进 C-state"这件事拆成三层,各管一段:

2.1 cpuidle_state:一个 C-state 的参数

每个 C-state 用一个 cpuidle_state 描述,最关键的是两个时间参数:

c 复制代码
// include/linux/cpuidle.h (v6.6, line 49)
struct cpuidle_state {
	char		name[CPUIDLE_NAME_LEN];
	char		desc[CPUIDLE_DESC_LEN];

	s64		exit_latency_ns;      // 退出这个 state 的延迟(多久能醒)
	s64		target_residency_ns;  // 要停留多久才"划算"(能量回本)
	unsigned int	flags;
	unsigned int	exit_latency;        // 旧版(us),已废弃
	int		power_usage;         // 功耗(mW)
	unsigned int	target_residency;    // 旧版(us),已废弃

	int (*enter)	(struct cpuidle_device *dev,
			struct cpuidle_driver *drv,
			int index);          // 真正钻进去的回调
	// ...
};
  • exit_latency_ns:从这个 state 醒过来要多久。决定"会不会醒得太慢"。
  • target_residency_ns:要在这个 state 里停留多久,省的能源才抵得上进出的开销(能量盈亏平衡点)。决定"值不值得进"。

这两个参数是理解 governor 选 state 的关键(见第三节)。

2.2 cpuidle_driver:这台 CPU 有哪些 state

driver 提供一组 state(每台 CPU 型号不同),按功耗递减排序:

c 复制代码
// include/linux/cpuidle.h (v6.6, line 152)
struct cpuidle_driver {
	const char		*name;
	/* states array must be ordered in decreasing power consumption */
	struct cpuidle_state	states[CPUIDLE_STATE_MAX];   // state 数组
	int			state_count;                 // 有几个 state
	int			safe_state_index;            // 最安全的浅 state
	struct cpumask		*cpumask;
	const char		*governor;
};

注意注释 states array must be ordered in decreasing power consumption------state 数组按功耗递减排序 ,即 states[0] 功耗最高(通常是 polling 或 C1),越往后越省电(C6 等),也越深。x86 上有两个 driver:老的 acpi_idle(读 ACPI _CST 表)和现代的 intel_idle(直接用 mwait,见第五节)。

2.3 cpuidle_device:每 CPU 的实例 + 统计

driver 是共享的(同型号 CPU 一份),device 是每 CPU 一份:

c 复制代码
// include/linux/cpuidle.h (v6.6, line 93)
struct cpuidle_device {
	unsigned int		registered:1;
	unsigned int		enabled:1;
	unsigned int		cpu;
	ktime_t			next_hrtimer;

	int			last_state_idx;       // 上次进了哪个 state
	u64			last_residency_ns;    // 上次实际停留了多久
	struct cpuidle_state_usage	states_usage[CPUIDLE_STATE_MAX];  // 每个 state 的统计
	// ...
};

states_usage[] 记录每个 state 的使用次数、实际停留时间、以及"选深了/选浅了"的次数 (above/below,见 4.3)------这些统计正是 governor 校正自己预测的依据。

2.4 cpuidle_governor:谁来决定进哪个 state

c 复制代码
// include/linux/cpuidle.h (v6.6, line 288)
struct cpuidle_governor {
	char			name[CPUIDLE_NAME_LEN];
	unsigned int		rating;               // 评分,高的优先

	int  (*select)		(struct cpuidle_driver *drv,
				struct cpuidle_device *dev,
				bool *stop_tick);       // 选 state
	void (*reflect)		(struct cpuidle_device *dev, int index);  // 反馈
};

governor 有两个动作:select(选 state)和 reflect(拿到实际停留时长后做校正)。v6.6 默认是 menu governor (第三节),老的 ladder 已基本不用。


三、menu governor:怎么预测"这轮空闲会持续多久"

menu governor 的核心是预测 idle 时长 ,然后选一个"停留时间回本、又不会醒太慢"的 state。它自己的注释(menu.c:31-109)把决策因素总结成三个:

3.1 三个决策因素

  1. 能量盈亏平衡点(Energy break even) :进/出 C-state 有能量开销,得停留足够久才划算------这个时长就是 target_residency。所以关键就是预测空闲时长。
  2. 性能影响(Performance impact):深 C-state 退出延迟大,会拖慢工作负载。越忙的系统越不能接受深的 state。
  3. 延迟容忍度(Latency tolerance):来自 pmqos 基础设施,用户/驱动可以声明"我能接受多长的延迟"。

3.2 怎么预测 idle 时长:两个预测器

menu 用一个 predicted_ns 作为预测,它来自两个预测器取最小值:

预测器 1:下一个定时器事件(next_timer_ns)× 校正因子

c 复制代码
// drivers/cpuidle/governors/menu.c (v6.6, line 292)
data->next_timer_ns = delta;                        // 最近的定时器/时钟事件
data->bucket = which_bucket(data->next_timer_ns, nr_iowaiters);
// 用历史校正因子修正这个估计
timer_us = div_u64((RESOLUTION * DECAY * NSEC_PER_USEC) / 2 +
			data->next_timer_ns *
				data->correction_factor[data->bucket],
		   RESOLUTION * DECAY * NSEC_PER_USEC);
predicted_ns = min((u64)timer_us * NSEC_PER_USEC, predicted_ns);

为什么需要校正因子?因为唤醒 CPU 的不只有定时器 ,还有中断。所以"下一个定时器"的估计偏乐观(实际往往更早被中断唤醒)。menu 用历史数据算一个校正因子(实际空闲时长 / 下一个定时器的比例),按"时长量级 + 是否有 IO 在等"分 12 个 bucket 分别维护。

预测器 2:重复间隔检测器(get_typical_interval)

c 复制代码
// drivers/cpuidle/governors/menu.c (v6.6, line 171)
static unsigned int get_typical_interval(struct menu_device *data)

有些场景"下一个定时器"完全不可用------比如鼠标、网络包这种固定间隔的硬件事件。menu 记录最近 8 次空闲间隔,如果这 8 次的标准差很小(很稳定),就用平均值作为预测。

拿到 predicted_ns 后,menu_select 遍历 state 数组,选最深的、且满足两个约束的:

c 复制代码
// drivers/cpuidle/governors/menu.c (v6.6, line 353)
for (i = 0; i < drv->state_count; i++) {
	struct cpuidle_state *s = &drv->states[i];

	if (dev->states_usage[i].disable)
		continue;

	if (idx == -1)
		idx = i; /* first enabled state */

	if (s->target_residency_ns > predicted_ns) {
		// 停留时间回不了本,这个 state 太深了,break
		// ...(细节略)
	}
	if (s->exit_latency_ns > latency_req)
		break;   // 退出延迟超过容忍度,太深了

	idx = i;     // 满足约束,继续往深处找
}

两个 break 条件正好对应两个参数:

  • target_residency_ns > predicted_ns:预测空闲不够长,停留回不了本 → 太深了,停。
  • exit_latency_ns > latency_req:退出延迟超过延迟容忍度 → 会醒得太慢,停。

而 latency_req 还会被 performance_multiplier 进一步收紧(menu.c:155):

c 复制代码
static inline int performance_multiplier(unsigned int nr_iowaiters)
{
	/* for IO wait tasks (per cpu!) we add 10x each */
	return 1 + 10 * nr_iowaiters;
}

实际是 predicted_ns 除以 multiplier 得到 interactivity_req,然后 latency_req = min(latency_req, interactivity_req)(menu.c:343)。所以 nr_iowaiters 越多(越忙),乘数越大,latency_req 越可能被压到 predicted_ns / multiplier 这个更小的值,越难选到深的 state------这就是"越忙越不能睡深"的实现。


四、进入/退出流程:从 idle 循环到 mwait

4.1 do_idle:idle 线程的主循环

每个 CPU 的 idle 线程在 do_idle(kernel/sched/idle.c:237)里无限循环,核心是:

c 复制代码
// kernel/sched/idle.c (v6.6, line 258)
while (!need_resched()) {
	rmb();
	local_irq_disable();
	// ...
	if (cpu_idle_force_poll || tick_check_broadcast_expired()) {
		tick_nohz_idle_restart_tick();
		cpu_idle_poll();          // 轮询模式(不进 C-state)
	} else {
		cpuidle_idle_call();      // 走 cpuidle 框架
	}
	// ...
}

每次循环:关中断 → 调 cpuidle_idle_call 进 C-state(阻塞直到被唤醒)→ 醒来后继续判断有没有活干。

4.2 cpuidle_idle_call:一次完整的"选 + 进 + 反馈"

c 复制代码
// kernel/sched/idle.c (v6.6, line 146)
static void cpuidle_idle_call(void)
{
	struct cpuidle_device *dev = cpuidle_get_device();
	struct cpuidle_driver *drv = cpuidle_get_cpu_driver(dev);
	int next_state, entered_state;

	if (need_resched()) {          // 又有活了,别进 idle
		local_irq_enable();
		return;
	}

	if (cpuidle_not_available(drv, dev)) {
		default_idle_call();       // 没有 cpuidle,兜底走 arch_cpu_idle
		goto exit_idle;
	}
	// ...
	next_state = cpuidle_select(drv, dev, &stop_tick);   // ① governor 选 state
	// ...
	entered_state = call_cpuidle(drv, dev, next_state);  // ② 进入 state
	cpuidle_reflect(dev, entered_state);                 // ③ 反馈给 governor
}

三步:选(select)→ 进(enter)→ 反馈(reflect) 。cpuidle_select 只是转调 governor 的 select(cpuidle.c:356)。

4.3 cpuidle_enter_state:真正钻进去 + 事后统计

cpuidle_enter(cpuidle.c:372)→ cpuidle_enter_state(cpuidle.c:211)是进入 state 的核心,也是"选深了还是选浅了"的统计发生地:

c 复制代码
// drivers/cpuidle/cpuidle.c (v6.6, line 246)
time_start = ns_to_ktime(local_clock_noinstr());
// ...
entered_state = target_state->enter(dev, drv, index);   // ← 真正钻进去(mwait)
// ...(醒来后)
time_end = ns_to_ktime(local_clock_noinstr());
// ...
diff = ktime_sub(time_end, time_start);                 // 实际停留了多久

dev->last_residency_ns = diff;
dev->states_usage[entered_state].time_ns += diff;
dev->states_usage[entered_state].usage++;

if (diff < drv->states[entered_state].target_residency_ns) {
	// 停留比预期短 → 这次"选深了"(above++)
	dev->states_usage[entered_state].above++;
} else if (diff > delay) {
	// 停留足够久,也许更深的 state 更合适 → 这次"选浅了"(below++)
	dev->states_usage[entered_state].below++;
}

关键:进入 state 前后的时间差 diff 就是"实际空闲时长"。它被用来:

  1. 更新 states_usage[] 的统计(time_ns、usage);
  2. 标记"选深了"(above,停留 < target_residency)还是"选浅了"(below,停留 > exit_latency 且够得上更深的 state)。

这些 above/below 统计会在下次 menu_update(menu.c:461)里被读出来,用于校正预测因子------这正是 cpuidle 反馈闭环的落点。


五、x86 底层:hlt 与 mwait

前面讲的都是框架,最后落地的还是 x86 的两条指令。

5.1 default_idle(hlt)与 mwait_idle(monitor/mwait)

c 复制代码
// arch/x86/kernel/process.c (v6.6, line 740)
void __cpuidle default_idle(void)
{
	raw_safe_halt();          // hlt 指令,进 C1
	raw_local_irq_disable();
}

mwait_idle 用的是 monitor/mwait 指令对:

c 复制代码
// arch/x86/kernel/process.c (v6.6, line 918)
static __cpuidle void mwait_idle(void)
{
	if (!current_set_polling_and_test()) {
		__monitor((void *)&current_thread_info()->flags, 0, 0);  // 监视 flags 地址
		if (!need_resched()) {
			__sti_mwait(0, 0);      // mwait 进 C-state(hint=0 即 C1)
			raw_local_irq_disable();
		}
	}
	__current_clr_polling();
}

arch_cpu_idle(process.c:777)通过 static_call 间接调用实际例程,select_idle_routine(process.c:936)在启动时决定用 mwait 还是 hlt。

5.2 intel_idle:用 mwait 的 hint 指定 C-state 深度

hlt/mwait_idle 只能进浅的 C1。要进更深的 C-state(C1E/C6),得用 intel_idle driver------它用 mwait 的 hint 参数指定深度:

c 复制代码
// drivers/idle/intel_idle.c (v6.6, line 124)
/*
 * MWAIT takes an 8-bit "hint" in EAX "suggesting"
 * the C-state (top nibble) and sub-state (bottom nibble)
 * 0x00 means "MWAIT(C1)", 0x10 means "MWAIT(C2)" etc.
 */
#define flg2MWAIT(flags) (((flags) >> 24) & 0xFF)   // 从 flags 取 hint
#define MWAIT2flg(eax) ((eax & 0xFF) << 24)

static __cpuidle int intel_idle(struct cpuidle_device *dev,
				struct cpuidle_driver *drv, int index)
{
	struct cpuidle_state *state = &drv->states[index];
	unsigned long eax = flg2MWAIT(state->flags);     // 这个 state 对应的 hint
	unsigned long ecx = 1; /* break on interrupt flag */

	mwait_idle_with_hints(eax, ecx);                 // mwait,eax = C-state hint
	return index;
}

mwait 的 hint 用一个字节表示,高 nibble 是主状态、低 nibble 是子状态:0x00=C1、0x01=C1E、0x10=C2、0x20=C6...... 这些 hint 编码在 intel_idle 的 state 表里(intel_idle.c:237 起),比如 C6 对应 MWAIT2flg(0x20),且带 CPUIDLE_FLAG_TLB_FLUSHED(C6 会 flush TLB)。

所以整条链是:governor 选中 C6 → cpuidle_enter_state 调 state->enter(= intel_idle)→ mwait(0x20) 钻进去。


六、为什么这样设计(Why 层)

6.1 为什么要有 governor,而不是一个固定阈值

"进多深"没有固定答案------这轮空闲可能 1 微秒(马上有活)也可能 100 毫秒(等磁盘)。固定阈值要么"太保守"(空闲长却只进 C1,浪费电),要么"太激进"(空闲短却钻进 C6,醒来慢还 flush 了 TLB)。所以必须预测 。governor 就是那个预测者,而预测不可能完美,于是又有了 above/below 统计 + 校正因子的反馈闭环(4.3 → 3.2),让预测越用越准。

6.2 为什么是 exit_latency 和 target_residency 两个参数,不是一个

这两个参数回答了两个不同的问题:

  • target_residency 回答"值不值":停留够久能量才回本,否则进出的能量开销比省的还多。
  • exit_latency 回答"允不允许":即使能量回本,如果醒来太慢拖累了关键路径(延迟敏感),也不能进。

一个是经济账(能量),一个是性能账(延迟)。menu 的两个 break 条件(3.3)正好对应这两本账。

6.3 为什么 mwait 比 hlt 好

hlt 只能进 C1,且每次唤醒都要走完整的中断流程。mwait 有三个优势:

  1. 能进更深的 state:hint 参数指定 C1E/C6 等,省电空间大得多;
  2. monitor 提供"地址监视" :可以在进入前检查监视地址,避免"刚要睡就来了活"的竞态(mwait_idle 里 __monitor 后 need_resched 再查一次);
  3. 自动唤醒:mwait 在"监视地址被写或中断到来"时自动醒来,不需要软件参与。

6.4 为什么"选深了/选浅了"要单独统计(above/below)

预测本质是对未来的猜测,总会错。above(选深了,实际停留 < target_residency)和 below(选浅了)把错误分类记录 下来,喂给 governor 的校正因子。这是整个 cpuidle 的精髓------它不追求一次预测准,而是靠反馈闭环让预测逐步收敛。没有这个统计,menu 就只是个"猜下一个定时器"的朴素预测器。


七、与相邻主题的边界

主题 边界
CPU hotplug(#3) hotplug 的 cpuhp_state 管"CPU 存不存在",cpuidle 的 C-state 管"空闲睡多深"------两套状态机,别混
中断与 IPI 中断是唤醒 idle 的手段(mwait 因中断而醒),IPI 里 reschedule 一类正是把睡着的 CPU 叫醒
cpufreq(下一篇) cpuidle 管"睡多深"(C-state),cpufreq 管"跑多快"(P-state 调频)

附:本篇关键源码索引

符号 位置
cpuidle_state(exit_latency / target_residency / enter) include/linux/cpuidle.h:49
cpuidle_driver(state 数组按功耗递减) include/linux/cpuidle.h:152
cpuidle_device(per-CPU + states_usage 统计) include/linux/cpuidle.h:93
cpuidle_governor(select / reflect) include/linux/cpuidle.h:288
cpuidle_idle_call(选 → 进 → 反馈) kernel/sched/idle.c:146
do_idle(idle 主循环) kernel/sched/idle.c:237
cpuidle_select / cpuidle_enter / cpuidle_reflect drivers/cpuidle/cpuidle.c:356 / :372 / :402
cpuidle_enter_state(进入 + above/below 统计) drivers/cpuidle/cpuidle.c:211
menu_select(预测 + 选 state) drivers/cpuidle/governors/menu.c:262
get_typical_interval / performance_multiplier drivers/cpuidle/governors/menu.c:171 / :155
default_idle(hlt) / mwait_idle(mwait) arch/x86/kernel/process.c:740 / :918
intel_idle(mwait hint 指定深度) drivers/idle/intel_idle.c:159
相关推荐
星夜夏空994 小时前
网络编程(5)—— Reactor实现(v3)
服务器·网络·php
程序猿老A4 小时前
从入门型到企业型:云服务器开放共享型到独享型规格升级
运维·服务器
殷色玫瑰4 小时前
C++ string类详解:常用接口、字符串操作与模拟实现
java·linux·c语言·开发语言·数据结构·c++
忆挽篱笙歌5 小时前
gcc,g++
linux
l1t5 小时前
修复WSL CreateInstance/E_UNEXPECTED和 mounted read-only 错误
linux·windows·wsl
Ruiery6 小时前
Linux 6.6内核 CPU 深度解析(九):时钟与 TSC — 内核怎么从 PIT/HPET 切到 TSC
linux·运维·服务器
吴声子夜歌6 小时前
Nginx应用与运维——Nginx日志管理
java·运维·nginx
打工仔折腾 AI6 小时前
用UU远程把家里电脑变成AI Agent常驻服务器:CLI、端口映射与代理实测
运维·服务器·人工智能·后端·python·电脑·ai agent 实战
91刘仁德6 小时前
Linux网络编程从入门到实战:UDP/TCP协议与socket编程全解析
linux·网络·笔记·tcp/ip·udp
guo_wen_qiang6 小时前
云服务器nacos搭建-集群&持久化
java·运维·服务器