Linux 6.6内核 CPU 启动深度解析(二):AP 拉起 — BSP 如何用 INIT-SIPI 唤醒其余 CPU

〇、全景:一个 CPU 唤醒另一个 CPU,靠"中断信号 + 低内存跳板"

上一篇讲的是 BSP(第一个 CPU)如何从实模式一路走到 kernel_init。但多核系统里,其余 CPU(AP,Application Processor)不是自己启动的,而是被 BSP 一个个"踢"醒的 。BSP 手里没有"启动 AP"的魔法指令,它只能靠 x86 的 INIT/SIPI 中断信号------向 AP 的 local APIC 发一个 STARTUP 中断,中断里捎带一个"入口地址",AP 收到后从那个地址(低内存里的 trampoline 跳板)开始执行,再一路爬到内核。
#mermaid-svg-HqeX4hHhTr4ERZ0k{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;fill:#333;}@keyframes edge-animation-frame{from{stroke-dashoffset:0;}}@keyframes dash{to{stroke-dashoffset:0;}}#mermaid-svg-HqeX4hHhTr4ERZ0k .edge-animation-slow{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 50s linear infinite;stroke-linecap:round;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edge-animation-fast{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 20s linear infinite;stroke-linecap:round;}#mermaid-svg-HqeX4hHhTr4ERZ0k .error-icon{fill:#552222;}#mermaid-svg-HqeX4hHhTr4ERZ0k .error-text{fill:#552222;stroke:#552222;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edge-thickness-normal{stroke-width:1px;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edge-thickness-thick{stroke-width:3.5px;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edge-pattern-solid{stroke-dasharray:0;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edge-thickness-invisible{stroke-width:0;fill:none;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edge-pattern-dashed{stroke-dasharray:3;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edge-pattern-dotted{stroke-dasharray:2;}#mermaid-svg-HqeX4hHhTr4ERZ0k .marker{fill:#333333;stroke:#333333;}#mermaid-svg-HqeX4hHhTr4ERZ0k .marker.cross{stroke:#333333;}#mermaid-svg-HqeX4hHhTr4ERZ0k svg{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;}#mermaid-svg-HqeX4hHhTr4ERZ0k p{margin:0;}#mermaid-svg-HqeX4hHhTr4ERZ0k .label{font-family:"trebuchet ms",verdana,arial,sans-serif;color:#333;}#mermaid-svg-HqeX4hHhTr4ERZ0k .cluster-label text{fill:#333;}#mermaid-svg-HqeX4hHhTr4ERZ0k .cluster-label span{color:#333;}#mermaid-svg-HqeX4hHhTr4ERZ0k .cluster-label span p{background-color:transparent;}#mermaid-svg-HqeX4hHhTr4ERZ0k .label text,#mermaid-svg-HqeX4hHhTr4ERZ0k span{fill:#333;color:#333;}#mermaid-svg-HqeX4hHhTr4ERZ0k .node rect,#mermaid-svg-HqeX4hHhTr4ERZ0k .node circle,#mermaid-svg-HqeX4hHhTr4ERZ0k .node ellipse,#mermaid-svg-HqeX4hHhTr4ERZ0k .node polygon,#mermaid-svg-HqeX4hHhTr4ERZ0k .node path{fill:#ECECFF;stroke:#9370DB;stroke-width:1px;}#mermaid-svg-HqeX4hHhTr4ERZ0k .rough-node .label text,#mermaid-svg-HqeX4hHhTr4ERZ0k .node .label text,#mermaid-svg-HqeX4hHhTr4ERZ0k .image-shape .label,#mermaid-svg-HqeX4hHhTr4ERZ0k .icon-shape .label{text-anchor:middle;}#mermaid-svg-HqeX4hHhTr4ERZ0k .node .katex path{fill:#000;stroke:#000;stroke-width:1px;}#mermaid-svg-HqeX4hHhTr4ERZ0k .rough-node .label,#mermaid-svg-HqeX4hHhTr4ERZ0k .node .label,#mermaid-svg-HqeX4hHhTr4ERZ0k .image-shape .label,#mermaid-svg-HqeX4hHhTr4ERZ0k .icon-shape .label{text-align:center;}#mermaid-svg-HqeX4hHhTr4ERZ0k .node.clickable{cursor:pointer;}#mermaid-svg-HqeX4hHhTr4ERZ0k .root .anchor path{fill:#333333!important;stroke-width:0;stroke:#333333;}#mermaid-svg-HqeX4hHhTr4ERZ0k .arrowheadPath{fill:#333333;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edgePath .path{stroke:#333333;stroke-width:2.0px;}#mermaid-svg-HqeX4hHhTr4ERZ0k .flowchart-link{stroke:#333333;fill:none;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edgeLabel{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-HqeX4hHhTr4ERZ0k .edgeLabel p{background-color:rgba(232,232,232, 0.8);}#mermaid-svg-HqeX4hHhTr4ERZ0k .edgeLabel rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-HqeX4hHhTr4ERZ0k .labelBkg{background-color:rgba(232, 232, 232, 0.5);}#mermaid-svg-HqeX4hHhTr4ERZ0k .cluster rect{fill:#ffffde;stroke:#aaaa33;stroke-width:1px;}#mermaid-svg-HqeX4hHhTr4ERZ0k .cluster text{fill:#333;}#mermaid-svg-HqeX4hHhTr4ERZ0k .cluster span{color:#333;}#mermaid-svg-HqeX4hHhTr4ERZ0k div.mermaidTooltip{position:absolute;text-align:center;max-width:200px;padding:2px;font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:12px;background:hsl(80, 100%, 96.2745098039%);border:1px solid #aaaa33;border-radius:2px;pointer-events:none;z-index:100;}#mermaid-svg-HqeX4hHhTr4ERZ0k .flowchartTitleText{text-anchor:middle;font-size:18px;fill:#333;}#mermaid-svg-HqeX4hHhTr4ERZ0k rect.text{fill:none;stroke-width:0;}#mermaid-svg-HqeX4hHhTr4ERZ0k .icon-shape,#mermaid-svg-HqeX4hHhTr4ERZ0k .image-shape{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-HqeX4hHhTr4ERZ0k .icon-shape p,#mermaid-svg-HqeX4hHhTr4ERZ0k .image-shape p{background-color:rgba(232,232,232, 0.8);padding:2px;}#mermaid-svg-HqeX4hHhTr4ERZ0k .icon-shape .label rect,#mermaid-svg-HqeX4hHhTr4ERZ0k .image-shape .label rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-HqeX4hHhTr4ERZ0k .label-icon{display:inline-block;height:1em;overflow:visible;vertical-align:-0.125em;}#mermaid-svg-HqeX4hHhTr4ERZ0k .node .label-icon path{fill:currentColor;stroke:revert;stroke-width:revert;}#mermaid-svg-HqeX4hHhTr4ERZ0k :root{--mermaid-font-family:"trebuchet ms",verdana,arial,sans-serif;} AP 侧(被唤醒者,走汇编跳板)
BSP 侧(唤醒者,走内核 C 代码)
STARTUP IPI 携带入口地址
smp_init()

kernel/smp.c:960
bringup_nonboot_cpus

kernel/cpu.c:1897
native_kick_ap

smpboot.c:1051
do_boot_cpu

smpboot.c:985
send_init_sequence +

2× STARTUP IPI

smpboot.c:812/879
trampoline_start(实模式)

trampoline_64.S:59
trampoline startup_32(保护模式)

trampoline_64.S:124
trampoline startup_64(长模式)

trampoline_64.S:207
secondary_startup_64

head_64.S:121(BSP/AP 共用)
start_secondary

smpboot.c:239
cpu_startup_entry

AP 变 idle

一句话主线:AP 拉起的本质,是 BSP 通过 INIT/SIPI 中断把"入口地址"丢给 AP,AP 从低内存的 trampoline(实模式→保护模式→长模式三段跳)爬到内核的 secondary_startup_64,再进入 start_secondary 完成自己的初始化。整条链路分两半:BSP 侧是"发信号 + 准备环境",AP 侧是"接信号 + 爬进内核"。


一、预备概念:四个术语先钉死

1.1 INIT / SIPI(STARTUP IPI)

这是 x86 用来远程启动另一个 CPU 的两个特殊中断(通过 local APIC 的 ICR 寄存器发送):

信号 全称 作用
INIT Initialize 让目标 CPU 复位到已知状态(清寄存器、回实模式),类似硬件复位但不改内存
SIPI / STARTUP Startup Inter-Processor Interrupt 让目标 CPU 从指定地址开始执行实模式代码

关键点:STARTUP IPI 的"vector"字段不是中断向量,而是"入口地址 >> 12" 。因为 vector 只有 8 位,放不下完整物理地址,所以约定放"入口地址右移 12 位"(要求入口 4KB 对齐),AP 收到后从 vector << 12 处取指。这正是 trampoline 必须待在低内存、且页对齐的原因。

1.2 trampoline(跳板)

一块被内核复制到低内存(1MB 以下)的实模式代码,作用相当于 AP 的"迷你启动器"。BSP 篇里 BSP 走的是 arch/x86/boot/ 那套复杂的实模式→保护模式→长模式流程;AP 不需要重复那套(BIOS 信息已经收集过了),所以用一段精简版的 trampoline 完成同样的模式跃迁。

c 复制代码
// arch/x86/realmode/init.c (v6.6, line 151)
trampoline_header->start = (u64) secondary_startup_64;  // trampoline 最终跳到这里

1.3 idle 线程

每个 CPU 都要有一个"没活干时待着"的 idle 线程。BSP 的 idle 是 rest_init 里自己变的(上一篇讲过),而 AP 的 idle 线程是 BSP 在唤醒 AP 之前替它建好的 (idle_threads_init),AP 醒来后直接用。

1.4 握手(sync state)

BSP 把 AP 踢醒后,不能立刻放手 ------AP 可能还没爬进内核。所以两者用一个 ap_sync_state 原子变量握手:AP 爬到 cpuhp_ap_sync_alive() 时置 SYNC_STATE_ALIVE,BSP 看到后置 SYNC_STATE_SHOULD_ONLINE 放行。这是"踢醒"和"真正开始初始化"之间的同步屏障。


二、BSP 侧:smp_init 到 INIT-SIPI-STARTUP

2.1 smp_init:给每个 CPU 建 idle 线程,然后唤醒所有 AP

上一篇的 kernel_init_freeable() 末尾会调 smp_init(),这是整个 AP 拉起的起点:

c 复制代码
// kernel/smp.c (v6.6, line 960)
void __init smp_init(void)
{
    int num_nodes, num_cpus;

    idle_threads_init();          // ① 为每个 present CPU 建 idle 线程
    cpuhp_threads_init();         // ② 初始化 hotplug 线程(状态机用)

    pr_info("Bringing up secondary CPUs ...\n");

    bringup_nonboot_cpus(setup_max_cpus);  // ③ 唤醒所有 non-boot CPU

    num_nodes = num_online_nodes();
    num_cpus  = num_online_cpus();
    pr_info("Brought up %d node%s, %d CPU%s\n",
            num_nodes, (num_nodes > 1 ? "s" : ""),
            num_cpus,  (num_cpus  > 1 ? "s" : ""));
    smp_cpus_done(setup_max_cpus);         // ④ 收尾(拓扑建立等)
}

bringup_nonboot_cpus(kernel/cpu.c:1897)负责把 cpu_present_mask 里的所有 non-boot CPU 唤醒。它不是简单的一个接一个踢,而是有三种组织模式(串行 / 并行 / 分批)------下一节先讲清楚这三种模式的区别和动机,2.3 / 2.4 再深入"踢醒单个 CPU"的具体动作。

2.2 多个 AP 的唤醒顺序:串行 / 并行 / 分批

c 复制代码
// kernel/cpu.c (v6.6, line 1897)
void __init bringup_nonboot_cpus(unsigned int setup_max_cpus)
{
    /* 优先尝试并行 bringup 优化 */
    if (cpuhp_bringup_cpus_parallel(setup_max_cpus))
        return;

    /* 否则逐个 CPU 串行 bringup */
    cpuhp_bringup_mask(cpu_present_mask, setup_max_cpus, CPUHP_ONLINE);
}

为什么要有三种模式?本质是一个"保底 → 加速 → 受约束"的演进,不是三种并列的设计:

  1. 串行是保底 :最简单可靠,逐个 CPU 完整 bringup,也永远是并行不可用(TDX/SEV 禁用、cpuhp.parallel=0)时的退回路径。
  2. 并行是加速 :串行的启动时间随 CPU 数线性增长 ------每个 AP 从收到 STARTUP IPI 到爬进内核(trampoline → start_secondary → cpuhp_ap_sync_alive,中间还夹着微码加载)都要几十毫秒,上百核的总启动时间就高达好几秒。并行把"发 SIPI"和"等 alive"拆开,让所有 AP 并行爬 ,启动时间从"每个 AP 时间之和"降到"最慢那个 AP 的时间"。源码注释直接点明动机(kernel/cpu.c:1855):This avoids waiting for each AP to respond to the startup IPI。
  3. 分批是并行之上的正确性约束 :并行也不能完全无序------硬件限制 primary thread 更新微码时 SMT sibling 必须停着 (kernel/cpu.c:1871 注释),所以 SMT 开启时,必须在并行框架里再拆"先 primary、后 sibling"。

三种模式的核心区别,在于"踢醒(kick)"和"上线(online)"是一步到位还是拆成两阶段:

模式 触发条件 组织方式 smpboot_control
串行 默认 / 无并行优化 / 平台禁用 逐个 CPU:kick+online 一气呵成 = cpu(逐个写 CPU 号)
并行 CONFIG_HOTPLUG_PARALLEL + 平台支持 两阶段:先逐个 kick(不等响应),再逐个 online = STARTUP_READ_APICID
分批 并行 + SMT(超线程) 先 primary threads,再 SMT siblings 同并行

(1)串行:一步到位,简单但慢

cpuhp_bringup_mask(cpu_present_mask, ncpus, CPUHP_ONLINE) 的"逐个"在源码里就是一个 for_each_cpu 循环:

c 复制代码
// kernel/cpu.c (v6.6, line 1807)
static void __init cpuhp_bringup_mask(const struct cpumask *mask, unsigned int ncpus,
                                      enum cpuhp_state target)
{
    unsigned int cpu;

    for_each_cpu(cpu, mask) {
        struct cpuhp_cpu_state *st = per_cpu_ptr(&cpuhp_state, cpu);

        if (cpu_up(cpu, target) && can_rollback_cpu(st)) {   // 逐个 CPU 拉到 target
            ...
        }
        if (!--ncpus)
            break;
    }
}

for_each_cpu 逐个调 cpu_up(cpu, CPUHP_ONLINE)------踢醒 CPU1、等它爬到 cpuhp_ap_sync_alive、放行、直到 online 全部完成后 才轮到 CPU2。每个 AP 都要串行等待它响应 STARTUP IPI,CPU 越多越慢。这也正是下面 2.3 节 do_boot_cpu 里逐个写 smpboot_control = cpu 的原因。

(2)并行:先逐个踢(不等响应),再逐个放行

c 复制代码
// kernel/cpu.c (v6.6, line 1858)
static bool __init cpuhp_bringup_cpus_parallel(unsigned int ncpus)
{
    const struct cpumask *mask = cpu_present_mask;

    if (__cpuhp_parallel_bringup)
        __cpuhp_parallel_bringup = arch_cpuhp_init_parallel_bringup();
    if (!__cpuhp_parallel_bringup)
        return false;                       // 平台不支持 → 退回串行

    if (cpuhp_smt_aware()) {                // SMT 时先起 primary threads(见下)
        ...
    }

    /* 阶段 1:逐个发 STARTUP IPI 踢醒所有 AP(不等响应) */
    cpuhp_bringup_mask(mask, ncpus, CPUHP_BP_KICK_AP);
    /* 阶段 2:逐个放行到 online */
    cpuhp_bringup_mask(mask, ncpus, CPUHP_ONLINE);
    return true;
}

并行的核心优化是把"踢醒"和"上线"拆成两个 hotplug 状态 (CPUHP_BP_KICK_AP 和 CPUHP_BRINGUP_CPU)。这个拆分在源码里看得一清二楚------对比并行模式和串行模式各自的 bringup 回调:

c 复制代码
// kernel/cpu.c (v6.6, line 800) ------ CPUHP_BP_KICK_AP 的 startup:只发 STARTUP,不等 alive
static int cpuhp_kick_ap_alive(unsigned int cpu)
{
    if (!cpuhp_can_boot_ap(cpu))
        return -EAGAIN;
    return arch_cpuhp_kick_ap_alive(cpu, idle_thread_get(cpu));  // 发完即返回,不等响应
}

// kernel/cpu.c (v6.6, line 808) ------ CPUHP_BRINGUP_CPU 的 startup:等 alive + 等 online
static int cpuhp_bringup_ap(unsigned int cpu)
{
    ...
    ret = cpuhp_bp_sync_alive(cpu);          // ① 等 AP 报告 alive
    ...
    ret = bringup_wait_for_ap_online(cpu);   // ② 等 AP online
    ...
}

// kernel/cpu.c (v6.6, line 840) ------ 串行模式的 CPUHP_BRINGUP_CPU:三件事合在一个回调里
static int bringup_cpu(unsigned int cpu)
{
    ...
    ret = __cpu_up(cpu, idle);               // 发 STARTUP IPI
    ...
    ret = cpuhp_bp_sync_alive(cpu);          // 等 alive
    ...
    ret = bringup_wait_for_ap_online(cpu);   // 等 online
    ...
}

看出区别了:串行模式的 bringup_cpu 把"发 SIPI + 等 alive + 等 online"三件事塞进同一个回调 ,所以每个 CPU 都是"发一个、等一个"一气呵成,处理完才轮到下一个;并行模式(CONFIG_HOTPLUG_PARALLEL 会 select CONFIG_HOTPLUG_SPLIT_STARTUP,见 arch/Kconfig:48)把"发 SIPI"拆成 cpuhp_kick_ap_alive(只发、不等),把"等 alive + 等 online"拆成 cpuhp_bringup_ap。

于是阶段 1(CPUHP_BP_KICK_AP)逐个发 STARTUP IPI、发完不等待,所有 AP 并行 爬 trampoline → secondary_startup_64 → start_secondary,最后在 cpuhp_ap_sync_alive 自旋等待;阶段 2(CPUHP_BRINGUP_CPU)再逐个等 alive + 放行。一句话概括:把"发 SIPI"和"等 alive"从每个 CPU 的原子操作里拆开------串行是"发一个、等一个",并行是"全发完、再一起等"。

并行模式要求 AP 能"自己识别自己是几号 CPU"(因为不再逐个写 smpboot_control = cpu),所以 x86 侧把 smpboot_control 改成 STARTUP_READ_APICID:

c 复制代码
// arch/x86/kernel/smpboot.c (v6.6, line 1181)
bool __init arch_cpuhp_init_parallel_bringup(void)
{
    if (!x86_cpuinit.parallel_bringup) {    // 默认 true(x86_init.c:129),TDX/SEV 禁用它
        pr_info("Parallel CPU startup disabled by the platform\n");
        return false;
    }
    smpboot_control = STARTUP_READ_APICID;  // bit31 = 1(0x80000000)
    return true;
}

这正是上一篇 head_64.S 里 secondary_startup_64 的 .Lread_apicid 分支(head_64.S:243)------smpboot_control 的 bit31 置位后,AP 不走"直接读 CPU 号"的默认路径,而是从 local APIC 读自己的 APICID,再查 cpuid_to_apicid 表得出 CPU 号。

(3)分批:SMT 时先起 primary threads

并行模式在 SMT(超线程)时还要分批------先起 primary threads,再起 SMT siblings:

c 复制代码
// kernel/cpu.c (v6.6, line 1867)
    if (cpuhp_smt_aware()) {
        const struct cpumask *pmask = cpuhp_get_primary_thread_mask();
        static struct cpumask tmp_mask __initdata;

        /* X86 要求:primary thread 做微码更新时,SMT sibling 必须停着 */
        cpumask_and(&tmp_mask, mask, pmask);
        cpuhp_bringup_mask(&tmp_mask, ncpus, CPUHP_BP_KICK_AP);  // 先踢 primary
        cpuhp_bringup_mask(&tmp_mask, ncpus, CPUHP_ONLINE);      // 先上 primary
        ncpus -= num_online_cpus();

        cpumask_andnot(&tmp_mask, mask, pmask);  // 剩下的是 SMT siblings
        mask = &tmp_mask;
    }
    /* 落到下面再踢 + 上 siblings */

cpuhp_smt_aware() 和"谁是 primary thread"在源码里都有明确定义:

c 复制代码
// kernel/cpu.c (v6.6, line 1838) ------ 系统是否启用 SMT(超线程)
static inline bool cpuhp_smt_aware(void)
{
    return cpu_smt_max_threads > 1;
}

// arch/x86/include/asm/topology.h (v6.6, line 146) ------ primary thread 掩码
extern struct cpumask __cpu_primary_thread_mask;
#define cpu_primary_thread_mask ((const struct cpumask *)&__cpu_primary_thread_mask)

// arch/x86/kernel/apic/apic.c (v6.6, line 2329) ------ 判定谁是 primary thread
static void cpu_mark_primary_thread(unsigned int cpu, unsigned int apicid)
{
    /* 隔离 APICID 里的 SMT 位,SMT 位全 0 的才是 primary thread */
    u32 mask = (1U << (fls(smp_num_siblings) - 1)) - 1;

    if (smp_num_siblings == 1 || !(apicid & mask))
        cpumask_set_cpu(cpu, &__cpu_primary_thread_mask);
}

"primary thread"就是 APICID 的 SMT 位(最低 log2(smp_num_siblings) 位)为 0 的那个逻辑线程------它对应物理核里的第一个线程。

分批的原因在源码注释里写得很清楚(kernel/cpu.c:1871):X86 要求防止 SMT sibling 在 primary thread 做微码更新时被停止------同一个物理核的两个逻辑线程共享微码状态,primary thread 更新微码时 sibling 不能同时在跑,所以必须先起 primary、更新完微码,再起 sibling。

小结:默认情况下,现代 x86(非 TDX/SEV)走的是并行 + SMT 分批 路径;cpuhp.parallel=0 或平台禁用时退回串行 。无论哪种模式,"踢醒单个 CPU"的动作都是下一节要讲的 native_kick_ap → do_boot_cpu → INIT-SIPI。

2.3 唤醒单个 CPU:native_kick_ap → do_boot_cpu

状态机执行到"踢醒"这一步时,最终落到 x86 的 native_kick_ap:

c 复制代码
// arch/x86/kernel/smpboot.c (v6.6, line 1051)
int native_kick_ap(unsigned int cpu, struct task_struct *tidle)
{
    int apicid = apic->cpu_present_to_apicid(cpu);  // 逻辑 CPU → APIC ID
    int err;

    if (apicid == BAD_APICID || !physid_isset(apicid, phys_cpu_present_map) ||
        !apic_id_valid(apicid)) {
        pr_err("%s: bad cpu %d\n", __func__, cpu);
        return -EINVAL;
    }

    mtrr_save_state();                            // 保存 MTRR 状态
    per_cpu(fpu_fpregs_owner_ctx, cpu) = NULL;    // FPU 上下文清空

    err = common_cpu_up(cpu, tidle);              // ① 准备 AP 的 current_task、栈金丝雀
    if (err)
        return err;

    err = do_boot_cpu(apicid, cpu, tidle);        // ② 真正发 INIT-SIPI
    if (err)
        pr_err("do_boot_cpu failed(%d) to wakeup CPU#%u\n", err, cpu);

    return err;
}

common_cpu_up(smpboot.c:957)做的是"软环境准备":把 pcpu_hot.current_task[cpu] 设成该 CPU 的 idle 线程、初始化栈金丝雀、建中断栈。这些是 AP 醒来后马上要用的。

do_boot_cpu(smpboot.c:985)才是"发信号"的核心:

c 复制代码
// arch/x86/kernel/smpboot.c (v6.6, line 985)
static int do_boot_cpu(int apicid, int cpu, struct task_struct *idle)
{
    unsigned long start_ip = real_mode_header->trampoline_start;  // ① trampoline 入口
    int ret;

    idle->thread.sp = (unsigned long)task_pt_regs(idle);          // ② 设 idle 栈
    initial_code = (unsigned long)start_secondary;                // ③ 改写 initial_code!

    if (IS_ENABLED(CONFIG_X86_32)) {
        early_gdt_descr.address = (unsigned long)get_cpu_gdt_rw(cpu);
        initial_stack  = idle->thread.sp;
    } else if (!(smpboot_control & STARTUP_PARALLEL_MASK)) {
        smpboot_control = cpu;                                    // ④ 把 CPU 号塞进 smpboot_control
    }

    init_espfix_ap(cpu);
    announce_cpu(cpu, apicid);

    /* ⑤ 挑一个唤醒方法,默认走 INIT-SIPI */
    if (apic->wakeup_secondary_cpu_64)
        ret = apic->wakeup_secondary_cpu_64(apicid, start_ip);
    else if (apic->wakeup_secondary_cpu)
        ret = apic->wakeup_secondary_cpu(apicid, start_ip);
    else
        ret = wakeup_secondary_cpu_via_init(apicid, start_ip);

    if (ret)
        arch_cpuhp_cleanup_kick_cpu(cpu);
    return ret;
}

这里有两个和上一篇呼应的关键动作:

  • ③ initial_code = start_secondary :上一篇讲过,initial_code 的默认值是 x86_64_start_kernel(head_64.S:495),BSP 走 secondary_startup_64 时靠它跳进 C。而这里在唤醒 AP 前把它改成了 start_secondary ------所以 AP 走同一个 secondary_startup_64 后,跳的是 AP 自己的 C 入口,不是 BSP 的。
  • ④ smpboot_control = cpu :上一篇的 secondary_startup_64 里,汇编读 smpboot_control 来"识别自己是几号 CPU"。BSP 时代它是默认值 0;现在 BSP 把目标 CPU 号写进去,AP 醒来一读就知道自己是几号。

2.4 INIT-INIT-STARTUP 序列:为什么发两次 STARTUP

c 复制代码
// arch/x86/kernel/smpboot.c (v6.6, line 838)
static int wakeup_secondary_cpu_via_init(int phys_apicid, unsigned long start_eip)
{
    unsigned long send_status = 0, accept_status = 0;
    int num_starts, j, maxlvt;

    preempt_disable();
    maxlvt = lapic_get_maxlvt();
    send_init_sequence(phys_apicid);      // ① 先发 INIT

    mb();

    /* ② 根据 APIC 是否 integrated 决定发几次 STARTUP */
    if (APIC_INTEGRATED(boot_cpu_apic_version))
        num_starts = 2;
    else
        num_starts = 0;

    /* ③ STARTUP IPI 循环 */
    for (j = 1; j <= num_starts; j++) {
        if (maxlvt > 3)                   /* Pentium erratum 3AP */
            apic_write(APIC_ESR, 0);
        apic_read(APIC_ESR);

        /* STARTUP IPI:vector 字段塞的是 start_eip >> 12 */
        apic_icr_write(APIC_DM_STARTUP | (start_eip >> 12), phys_apicid);

        udelay(init_udelay == 0 ? 10 : 300);   // 给 AP 时间接受
        send_status = safe_apic_wait_icr_idle();
        ...
    }
    ...
}

而 send_init_sequence(smpboot.c:812)做的是"INIT 的一次完整脉冲":

c 复制代码
// arch/x86/kernel/smpboot.c (v6.6, line 812)
static void send_init_sequence(int phys_apicid)
{
    /* Assert INIT:拉高 INIT 电平 */
    apic_icr_write(APIC_INT_LEVELTRIG | APIC_INT_ASSERT | APIC_DM_INIT, phys_apicid);
    safe_apic_wait_icr_idle();

    udelay(init_udelay);

    /* Deassert INIT:拉低 INIT 电平 */
    apic_icr_write(APIC_INT_LEVELTRIG | APIC_DM_INIT, phys_apicid);
    safe_apic_wait_icr_idle();
}

把信号语义拆开看:

  1. INIT(assert → deassert):让 AP 复位到实模式、清寄存器。这是"先按一下复位键",保证 AP 从一个干净、已知的状态开始。
  2. STARTUP × 2 :让 AP 从 start_eip(trampoline 入口)开始执行。发两次是历史兼容(老 CPU 对单次 STARTUP 不可靠,多发一次无害)。

完整的信号序列是:一次 INIT(assert → deassert)+ 两次 STARTUP 。源码注释写作 "INIT, INIT, STARTUP"------前两个 "INIT" 就是 assert 和 deassert 两个动作。AP 收到 STARTUP 后,从 (start_eip >> 12) << 12(即 trampoline 入口,因 4KB 对齐所以等于 start_eip)开始取指。


三、AP 侧:trampoline 三连跳

AP 被 STARTUP 唤醒后,从低内存的 trampoline 开始执行。这一段和 BSP 篇的 compressed 自解压目的相同 (实模式→保护模式→长模式),但精简得多------因为内存映射、CPU 特性这些 BSP 早就探测好了,AP 直接复用。

3.1 trampoline_start:实模式,开门进保护模式

c 复制代码
// arch/x86/realmode/rm/trampoline_64.S (v6.6, line 59)
SYM_CODE_START(trampoline_start)
    cli
    wbinvd

    LJMPW_RM(1f)
1:
    mov     %cs, %ax        # 代码和数据在同一段
    mov     %ax, %ds
    mov     %ax, %es
    mov     %ax, %ss

    LOCK_AND_LOAD_REALMODE_ESP

    call    verify_cpu      # ① 校验 CPU 支持长模式
    testl   %eax, %eax
    jnz     no_longmode

.Lswitch_to_protected:
    lidtl   tr_idt          # ② 加载 IDT(空表)+ GDT
    lgdtl   tr_gdt

    movw    $__KERNEL_DS, %dx

    movl    $(CR0_STATE & ~X86_CR0_PG), %eax
    movl    %eax, %cr0      # ③ 开保护模式(不开分页,因为还没页表)

    ljmpl   $__KERNEL32_CS, $pa_startup_32   # ④ 跳 32 位入口
SYM_CODE_END(trampoline_start)

对比 BSP 篇:BSP 的 main() 在实模式干了 12 件事(收集 BIOS 信息),而 AP 的 trampoline_start 只干 1 件事------verify_cpu 校验后直接切保护模式。因为 BIOS 信息 BSP 已经收集完放进 boot_params 了,AP 不需要再问 BIOS。

3.2 trampoline startup_32:开 PAE、装页表、进长模式

c 复制代码
// arch/x86/realmode/rm/trampoline_64.S (v6.6, line 124)
SYM_CODE_START(startup_32)
    movl    %edx, %ss
    addl    $pa_real_mode_base, %esp
    movl    %edx, %ds / %es / %fs / %gs

    ...  // SME 内存加密位检查(省略)

    movl    pa_tr_cr4, %eax
    movl    %eax, %cr4        # ① 开 PAE

    movl    $pa_trampoline_pgd, %eax
    movl    %eax, %cr3        # ② 装 trampoline 页表

    movl    $MSR_EFER, %ecx
    rdmsr
    ...  // 写 EFER(含 LME)   ③ 使能长模式

    movl    $CR0_STATE, %eax
    movl    %eax, %cr0        # ④ 开分页,激活长模式

    ljmpl   $__KERNEL_CS, $pa_startup_64   # ⑤ 跳 64 位入口
SYM_CODE_END(startup_32)

这套操作和 BSP 篇 compressed 的 startup_32 几乎一模一样(开 PAE → 装页表 → EFER.LME → CR0.PG → 跳 64 位),区别只有两个:

  1. 页表不同 :AP 装的是 pa_trampoline_pgd------一块专门为 AP 准备的、低内存可寻址的临时页表。init.c:170-171 把 init_top_pgt 的高地址内核映射 (__PAGE_OFFSET 以上)复制进来,trampoline_pgd[0] 则单独映射低地址的 real mode stub,所以 AP 跑在它上面既能访问内核高地址、又能访问 trampoline 代码。
  2. 不用现搓页表 :BSP 篇的 compressed startup_32 要手写 4GB 恒等映射页表;AP 直接复用 BSP 早就建好的 init_top_pgt。

3.3 trampoline startup_64:一跳进内核

c 复制代码
// arch/x86/realmode/rm/trampoline_64.S (v6.6, line 207)
SYM_CODE_START(startup_64)
    jmpq    *tr_start(%rip)   # tr_start = secondary_startup_64
SYM_CODE_END(startup_64)

tr_start 就是 init.c:151 设好的 secondary_startup_64。所以 AP 经过 trampoline 三连跳,最终落在上一篇详讲过的 secondary_startup_64 (head_64.S:121)------那个 BSP/AP 共用入口。

从 secondary_startup_64 开始,AP 走的流程和 BSP 后半段完全一致(上一篇已详讲):读 smpboot_control 识别自己是几号 CPU → 换栈 → 装 GDT/IDT → 跳 initial_code。区别是:BSP 时代 initial_code 还是 x86_64_start_kernel,而现在 do_boot_cpu 已把它改成了 start_secondary。于是 AP 跳进了 AP 专属的 C 入口。


四、start_secondary:AP 的 C 代码,从握手到变 idle

4.1 完整流程

c 复制代码
// arch/x86/kernel/smpboot.c (v6.6, line 239)
static void notrace start_secondary(void *unused)
{
    cr4_init();                          // ① 初始化 CR4
    if (IS_ENABLED(CONFIG_X86_32)) {     // 64 位不用(页表已对)
        load_cr3(swapper_pg_dir);
        __flush_tlb_all();
    }

    cpu_init_exception_handling();       // ② 建 CPU 专属异常处理
    if (IS_ENABLED(CONFIG_X86_64))
        load_ucode_ap();                 // ③ 加载微码(AP 版)

    cpuhp_ap_sync_alive();               // ④ 握手:置 ALIVE,等 BSP 放行

    cpu_init();                          // ⑤ 初始化 CPU 状态(TSS、GDT、段...)
    fpu__init_cpu();
    rcu_cpu_starting(raw_smp_processor_id());
    x86_cpuinit.early_percpu_clock_init();

    ap_starting();                       // ⑥ 通知 hotplug 状态机"AP 在启动"

    check_tsc_sync_target();             // ⑦ 检查 TSC 是否和 BSP 同步
    ap_calibrate_delay();                // ⑧ 校准延时循环

    speculative_store_bypass_ht_init();

    lock_vector_lock();
    set_cpu_online(smp_processor_id(), true);  // ⑨ 标记自己 online!
    lapic_online();
    unlock_vector_lock();
    x86_platform.nmi_init();

    local_irq_enable();                  // ⑩ 开中断

    x86_cpuinit.setup_percpu_clockev();

    wmb();
    cpu_startup_entry(CPUHP_AP_ONLINE_IDLE);  // ⑪ 进入 idle 循环
}

对照 BSP 篇的 x86_64_start_kernel + start_kernel(BSP 的 C 入口),AP 的 start_secondary 是高度精简版 :BSP 要初始化内存、中断、调度器、时钟等整个世界(start_kernel 几十个调用),而 AP 只初始化"自己这台 CPU 需要的"东西(CR4、异常处理、微码、TSS、FPU、local APIC),最后直接进 idle。

4.2 关键点一:握手 cpuhp_ap_sync_alive

c 复制代码
// kernel/cpu.c (v6.6, line 392)
void cpuhp_ap_sync_alive(void)
{
    atomic_t *st = this_cpu_ptr(&cpuhp_state.ap_sync_state);

    cpuhp_ap_update_sync_state(SYNC_STATE_ALIVE);   // ① 报告"我活着,爬到这了"

    /* ② 自旋等 BSP 放行 */
    while (atomic_read(st) != SYNC_STATE_SHOULD_ONLINE)
        cpu_relax();
}

AP 在 cpu_init() 之前先在这里自旋等待 。这是一道"确认 AP 活着"的同步屏障------AP 报告 SYNC_STATE_ALIVE 后自旋等待,BSP 侧在 cpuhp_bp_sync_alive(kernel/cpu.c:437)里等 AP 报告 ALIVE 后,把它改成 SYNC_STATE_SHOULD_ONLINE 放行。之所以用原子变量自旋而不是 complete(),源码注释(kernel/cpu.c:433)写得很明白:bringup 早期 AP 还无法调用 complete(),所以只能用最原始的原子变量轮询。

4.3 关键点二:set_cpu_online 让 AP 正式"上线"

c 复制代码
set_cpu_online(smp_processor_id(), true);   // ⑨ 标记自己 online

上一篇讲过,BSP 在 boot_cpu_init(kernel/cpu.c:3155)里一次性把自己 online/active/present/possible 四态全置位。而 AP 是逐个态 推进的:present 在早期 ACPI 探测时就被置好,online 要到 start_secondary 这里才置上。这一行置位后,调度器、中断路由才真正把这台 CPU 当"可用"看。

4.4 关键点三:cpu_startup_entry 让 AP 也变 idle

c 复制代码
cpu_startup_entry(CPUHP_AP_ONLINE_IDLE);   // ⑪ 进入 idle 循环

和 BSP 篇 rest_init 末尾的 cpu_startup_entry(CPUHP_ONLINE) 对应------AP 完成初始化后,也变成 idle 线程 ,没活干时 schedule() 让出 CPU。区别只在于传入的状态参数:BSP 是 CPUHP_ONLINE(已经 fully online),AP 是 CPUHP_AP_ONLINE_IDLE(AP 侧 online idle)。

到这里,一台 AP 就拉起来了:从 INIT-SIPI 被踢醒,到 trampoline 三连跳,到 secondary_startup_64,到 start_secondary 完成初始化,最后进 idle。BSP 会重复这个流程,直到 cpu_present_mask 里的所有 CPU 都上线。


五、BSP vs AP:一张表看清两套启动路径

维度 BSP(第一篇) AP(本篇)
谁启动 自己(BIOS 指定) 被 BSP 踢醒(INIT-SIPI)
实模式入口 main()(arch/x86/boot/main.c) trampoline_start(trampoline_64.S)
实模式干什么 收集 BIOS 信息(12 步) 只 verify_cpu(信息 BSP 已收集)
模式跃迁 compressed 自解压(长流程) trampoline 三连跳(精简版)
页表 early_top_pgt(保留恒等映射) trampoline_pgd → init_top_pgt
内核入口 startup_64(head_64.S:46) secondary_startup_64(head_64.S:121)
C 入口 x86_64_start_kernel(head64.c:474) start_secondary(smpboot.c:239)
C 代码干什么 start_kernel 初始化整个世界 只初始化自己这台 CPU
online 时机 boot_cpu_init 四态一次置位 start_secondary 里逐个置位
终点 rest_init → idle(pid1/2 已建) cpu_startup_entry → idle

核心规律:AP 复用 BSP 已经建好的一切 (页表、内存映射、boot_params),所以它的启动路径是 BSP 的"精简 + 复用"版。


六、总结:记住三点

  1. AP 不是自启,是被"中断信号 + 低内存跳板"拉起来的 :BSP 通过 INIT-SIPI 序列(INIT 复位 + STARTUP 捎带入口地址)把 AP 踢醒;STARTUP 的 vector 字段放的是 入口地址 >> 12,所以 trampoline 必须待在低内存且页对齐。

  2. trampoline 三连跳 = BSP compressed 的"精简复用版" :trampoline_start(实模式)→ startup_32(开 PAE + 装 trampoline_pgd)→ startup_64(跳 secondary_startup_64)。它不开分页前不收集 BIOS、不手搓页表,直接复用 BSP 的 init_top_pgt。

  3. initial_code 和 smpboot_control 是两个"交接暗号" :BSP 唤醒 AP 前,把 initial_code 从默认的 x86_64_start_kernel 改成 start_secondary;smpboot_control 则按模式不同塞不同值------串行塞 CPU 号、并行塞 STARTUP_READ_APICID(让 AP 自己查表)。于是 AP 和 BSP 走同一个 secondary_startup_64,却跳进不同 的 C 入口、识别出不同的 CPU 号------这是 BSP/AP 共用入口的精妙之处。


附录:AP 拉起的完整调用链速查

BSP 侧(唤醒者)

复制代码
smp_init                          kernel/smp.c:960
  ├─ idle_threads_init            (每 CPU 建 idle 线程)
  └─ bringup_nonboot_cpus         kernel/cpu.c:1897
       ├─ [并行] cpuhp_bringup_cpus_parallel  kernel/cpu.c:1858(两阶段)
       └─ [串行] cpuhp_bringup_mask            (遍历 cpu_present_mask)
            └─ _cpu_up → cpuhp_up_callbacks(hotplug 状态机)
                 ├─ CPUHP_BP_KICK_AP → cpuhp_kick_ap_alive  kernel/cpu.c:800
                 │      └─ native_kick_ap                     smpboot.c:1051
                 │           ├─ common_cpu_up                 smpboot.c:957
                 │           └─ do_boot_cpu                   smpboot.c:985
                 │                └─ wakeup_secondary_cpu_via_init  smpboot.c:838
                 │                     ├─ send_init_sequence  smpboot.c:812(INIT)
                 │                     └─ STARTUP IPI × 2      smpboot.c:879
                 └─ CPUHP_BRINGUP_CPU → 等 AP alive + 放行(kernel/cpu.c:2148)

AP 侧(被唤醒者)

复制代码
trampoline_start      trampoline_64.S:59   (实模式)
trampoline startup_32 trampoline_64.S:124  (保护模式,开 PAE/长模式)
trampoline startup_64 trampoline_64.S:207  (长模式)
secondary_startup_64  head_64.S:121        (BSP/AP 共用入口,上一篇详讲)
start_secondary       smpboot.c:239        (AP 的 C 入口)
  ├─ cpuhp_ap_sync_alive   kernel/cpu.c:392(握手)
  ├─ cpu_init              (TSS/GDT/段)
  ├─ set_cpu_online        (正式上线)
  └─ cpu_startup_entry     (变 idle)

版本边界:本文只覆盖 x86、64 位 (CONFIG_X86_64)的 AP 拉起。ARM64 用 PSCI、RISC-V 用 SBI 的 HSM 扩展唤醒 AP,机制完全不同,不在本文范围。

相关推荐
Lsetea1 小时前
OpenSSL verify报error 62:证书主机名不匹配与-verify_hostname排查
运维·https·ssl证书·openssl·san
xcLeigh1 小时前
【KingbaseES数据库教程】国产化信创背景下的数据库选型与初识
linux·数据库·windows·kes·数据选型
rm -rf * && haha.sh1 小时前
【Linux】Rocky 9.8 操作系统磁盘分区与MySQL数据目录迁移
linux·运维·mysql
anew___1 小时前
《从零手写操作系统 (20):ELF加载器——让OS读懂现代编译器》
java·服务器·前端
高山有多高1 小时前
【Linux笔记】Socket编程基础
linux
吴声子夜歌1 小时前
Nginx应用与运维——Nginx Web服务应用实战(Python网站的搭建)
运维·前端·nginx
资深技术分享员2 小时前
Geejing WebBuilder 数据库连接配置与在线 SQL 工具,运维的左右手
运维·数据库·sql·低代码
高山有多高2 小时前
【Linux笔记】UpdSocket
linux·运维·笔记
皓月盈江2 小时前
Linux系统grep、sed 、awk介绍下,使用方法及区别?
linux·运维·服务器·sed·grep·awk