并发编程(1.5):操作系统层——Thread、Scheduler 与 Runtime 调度

目录

  • [0. 从语言执行单元到 CPU](#0. 从语言执行单元到 CPU)
  • [1. Java 的 Thread 是怎么调度的?](#1. Java 的 Thread 是怎么调度的?)
    • [1.1 模型演进:从 Platform Thread 到 Virtual Thread](#1.1 模型演进:从 Platform Thread 到 Virtual Thread)
      • [1.1.1 最简单的模型:一个 Java Thread 对应一个 OS Thread](#1.1.1 最简单的模型:一个 Java Thread 对应一个 OS Thread)
      • [1.1.2 问题:大量 OS Thread 的资源与调度成本](#1.1.2 问题:大量 OS Thread 的资源与调度成本)
      • [1.1.3 解决方案:让 JVM 自己调度轻量 Thread](#1.1.3 解决方案:让 JVM 自己调度轻量 Thread)
      • [1.1.4 最终模型:Platform Thread 与 Virtual Thread](#1.1.4 最终模型:Platform Thread 与 Virtual Thread)
    • [1.2 分层](#1.2 分层)
    • [1.3 Java Thread:从生命周期到真实实现](#1.3 Java Thread:从生命周期到真实实现)
      • [1.3.1 创建与启动:Created → Runnable](#1.3.1 创建与启动:Created → Runnable)
      • [1.3.2 等待调度:Runnable](#1.3.2 等待调度:Runnable)
      • [1.3.3 真正运行:Runnable → Running](#1.3.3 真正运行:Runnable → Running)
      • [1.3.4 暂停与等待:Running → Waiting](#1.3.4 暂停与等待:Running → Waiting)
      • [1.3.5 唤醒与恢复:Waiting → Runnable → Running](#1.3.5 唤醒与恢复:Waiting → Runnable → Running)
      • [1.3.6 结束与销毁:Running → Terminated](#1.3.6 结束与销毁:Running → Terminated)
    • [1.4 完整时序图](#1.4 完整时序图)
    • [1.5 Platform Thread 和 Virtual Thread 的区别](#1.5 Platform Thread 和 Virtual Thread 的区别)
    • [1.6 回到最开始提到的问题](#1.6 回到最开始提到的问题)
  • [2. Go 的 Goroutine 是怎么调度的?](#2. Go 的 Goroutine 是怎么调度的?)
    • [2.1 模型演进:从 G-M 到 G-M-P](#2.1 模型演进:从 G-M 到 G-M-P)
      • [2.1.1 最简单的模型:Goroutine 直接交给 OS Thread](#2.1.1 最简单的模型:Goroutine 直接交给 OS Thread)
      • [2.1.2 问题:大量 Goroutine 如何运行在有限 OS Thread 上?](#2.1.2 问题:大量 Goroutine 如何运行在有限 OS Thread 上?)
      • [2.1.3 解决方案:从 G-M 到 G-M-P](#2.1.3 解决方案:从 G-M 到 G-M-P)
      • [2.1.4 最终模型:G、M、P](#2.1.4 最终模型:G、M、P)
    • [2.2 分层](#2.2 分层)
    • [2.3 Goroutine:从生命周期到真实实现](#2.3 Goroutine:从生命周期到真实实现)
      • [2.3.1 创建:Created → Runnable](#2.3.1 创建:Created → Runnable)
      • [2.3.2 等待调度:Runnable](#2.3.2 等待调度:Runnable)
      • [2.3.3 真正运行:Runnable → Running](#2.3.3 真正运行:Runnable → Running)
      • [2.3.4 暂停与等待:Running → Waiting](#2.3.4 暂停与等待:Running → Waiting)
      • [2.3.5 唤醒与恢复:Waiting → Runnable → Running](#2.3.5 唤醒与恢复:Waiting → Runnable → Running)
      • [2.3.6 System Call 与 Preemption](#2.3.6 System Call 与 Preemption)
      • [2.3.7 结束与销毁:Running → Dead](#2.3.7 结束与销毁:Running → Dead)
    • [2.4 完整时序图](#2.4 完整时序图)
    • [2.5 回到最开始提到的问题](#2.5 回到最开始提到的问题)
  • [3. CPython 的 Thread 是怎么调度的?](#3. CPython 的 Thread 是怎么调度的?)
    • [3.1 模型演进:从 OS Thread 到 GIL / Free-Threaded](#3.1 模型演进:从 OS Thread 到 GIL / Free-Threaded)
      • [3.1.1 最简单的模型:Python Thread 对应 OS Thread](#3.1.1 最简单的模型:Python Thread 对应 OS Thread)
      • [3.1.2 问题:多个 OS Thread 如何共享 Interpreter Runtime State?](#3.1.2 问题:多个 OS Thread 如何共享 Interpreter Runtime State?)
      • [3.1.3 默认方案:用 GIL 约束 Interpreter Execution](#3.1.3 默认方案:用 GIL 约束 Interpreter Execution)
      • [3.1.4 进一步演进:Free-Threaded CPython](#3.1.4 进一步演进:Free-Threaded CPython)
      • [3.1.5 最终模型:Thread、PyThreadState 与 GIL](#3.1.5 最终模型:Thread、PyThreadState 与 GIL)
    • [3.2 分层](#3.2 分层)
    • [3.3 CPython Thread:从生命周期到真实实现](#3.3 CPython Thread:从生命周期到真实实现)
      • [3.3.1 创建与启动:Created → Runnable](#3.3.1 创建与启动:Created → Runnable)
      • [3.3.2 等待调度:Runnable](#3.3.2 等待调度:Runnable)
      • [3.3.3 真正运行:Runnable → Running](#3.3.3 真正运行:Runnable → Running)
      • [3.3.4 暂停与等待:Running → Waiting](#3.3.4 暂停与等待:Running → Waiting)
      • [3.3.5 唤醒与恢复:Waiting → Runnable → Running](#3.3.5 唤醒与恢复:Waiting → Runnable → Running)
      • [3.3.6 结束与销毁:Running → Terminated](#3.3.6 结束与销毁:Running → Terminated)
    • [3.4 完整时序图](#3.4 完整时序图)
    • [3.5 回到最开始提到的问题](#3.5 回到最开始提到的问题)
  • [4. 三种实现的共同套路](#4. 三种实现的共同套路)
    • [4.1 执行单元如何获得 CPU?------确定执行载体](#4.1 执行单元如何获得 CPU?——确定执行载体)
    • [4.2 谁决定下一个执行单元?------区分两层调度](#4.2 谁决定下一个执行单元?——区分两层调度)
    • [4.3 等待之后如何恢复?------重新获得执行机会](#4.3 等待之后如何恢复?——重新获得执行机会)
  • [5. 下一篇:Language Memory Model](#5. 下一篇:Language Memory Model)

0. 从语言执行单元到 CPU

上一篇从硬件层分析了 count++ 的执行过程,并介绍 Atomic Instruction、Cache Coherence 和 Fence 如何为原子性、可见性与有序性提供基础能力。

本篇把视角移到硬件的上一层,操作系统层。主要回答一个核心问题:

Java Thread、Go Goroutine 和 CPython Thread 这些执行单元如何通过 Runtime 与 Operating System,最终在 CPU 上执行?

要回答这个问题,需要理清三个环节:

  1. 执行映射:语言中的执行单元如何映射到 OS Thread?
  2. 调度分工:Runtime Scheduler 与 Operating System Scheduler 分别负责什么?
  3. 等待与恢复:执行单元等待、唤醒或恢复时,哪些对象被挂起,哪些线程还能继续执行?

接下来,我们分别以 Java、Go 和 CPython 为例,从线程模型的演进切入,结合执行分层、生命周期与时序,逐一回答上述三个问题;最后再提炼出语言无关的共同机制,回答开篇的核心问题。


1. Java 的 Thread 是怎么调度的?

1.1 模型演进:从 Platform Thread 到 Virtual Thread

1.1.1 最简单的模型:一个 Java Thread 对应一个 OS Thread

如果程序里只有少量并发任务,最直接的办法就是:

一个Java Thread对应一个OS Thread,由Java 负责创建 Thread,Operating System 负责调度对应的 OS Thread。

线程规模不大时,这种 1:1 映射足够简单,也容易理解和调试。

1.1.2 问题:大量 OS Thread 的资源与调度成本

假设服务里有大量请求,每个请求都使用一个 Thread:

text 复制代码
Request 1 → Thread 1 → OS Thread 1
Request 2 → Thread 2 → OS Thread 2
Request 3 → Thread 3 → OS Thread 3
...
Request N → Thread N → OS Thread N

很多服务任务的大部分时间都在等待:

text 复制代码
读取数据库
等待网络响应
等待消息
等待其他 I/O

这种情况下,如果 Java Thread 和 OS Thread 还一一对应,

会导致两个问题:

  1. 线程本身的资源占用开销:即使任务正在等待,Operating System 仍需要保留对应的 Kernel Thread、Native Stack 等资源,线程数量越多,这部分资源占用越高。
  2. 线程调度的上下文切换开销: 当大量 Thread 被频繁唤醒并进入 Runnable 时,Operating System 还要在更多 Runnable Thread 之间调度,Context Switch 和 Scheduler 本身也会消耗更多 CPU 时间。

1.1.3 解决方案:让 JVM 自己调度轻量 Thread

如何解决呢?

这种场景下,每个逻辑任务往往大部分时间都在等待 I/O,占用 CPU 的时间很短,

那就没必要让每个逻辑任务都占用一个OS Thread,

也就是把逻辑执行单元和OS Thread解耦。

两者解耦后,JVM 就可以在中间增加自己的 Scheduler:

text 复制代码
逻辑 Thread A ─┐
逻辑 Thread B ─┼─→ JVM Scheduler ─→ Carrier Thread 1 ⇔ OS Thread 1
逻辑 Thread C ─┤                  └→ Carrier Thread 2 ⇔ OS Thread 2
逻辑 Thread D ─┘

其中,Carrier Thread 是承担承载角色的 Platform Thread;⇔ 表示它与 OS Thread 的 1:1 原生映射,而非调用关系。

JVM 可以维护大量逻辑 Thread,而 Operating System 只需要调度较少 OS Thread。

这种逻辑Thread在Java中叫做:Virtual Thread

1.1.4 最终模型:Platform Thread 与 Virtual Thread

Platform Thread 与 OS Thread 一对一映射;Virtual Thread 由 JVM Scheduler 分派到作为 Carrier 的 Platform Thread 上执行,底层 OS Thread 均由 OS Scheduler 调度。


1.2 分层

1.3 Java Thread:从生命周期到真实实现

Java Thread 的生命周期可以抽象为:

注意,这里描述的是线程的概念生命周期,与 Java 的 Thread.State 并不完全对应。

例如,Java 的 RUNNABLE 既包括等待 CPU 调度的线程,也包括正在运行的线程。

1.3.1 创建与启动:Created → Runnable

Platform Thread 在 start() 时创建 OS Thread;Virtual Thread 则将 Continuation 提交给 JVM Scheduler。

(一)Platform Thread

(二)Virtual Thread

1.3.2 等待调度:Runnable

Platform Thread 等待 OS Scheduler 分配 CPU;Virtual Thread 等待 JVM Scheduler 将任务分派给 Carrier Thread。

(一)Platform Thread

(二)Virtual Thread

1.3.3 真正运行:Runnable → Running

Platform Thread 由 OS Thread 执行 Thread.run();Virtual Thread 则在 Carrier Thread 上挂载并执行 Continuation。

(一)Platform Thread

(二)Virtual Thread

1.3.4 暂停与等待:Running → Waiting

Platform Thread 等待时通常占用对应的 OS Thread;Virtual Thread 在允许卸载时可暂停自身并释放 Carrier。

(一)Platform Thread

(二)Virtual Thread

1.3.5 唤醒与恢复:Waiting → Runnable → Running

Platform Thread 由原 OS Thread 恢复执行;Virtual Thread 重新进入 JVM Scheduler,恢复时可以使用另一个 Carrier。

(一)Platform Thread

(二)Virtual Thread

1.3.6 结束与销毁:Running → Terminated

Platform Thread 结束后清理 JavaThread 和原生线程;Virtual Thread 结束时 Carrier Thread 仍可复用。

(一)Platform Thread

(二)Virtual Thread

1.4 完整时序图

下面分别展示 Platform Thread 的 Native 启动链路,以及 Virtual Thread 的 Mount / Unmount / Resume 关键路径。Virtual Thread 图从 Carrier 已获得 CPU 的条件开始,不重复 OS Scheduler 的调度时序。

1.4.1 Platform Thread

1.4.2 Virtual Thread

1.5 Platform Thread 和 Virtual Thread 的区别

按生命周期比较,两种 Thread 的差异如下:

生命周期阶段 Platform Thread Virtual Thread
Created → Runnable JVM 创建对应 OS Thread JVM 把 Continuation 提交给 Scheduler
Runnable 等待 OS Scheduler 由 JVM Scheduler 分派到 Carrier;Carrier 的 OS Thread 由 OS Scheduler 调度
Running OS Thread 直接执行 Java Code Virtual Thread 挂载到 Carrier Thread 后执行
Waiting OS Thread 通常一起等待 可卸载时 Virtual Thread 等待,Carrier 可执行其他任务
Resume OS Thread 重新 Runnable Virtual Thread 重新提交 JVM Scheduler
Terminated JavaThread / OS Thread Exit Continuation 完成,afterDone() 进入 TERMINATED

上述差异可以归结为承载关系:Platform Thread 长期绑定对应的 OS Thread;Virtual Thread 由 JVM Scheduler 分派给 Carrier,等待时如能卸载,就能释放 Carrier 去执行其他任务。

实现参考 OpenJDK 的 Thread.java、jvm.cpp、javaThread.cpp 和 VirtualThread.java;Virtual Thread 的设计见 JEP 444。

1.6 回到最开始提到的问题

由 [1.2 分层](#1.2 分层)可知,Platform Thread 与 OS Thread 一对一映射,而 Virtual Thread 通过 Carrier Thread 间接映射到 OS Thread。因此,两种线程最终都依赖 OS Thread 执行。

由 [1.3.2 等待调度](#1.3.2 等待调度) 和 [1.3.3 真正运行](#1.3.3 真正运行) 可知,JVM Scheduler 负责将 Virtual Thread 分派给 Carrier Thread,而 OS Scheduler 负责调度底层 OS Thread。因此,两种线程的调度路径不同,但最终都由 OS Scheduler 分配 CPU 执行机会。

再由 [1.3.4 暂停与等待](#1.3.4 暂停与等待) 和 [1.3.5 唤醒与恢复](#1.3.5 唤醒与恢复) 可知,Virtual Thread 在允许卸载的等待中可以释放 Carrier Thread,唤醒后重新参与 JVM 调度。因此,大量等待任务不必各自长期占用 OS Thread。

最终回答开篇的问题: Java 保留了 Platform Thread 的原生线程模型,同时通过 Virtual Thread 复用 Carrier Thread,让大量并发任务最终通过有限的 OS Thread 在 CPU 上执行。


2. Go 的 Goroutine 是怎么调度的?

2.1 模型演进:从 G-M 到 G-M-P

2.1.1 最简单的模型:Goroutine 直接交给 OS Thread

如果程序里只有少量并发任务,最直接的办法就是:

text 复制代码
Goroutine
    ↓
OS Thread
    ↓
Operating System Scheduler
    ↓
CPU Core

一个 Goroutine 对应一个 OS Thread,由 Go Runtime 管理 Goroutine,Operating System 负责调度对应的 OS Thread。

线程规模不大时,这种 1:1 映射足够简单,也容易理解和调试。

2.1.2 问题:大量 Goroutine 如何运行在有限 OS Thread 上?

假设服务里有大量请求,每个请求都使用一个 Goroutine:

text 复制代码
Request 1 → Goroutine 1 → OS Thread 1
Request 2 → Goroutine 2 → OS Thread 2
Request 3 → Goroutine 3 → OS Thread 3
...
Request N → Goroutine N → OS Thread N

很多服务任务的大部分时间都在等待数据库、网络响应或其他 I/O。

如果 Goroutine 和 OS Thread 仍然一一对应,就会遇到与 Java Platform Thread 相同的两个问题:

  1. 资源占用:等待中的任务仍然占用 Kernel Thread、Native Stack 等资源。
  2. 调度开销:大量 OS Thread 频繁变成 Runnable 时,Context Switch 和 OS Scheduler 的开销增加。

2.1.3 解决方案:从 G-M 到 G-M-P

如何解决呢?

这些逻辑任务多数时间都在等待 I/O,占用 CPU 的时间很短,因此没有必要让每个 Goroutine 长期占用一个 OS Thread。

也就是把 逻辑执行单元 和 OS Thread 解耦:Go Runtime 负责在多个 Goroutine 之间调度,让有限的 OS Thread 复用执行机会。

2.1.3.1 第一步:G-M + Global Run Queue

先从 G、M 和 Global Run Queue 的关系看起:

其中 G 表示 Goroutine,M 表示 OS Thread。

所有 Runnable G 主要进入同一个 Global Run Queue,多个 M 从这里获取工作。

随着并发增加,会出现两个问题:

text 复制代码
多个 M
  ↓
同时访问 Global Run Queue
  ↓
更多同步竞争

以及:

text 复制代码
G 的调度状态集中在全局
        ↓
局部性较差

要降低全局竞争,可以把一部分调度状态和 Run Queue 分散到各个执行上下文:

能不能给每个执行上下文一份自己的调度资源和 Local Run Queue?

2.1.3.2 第二步:引入 P + Local Run Queue

为减少对 Global Run Queue 的竞争,Go Runtime 引入 P,让每个 P 维护自己的 Local Run Queue:

P 持有执行 User Go Code 所需的 Runtime 资源,并维护自己的 Local Run Queue。实际的 G-M-P 调度仍保留 Global Run Queue,G 也不固定属于某个 P;后文再介绍其他任务来源。

于是得到 G-M-P 最核心的约束:

M 必须获得 P,才能执行 User Go Code。

这样调度状态从"所有 M 竞争一个全局中心",变成由各个 P 分散维护自己的 Local Run Queue。

2.1.3.3 第三个问题:P 的 Local Run Queue 没有工作怎么办?

Local Run Queue 分散了竞争,也带来负载不均:有的 P 很忙,有的 P 的队列已经为空。此时,持有空闲 P 的 M 可以通过 Work Stealing,从其他 P 的队列中取得 Runnable G。

除了 Work Stealing,Scheduler 还会从其他来源寻找可运行的 G:

真实 findRunnable() 会在这些来源之间结合公平性、Timer、Network Poller 等条件寻找可运行的 G,并不是机械地固定按一条顺序执行。

2.1.3.4 第四个问题:M 阻塞在 System Call 怎么办?

另一个约束来自 Blocking System Call:

持有 P 的 M 被 OS 阻塞时,P 是否也要一起闲置?

如果 P 永远绑定在这个 M 上,那么 M 被 OS 阻塞时,P 也无法继续让其他 G 运行。

P 和 M 因此必须能够解耦:

阻塞的 G0 仍由 M0 对应的 OS Thread 执行系统调用;P0 则可以被 Runtime 重新取得,让 M1 执行其他 Runnable G。M0 从系统调用返回后,需要尝试重新获得 P 才能继续执行 Go Code。这说明 P 与 M 必须能够解除绑定并重新组合。

2.1.3.5 第五个问题:一个 G 长时间占着 M 怎么办?

即使没有 I/O,如果某个 G 长时间持续执行,也会压缩其他 Runnable G 的执行机会。

因此 Runtime 还需要 Preemption:

text 复制代码
G Running
    ↓
Preempt
    ↓
G Runnable
    ↓
Scheduler
    ↓
其他 G 获得执行机会

现代 Go Runtime 支持抢占,包括异步抢占,用于降低长时间运行的 G 阻碍其他任务执行的风险;抢占本身不构成严格的公平性保证。

2.1.4 最终模型:G、M、P

2.1.4.1 G:Goroutine

G 是 Goroutine,也就是 Go Runtime 管理的逻辑执行单元。

它保存:

  • Stack;
  • PC / Scheduler Context;
  • Runnable、Running、Waiting 等调度状态。

go f() 创建的是 G,不是一个新的 OS Thread。

2.1.4.2 M:Machine

M 对应一个真实 OS Thread。

text 复制代码
M
↓
OS Thread
↓
OS Scheduler
↓
CPU

Operating System Scheduler 真正调度到 CPU 的对象是 M 对应的 OS Thread。

每个 M 还带有一个特殊的 g0。g0 使用 M 的 System Stack,Runtime 会在这里执行调度、栈管理等底层工作。

2.1.4.3 P:Processor

P 表示执行 User Go Code 所需的 Runtime 资源和执行资格。

它包含 Scheduler / Allocator State、Local Run Queue 和用于优先调度的 runnext。

P 的数量由 GOMAXPROCS 决定:

text 复制代码
GOMAXPROCS
    ↓
P 的数量
    ↓
同一时刻执行 User Go Code 的并行度上限

三者关系可以概括为:

G 是要运行的任务,M 是真实 OS Thread,P 是让 M 能够执行 User Go Code 的 Runtime 资源。

2.2 分层

图中的 Go Runtime Scheduler 与 OS Scheduler 是两个独立层级。Goroutine 在同一个 M 上切换时,不要求发生一次 OS Thread Context Switch。

2.3 Goroutine:从生命周期到真实实现

Goroutine 的生命周期可以抽象为:

图中用 Created、Waiting 等概念状态描述 Goroutine 的运行过程。Go Runtime 则使用 _Grunnable、_Gwaiting 等内部状态。

2.3.1 创建:Created → Runnable

go f() 创建一个 G,将它置为 Runnable 并加入调度队列。

2.3.2 等待调度:Runnable

G 已进入 Run Queue,但尚未被持有 P 的 M 选中执行。

2.3.3 真正运行:Runnable → Running

持有 P 的 M 通过 Runtime Scheduler 选取 G,随后执行 execute(G) 和 gogo;OS Scheduler 独立调度该 M 对应的 OS Thread。

2.3.4 暂停与等待:Running → Waiting

gopark() 只挂起 G,持有 P 的 M 可以继续执行其他 Runnable G。

2.3.5 唤醒与恢复:Waiting → Runnable → Running

等待结束后,goready() 使 G 重新入队,Scheduler 再将它交给可用的 M 执行。

2.3.6 System Call 与 Preemption

System Call 可能阻塞 M,此时 P 可被交给其他 M;Preemption 则使长时间运行的 G 重新参与调度。

2.3.7 结束与销毁:Running → Dead

G 的函数返回后,Runtime 清理 G 并将其标记为 Dead,M 继续调度其他 G。

2.4 完整时序图

以下时序用一个 G 展示创建、运行、等待和结束。OS 调度与 G 的每次切换是两个独立层级:图中只示意 M 获得 CPU 的一次情形,后续在同一 M 上切换 G 不需要再次经过 OS 调度。

实现参考 Go Runtime 的 proc.go 和 runtime2.go。

2.5 回到最开始提到的问题

由 [2.1.4 最终模型](#2.1.4 最终模型) 和 [2.2 分层](#2.2 分层)可知,G 是 Runtime 管理的逻辑执行单元,M 对应 OS Thread,而 M 必须持有 P 才能执行 Go Code。因此,G 不直接对应 OS Thread,而是通过持有 P 的 M 获得执行载体。

由 [2.1.3.3 任务来源](#2.1.3.3 任务来源) 和 [2.3.2 等待调度](#2.3.2 等待调度)、[2.3.3 真正运行](#2.3.3 真正运行) 可知,Go Scheduler 从 Run Queue、Work Stealing、Netpoll 等来源选择 G,由 M 执行;OS Scheduler 则调度 M 对应的 OS Thread。因此,两级 Scheduler 分别决定哪个 G 运行,以及哪个 OS Thread 获得 CPU 执行机会。

再由 [2.3.4 暂停与等待](#2.3.4 暂停与等待) 和 [2.3.5 唤醒与恢复](#2.3.5 唤醒与恢复) 可知,G 等待时 M 可以继续运行其他 G;[2.3.6 System Call 与 Preemption](#2.3.6 System Call 与 Preemption) 进一步说明,M 遇到阻塞系统调用时 P 可被其他 M 使用,长时间运行的 G 也可以被抢占。因此,等待或阻塞不必让对应的调度资源一直闲置。

最终回答开篇的问题: Go 通过 G-M-P 将大量 Goroutine 调度到持有 P 的 M 上,再由 OS Scheduler 将 M 对应的 OS Thread 调度到 CPU,使并发任务能够在有限的 CPU Core 上持续推进。


3. CPython 的 Thread 是怎么调度的?

3.1 模型演进:从 OS Thread 到 GIL / Free-Threaded

3.1.1 最简单的模型:Python Thread 对应 OS Thread

如果程序里只有少量并发任务,最直接的执行关系是:

text 复制代码
Python Thread
      ↓
OS Thread
      ↓
Operating System Scheduler
      ↓
CPU Core

OS Scheduler 决定哪个 OS Thread 获得 CPU。

3.1.2 问题:多个 OS Thread 如何共享 Interpreter Runtime State?

例如,从 OS 调度角度看,两个线程可能同时运行在不同 CPU Core 上:

text 复制代码
Python Thread A → OS Thread A → CPU Core A
Python Thread B → OS Thread B → CPU Core B

从 Operating System 的角度看,两个 OS Thread 都可以被并行调度。但 CPython Interpreter 内部还存在由多个 Thread 共同访问的 Runtime State,例如对象状态、引用计数以及 Interpreter 自身的数据结构。

因此这里多了一个 Runtime 层的问题:

多个 OS Thread 进入 Python Interpreter 执行时,如何同步这些共享 Runtime State?

默认开启 GIL 的 CPython 使用 GIL 保护解释器内部的共享状态。因此,线程能否执行 Python 解释器代码,需要分别看 OS 调度和 GIL:

text 复制代码
OS Scheduler
决定哪个 OS Thread 获得 CPU

GIL
约束哪个 Thread 可以进入受保护的 Python Interpreter Execution

3.1.3 默认方案:用 GIL 约束 Interpreter Execution

在默认 GIL-enabled CPython 中,Thread 执行受 GIL 保护的 Interpreter Code 时必须持有 GIL:

text 复制代码
OS Thread
    ↓
OS Scheduler
    ↓
CPU Time

Python Thread
    ↓
Acquire GIL
    ↓
Python Interpreter Execution

因此,一个 Thread 能执行 Python Bytecode,需要同时满足两类条件:OS Thread 获得 CPU,并且该 Thread 获得 GIL。

如果另一个 Thread 正持有 GIL:

text 复制代码
Thread A
  ↓
持有 GIL
  ↓
执行 Python Bytecode

Thread B
  ↓
尝试 acquire GIL
  ↓
等待
  ↓
GIL 被释放
  ↓
重新竞争

两者职责独立:

  • Operating System Scheduler 管 OS Thread → CPU;
  • GIL 管默认 CPython 中 Thread → Python Interpreter Execution。

3.1.4 进一步演进:Free-Threaded CPython

如果希望多个 Python Thread 并行执行 Python Code,就需要让正常执行路径不再依赖这个全局互斥门槛:

text 复制代码
Python Thread A → OS Thread A → CPU Core A
Python Thread B → OS Thread B → CPU Core B

Free-Threaded Build 不再把 GIL 作为正常执行路径上的全局互斥门槛,改用更细粒度的 Runtime 同步机制维护 CPython 内部状态,使多个 Thread 可以并行执行 Python Code。这改变的是 Interpreter Execution 的同步方式,OS Thread 仍由 Operating System Scheduler 调度。

CPython 当前有两种主要执行模型:

某些不兼容扩展仍可能让 Free-Threaded Runtime 重新启用 GIL。

3.1.5 最终模型:Thread、PyThreadState 与 GIL

3.1.5.1 Python Thread

Python 程序通过 threading.Thread 创建和管理线程。

在 CPython 中,它最终通过 _thread 创建真实 OS Thread。

3.1.5.2 OS Thread

OS Thread 是 Operating System Scheduler 真正调度到 CPU 的执行单元。

CPython 没有像 Go G-M-P 或 Java Virtual Thread 那样,把大量 threading.Thread 复用到少量 OS Thread 上。

3.1.5.3 PyThreadState

PyThreadState 保存一个 Python Thread 在 Interpreter 中的 Runtime State。

这个 Thread 在 Runtime 中的执行上下文由两部分共同构成:

text 复制代码
OS Thread
    +
PyThreadState
    ↓
CPython Runtime 中这个 Thread 的执行上下文
3.1.5.4 GIL

在默认开启 GIL 的 CPython 中,线程必须持有 GIL,才能执行受保护的 Python 解释器代码。

因此,即使 OS Thread 已经获得 CPU,仍需满足 GIL 条件才能继续执行相应的 Python Bytecode。

3.2 分层

3.3 CPython Thread:从生命周期到真实实现

CPython 的 threading.Thread 仍由 OS Scheduler 调度;默认 GIL-enabled 构建还要求获取 GIL,Free-Threaded 构建则改变解释器内部的同步方式。

3.3.1 创建与启动:Created → Runnable

Thread.start() 创建 OS Thread 和对应的 PyThreadState;CPU 执行机会仍需等待 OS Scheduler。

3.3.2 等待调度:Runnable

Python Thread 对应的 OS Thread 直接等待 OS Scheduler;默认构建还需要满足 GIL 执行条件。

3.3.3 真正运行:Runnable → Running

OS Scheduler 让原生线程获得 CPU;在默认 GIL-enabled 构建中,线程取得 GIL 后执行 Python Code。

3.3.4 暂停与等待:Running → Waiting

阻塞 I/O、Lock 或 Condition 等待可能挂起 OS Thread;默认构建的可释放 GIL 路径也会让出解释器执行权。

3.3.5 唤醒与恢复:Waiting → Runnable → Running

等待结束后 OS Thread 重新参与 OS 调度;默认构建进入受保护的 Python 代码前还需重新取得 GIL。

3.3.6 结束与销毁:Running → Terminated

Thread.run() 返回后 Runtime 清理 PyThreadState,原生线程结束,ThreadHandle 完成。

3.4 完整时序图

默认 GIL-enabled CPython 的完整生命周期时序如下:

实现参考 CPython 的 _threadmodule.c、pystate.c 和 ceval_gil.c;Free-Threaded 模式见官方文档。

3.5 回到最开始提到的问题

由 [3.1.5 最终模型](#3.1.5 最终模型) 和 [3.2 分层](#3.2 分层)可知,threading.Thread 对应 Native OS Thread,PyThreadState 保存解释器中的线程状态。因此,CPython 的 Thread 直接由 OS Scheduler 获得 CPU 执行机会,并不存在 Go G-M-P 那样的用户态线程调度器。

由 [3.1.3 GIL](#3.1.3 GIL) 和 [3.3.2 等待调度](#3.3.2 等待调度)、[3.3.3 真正运行](#3.3.3 真正运行) 可知,默认 GIL-enabled 构建中的线程除了等待 OS 调度,还要取得 GIL 才能执行受保护的解释器代码。因此,OS Thread 获得 CPU 不等于立刻可以执行 Python Code;GIL 是执行约束,不是第二个 Thread Scheduler。

再由 [3.3.4 暂停与等待](#3.3.4 暂停与等待)、[3.3.5 唤醒与恢复](#3.3.5 唤醒与恢复) 和 [3.1.4 Free-Threaded](#3.1.4 Free-Threaded) 可知,阻塞等待可能挂起 OS Thread、释放 GIL;Free-Threaded 则改变了 Python Code 的同步条件,却没有改变 threading.Thread 与 OS Thread 的映射。因此,改变的是解释器的并发约束,而非线程的执行载体。

最终回答开篇的问题: CPython 由 OS Scheduler 调度原生线程;默认构建通过 GIL 限制同一解释器内 Python Code 的并行执行,Free-Threaded 构建则允许多个线程并行执行 Python Code,但两者都不将大量 threading.Thread 复用到少量 OS Thread。


4. 三种实现的共同套路

看完 Java、Go 和 CPython,先把具体 API 和 Runtime 名称拿掉,只看三个共同问题:执行单元由谁承载、谁决定它何时运行,以及等待之后如何继续执行?

4.1 执行单元如何获得 CPU?------确定执行载体

先区分两个对象:语言或 Runtime 定义的执行单元 负责承载任务,OS Scheduler 真正调度的则是 OS Thread。它们不一定一对一映射。

假设程序有三个执行单元 A、B、C。最直接的承载方式是各自使用一个 OS Thread:

text 复制代码
执行单元 A → OS Thread 1
执行单元 B → OS Thread 2
执行单元 C → OS Thread 3

但任务可能大部分时间都在等待 I/O。如果每个任务始终独占一个 OS Thread,就需要为大量等待任务保留原生线程资源。

另一种办法是将任务 与承载线程解耦:

text 复制代码
Runnable Task A ─┐
Runnable Task B ─┼─→ Runtime 调度 → OS Thread 1
Runnable Task C ─┘               (分时复用)

同一个 OS Thread 可以在不同时刻执行不同任务;不需要用户态调度器的模型,则直接建立一对一映射。两种方式最终都落到同一条路径:

因此,语言执行单元可以有不同的承载方式,但真正被 OS Scheduler 调度到 CPU 的始终是 OS Thread。

4.2 谁决定下一个执行单元?------区分两层调度

有了承载关系,并不代表任务立刻开始执行,还需要回答两个不同的问题:

  • 任务选择:如果 Runtime 管理可运行的轻量任务,它先选出一个任务交给承载线程。
  • 线程调度:OS Scheduler 从可运行的 OS Thread 中选出一个,让它获得 CPU 时间。

例如只有一个 CPU Core,T1、T2 都处于 Runnable,而 Runtime 已经选择任务 A 交给 T1:

text 复制代码
Runtime: 选中任务 A → 交给 T1

OS Scheduler: T1、T2 都是 Runnable
                      ↓
                    先选 T2
                      ↓
               CPU Core 执行 T2

此时任务 A 虽然已经被 Runtime 选中,仍要等 T1 获得 CPU 才能继续执行。反过来,即使 OS Thread 获得 CPU,语言层的同步条件也可能尚未满足;同步约束不等于另一套线程调度。

直接一对一映射的模型不一定需要 Runtime Thread Scheduler,但 OS Scheduler 仍然存在。对于长时间运行的任务,还可能需要让出或被抢占执行机会,让其他任务获得运行机会。

因此,Runtime 的任务分派和 OS 的线程调度是两个不同层次的选择;Runnable 不等于 Running。

4.3 等待之后如何恢复?------重新获得执行机会

假设任务 A 正在运行,随后发起一次 I/O,而任务 B 已经处于 Runnable。最关键的问题是:A 等待时,承载它的 OS Thread 能不能去执行 B?

text 复制代码
任务 A 运行 → 遇到 I/O 等待
                   │
                   ├─ 只挂起逻辑任务 A
                   │       └─ OS Thread 可执行任务 B
                   │
                   └─ 阻塞承载它的 OS Thread
                           └─ 该线程等待,其他 OS Thread 仍可运行

两条路径都可能出现,取决于执行模型和这次等待能否让出承载资源。挂起逻辑任务不必然阻塞 OS Thread;反过来,阻塞一个 OS Thread 也不意味着整个程序停止。

当 I/O 完成,A 获得的不是"立即执行"的保证,而是重新参与调度的机会:

text 复制代码
Waiting
   ↓ 条件满足 / 唤醒
Runnable
   ↓ 必要时进入 Runtime 队列,等待承载线程
   ↓ OS Thread 获得 CPU
Running

如果等待只挂起逻辑任务,原来的 OS Thread 可以在此期间推进其他任务;如果 OS Thread 自身阻塞,它需要重新就绪并获得 CPU。等待与恢复决定了执行资源能否复用;唤醒不等于马上运行。

由此回到开篇的总问题:语言执行单元首先建立与 OS Thread 的承载关系;可选的 Runtime 调度负责选择任务,OS Scheduler 负责分配 CPU 时间;等待与恢复决定承载资源何时能继续服务其他任务。三个机制串起来,才是执行单元最终在 CPU 上运行的完整路径。


5. 下一篇:Language Memory Model

前面讨论了 CPU 如何执行指令,但多个并发执行单元如何获得 CPU、又如何在等待后恢复执行,仍需要从 Runtime 与 Operating System 层寻找答案,这也是本篇讨论线程调度的原因。

在理解执行过程之后,下一篇将进入 Language Memory Model,讨论并发读写的语义保证,并为后续 Mutex、Atomic、Volatile 等同步机制建立基础。


本文首发于 ThinkerQAQ 的个人博客,由作者本人同步发布。原文可能持续修订,最新版本请以个人博客为准。