上篇回顾 :3.3-01 讲了 Python OOP 的魔法方法与元类,揭秘了 Django Model 的 ORM 魔法。本篇进入 Python 并发------这是 Python 最被误解的领域。GIL 的存在让很多人以为「Python 多线程没用」,但实际情况复杂得多。
一、开篇:一个性能对比实验
爬取 100 个网页,三种方案:
python
# 方案一:串行
def crawl_serial(urls):
results = []
for url in urls:
results.append(requests.get(url).text)
return results
# 耗时:100s(每个 1s)
# 方案二:多线程
def crawl_threading(urls):
with ThreadPoolExecutor(max_workers=10) as executor:
results = list(executor.map(lambda u: requests.get(u).text, urls))
return results
# 耗时:10s(10 个并发)
# 方案三:asyncio
async def crawl_asyncio(urls):
async with aiohttp.ClientSession() as session:
tasks = [fetch(session, url) for url in urls]
return await asyncio.gather(*tasks)
# 耗时:1s(100 个并发,单线程非阻塞 IO)
串行 100s,多线程 10s,asyncio 1s。如果 GIL 让多线程没用,为什么多线程比串行快 10 倍?
答案:GIL 只限制 CPU 计算,不限制 IO 等待。爬虫是 IO 密集型,多线程在 IO 等待时释放 GIL,其他线程可以跑。
二、GIL 的本质
2.1 GIL 是什么
GIL(Global Interpreter Lock)是 CPython 解释器的一把全局锁------同一时刻只有一个线程能执行 Python 字节码。
2.2 为什么有 GIL
CPython 的内存管理(引用计数)不是线程安全的。加锁比做细粒度锁简单得多。Python 1991 年诞生时选了简单方案------一把全局锁。
2.3 GIL 的真实影响范围
| 场景 | GIL 影响 | 并发方案 |
|---|---|---|
| CPU 密集(纯 Python 计算) | 严重(多线程=串行) | 多进程 |
| CPU 密集(C 扩展如 numpy) | 无(扩展内释放 GIL) | 多线程 |
| IO 密集(网络/文件/DB) | 无(IO 等待时释放 GIL) | 多线程/asyncio |
| 混合型 | 部分 | 看瓶颈 |
关键 :GIL 在 IO 等待时会释放------recv()/read()/sleep() 等 IO 操作会让出 GIL。所以 IO 密集场景多线程仍然有效。
2.4 GIL 释放的时机
- IO 操作(网络请求、文件读写)
time.sleep()- 每执行 100 条字节码(Python 2)/ 每 5ms(Python 3.2+)自动切换
- C 扩展主动释放(如 numpy 的矩阵运算)
2.5 Python 3.13 的 GIL 可选
Python 3.13 引入 --no-gil 模式(PEP 703),可以禁用 GIL 做真正的多线程并行。但兼容性还在完善,生产环境建议等 3.14+。
三、多线程:IO 密集场景的选择
3.1 threading 模块
python
import threading
def fetch(url):
return requests.get(url).text
threads = []
for url in urls:
t = threading.Thread(target=fetch, args=(url,))
threads.append(t)
t.start()
for t in threads:
t.join()
3.2 ThreadPoolExecutor(推荐)
python
from concurrent.futures import ThreadPoolExecutor, as_completed
with ThreadPoolExecutor(max_workers=10) as executor:
futures = {executor.submit(fetch, url): url for url in urls}
for future in as_completed(futures):
result = future.result()
# 处理结果
ThreadPoolExecutor 比 threading 更优雅:自动管理线程池、支持 future 取结果、支持上下文管理器自动关闭。
3.3 线程数怎么设
IO 密集型:线程数 = CPU 核数 * (1 + IO时间/CPU时间)。经验值 5~20 倍 CPU 核数。
但 Python 多线程受 GIL 限制,不是越多越好------线程切换也有开销。一般 10~50 个线程够用。
四、多进程:CPU 密集场景的选择
4.1 为什么 CPU 密集用多进程
python
# CPU 密集:计算 100 万个数的平方和
def cpu_task(n):
total = 0
for i in range(n):
total += i ** 2
return total
# 多线程:GIL 限制,= 串行
# 多进程:真正并行
多进程每个进程有独立 GIL,可以真正并行利用多核。
4.2 multiprocessing
python
from multiprocessing import Pool
with Pool(processes=4) as pool:
results = pool.map(cpu_task, [250000] * 4)
4.3 ProcessPoolExecutor
python
from concurrent.futures import ProcessPoolExecutor
with ProcessPoolExecutor(max_workers=4) as executor:
results = list(executor.map(cpu_task, [250000] * 4))
4.4 多进程的代价
| 代价 | 说明 |
|---|---|
| 内存大 | 每个进程独立内存空间 |
| 启动慢 | 进程创建比线程慢 |
| 通信复杂 | 队列/管道/共享内存 |
五、asyncio:单线程高并发
5.1 asyncio 的核心思想
单线程 + 事件循环 + 协程
↓
IO 等待时 → 切到其他协程
↓
IO 完成后 → 切回来继续
与 JS 的事件循环同源------单线程 + 异步回调。但 Python 用 async/await 语法,比 JS 的回调/Promise 更清晰。
5.2 async/await
python
import asyncio
import aiohttp
async def fetch(session, url):
async with session.get(url) as response:
return await response.text()
async def main():
async with aiohttp.ClientSession() as session:
tasks = [fetch(session, url) for url in urls]
results = await asyncio.gather(*tasks)
return results
results = asyncio.run(main())
5.3 asyncio 的事件循环
python
# 事件循环
loop = asyncio.get_event_loop()
# 创建任务
task = loop.create_task(coroutine())
# 等待完成
await task
5.4 asyncio vs JS 事件循环
| 维度 | Python asyncio | JS 事件循环 |
|---|---|---|
| 并发单元 | 协程(coroutine) | 回调/Promise |
| 语法 | async/await | async/await |
| 调度 | 事件循环(asyncio) | 事件循环(V8) |
| 阻塞 | await 才让出 | 自动让出 |
| 生态 | aiohttp/httpx | fetch/axios |
关键差异 :Python 的 asyncio 是显式协作式 ------必须用 await 才让出控制权。JS 是隐式------Promise resolved 后自动调度。所以 Python 的 async 代码如果忘了 await,会变成同步阻塞。
5.5 asyncio 的陷阱
python
# 错误:忘了 await
async def fetch(url):
response = aiohttp.get(url) # 忘了 await,返回 coroutine 而非结果
return response
# 正确
async def fetch(url):
async with aiohttp.ClientSession() as session:
async with session.get(url) as response:
return await response.text()
六、三种方案选择决策
任务类型?
├─ CPU 密集
│ └─ 多进程(multiprocessing)
├─ IO 密集
│ ├─ 并发量 <100 → 多线程(ThreadPoolExecutor)
│ ├─ 并发量 >100 → asyncio
│ └─ 已有同步库 → 多线程(asyncio 要全链路异步)
└─ 混合型
└─ 多进程 + 多线程(每进程内多线程)
6.1 asyncio 的全链路问题
asyncio 要求所有 IO 操作都是异步的 。如果链路中有一个同步阻塞调用(如 requests.get),整个事件循环卡住。
python
# 错误:在 async 函数里用同步 requests
async def fetch(url):
return requests.get(url).text # 阻塞整个事件循环!
# 正确:用 aiohttp
async def fetch(url):
async with aiohttp.ClientSession() as session:
async with session.get(url) as resp:
return await resp.text()
这也是 asyncio 推广难的原因------生态要全异步改造。已有同步代码迁移成本高。
七、线程安全与锁
7.1 GIL 不等于线程安全
GIL 保证字节码执行是原子的,但多步操作不原子:
python
# 不安全
counter = 0
def increment():
global counter
counter += 1 # 读取+加1+写入,三步非原子
# 两个线程同时执行 → counter 可能丢失增量
7.2 用锁保护
python
from threading import Lock
counter = 0
lock = Lock()
def increment():
global counter
with lock:
counter += 1 # 原子操作
7.3 Queue 天然线程安全
python
from queue import Queue
from threading import Thread
q = Queue()
for url in urls:
q.put(url)
def worker():
while True:
url = q.get()
if url is None:
break
fetch(url)
q.task_done()
threads = [Thread(target=worker) for _ in range(10)]
for t in threads:
t.start()
q.join() # 等所有任务完成
八、小结表
| 方案 | 适用 | 优势 | 劣势 |
|---|---|---|---|
| 多线程 | IO 密集、并发量中 | 简单、生态好 | GIL 限制 CPU |
| 多进程 | CPU 密集 | 真正并行 | 内存大、通信复杂 |
| asyncio | IO 密集、并发量高 | 单线程高并发 | 全链路异步 |
