不为失败找理由,只为成功找方法。所有的不甘,因为还心存梦想,所以在你放弃之前,好好拼一把,只怕心老,不怕路长。
python系列之多线程编程
-
- 一、前言:为什么需要多线程?
- 二、基本概念:线程是什么?
- 三、创建你的第一个多线程程序
-
- [3.1 使用threading模块](#3.1 使用threading模块)
- 四、线程的创建方式
-
- [4.1 方法一:使用函数创建线程](#4.1 方法一:使用函数创建线程)
- [4.2 方法二:使用类创建线程](#4.2 方法二:使用类创建线程)
- 五、线程同步:避免"厨房混乱"
-
- [5.1 使用Lock(锁)进行同步](#5.1 使用Lock(锁)进行同步)
- [5.2 使用with语句简化锁操作](#5.2 使用with语句简化锁操作)
- 六、生产者-消费者模式
- 七、线程间通信:Event和Queue
-
- [7.1 使用Event进行线程通信](#7.1 使用Event进行线程通信)
- [7.2 使用Queue进行线程安全的数据交换](#7.2 使用Queue进行线程安全的数据交换)
- 八、线程池:更高效的多线程管理
- 九、GIL(全局解释器锁)的限制
-
- [9.1 CPU密集型 vs I/O密集型](#9.1 CPU密集型 vs I/O密集型)
- 十、实际应用案例
-
- [10.1 多线程下载器](#10.1 多线程下载器)
- [10.2 网页内容抓取](#10.2 网页内容抓取)
- 十一、多线程编程的最佳实践
- 十二、调试多线程程序
- 十三、总结
- 十四、练习挑战
python系列前期章节
- python系列之注释与变量
- python系列之输入输出语句与数据类型
- python系列之运算符
- python系列之控制流程语句
- python系列之字符串
- python系列之列表
- python系列之元组
- python系列之字典
- python系列之集合
- python系列之函数基础
- python系列之函数进阶
- python系统之综合案例1:用python打造智能诗词生成助手
- python系列之综合案例2:用python开发《魔法学院入学考试》文字冒险游戏
- python系列之类与对象:面向对象编程(用造人计划秒懂面向对象)
- python系列之类与对象:python系列之详解面向对象的属性)
- python系列之详解面向对象的函数
- python系列之面向对象的三大特性
- python系列之异常处理:给代码穿上"防弹衣"
- python系列之文件操作:让程序拥有"记忆"的超能力!"
- python系列之综合项目:智能个人任务管理系统"
本文最佳阅读姿势:想象自己是一位餐厅经理,学习如何同时服务多位客人(任务),而不是让客人排长队等待!
一、前言:为什么需要多线程?
想象一下:你在餐厅点餐,厨师却坚持要做完第一道菜才做第二道菜。后面的客人会饿疯的!
程序也是如此。单线程就像只有一个厨师的餐厅,而多线程就像有多个厨师同时工作:
- 单线程:任务一个接一个执行
- 多线程:多个任务同时进行(看起来同时,实际是快速切换)
二、基本概念:线程是什么?
线程是程序中的执行流程。一个进程(程序)可以有多个线程,共享相同的内存空间。
简单比喻:
- 进程 = 整个餐厅
- 线程 = 餐厅里的多个厨师
三、创建你的第一个多线程程序
3.1 使用threading模块
Python通过threading模块提供多线程支持:
python
import threading
import time
# 定义一个简单的任务函数
def print_numbers():
for i in range(1, 6):
time.sleep(1) # 模拟耗时操作
print(f"数字: {i}")
def print_letters():
for letter in ['A', 'B', 'C', 'D', 'E']:
time.sleep(1.5) # 模拟耗时操作
print(f"字母: {letter}")
# 单线程执行 - 顺序执行
print("=== 单线程执行 ===")
start_time = time.time()
print_numbers()
print_letters()
end_time = time.time()
print(f"单线程耗时: {end_time - start_time:.2f}秒")
# 多线程执行 - 同时执行
print("\n=== 多线程执行 ===")
start_time = time.time()
# 创建线程对象
thread1 = threading.Thread(target=print_numbers)
thread2 = threading.Thread(target=print_letters)
# 启动线程
thread1.start()
thread2.start()
# 等待线程结束
thread1.join()
thread2.join()
end_time = time.time()
print(f"多线程耗时: {end_time - start_time:.2f}秒")
运行这个例子,你会发现多线程版本花费的时间更少!
四、线程的创建方式
4.1 方法一:使用函数创建线程
python
import threading
def worker(task_name, duration):
print(f"{task_name} 开始执行")
time.sleep(duration)
print(f"{task_name} 完成,耗时 {duration}秒")
# 创建多个线程
threads = []
tasks = [("任务1", 2), ("任务2", 3), ("任务3", 1)]
for task_name, duration in tasks:
thread = threading.Thread(target=worker, args=(task_name, duration))
threads.append(thread)
thread.start()
# 等待所有线程完成
for thread in threads:
thread.join()
print("所有任务完成!")
4.2 方法二:使用类创建线程
python
import threading
class MyThread(threading.Thread):
def __init__(self, task_name, duration):
super().__init__()
self.task_name = task_name
self.duration = duration
def run(self):
"""线程运行时执行的方法"""
print(f"{self.task_name} 开始执行")
time.sleep(self.duration)
print(f"{self.task_name} 完成,耗时 {self.duration}秒")
# 使用自定义线程类
threads = []
for i in range(3):
thread = MyThread(f"任务{i+1}", i+1)
threads.append(thread)
thread.start()
for thread in threads:
thread.join()
print("所有任务完成!")
五、线程同步:避免"厨房混乱"
当多个线程同时访问共享资源时,可能会出现问题。就像多个厨师同时使用同一个锅!
5.1 使用Lock(锁)进行同步
python
import threading
# 共享资源
counter = 0
lock = threading.Lock()
def increment_counter():
global counter
for _ in range(100000):
# 获取锁
lock.acquire()
try:
counter += 1 # 临界区代码
finally:
# 释放锁
lock.release()
# 创建多个线程增加计数器
threads = []
for _ in range(5):
thread = threading.Thread(target=increment_counter)
threads.append(thread)
thread.start()
for thread in threads:
thread.join()
print(f"最终计数器值: {counter}") # 应该是500000
5.2 使用with语句简化锁操作
python
def increment_counter_safe():
global counter
for _ in range(100000):
with lock: # 自动获取和释放锁
counter += 1
# 测试安全的计数器增加
counter = 0
threads = []
for _ in range(5):
thread = threading.Thread(target=increment_counter_safe)
threads.append(thread)
thread.start()
for thread in threads:
thread.join()
print(f"安全的最终计数器值: {counter}") # 应该是500000
六、生产者-消费者模式
这是一个经典的多线程模式,就像餐厅的厨师(生产者)和服务员(消费者)的关系:
python
import threading
import time
import random
from collections import deque
# 共享的队列
queue = deque()
max_size = 5
lock = threading.Lock()
# 条件变量用于线程间通信
condition = threading.Condition(lock)
def producer(name):
"""生产者函数"""
for i in range(10):
time.sleep(random.uniform(0.1, 0.5)) # 模拟生产时间
with condition:
# 如果队列已满,等待
while len(queue) >= max_size:
print(f"{name} 等待:队列已满")
condition.wait()
# 生产物品
item = f"物品-{name}-{i}"
queue.append(item)
print(f"{name} 生产了: {item} (队列大小: {len(queue)})")
# 通知消费者
condition.notify_all()
def consumer(name):
"""消费者函数"""
for i in range(10):
time.sleep(random.uniform(0.2, 0.7)) # 模拟消费时间
with condition:
# 如果队列为空,等待
while len(queue) == 0:
print(f"{name} 等待:队列为空")
condition.wait()
# 消费物品
item = queue.popleft()
print(f"{name} 消费了: {item} (队列大小: {len(queue)})")
# 通知生产者
condition.notify_all()
# 创建生产者和消费者线程
producers = [threading.Thread(target=producer, args=(f"生产者-{i}",)) for i in range(2)]
consumers = [threading.Thread(target=consumer, args=(f"消费者-{i}",)) for i in range(3)]
# 启动所有线程
for p in producers:
p.start()
for c in consumers:
c.start()
# 等待所有线程完成
for p in producers:
p.join()
for c in consumers:
c.join()
print("生产-消费过程结束")
七、线程间通信:Event和Queue
7.1 使用Event进行线程通信
python
import threading
# 创建事件对象
start_event = threading.Event()
stop_event = threading.Event()
def worker(name):
print(f"{name} 等待开始信号...")
start_event.wait() # 等待事件信号
while not stop_event.is_set():
print(f"{name} 正在工作...")
time.sleep(1)
print(f"{name} 收到停止信号,结束工作")
# 创建工作线程
threads = []
for i in range(3):
thread = threading.Thread(target=worker, args=(f"工人-{i}",))
threads.append(thread)
thread.start()
# 主线程控制
time.sleep(2)
print("主线程:发送开始信号!")
start_event.set() # 设置事件,唤醒所有等待的线程
time.sleep(3)
print("主线程:发送停止信号!")
stop_event.set() # 设置停止事件
for thread in threads:
thread.join()
print("所有工作线程已结束")
7.2 使用Queue进行线程安全的数据交换
python
import threading
import queue
import time
import random
# 创建线程安全的队列
task_queue = queue.Queue()
result_queue = queue.Queue()
def task_producer():
"""任务生产者"""
for i in range(10):
task = f"任务-{i}"
task_queue.put(task)
print(f"生产了: {task}")
time.sleep(random.uniform(0.1, 0.3))
# 发送结束信号
for _ in range(3): # 有3个消费者
task_queue.put(None)
def task_consumer(name):
"""任务消费者"""
while True:
task = task_queue.get()
if task is None: # 收到结束信号
task_queue.put(None) # 让其他消费者也能收到
break
# 处理任务
print(f"{name} 开始处理: {task}")
time.sleep(random.uniform(0.5, 1.5))
result = f"{task}的处理结果"
result_queue.put(result)
print(f"{name} 完成了: {task}")
task_queue.task_done() # 标记任务完成
# 创建生产者和消费者
producer_thread = threading.Thread(target=task_producer)
consumer_threads = [
threading.Thread(target=task_consumer, args=(f"消费者-{i}",))
for i in range(3)
]
# 启动所有线程
producer_thread.start()
for thread in consumer_threads:
thread.start()
# 等待生产者完成
producer_thread.join()
# 等待所有任务完成
task_queue.join()
# 等待消费者完成
for thread in consumer_threads:
thread.join()
# 输出结果
print("\n=== 处理结果 ===")
while not result_queue.empty():
print(result_queue.get())
八、线程池:更高效的多线程管理
使用concurrent.futures模块可以更方便地管理线程:
python
import concurrent.futures
import time
import random
def process_task(task_id):
"""模拟处理任务"""
print(f"开始处理任务 {task_id}")
processing_time = random.uniform(0.5, 2.0)
time.sleep(processing_time)
print(f"完成任务 {task_id}, 耗时 {processing_time:.2f}秒")
return f"任务{task_id}结果"
# 使用ThreadPoolExecutor
with concurrent.futures.ThreadPoolExecutor(max_workers=3) as executor:
# 提交任务
future_to_task = {
executor.submit(process_task, i): i
for i in range(10)
}
# 获取结果
for future in concurrent.futures.as_completed(future_to_task):
task_id = future_to_task[future]
try:
result = future.result()
print(f"收到结果: {result}")
except Exception as e:
print(f"任务 {task_id} 产生异常: {e}")
print("所有任务处理完成!")
九、GIL(全局解释器锁)的限制
Python有一个著名的GIL(Global Interpreter Lock),它限制同一时间只能有一个线程执行Python字节码。
重要理解:
- GIL不是Python语言的特性,而是CPython解释器的实现细节
- GIL主要影响CPU密集型任务
- I/O密集型任务仍然能从多线程中受益
9.1 CPU密集型 vs I/O密集型
python
import threading
import time
import math
# CPU密集型任务
def cpu_intensive_task(n):
print(f"开始CPU密集型任务 {n}")
result = 0
for i in range(10**7):
result += math.sqrt(i)
print(f"完成CPU密集型任务 {n}")
return result
# I/O密集型任务
def io_intensive_task(n):
print(f"开始I/O密集型任务 {n}")
time.sleep(2) # 模拟I/O操作
print(f"完成I/O密集型任务 {n}")
return f"任务{n}结果"
# 测试CPU密集型任务
print("=== CPU密集型任务测试 ===")
start_time = time.time()
threads = []
for i in range(3):
thread = threading.Thread(target=cpu_intensive_task, args=(i,))
threads.append(thread)
thread.start()
for thread in threads:
thread.join()
cpu_time = time.time() - start_time
print(f"CPU密集型任务总耗时: {cpu_time:.2f}秒")
# 测试I/O密集型任务
print("\n=== I/O密集型任务测试 ===")
start_time = time.time()
threads = []
for i in range(3):
thread = threading.Thread(target=io_intensive_task, args=(i,))
threads.append(thread)
thread.start()
for thread in threads:
thread.join()
io_time = time.time() - start_time
print(f"I/O密集型任务总耗时: {io_time:.2f}秒")
你会发现I/O密集型任务从多线程中获益更多!
十、实际应用案例
10.1 多线程下载器
python
import threading
import requests
import os
from urllib.parse import urlparse
class Downloader:
def __init__(self, max_workers=5):
self.max_workers = max_workers
self.lock = threading.Lock()
def download_file(self, url, save_path):
"""下载单个文件"""
try:
response = requests.get(url, stream=True)
response.raise_for_status()
# 创建保存目录
os.makedirs(os.path.dirname(save_path), exist_ok=True)
with open(save_path, 'wb') as file:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
file.write(chunk)
with self.lock:
print(f"下载完成: {os.path.basename(save_path)}")
return True
except Exception as e:
with self.lock:
print(f"下载失败 {url}: {e}")
return False
def download_files(self, url_list, save_dir):
"""多线程下载多个文件"""
with concurrent.futures.ThreadPoolExecutor(max_workers=self.max_workers) as executor:
# 准备下载任务
future_to_url = {}
for url in url_list:
filename = os.path.basename(urlparse(url).path)
save_path = os.path.join(save_dir, filename)
future = executor.submit(self.download_file, url, save_path)
future_to_url[future] = url
# 处理结果
success_count = 0
for future in concurrent.futures.as_completed(future_to_url):
url = future_to_url[future]
if future.result():
success_count += 1
print(f"\n下载完成: {success_count}/{len(url_list)} 个文件成功")
# 使用示例
if __name__ == "__main__":
# 示例URL列表(需要替换为真实的URL)
urls = [
"https://example.com/file1.jpg",
"https://example.com/file2.jpg",
"https://example.com/file3.jpg",
# 添加更多URL...
]
downloader = Downloader(max_workers=3)
downloader.download_files(urls, "downloads")
10.2 网页内容抓取
python
import threading
import requests
from bs4 import BeautifulSoup
import time
class WebCrawler:
def __init__(self, max_threads=5):
self.max_threads = max_threads
self.visited = set()
self.lock = threading.Lock()
self.results = []
def crawl_page(self, url, depth=0, max_depth=2):
"""抓取单个页面"""
if depth > max_depth or url in self.visited:
return
with self.lock:
self.visited.add(url)
try:
print(f"抓取: {url} (深度: {depth})")
response = requests.get(url, timeout=5)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')
title = soup.title.string if soup.title else "无标题"
# 提取页面信息
page_info = {
'url': url,
'title': title.strip(),
'depth': depth,
'links': []
}
# 提取链接(用于进一步抓取)
if depth < max_depth:
for link in soup.find_all('a', href=True):
href = link['href']
if href.startswith('http'):
page_info['links'].append(href)
with self.lock:
self.results.append(page_info)
print(f"完成: {url} - 找到 {len(page_info['links'])} 个链接")
return page_info
except Exception as e:
with self.lock:
print(f"抓取失败 {url}: {e}")
return None
def start_crawling(self, start_url, max_depth=2):
"""开始多线程爬取"""
from concurrent.futures import ThreadPoolExecutor
to_crawl = [start_url]
current_depth = 0
with ThreadPoolExecutor(max_workers=self.max_threads) as executor:
while to_crawl and current_depth <= max_depth:
print(f"\n深度 {current_depth}: 处理 {len(to_crawl)} 个URL")
# 提交当前深度的所有URL
future_to_url = {
executor.submit(self.crawl_page, url, current_depth, max_depth): url
for url in to_crawl
}
# 收集下一层要抓取的URL
to_crawl = []
for future in concurrent.futures.as_completed(future_to_url):
result = future.result()
if result and current_depth < max_depth:
to_crawl.extend(result['links'])
current_depth += 1
print(f"\n爬取完成! 总共访问了 {len(self.visited)} 个页面")
return self.results
# 使用示例
if __name__ == "__main__":
crawler = WebCrawler(max_threads=3)
results = crawler.start_crawling("https://example.com", max_depth=1)
print("\n=== 抓取结果 ===")
for result in results:
print(f"{result['url']} - {result['title']}")
十一、多线程编程的最佳实践
- 避免过度使用线程:线程不是越多越好,创建和管理线程有开销
- 使用线程池:对于大量短期任务,使用线程池更高效
- 注意线程安全:对共享资源使用适当的同步机制
- 避免死锁:按固定顺序获取锁,使用超时机制
- 使用队列通信:优先使用队列而不是共享变量进行线程间通信
- 处理异常:线程中的异常不会传播到主线程,需要在线程内部处理
- 考虑GIL的影响:对于CPU密集型任务,考虑使用多进程 instead
十二、调试多线程程序
调试多线程程序可能比较困难,以下是一些技巧:
python
import threading
import logging
# 配置日志以便调试
logging.basicConfig(
level=logging.DEBUG,
format='%(asctime)s - %(threadName)s - %(levelname)s - %(message)s'
)
def debug_example():
"""调试示例"""
logging.debug("线程启动")
# 获取当前线程信息
current_thread = threading.current_thread()
logging.info(f"线程名称: {current_thread.name}")
logging.info(f"线程ID: {current_thread.ident}")
# 模拟工作
try:
logging.debug("开始工作")
# ... 一些代码
logging.debug("工作完成")
except Exception as e:
logging.error(f"工作出错: {e}")
logging.debug("线程结束")
# 创建并启动线程
thread = threading.Thread(target=debug_example, name="工作线程")
thread.start()
thread.join()
十三、总结
通过本文,你已经学习了:
- 多线程基础:线程的概念和创建方式
- 线程同步:使用Lock、Condition等同步机制
- 线程通信:使用Event、Queue进行线程间通信
- 线程池:使用ThreadPoolExecutor管理线程
- GIL限制:理解Python的全局解释器锁
- 实际应用:文件下载、网页抓取等实际案例
- 最佳实践:多线程编程的注意事项和技巧
记住:多线程是一把双刃剑。正确使用时可以提高程序性能,但使用不当会导致难以调试的问题。始终优先考虑代码的清晰性和正确性,只有在确实需要时才使用多线程。
十四、练习挑战
- 修改下载器程序,添加下载进度显示功能
- 实现一个多线程的图片缩略图生成器
- 创建一个多线程的日志分析工具
- 实现一个生产者-消费者模式的数据处理管道
当你完成这些练习,恭喜你!你已经掌握了Python多线程编程的核心技能!🎉
现在去创建高效、并发的Python应用程序吧!