python系列之多线程编程

不为失败找理由,只为成功找方法。所有的不甘,因为还心存梦想,所以在你放弃之前,好好拼一把,只怕心老,不怕路长。

python系列之多线程编程

python系列前期章节

  1. python系列之注释与变量
  2. python系列之输入输出语句与数据类型
  3. python系列之运算符
  4. python系列之控制流程语句
  5. python系列之字符串
  6. python系列之列表
  7. python系列之元组
  8. python系列之字典
  9. python系列之集合
  10. python系列之函数基础
  11. python系列之函数进阶
  12. python系统之综合案例1:用python打造智能诗词生成助手
  13. python系列之综合案例2:用python开发《魔法学院入学考试》文字冒险游戏
  14. python系列之类与对象:面向对象编程(用造人计划秒懂面向对象)
  15. python系列之类与对象:python系列之详解面向对象的属性)
  16. python系列之详解面向对象的函数
  17. python系列之面向对象的三大特性
  18. python系列之异常处理:给代码穿上"防弹衣"
  19. python系列之文件操作:让程序拥有"记忆"的超能力!"
  20. python系列之综合项目:智能个人任务管理系统"

本文最佳阅读姿势:想象自己是一位餐厅经理,学习如何同时服务多位客人(任务),而不是让客人排长队等待!

一、前言:为什么需要多线程?

想象一下:你在餐厅点餐,厨师却坚持要做完第一道菜才做第二道菜。后面的客人会饿疯的!

程序也是如此。单线程就像只有一个厨师的餐厅,而多线程就像有多个厨师同时工作:

  • 单线程:任务一个接一个执行
  • 多线程:多个任务同时进行(看起来同时,实际是快速切换)

二、基本概念:线程是什么?

线程是程序中的执行流程。一个进程(程序)可以有多个线程,共享相同的内存空间。

简单比喻:

  • 进程 = 整个餐厅
  • 线程 = 餐厅里的多个厨师

三、创建你的第一个多线程程序

3.1 使用threading模块

Python通过threading模块提供多线程支持:

python 复制代码
import threading
import time

# 定义一个简单的任务函数
def print_numbers():
    for i in range(1, 6):
        time.sleep(1)  # 模拟耗时操作
        print(f"数字: {i}")

def print_letters():
    for letter in ['A', 'B', 'C', 'D', 'E']:
        time.sleep(1.5)  # 模拟耗时操作
        print(f"字母: {letter}")

# 单线程执行 - 顺序执行
print("=== 单线程执行 ===")
start_time = time.time()
print_numbers()
print_letters()
end_time = time.time()
print(f"单线程耗时: {end_time - start_time:.2f}秒")

# 多线程执行 - 同时执行
print("\n=== 多线程执行 ===")
start_time = time.time()

# 创建线程对象
thread1 = threading.Thread(target=print_numbers)
thread2 = threading.Thread(target=print_letters)

# 启动线程
thread1.start()
thread2.start()

# 等待线程结束
thread1.join()
thread2.join()

end_time = time.time()
print(f"多线程耗时: {end_time - start_time:.2f}秒")

运行这个例子,你会发现多线程版本花费的时间更少!

四、线程的创建方式

4.1 方法一:使用函数创建线程

python 复制代码
import threading

def worker(task_name, duration):
    print(f"{task_name} 开始执行")
    time.sleep(duration)
    print(f"{task_name} 完成,耗时 {duration}秒")

# 创建多个线程
threads = []
tasks = [("任务1", 2), ("任务2", 3), ("任务3", 1)]

for task_name, duration in tasks:
    thread = threading.Thread(target=worker, args=(task_name, duration))
    threads.append(thread)
    thread.start()

# 等待所有线程完成
for thread in threads:
    thread.join()

print("所有任务完成!")

4.2 方法二:使用类创建线程

python 复制代码
import threading

class MyThread(threading.Thread):
    def __init__(self, task_name, duration):
        super().__init__()
        self.task_name = task_name
        self.duration = duration
    
    def run(self):
        """线程运行时执行的方法"""
        print(f"{self.task_name} 开始执行")
        time.sleep(self.duration)
        print(f"{self.task_name} 完成,耗时 {self.duration}秒")

# 使用自定义线程类
threads = []
for i in range(3):
    thread = MyThread(f"任务{i+1}", i+1)
    threads.append(thread)
    thread.start()

for thread in threads:
    thread.join()

print("所有任务完成!")

五、线程同步:避免"厨房混乱"

当多个线程同时访问共享资源时,可能会出现问题。就像多个厨师同时使用同一个锅!

5.1 使用Lock(锁)进行同步

python 复制代码
import threading

# 共享资源
counter = 0
lock = threading.Lock()

def increment_counter():
    global counter
    for _ in range(100000):
        # 获取锁
        lock.acquire()
        try:
            counter += 1  # 临界区代码
        finally:
            # 释放锁
            lock.release()

# 创建多个线程增加计数器
threads = []
for _ in range(5):
    thread = threading.Thread(target=increment_counter)
    threads.append(thread)
    thread.start()

for thread in threads:
    thread.join()

print(f"最终计数器值: {counter}")  # 应该是500000

5.2 使用with语句简化锁操作

python 复制代码
def increment_counter_safe():
    global counter
    for _ in range(100000):
        with lock:  # 自动获取和释放锁
            counter += 1

# 测试安全的计数器增加
counter = 0
threads = []

for _ in range(5):
    thread = threading.Thread(target=increment_counter_safe)
    threads.append(thread)
    thread.start()

for thread in threads:
    thread.join()

print(f"安全的最终计数器值: {counter}")  # 应该是500000

六、生产者-消费者模式

这是一个经典的多线程模式,就像餐厅的厨师(生产者)和服务员(消费者)的关系:

python 复制代码
import threading
import time
import random
from collections import deque

# 共享的队列
queue = deque()
max_size = 5
lock = threading.Lock()
# 条件变量用于线程间通信
condition = threading.Condition(lock)

def producer(name):
    """生产者函数"""
    for i in range(10):
        time.sleep(random.uniform(0.1, 0.5))  # 模拟生产时间
        
        with condition:
            # 如果队列已满,等待
            while len(queue) >= max_size:
                print(f"{name} 等待:队列已满")
                condition.wait()
            
            # 生产物品
            item = f"物品-{name}-{i}"
            queue.append(item)
            print(f"{name} 生产了: {item} (队列大小: {len(queue)})")
            
            # 通知消费者
            condition.notify_all()

def consumer(name):
    """消费者函数"""
    for i in range(10):
        time.sleep(random.uniform(0.2, 0.7))  # 模拟消费时间
        
        with condition:
            # 如果队列为空,等待
            while len(queue) == 0:
                print(f"{name} 等待:队列为空")
                condition.wait()
            
            # 消费物品
            item = queue.popleft()
            print(f"{name} 消费了: {item} (队列大小: {len(queue)})")
            
            # 通知生产者
            condition.notify_all()

# 创建生产者和消费者线程
producers = [threading.Thread(target=producer, args=(f"生产者-{i}",)) for i in range(2)]
consumers = [threading.Thread(target=consumer, args=(f"消费者-{i}",)) for i in range(3)]

# 启动所有线程
for p in producers:
    p.start()
for c in consumers:
    c.start()

# 等待所有线程完成
for p in producers:
    p.join()
for c in consumers:
    c.join()

print("生产-消费过程结束")

七、线程间通信:Event和Queue

7.1 使用Event进行线程通信

python 复制代码
import threading

# 创建事件对象
start_event = threading.Event()
stop_event = threading.Event()

def worker(name):
    print(f"{name} 等待开始信号...")
    start_event.wait()  # 等待事件信号
    
    while not stop_event.is_set():
        print(f"{name} 正在工作...")
        time.sleep(1)
    
    print(f"{name} 收到停止信号,结束工作")

# 创建工作线程
threads = []
for i in range(3):
    thread = threading.Thread(target=worker, args=(f"工人-{i}",))
    threads.append(thread)
    thread.start()

# 主线程控制
time.sleep(2)
print("主线程:发送开始信号!")
start_event.set()  # 设置事件,唤醒所有等待的线程

time.sleep(3)
print("主线程:发送停止信号!")
stop_event.set()   # 设置停止事件

for thread in threads:
    thread.join()

print("所有工作线程已结束")

7.2 使用Queue进行线程安全的数据交换

python 复制代码
import threading
import queue
import time
import random

# 创建线程安全的队列
task_queue = queue.Queue()
result_queue = queue.Queue()

def task_producer():
    """任务生产者"""
    for i in range(10):
        task = f"任务-{i}"
        task_queue.put(task)
        print(f"生产了: {task}")
        time.sleep(random.uniform(0.1, 0.3))
    
    # 发送结束信号
    for _ in range(3):  # 有3个消费者
        task_queue.put(None)

def task_consumer(name):
    """任务消费者"""
    while True:
        task = task_queue.get()
        
        if task is None:  # 收到结束信号
            task_queue.put(None)  # 让其他消费者也能收到
            break
        
        # 处理任务
        print(f"{name} 开始处理: {task}")
        time.sleep(random.uniform(0.5, 1.5))
        result = f"{task}的处理结果"
        result_queue.put(result)
        print(f"{name} 完成了: {task}")
        
        task_queue.task_done()  # 标记任务完成

# 创建生产者和消费者
producer_thread = threading.Thread(target=task_producer)
consumer_threads = [
    threading.Thread(target=task_consumer, args=(f"消费者-{i}",))
    for i in range(3)
]

# 启动所有线程
producer_thread.start()
for thread in consumer_threads:
    thread.start()

# 等待生产者完成
producer_thread.join()

# 等待所有任务完成
task_queue.join()

# 等待消费者完成
for thread in consumer_threads:
    thread.join()

# 输出结果
print("\n=== 处理结果 ===")
while not result_queue.empty():
    print(result_queue.get())

八、线程池:更高效的多线程管理

使用concurrent.futures模块可以更方便地管理线程:

python 复制代码
import concurrent.futures
import time
import random

def process_task(task_id):
    """模拟处理任务"""
    print(f"开始处理任务 {task_id}")
    processing_time = random.uniform(0.5, 2.0)
    time.sleep(processing_time)
    print(f"完成任务 {task_id}, 耗时 {processing_time:.2f}秒")
    return f"任务{task_id}结果"

# 使用ThreadPoolExecutor
with concurrent.futures.ThreadPoolExecutor(max_workers=3) as executor:
    # 提交任务
    future_to_task = {
        executor.submit(process_task, i): i 
        for i in range(10)
    }
    
    # 获取结果
    for future in concurrent.futures.as_completed(future_to_task):
        task_id = future_to_task[future]
        try:
            result = future.result()
            print(f"收到结果: {result}")
        except Exception as e:
            print(f"任务 {task_id} 产生异常: {e}")

print("所有任务处理完成!")

九、GIL(全局解释器锁)的限制

Python有一个著名的GIL(Global Interpreter Lock),它限制同一时间只能有一个线程执行Python字节码。

重要理解:

  • GIL不是Python语言的特性,而是CPython解释器的实现细节
  • GIL主要影响CPU密集型任务
  • I/O密集型任务仍然能从多线程中受益

9.1 CPU密集型 vs I/O密集型

python 复制代码
import threading
import time
import math

# CPU密集型任务
def cpu_intensive_task(n):
    print(f"开始CPU密集型任务 {n}")
    result = 0
    for i in range(10**7):
        result += math.sqrt(i)
    print(f"完成CPU密集型任务 {n}")
    return result

# I/O密集型任务
def io_intensive_task(n):
    print(f"开始I/O密集型任务 {n}")
    time.sleep(2)  # 模拟I/O操作
    print(f"完成I/O密集型任务 {n}")
    return f"任务{n}结果"

# 测试CPU密集型任务
print("=== CPU密集型任务测试 ===")
start_time = time.time()

threads = []
for i in range(3):
    thread = threading.Thread(target=cpu_intensive_task, args=(i,))
    threads.append(thread)
    thread.start()

for thread in threads:
    thread.join()

cpu_time = time.time() - start_time
print(f"CPU密集型任务总耗时: {cpu_time:.2f}秒")

# 测试I/O密集型任务
print("\n=== I/O密集型任务测试 ===")
start_time = time.time()

threads = []
for i in range(3):
    thread = threading.Thread(target=io_intensive_task, args=(i,))
    threads.append(thread)
    thread.start()

for thread in threads:
    thread.join()

io_time = time.time() - start_time
print(f"I/O密集型任务总耗时: {io_time:.2f}秒")

你会发现I/O密集型任务从多线程中获益更多!

十、实际应用案例

10.1 多线程下载器

python 复制代码
import threading
import requests
import os
from urllib.parse import urlparse

class Downloader:
    def __init__(self, max_workers=5):
        self.max_workers = max_workers
        self.lock = threading.Lock()
    
    def download_file(self, url, save_path):
        """下载单个文件"""
        try:
            response = requests.get(url, stream=True)
            response.raise_for_status()
            
            # 创建保存目录
            os.makedirs(os.path.dirname(save_path), exist_ok=True)
            
            with open(save_path, 'wb') as file:
                for chunk in response.iter_content(chunk_size=8192):
                    if chunk:
                        file.write(chunk)
            
            with self.lock:
                print(f"下载完成: {os.path.basename(save_path)}")
                
            return True
        except Exception as e:
            with self.lock:
                print(f"下载失败 {url}: {e}")
            return False
    
    def download_files(self, url_list, save_dir):
        """多线程下载多个文件"""
        with concurrent.futures.ThreadPoolExecutor(max_workers=self.max_workers) as executor:
            # 准备下载任务
            future_to_url = {}
            for url in url_list:
                filename = os.path.basename(urlparse(url).path)
                save_path = os.path.join(save_dir, filename)
                future = executor.submit(self.download_file, url, save_path)
                future_to_url[future] = url
            
            # 处理结果
            success_count = 0
            for future in concurrent.futures.as_completed(future_to_url):
                url = future_to_url[future]
                if future.result():
                    success_count += 1
            
            print(f"\n下载完成: {success_count}/{len(url_list)} 个文件成功")

# 使用示例
if __name__ == "__main__":
    # 示例URL列表(需要替换为真实的URL)
    urls = [
        "https://example.com/file1.jpg",
        "https://example.com/file2.jpg",
        "https://example.com/file3.jpg",
        # 添加更多URL...
    ]
    
    downloader = Downloader(max_workers=3)
    downloader.download_files(urls, "downloads")

10.2 网页内容抓取

python 复制代码
import threading
import requests
from bs4 import BeautifulSoup
import time

class WebCrawler:
    def __init__(self, max_threads=5):
        self.max_threads = max_threads
        self.visited = set()
        self.lock = threading.Lock()
        self.results = []
    
    def crawl_page(self, url, depth=0, max_depth=2):
        """抓取单个页面"""
        if depth > max_depth or url in self.visited:
            return
        
        with self.lock:
            self.visited.add(url)
        
        try:
            print(f"抓取: {url} (深度: {depth})")
            response = requests.get(url, timeout=5)
            response.raise_for_status()
            
            soup = BeautifulSoup(response.text, 'html.parser')
            title = soup.title.string if soup.title else "无标题"
            
            # 提取页面信息
            page_info = {
                'url': url,
                'title': title.strip(),
                'depth': depth,
                'links': []
            }
            
            # 提取链接(用于进一步抓取)
            if depth < max_depth:
                for link in soup.find_all('a', href=True):
                    href = link['href']
                    if href.startswith('http'):
                        page_info['links'].append(href)
            
            with self.lock:
                self.results.append(page_info)
                print(f"完成: {url} - 找到 {len(page_info['links'])} 个链接")
            
            return page_info
            
        except Exception as e:
            with self.lock:
                print(f"抓取失败 {url}: {e}")
            return None
    
    def start_crawling(self, start_url, max_depth=2):
        """开始多线程爬取"""
        from concurrent.futures import ThreadPoolExecutor
        
        to_crawl = [start_url]
        current_depth = 0
        
        with ThreadPoolExecutor(max_workers=self.max_threads) as executor:
            while to_crawl and current_depth <= max_depth:
                print(f"\n深度 {current_depth}: 处理 {len(to_crawl)} 个URL")
                
                # 提交当前深度的所有URL
                future_to_url = {
                    executor.submit(self.crawl_page, url, current_depth, max_depth): url
                    for url in to_crawl
                }
                
                # 收集下一层要抓取的URL
                to_crawl = []
                for future in concurrent.futures.as_completed(future_to_url):
                    result = future.result()
                    if result and current_depth < max_depth:
                        to_crawl.extend(result['links'])
                
                current_depth += 1
        
        print(f"\n爬取完成! 总共访问了 {len(self.visited)} 个页面")
        return self.results
        # 使用示例
if __name__ == "__main__":
    crawler = WebCrawler(max_threads=3)
    results = crawler.start_crawling("https://example.com", max_depth=1)
    
    print("\n=== 抓取结果 ===")
    for result in results:
        print(f"{result['url']} - {result['title']}")

十一、多线程编程的最佳实践

  1. 避免过度使用线程:线程不是越多越好,创建和管理线程有开销
  2. 使用线程池:对于大量短期任务,使用线程池更高效
  3. 注意线程安全:对共享资源使用适当的同步机制
  4. 避免死锁:按固定顺序获取锁,使用超时机制
  5. 使用队列通信:优先使用队列而不是共享变量进行线程间通信
  6. 处理异常:线程中的异常不会传播到主线程,需要在线程内部处理
  7. 考虑GIL的影响:对于CPU密集型任务,考虑使用多进程 instead

十二、调试多线程程序

调试多线程程序可能比较困难,以下是一些技巧:

python 复制代码
import threading
import logging

# 配置日志以便调试
logging.basicConfig(
    level=logging.DEBUG,
    format='%(asctime)s - %(threadName)s - %(levelname)s - %(message)s'
)

def debug_example():
    """调试示例"""
    logging.debug("线程启动")
    
    # 获取当前线程信息
    current_thread = threading.current_thread()
    logging.info(f"线程名称: {current_thread.name}")
    logging.info(f"线程ID: {current_thread.ident}")
    
    # 模拟工作
    try:
        logging.debug("开始工作")
        # ... 一些代码
        logging.debug("工作完成")
    except Exception as e:
        logging.error(f"工作出错: {e}")
    
    logging.debug("线程结束")

# 创建并启动线程
thread = threading.Thread(target=debug_example, name="工作线程")
thread.start()
thread.join()

十三、总结

通过本文,你已经学习了:

  1. 多线程基础:线程的概念和创建方式
  2. 线程同步:使用Lock、Condition等同步机制
  3. 线程通信:使用Event、Queue进行线程间通信
  4. 线程池:使用ThreadPoolExecutor管理线程
  5. GIL限制:理解Python的全局解释器锁
  6. 实际应用:文件下载、网页抓取等实际案例
  7. 最佳实践:多线程编程的注意事项和技巧

记住:多线程是一把双刃剑。正确使用时可以提高程序性能,但使用不当会导致难以调试的问题。始终优先考虑代码的清晰性和正确性,只有在确实需要时才使用多线程。

十四、练习挑战

  1. 修改下载器程序,添加下载进度显示功能
  2. 实现一个多线程的图片缩略图生成器
  3. 创建一个多线程的日志分析工具
  4. 实现一个生产者-消费者模式的数据处理管道

当你完成这些练习,恭喜你!你已经掌握了Python多线程编程的核心技能!🎉

现在去创建高效、并发的Python应用程序吧!

相关推荐
梅雅达编程笔记1 小时前
03_网页表格数据抓取
爬虫·python·beautifulsoup·pandas·数据采集
迅猛龙办公室1 小时前
Python实现求两个数字之和
java·开发语言·python
企业数字化笔记1 小时前
视频成片怎么稳定导出?编码选择、GPU/CPU降级与媒体交付验收
java·python·ffmpeg
szial2 小时前
网络爬虫与 CDP 实战(一):网页数据到底在哪?从 HTTP 到分页采集
网络·爬虫·python·网络爬虫
DanCheng-studio2 小时前
智能科学与技术毕设易上手题目建议
python·毕业设计·毕设
醇氧2 小时前
nohup 后台启动 uvicorn(仅测试 / 临时运行)
linux·运维·网络·python·python3.11
2601_957883843 小时前
2026年9月:专业解决雷神笔记本维修难题
python·电脑
Omics Pro3 小时前
预测性虚拟细胞中显式机理算子
大数据·数据库·人工智能·python·算法·机器学习·自然语言处理
秋名RG3 小时前
Java历代LTS版本新特性详解(代码增强版)
java·开发语言·python