Python 爬虫进阶实战指南(Scrapy 与反爬对抗)
爬虫是 Python 最经典的实战方向,也是 CSDN 流量最大的话题之一。本文从爬虫基础讲起,覆盖 Requests 与解析、Scrapy 框架、常见反爬机制(UA、IP、验证码、动态渲染)与应对、数据存储、爬虫规范与合规边界,全程带可运行代码。
一、爬虫是什么
1.1 定义
爬虫(Web Scraping):用程序自动获取网页数据------把"人肉复制粘贴"变成"程序自动采集"。
应用场景:
- 价格监控(电商比价);
- 舆情监测(新闻、评论);
- 数据采集(招聘、房源、商品);
- 搜索引擎(索引网页)。
1.2 爬虫 vs API
优先用 API:很多网站提供官方 API(更稳定、合法、结构清晰)。只有没有 API 或 API 不全时才爬。
1.3 爬虫三件套
发送请求(Requests/httpx)→ 解析内容(BeautifulSoup/XPath/正则)→ 存储(CSV/JSON/数据库)
二、Requests 与解析基础
2.1 基本请求
python
import requests
# GET
resp = requests.get(
'https://example.com/api/data',
headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64)'},
timeout=10
)
print(resp.status_code) # 200
print(resp.text) # 原始文本
print(resp.json()) # JSON 响应
# POST
resp = requests.post(
'https://example.com/api/login',
json={'username': 'u', 'password': 'p'},
headers={'User-Agent': '...'}
)
# 带 Session(保持 Cookie)
session = requests.Session()
session.get('https://example.com/login')
resp = session.get('https://example.com/profile')
基本规范:
- 必须带 User-Agent(否则很多站直接拒绝);
- 设置 timeout(别让程序挂死);
- 错误处理(状态码、超时、连接错误)。
2.2 解析 HTML:BeautifulSoup
python
from bs4 import BeautifulSoup
html = resp.text
soup = BeautifulSoup(html, 'html.parser')
# 选择器
title = soup.title.text # 标题
items = soup.select('.product-item') # CSS 选择器
for item in items:
name = item.select_one('.name').text.strip()
price = item.select_one('.price').text.strip()
print(name, price)
2.3 解析 JSON(接口直取)
很多网站数据在接口里(JSON)------比解析 HTML 高效:
python
# 浏览器 Network 里找到 XHR 请求,直接请求接口
resp = requests.get('https://example.com/api/products?page=1', headers=headers)
data = resp.json()
for item in data['data']['list']:
print(item['title'], item['price'])
判断:优先找接口(稳定、字段全),找不到再解析 HTML。
三、Scrapy 框架
3.1 为什么用 Scrapy
Requests 写小爬虫够,但大规模采集会碰到:
- 并发控制、去重、调度;
- 中间件(代理、UA、Cookie 管理);
- 管道(清洗、存储);
- 断点续爬。
Scrapy 是爬虫框架的事实标准。
3.2 项目结构
bash
scrapy startproject my_spider
cd my_spider
scrapy genspider quotes quotes.example.com
my_spider/
├── spiders/
│ └── quotes.py # 爬虫逻辑
├── items.py # 数据模型
├── middlewares.py # 中间件(代理/UA)
├── pipelines.py # 数据管道(清洗/存储)
└── settings.py # 配置
3.3 一个完整爬虫
python
# spiders/quotes.py
import scrapy
class QuotesSpider(scrapy.Spider):
name = 'quotes'
start_urls = ['https://quotes.example.com/page/1/']
def parse(self, response):
# 提取当前页数据
for quote in response.css('div.quote'):
yield {
'text': quote.css('span.text::text').get(),
'author': quote.css('small.author::text').get(),
}
# 翻页(下一页)
next_page = response.css('li.next a::attr(href)').get()
if next_page:
yield response.follow(next_page, self.parse)
bash
scrapy crawl quotes -o quotes.json # 输出 JSON
3.4 Item 与 Pipeline(清洗存储)
python
# items.py
import scrapy
class QuoteItem(scrapy.Item):
text = scrapy.Field()
author = scrapy.Field()
tags = scrapy.Field()
python
# pipelines.py
class SavePipeline:
def open_spider(self, spider):
self.file = open('quotes.csv', 'w', encoding='utf-8-sig')
self.file.write('text,author\n')
def process_item(self, item, spider):
# 清洗 + 写文件
line = f"{item['text']},{item['author']}\n"
self.file.write(line)
return item
def close_spider(self, spider):
self.file.close()
python
# settings.py
ITEM_PIPELINES = {
'my_spider.pipelines.SavePipeline': 300, # 数字控制执行顺序
}
3.5 并发与限速配置
python
# settings.py
CONCURRENT_REQUESTS = 8 # 并发请求数
DOWNLOAD_DELAY = 1.0 # 请求间隔(礼貌爬取,防封)
DOWNLOAD_TIMEOUT = 15
USER_AGENT = 'Mozilla/5.0 ...'
COOKIES_ENABLED = True
四、常见反爬机制与应对
4.1 UA 检测与应对
反爬:拒绝无 UA 或非浏览器 UA 的请求。
应对:
python
# 1. 固定浏览器 UA
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/120.0.0.0'}
# 2. 随机 UA(fake-useragent 库)
from fake_useragent import UserAgent
ua = UserAgent()
headers = {'User-Agent': ua.random}
4.2 IP 封禁与应对
反爬:单位时间请求过多 → 封 IP。
应对:
- 限速 (首要):
time.sleep(1)或DOWNLOAD_DELAY; - 代理池:换 IP。
python
# 代理(Requests)
proxies = {
'http': 'http://user:pass@proxy_ip:port',
'https': 'http://user:pass@proxy_ip:port'
}
resp = requests.get(url, headers=headers, proxies=proxies, timeout=10)
# Scrapy 中间件切换代理
class ProxyMiddleware:
def process_request(self, request, spider):
request.meta['proxy'] = get_proxy() # 从代理池取
代理来源:付费代理服务(稳定)、自建代理池、免费代理(不稳定,慎用)。
4.3 Cookie/Session 校验
反爬:需要登录态或先访问首页种 Cookie。
应对:用 Session 保持、模拟登录。
python
session = requests.Session()
# 先访问首页拿 Cookie
session.get('https://example.com')
# 再请求目标页
resp = session.get('https://example.com/data')
4.4 动态渲染(JS 渲染页面)
反爬:数据是 JS 异步加载的------直接请求 HTML 拿不到。
应对(按优先级):
- 找接口(最优先):浏览器 Network 里找 XHR,直接请求 JSON 接口;
- 无头浏览器:Selenium/Playwright 模拟浏览器执行 JS。
python
# Playwright(现代推荐)
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto('https://example.com', timeout=30000)
page.wait_for_selector('.product-item') # 等 JS 渲染完成
html = page.content()
browser.close()
4.5 验证码
应对原则:
- 降低触发概率:限速、像人一样访问(随机延迟、滚动页面);
- 打码平台(收费,合规前提下);
- 有验证码的站点直接放弃或换数据源------强攻性价比极低。
4.6 请求签名/加密参数
部分站(如某些电商)接口带签名参数(sign、token)------逆向成本高:
- 优先找公开接口(移动端/旧版);
- 评估 ROI:签名逆向几小时,数据价值是否值得。
五、数据存储
5.1 CSV
python
import csv
with open('data.csv', 'w', encoding='utf-8-sig', newline='') as f:
writer = csv.DictWriter(f, fieldnames=['title', 'price'])
writer.writeheader()
writer.writerows(items)
5.2 JSON
python
import json
with open('data.json', 'w', encoding='utf-8') as f:
json.dump(items, f, ensure_ascii=False, indent=2)
5.3 数据库
python
import pymysql
conn = pymysql.connect(host='localhost', user='root', password='123456', db='spider')
with conn.cursor() as cursor:
cursor.executemany(
'INSERT INTO products(title, price) VALUES (%s, %s)',
[(i['title'], i['price']) for i in items]
)
conn.commit()
conn.close()
六、爬虫规范与合规边界(重要)
6.1 遵守的规则
- robots.txt :先看
https://site.com/robots.txt允许哪些路径; - 限速:控制请求频率,不给对方服务器造成压力;
- 声明用途:注明爬虫 UA、提供联系方式;
- 只取需要的数据:别整站下载。
6.2 法律与合规风险(必须知道)
- 个人信息:爬取个人信息(手机号、身份证)涉及个人信息保护法,风险极高;
- 著作权:爬取并转载他人内容可能侵权;
- 突破技术措施:绕过登录验证、破解加密可能违法;
- 影响服务:高频请求导致对方服务器崩溃属于破坏计算机信息系统。
边界建议:
- 只爬公开数据、公开接口;
- 个人学习、研究用途为主;
- 商用爬虫务必做法律合规评估。
6.3 判刑案例背景(了解即可)
国内已有爬虫从业者因**侵入性爬取(绕过验证、高频攻击、爬取付费数据)**被判刑的案例------这不是危言耸听,是真实司法实践。爬虫本身不违法,违法的是手段(绕过防护)和目的(侵害权益)。
七、实战:招聘信息采集(公开数据示例)
python
import requests
import time
import csv
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) ...',
'Accept': 'application/json'
}
all_jobs = []
for page in range(1, 6): # 前 5 页
resp = requests.get(
f'https://api.example.com/jobs?page={page}&size=20',
headers=headers,
timeout=10
)
if resp.status_code != 200:
print(f'第{page}页请求失败: {resp.status_code}')
time.sleep(2)
continue
jobs = resp.json().get('data', {}).get('list', [])
if not jobs:
break
for job in jobs:
all_jobs.append({
'title': job['title'],
'company': job['company'],
'salary': job.get('salary', ''),
'city': job.get('city', ''),
'url': job.get('url', '')
})
print(f'第{page}页: {len(jobs)} 条')
time.sleep(1) # 礼貌限速
# 存储
with open('jobs.csv', 'w', encoding='utf-8-sig', newline='') as f:
writer = csv.DictWriter(f, fieldnames=['title', 'company', 'salary', 'city', 'url'])
writer.writeheader()
writer.writerows(all_jobs)
print(f'共采集 {len(all_jobs)} 条')
八、爬虫工程质量
8.1 断点续爬
python
# 记录已爬页码
def load_progress():
try:
with open('progress.txt') as f:
return int(f.read())
except FileNotFoundError:
return 1
def save_progress(page):
with open('progress.txt', 'w') as f:
f.write(str(page))
8.2 异常与重试
python
from tenacity import retry, stop_after_attempt, wait_exponential
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, max=10))
def fetch(url):
resp = requests.get(url, headers=headers, timeout=10)
resp.raise_for_status()
return resp
8.3 去重
python
seen = set()
for item in items:
if item['url'] in seen:
continue
seen.add(item['url'])
# 处理
九、常见坑速查
| 报错/现象 | 原因 | 解决 |
|---|---|---|
| 403 Forbidden | UA 被识别 | 浏览器 UA + 随机 |
| 429 Too Many Requests | 请求太快 | 限速、代理 |
| 空数据 | JS 渲染 | 找接口 / 无头浏览器 |
| 中文乱码 | 编码不对 | resp.encoding = 'utf-8' |
| 连接超时 | 网络/目标慢 | 重试 + 更长 timeout |
| 反爬验证码 | 触发风控 | 降频、像人一样 |
| SSL 证书报错 | 证书验证 | verify=False(慎用) |
| 数据存不进数据库 | 类型/字段问题 | 打印日志逐条排查 |
本章小结
爬虫进阶的核心:先找接口、再解析、框架用 Scrapy、反爬要礼貌。记住优先级------API > 接口 > HTML 解析 > 无头浏览器。同时牢记合规红线:公开数据、限速、不绕防护、不爬个人信息。
下一篇讲操作系统核心------大厂面试的硬核基础。
(全文完)