python (第十一章 动态网页爬取和反爬机制)

第10周学习计划:动态网页爬取和反爬机制

目标:掌握动态网页爬取工具(Selenium)和应对常见反爬措施。

学习内容总览
  1. 动态网页爬取:Selenium 基础。
  2. 反爬机制:识别和应对(如请求头、延迟、代理)。
  3. 实践任务:爬取动态电商页面(如京东商品价格)。

第一部分:动态网页爬取 - Selenium

1. Selenium 简介
  • 功能:模拟浏览器操作,加载 JavaScript 动态内容。
  • 安装 :
    • pip install selenium
    • 下载浏览器驱动(如 ChromeDriver),与你的 Chrome 版本匹配。
    • 将驱动放入 PATH,或在代码中指定路径。
2. 基本用法
  • 示例:
python 复制代码
from selenium import webdriver
from selenium.webdriver.chrome.service import Service

# 指定 ChromeDriver 路径(替换为你的路径)
service = Service(executable_path="path/to/chromedriver")
driver = webdriver.Chrome(service=service)

driver.get("https://www.example.com")
print(driver.title)  # 输出页面标题
driver.quit()  # 关闭浏览器
3. 元素定位
  • 用 find_element 或 find_elements 提取网页元素。
  • 常用方法 :
    • By.ID
    • By.CLASS_NAME
    • By.CSS_SELECTOR
  • 示例:
python 复制代码
from selenium.webdriver.common.by import By

driver.get("http://localhost:8000/test_shop.html")
products = driver.find_elements(By.CLASS_NAME, "product")
for product in products:
    name = product.find_element(By.CLASS_NAME, "name").text
    price = product.find_element(By.CLASS_NAME, "price").text
    print(f"{name} - {price}")
driver.quit()

第二部分:反爬机制

1. 常见反爬手段
  • User-Agent 检查:网站检测请求头。
  • IP 限制:频繁请求被封。
  • JavaScript 验证:数据动态加载。
  • 验证码:需要人工干预。
2. 应对方法
  • 伪装请求头 :

    python 复制代码
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/91.0.4472.124"}
  • 延迟请求 :

    python 复制代码
    import time
    time.sleep(2)  # 每次请求间隔2秒
  • 代理IP :

    • 使用免费代理或代理服务(如 proxyscrape.com)。
    python 复制代码
    proxies = {"http": "http://代理IP:端口", "https": "https://代理IP:端口"}
    response = requests.get(url, proxies=proxies)

实践任务:爬取动态电商页面

目标

用 Selenium 爬取京东(https://www.jd.com/)首页的商品名称和价格(动态加载部分),并保存到文件。

代码实现
python 复制代码
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
import json
from datetime import datetime
import time

class JDPriceTracker:
    def __init__(self, filename="jd_prices.json"):
        self.prices = {}
        self.filename = filename
        self.load_prices()
        # 配置 Selenium
        self.service = Service(executable_path="path/to/chromedriver")  # 替换为你的路径
        self.driver = webdriver.Chrome(service=self.service)

    def fetch_prices(self, url):
        """抓取京东首页商品价格"""
        self.driver.get(url)
        time.sleep(3)  # 等待页面加载
        
        try:
            # 定位商品元素(京东的动态商品在 .gl-i-wrap 中)
            products = self.driver.find_elements(By.CLASS_NAME, "gl-i-wrap")
            current_prices = {}
            for product in products[:5]:  # 只抓前5个
                try:
                    name = product.find_element(By.CSS_SELECTOR, ".p-name a").get_attribute("title")
                    price = product.find_element(By.CSS_SELECTOR, ".p-price i").text
                    current_prices[name] = float(price)
                    print(f"抓到:{name} - ¥{price}")
                except Exception as e:
                    print(f"解析单个商品失败:{e}")
            self.update_prices(current_prices)
        except Exception as e:
            print(f"抓取失败:{e}")
        finally:
            self.driver.quit()

    def update_prices(self, current_prices):
        now = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
        for name, price in current_prices.items():
            if name not in self.prices:
                self.prices[name] = [{"time": now, "price": price}]
            elif self.prices[name][-1]["price"] != price:
                self.prices[name].append({"time": now, "price": price})
        self.save_prices()

    def view_prices(self):
        if not self.prices:
            print("暂无记录!")
        else:
            print("\n价格记录:")
            for name, history in self.prices.items():
                print(f"{name}:")
                for entry in history:
                    print(f"  {entry['time']} - ¥{entry['price']}")

    def save_prices(self):
        with open(self.filename, "w", encoding="utf-8") as f:
            json.dump(self.prices, f, ensure_ascii=False, indent=2)

    def load_prices(self):
        try:
            with open(self.filename, "r", encoding="utf-8") as f:
                self.prices = json.load(f)
        except FileNotFoundError:
            self.prices = {}

def main():
    tracker = JDPriceTracker()
    url = "https://www.jd.com/"
    
    while True:
        print("\n=== 京东价格监控 ===")
        print("1. 抓取当前价格")
        print("2. 查看记录")
        print("3. 退出")
        
        choice = input("请选择操作(1-3):")
        
        if choice == "1":
            tracker.fetch_prices(url)
        
        elif choice == "2":
            tracker.view_prices()
        
        elif choice == "3":
            print("谢谢使用!")
            break
        
        else:
            print("无效选择,请输入 1-3!")

if __name__ == "__main__":
    main()

代码讲解
  1. Selenium 配置:

    • 用 Service 指定 ChromeDriver 路径。
    • time.sleep(3) 等待页面动态内容加载。
  2. 元素定位:

    • .gl-i-wrap 是京东商品的容器。
    • .p-name a 获取标题,.p-price i 获取价格。
  3. 异常处理:

    • 捕获单个商品解析失败,避免程序崩溃。
    • 用 finally 确保浏览器关闭。

动手实践
  1. 准备环境 :
    • 安装 selenium:pip install selenium。
    • 下载 ChromeDriver,替换代码中的路径。
  2. 运行程序 :
    • 执行代码,选 1 抓取京东首页。
    • 选 2 查看记录。
    • 选 3 退出。
  3. 检查结果 :
    • 查看 jd_prices.json,确认数据。

预期输出
  • 抓取(示例,实际结果随京东变化):

    抓到:Apple iPhone 14 Pro Max - ¥7999.0
    抓到:华为 Mate 60 Pro - ¥6499.0
    ...

  • 查看记录:

    价格记录:
    Apple iPhone 14 Pro Max:
    2025-02-25 22:00:00 - ¥7999.0
    华为 Mate 60 Pro:
    2025-02-25 22:00:00 - ¥6499.0


小挑战
  1. 多页爬取:滚动页面加载更多商品。
  2. 代理支持:加入代理池应对 IP 限制。
  3. 价格变化提醒:如果价格变化,打印通知。
相关推荐
个 人 练 习 生几秒前
C++ string 类模拟实现:从底层理解字符串(上)
开发语言·c++·经验分享·学习·程序人生
2601_966949659 分钟前
多市场量化策略的数据接口应该如何设计:从数据层架构到策略接入
开发语言·python·数据分析·pandas·量化交易·股票数据·quantdash
坤坤子吖12 分钟前
C++智能指针:RAII、shared_ptr与内存泄漏
开发语言·c++·笔记·学习
全栈弄潮儿18 分钟前
Python实战第1期:Python环境搭建与第一个程序
python
奋斗的阿狸_198622 分钟前
ESP32-S3 + ES7210 四通道麦克风录音上传方案
c语言·开发语言·fpga开发
心易行者24 分钟前
Agent应用+API端点商业化进阶实战:从单体智能体到可付费调用的API全流程
运维·服务器·人工智能·python·apache
weixin1997010801637 分钟前
《1688图片空间API踩坑:img.upload 与 album.* 的防盗链与CDN缓存问题》(附Python源码)
开发语言·python·缓存
ACP广源盛139246256731 小时前
Type‑C 扩展坞方案选型笔记|国产 DP1.4 转 HDMI2.0 桥接芯片 GSV2201S @ACP评估
c语言·开发语言·笔记·硬件架构·硬件工程·国产芯片
布吉岛的石头1 小时前
Java 程序员第 49 阶段11:Java 接 BERT:用 ONNX Runtime 做句向量推理
java·人工智能·python·深度学习·bert·transformer
第25小时记录1 小时前
数据结构与算法 -第 3 章 常用算法-动态规划
数据结构·python