爬虫爬取豆瓣电影、价格、书名

1、爬取豆瓣电影top250

bash 复制代码
import requests
from bs4 import BeautifulSoup

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

for i in range(0, 250, 25):
    print(f"--------第{i+1}到{i+25}个电影------------")
    response = requests.get(f"https://movie.douban.com/top250?start={i}", headers=headers)

    if response.ok:
        html = response.text
        soup = BeautifulSoup(html, "html.parser")
        all_titles = soup.findAll("span", attrs={"class": "title"})
        j = i
        for title in all_titles:
            title_string = title.string
            if "/" not in title_string:
                j += 1
                print(f"{j}、{title_string}")
    else:
        print("请求失败")

2、爬取价格

bash 复制代码
import requests
from bs4 import BeautifulSoup

content = requests.get("http://books.toscrape.com/").text
soup = BeautifulSoup(content, "html.parser")
# 因为价格在标签为p的里面,所以写p,它的属性为class="price_color"
all_prices = soup.findAll("p", attrs={"class": "price_color"})
print(all_prices)
for price in all_prices:
    print(price.string[2:])

3、爬取书名

bash 复制代码
import requests
from bs4 import BeautifulSoup

content = requests.get("http://books.toscrape.com/").text
soup = BeautifulSoup(content, "html.parser")
# 因为书名在h3中,又包了一层a,所以先找h3,再找a
all_titles = soup.findAll("h3")
for title in all_titles:
    all_links = title.findAll("a")
    for link in all_links:
        print(link.string)
相关推荐
szial6 小时前
网络爬虫与 CDP 实战(二):浏览器能访问,Python 却被拒绝?拆开登录态与 CSRF
爬虫·python·csrf
梅雅达编程笔记18 小时前
04-Python CSV数据保存与翻页抓取
开发语言·爬虫·python·pandas·数据采集·csv
梅雅达编程笔记20 小时前
03_网页表格数据抓取
爬虫·python·beautifulsoup·pandas·数据采集
傻啦嘿哟20 小时前
爬虫代理IP池从0到1:构建高可用代理池,彻底解决IP被封问题
网络·爬虫·tcp/ip
szial20 小时前
网络爬虫与 CDP 实战(一):网页数据到底在哪?从 HTTP 到分页采集
网络·爬虫·python·网络爬虫
深蓝电商API21 小时前
反检测浏览器(比特/AdsPower)原理剖析与自建替代方案
爬虫·反检测浏览器
深蓝电商API2 天前
Puppeteer vs Playwright:多维度压测后的选型建议
爬虫·puppeteer·playwright
压码路2 天前
BOSS直聘网页端研究【2026年9月】
爬虫·语言模型·boss直聘
深蓝电商API2 天前
用 CDP 协议直连浏览器:绕过框架做更底层的控制
爬虫·cdp
深蓝电商API2 天前
从元素定位到等待策略:写出稳定不出错的自动化脚本
爬虫