爬虫爬取豆瓣电影、价格、书名

1、爬取豆瓣电影top250

bash 复制代码
import requests
from bs4 import BeautifulSoup

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

for i in range(0, 250, 25):
    print(f"--------第{i+1}到{i+25}个电影------------")
    response = requests.get(f"https://movie.douban.com/top250?start={i}", headers=headers)

    if response.ok:
        html = response.text
        soup = BeautifulSoup(html, "html.parser")
        all_titles = soup.findAll("span", attrs={"class": "title"})
        j = i
        for title in all_titles:
            title_string = title.string
            if "/" not in title_string:
                j += 1
                print(f"{j}、{title_string}")
    else:
        print("请求失败")

2、爬取价格

bash 复制代码
import requests
from bs4 import BeautifulSoup

content = requests.get("http://books.toscrape.com/").text
soup = BeautifulSoup(content, "html.parser")
# 因为价格在标签为p的里面,所以写p,它的属性为class="price_color"
all_prices = soup.findAll("p", attrs={"class": "price_color"})
print(all_prices)
for price in all_prices:
    print(price.string[2:])

3、爬取书名

bash 复制代码
import requests
from bs4 import BeautifulSoup

content = requests.get("http://books.toscrape.com/").text
soup = BeautifulSoup(content, "html.parser")
# 因为书名在h3中,又包了一层a,所以先找h3,再找a
all_titles = soup.findAll("h3")
for title in all_titles:
    all_links = title.findAll("a")
    for link in all_links:
        print(link.string)
相关推荐
Yan-英杰5 小时前
从“投屏“到“办公“,ToDesk鸿蒙版4.8.0.0补齐远控“进入→处理→离开“全流程
服务器·人工智能·爬虫·机器学习·ai开发工具
深蓝电商API6 小时前
Redis 在分布式爬虫中的作用
redis·分布式·爬虫
K哥爬虫1 天前
【验证码逆向专栏】某美世界卡状态查询滑块逆向分析
爬虫
万象观察录1 天前
API面临越权、爬虫、撞库等风险,有什么解决方案?
爬虫
CDN3601 天前
爬虫把价格库扒光、SQL注入打穿数据库?360CDN WAF规则引擎深度定制,把防护精度拉到逐字段级
数据库·爬虫·sql
深蓝电商API1 天前
Scrapy 架构源码解析
爬虫·scrapy
IT毕设实战小研2 天前
基于大数据的WTA职业网球赛事演变与竞技格局数据可视化分析
大数据·后端·爬虫·python·信息可视化·课程设计
2601_962382432 天前
用“爬虫”抓取小红书内容侵权吗 福建法院判明边界
爬虫·数字经济·知识产权·不正当竞争·福建法院
Mycdn_WD2 天前
从 AI 爬虫变多了到 AI 抓取开始分流,CDN 行业进入下一阶段
人工智能·爬虫·cdn·pcdn·城域网·pcdn资源招募
该逃避避3 天前
爬虫代理配置指南:从0搭建Crawl4AI网页爬虫教程
爬虫