Python 爬虫基础

Python 爬虫基础

1.1 理论

在浏览器通过网页拼接【/robots.txt】来了解可爬取的网页路径范围

例如访问: https://www.csdn.net/robots.txt

User-agent: *

Disallow: /scripts

Disallow: /public

Disallow: /css/

Disallow: /images/

Disallow: /content/

Disallow: /ui/

Disallow: /js/

Disallow: /scripts/

Disallow: /article_preview.html*

Disallow: /tag/

Disallow: /?

Disallow: /link/

Disallow: /tags/

Disallow: /news/

Disallow: /xuexi/

通过Python Requests 库发送HTTP【Hypertext Transfer Protocol "超文本传输协议"】请求

通过Python Beautiful Soup 库来解析获取到的HTML内容

HTTP请求

HTTP响应

1.2 实践代码 【获取价格&书名】

python 复制代码
import requests
# 解析HTML
from bs4 import BeautifulSoup

# 将程序伪装成浏览器请求
head = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)"
}
requests = requests.get("http://books.toscrape.com/",headers= head)
# 指定编码
# requests.encoding= 'gbk'
if requests.ok:
    # file = open(r'C:\Users\root\Desktop\Bug.html', 'w')
    # file.write(requests.text)
    # file.close
    content =  requests.text
    ## html.parser 指定当前解析 HTML 元素
    soup = BeautifulSoup(content, "html.parser")
    
    ## 获取价格
    all_prices = soup.findAll("p", attrs={"class":"price_color"})
    for price in all_prices:
        print(price.string[2:])

    ## 获取名称
    all_title = soup.findAll("h3")
    for title in all_title:
        ## 获取h3下面的第一个a元素
        print(title.find("a").string)
else:
    print(requests.status_code)

1.3 实践代码 【获取 Top250 的电影名】

python 复制代码
import requests
# 解析HTML
from bs4 import BeautifulSoup
head = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)"
}
# 获取 TOP 250个电影名
for i in range(0,250,25):
    response = requests.get(f"https://movie.douban.com/top250?start={i}", headers= head)
    if response.ok:
        content =  response.text
        soup = BeautifulSoup(content, "html.parser")
        all_titles = soup.findAll("span", attrs={"class": "title"})
        for title in all_titles:
            if "/" not in title.string:
                print(title.string) 
    else:
        print(response.status_code)

1.4 实践代码 【下载图片】

python 复制代码
import requests
# 解析HTML
from bs4 import BeautifulSoup

head = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)"
}
response = requests.get(f"https://www.maoyan.com/", headers= head)
if response.ok:
    soup = BeautifulSoup(response.text, "html.parser")
    for img in soup.findAll("img", attrs={"class": "movie-poster-img"}):
        img_url = img.get('data-src')
        alt = img.get('alt')
        path = 'img/' + alt + '.jpg'
        res = requests.get(img_url)
        with open(path, 'wb') as f:
            f.write(res.content)
else:
    print(response.status_code)

1.5 实践代码 【千图网图片 - 爬取 - 下载图片】

python 复制代码
import requests
# 解析HTML
from bs4 import BeautifulSoup


# 千图网图片 - 爬取
head = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)"
}
# response = requests.get(f"https://www.58pic.com/piccate/53-0-0.html", headers= head)
# response = requests.get(f"https://www.58pic.com/piccate/53-598-2544.html", headers= head)
response = requests.get(f"https://www.58pic.com/piccate/53-527-1825.html", headers= head)
if response.ok:
    soup = BeautifulSoup(response.text, "html.parser")
    for img in soup.findAll("img", attrs={"class": "lazy"}):
        img_url = "https:" + img.get('data-original')
        alt = img.get('alt')
        path = 'imgqiantuwang/' + str(alt) + '.jpg'
        res = requests.get(img_url)
        with open(path, 'wb') as f:
            f.write(res.content)
else:
    print(response.status_code)
相关推荐
开源量化GO3 小时前
看到“2026年 Python 与 API”时,用示例和拆解看清关系
人工智能·python
高洁013 小时前
大模型是怎么“学会“的:预训练、微调、对齐三段论
python·深度学习·transformer·知识图谱·tornado
529宝宝起名网3 小时前
用 Python 开发名字五行八字匹配工具:从八字排盘到喜用神起名的全流程实现
python
CodeStats3 小时前
【Java类加载器】Java 类加载器完整体系深度拆解(下):自定义加载器实战与 SPI 破坏双亲委派(MySQL 驱动揭秘)
java·开发语言·jvm·classloader·双亲委派
砚底藏山河3 小时前
并发与限频工程:把20只的2秒压到0.5秒不封号(魔码量化实战 #03)
java·开发语言·数据库·python·金融
论文复现现场3 小时前
本地 PyTorch 训练 OOM,第一次租 RTX 4090 云 GPU 怎么迁移项目?从环境检查到 100 Step 跑通
人工智能·pytorch·python·深度学习·cuda
云雀衔光3 小时前
MCP 协议全景:为什么它是 AI 连接工具的「USB-C」
java·开发语言·数据库·人工智能·ai编程
geovindu3 小时前
go: Task Scheduler
开发语言·后端·golang
会飞的拖把4 小时前
Python 多任务编程:并发、进程、线程与线程安全详解
开发语言·python
君顾14 小时前
折扣卡 CPS 小程序定制:从业务流程到系统架构的实践指南
java·开发语言·折扣卡