利用python抓取小说,爬虫抓取小说

1.https://www.bqg70.com/ 首先进入这个网址,进入笔趣阁官网

2.搜索你想要看的小说

3.选择你想看的小说后,在地址栏会出现一个数字,举例:"https://www.bqg70.com/book/3315/"

那个数字请复制好,例如:"3315"

4.运行代码,输入这个数字 ,即可下载对应的小说

5.安装包

pip install requests

pip install parsel

pip install prettytable

代码如下:

python 复制代码
import requests  # 第三方的模块
import parsel  # 第三方的模块
import os  # 内置模块 文件或文件夹

filename = '小说\\'
if not os.path.exists(filename):
    os.mkdir(filename)

headers = {
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36',
}

    

rid = input('输入书名ID:')
link = f'https://www.bqg70.com/book/{rid}/'

html_data = requests.get(url=link, headers=headers).text
# print(html_data)
selector_2 = parsel.Selector(html_data)
divs = selector_2.css('.listmain dd')
for div in divs:
    title = div.css('a::text').get()
    href = div.css('a::attr(href)').get()
    url = 'https://www.bqg70.com' + href

    try:
        response = requests.get(url=url, headers=headers)
        selector = parsel.Selector(response.text)
        # getall 返回的是一个列表 []
        book = selector.css('#chaptercontent::text').getall()
        book = '\n'.join(book)
        # 数据保存
        with open(filename + title + '.txt', mode='a', encoding='utf-8') as f:
            f.write(book)
            print('正在下载章节:  ', title)
    except Exception as e:
        print(e)
相关推荐
朝朝辞暮i7 小时前
C++ 第 40 章:ROS2 Publisher + Timer + Subscriber + Callback + Executor 完整闭环
开发语言·c++·算法·ros2
专注仿真7 小时前
决策树的算法
开发语言·高性能算法
Evand J7 小时前
【电机滤波例程2】负载转矩增广扩展卡尔曼滤波(EKF)原理与MATLAB例程:五维PMSM状态与负载阶跃估计。订阅专栏后可查看完整代码
开发语言·matlab·电机·ekf·卡尔曼滤波
迅猛龙办公室7 小时前
python实现简单进度条
java·前端·python
郝学胜_神的一滴7 小时前
AI 编程智能体 06:用Anaconda搞定Python多环境,彻底告别版本兼容灾难
人工智能·python
清水白石0087 小时前
Python 对象模型深度解析:从“一切皆对象”到 id、type、isinstance 底层机制与小整数缓存原理
开发语言·python·缓存
for_ever_love__7 小时前
Python 进阶——迭代器、生成器、装饰器与 asyncio
python·装饰器·生成器·大模型开发
燕落南枝7 小时前
c语言 洛谷P1765手机题解
c语言·开发语言
Rosanci7 小时前
Codex 下载与本地部署实战:从安装到运行全流程指南
开发语言·前端·算法·chatgpt·codex
5G微创业7 小时前
去水印API如何实现技术变现
python·sdk·短视频