爬虫案例(读书网)(下)

上篇链接:

CSDN-读书网https://mp.csdn.net/mp_blog/creation/editor/139306808

可以看见基本的全部信息:如(author、bookname、link.....)

写下代码如下:

python 复制代码
import requests
from bs4 import BeautifulSoup
from lxml import etree

headers={'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36'}
link="https://www.dushu.com/"
r=requests.get(link,headers=headers)
r.encoding='utf-8'

soup=BeautifulSoup(r.text,'lxml')
house_list=soup.find_all('div',class_="border books-center")
html=etree.HTML(r.text)
    # name=html.xpath('//div[@class="property-content-title"]/h3/text()')
# for house in house_list:
#     name=soup.find('div',class_="nlist").a.strong.text()
#
#     print(name)
name=html.xpath('//div[@class="bookname"]/a/text()')
author=html.xpath('//div[@class="bookauthor"]/text()')
# href=html.xpath('//div[@class="nlist"]/div/ul/li/a/@href')

#print(type(author))
for i,o in zip(name,author):
    print('<<'+i+'>>',o)

运行结果:

接下来添加link链接:

可以看见现在网站设置了反爬,我们现在通过检查浏览器能正常爬取还是有反爬:

python 复制代码
# 请用 python+selenium  爬取 XXX 网站上的所有a链接的 href属性并访问,输出访问地址和状态码
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import requests

driver = webdriver.Chrome()
# 这里以百度为例
driver.get("https://www.dushu.com/")

wait = WebDriverWait(driver, 10)
links = wait.until(EC.presence_of_all_elements_located((By.XPATH, "//a")))

# 遍历所有的链接元素,并输出href属性值
for link in links:
    href = link.get_attribute("href")
    if href.startswith("http"):
        response = requests.get(href)
        print(href, response.status_code)
    else:
        link.click()
        print(driver.current_url, driver.execute_script('return document.readyState'),
              requests.get(driver.current_url).status_code)

# 关闭浏览器
driver.quit()

运行结果:

现在可以看出是反爬。

最后我们的解析反爬,在下一篇文章详细介绍几个方法和使用效果。

相关推荐
belldeep1 分钟前
python:wps2md
python·wps2md
IvanCodes5 分钟前
Python 正则表达式(十四):文本匹配、查找与替换
开发语言·python·正则表达式
夜雪一千5 分钟前
Python URL编码踩坑:单层编码、双重编码实战解析
开发语言·python
茶栀(*´I`*)8 分钟前
【Python数据可视化】Matplotlib从零到精通:图表核心解剖、多子图布局与三大经典图表实战
python·信息可视化·matplotlib
孙启超10 分钟前
【AI开发之Rust】第 21 课:双端集成与出包 —— Android(.so→AAR)与 iOS(xcframework)
开发语言·后端·rust
SelectDB技术团队15 分钟前
Agent Trace 数据底座建设:宽表建模、全文检索与成本聚合的配置与验证步骤
大数据·python·clickhouse·elk·elasticsearch·全文检索·复杂查询
Brilliantwxx20 分钟前
【STM32】 __weak 弱引用、中断标志位、异常处理与 HAL_Delay 优先级
开发语言·stm32·单片机·嵌入式硬件
驭渊的小故事21 分钟前
SpringBoot 配置文件详解:properties 与 yml 从入门到实战
java·开发语言·笔记
计算机源码社24 分钟前
【27届大数据毕设】基于Spark的TikTok短视频属性分享与互动比率可视化分析系统 基于机器学习的TikTok短视频认证属性与互动效能关联分析
大数据·hadoop·python·机器学习·spark·毕业设计·课程设计
北冥有鱼被烹32 分钟前
Linux timeout 命令完全指南:指定固定时间并发送固定信号
python