XPath与lxml解析库

test.xml

XML 复制代码
<?xml version="1.0" encoding="utf-8"?>

<bookstore>

    <book name="halibote">
        <title lang="en">Harry Potter</title>
        <author>J K. Rowling</author>
        <year>2005</year>
        <price>29.99</price>
        <abc>
            <book lang="中文">neibu</book>
        </abc>
    </book>

    <book name="hongloumeng">
        红楼梦
    </book>

</bookstore>

hello.html

html 复制代码
<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <title>Title</title>
</head>
<body>
<!-- hello.html -->
<div>
    <ul>
        <li class="item-0">meiguo<a href="link1.html">first item</a></li>
        <li class="item-1"><a href="link2.html">second item</a></li>
        <li class="item-inactive"><a href="link3.html"><span
                class="bold">third item</span></a></li>
        <li class="item-1"><a href="link4.html">fourth item</a></li>
        <li class="item-0"><a href="link5.html">fifth item</a></li>
    </ul>
</div>

</body>
</html>

选取节点

python 复制代码
from lxml import etree

tree = etree.parse("test.xml")
list_node = tree.xpath("book/@name")
print(list_node[0])

list_node = tree.xpath("/bookstore")
print(list_node[0])
list_node = tree.xpath("book/title")
print(list_node[0].text)
list_node = tree.xpath("book//book")
print(list_node)
list_node = tree.xpath("//@lang")
print(list_node)

谓语:指路径表达式的附加条件

python 复制代码
from lxml import etree

tree = etree.parse("test.xml")
list_node = tree.xpath("book[2]")
print(list_node[0].text)

选取未知节点

python 复制代码
from lxml import etree

tree = etree.parse("test.xml")
list_node = tree.xpath("/bookstore/*")
print(list_node)

选取若干路径

python 复制代码
from lxml import etree

tree = etree.parse("test.xml")
list_node = tree.xpath("//book/title | //book/price")
print(list_node)

通过轴限定

python 复制代码
from lxml import etree

tree = etree.parse("test.xml")
list_node = tree.xpath("descendant::book")
print(list_node)

操作XML节点

python 复制代码
from lxml import etree

root = etree.Element("root",a="1")
child = etree.SubElement(root, "child")
root.set("b", "2")
root.text = "yilang"
print(etree.tostring(root))
print(root.tag)

print(root.text)

# 从字符串中解析XML,返回根节点
root = etree.XML("<root>"
                    "<a x='123'>aText"
                        "<b/>"
                        "<c/>"
                        "<b/>"
                    "</a>"
                 "</root>")
# 从根节点查找,返回匹配到的节点名称
print(root.find("a").tag)
# 从根节点开始查找,返回匹配到的第一个节点的名称
print(root.findall(".//a[@x]")[0].tag)

在XML中搜索

python 复制代码
from lxml import etree

tree = etree.parse("hello.html",parser=etree.HTMLParser())
list_node = tree.xpath("//li")
print(list_node[0].text)
相关推荐
Gu Gu Study15 小时前
ScoutLoop开放域深度研究引擎(agent的初步设计想法)
人工智能·python
卷无止境16 小时前
写代码这件事,到底该讲究点什么?
后端·python
卷无止境16 小时前
循环复杂度到底在算什么,Python 代码怎么才能写得让人一看就懂
后端·python
lpfasd12316 小时前
MediaCrawler 项目深度分析
chrome·python·chrome devtools
Dxy123931021616 小时前
Python项目打包成EXE完整教程(PyInstaller实战避坑)
开发语言·python
bamb0017 小时前
一个项目带你入门AI应用开发01
python
05664617 小时前
Python康复训练——常用标准库
开发语言·python·学习
hyf32663317 小时前
泛程序:从零开始搭建稳定程序项目框架
运维·服务器·爬虫·百度·seo
昆曲之源_娄江河畔17 小时前
Python如何安装flask, pymssql
开发语言·python·flask·pymssql
05664618 小时前
Python康复训练——控制流与函数
开发语言·python·学习