HTML Document Loaders in LangChain

https://python.langchain.com.cn/docs/modules/data_connection/document_loaders/how_to/html

HTML Document Loaders in LangChain

This content is based on LangChain's official documentation (langchain.com.cn) and explains two HTML loaders ---tools to extract text and metadata from HTML files into LangChain Document objects---in simplified terms. It strictly preserves original source codes, examples, and knowledge points without arbitrary additions or modifications.

Key Note: HTML (HyperText Markup Language) is the standard language for web documents. LangChain's HTML loaders strip away HTML tags to extract usable text, with optional metadata (e.g., page title).

1. What Are HTML Loaders?

HTML loaders convert raw HTML files into structured Document objects for LangChain workflows.

  • Core function: Extract text content from HTML (removing tags like <h1>, <p>) and attach metadata (e.g., file source).
  • Two supported loaders:
    • UnstructuredHTMLLoader: Simple loader for basic text extraction.
    • BSHTMLLoader: Uses the BeautifulSoup4 library to extract text + page title (stored in metadata).

2. Prerequisites

  • For BSHTMLLoader, install the BeautifulSoup4 library first (required for HTML parsing):

    bash 复制代码
    pip install beautifulsoup4

3. Loader 1: UnstructuredHTMLLoader (Basic Text Extraction)

This loader extracts plain text from HTML, ignoring complex metadata (e.g., page title).

Step 3.1: Import the Loader

python 复制代码
from langchain.document_loaders import UnstructuredHTMLLoader

Step 3.2: Initialize and Load the HTML File

python 复制代码
# Initialize loader with the path to your HTML file
loader = UnstructuredHTMLLoader("example_data/fake-content.html")

# Load the HTML into a Document object
data = loader.load()

Step 3.3: View the Result

python 复制代码
data

Output (Exact as Original):

python 复制代码
[Document(page_content='My First Heading\n\nMy first paragraph.', lookup_str='', metadata={'source': 'example_data/fake-content.html'}, lookup_index=0)]

4. Loader 2: BSHTMLLoader (Text + Title Extraction)

This loader uses BeautifulSoup4 to extract both text content and the HTML page's title (stored in the title field of metadata).

Step 4.1: Import the Loader

python 复制代码
from langchain.document_loaders import BSHTMLLoader

Step 4.2: Initialize and Load the HTML File

python 复制代码
# Initialize loader with the path to your HTML file
loader = BSHTMLLoader("example_data/fake-content.html")

# Load the HTML into a Document object
data = loader.load()

Step 4.3: View the Result

python 复制代码
data

Output (Exact as Original):

python 复制代码
[Document(page_content='\n\nTest Title\n\n\nMy First Heading\nMy first paragraph.\n\n\n', metadata={'source': 'example_data/fake-content.html', 'title': 'Test Title'})]

Key Takeaways

  • UnstructuredHTMLLoader: Extracts basic text from HTML (no title metadata).
  • BSHTMLLoader: Requires BeautifulSoup4, extracts text + page title (stored in metadata["title"]).
  • Both loaders return Document objects with page_content (extracted text) and metadata["source"] (file path).
相关推荐
骑着蜗牛撵大象32712 小时前
多 Agent 串行流水线:把一个任务拆成可重试、可续跑的 Pipeline 节点链
前端·langchain
用户2986985301412 小时前
Python 中 HTML 转 PDF 的实现方法
python·html·api
Maiko Star14 小时前
* LangChain 文档切分器:切分策略、源码分析与常用切分器
langchain
Frag0ut15 小时前
Chrome安全隐私设置指南:5个原生功能筑牢浏览器数字防线
chrome·google·谷歌浏览器·权限管理·隐私保护·安全设置·安全浏览
我就是不信17 小时前
用 C 语言实现一个支持 HTML 与视频的 HTTP(S) 服务器
c语言·html·音视频
winfredzhang17 小时前
用 Chrome 插件 + 局域网 Qwen2.5-VL 打造视频截屏 OCR 工具
人工智能·chrome·ocr·api·plugin
~kiss~1 天前
LangChain 基础学习
学习·langchain
制造数据与AI践行者老蒋1 天前
排坑笔记:LangChain 多工具 Agent 完整性校验 return_intermediate_steps 事后核对方案
langchain·ai agent·工具调用·agent开发·排坑笔记·多工具协同·工程化 质量保障
yu俞娥宝1 天前
Chrome插件开发实战进阶(Manifest V3):通信原理、高级API、工程优化与上架避坑(第二篇)
chrome