面对复杂 PDF 的解析,不是简单地把文件读成文本,而是把版面、图片、表格和扫描件重新还原成可检索、可溯源、可问答的知识, 特别是一个pdf中存在段落、图片(图片中有文字,有公司logo等标识)、表格、PPT 页面甚至整页扫描件 的混合情况 ,相对可靠的步骤是:逐页判别文本型、扫描型与混合型页面,原生文本优先,大图与扫描区域走RapidOCR,表格用pdfplumber 或 PP-Structure保留结构,最后按内容类型切分并写入 Milvus,构建面向 RAG 的复杂文档知识库;
这里给出了一套按"复杂 PDF 解析 → 扫描件 OCR → 结构化存储 → 切分入库 "的可落地方案, 核心思想是:不要用一个 Loader 硬解所有 PDF,而是先判断页面类型,再分流处理, 文本、图片 OCR、表格、PPT 页面分别产出结构化 Document,最后统一切分和入库
一、总体方案
1. PDF 页面分类
用PyMuPDF逐页判断:
文本型页面 :
page.get_text("text")能提取到较多文字扫描型页面:原生文字很少,页面大部分是图片
混合型页面:既有原生文本,又有大图、表格、扫描区域
判断逻辑可参考:
Go
native_text = page.get_text("text").strip()
image_infos = page.get_image_info(xrefs=True)
page_area = page.rect.width * page.rect.height
img_area = sum(
(img["bbox"][2] - img["bbox"][0]) * (img["bbox"][3] - img["bbox"][1])
for img in image_infos
)
is_scan = len(native_text) < 50 and img_area / page_area > 0.5
分类判断说明: -当页面原生文本量极少(少于阈值)且图片面积占比超过 50% 时,判定为扫描型页面; 反之则为文本型页面,优先使用原生文本提取 - 混合型页面(部分文本+部分图片)将在后续分流中分别处理
2. 不同内容处理策略
| 内容类型 | 处理方式 |
|---|---|
| 文案段落 | page.get_text("text") 或 page.get_text("dict") |
| 图片中的文字 | 提取图片,尺寸超过 PDF_OCR_THRESHOLD 才 OCR,调用 get_ocr() |
| 表格 | 文本型 PDF 用 pdfplumber.extract_tables();扫描件用 PP-Structure / PaddleOCR表格识别 |
| PPT 页面 | 若 PDF 由 PPT 导出,按页面解析;若 PDF 内嵌 PPT,用 doc.embfile_get() 提取后走 OCRPPTLoader |
| 扫描件整页 | 整页 300 DPI 渲染,预处理,RapidOCR 整页 OCR,PP-Structure 做版面与表格 |
| 复杂版面 | 用 PP-Structure 做 layout:标题、正文、表格、图片区域,再按阅读顺序合并 |
3. 存储设计
建议分层存储:
原始PDF/图片:MinIO、OSS、S3
文本块、表格 HTML、OCR 文本:Elasticsearch / PostgreSQL / JSON
向量:Milvus
元数据 :
source、page、type、bbox、table_html、ocr_engine、confidence
每个 Document 示例:
Go
{
"page_content": "识别出的文本或表格 Markdown/HTML",
"metadata": {
"source": "complex.pdf",
"page": 3,
"type": "image_ocr",
"bbox": [100, 200, 500, 800],
"engine": "rapidocr_paddle"
}
}
metadata 字段说明:
- source:原始文件路径,用于溯源
- page:来源页码,支持定位到原始 PDF 的具体位置 -
- type:内容类型(text / image_ocr / table / scanned_text),用于后续按类型差异化切分
- bbox:内容在页面中的坐标区域,支持可视化高亮定位
- engine:使用的解析引擎,便于排查和替换
二、复杂 PDF 解析代码案例
下面是一个增强版 ComplexPDFLoader,整合:
PyMuPDF 文本提取
RapidOCR 图片 OCR
pdfplumber 表格提取
扫描页整页 OCR
可选 PP-Structure 版面/表格识别
注意:
fitz是 PyMuPDF,不是pip install fitz
Go
import os
import re
import logging
from typing import Iterator, List
import cv2
import fitz # PyMuPDF
import numpy as np
import pdfplumber
from tqdm import tqdm
from langchain_core.documents import Document
from langchain_core.document_loaders import BaseLoader
# 复用项目中的 OCR 工厂
from rag_qa.edu_document_loaders.edu_ocr import get_ocr
PDF_OCR_THRESHOLD = (0.6, 0.6)
def rotate_img(img, angle):
h, w = img.shape[:2]
center = (w / 2, h / 2)
M = cv2.getRotationMatrix2D(center, angle, 1.0)
new_w = int(h * np.abs(M[0, 1]) + w * np.abs(M[0, 0]))
new_h = int(h * np.abs(M[0, 0]) + w * np.abs(M[0, 1]))
M[0, 2] += (new_w - w) / 2
M[1, 2] += (new_h - h) / 2
return cv2.warpAffine(img, M, (new_w, new_h))
def preprocess_image(img):
"""扫描件预处理:灰度、去噪、自适应二值化。"""
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
gray = cv2.fastNlMeansDenoising(gray, None, 10, 7, 21)
binary = cv2.adaptiveThreshold(
gray, 255,
cv2.ADAPTIVE_THRESH_GAUSSIAN_C,
cv2.THRESH_BINARY,
31, 15
)
return cv2.cvtColor(binary, cv2.COLOR_GRAY2BGR)
def table_to_markdown(table):
"""把 pdfplumber 的表格转成 Markdown,便于检索和展示。"""
if not table:
return ""
rows = []
for row in table:
cells = [
"" if c is None else str(c).replace("\n", " ").strip()
for c in row
]
rows.append("| " + " | ".join(cells) + " |")
if len(rows) >= 2:
sep = "| " + " | ".join(["---"] * len(table[0])) + " |"
return rows[0] + "\n" + sep + "\n" + "\n".join(rows[1:])
return "\n".join(rows)
def extract_tables_pdfplumber(pdf_path, page_idx):
"""文本型 PDF 表格提取,page_idx 从 0 开始。"""
try:
with pdfplumber.open(pdf_path) as pdf:
page = pdf.pages[page_idx]
tables = page.extract_tables()
return [table_to_markdown(t) for t in tables if t]
except Exception as e:
logging.warning(f"pdfplumber 表格提取失败 page={page_idx}: {e}")
return []
def parse_ppstructure(img, use_cuda, source, page_idx):
"""
可选:用 PaddleOCR PP-Structure 做版面分析和表格识别。
适合扫描件、复杂版面。
"""
try:
from paddleocr import PPStructure
engine = PPStructure(
show_log=False,
image_orientation=True,
layout=True,
table=True,
ocr=True,
use_gpu=use_cuda
)
regions = engine(img)
for r in regions:
r_type = r.get("type")
if r_type == "table":
html = r.get("res", {}).get("html", "")
if html:
yield Document(
page_content=html,
metadata={
"source": source,
"page": page_idx,
"type": "table",
"format": "html",
"engine": "ppstructure"
}
)
elif r_type == "text":
res = r.get("res", [])
text = "\n".join(
item.get("text", "")
for item in res
if isinstance(item, dict)
)
if text.strip():
yield Document(
page_content=text,
metadata={
"source": source,
"page": page_idx,
"type": "text",
"engine": "ppstructure"
}
)
except Exception as e:
logging.warning(f"PP-Structure 解析失败 page={page_idx}: {e}")
return
class ComplexPDFLoader(BaseLoader):
def __init__(
self,
file_path: str,
use_cuda: bool = True,
scan_text_threshold: int = 50,
dpi: int = 300,
):
self.file_path = file_path
self.use_cuda = use_cuda
self.scan_text_threshold = scan_text_threshold
self.dpi = dpi
self.ocr = get_ocr(use_cuda=use_cuda)
def lazy_load(self) -> Iterator[Document]:
doc = fitz.open(self.file_path)
pbar = tqdm(total=doc.page_count, desc="ComplexPDF")
for page_idx, page in enumerate(doc):
pbar.set_description(
f"ComplexPDF page {page_idx + 1}/{doc.page_count}"
)
native_text = page.get_text("text").strip()
image_infos = page.get_image_info(xrefs=True)
page_area = page.rect.width * page.rect.height
img_area = sum(
(img["bbox"][2] - img["bbox"][0])
* (img["bbox"][3] - img["bbox"][1])
for img in image_infos
)
is_scan = (
len(native_text) < self.scan_text_threshold
and page_area > 0
and img_area / page_area > 0.5
)
if is_scan:
yield from self._parse_scan_page(page, page_idx)
else:
yield from self._parse_text_page(
doc, page, page_idx, image_infos
)
pbar.update(1)
pbar.close()
def _pix_to_array(self, pix, rotation=0):
arr = np.frombuffer(
pix.samples, dtype=np.uint8
).reshape(pix.height, pix.width, pix.n)
if pix.n == 4:
arr = cv2.cvtColor(arr, cv2.COLOR_RGBA2RGB)
elif pix.n == 1:
arr = cv2.cvtColor(arr, cv2.COLOR_GRAY2RGB)
if int(rotation) != 0:
arr = rotate_img(arr, 360 - rotation)
return arr
def _parse_text_page(self, doc, page, page_idx, image_infos):
source = self.file_path
# 1. 原生文本
native_text = page.get_text("text").strip()
if native_text:
yield Document(
page_content=native_text,
metadata={
"source": source,
"page": page_idx,
"type": "text"
}
)
# 2. 页面内大图 OCR
for img in image_infos:
xref = img.get("xref")
if not xref:
continue
bbox = img["bbox"]
# 小图跳过,避免干扰
if (
(bbox[2] - bbox[0]) / page.rect.width < PDF_OCR_THRESHOLD[0]
or (bbox[3] - bbox[1]) / page.rect.height < PDF_OCR_THRESHOLD[1]
):
continue
pix = fitz.Pixmap(doc, xref)
arr = self._pix_to_array(pix, page.rotation)
result, _ = self.ocr(arr)
if result:
text = "\n".join([line[1] for line in result])
yield Document(
page_content=text,
metadata={
"source": source,
"page": page_idx,
"type": "image_ocr",
"bbox": bbox,
"engine": "rapidocr"
}
)
# 3. 表格提取
for t_idx, table_md in enumerate(
extract_tables_pdfplumber(self.file_path, page_idx)
):
if table_md.strip():
yield Document(
page_content=table_md,
metadata={
"source": source,
"page": page_idx,
"type": "table",
"format": "markdown",
"table_index": t_idx
}
)
def _parse_scan_page(self, page, page_idx):
source = self.file_path
# 整页 300 DPI 渲染
matrix = fitz.Matrix(self.dpi / 72, self.dpi / 72)
pix = page.get_pixmap(matrix=matrix, alpha=False)
img = np.frombuffer(
pix.samples, dtype=np.uint8
).reshape(pix.height, pix.width, pix.n)
if pix.n == 4:
img = cv2.cvtColor(img, cv2.COLOR_RGBA2RGB)
elif pix.n == 3:
img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR)
# 预处理:去噪、二值化
img = preprocess_image(img)
# 整页 OCR
result, _ = self.ocr(img)
if result:
# 按从上到下、从左到右排序
lines = sorted(
result,
key=lambda x: (x[0][0][1], x[0][0][0])
)
text = "\n".join([line[1] for line in lines])
yield Document(
page_content=text,
metadata={
"source": source,
"page": page_idx,
"type": "scanned_text",
"engine": "rapidocr"
}
)
# 可选:PP-Structure 识别表格和版面
yield from parse_ppstructure(
img, self.use_cuda, source, page_idx
)
使用:
Go
loader = ComplexPDFLoader(
file_path="./complex.pdf",
use_cuda=True,
dpi=300
)
docs = loader.load()
for d in docs[:5]:
print(d.metadata)
print(d.page_content[:200])
print("-" * 80)
三、扫描件 PDF 提取方案
如果 PDF 是复杂扫描件,关键步骤:
(1).判断是否为扫描件
原生文本极少,图片面积占页面大部分
(2).整页高分辨率渲染
用
page.get_pixmap(matrix=fitz.Matrix(300/72, 300/72))(3).图像预处理
灰度化
去噪
自适应二值化
倾斜校正
方向分类,RapidOCR 可开启
cls_use_cuda(4).OCR
GPU 环境优先
rapidocr_paddleCPU 或跨平台回退
rapidocr_onnxruntime(5).表格和版面
文本 OCR:RapidOCR进行全文字识别
表格识别:PaddleOCR PP-Structure
版面区域:标题、正文、表格、图片
按阅读顺序合并,保留 bbox
(6).后处理
合并断行
去掉页眉页脚
表格转 HTML/Markdown
图片 OCR 文本单独成块
扫描件整页 OCR 核心代码就是上面的
_parse_scan_page(),如果扫描件表格很多,建议额外启用PPStructure:
Go
engine = PPStructure(
show_log=False,
image_orientation=True,
layout=True,
table=True,
ocr=True,
use_gpu=True
)
regions = engine(img)
四、PPT 的处理
分三种情况:
1. 独立 PPTX
直接你项目里的 OCRPPTLoader:
Go
from rag_qa.edu_document_loaders.edu_pptloader import OCRPPTLoader
ppt_loader = OCRPPTLoader(filepath="./demo.pptx")
ppt_docs = ppt_loader.load()
它会提取:
文本框
表格
图片 OCR
组合形状递归解析
2. PDF 由 PPT 导出
把每一页当成普通 PDF 页面处理即可,走
ComplexPDFLoader,此类 PDF 的每页本质上是一张 PPT 幻灯片渲染图, 处理方式:
- 若导出时保留了文本层(可选中文字),走文本型页面逻辑(page.get_text + pdfplumber)
- 若导出时仅保留图片(不可选中文字),走扫描型页面逻辑(300 DPI 渲染 → OCR → PP-Structure)
- 可通过检查 native_text 是否为空来自动判断走哪条路径
3. PDF 内嵌 PPT 附件
用 PyMuPDF 提取附件,部分 PDF 会将原始 PPTX 文件作为附件嵌入(embfile),需要提取后再处理:
Go
import os
import fitz
def extract_embedded_pptx(pdf_path, out_dir):
doc = fitz.open(pdf_path)
os.makedirs(out_dir, exist_ok=True)
for name in doc.embfile_names():
data = doc.embfile_get(name)
if name.lower().endswith((".ppt", ".pptx")):
path = os.path.join(out_dir, name)
with open(path, "wb") as f:
f.write(data)
print("提取:", path)
然后再用**
OCRPPTLoader**解析
五、复杂PDF图片不仅存在文案, 还存在特殊图片, 如国旗/公司logo等非常重要的标记, 处理方案
如果 PDF 中一张图片里既有文案,又有国旗、公司 logo 等非常重要的标记 ,就不能再把整张图当一个整体处理了, 正确做法是:先对图片做区域级版面分析,把"文字区域"和"特殊标记区域"拆开,分别处理,再按空间关系合并,保留它们之间的关联 , 也就是说,要处理的是"图片中的子区域",而不是"整张图片"
1. 核心思路
一张混合图片可以拆成几类区域:
| 区域类型 | 处理方式 |
|---|---|
| 文字区域 | RapidOCR 识别文字 |
| 国旗区域 | 国旗识别/CLIP,生成"中国国旗" |
| 公司 logo 区域 | Logo 识别/CLIP,生成"苹果公司Logo" |
| 表格区域 | PP-Structure,保留 HTML |
| 照片/图表区域 | 生成 caption |
| 二维码区域 | 解码 |
关键点:
不要整图 OCR:整图 OCR 会把国旗、logo 识别成乱码
先检测子区域 :用 PP-Structure、YOLO、或 CLIP 滑窗
区域分类后分流处理 :文字走 OCR,标记走识别
保留空间关系 :记录每个区域的 bbox,按阅读顺序合并
保留关联 :文字和标记可能共同表达一个意思,比如"国旗 + 国家名称",要能一起被检索到
2. 整体流程
Go
PDF 页面
↓
提取图片
↓
图片区域检测(PP-Structure / YOLO / 连通域)
↓
子区域分类
├── 文字区域 → RapidOCR → 文本
├── 国旗区域 → 国旗识别 → "中国国旗"
├── Logo区域 → Logo识别 → "苹果公司Logo"
├── 表格区域 → PP-Structure → HTML
├── 二维码区域 → 解码
└── 照片区域 → Caption
↓
按 bbox 空间顺序合并
↓
生成结构化 Document
page_content = 区域描述 + OCR文本
metadata = 区域列表、bbox、标签、图片URL
↓
原图存 MinIO/OSS
文本描述存 ES
文本向量 + 图像向量存 Milvus
3. 区域检测方案
方案 A:PP-Structure(推荐)
PaddleOCR 的 PP-Structure 可以直接做版面分析:
Go
from paddleocr import PPStructure
engine = PPStructure(
show_log=False,
image_orientation=True,
layout=True,
table=True,
ocr=True,
use_gpu=True
)
regions = engine(img)
返回的 regions 会包含:
Go
[
{"type": "text", "bbox": [...], "res": [...]},
{"type": "figure", "bbox": [...], "res": [...]},
{"type": "table", "bbox": [...], "res": {...}},
{"type": "title", "bbox": [...], "res": [...]},
]
figure类型就是图片区域,可以进一步分类成国旗、logo、照片等
方案 B:CLIP 滑窗
如果没有 PP-Structure,可以用 CLIP 在图片上滑窗:
Go
def sliding_window_classify(pil_img, window=224, stride=112):
results = []
w, h = pil_img.size
for y in range(0, h, stride):
for x in range(0, w, stride):
box = (x, y, min(x + window, w), min(y + window, h))
crop = pil_img.crop(box)
label, score = classify_image(crop)
if score > 0.6:
results.append({
"bbox": box,
"label": label,
"score": score
})
return results
方案 C:YOLO 目标检测
训练或使用预训练 YOLO 检测国旗、logo:
Go
from ultralytics import YOLO
model = YOLO("yolov8n.pt")
results = model(np.array(pil_img))
for r in results:
for box in r.boxes:
cls = int(box.cls)
conf = float(box.conf)
xyxy = box.xyxy.tolist()[0]
print(cls, conf, xyxy)
4. 完整代码案例
下面是一个处理"图片中既有文案又有国旗/logo"的完整函数
(1). 图片区域检测 + 分类
Go
import cv2
import numpy as np
from PIL import Image
from langchain_core.documents import Document
from rag_qa.edu_document_loaders.edu_ocr import get_ocr
import torch
import clip
device = "cuda" if torch.cuda.is_available() else "cpu"
clip_model, clip_preprocess = clip.load("ViT-B/32", device=device)
IMAGE_LABELS = [
"文字区域",
"国旗",
"公司logo",
"表格",
"二维码",
"自然照片",
"图表",
"普通插图"
]
text_tokens = clip.tokenize(IMAGE_LABELS).to(device)
FLAG_LABELS = [
"中国国旗", "美国国旗", "日本国旗", "英国国旗",
"法国国旗", "德国国旗", "俄罗斯国旗", "印度国旗",
"巴西国旗", "其他国家国旗"
]
flag_tokens = clip.tokenize(FLAG_LABELS).to(device)
LOGO_LABELS = [
"苹果公司Logo", "微软Logo", "谷歌Logo", "亚马逊Logo",
"腾讯Logo", "阿里巴巴Logo", "字节跳动Logo", "其他公司Logo"
]
logo_tokens = clip.tokenize(LOGO_LABELS).to(device)
def classify_region(pil_crop):
image_input = clip_preprocess(pil_crop).unsqueeze(0).to(device)
with torch.no_grad():
logits, _ = clip_model(image_input, text_tokens)
probs = logits.softmax(dim=-1).cpu().numpy()[0]
idx = int(probs.argmax())
return IMAGE_LABELS[idx], float(probs[idx])
def recognize_flag(pil_crop):
image_input = clip_preprocess(pil_crop).unsqueeze(0).to(device)
with torch.no_grad():
logits, _ = clip_model(image_input, flag_tokens)
probs = logits.softmax(dim=-1).cpu().numpy()[0]
return FLAG_LABELS[int(probs.argmax())]
def recognize_logo(pil_crop):
image_input = clip_preprocess(pil_crop).unsqueeze(0).to(device)
with torch.no_grad():
logits, _ = clip_model(image_input, logo_tokens)
probs = logits.softmax(dim=-1).cpu().numpy()[0]
return LOGO_LABELS[int(probs.argmax())]
(2). 区域检测
Go
def detect_regions(pil_img):
"""
用 PP-Structure 或滑窗检测图片中的子区域。
返回 [{"bbox": (x1,y1,x2,y2), "crop": PIL.Image}, ...]
"""
try:
from paddleocr import PPStructure
engine = PPStructure(
show_log=False,
layout=True,
table=False,
ocr=False,
use_gpu=torch.cuda.is_available()
)
regions = engine(np.array(pil_img))
result = []
for r in regions:
x1, y1, x2, y2 = r["bbox"]
crop = pil_img.crop((x1, y1, x2, y2))
result.append({
"bbox": (x1, y1, x2, y2),
"crop": crop,
"pp_type": r.get("type", "unknown")
})
return result
except Exception:
# 回退:CLIP 滑窗
return sliding_window_detect(pil_img)
def sliding_window_detect(pil_img, window=224, stride=112):
w, h = pil_img.size
regions = []
for y in range(0, h, stride):
for x in range(0, w, stride):
x2 = min(x + window, w)
y2 = min(y + window, h)
crop = pil_img.crop((x, y, x2, y2))
label, score = classify_region(crop)
if score > 0.6:
regions.append({
"bbox": (x, y, x2, y2),
"crop": crop,
"pp_type": label
})
return regions
(3). 区域分流处理
Go
def process_region(region, ocr, page_idx, source):
crop = region["crop"]
bbox = region["bbox"]
label, score = classify_region(crop)
content_parts = []
tags = []
if label == "文字区域":
result, _ = ocr(np.array(crop))
if result:
text = "\n".join([line[1] for line in result])
content_parts.append(text)
tags.append("文字")
elif label == "国旗":
country = recognize_flag(crop)
content_parts.append(f"图片中有一面{country}。")
tags.extend(["国旗", country])
elif label == "公司logo":
brand = recognize_logo(crop)
content_parts.append(f"图片中有一个{brand}。")
tags.extend(["公司logo", brand])
elif label == "二维码":
try:
from pyzbar.pyzbar import decode
decoded = decode(np.array(crop))
qr_text = "\n".join([d.data.decode("utf-8") for d in decoded])
content_parts.append(f"二维码内容:{qr_text}")
tags.append("二维码")
except Exception:
content_parts.append("图片包含二维码,但解码失败。")
tags.append("二维码")
elif label in ["自然照片", "图表", "普通插图"]:
caption = generate_caption(crop)
content_parts.append(f"图片描述:{caption}")
tags.append(label)
else:
result, _ = ocr(np.array(crop))
if result:
text = "\n".join([line[1] for line in result])
content_parts.append(text)
tags.append("未知区域")
return {
"bbox": bbox,
"label": label,
"score": score,
"content": "\n".join(content_parts),
"tags": tags
}
4. 合并区域,生成 Document
Go
def process_mixed_image(pil_img, ocr, page_idx, source, image_url):
regions = detect_regions(pil_img)
# 按从上到下、从左到右排序
regions_sorted = sorted(
regions,
key=lambda r: (r["bbox"][1], r["bbox"][0])
)
processed = []
for r in regions_sorted:
processed.append(process_region(r, ocr, page_idx, source))
# 合并内容
full_text = []
all_tags = []
region_meta = []
for p in processed:
if p["content"].strip():
full_text.append(p["content"])
all_tags.extend(p["tags"])
region_meta.append({
"bbox": p["bbox"],
"label": p["label"],
"score": p["score"],
"tags": p["tags"]
})
page_content = "\n".join(full_text)
metadata = {
"source": source,
"page": page_idx,
"type": "mixed_image",
"image_url": image_url,
"tags": list(set(all_tags)),
"regions": region_meta
}
return Document(page_content=page_content, metadata=metadata)
(5). 在 ComplexPDFLoader 中调用
Go
def _parse_text_page(self, doc, page, page_idx, image_infos):
source = self.file_path
native_text = page.get_text("text").strip()
if native_text:
yield Document(
page_content=native_text,
metadata={"source": source, "page": page_idx, "type": "text"}
)
for img in image_infos:
xref = img.get("xref")
if not xref:
continue
bbox = img["bbox"]
if (
(bbox[2] - bbox[0]) / page.rect.width < PDF_OCR_THRESHOLD[0]
or (bbox[3] - bbox[1]) / page.rect.height < PDF_OCR_THRESHOLD[1]
):
continue
pix = fitz.Pixmap(doc, xref)
img_array = self._pix_to_array(pix, page.rotation)
pil_img = Image.fromarray(
cv2.cvtColor(img_array, cv2.COLOR_BGR2RGB)
)
image_url = save_image_to_storage(pil_img, page_idx, xref)
doc = process_mixed_image(
pil_img=pil_img,
ocr=self.ocr,
page_idx=page_idx,
source=source,
image_url=image_url
)
yield doc
5. 生成结果示例
假设 PDF 第 5 页有一张图,里面有:
一段文案:"本报告由 ABC 公司发布"
一面中国国旗
一个 ABC 公司 logo
处理后的 Document:
Go
Document(
page_content="""
图片中有一面中国国旗,
图片中有一个ABC公司Logo,
本报告由 ABC 公司发布
""",
metadata={
"source": "report.pdf",
"page": 5,
"type": "mixed_image",
"image_url": "http://minio/rag/report_p5_img1.png",
"tags": ["国旗", "中国国旗", "公司logo", "ABC公司Logo", "文字"],
"regions": [
{"bbox": [10, 10, 80, 60], "label": "国旗", "tags": ["中国国旗"]},
{"bbox": [100, 10, 200, 80], "label": "公司logo", "tags": ["ABC公司Logo"]},
{"bbox": [10, 100, 500, 200], "label": "文字区域", "tags": ["文字"]}
]
}
)
用户检索时:
问"文档里出现了哪国国旗?" → 命中
中国国旗问"ABC 公司 logo 出现在哪?" → 命中
公司logo + ABC公司Logo问"报告是谁发布的?" → 命中 OCR 文本
本报告由 ABC 公司发布想看原图 → 返回
image_url
6.存储与检索设计
(1). 分层存储
| 内容 | 存储 |
|---|---|
| 原图 | MinIO / OSS / S3 |
| 文本内容 | Elasticsearch / PostgreSQL |
| 区域元数据 | JSON 字段或独立表 |
| 文本向量 | Milvus |
| 图像向量 | Milvus 多模态集合 |
(2). Milvus 多向量策略
可以建两个集合:
text_collection:存 OCR 文本、caption 的文本向量
image_collection:存整图 CLIP 向量,支持"以图搜图"检索时:
文本问题 → 查
text_collection图像相似 → 查
image_collection混合问题 → 两路召回后融合
(3). 元数据过滤
Milvus 支持按 metadata 过滤:
Go
expr = 'tags like "%国旗%" and page == 5'
这样用户问"第 5 页的国旗是什么",可以直接过滤
7.关键注意事项
不要整图 OCR
国旗、logo 会被识别成乱码,污染正文
先检测区域,再分类
区域级处理比整图分类更准确
保留空间关系
记录 bbox,按阅读顺序合并,避免语义错乱
文字和标记要关联
比如"国旗 + 国家名称"放在同一个 Document 里,检索时才能一起命中
原图必须单独存储
返回
image_url,方便溯源和展示多模态检索
文本向量 + 图像向量,支持"以文搜图"和"以图搜图"
置信度过滤
低置信度的 OCR 和分类结果不要直接入库,避免噪声
当 PDF 图片中同时存在文案和国旗/logo 等重要标记时,正确方案是:
Go
整图 → 区域检测 → 区域分类 → 分流处理 → 空间合并 → 结构化存储
文字区域 → RapidOCR
国旗区域 → 国旗识别
Logo 区域 → Logo 识别
表格区域 → PP-Structure
二维码 → 解码
照片/图表 → Caption
最终生成带
bbox、tags、image_url的结构化 Document,文本向量和图像向量分别入 Milvus,既保留文字信息,也保留国旗、logo 这类重要标记的语义,支持精准检索和多模态问答
六、文本切分与 Milvus 入库
解析后得到 Document 列表,建议按类型切分:
普通文本、OCR 文本:用
ChineseRecursiveTextSplitter按语义段落切分,保留分隔符表格:保留整块,或按行转 Markdown, 避免破坏表格行列结构
图片 OCR: 按图片粒度保留,不跨图切分
长文档:可考虑
AliTextSplitter语义切分,但 CPU 较慢
Go
from rag_qa.edu_text_spliter.edu_chinese_recursive_text_splitter import (
ChineseRecursiveTextSplitter
)
loader = ComplexPDFLoader("./complex.pdf", use_cuda=True)
docs = loader.load()
splitter = ChineseRecursiveTextSplitter(
keep_separator=True,
is_separator_regex=True,
chunk_size=500,
chunk_overlap=80
)
chunks = []
for d in docs:
if d.metadata.get("type") in ("text", "scanned_text", "image_ocr"):
chunks.extend(splitter.split_documents([d]))
else:
chunks.append(d) # 表格整块保留
存 Milvus:
Go
from langchain_milvus import Milvus
from langchain_community.embeddings import HuggingFaceBgeEmbeddings
embeddings = HuggingFaceBgeEmbeddings(
model_name="BAAI/bge-base-zh-v1.5",
encode_kwargs={"normalize_embeddings": True}
)
vector_store = Milvus.from_documents(
chunks,
embeddings,
connection_args={"host": "localhost", "port": "19530"},
collection_name="complex_pdf_rag"
)
原始 PDF 和图片可放 MinIO:
Go
# from minio import Minio
# client = Minio("localhost:9000", access_key="minio", secret_key="minio123", secure=False)
# client.fput_object("rag", "complex.pdf", "./complex.pdf")
七、推荐最终流程
Go
复杂 PDF
↓
逐页判断:文本型 / 扫描型 / 混合型
↓
文本型:
page.get_text() + 大图 OCR + pdfplumber 表格
↓
扫描型:
300 DPI 整页渲染 → 预处理 → RapidOCR → PP-Structure 表格/版面
↓
混合型:
文本 + 图片 OCR + 表格 + 扫描区域 OCR 合并
↓
生成结构化 Document
metadata: source/page/type/bbox/table_html/engine
↓
按类型切分:中文递归切分 / 语义切分 / 表格不切
↓
存储:
原始文件 → MinIO/OSS
文本表格 → ES/PostgreSQL
向量 → Milvus
关键点:
文本型 PDF 不要整页 OCR,优先原生文本,只对大图 OCR
扫描件必须整页 OCR,并做预处理和版面分析
表格不要只 OCR 成纯文本,尽量保留 HTML/Markdown
图片 OCR 文本要带 page、bbox、engine 元数据,方便溯源
切分时按内容类型区别处理,表格整块保留,文本用中文递归切分器。
最终存入 Milvus 时保留 metadata,支持按来源、页码、类型过滤检索
八、核心设计思想总结
1. 为什么不能用一个 Loader 硬解所有 PDF?
复杂 PDF 的本质是多模态文档的容器------同一份文件中可能同时包含:
- 原生文本段落(可复制搜索)
- 图片中的文字(需 OCR 还原)
- 结构化表格(需保留行列关系)
- PPT 幻灯片(需特殊解析)
- 整页扫描件(需完整 OCR 流水线)
用一个通用的 Loader 强行统一处理,必然导致:文本型页面被错误 OCR(浪费算力且丢失精确文本)、表格结构被破坏、扫描件无法识别
2. 核心原则
|--------------|----------------------------------------------------|
| 原则 | 说明 |
| 逐页判别 | 先判断页面类型(文本/扫描/混合),再选择处理策略 |
| 原生文本优先 | 能直接提取文本就不走 OCR,保证 100% 准确率 |
| 大图与扫描区域走 OCR | 图片中的文字和扫描区域用 RapidOCR 还原 |
| 表格保留结构 | 用 pdfplumber(文本型)或 PP-Structure(扫描型)保留表格结构 |
| 按类型切分 | 文本切分、表格不切、图片不跨图切分 |
| 完整溯源 | 每个 chunk 携带 source/page/type/bbox 元数据,可定位到原始位置 |
3. 最终目标
构建面向 RAG 的复杂文档知识库,使得:
- 可检索:文本、表格、OCR 内容统一向量化,支持语义搜索
- 可溯源:每个 chunk 可定位到原始 PDF 的页码和坐标区域
- 可问答:结构化表格和段落保持语义完整,支持精准问答
- 可扩展:模块化设计,新增页面类型或 OCR 引擎只需扩展对应分支