1.3 手写 PDF 内容流解析器:从字节流到结构化水印候选
清印 ClearMark 系列第 3 篇 · 共 12 篇
前两篇讲了系统的整体架构与处理器路由。从这篇开始,我们要进入项目最硬核的部分------PDF 矢量水印的检测与移除。
这一篇专门讲 内容流解析器 (content stream parser)。它是整个项目的算法引擎,也是技术难度最高的一块。市面上讲 PDF 内容流的中文资料极少,能找到的也多是 PDF 规范的翻译,缺乏实战视角。我会从 PDF 内容流的语法讲起,一步步拆解我们的
core/content_stream.py是怎么把一段字节流解析成结构化文字块和图片块的,又怎么从中提取旋转角、透明度、颜色等水印特征。读完这篇,你应该能自己写一个基础的内容流解析器,理解 PDF 是怎么把"在 (100, 200) 画一个 36 号字号的'机密'"这种指令编码成字节流的。
一、为什么必须自己写内容流解析器
先回答一个绕不开的问题:PyMuPDF 不是已经提供了 page.get_texttrace()、page.get_text("dict") 这些 API 吗?为什么还要自己写?
答案是 字节偏移。
我们的移除策略之一是 内容流区间擦除 ------找到水印文字在内容流字节流中的 [start, end) 区间,直接从字节流里删掉这段,然后写回 PDF。这是无损移除,不会破坏 PDF 其他结构。要做到这件事,必须知道每个文字块在原始字节流里的精确字节偏移。
PyMuPDF 的 get_texttrace() 返回的是高级结构化数据(文字、坐标、字体),但 不暴露字节偏移 。get_text("dict") 更不行,它甚至做了进一步的语义化封装。
至于直接用 page.get_contents() 拿到原始字节流然后自己解析------这才是唯一可行的路线。这也就是 core/content_stream.py 存在的原因。
二、PDF 内容流是什么
PDF 的页面内容(文字、图片、矢量绘图)存放在 内容流(content stream) 里。它是一段按 PDF 操作符语法编码的字节流,通常经过 FlateDecode 压缩。PyMuPDF 的 page.read_contents() 会自动解压并返回原始字节流。
内容流的语法非常简单------操作数 + 操作符 的序列。例如:
BT
/F1 36 Tf
1 0 0 1 100 200 Tm
(机密) Tj
ET
这段内容流的含义是:
BT:Begin Text,开始一个文字块/F1 36 Tf:设置字体为 F1,字号 361 0 0 1 100 200 Tm:设置文字矩阵(translate 到 (100, 200))(机密) Tj:在当前位置画字符串"机密"ET:End Text,结束文字块
操作符都是字母(BT、Tf、Tj 等),操作数在前、操作符在后。这与 PostScript 一脉相承。常见的水印相关操作符有:
| 操作符 | 含义 | 与水印的关系 |
|---|---|---|
BT / ET |
文字块起止 | 文字水印的基本容器 |
Tf |
设置字体和字号 | 字号是判断水印的重要依据 |
Tm |
设置文字矩阵 | 旋转角、缩放都藏在矩阵里 |
Td / TD |
文字位置偏移 | 多行水印的行距 |
Tj / TJ / ' / " |
画文字 | 实际的文字内容 |
rg / RG |
设置填充/描边颜色 | 灰色半透明水印的特征 |
gs |
设置图形状态 | 透明度(ca/CA)藏在 ExtGState 里 |
ca / CA |
直接透明度 | 某些 PDF 会内联透明度 |
q / Q |
保存/恢复图形状态 | 图片块的容器 |
cm |
变换矩阵 | 图片位置、旋转、缩放 |
Do |
调用 XObject | 图片绘制 |
理解这些操作符,是写解析器的前提。完整的 PDF 规范有 700 多页,但内容流相关操作符只占其中一小部分。
三、分词器(Tokenizer)
字节流解析的第一步是分词。我们的分词器在 core/content_stream.py:
python
# 操作符由字母组成
_OP_CHARS = set(
"abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ*'-+\""
)
_NUM_RE = re.compile(rb'[+-]?(?:\d+\.?\d*|\.\d+)(?:[eE][+-]?\d+)?')
_NAME_RE = re.compile(rb'/[^\x00\t\n\r\f ()<>\[\]{}/%]+')
_HEX_STR_RE = re.compile(rb'<[0-9A-Fa-f\s]*>')
_OP_RE = re.compile(rb"[a-zA-Z*'\"]+")
def _tokenize(data: bytes):
"""生成 (token_type, value, start, end) 元组"""
i = 0
n = len(data)
while i < n:
ch = data[i:i + 1]
# 跳过空白
if ch in b' \t\n\r\f\x00':
i += 1
continue
# 跳过注释
if ch == b'%':
nl = data.find(b'\n', i)
if nl == -1:
break
i = nl + 1
continue
start = i
# 数字
if ch in b'+-.' or ch.isdigit():
m = _NUM_RE.match(data, i)
if m:
raw = m.group()
try:
val = float(raw)
if b'.' not in raw and b'e' not in raw and b'E' not in raw:
val = int(raw)
except ValueError:
val = 0.0
yield ('num', val, start, m.end())
i = m.end()
continue
# 名称 /Name
if ch == b'/':
m = _NAME_RE.match(data, i)
if m:
yield ('name', m.group()[1:], start, m.end())
i = m.end()
continue
# 十六进制字符串 <...>
if ch == b'<':
if data[i:i + 2] == b'<<':
yield ('dict_open', b'<<', start, start + 2)
i += 2
continue
m = _HEX_STR_RE.match(data, i)
if m:
yield ('hexstr', m.group(), start, m.end())
i = m.end()
continue
yield ('lt', b'<', start, start + 1)
i += 1
continue
if ch == b'>':
if data[i:i + 2] == b'>>':
yield ('dict_close', b'>>', start, start + 2)
i += 2
continue
yield ('gt', b'>', start, start + 1)
i += 1
continue
# 括号字符串 (...)
if ch == b'(':
depth = 1
j = i + 1
while j < n and depth > 0:
if data[j:j + 1] == b'\\':
j += 2
continue
if data[j:j + 1] == b'(':
depth += 1
elif data[j:j + 1] == b')':
depth -= 1
j += 1
yield ('str', data[i + 1:j - 1], start, j)
i = j
continue
# 数组
if ch == b'[':
yield ('arr_open', b'[', start, start + 1)
i += 1
continue
if ch == b']':
yield ('arr_close', b']', start, start + 1)
i += 1
continue
# 操作符(字母序列)
if ch in _OP_CHARS or ch.isalpha():
m = _OP_RE.match(data, i)
if m and m.group():
op = m.group().decode('latin-1')
yield ('op', op, start, m.end())
i = m.end()
continue
# 未知字符,跳过
i += 1
分词器是一个生成器,逐个产出 token,每个 token 是 (type, value, start, end) 四元组。start/end 是字节偏移,后续移除时要用。
3.1 几个分词细节
字符串 (...) :PDF 的字符串用圆括号包裹,支持嵌套 (((a)) 是合法的),支持转义 (\(、\)、\n 等)。所以分词时必须做深度计数和转义跳过:
python
if ch == b'(':
depth = 1
j = i + 1
while j < n and depth > 0:
if data[j:j + 1] == b'\\':
j += 2 # 跳过转义字符
continue
if data[j:j + 1] == b'(':
depth += 1
elif data[j:j + 1] == b')':
depth -= 1
j += 1
yield ('str', data[i + 1:j - 1], start, j)
如果只匹配第一个 ),遇到 (1(2)3) 这种字符串会截断错位,后续整个解析全崩。
字典 <<...>> :PDF 的字典用双尖括号包裹,必须先检查 << 再考虑单个 <(单个 < 是十六进制字符串的开始)。这个顺序在分词器里很关键,写反了会把所有字典都识别成 hexstr。
数字 vs 操作符 :数字以 +-. 或数字开头,操作符以字母开头。但 PDF 规范里有些操作符包含非字母字符(如 '、"),所以 _OP_CHARS 把它们也加进来了。
四、操作分组
分词器产出的是 token 序列,但 PDF 操作是"操作数 + 操作符"的组合。我们需要把连续的数字、字符串、名称这些"操作数"和后续的"操作符"组合成一条 Operation:
python
@dataclass
class Operation:
"""内容流中的一条指令:操作数 + 操作符"""
operator: str
operands: list
start: int # 在原始字节流中的起始偏移
end: int # 在原始字节流中的结束偏移
def parse_content_stream(data: bytes) -> ParsedContent:
"""解析内容流,返回结构化的文字块和图片块列表"""
tokens = list(_tokenize(data))
operations = []
operands = []
op_start = 0
for ttype, tval, tstart, tend in tokens:
if ttype == 'op':
operations.append(Operation(
operator=tval,
operands=operands[:],
start=op_start if operands else tstart,
end=tend,
))
operands = []
op_start = tend
else:
if not operands:
op_start = tstart
operands.append((ttype, tval))
...
关键点:Operation.start 是第一个操作数的起始偏移 ,end 是操作符的结束偏移。这样定义后,一条操作的字节范围就是 [op.start, op.end)。
后续移除水印时,要删除的是整个 BT...ET 文字块或 q...Q 图片块,所以还要把它们再组合成块。
五、文字块提取:从 BT...ET 到 TextBlock
文字块的提取在 _extract_text_blocks:
python
@dataclass
class TextBlock:
"""一个 BT...ET 文字块"""
start: int # BT 的字节偏移
end: int # ET 之后的字节偏移
origin: tuple # (x, y) 文字基点(用户空间坐标)
bbox: tuple # (x0, y0, x1, y1) 近似包围盒
text: str # 解码出的文本
font: str # 字体名
font_size: float # 字号
color: int = 0 # 颜色值(0xRRGGBB)
raw_bytes: bytes = b'' # 原始字节
rotation: float = 0.0 # 旋转角度(度),从 Tm 矩阵提取
alpha: float = 1.0 # 透明度 0~1(从 gs/ca 指令提取)
color_rgb: tuple = (0.0, 0.0, 0.0) # RGB 分量 0~1
is_gray: bool = False # 是否为灰色/低饱和(水印特征)
tm: tuple = (1.0, 0.0, 0.0, 1.0, 0.0, 0.0) # 完整 Tm 矩阵
提取逻辑核心循环:
python
def _extract_text_blocks(ops: list, data: bytes) -> list:
blocks = []
i = 0
while i < len(ops):
if ops[i].operator == 'BT':
bt_start = ops[i].start
bt_idx = i
et_idx = -1
for j in range(i + 1, len(ops)):
if ops[j].operator == 'ET':
et_idx = j
break
if et_idx == -1:
i += 1
continue
text_parts = []
font_name = ''
font_size = 0.0
color_val = 0
color_rgb = (0.0, 0.0, 0.0)
origin = (0.0, 0.0)
tm = (1.0, 0.0, 0.0, 1.0, 0.0, 0.0)
tx, ty = 0.0, 0.0
alpha = 1.0
gs_name = ''
for k in range(bt_idx + 1, et_idx):
op = ops[k]
if op.operator == 'Tf':
if len(op.operands) >= 2:
font_name = str(op.operands[0][1])
font_size = float(op.operands[1][1])
elif op.operator == 'Tm':
nums = [float(o[1]) for o in op.operands if o[0] == 'num']
if len(nums) >= 6:
tm = tuple(nums[:6])
tx, ty = 0.0, 0.0
origin = (tm[4], tm[5])
elif op.operator in ('Td', 'TD'):
nums = [float(o[1]) for o in op.operands if o[0] == 'num']
if len(nums) >= 2:
tx += nums[0]
ty += nums[1] if op.operator == 'Td' else -nums[1]
origin = (tm[4] + tx, tm[5] + ty)
elif op.operator in ('Tj', "'"):
text_parts.append(_decode_text_operand(op.operands[-1]))
elif op.operator == 'TJ':
for operand in op.operands:
if operand[0] == 'arr_open':
pass
elif operand[0] == 'str':
text_parts.append(_decode_text_operand(operand))
elif operand[0] == 'hexstr':
text_parts.append(_decode_text_operand(operand))
elif op.operator == '"':
text_parts.append(_decode_text_operand(op.operands[-1]))
elif op.operator in ('rg', 'RG', 'scn', 'SCN', 'sc', 'SC'):
nums = [float(o[1]) for o in op.operands if o[0] == 'num']
color_val = _nums_to_color(nums)
if len(nums) >= 3:
color_rgb = (nums[0], nums[1], nums[2])
elif len(nums) == 1:
color_rgb = (nums[0], nums[0], nums[0])
elif op.operator == 'gs':
if op.operands:
gs_name = str(op.operands[-1][1])
elif op.operator in ('ca', 'CA'):
nums = [float(o[1]) for o in op.operands if o[0] == 'num']
if nums:
alpha = nums[0]
...
这段代码做的事:扫描 BT 到 ET 之间的所有操作,累积出 text_parts、font_name、font_size、tm、color_rgb、alpha 等属性。下面分别讲几个关键点。
5.1 从 Tm 矩阵提取旋转角
PDF 的文字矩阵 Tm 是 6 个数 (a, b, c, d, e, f),对应 3×3 仿射变换矩阵:
| a b 0 |
| c d 0 |
| e f 1 |
其中 (e, f) 是平移(即文字基点 origin),(a, b, c, d) 是旋转 + 缩放。
对于纯旋转(无缩放)的情况,矩阵是:
| cos(θ) sin(θ) 0 |
| -sin(θ) cos(θ) 0 |
| e f 1 |
所以旋转角 θ = atan2(b, a)。代码:
python
# 从 Tm 提取旋转角
a, b = tm[0], tm[1]
rotation = math.degrees(math.atan2(b, a)) if (a or b) else 0.0
这个角度对水印检测至关重要。正文文字通常旋转角是 0(轴对齐),而水印文字经常是 30°/45°/60° 斜铺。所以"非轴对齐的旋转角"是一个强水印特征。
我们定义了一个辅助函数判断:
python
def _is_non_axis_rotation(rotation: float) -> bool:
"""判断旋转角是否非轴对齐(不是 0/90/180/270 附近)"""
norm = abs(rotation) % 180
# 接近 0 或 90 视为轴对齐
return not (norm < 5 or norm > 175 or 85 < norm < 95)
容差 5°------接近 0° 或 90° 的都视为轴对齐,不当作水印旋转特征。这个阈值是经验值,太严会漏检,太松会把倾斜扫描页误判。
5.2 灰色与透明度:水印的视觉特征
水印文字的视觉特征通常是:灰色(低饱和)+ 半透明。我们提取这两个特征:
python
# 灰色/低饱和判断(水印常见特征)
r, g, bb = color_rgb
is_gray = (abs(r - g) < 0.1 and abs(g - bb) < 0.1 and
0.3 < r < 0.9) # 灰色且非纯黑白
abs(r - g) < 0.1 检查 RGB 三通道差异------纯灰色的 R=G=B。0.3 < r < 0.9 排除纯黑(正文)和纯白(背景)。
透明度 alpha 从两个地方提取:
gs操作符:引用一个 ExtGState 字典,字典里有ca字段。这种情况需要额外查 ExtGState 字典。ca/CA操作符:直接设置透明度。某些 PDF 不走 ExtGState,直接内联。
我们的解析器只处理了内联的 ca/CA。对于走 gs 的情况,在检测阶段再补充查询(PdfVectorProcessor._boost_by_visual_features 里会查 ExtGState)。这是分层的------解析器只做内容流本身能解决的部分,跨对象查询留给处理器。
5.3 文字宽度估算:中文水印红框框不全的坑
这是踩过的一个坑。早期版本用 0.6 * font_size * len(text) 估算文字宽度,对中文水印红框框不全------3 个汉字只框到 2 个。
原因是 中文是全角字符,宽度约等于 font_size 本身,不是 0.6 倍。修复后的版本:
python
def _estimate_text_width(text: str, fs: float) -> float:
"""按字符类别估算文字宽度
旧实现统一按 0.6fs/字,对全角中文(约 1.0fs/字)严重偏窄,
导致中文水印候选红框框不全。
"""
import unicodedata
w = 0.0
for ch in text:
if ch in (' ', '\u00a0'):
w += fs * 0.3
elif unicodedata.east_asian_width(ch) in ('W', 'F'):
# 全角字符:CJK 汉字、假名、全角标点等
w += fs * 1.0
elif ch.isdigit():
w += fs * 0.55
elif ch.isalpha():
w += fs * 0.55
else:
w += fs * 0.5
return w
用 unicodedata.east_asian_width() 判断字符是全角(W/F)还是半角。中文按 1.0fs、英文/数字按 0.55fs、空格按 0.3fs。
不过这个估算终究是粗略的,TJ 字距、字体内置宽度等无法精确还原。所以 PdfVectorProcessor._refine_text_bboxes 会在解析完成后,调用 page.search_for(text) 用 PyMuPDF 的字形布局信息修正 bbox。这是一个"先粗略估算,再精确修正"的两阶段策略。
5.4 文本解码:UTF-8/GBK/Latin-1 三连试
PDF 文本字符串的编码是个老大难问题。PDF 规范允许字符串使用:
- PDFDocEncoding(基本上是 Latin-1 的超集)
- UTF-16BE(带 BOM)
- 字体自定义编码(Type1 字体可能有自己的 Encoding 字典)
实际遇到的水印文本,90% 是中文,编码可能是 UTF-8、GBK、UTF-16BE、Latin-1。我们的解码逻辑:
python
def _decode_text_operand(operand) -> str:
"""解码文字操作数(字符串或十六进制字符串)"""
ttype, tval = operand[0], operand[1]
if ttype == 'str':
# 括号字符串:尝试 latin-1 和 gbk 解码
if isinstance(tval, bytes):
for enc in ('utf-8', 'gbk', 'latin-1'):
try:
return tval.decode(enc)
except (UnicodeDecodeError, ValueError):
continue
return tval.decode('latin-1', errors='replace')
return str(tval)
elif ttype == 'hexstr':
# 十六进制字符串:<C9A8C3E8...>
if isinstance(tval, bytes):
hex_str = tval[1:-1] if tval[:1] == b'<' else tval
hex_str = hex_str.replace(b' ', b'').replace(b'\n', b'').replace(b'\r', b'')
try:
raw = bytes.fromhex(hex_str.decode('ascii'))
except ValueError:
return ''
for enc in ('gbk', 'utf-16-be', 'latin-1'):
try:
return raw.decode(enc)
except (UnicodeDecodeError, ValueError):
continue
return raw.decode('latin-1', errors='replace')
return str(tval)
return ''
依次尝试 UTF-8 → GBK → Latin-1,谁先成功用谁。UTF-16BE 单独处理(hexstr 形式)。
这个策略不完美------如果一段 UTF-8 字符串碰巧也能用 GBK 解码(虽然结果是错的),会选错编码。但在实际水印场景里,绝大多数中文 PDF 都是 GBK 或 UTF-8,这个启发式已经够用。完美方案需要查字体的 Encoding 字典,但成本太高。
六、图片块提取:从 q...Q 到 ImageBlock
图片块的提取逻辑类似,但容器是 q...Q(保存/恢复图形状态):
python
@dataclass
class ImageBlock:
"""一个 q...Q 图片绘制块"""
start: int # q 的字节偏移
end: int # Q 之后的字节偏移
origin: tuple # (x, y) 图片左下角
size: tuple # (width, height) 图片在页面上的尺寸
xobj_name: str = '' # XObject 名称
raw_bytes: bytes = b''
rotation: float = 0.0
alpha: float = 1.0
cm: tuple = (1.0, 0.0, 0.0, 1.0, 0.0, 0.0) # 完整变换矩阵
提取循环:
python
def _extract_image_blocks(ops: list, data: bytes) -> list:
blocks = []
i = 0
while i < len(ops):
if ops[i].operator == 'q':
q_start = ops[i].start
q_idx = i
depth = 1
q_end_idx = -1
has_do = False
for j in range(i + 1, len(ops)):
if ops[j].operator == 'q':
depth += 1
elif ops[j].operator == 'Q':
depth -= 1
if depth == 0:
q_end_idx = j
break
elif ops[j].operator == 'Do':
has_do = True
if q_end_idx == -1:
i += 1
continue
if has_do:
xobj_name = ''
cur_matrix = (1.0, 0.0, 0.0, 1.0, 0.0, 0.0)
alpha = 1.0
for k in range(q_idx + 1, q_end_idx):
op = ops[k]
if op.operator == 'cm':
nums = [float(o[1]) for o in op.operands if o[0] == 'num']
if len(nums) >= 6:
new_m = tuple(nums[:6])
cur_matrix = _mat_mul(new_m, cur_matrix)
elif op.operator == 'Do':
if op.operands:
xobj_name = str(op.operands[-1][1])
elif op.operator in ('ca', 'CA'):
nums = [float(o[1]) for o in op.operands if o[0] == 'num']
if nums:
alpha = nums[0]
origin = (cur_matrix[4], cur_matrix[5])
size = (abs(cur_matrix[0]), abs(cur_matrix[3]))
a, b = cur_matrix[0], cur_matrix[1]
rotation = math.degrees(math.atan2(b, a)) if (a or b) else 0.0
if xobj_name:
blocks.append(ImageBlock(
start=q_start,
end=ops[q_end_idx].end,
origin=origin,
size=size,
xobj_name=xobj_name,
raw_bytes=data[q_start:ops[q_end_idx].end],
rotation=rotation,
alpha=alpha,
cm=cur_matrix,
))
i = q_end_idx + 1
else:
i += 1
return blocks
几个关键点:
6.1 q...Q 是可嵌套的
q 保存图形状态,Q 恢复。它们可以嵌套(q q q ... Q Q Q)。所以分词时要做深度计数:
python
depth = 1
for j in range(i + 1, len(ops)):
if ops[j].operator == 'q':
depth += 1
elif ops[j].operator == 'Q':
depth -= 1
if depth == 0:
q_end_idx = j
break
如果不做深度计数,遇到嵌套 q...Q 会过早闭合,把外层的 Do 漏掉。
6.2 矩阵乘法累积
图片块内可能有多个 cm 操作(变换矩阵),它们是 累积 的:
python
if op.operator == 'cm':
nums = [float(o[1]) for o in op.operands if o[0] == 'num']
if len(nums) >= 6:
new_m = tuple(nums[:6])
cur_matrix = _mat_mul(new_m, cur_matrix)
_mat_mul 实现 3×3 仿射变换矩阵乘法:
python
def _mat_mul(m1: tuple, m2: tuple) -> tuple:
"""3×3 仿射变换矩阵乘法 (a b c d e f 表示)
| a b 0 |
| c d 0 |
| e f 1 |
返回 m1 × m2 的结果 (a, b, c, d, e, f)
"""
a1, b1, c1, d1, e1, f1 = m1
a2, b2, c2, d2, e2, f2 = m2
return (
a1 * a2 + b1 * c2,
a1 * b2 + b1 * d2,
c1 * a2 + d1 * c2,
c1 * b2 + d1 * d2,
e1 * a2 + f1 * c2 + e2,
e1 * b2 + f1 * d2 + f2,
)
PDF 规范里,多个 cm 的累积效果是从右到左应用到坐标的------即 M_total = M_n × ... × M_2 × M_1。所以我们的累积顺序是 cur_matrix = new_m × cur_matrix(新矩阵左乘),这是规范要求。
6.3 只关注含 Do 的块
q...Q 块除了画图,还可能保存颜色设置、裁剪路径等。我们只关心含 Do 操作(调用 XObject)的块,因为只有这种才会绘制图片。其他 q...Q 块对水印检测无意义,跳过。
七、字节流编辑:移除区间
解析器的最后一个核心功能是 apply_removals,它根据待移除区间列表,从原始字节流中删除这些区间:
python
def apply_removals(parsed: ParsedContent) -> bytes:
"""根据 remove_ranges 从原始内容流中移除指定区间,返回新内容流"""
if not parsed.remove_ranges:
return parsed.raw
# 合并重叠区间
ranges = sorted(parsed.remove_ranges)
merged = [list(ranges[0])]
for s, e in ranges[1:]:
if s <= merged[-1][1]:
merged[-1][1] = max(merged[-1][1], e)
else:
merged.append([s, e])
# 从原始字节中删除这些区间
result = bytearray()
prev = 0
for s, e in merged:
result.extend(parsed.raw[prev:s])
prev = e
result.extend(parsed.raw[prev:])
return bytes(result)
两步:
- 合并重叠区间:把多个待删区间合并,避免区间相交导致字节错位。
- 保留区间外的字节:遍历合并后的区间,把每段区间之前的字节拼到结果里,跳过区间本身。
这个实现是 无损 的------只删除指定字节区间,不动其他任何字节。PDF 解析器对内容流的语法很宽容,删除一整个 BT...ET 块后剩下的内容流依然合法。这是矢量 PDF 无损移除水印的根基。
八、ParsedContent:解析结果汇总
最终所有解析结果汇成一个 ParsedContent:
python
@dataclass
class ParsedContent:
"""一页内容流的解析结果"""
text_blocks: list # [TextBlock]
image_blocks: list # [ImageBlock]
raw: bytes # 原始内容流字节
remove_ranges: list = field(default_factory=list) # 待移除的 (start, end) 区间
调用方(PdfVectorProcessor)拿到 ParsedContent 后:
- 遍历
text_blocks找水印候选(跨页重复、平铺网格、旋转大字、品牌签名)。 - 遍历
image_blocks找图片水印候选(小尺寸 Logo、二维码)。 - 决定移除时,把候选的
(start, end)加到remove_ranges,调apply_removals生成新内容流,再用doc.update_stream写回 PDF。
下一篇我们会详细讲这些检测策略和移除流程。
九、性能考量
内容流解析器的性能很关键。一份 100 页的 PDF,每页内容流平均 10KB,总共 1MB 字节流。我们的实现是:
- 分词器是生成器 :不一次性返回所有 token,而是逐个 yield。但
parse_content_stream里tokens = list(_tokenize(data))又把它转成列表了,这是为了能多次遍历。后续优化可以改成生成器 + 索引访问。 - 文字块提取是 O(n):一次遍历所有 operations,找出所有 BT...ET。
- 图片块提取是 O(n):同上,找 q...Q。
- 总复杂度 O(n):n 是 token 数量。100 页 PDF 解析总耗时约 200ms,可接受。
最大的性能瓶颈不在解析,而在 page.search_for(text)------PyMuPDF 的字形布局查询。它要解析字体子集,对每个文字块都调一次会很慢。所以 _refine_text_bboxes 只对 置信度 ≥0.55 的文字候选 调用,跳过低置信度的(反正用户也不会勾选它们)。
十、本篇总结
这一篇我们深入讲了了自己写的 PDF 内容流解析器:
- 为什么自己写:PyMuPDF 不暴露字节偏移,无法做区间擦除。
- 分词器:处理字符串嵌套、字典/十六进制区分、数字/操作符区分。
- 操作分组 :操作数 + 操作符组合成
Operation,记录字节偏移。 - 文字块提取 :从
BT...ET累积字体、字号、矩阵、颜色、透明度、文字内容。 - Tm 矩阵提取旋转角 :
atan2(b, a)算出旋转角,是水印检测的核心特征。 - 灰色与透明度判断 :
abs(r-g) < 0.1判灰色,ca/CA提取透明度。 - 中文宽度估算 :用
unicodedata.east_asian_width区分全角/半角,避免红框框不全。 - 文本解码:UTF-8/GBK/Latin-1 三连试,覆盖 90% 中文水印场景。
- 图片块提取 :
q...Q嵌套深度计数,cm矩阵累积,只关注含Do的块。 - 字节流编辑:区间合并 + 字节拼接,实现无损移除。
下一篇我们继续讲 PdfVectorProcessor,看它怎么用这个解析器做 多策略水印检测 ------跨页重复、同页平铺网格、旋转大字、品牌签名、共享 XObject、OCG 图层、注释水印------以及对应的 多策略移除------内容流擦除、XObject 清空、OCG 禁用、注释删除。
📥 完整源码与可执行程序已上传 CSDN 资源,搜索"清印 ClearMark 智能文档去水印工作台"。本系列所有代码片段均来自该项目实际实现,对应文件为
core/content_stream.py,全文 510 行,无任何第三方依赖。