1.2 多格式文档处理系统分层架构与处理器路由设计
清印 ClearMark 系列第 2 篇 · 共 12 篇
上一篇我们讲了为什么要做这个工具,以及它的整体定位。这篇我们深入到代码层面,聊一聊整个系统的分层架构------为什么是三层、
IProcessor抽象接口怎么设计、format_probe如何做格式路由、SessionManager与TaskOrchestrator如何分工、统一数据模型如何让 UI 层与处理层解耦。这套架构不是一开始就定型的,是经过两轮重构之后的产物。我会把每一处设计取舍的"为什么"讲清楚。
源码与可执行程序下载链接将在系列文章全部发布后统一更新。
一、为什么需要分层
最早的原型是单文件脚本,所有逻辑塞在一个 1500 行的 main.py 里:开 PDF、解析内容流、画 UI、移除水印------全部串在一起。后果是:
- 改一处崩三处。修一个 PDF 解析的 bug,UI 也跟着崩,因为 UI 直接读了 PDF 内部状态。
- 批量无法实现。单文件脚本是过程式的,没法把"检测一个文件"封装成可并发执行的任务。
- 测试不可写。任何一个函数都依赖全局 QApplication,单元测试根本起不来。
第一轮重构拆成两层:UI 层 + PDF 处理层。但很快发现不够------图片格式根本进不来,因为图片的处理逻辑和 PDF 完全不一样(一个走内容流,一个走像素),但 UI 层希望用同一套接口调用。
第二轮重构引入 三层 + 策略模式 :UI 层只消费统一数据模型;处理层通过 IProcessor 抽象基类被 UI 调用;处理层内部又拆成检测器、修复引擎、规则库三组组件。这一版稳定下来了,后续增加 Word、PPT 也只是新增一个 IProcessor 子类的事。
二、三层架构全景
┌────────────────────────────────────────────────────────────┐
│ UI 层(ui/) │
│ ┌────────────────┬────────────────┬──────────────────┐ │
│ │ MainWindow │ PreviewWidget │ ResultPanel │ │
│ │ 主窗口三栏布局 │ QPainter预览 │ 候选列表+勾选 │ │
│ ├────────────────┼────────────────┼──────────────────┤ │
│ │ SettingsDialog │ Icons 图标系统 │ Styles 主题 │ │
│ ├────────────────┴────────────────┴──────────────────┤ │
│ │ Worker 线程:DetectWorker / RemoveWorker │ │
│ │ BatchDetectWorker / BatchRemoveWorker│ │
│ └────────────────────────────────────────────────────┘ │
└─────────────────────────┬────────────────────────────────┘
│ 通过 SessionManager / TaskOrchestrator 调用
┌─────────────────────────┴────────────────────────────────┐
│ 会话与编排层(core/) │
│ ┌─────────────────┬──────────────────┐ │
│ │ SessionManager │ TaskOrchestrator │ │
│ │ 单文件会话: │ 批量任务: │ │
│ │ open/detect/ │ add_files/ │ │
│ │ remove/close │ detect_all/ │ │
│ │ + 备份逻辑 │ remove_all │ │
│ └─────────────────┴──────────────────┘ │
│ ┌─────────────────┬──────────────────┐ │
│ │ FormatProbe │ ContentStream │ │
│ │ 格式探测+路由 │ PDF内容流解析器 │ │
│ └─────────────────┴──────────────────┘ │
│ ┌──────────────────────────────────────┐ │
│ │ Processor 层(策略模式) │ │
│ │ ┌──────────────┬──────────────────┐ │ │
│ │ │ PdfVector │ PdfScanned │ │ │
│ │ │ 矢量PDF │ 扫描PDF │ │ │
│ │ ├──────────────┼──────────────────┤ │ │
│ │ │ ImageInpaint│ (预留Docx/Pptx) │ │ │
│ │ │ 图片 │ │ │ │
│ │ └──────────────┴──────────────────┘ │ │
│ └──────────────────────────────────────┘ │
│ ┌────────────────┬────────────────┬────────────────┐ │
│ │ Detect/ │ Inpaint/ │ Rules/ │ │
│ │ OCR/重复/模板/ │ cv_engine │ rule_loader │ │
│ │ 视觉/图层 │ OpenCV修复引擎 │ 水印规则库 │ │
│ └────────────────┴────────────────┴────────────────┘ │
└─────────────────────────┬────────────────────────────────┘
│ 数据流通过统一模型
┌─────────────────────────┴────────────────────────────────┐
│ 数据层(core/model.py) │
│ WatermarkCandidate · DetectionResult · TaskItem │
│ OutputStrategy · UserMark · BatchReport │
└──────────────────────────────────────────────────────────┘
三层职责清晰:
- UI 层:只管渲染和交互,不感知文件格式。
- 会话与编排层:管理"当前打开的文件"和"批量任务队列",调用处理器完成实际工作。
- 处理器层:每种格式一个具体处理器,内部组合使用检测器、修复引擎、规则库。
- 数据层:贯穿三层的统一数据结构,是层间通信的"通用货币"。
三、IProcessor:抽象基类的设计
整个架构最关键的一个抽象是 core/processor/base.py 里的 IProcessor:
python
class IProcessor(ABC):
"""统一处理器接口
子类需实现:
open(path) --- 打开文件
detect_auto() --- 智能检测
detect_region(marks)--- 框选/涂抹检测
detect_preset(rule)--- 规则库检测
remove(candidates, output) --- 移除水印并保存
render_page(idx, zoom) --- 渲染预览图
"""
format_type: FormatType = FormatType.UNKNOWN
def __init__(self):
self.file_path: Optional[str] = None
self.result: Optional[DetectionResult] = None
@abstractmethod
def open(self, path: str): ...
@abstractmethod
def close(self): ...
@property
@abstractmethod
def page_count(self) -> int: ...
@property
@abstractmethod
def is_open(self) -> bool: ...
@abstractmethod
def detect_auto(self) -> DetectionResult: ...
@abstractmethod
def detect_region(self, marks: List[UserMark]) -> DetectionResult: ...
@abstractmethod
def detect_preset(self, rule_name: str) -> DetectionResult: ...
@abstractmethod
def remove(self, candidates: List[WatermarkCandidate],
output_path: str,
progress_callback: Optional[Callable] = None) -> int: ...
@abstractmethod
def render_page(self, page_index: int, zoom: float = 1.0) -> bytes: ...
这个接口有几个细节值得拆解。
3.1 三种检测模式分开
为什么 detect_auto / detect_region / detect_preset 要分开?因为它们的输入完全不同:
detect_auto:无外部输入,完全靠启发式和算法自动识别。detect_region:输入是用户画的UserMark列表(矩形/画笔/套索)。detect_preset:输入是规则名(如camscanner),从规则库加载关键词。
把它们合并成一个 detect(mode, **kwargs) 也可以------早期版本就是这么写的------但很快发现两个问题:
- UI 层要根据当前模式禁用/启用按钮,合并接口后还得查
mode字符串。 - 不同模式返回的
DetectionResult.summary文案不同,分开写更清晰。
所以最终还是拆成了三个独立方法。这是 接口职责单一原则 优先于"减少方法数"的典型例子。
3.2 render_page 返回 bytes 而不是 QPixmap
render_page 返回 PNG bytes,而不是 QPixmap。这是有意的------core 层不依赖 Qt,所有依赖项都在 ui/ 层。这样设计的好处是:
core可以独立单元测试(不需要 QApplication)- 后续要做 CLI 版本(无 GUI 批处理)很容易
- 跨进程调用(比如做成 HTTP 服务)也只需要换序列化协议
UI 层拿到 bytes 后自己用 QPixmap.loadFromData() 转 pixmap。
3.3 remove 返回移除数量,不返回输出路径
python
def remove(self, candidates, output_path, progress_callback=None) -> int:
"""移除水印并保存
Returns: 实际移除的元素数
"""
输出路径由调用方(SessionManager.remove)传入,而不是处理器自己决定。这又是另一个解耦:输出策略(位置、命名、冲突处理)是 UI 层的职责,处理器只负责"按指定路径写文件"。
具体策略实现在 core/model.py 的 OutputStrategy:
python
class OutputStrategy:
"""输出文件策略
默认策略:
- 输出到源文件同目录下的 output/ 子文件夹
- 文件名: 原名_clean.扩展名
- 冲突: 自动编号
- 源目录不可写时,自动回退到 ~/.clearmark/output/
"""
DEFAULT_OUTPUT_DIR_NAME = 'output'
DEFAULT_SUFFIX = '_clean'
CONFLICT_AUTO_NUMBER = 'auto_number'
def resolve_output_path(self, source_path: str) -> str:
"""计算输出文件路径,自动选择第一个可写的候选目录"""
dirs = self._candidate_dirs(source_path)
for d in dirs:
if ensure_writable_dir(d):
base = os.path.basename(source_path)
name, ext = os.path.splitext(base)
out_name = f'{name}{self.suffix}{ext}'
out_path = os.path.join(d, out_name)
# 冲突自动编号
if os.path.exists(out_path):
i = 2
while True:
cand = os.path.join(d, f'{name}{self.suffix}_{i}{ext}')
if not os.path.exists(cand):
out_path = cand
break
i += 1
self.last_output_dir = d
self.last_fallback = (d != dirs[0])
return out_path
# 全部候选都不可写 - 抛异常由调用方处理
raise OSError("无可写输出目录")
ensure_writable_dir 是个坑点:你不能只用 os.access(path, os.W_OK) 判断,因为权限位判断在沙箱/ACL/挂载掩码场景下会误报可写。必须真实创建临时文件探测:
python
def ensure_writable_dir(path: str) -> bool:
try:
os.makedirs(path, exist_ok=True)
probe = os.path.join(path, f'.cm_write_probe_{os.getpid()}')
with open(probe, 'wb') as f:
f.write(b'x')
os.remove(probe)
return True
except OSError:
return False
这个函数是踩过坑才写出来的。早期版本只查 os.access,结果在 U盘挂载的 FAT32 目录上跑挂了------权限位显示可写,实际写不进去。
3.4 remove_to_output:恢复处理器原始状态
基类还提供了一个模板方法 remove_to_output,所有子类共用:
python
def remove_to_output(self, candidates, output_strategy, progress_callback=None) -> str:
"""按输出策略移除并保存
每次都从原始文件重新操作,确保多次移除互不影响。
移除完毕后恢复处理器到原始文档状态。
"""
if not self.file_path:
raise ValueError("未打开文件")
out_path = output_strategy.resolve_output_path(self.file_path)
os.makedirs(os.path.dirname(out_path), exist_ok=True)
original_path = self.file_path
try:
self.remove(candidates, out_path, progress_callback)
finally:
# 无论成功或失败,都恢复到原始文档状态
try:
self.close()
self.open(original_path)
except Exception:
pass
return out_path
这里 finally 块的 close() + open(original_path) 是 关键 bug 修复 。原因:remove 内部会修改 PDF 文档对象(删除内容流区间、清空 XObject、禁用 OCG),如果直接拿当前文档对象做下一次 remove,会得到空白页或损坏文件。这个坑我们会在第 12 篇详细讲。
四、格式探测器:路由的设计
core/format_probe.py 是整个路由的入口。它做的事很简单------给一个文件路径,返回 FormatType 枚举:
python
class FormatType(IntEnum):
PDF_VECTOR = 1 # 矢量 PDF
PDF_SCANNED = 2 # 扫描版 PDF(整页位图)
PDF_MIXED = 3 # 混合型 PDF
IMAGE = 4 # 位图图片
DOCX = 5 # Word 文档(V2.1)
PPTX = 6 # PowerPoint(V2.1)
UNKNOWN = 99
这里有一个关键设计:PDF 被细分成三类。为什么?因为同样是 .pdf 后缀,里面可能是完全不同的结构:
- 矢量 PDF:内容流里有文字指令(BT...ET)和矢量绘图指令,可以直接解析文本。常见于 Word/LaTeX 直接导出的 PDF。
- 扫描 PDF:每页只有一张大图,整页是位图,内容流里几乎没有文字指令。常见于扫描全能王、打印机扫描导出。
- 混合 PDF:既有大图,又有文字指令。常见于先扫描再加页眉页脚的文档。
三类 PDF 的水印检测策略完全不同。矢量型可以走内容流解析无损擦除;扫描型必须走像素级 OCR + 图像修复;混合型则需要两者结合。
4.1 探测算法
python
def _probe_pdf_type(path: str) -> FormatType:
"""判断 PDF 类型:矢量 / 扫描 / 混合"""
try:
import pymupdf as fitz
doc = fitz.open(path)
if doc.page_count == 0:
doc.close()
return FormatType.PDF_VECTOR
has_text = False
has_full_page_image = False
has_vector = False
check_pages = min(doc.page_count, 5)
for i in range(check_pages):
page = doc[i]
text = page.get_text("text").strip()
if text:
has_text = True
images = page.get_images(full=True)
for img in images:
xref = img[0]
try:
pix = fitz.Pixmap(doc, xref)
# 图片面积接近页面面积 → 扫描件
if pix.width * pix.height > page.rect.width * page.rect.height * 0.8:
has_full_page_image = True
pix = None
except Exception:
pass
drawings = page.get_drawings()
if drawings:
has_vector = True
doc.close()
if has_full_page_image and not has_text:
return FormatType.PDF_SCANNED
elif has_full_page_image and has_text:
return FormatType.PDF_MIXED
else:
return FormatType.PDF_VECTOR
except Exception:
return FormatType.PDF_VECTOR
只查前 5 页,是为了性能。曾经遇到过 1000+ 页的扫描合集 PDF,全查的话开文件就要 10 秒。
"图片面积 > 页面面积 × 0.8" 这个阈值是经验值------扫描件的内嵌图通常占页面 95% 以上,但有些 PDF 加了页边距,图片会缩到 85% 左右。0.8 是为了容忍这种情况,再低就会把带图章的矢量 PDF 误判成扫描件。
4.2 路由器
python
def get_processor(format_type: FormatType):
"""根据格式类型获取对应的处理器实例"""
if format_type in (FormatType.PDF_VECTOR, FormatType.PDF_MIXED):
from .processor.pdf_vector import PdfVectorProcessor
return PdfVectorProcessor()
elif format_type == FormatType.PDF_SCANNED:
from .processor.pdf_scanned import PdfScannedProcessor
return PdfScannedProcessor()
elif format_type == FormatType.IMAGE:
from .processor.image_inpaint import ImageInpaintProcessor
return ImageInpaintProcessor()
elif format_type == FormatType.DOCX:
raise NotImplementedError("Word 文档支持将在 V2.1 实现")
elif format_type == FormatType.PPTX:
raise NotImplementedError("PowerPoint 支持将在 V2.1 实现")
else:
raise ValueError(f"不支持的格式类型: {format_type}")
注意 import 是延迟到函数内部,而不是模块顶部。这有两个好处:
- 启动加速:如果用户只处理图片,PyMuPDF 不会被打包进来(其实没法不打包,但至少不会被 import)。
- 可选依赖 :HEIC 支持需要
pillow_heif,OCR 需要rapidocr_onnxruntime,这些都是可选的,缺失时只在调用对应功能时报错,不影响其他功能。
五、SessionManager:单文件会话
core/session.py 的 SessionManager 管理当前打开的文件。它做的事看起来简单,但有几个坑点:
python
class SessionManager:
def __init__(self):
self.processor = None
self.file_path: Optional[str] = None
self.detection_result: Optional[DetectionResult] = None
self._backup_dir: str = 'backup'
self.auto_backup: bool = True
self.backup_path: Optional[str] = None
self.backup_note: str = ''
def open(self, path: str):
"""打开文件,自动选择处理器
备份是辅助动作:任何备份失败都不会阻断文件打开
"""
fmt = probe_format(path)
self.processor = get_processor(fmt)
self.processor.open(path)
self.file_path = path
self.detection_result = None
self.backup_path = None
self.backup_note = ''
if self.auto_backup:
self._backup(path)
5.1 备份是辅助,不是阻塞
_backup 方法的核心设计是 所有异常都被吞掉:
python
def _backup(self, path: str):
"""备份源文件
位置优先级:
1. 源文件目录/backup/
2. ~/.clearmark/backup/<源目录哈希>/
任何异常都被吞掉并记录到 backup_note,绝不向上抛出
"""
try:
name = os.path.basename(path)
src_dir = os.path.dirname(os.path.abspath(path)) or os.getcwd()
tag = hashlib.md5(src_dir.encode('utf-8')).hexdigest()[:8]
candidates = [
os.path.join(src_dir, self._backup_dir),
os.path.join(user_data_dir(), 'backup', tag),
]
last_error = None
for backup_dir in candidates:
try:
os.makedirs(backup_dir, exist_ok=True)
if not os.access(backup_dir, os.W_OK):
last_error = PermissionError(f'目录不可写: {backup_dir}')
continue
backup_path = os.path.join(backup_dir, name)
if not os.path.exists(backup_path):
shutil.copy2(path, backup_path)
self.backup_path = backup_path
...
return
except OSError as e:
last_error = e
continue
self.backup_note = (...)
except Exception as e:
self.backup_note = f'自动备份跳过: {e}'
为什么这么设计?因为 打开文件是用户主动行为,不能因为备份失败让用户看不到自己的文件。哪怕备份目录不可写(U盘只读、网络盘权限),文件还是要能打开、能检测、能另存到别处。
backup_note 会在 UI 上以友好提示形式展示,让用户知道"备份没成功,但你可以继续处理"。
5.2 auto_backup 是公开属性
python
self.auto_backup: bool = True
注意是 auto_backup 不是 _auto_backup。早期版本写成 _auto_backup(Python 私有约定),结果 UI 层要开关它时只能 session._auto_backup = False,访问私有属性既不规范也不安全。改成公开属性后,SettingsDialog 可以直接:
python
session.auto_backup = self.cb_backup.isChecked()
这是一个看似细节但很关键的工程规范。Python 的下划线前缀只是约定,但跨模块访问 _xxx 会让代码评审报警。
六、TaskOrchestrator:批量任务编排
core/task_orchestrator.py 是批量场景的核心。它的设计目标有四个:
- 并发:多文件同时检测/移除,CPU 核心数自动适配。
- 可取消:用户点取消按钮立即停止后续任务。
- 状态追踪:每个文件有独立状态机,UI 能精确显示进度。
- 失败隔离:一个文件失败不影响其他文件。
python
class TaskOrchestrator:
def __init__(self, max_workers: int = 0):
if max_workers <= 0:
cpu = os.cpu_count() or 2
max_workers = max(1, cpu - 1)
self._max_workers = max_workers
self._tasks: List[TaskItem] = []
self._cancelled = False
self.auto_backup: bool = True
max_workers 默认是 cpu_count() - 1。留一个核给 UI 线程,否则用户操作会卡顿。
6.1 状态机
每个任务项 TaskItem 有自己的状态:
python
class TaskStatus(IntEnum):
PENDING = 1
DETECTING = 2
DETECTED = 3
REMOVING = 4
COMPLETED = 5
FAILED = 6
SKIPPED = 7
状态转移图:
PENDING ──detect──> DETECTING ──ok──> DETECTED ──remove──> REMOVING ──ok──> COMPLETED
│ │ │ │
│ └──fail──> FAILED └──fail──> FAILED └──skip──> SKIPPED
│
└──cancel──> SKIPPED
SKIPPED 状态有两种触发条件:
- 用户取消时未开始的任务(
PENDING → SKIPPED)。 - 移除时发现没有勾选候选(
REMOVING → SKIPPED)。
UI 上的列表项按状态显示后缀,比如:
- "已检测:3 候选"
- "已完成:2 移除"
- "失败:权限不足"
- "跳过:无勾选"
6.2 并发与取消
python
def detect_all(self, progress_callback=None, cancel_check=None) -> None:
self._cancelled = False
pending = [t for t in self._tasks
if t.status in (TaskStatus.PENDING, TaskStatus.FAILED)]
total = len(pending)
if total == 0:
return
done = 0
with ThreadPoolExecutor(max_workers=self._max_workers) as executor:
future_map = {executor.submit(self._detect_one, t): t for t in pending}
for future in as_completed(future_map):
if (cancel_check and cancel_check()) or self._cancelled:
self._cancelled = True
for f in future_map:
f.cancel()
break
task = future_map[future]
try:
future.result()
except Exception as e:
task.status = TaskStatus.FAILED
task.error = f"检测异常: {e}"
done += 1
if progress_callback:
progress_callback(done, total, task.file_name)
这里有几个关键点:
- 只重试
PENDING和FAILED状态 。已检测过的(DETECTED)不重复检测,已完成的(COMPLETED)不重复处理。这样用户可以"失败重试"而不影响已成功的文件。 - 取消是协作式的 。
cancel_check是 UI 传入的回调,返回True表示用户点了取消。线程池里已提交的任务无法强制中断,但可以future.cancel()取消未启动的,并跳过已完成的。 progress_callback(done, total, filename)把进度回报给 UI。UI 在主线程接收,更新进度条与状态栏。
批量移除 remove_all 的结构完全一样,只是调 self._remove_one 而不是 self._detect_one。_remove_one 里有个细节------如果文件还没检测过,会先做一次检测:
python
def _remove_one(self, task: TaskItem, output_strategy: OutputStrategy) -> TaskItem:
start = time.time()
processor = None
try:
task.status = TaskStatus.REMOVING
processor = get_processor(task.format_type)
processor.open(task.file_path)
# 如果还未检测,先检测
if task.result is None:
task.result = processor.detect_auto()
task.candidate_count = len(task.result.candidates)
selected = [c for c in task.result.candidates if c.selected]
if not selected:
task.status = TaskStatus.SKIPPED
return task
out_path = processor.remove_to_output(
selected, output_strategy, None)
task.output_path = out_path
task.removed_count = len(selected)
task.status = TaskStatus.COMPLETED
except Exception as e:
task.status = TaskStatus.FAILED
task.error = f"移除失败: {e}"
finally:
if processor is not None:
try:
processor.close()
except Exception:
pass
task.duration = time.time() - start
return task
这个"未检测则先检测"的逻辑允许用户跳过检测步骤直接点"全部移除"------某些场景下用户已经知道要移除所有候选,不想等检测完再操作。
七、统一数据模型:层间通信的"通用货币"
最后讲一下数据层。core/model.py 定义了所有跨层数据结构。其中最核心的是 WatermarkCandidate,上一篇已经详细展示过。这里再补一个细节------置信度自动分级:
python
@dataclass
class WatermarkCandidate:
...
confidence: float = 0.0
confidence_level: ConfidenceLevel = ConfidenceLevel.NONE
selected: bool = True
def __post_init__(self):
"""根据置信度自动设置勾选状态"""
if self.confidence_level == ConfidenceLevel.NONE:
self.confidence_level = self._calc_level()
self.selected = self.confidence_level >= ConfidenceLevel.MEDIUM
def _calc_level(self) -> ConfidenceLevel:
if self.confidence >= 0.85:
return ConfidenceLevel.HIGH
elif self.confidence >= 0.55:
return ConfidenceLevel.MEDIUM
else:
return ConfidenceLevel.LOW
confidence 是 0~1 的浮点数,由检测器根据多种特征加权得出。confidence_level 是离散的等级:
HIGH(≥0.85):自动勾选(用户可手动取消)MEDIUM(0.55~0.85):默认勾选LOW(<0.55):默认不勾选,让用户决定
这套分级让 UI 能默认勾选高置信度候选,把不确定的交给用户判断。这对实际体验非常关键------如果所有候选都默认勾选,误删风险大;都不勾选,用户得一个个点。中间值就是"高置信度自动勾选,低置信度待确认"。
八、依赖关系总结
最后梳理一下整个项目的依赖关系:
main.py
└─ ui/main_window.py
├─ ui/preview_widget.py
├─ ui/result_panel.py
├─ ui/settings_dialog.py
├─ ui/styles.py
├─ ui/icons.py
└─ ui/worker.py
├─ core/session.py ──── core/format_probe.py ──── core/processor/*
│ └─ core/content_stream.py
│ └─ core/rules/rule_loader.py
└─ core/task_orchestrator.py
├─ core/format_probe.py
└─ core/processor/*
├─ core/detect/* (OCR/重复/模板/视觉/图层)
├─ core/inpaint/cv_engine.py
└─ core/model.py ←─── 所有模块都依赖
core/model.py没有任何项目内依赖,是最底层core/format_probe.py只依赖modelcore/content_stream.py没有项目内依赖,是独立的纯算法模块core/processor/*依赖base、content_stream、inpaint、rules、modelui/*依赖core/session、core/task_orchestrator、core/model,但不直接依赖core/processor/*main.py只依赖ui/main_window和ui/icons
这种"单向无环依赖"让任何一层都可以单独替换:换 UI(比如改成 Web),只需要重写 ui/;换处理引擎(比如改成基于 PaddleOCR),只需要重写 core/processor/*。
九、本篇总结
这一篇我们讲了:
- 三层架构:UI 层 / 会话与编排层 / 处理器层,单向依赖,互不耦合。
IProcessor抽象基类:6 个核心方法,UI 层通过它统一调用,不感知格式差异。format_probe路由:扩展名 + 内容探测,把 PDF 细分成矢量/扫描/混合三类。SessionManager单文件会话:备份是辅助,失败不影响打开。TaskOrchestrator批量编排:状态机 + ThreadPoolExecutor + 协作式取消。WatermarkCandidate统一数据模型:所有格式的检测结果都被统一表达,置信度自动分级。OutputStrategy输出策略:源目录优先、不可写回退到用户目录、冲突自动编号。
下一篇我们开始进入 PDF 矢量处理器的内部------手写 PDF 内容流解析器。这是整个项目技术难度最高的一块,也是最值得学习的一块。我们会从 PDF 的内容流语法讲起,看怎么把一段字节流解析成结构化的文字块和图片块,又怎么从中提取旋转角、透明度、颜色等水印特征。
📥 完整源码与可执行程序已上传 CSDN 资源,搜索"清印 ClearMark 智能文档去水印工作台"。本系列所有代码片段均来自该项目实际实现,对应文件路径已在文中标注。