1.2 多格式文档处理系统分层架构与处理器路由设计

1.2 多格式文档处理系统分层架构与处理器路由设计

清印 ClearMark 系列第 2 篇 · 共 12 篇

上一篇我们讲了为什么要做这个工具,以及它的整体定位。这篇我们深入到代码层面,聊一聊整个系统的分层架构------为什么是三层、IProcessor 抽象接口怎么设计、format_probe 如何做格式路由、SessionManager 与 TaskOrchestrator 如何分工、统一数据模型如何让 UI 层与处理层解耦。

这套架构不是一开始就定型的,是经过两轮重构之后的产物。我会把每一处设计取舍的"为什么"讲清楚。

源码与可执行程序下载链接将在系列文章全部发布后统一更新。

一、为什么需要分层

最早的原型是单文件脚本,所有逻辑塞在一个 1500 行的 main.py 里:开 PDF、解析内容流、画 UI、移除水印------全部串在一起。后果是:

  • 改一处崩三处。修一个 PDF 解析的 bug,UI 也跟着崩,因为 UI 直接读了 PDF 内部状态。
  • 批量无法实现。单文件脚本是过程式的,没法把"检测一个文件"封装成可并发执行的任务。
  • 测试不可写。任何一个函数都依赖全局 QApplication,单元测试根本起不来。

第一轮重构拆成两层:UI 层 + PDF 处理层。但很快发现不够------图片格式根本进不来,因为图片的处理逻辑和 PDF 完全不一样(一个走内容流,一个走像素),但 UI 层希望用同一套接口调用。

第二轮重构引入 三层 + 策略模式 :UI 层只消费统一数据模型;处理层通过 IProcessor 抽象基类被 UI 调用;处理层内部又拆成检测器、修复引擎、规则库三组组件。这一版稳定下来了,后续增加 Word、PPT 也只是新增一个 IProcessor 子类的事。

二、三层架构全景

复制代码
┌────────────────────────────────────────────────────────────┐
│  UI 层(ui/)                                              │
│  ┌────────────────┬────────────────┬──────────────────┐  │
│  │ MainWindow     │ PreviewWidget   │ ResultPanel      │  │
│  │ 主窗口三栏布局  │ QPainter预览    │ 候选列表+勾选     │  │
│  ├────────────────┼────────────────┼──────────────────┤  │
│  │ SettingsDialog │ Icons 图标系统  │ Styles 主题       │  │
│  ├────────────────┴────────────────┴──────────────────┤  │
│  │ Worker 线程:DetectWorker / RemoveWorker           │  │
│  │                BatchDetectWorker / BatchRemoveWorker│  │
│  └────────────────────────────────────────────────────┘  │
└─────────────────────────┬────────────────────────────────┘
                          │ 通过 SessionManager / TaskOrchestrator 调用
┌─────────────────────────┴────────────────────────────────┐
│  会话与编排层(core/)                                    │
│  ┌─────────────────┬──────────────────┐                 │
│  │ SessionManager  │ TaskOrchestrator  │                 │
│  │ 单文件会话:     │ 批量任务:         │                 │
│  │ open/detect/    │ add_files/        │                 │
│  │ remove/close    │ detect_all/       │                 │
│  │ + 备份逻辑      │ remove_all        │                 │
│  └─────────────────┴──────────────────┘                 │
│  ┌─────────────────┬──────────────────┐                 │
│  │ FormatProbe     │ ContentStream    │                 │
│  │ 格式探测+路由    │ PDF内容流解析器   │                 │
│  └─────────────────┴──────────────────┘                 │
│  ┌──────────────────────────────────────┐                │
│  │ Processor 层(策略模式)              │                │
│  │ ┌──────────────┬──────────────────┐ │                │
│  │ │ PdfVector    │ PdfScanned       │ │                │
│  │ │ 矢量PDF      │ 扫描PDF          │ │                │
│  │ ├──────────────┼──────────────────┤ │                │
│  │ │ ImageInpaint│ (预留Docx/Pptx) │ │                │
│  │ │ 图片         │                 │ │                │
│  │ └──────────────┴──────────────────┘ │                │
│  └──────────────────────────────────────┘                │
│  ┌────────────────┬────────────────┬────────────────┐   │
│  │ Detect/        │ Inpaint/       │ Rules/         │   │
│  │ OCR/重复/模板/ │ cv_engine      │ rule_loader     │   │
│  │ 视觉/图层       │ OpenCV修复引擎  │ 水印规则库      │   │
│  └────────────────┴────────────────┴────────────────┘   │
└─────────────────────────┬────────────────────────────────┘
                          │ 数据流通过统一模型
┌─────────────────────────┴────────────────────────────────┐
│  数据层(core/model.py)                                  │
│  WatermarkCandidate · DetectionResult · TaskItem        │
│  OutputStrategy · UserMark · BatchReport                │
└──────────────────────────────────────────────────────────┘

三层职责清晰:

  • UI 层:只管渲染和交互,不感知文件格式。
  • 会话与编排层:管理"当前打开的文件"和"批量任务队列",调用处理器完成实际工作。
  • 处理器层:每种格式一个具体处理器,内部组合使用检测器、修复引擎、规则库。
  • 数据层:贯穿三层的统一数据结构,是层间通信的"通用货币"。

三、IProcessor:抽象基类的设计

整个架构最关键的一个抽象是 core/processor/base.py 里的 IProcessor:

python 复制代码
class IProcessor(ABC):
    """统一处理器接口

    子类需实现:
        open(path)          --- 打开文件
        detect_auto()      --- 智能检测
        detect_region(marks)--- 框选/涂抹检测
        detect_preset(rule)--- 规则库检测
        remove(candidates, output) --- 移除水印并保存
        render_page(idx, zoom) --- 渲染预览图
    """

    format_type: FormatType = FormatType.UNKNOWN

    def __init__(self):
        self.file_path: Optional[str] = None
        self.result: Optional[DetectionResult] = None

    @abstractmethod
    def open(self, path: str): ...

    @abstractmethod
    def close(self): ...

    @property
    @abstractmethod
    def page_count(self) -> int: ...

    @property
    @abstractmethod
    def is_open(self) -> bool: ...

    @abstractmethod
    def detect_auto(self) -> DetectionResult: ...

    @abstractmethod
    def detect_region(self, marks: List[UserMark]) -> DetectionResult: ...

    @abstractmethod
    def detect_preset(self, rule_name: str) -> DetectionResult: ...

    @abstractmethod
    def remove(self, candidates: List[WatermarkCandidate],
               output_path: str,
               progress_callback: Optional[Callable] = None) -> int: ...

    @abstractmethod
    def render_page(self, page_index: int, zoom: float = 1.0) -> bytes: ...

这个接口有几个细节值得拆解。

3.1 三种检测模式分开

为什么 detect_auto / detect_region / detect_preset 要分开?因为它们的输入完全不同:

  • detect_auto:无外部输入,完全靠启发式和算法自动识别。
  • detect_region:输入是用户画的 UserMark 列表(矩形/画笔/套索)。
  • detect_preset:输入是规则名(如 camscanner),从规则库加载关键词。

把它们合并成一个 detect(mode, **kwargs) 也可以------早期版本就是这么写的------但很快发现两个问题:

  1. UI 层要根据当前模式禁用/启用按钮,合并接口后还得查 mode 字符串。
  2. 不同模式返回的 DetectionResult.summary 文案不同,分开写更清晰。

所以最终还是拆成了三个独立方法。这是 接口职责单一原则 优先于"减少方法数"的典型例子。

3.2 render_page 返回 bytes 而不是 QPixmap

render_page 返回 PNG bytes,而不是 QPixmap。这是有意的------core 层不依赖 Qt,所有依赖项都在 ui/ 层。这样设计的好处是:

  • core 可以独立单元测试(不需要 QApplication)
  • 后续要做 CLI 版本(无 GUI 批处理)很容易
  • 跨进程调用(比如做成 HTTP 服务)也只需要换序列化协议

UI 层拿到 bytes 后自己用 QPixmap.loadFromData() 转 pixmap。

3.3 remove 返回移除数量,不返回输出路径

python 复制代码
def remove(self, candidates, output_path, progress_callback=None) -> int:
    """移除水印并保存
    Returns: 实际移除的元素数
    """

输出路径由调用方(SessionManager.remove)传入,而不是处理器自己决定。这又是另一个解耦:输出策略(位置、命名、冲突处理)是 UI 层的职责,处理器只负责"按指定路径写文件"。

具体策略实现在 core/model.py 的 OutputStrategy:

python 复制代码
class OutputStrategy:
    """输出文件策略

    默认策略:
    - 输出到源文件同目录下的 output/ 子文件夹
    - 文件名: 原名_clean.扩展名
    - 冲突: 自动编号
    - 源目录不可写时,自动回退到 ~/.clearmark/output/
    """
    DEFAULT_OUTPUT_DIR_NAME = 'output'
    DEFAULT_SUFFIX = '_clean'
    CONFLICT_AUTO_NUMBER = 'auto_number'

    def resolve_output_path(self, source_path: str) -> str:
        """计算输出文件路径,自动选择第一个可写的候选目录"""
        dirs = self._candidate_dirs(source_path)
        for d in dirs:
            if ensure_writable_dir(d):
                base = os.path.basename(source_path)
                name, ext = os.path.splitext(base)
                out_name = f'{name}{self.suffix}{ext}'
                out_path = os.path.join(d, out_name)
                # 冲突自动编号
                if os.path.exists(out_path):
                    i = 2
                    while True:
                        cand = os.path.join(d, f'{name}{self.suffix}_{i}{ext}')
                        if not os.path.exists(cand):
                            out_path = cand
                            break
                        i += 1
                self.last_output_dir = d
                self.last_fallback = (d != dirs[0])
                return out_path
        # 全部候选都不可写 - 抛异常由调用方处理
        raise OSError("无可写输出目录")

ensure_writable_dir 是个坑点:你不能只用 os.access(path, os.W_OK) 判断,因为权限位判断在沙箱/ACL/挂载掩码场景下会误报可写。必须真实创建临时文件探测:

python 复制代码
def ensure_writable_dir(path: str) -> bool:
    try:
        os.makedirs(path, exist_ok=True)
        probe = os.path.join(path, f'.cm_write_probe_{os.getpid()}')
        with open(probe, 'wb') as f:
            f.write(b'x')
        os.remove(probe)
        return True
    except OSError:
        return False

这个函数是踩过坑才写出来的。早期版本只查 os.access,结果在 U盘挂载的 FAT32 目录上跑挂了------权限位显示可写,实际写不进去。

3.4 remove_to_output:恢复处理器原始状态

基类还提供了一个模板方法 remove_to_output,所有子类共用:

python 复制代码
def remove_to_output(self, candidates, output_strategy, progress_callback=None) -> str:
    """按输出策略移除并保存

    每次都从原始文件重新操作,确保多次移除互不影响。
    移除完毕后恢复处理器到原始文档状态。
    """
    if not self.file_path:
        raise ValueError("未打开文件")
    out_path = output_strategy.resolve_output_path(self.file_path)
    os.makedirs(os.path.dirname(out_path), exist_ok=True)

    original_path = self.file_path
    try:
        self.remove(candidates, out_path, progress_callback)
    finally:
        # 无论成功或失败,都恢复到原始文档状态
        try:
            self.close()
            self.open(original_path)
        except Exception:
            pass
    return out_path

这里 finally 块的 close() + open(original_path) 是 关键 bug 修复 。原因:remove 内部会修改 PDF 文档对象(删除内容流区间、清空 XObject、禁用 OCG),如果直接拿当前文档对象做下一次 remove,会得到空白页或损坏文件。这个坑我们会在第 12 篇详细讲。

四、格式探测器:路由的设计

core/format_probe.py 是整个路由的入口。它做的事很简单------给一个文件路径,返回 FormatType 枚举:

python 复制代码
class FormatType(IntEnum):
    PDF_VECTOR = 1       # 矢量 PDF
    PDF_SCANNED = 2      # 扫描版 PDF(整页位图)
    PDF_MIXED = 3        # 混合型 PDF
    IMAGE = 4            # 位图图片
    DOCX = 5             # Word 文档(V2.1)
    PPTX = 6             # PowerPoint(V2.1)
    UNKNOWN = 99

这里有一个关键设计:PDF 被细分成三类。为什么?因为同样是 .pdf 后缀,里面可能是完全不同的结构:

  • 矢量 PDF:内容流里有文字指令(BT...ET)和矢量绘图指令,可以直接解析文本。常见于 Word/LaTeX 直接导出的 PDF。
  • 扫描 PDF:每页只有一张大图,整页是位图,内容流里几乎没有文字指令。常见于扫描全能王、打印机扫描导出。
  • 混合 PDF:既有大图,又有文字指令。常见于先扫描再加页眉页脚的文档。

三类 PDF 的水印检测策略完全不同。矢量型可以走内容流解析无损擦除;扫描型必须走像素级 OCR + 图像修复;混合型则需要两者结合。

4.1 探测算法

python 复制代码
def _probe_pdf_type(path: str) -> FormatType:
    """判断 PDF 类型:矢量 / 扫描 / 混合"""
    try:
        import pymupdf as fitz
        doc = fitz.open(path)
        if doc.page_count == 0:
            doc.close()
            return FormatType.PDF_VECTOR

        has_text = False
        has_full_page_image = False
        has_vector = False

        check_pages = min(doc.page_count, 5)
        for i in range(check_pages):
            page = doc[i]
            text = page.get_text("text").strip()
            if text:
                has_text = True

            images = page.get_images(full=True)
            for img in images:
                xref = img[0]
                try:
                    pix = fitz.Pixmap(doc, xref)
                    # 图片面积接近页面面积 → 扫描件
                    if pix.width * pix.height > page.rect.width * page.rect.height * 0.8:
                        has_full_page_image = True
                    pix = None
                except Exception:
                    pass

            drawings = page.get_drawings()
            if drawings:
                has_vector = True

        doc.close()

        if has_full_page_image and not has_text:
            return FormatType.PDF_SCANNED
        elif has_full_page_image and has_text:
            return FormatType.PDF_MIXED
        else:
            return FormatType.PDF_VECTOR
    except Exception:
        return FormatType.PDF_VECTOR

只查前 5 页,是为了性能。曾经遇到过 1000+ 页的扫描合集 PDF,全查的话开文件就要 10 秒。

"图片面积 > 页面面积 × 0.8" 这个阈值是经验值------扫描件的内嵌图通常占页面 95% 以上,但有些 PDF 加了页边距,图片会缩到 85% 左右。0.8 是为了容忍这种情况,再低就会把带图章的矢量 PDF 误判成扫描件。

4.2 路由器

python 复制代码
def get_processor(format_type: FormatType):
    """根据格式类型获取对应的处理器实例"""
    if format_type in (FormatType.PDF_VECTOR, FormatType.PDF_MIXED):
        from .processor.pdf_vector import PdfVectorProcessor
        return PdfVectorProcessor()
    elif format_type == FormatType.PDF_SCANNED:
        from .processor.pdf_scanned import PdfScannedProcessor
        return PdfScannedProcessor()
    elif format_type == FormatType.IMAGE:
        from .processor.image_inpaint import ImageInpaintProcessor
        return ImageInpaintProcessor()
    elif format_type == FormatType.DOCX:
        raise NotImplementedError("Word 文档支持将在 V2.1 实现")
    elif format_type == FormatType.PPTX:
        raise NotImplementedError("PowerPoint 支持将在 V2.1 实现")
    else:
        raise ValueError(f"不支持的格式类型: {format_type}")

注意 import 是延迟到函数内部,而不是模块顶部。这有两个好处:

  1. 启动加速:如果用户只处理图片,PyMuPDF 不会被打包进来(其实没法不打包,但至少不会被 import)。
  2. 可选依赖 :HEIC 支持需要 pillow_heif,OCR 需要 rapidocr_onnxruntime,这些都是可选的,缺失时只在调用对应功能时报错,不影响其他功能。

五、SessionManager:单文件会话

core/session.py 的 SessionManager 管理当前打开的文件。它做的事看起来简单,但有几个坑点:

python 复制代码
class SessionManager:
    def __init__(self):
        self.processor = None
        self.file_path: Optional[str] = None
        self.detection_result: Optional[DetectionResult] = None
        self._backup_dir: str = 'backup'
        self.auto_backup: bool = True
        self.backup_path: Optional[str] = None
        self.backup_note: str = ''

    def open(self, path: str):
        """打开文件,自动选择处理器

        备份是辅助动作:任何备份失败都不会阻断文件打开
        """
        fmt = probe_format(path)
        self.processor = get_processor(fmt)
        self.processor.open(path)
        self.file_path = path
        self.detection_result = None
        self.backup_path = None
        self.backup_note = ''

        if self.auto_backup:
            self._backup(path)

5.1 备份是辅助,不是阻塞

_backup 方法的核心设计是 所有异常都被吞掉:

python 复制代码
def _backup(self, path: str):
    """备份源文件

    位置优先级:
    1. 源文件目录/backup/
    2. ~/.clearmark/backup/<源目录哈希>/

    任何异常都被吞掉并记录到 backup_note,绝不向上抛出
    """
    try:
        name = os.path.basename(path)
        src_dir = os.path.dirname(os.path.abspath(path)) or os.getcwd()
        tag = hashlib.md5(src_dir.encode('utf-8')).hexdigest()[:8]
        candidates = [
            os.path.join(src_dir, self._backup_dir),
            os.path.join(user_data_dir(), 'backup', tag),
        ]

        last_error = None
        for backup_dir in candidates:
            try:
                os.makedirs(backup_dir, exist_ok=True)
                if not os.access(backup_dir, os.W_OK):
                    last_error = PermissionError(f'目录不可写: {backup_dir}')
                    continue
                backup_path = os.path.join(backup_dir, name)
                if not os.path.exists(backup_path):
                    shutil.copy2(path, backup_path)
                self.backup_path = backup_path
                ...
                return
            except OSError as e:
                last_error = e
                continue

        self.backup_note = (...)
    except Exception as e:
        self.backup_note = f'自动备份跳过: {e}'

为什么这么设计?因为 打开文件是用户主动行为,不能因为备份失败让用户看不到自己的文件。哪怕备份目录不可写(U盘只读、网络盘权限),文件还是要能打开、能检测、能另存到别处。

backup_note 会在 UI 上以友好提示形式展示,让用户知道"备份没成功,但你可以继续处理"。

5.2 auto_backup 是公开属性

python 复制代码
self.auto_backup: bool = True

注意是 auto_backup 不是 _auto_backup。早期版本写成 _auto_backup(Python 私有约定),结果 UI 层要开关它时只能 session._auto_backup = False,访问私有属性既不规范也不安全。改成公开属性后,SettingsDialog 可以直接:

python 复制代码
session.auto_backup = self.cb_backup.isChecked()

这是一个看似细节但很关键的工程规范。Python 的下划线前缀只是约定,但跨模块访问 _xxx 会让代码评审报警。

六、TaskOrchestrator:批量任务编排

core/task_orchestrator.py 是批量场景的核心。它的设计目标有四个:

  1. 并发:多文件同时检测/移除,CPU 核心数自动适配。
  2. 可取消:用户点取消按钮立即停止后续任务。
  3. 状态追踪:每个文件有独立状态机,UI 能精确显示进度。
  4. 失败隔离:一个文件失败不影响其他文件。
python 复制代码
class TaskOrchestrator:
    def __init__(self, max_workers: int = 0):
        if max_workers <= 0:
            cpu = os.cpu_count() or 2
            max_workers = max(1, cpu - 1)
        self._max_workers = max_workers
        self._tasks: List[TaskItem] = []
        self._cancelled = False
        self.auto_backup: bool = True

max_workers 默认是 cpu_count() - 1。留一个核给 UI 线程,否则用户操作会卡顿。

6.1 状态机

每个任务项 TaskItem 有自己的状态:

python 复制代码
class TaskStatus(IntEnum):
    PENDING = 1
    DETECTING = 2
    DETECTED = 3
    REMOVING = 4
    COMPLETED = 5
    FAILED = 6
    SKIPPED = 7

状态转移图:

复制代码
PENDING ──detect──> DETECTING ──ok──> DETECTED ──remove──> REMOVING ──ok──> COMPLETED
   │                   │                  │                   │
   │                   └──fail──> FAILED  └──fail──> FAILED   └──skip──> SKIPPED
   │
   └──cancel──> SKIPPED

SKIPPED 状态有两种触发条件:

  1. 用户取消时未开始的任务(PENDING → SKIPPED)。
  2. 移除时发现没有勾选候选(REMOVING → SKIPPED)。

UI 上的列表项按状态显示后缀,比如:

  • "已检测:3 候选"
  • "已完成:2 移除"
  • "失败:权限不足"
  • "跳过:无勾选"

6.2 并发与取消

python 复制代码
def detect_all(self, progress_callback=None, cancel_check=None) -> None:
    self._cancelled = False
    pending = [t for t in self._tasks
               if t.status in (TaskStatus.PENDING, TaskStatus.FAILED)]
    total = len(pending)
    if total == 0:
        return
    done = 0
    with ThreadPoolExecutor(max_workers=self._max_workers) as executor:
        future_map = {executor.submit(self._detect_one, t): t for t in pending}
        for future in as_completed(future_map):
            if (cancel_check and cancel_check()) or self._cancelled:
                self._cancelled = True
                for f in future_map:
                    f.cancel()
                break
            task = future_map[future]
            try:
                future.result()
            except Exception as e:
                task.status = TaskStatus.FAILED
                task.error = f"检测异常: {e}"
            done += 1
            if progress_callback:
                progress_callback(done, total, task.file_name)

这里有几个关键点:

  1. 只重试 PENDING 和 FAILED 状态 。已检测过的(DETECTED)不重复检测,已完成的(COMPLETED)不重复处理。这样用户可以"失败重试"而不影响已成功的文件。
  2. 取消是协作式的 。cancel_check 是 UI 传入的回调,返回 True 表示用户点了取消。线程池里已提交的任务无法强制中断,但可以 future.cancel() 取消未启动的,并跳过已完成的。
  3. progress_callback(done, total, filename) 把进度回报给 UI。UI 在主线程接收,更新进度条与状态栏。

批量移除 remove_all 的结构完全一样,只是调 self._remove_one 而不是 self._detect_one。_remove_one 里有个细节------如果文件还没检测过,会先做一次检测:

python 复制代码
def _remove_one(self, task: TaskItem, output_strategy: OutputStrategy) -> TaskItem:
    start = time.time()
    processor = None
    try:
        task.status = TaskStatus.REMOVING
        processor = get_processor(task.format_type)
        processor.open(task.file_path)
        # 如果还未检测,先检测
        if task.result is None:
            task.result = processor.detect_auto()
            task.candidate_count = len(task.result.candidates)
        selected = [c for c in task.result.candidates if c.selected]
        if not selected:
            task.status = TaskStatus.SKIPPED
            return task
        out_path = processor.remove_to_output(
            selected, output_strategy, None)
        task.output_path = out_path
        task.removed_count = len(selected)
        task.status = TaskStatus.COMPLETED
    except Exception as e:
        task.status = TaskStatus.FAILED
        task.error = f"移除失败: {e}"
    finally:
        if processor is not None:
            try:
                processor.close()
            except Exception:
                pass
        task.duration = time.time() - start
    return task

这个"未检测则先检测"的逻辑允许用户跳过检测步骤直接点"全部移除"------某些场景下用户已经知道要移除所有候选,不想等检测完再操作。

七、统一数据模型:层间通信的"通用货币"

最后讲一下数据层。core/model.py 定义了所有跨层数据结构。其中最核心的是 WatermarkCandidate,上一篇已经详细展示过。这里再补一个细节------置信度自动分级:

python 复制代码
@dataclass
class WatermarkCandidate:
    ...
    confidence: float = 0.0
    confidence_level: ConfidenceLevel = ConfidenceLevel.NONE
    selected: bool = True

    def __post_init__(self):
        """根据置信度自动设置勾选状态"""
        if self.confidence_level == ConfidenceLevel.NONE:
            self.confidence_level = self._calc_level()
        self.selected = self.confidence_level >= ConfidenceLevel.MEDIUM

    def _calc_level(self) -> ConfidenceLevel:
        if self.confidence >= 0.85:
            return ConfidenceLevel.HIGH
        elif self.confidence >= 0.55:
            return ConfidenceLevel.MEDIUM
        else:
            return ConfidenceLevel.LOW

confidence 是 0~1 的浮点数,由检测器根据多种特征加权得出。confidence_level 是离散的等级:

  • HIGH(≥0.85):自动勾选(用户可手动取消)
  • MEDIUM(0.55~0.85):默认勾选
  • LOW(<0.55):默认不勾选,让用户决定

这套分级让 UI 能默认勾选高置信度候选,把不确定的交给用户判断。这对实际体验非常关键------如果所有候选都默认勾选,误删风险大;都不勾选,用户得一个个点。中间值就是"高置信度自动勾选,低置信度待确认"。

八、依赖关系总结

最后梳理一下整个项目的依赖关系:

复制代码
main.py
  └─ ui/main_window.py
       ├─ ui/preview_widget.py
       ├─ ui/result_panel.py
       ├─ ui/settings_dialog.py
       ├─ ui/styles.py
       ├─ ui/icons.py
       └─ ui/worker.py
            ├─ core/session.py ──── core/format_probe.py ──── core/processor/*
            │                                          └─ core/content_stream.py
            │                                          └─ core/rules/rule_loader.py
            └─ core/task_orchestrator.py
                  ├─ core/format_probe.py
                  └─ core/processor/*
                       ├─ core/detect/* (OCR/重复/模板/视觉/图层)
                       ├─ core/inpaint/cv_engine.py
                       └─ core/model.py  ←─── 所有模块都依赖
  • core/model.py 没有任何项目内依赖,是最底层
  • core/format_probe.py 只依赖 model
  • core/content_stream.py 没有项目内依赖,是独立的纯算法模块
  • core/processor/* 依赖 base、content_stream、inpaint、rules、model
  • ui/* 依赖 core/session、core/task_orchestrator、core/model,但不直接依赖 core/processor/*
  • main.py 只依赖 ui/main_window 和 ui/icons

这种"单向无环依赖"让任何一层都可以单独替换:换 UI(比如改成 Web),只需要重写 ui/;换处理引擎(比如改成基于 PaddleOCR),只需要重写 core/processor/*。

九、本篇总结

这一篇我们讲了:

  1. 三层架构:UI 层 / 会话与编排层 / 处理器层,单向依赖,互不耦合。
  2. IProcessor 抽象基类:6 个核心方法,UI 层通过它统一调用,不感知格式差异。
  3. format_probe 路由:扩展名 + 内容探测,把 PDF 细分成矢量/扫描/混合三类。
  4. SessionManager 单文件会话:备份是辅助,失败不影响打开。
  5. TaskOrchestrator 批量编排:状态机 + ThreadPoolExecutor + 协作式取消。
  6. WatermarkCandidate 统一数据模型:所有格式的检测结果都被统一表达,置信度自动分级。
  7. OutputStrategy 输出策略:源目录优先、不可写回退到用户目录、冲突自动编号。

下一篇我们开始进入 PDF 矢量处理器的内部------手写 PDF 内容流解析器。这是整个项目技术难度最高的一块,也是最值得学习的一块。我们会从 PDF 的内容流语法讲起,看怎么把一段字节流解析成结构化的文字块和图片块,又怎么从中提取旋转角、透明度、颜色等水印特征。

📥 完整源码与可执行程序已上传 CSDN 资源,搜索"清印 ClearMark 智能文档去水印工作台"。本系列所有代码片段均来自该项目实际实现,对应文件路径已在文中标注。