Linux笔记--基于OCRmyPDF将扫描件PDF转换为可搜索的PDF

1--官方仓库

https://github.com/ocrmypdf/OCRmyPDF

2--基本步骤

bash 复制代码
# 安装ocrmypdf库
sudo apt install ocrmypdf

# 安装简体中文库
sudo apt-get install tesseract-ocr-chi-sim

# 转换
# -l 表示使用的语言
# --force-ocr 防止出现以下错误:ERROR - PriorOcrFoundError: page already has text! - aborting (use --force-ocr to force OCR)
# input.pdf 表示待转换的pdf
# output.pdf 表示转换后保存的pdf
ocrmypdf -l chi_sim input.pdf output.pdf --force-ocr

3--常见错误

Error1:

ERROR - PriorOcrFoundError: page already has text! - aborting (use --force-ocr to force OCR)

Solution:

添加--force-ocr

ocrmypdf -l chi_sim input.pdf output3.pdf --force-ocr

相关推荐
盐焗鹌鹑蛋4 分钟前
【Linux】基本指令
linux
海阔天空任鸟飞~7 小时前
Linux 权限 777
linux·运维·服务器
Kina_C10 小时前
Linux iptables 防火墙原理与实操——从四表五链到 NAT 配置
linux·运维·服务器·iptables
wdfk_prog10 小时前
嵌入式面试真题第 15 题:不可恢复异常后的通用崩溃快照、调用栈保存与离线分析架构
linux·开发语言·面试·架构
雷工笔记14 小时前
什么是Ubuntu?
linux·运维·ubuntu
jsons115 小时前
autofs挂载
linux·服务器·网络
liwulin050615 小时前
【ollama】自定义结构化输出
linux·前端·数据库·ollama
gs8014016 小时前
【实战】记一次 Linux 服务器入侵排查:grok-agent 后门木马深度剖析与彻底清理
linux·运维·服务器
时空无限16 小时前
vllm 大模型启动缓存相关环境变量 export
linux·缓存·vllm
Echo_cy_17 小时前
基于ZYNQ-7000的Ethernet驱动配置
linux·驱动开发·嵌入式硬件·fpga开发·ethernet·zynq