python 提取PDF文字

使用pdfplumber,不能提取扫描的pdf和插入的图片。

python 复制代码
import pdfplumber

file_path = r'D:\UserData\admindesktop\官方文档\1903_Mesh-Models-Overview_FINAL.pdf'
with pdfplumber.open(file_path) as pdf:
    page = pdf.pages[0]
    print(page.extract_text()) # 所以文字
    print([word["text"] for word in page.extract_words()]) # 提取存在的文字
相关推荐
我的xiaodoujiao17 分钟前
Django 基础知识详细图文教程 12-Django 模型定义与使用 2
后端·python·django
happylifetree8 小时前
Python09:核心语法-数据存储与运算-字面量
python
心之语歌8 小时前
Tkinter 画布基本梳理
运维·服务器·python
Wx-bishekaifayuan9 小时前
django个性化旅游路线推荐平台49005-计算机课程设计、毕业设计
spring boot·后端·python·django·课程设计·express·旅游
小白快快跑哦10 小时前
python-字符串全解(六):正则表达式-量词
python·正则表达式·字符串
禹凕10 小时前
机器学习之Selenium(Machina Learning about Selenium)
爬虫·python·selenium·测试工具·机器学习
ebiobiz12 小时前
基于 GD32 Embedded Builder (GEB) 与 Nimmake 的 MCU 工程搭建指南
c++·python·单片机·嵌入式硬件·mcu
老歌老听老掉牙13 小时前
穿越时空的物流密码:VRP与VRPTW的深度剖析及实战求解
python·物流·vrp
FYKJ_201013 小时前
django电影推荐系统55066-计算机课程设计、毕业设计
vue.js·spring boot·python·mysql·typescript·spark·django
ShyanZh13 小时前
【Python3基础】13-Python 常用设计模式
开发语言·python·设计模式