import pdfplumber
import pandas as pd
def extract_tables_to_excel(pdf_path, excel_path):
# 打开PDF文件
with pdfplumber.open(pdf_path) as pdf:
# 创建一个空的DataFrame列表,用于存储所有表格数据
all_tables = []
# 遍历PDF的每一页
for page in pdf.pages:
# 提取当前页的表格
tables = page.extract_tables()
# 将每页的表格转换为DataFrame,并添加到all_tables列表中
for table in tables:
df = pd.DataFrame(table)
all_tables.append(df)
# 将所有表格数据合并为一个DataFrame
combined_tables = pd.concat(all_tables, ignore_index=True)
# 将合并后的表格数据保存到Excel文件中
combined_tables.to_excel(excel_path, index=False)
# PDF文件路径
pdf_path = '1.pdf'
# Excel文件路径
excel_path = 'output_tables.xlsx'
# 调用函数
extract_tables_to_excel(pdf_path, excel_path)
Python应用—从pdf文件中提取表格,并且保存在excel中
翻车吧奥斯卡2024-07-21 21:45
相关推荐
kobe_OKOK_6 小时前
__getattr__和__getattribute__如何用全栈弄潮儿7 小时前
Python实战第2期: 变量与数据类型长沙三为智能科技7 小时前
家政小程序门店运营模块设计:会员、下单、派单、结算四段链路的落地拆解Cc.Y7 小时前
Java零基础入门:可变字符串与包装类:StringBuilder、StringBuffer 与 通讯录管理系统实战启观川8 小时前
数据结构与算法 -第 3 章 常用算法-动态规划码爸8 小时前
排序算法介绍笔墨登场说说8 小时前
flink bin/start-cluster.sh 帮我做成开机启动鲲穹AI种草8 小时前
PDF 编辑工具怎么选,多款工具实际使用情况整理MoRanzhi12039 小时前
第 3 篇 · 数据清洗与特征工程:从问题判定到特征构建阿松爱学习9 小时前
【Unity开发】Aspose.PDF的使用