[解决方案] Word转PDF

背景:

之前做过一些pdf导出, 客户提了一个特别急的需求, 要求根据一个模版跟一个csv的数据源, 批量生成PDF, 因为之前用过FOP, 知道调整样式需要特别长的时间, 这个需求又特别急, 所以寻找了一个其他的方案。

优点:

生成快捷,代码简单, 样式依赖模版,所见即所得

缺点:

模版难以调整

思路:

既然已经放弃FOP,那么就直接从模版生成新的word文档, 并且将word文档直接导出

第一版思路:

复制代码
<dependency>
      <groupId>org.docx4j</groupId>
      <artifactId>docx4j-core</artifactId>
      <version>8.3.9</version>
    </dependency> <!-- :contentReference[oaicite:0]{index=0} -->

    <!-- 内置 MOXy JAXB 实现 -->
    <dependency>
      <groupId>org.docx4j</groupId>
      <artifactId>docx4j-JAXB-MOXy</artifactId>
      <version>8.3.9</version>
    </dependency> <!-- :contentReference[oaicite:1]{index=1} -->

<!--    &lt;!&ndash; FO 导出,用于生成 XSL-FO &ndash;&gt;-->
    <dependency>
      <groupId>org.docx4j</groupId>
      <artifactId>docx4j-export-fo</artifactId>
      <version>8.3.9</version>
    </dependency> <!-- :contentReference[oaicite:2]{index=2} -->


public static void main(String[] args) throws Exception {

        // 1. 加载模板
        InputStream tpl = Word2PDF.class
                .getResourceAsStream("/template.docx");
        if (tpl == null) {
            throw new RuntimeException("未找到模板 template.docx");
        }

        //这部分非必须, 是为了多次导出,不重复读模版
        byte[] template = tpl.readAllBytes();

        // 2. 准备多条替换数据
        List<Map<String,String>> dataList = new ArrayList<>();
        Map<String,String> maps = new HashMap<>();

        maps.put("firstName","Alice");
        maps.put("context","测试测试测试测试测试测试测试测试测试测试测试测试测试测试测试测试测");
        maps.put("lastName","Wang");
        maps.put("date","2025-05-20");
        dataList.add(maps);
        maps = new HashMap<>();
        maps.put("firstName","Bob");
        maps.put("lastName","Li");
        maps.put("date","2025-05-21");

        dataList.add(maps);
        maps = new HashMap<>();
        maps.put("firstName","Carol");
        maps.put("lastName","Zhang");
        maps.put("date","2025-05-22");
        dataList.add(maps);


        // 3. 循环生成
        for (Map<String,String> row : dataList) {
            // 3.1 重新加载模板
            WordprocessingMLPackage pkg;
            try (InputStream tplStream = new ByteArrayInputStream(template)) {
                pkg = WordprocessingMLPackage.load(tplStream);
            }

            // 3.2 执行替换 (${key})
            MainDocumentPart mdp = pkg.getMainDocumentPart();
            mdp.variableReplace(row);
            // 替换 ${firstName}、${lastName}、${date} :contentReference[oaicite:2]{index=2}

            // 3.3 保存为 DOCX
            String name = row.get("firstName");
            String docxPath = "/Users/Documents/" + name + ".docx";
            pkg.save(new File(docxPath));

            try(OutputStream os = new FileOutputStream("/Users/Documents/" + name + ".pdf"))  {
                Docx4J.toPDF(pkg, os);
            }
        }

    }

这种方式全部依赖docx4j的jar包,进行导出。

缺点, 当模版有复杂模型,比如侧边栏时这种方式是无法导出的, 在网上找到的解决方案也是无效的。可能是因为JDK版本的升级。

版本2:

上面代码的逻辑一样,额外使用了documents4j的jar

复制代码
<dependency>
      <groupId>org.docx4j</groupId>
      <artifactId>docx4j-core</artifactId>
      <version>8.3.9</version>
    </dependency> <!-- :contentReference[oaicite:0]{index=0} -->

    <!-- 内置 MOXy JAXB 实现 -->
    <dependency>
      <groupId>org.docx4j</groupId>
      <artifactId>docx4j-JAXB-MOXy</artifactId>
      <version>8.3.9</version>
    </dependency> <!-- :contentReference[oaicite:1]{index=1} -->

<!--    &lt;!&ndash; FO 导出,用于生成 XSL-FO &ndash;&gt;-->
    <dependency>
      <groupId>org.docx4j</groupId>
      <artifactId>docx4j-export-fo</artifactId>
      <version>8.3.9</version>
    </dependency> <!-- :contentReference[oaicite:2]{index=2} -->

//转化为PDF的代码使用

//声明转换器,可重用
IConverter converter = LocalConverter.builder()
                        .baseFolder(new File(targetPath))
                        .workerPool(5,15,30, TimeUnit.SECONDS)
                        .processTimeout (60, TimeUnit.SECONDS)
                        .build();
//声明 转换, 最后一步有schedule excute 两种写法, excute是直接生成,结果是boolean,是单条生成的,这种是为了批量运行
Future<Boolean> future = converter
                        .convert(word)
                        .as(DocumentType.MS_WORD)
                        .to(new File(wordName + ".pdf"))
                        .as(DocumentType.PDF)
                        .schedule();


//对应 schedule的运行
future.get();

这种方式可以达成所见即所得。

PS:

之前提出了模版难以修改,是因为模版中要使用{替换名称}的方式, 但是word有时会自动截断一个字符串, 导致实际上变成了{替 换名称 }的样式, 需要多改几次试下,连续输入试一下。

有一种比较简单的方式,就是将word文件的后缀名改成zip ,然后拿出document.xml 可以在这个里面直接改,名称改回后记得打开看是否报错, 如果报错,另存一下,就可以去掉报错。

相关推荐
~烈19 小时前
Acrobat Pro DC 便携精简版 PDF 专业处理工具安装教程
pdf·acrobat pro dc·pdf 专业处理工具
我不是QI19 小时前
Master PDF Editor 逆向复现
汇编·安全·pdf
蓝创工坊Blue Foundry20 小时前
多个同模板 PDF,怎样批量提取同一类字段到 Excel
运维·数据库·pdf·自动化·ocr·excel
asdzx672 天前
Python PDF 拆分实战指南:单页拆分与按需页码范围拆分
开发语言·python·pdf
慧都小妮子3 天前
Aspose.CAD for .NET 26.6 更新发布
pdf·.net·3d渲染·cad·pdf导出·ifc编辑
蝉蜕日记3 天前
AccumuPDF高级版 v2.57 | 视图文转换PDF,合并分割压缩
android·智能手机·pdf·生活·软件需求
蓝创工坊Blue Foundry3 天前
PDF 批量提取指定内容到 Excel:按字段整理多个 PDF 的方法
pdf·ocr·excel·文心一言·paddlepaddle·paddle
Am-Chestnuts3 天前
AI长回答批量导出PDF与长图:多轮内容整理和免费Markdown备份
人工智能·pdf·aigc
zyplayer-doc3 天前
研发接口文档怎么长期维护:zyplayer-doc把API、Markdown和变更记录放进同一个知识库
大数据·数据库·人工智能·笔记·pdf·ocr
蜡台3 天前
uniapp pdf文件预览组件
pdf·uni-app·合同·pdfh5