Papero是一款轻量级开源PDF解析工具,可将PDF文档转换为Markdown、JSON、Word和Excel等多种格式,同时完整保留原始布局、表格、公式、图形及各元素的位置信息1。该项目由Beatriz Almeida开发,采用MIT许可证发布,GitHub仓库为beatrizalmeidaf/papero-pdf-text-extractor1。
该工具运行于CPU之上,无需任何机器学习模型支持,提供浏览器应用、Python库和API三种使用方式1。在笔记本CPU上处理密集型arXiv论文时表现稳定,54篇论文的测试中失败率为0,中位处理速度为39毫秒/页1。Papero采用双引擎架构同时运行,包括PDFium上的布局引擎和Apache Tika(支持光学字符识别、元数据提取及非PDF格式处理)1。用户可通过pip install pdf-text-api命令安装,或使用浏览器应用和Docker部署等方式1。
Papero is a lightweight open-source PDF parsing tool that converts documents into multiple formats while maintaining their original structure and content details 1. The parser supports output in Markdown, JSON, Word, and Excel formats, preserving reading order, tables, formulas, graphics, and position information for each extracted element 1.
The tool operates on CPU-only infrastructure without requiring machine learning models, and can be deployed in three different environments: web browser, Python, and API 1. Papero uses a dual-engine architecture, combining a layout engine built on PDFium with Apache Tika for optical character recognition, metadata extraction, and support for non-PDF formats 1. Installation is available through pip package management or via browser-based and Docker deployment options 1.
Performance testing on complex multi-column documents such as arXiv papers demonstrates reliable operation, with zero failure rate across 54 papers tested and a median processing speed of 39 milliseconds per page on laptop CPU hardware 1. The project is released under the MIT license and hosted on GitHub as beatrizalmeidaf/papero-pdf-text-extractor 1.
评论
还没有评论,欢迎留下第一条。