New open-source tool 'papero' extracts PDF structure without ML models
A new PDF parsing tool called papero converts PDFs into Markdown, JSON, Word or Excel while preserving reading order, tables, formulas, figures and the bounding-box position of every block. It runs entirely on CPU using geometric analysis rather than machine learning models, and is available as a browser app, a Python package (pip install papero-extract), or an API. The browser version processes files locally so the PDF does not leave the user's machine.
GoKawiil's interpretation of the reporting above, not reported fact.
Extracting a PDF's structure, not just its raw text, is often what makes documents usable for retrieval-augmented generation, search indexing and citation-grounded LLM answers, since knowing where a table or formula sits lets systems reference exact locations. By avoiding ML models, papero could appeal to developers wanting faster, cheaper, and privacy-preserving document processing on ordinary hardware. Its handling of multi-column academic papers, LaTeX formulas and scanned OCR pages suggests a target audience in research and technical document workflows.
- papero converts PDFs to Markdown, JSON, Word, or Excel while preserving layout, tables, formulas and bounding boxes.
- It uses CPU-only geometric methods instead of machine learning, and can run in-browser, via Python, or as an API.
- Supports complex features like multi-column reading order, LaTeX formula conversion, OCR for scans, and reading DOCX/PPTX/XLSX/EPUB/HTML via Apache Tika.
Source: github.com, 2026-10-01
Published there as: “Lightweight PDF parser with layout, tables, formulas and bounding boxes”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.