New open-source tool 'papero' extracts PDF structure without ML models
A new PDF parsing tool called papero converts PDFs into Markdown, JSON, Word or Excel while preserving reading order, tables, formulas, figures and the bounding-box position of every block. It runs entirely on CPU using geometric analysis rather than machine learning models, and is available as a browser app, a Python package (pip install papero-extract), or an API. The browser version processes files locally so the PDF does not leave the user's machine.