Lightweight PDF parser with layout, tables, formulas and bounding boxes Developer Beatriz Almeida released papero, an open-source PDF parser that extracts reading order, tables, formulas, figures and bounding boxes from PDFs using geometry alone, with no ML models and CPU-only execution. The tool installs via `pip install pdf-text-api`, runs in the browser, Python or as a REST API, and exports to Markdown, JSON, Word, Excel, HTML and CSV, with OCR for scanned pages and Apache Tika support for DOCX, PPTX, XLSX, EPUB and HTML. papero is aimed at RAG, search and LLM pipelines that need document structure rather than raw text. PDF → Markdown · JSON · Word · Excel — with reading order, tables, formulas, figures and the position of every block. CPU only. No ML models. Runs in your browser, in Python, or as an API. ▶ Try it in your browser https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/ · Quick start quick-start · Benchmarks