# Lightweight PDF parser with layout, tables, formulas and bounding boxes

> Source: <https://github.com/beatrizalmeidaf/papero-pdf-text-extractor>
> Published: 2026-10-01 16:14:40+00:00

PDF → Markdown · JSON · Word · Excel — with reading order, tables, formulas, figures and the position of every block.

  CPU only. No ML models. Runs in your browser, in Python, or as an API.

[**▶ Try it in your browser**](https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/)
   · 
  [**Quick start**](#quick-start)
   · 
  **Benchmarks**

<sub>30 seconds in the browser app: load a PDF, inspect any block, check tables and formulas, export to Word. Your PDF never leaves your machine. ([MP4](https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/blob/main/web/assets/demo.mp4))</sub>

Getting the *text* out of a PDF is easy. Getting its **structure** back — which column comes first, which lines are a table, where the formula is — is what makes the output usable for RAG, search and LLMs. papero does that with plain geometry, so it stays fast on a laptop CPU.

| **📖 Reading order** Two- and three-column papers read column by column. Headers, footers, page numbers and repeated logos are set aside. | **▦ Real tables** Ruled, borderless and LaTeX *booktabs* tables come back as rows and columns — multi-line cells included. Export to CSV or Excel. | **∑ Formulas** Superscripts, subscripts and math symbols become LaTeX ( `E = mc^{2}` ), plus a cropped image of the formula. | 
| **📍 Position of everything** Every block has a bounding box — cite the exact spot in a RAG answer, draw over the page, or crop it. | **🖼 Figures & charts** Images and vector charts are cropped to PNG, with their caption, axis labels and legend kept together. | **📝 Back to Word** Alignment, indents, line spacing, bold runs and fonts are kept, so a `.docx` export looks like the original page. | 

Also: accents drawn as separate glyphs in LaTeX PDFs (`Computa¸ca˜o` → `Computação`), invisible white text used by form generators is dropped, scanned pages go through OCR, and DOCX/PPTX/XLSX/EPUB/HTML are read through Apache Tika.

```
pip install pdf-text-api
python
from pdf_text_api import extract

doc = extract("paper.pdf")
print(doc.to_markdown())
```

Or skip the install: **[open the browser app](https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/)**, drop a PDF, export to the format you need.

## **More Python** — tables, formulas, positions, images, options

``` python
from pdf_text_api import extract, extract_text

doc = extract("paper.pdf", images=True)

doc.tables[0].rows  # [["Model", "Accuracy"], ["Base", "0.81"], ...]
doc.formulas[0].latex  # "E = mc^{2}"
doc.figures[0].image.data  # PNG bytes

for block in doc.pages[0].blocks:  # reading order, with positions
    print(block.type, block.bbox, block.text[:60])

doc.to_html()  # keeps alignment and indents
doc.to_dict()  # the full JSON

extract("slides.pptx").to_markdown()  # any format Apache Tika reads
extract_text("contract.pdf").text  # fastest: clean text only
```

| Option | Default |  | 
|---|---|---|
| `pages` | all | `"1-3,5,10-"` | 
| `images` | `False` | crop figures, tables and formulas to PNG | 
| `tables` /`formulas` | `True` | detection on/off | 
| `ocr` | `"auto"` | `"auto"` (scanned pages only),`"force"` ,`"off"` | 
| `ocr_language` | `"por+eng"` | Tesseract languages | 
| `tika` | `True` | `False` runs the layout engine alone (no Java) | 
| `workers` | `1` | processes for long documents | 

## **CLI**

```
pdf-text-api extract paper.pdf -o paper.md --images   # Markdown + images/ folder
pdf-text-api extract paper.pdf -o paper.json          # format from the extension
pdf-text-api extract paper.pdf -f csv -o tables.csv   # tables only
pdf-text-api extract paper.pdf -p 1-5 -f html
pdf-text-api extract paper.pdf --fast                 # clean text only
pdf-text-api serve --port 8000                        # API + browser app
```

## **REST API & Docker**

```
docker compose up        # API + Apache Tika + Tesseract + browser app on :8000
curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=markdown"
curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=zip&images=true" -o paper.zip
curl -F "file=@paper.pdf" "localhost:8000/v1/extract?per_page=true"     # blocks + positions
```

One endpoint, `POST /v1/extract`; interactive docs at `/docs`.

| Parameter | Default |  | 
|---|---|---|
| `mode` | `structured` | `structured` (layout + Tika) or`fast` (text only) | 
| `format` | `json` | `json` ,`markdown` ,`text` ,`html` ,`csv` ,`zip` | 
| `pages` | all | `1-3,5,10-` | 
| `per_page` | `false` | include pages, blocks and positions in the JSON | 
| `images` | `false` | crop figures, tables and formulas | 
| `ocr` | `auto` | `auto` ,`force` ,`off` | 

Configuration through environment variables — see [`.env.example`](https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/blob/main/.env.example).

Every block knows what it is and where it was:

```
{
  "type": "table",
  "bbox": [56.7, 294.8, 481.9, 374.2],
  "rows": [["Model", "Accuracy"], ["Base", "0.81"]],
  "caption": "Table 1: Comparison between models."
}
```

| Output | Python · CLI · API | Browser app | 
|---|---|---|
| Markdown, plain text, JSON | ✓ | ✓ | 
| HTML (keeps alignment and indents) | ✓ | ✓ | 
| CSV of the tables, ZIP with images | ✓ | ✓ | 
| Word `.docx` that keeps the page's look | — | ✓ | 
| Excel `.xlsx` , one sheet per table | — | ✓ | 

## **Full JSON schema and block types**

```
{
  "schema": "pdf-text-api/document@1",
  "engine": "tika+pdfium",
  "page_count": 12,
  "metadata": { "title": "...", "author": "...", "language": "en" },
  "pages": [{
    "number": 1, "width": 595.3, "height": 841.9,
    "blocks": [{
      "id": "p1-b4", "type": "paragraph", "bbox": [74.0, 217.0, 522.0, 275.0],
      "text": "Atestamos que a estudante ...",
      "style": { "pt": 11.0, "font": "Arial", "bold": false },
      "format": { "align": "justify", "first_line": 42.7, "line_spacing": 1.8 },
      "runs": [{ "text": "FULANA DE TAL", "bold": true, "italic": false, "script": null }]
    }]
  }]
}
```

Block types: `heading` (with `level`), `paragraph`, `list_item` (with `marker`), `table` (with `rows`), `figure`, `formula` (with `latex`), `caption`, `code`, and — kept apart from the text — `header`, `footer`, `page_number`. Bounding boxes are `[x0, y0, x1, y1]` in points, origin at the top-left of the page.

Dense arXiv papers (multi-column, formulas, tables, figures) on one laptop CPU, no GPU. `papero · fast` returns clean text; `papero · structured` also rebuilds reading order, tables, formulas and figures — **0 failures on 54 papers, 39 ms per page (median)**. Reproduce with [`benchmarks/`](https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/blob/main/benchmarks).

|  | papero | PyMuPDF | pdfplumber | pypdf | Docling | Marker | 
|---|---|---|---|---|---|---|
| License | MIT | AGPL | MIT | BSD | MIT | GPL | 
| Needs ML models / PyTorch | no | no | no | no | yes | yes | 
| Multi-column reading order | ✓ | partial | — | — | ✓ | ✓ | 
| Structured tables | ✓ | ✓ | ✓ | — | ✓ | ✓ | 
| Formulas | LaTeX from glyphs + image | — | — | — | ✓ | ✓ | 
| Bounding boxes | ✓ | ✓ | ✓ | — | ✓ | ✓ | 
| DOCX / PPTX / XLSX / EPUB | ✓ | partial | — | — | ✓ | partial | 
| Runs entirely in the browser | ✓ | — | — | — | — | — | 

ML-based tools still win on very irregular layouts and complex math (stacked fractions, matrices) — papero gives you the formula as approximate LaTeX **and** as an image so nothing is lost.

Two engines run on the same file **at the same time**:

- **A layout engine on PDFium** reads every glyph with its position, font and size, plus every rule and image, and rebuilds columns, tables, formulas, lists and figures with a column-aware XY-cut.
- **Apache Tika** adds metadata, tagged-PDF headings, OCR (Tesseract) and every non-PDF format.

The browser app runs the same algorithm ported to JavaScript on pdf.js, and CI checks block by block that both engines agree.

## **Limitations**

- **Math:** LaTeX is rebuilt from glyphs — stacked fractions, matrices and big radicals come out linear (the cropped image is always there).
- **Borderless tables** with very narrow gaps between columns can read as text.
- **Scanned PDFs** need OCR, which runs on the server path (Tesseract is in the Docker image).
- **Word/Excel export** is in the browser app for now.

## **Development**

```
git clone https://github.com/beatrizalmeidaf/papero-pdf-text-extractor.git && cd pdf-text-extractor
pip install -e ".[dev]"
pytest -q                                   # includes real-world regressions
ruff check src tests && ruff format --check src tests
npm install --prefix tests/js && python tests/js/expected.py tests/js/out && node tests/js/parity.mjs tests/js/out
python -m http.server -d web                # browser app at http://localhost:8000
```

`src/pdf_text_api/` is the Python engine, API and CLI · `web/` is the browser app (GitHub Pages) · `tests/js/` checks the two engines agree · `benchmarks/` downloads the dataset and draws the chart.

Found a PDF papero gets wrong? **That's the most useful issue you can open** — attach the file (or a page of it) and say what you expected. Reading order, tables, formulas, encoding, OCR and browser/server differences are all fair game.

If papero saves you time, **a ⭐ helps other people find it.**

**Keywords:** PDF to Markdown · PDF to JSON · PDF to Word · PDF to Excel · PDF table extraction · PDF parser · document parsing · layout analysis · reading order · multi-column PDF · formula extraction · LaTeX · bounding boxes · OCR · Apache Tika · PDFium · pdf.js · RAG preprocessing · LLM document loader · Docling alternative · PyMuPDF alternative · converter PDF para Markdown, Word e Excel · extrair tabelas de PDF · extrair texto de PDF mantendo a formatação · OCR de PDF escaneado

<sub>MIT © Beatriz Almeida · package and imports keep the name `pdf-text-api` / `pdf_text_api` for compatibility.</sub>
