# Turn PDFs into clean Markdown chunks for your RAG pipeline (without writing a parser)

> Source: <https://dev.to/phenixik/turn-pdfs-into-clean-markdown-chunks-for-your-rag-pipeline-without-writing-a-parser-3ih8>
> Published: 2026-10-09 12:14:43+00:00

PDF parsing is the boring part of every RAG project. Line breaks in the middle of sentences, lost headings, headers and footers mixed into the text, no page numbers to cite. You can spend days tuning `pypdf` or `pdfplumber`, or you can skip that part.

Here's a setup-free way to get LLM-ready text from PDFs, including PDFs you haven't found yet.

``` python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("digitalni.produkty.pro.zivot/pdf-text-extractor").call(
    run_input={
        "urls": ["https://arxiv.org/pdf/1706.03762"],
        "includeMarkdown": True,
        "chunkSize": 1000,
        "chunkOverlap": 100,
    }
)
for pdf in client.dataset(run["defaultDatasetId"]).iterate_items():
    for chunk in pdf.get("chunks", []):
        ...  # embed and upsert into your vector DB
```

It also works as a tool for AI agents through the Apify MCP server, so Claude or ChatGPT can read PDFs on demand.

A test batch of 5 PDFs with 131 pages and ~78k words was processed in about 4 seconds. Price: **$0.003 per PDF**, any number of pages. Scanned (image-only) PDFs are detected and skipped, and you are not charged for them.

*Disclosure: I built this tool. Feature requests welcome in the Issues tab.*
