PDF parsing is the boring part of every RAG project. Line breaks in the middle of sentences, lost headings, headers and footers mixed into the text, no page numbers to cite. You can spend days tuning pypdf or pdfplumber, or you can skip that part.
Here's a setup-free way to get LLM-ready text from PDFs, including PDFs you haven't found yet.
from apify_client import ApifyClient
client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("digitalni.produkty.pro.zivot/pdf-text-extractor").call(
run_input={
"urls": ["https://arxiv.org/pdf/1706.03762"],
"includeMarkdown": True,
"chunkSize": 1000,
"chunkOverlap": 100,
}
)
for pdf in client.dataset(run["defaultDatasetId"]).iterate_items():
for chunk in pdf.get("chunks", []):
... # embed and upsert into your vector DB
It also works as a tool for AI agents through the Apify MCP server, so Claude or ChatGPT can read PDFs on demand.
A test batch of 5 PDFs with 131 pages and ~78k words was processed in about 4 seconds. Price: $0.003 per PDF, any number of pages. Scanned (image-only) PDFs are detected and skipped, and you are not charged for them.
Disclosure: I built this tool. Feature requests welcome in the Issues tab.