{"slug": "turn-pdfs-into-clean-markdown-chunks-for-your-rag-pipeline-without-writing-a", "title": "Turn PDFs into clean Markdown chunks for your RAG pipeline (without writing a parser)", "summary": "A developer built an Apify actor that converts PDFs into clean, chunked Markdown for RAG pipelines without requiring custom parsing code, handling headers, footers, and mid-sentence line breaks automatically. A test batch of 5 PDFs totaling 131 pages and roughly 78,000 words was processed in about 4 seconds at $0.003 per PDF, with image-only scanned PDFs detected and skipped at no charge. The tool is also exposed as an agent tool via the Apify MCP server so assistants like Claude or ChatGPT can read PDFs on demand.", "body_md": "PDF parsing is the boring part of every RAG project. Line breaks in the middle of sentences, lost headings, headers and footers mixed into the text, no page numbers to cite. You can spend days tuning `pypdf` or `pdfplumber`, or you can skip that part.\n\nHere's a setup-free way to get LLM-ready text from PDFs, including PDFs you haven't found yet.\n\n``` python\nfrom apify_client import ApifyClient\n\nclient = ApifyClient(\"<YOUR_API_TOKEN>\")\nrun = client.actor(\"digitalni.produkty.pro.zivot/pdf-text-extractor\").call(\n    run_input={\n        \"urls\": [\"https://arxiv.org/pdf/1706.03762\"],\n        \"includeMarkdown\": True,\n        \"chunkSize\": 1000,\n        \"chunkOverlap\": 100,\n    }\n)\nfor pdf in client.dataset(run[\"defaultDatasetId\"]).iterate_items():\n    for chunk in pdf.get(\"chunks\", []):\n        ...  # embed and upsert into your vector DB\n```\n\nIt also works as a tool for AI agents through the Apify MCP server, so Claude or ChatGPT can read PDFs on demand.\n\nA test batch of 5 PDFs with 131 pages and ~78k words was processed in about 4 seconds. Price: **$0.003 per PDF**, any number of pages. Scanned (image-only) PDFs are detected and skipped, and you are not charged for them.\n\n*Disclosure: I built this tool. Feature requests welcome in the Issues tab.*", "url": "https://wpnews.pro/news/turn-pdfs-into-clean-markdown-chunks-for-your-rag-pipeline-without-writing-a", "canonical_source": "https://dev.to/phenixik/turn-pdfs-into-clean-markdown-chunks-for-your-rag-pipeline-without-writing-a-parser-3ih8", "published_at": "2026-10-09 12:14:43+00:00", "updated_at": "2026-10-09 12:21:31.910970+00:00", "lang": "en", "topics": ["ai-tools", "ai-agents", "agent-protocols", "developer-tools"], "entities": ["Apify", "Claude", "ChatGPT", "Apify MCP server"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/turn-pdfs-into-clean-markdown-chunks-for-your-rag-pipeline-without-writing-a", "markdown": "https://wpnews.pro/news/turn-pdfs-into-clean-markdown-chunks-for-your-rag-pipeline-without-writing-a.md", "text": "https://wpnews.pro/news/turn-pdfs-into-clean-markdown-chunks-for-your-rag-pipeline-without-writing-a.txt", "jsonld": "https://wpnews.pro/news/turn-pdfs-into-clean-markdown-chunks-for-your-rag-pipeline-without-writing-a.jsonld"}}