cd/entity/PyMuPDF· home entities PyMuPDF
grep -l @pymupdf /news/*.json | wc -l → 23

PyMuPDF

mentions 23 type Organization page 1/2 feed RSS

// recent coverage 23 mentions

16:12
2026-08-20
promptcube3.com
developer-tools

Exit codes lie when PDF extraction yields nothing

A developer's integration with Claude Code and markitdown falsely reported a PDF as empty because exit codes only indicate process completion, not content yield, and byte thresholds miss near-misses l…

14:51
2026-08-19
dev.to
developer-tools

My AI said the PDF was empty. The PDF was not empty.

A developer using Microsoft's markitdown converter discovered that a PDF containing only raster images produced a zero-byte output with exit code 0, leading an AI model to incorrectly report the docum…

00:48
2026-07-25
promptcube3.com
artificial-intelligence

RAG pipeline crashes on scientific characters: BGE-small-en-v1.5

A RAG pipeline using LlamaIndex and the BGE-small-en-v1.5 embedding model crashes with a 'TextEncodeInput must be Union[TextInputSequence, Tuple[InputSequence, InputSequence]]' error when processing s…

18:05
2026-07-24
promptcube3.com
artificial-intelligence

AI Workflow: Optimizing Resume Parsing with LLM Agents

A developer improved resume parsing accuracy from 65% to 94% by shifting date arithmetic from an LLM prompt to a Python post-processing script, using a two-step chain that first extracts raw date rang…

07:15
2026-07-15
github.com
artificial-intelligence

PDF Batch Translator

A new open-source CLI tool, PDF Batch Translator, translates English PDFs into Japanese Markdown while stripping non-body content and preserving structure. Developed by an unnamed creator, the tool us…

18:08
2026-07-14
machinebrief.com
artificial-intelligence

AI Textbook Auditor: The Future of Educational Quality Assurance

The AI Textbook Auditor, a system using advanced LLM agents, has been tested on Romanian upper-secondary textbooks, identifying 56 technical findings in a computer science textbook with 62.5% expert-v…

10:00
2026-06-28
dev.to
large-language-models

From Regex Hell to AI: How I Finally Tamed Messy PDF Invoices

A developer built an AI-powered system to extract structured data from messy PDF invoices, replacing unreliable regex and rule-based parsers. By using a large language model with few-shot prompting an…

03:11
2026-06-24
byteiota.com
large-language-models

Baidu Unlimited-OCR: One-Shot PDF Parsing Is Here

Baidu released Unlimited-OCR, a new model that uses Reference Sliding Window Attention to parse up to 40 PDF pages in a single inference pass, eliminating the linear memory growth of traditional LLM-b…

11:35
2026-06-23
github.com
artificial-intelligence

Unlimited OCR: One-shot long-horizon parsing

Baidu released Unlimited-OCR, a one-shot long-horizon parsing model that extends DeepSeek-OCR, on June 22, 2026. The open-source model supports single-image and multi-page PDF parsing with configurabl…

16:32
2026-06-12
sgaud.com
large-language-models

A PDF that changes based on who is reading

A developer created a PDF that renders identically to human readers but extracts as clean markdown for machines, using a 2001 PDF specification property that allows replacement text for marked content…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics