cd /news/natural-language-processing/invisible-soft-hyphens-wrecked-rag-s… · home › topics › natural-language-processing › article
[ARTICLE · art-143434] src=dev.to ↗ pub= topic=natural-language-processing verified=true sentiment=↑ positive

Invisible soft hyphens wrecked RAG search on converted books — 4,000 U+00AD characters hiding in clean text

A developer building a RAG pipeline over converted EPUB books found that invisible U+00AD soft hyphens inherited from print typography were breaking full-text search and embeddings, with 4,213 such characters counted in a single 400-page technical manual. A benchmark of 300 extracted terms showed roughly 38% failed to match before normalization, but stripping soft hyphens and zero-width characters plus NFC-normalizing brought recall to 100%. The fix is now a normalization stage in the developer's EPUB-to-Markdown API that cleans every heading, paragraph and table cell before Markdown is written.

by read1 min views2 publishedOct 1, 2026

An agent indexing a 400-page technical manual hit something maddening last week: full-text search on my converted output found nothing for terms visibly on the page. "Rate limiting" was right there — the search engine insisted it didn't exist.

The Markdown looked flawless. I read five chapters and saw nothing wrong. Then I stopped trusting my eyes and checked the bytes.

The source EPUB had inherited print-shop typography. Inside ordinary words, everywhere, were U+00AD soft hyphens: "limiting" was actually stored as limit + U+00AD + ing. A script counted 4,213 of them in that one book. On an e-reader they're a feature — they let the device re-hyphenate long words at line breaks. Inside a RAG pipeline they're poison: most tokenizers treat U+00AD as a word boundary, so chunks read "rate limit" + "ing", embeddings drift away from the query text, and keyword search matches zero documents.

Quick benchmark: I extracted 300 terms from the book and searched the converted corpus. About 38% of terms failed to match before normalization. After stripping U+00AD and its cousins — zero-width space U+200B, zero-width joiner — plus NFC-normalizing, recall hit 100%.

The fix is now a normalization stage that runs in my EPUB-to-Markdown API (https://x402.freeq.one/tools/epub_to_markdown.html) on every heading, paragraph and table cell before the Markdown gets written: strip soft hyphens, drop zero-width characters, normalize Unicode form, and merge lines that were hyphen-split.

Lesson I keep relearning in document conversion: "it renders correctly" proves nothing. Rendering, diffing, even copy-paste all hide invisible codepoints. If your pipeline ingests converted books or manuals, grep your extracted text for U+00AD before blaming your embedding model — five seconds of grep $'\u00ad' would have saved me a whole afternoon of debugging.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @u+00ad 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/invisible-soft-hyphe…] indexed:0 read:1min 2026-10-01 · —