{"slug": "invisible-soft-hyphens-wrecked-rag-search-on-converted-books-4000-u-00ad-hiding", "title": "Invisible soft hyphens wrecked RAG search on converted books — 4,000 U+00AD characters hiding in clean text", "summary": "A developer building a RAG pipeline over converted EPUB books found that invisible U+00AD soft hyphens inherited from print typography were breaking full-text search and embeddings, with 4,213 such characters counted in a single 400-page technical manual. A benchmark of 300 extracted terms showed roughly 38% failed to match before normalization, but stripping soft hyphens and zero-width characters plus NFC-normalizing brought recall to 100%. The fix is now a normalization stage in the developer's EPUB-to-Markdown API that cleans every heading, paragraph and table cell before Markdown is written.", "body_md": "An agent indexing a 400-page technical manual hit something maddening last week: full-text search on my converted output found nothing for terms visibly on the page. \"Rate limiting\" was right there — the search engine insisted it didn't exist.\n\nThe Markdown looked flawless. I read five chapters and saw nothing wrong. Then I stopped trusting my eyes and checked the bytes.\n\nThe source EPUB had inherited print-shop typography. Inside ordinary words, everywhere, were U+00AD soft hyphens: \"limiting\" was actually stored as `limit` + U+00AD + `ing`. A script counted 4,213 of them in that one book. On an e-reader they're a feature — they let the device re-hyphenate long words at line breaks. Inside a RAG pipeline they're poison: most tokenizers treat U+00AD as a word boundary, so chunks read \"rate limit\" + \"ing\", embeddings drift away from the query text, and keyword search matches zero documents.\n\nQuick benchmark: I extracted 300 terms from the book and searched the converted corpus. About 38% of terms failed to match before normalization. After stripping U+00AD and its cousins — zero-width space U+200B, zero-width joiner — plus NFC-normalizing, recall hit 100%.\n\nThe fix is now a normalization stage that runs in my EPUB-to-Markdown API ([https://x402.freeq.one/tools/epub_to_markdown.html](https://x402.freeq.one/tools/epub_to_markdown.html)) on every heading, paragraph and table cell before the Markdown gets written: strip soft hyphens, drop zero-width characters, normalize Unicode form, and merge lines that were hyphen-split.\n\nLesson I keep relearning in document conversion: \"it renders correctly\" proves nothing. Rendering, diffing, even copy-paste all hide invisible codepoints. If your pipeline ingests converted books or manuals, grep your extracted text for U+00AD before blaming your embedding model — five seconds of `grep $'\\u00ad'` would have saved me a whole afternoon of debugging.", "url": "https://wpnews.pro/news/invisible-soft-hyphens-wrecked-rag-search-on-converted-books-4000-u-00ad-hiding", "canonical_source": "https://dev.to/imapphelp/invisible-soft-hyphens-wrecked-rag-search-on-converted-books-4000-u00ad-characters-hiding-in-nn9", "published_at": "2026-10-01 19:30:47+00:00", "updated_at": "2026-10-01 21:30:56.381927+00:00", "lang": "en", "topics": ["natural-language-processing", "ai-agents", "developer-tools", "ai-infrastructure"], "entities": ["U+00AD", "U+200B", "EPUB", "Markdown", "x402.freeq.one"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/invisible-soft-hyphens-wrecked-rag-search-on-converted-books-4000-u-00ad-hiding", "markdown": "https://wpnews.pro/news/invisible-soft-hyphens-wrecked-rag-search-on-converted-books-4000-u-00ad-hiding.md", "text": "https://wpnews.pro/news/invisible-soft-hyphens-wrecked-rag-search-on-converted-books-4000-u-00ad-hiding.txt", "jsonld": "https://wpnews.pro/news/invisible-soft-hyphens-wrecked-rag-search-on-converted-books-4000-u-00ad-hiding.jsonld"}}