Grep Without Word Boundaries: 70 Tokens Across 7.68M Words, and os Is Real 0.1% of the Time A developer's scan of a 7.68-million-word knowledge base found that the token 'os' appears as a real word only 0.1% of the time, highlighting the challenges of tokenization without word boundaries. The analysis also cites OpenAI's GPT-4 contamination report and a proposed method to detect citation laundering in retrieval systems. Sources · OpenAI, "GPT-4 Technical Report," arXiv:2303.08774, Appendix C "Contamination on professional and academic exams" — contamination methodology and its stated limitations, quoted verbatim from the report. · Liu, Z., et al., "HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards," arXiv:2608.06012 submitted August 6, 2026 — abstract quoted verbatim; the citation-laundering finding and repair. · Corpus statistics and token measurements are from our own knowledge-base scan of August 9, 2026 2,011 files; 7,683,956 words , reproducible with the measurement script described above; carrier examples are quoted from that corpus.