grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP Researchers from the University of Moratuwa have released grapheme-kit, an open-source Python library that computes lexical distance, similarity, and evaluation metrics on grapheme clusters rather than Unicode code points, providing more accurate text processing for complex scripts like Tamil and Sinhala. An OCR case study shows grapheme-level metrics yield a more faithful evaluation of such writing systems. grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP By Izzath Nisfer, Ashini Kavindya, Ovindu Atukorala, Purushoth Velayuthan, Menan VelayuthanSource: arXiv cs.CL https://arxiv.org/list/cs.CL/recent arXiv:2607.22456v1 Announce Type: new Abstract: Existing lexical distance, similarity, and evaluation /glossary/evaluation metrics operate on Unicode code points, which can misrepresent errors in writing systems where a single grapheme is represented by multiple Unicode code points. We introduce grapheme-kit, an open-source Python library that extends these metrics to operate on grapheme clusters instead. The library also provides improved grapheme processing for Tamil and Sinhala, including accurate grapheme cluster identification and grapheme composition/decomposition utilities. Through an OCR case study, we demonstrate that grapheme-level metrics provide a more faithful evaluation of complex scripts.Get AI news in your inbox Daily digest of what matters in AI.