{"slug": "a-computational-approach-to-measuring-semantic-change-in-sanskrit-literature", "title": "A Computational Approach to Measuring Semantic Change in Sanskrit Literature", "summary": "A new arXiv paper (arXiv:2609.25012v1) reports that diachronic word embeddings can track semantic change in Sanskrit, an ancient low-resource language, with 19 of 21 testable shifts moving in the philologically attested direction (sign test, p=0.00011). The author assembled a 2.7M-token corpus spanning four canonical periods and recovered word boundaries using a neural byte-level sandhi splitter and lemmatizer, then trained per-period embeddings across configurations. The work tests whether a paradigm largely validated on modern, high-resource, well-segmented languages transfers to a language whose sandhi, morphological inflection, compounding, and polysemy pose distinct challenges.", "body_md": "arXiv:2609.25012v1 Announce Type: new \nAbstract: Diachronic word embeddings have become the modern standard for tracking semantic change, yet they have been largely validated on modern, high-resource, and well-segmented languages. This paper tests whether the paradigm transfers to Sanskrit, an ancient, low-resource language whose phonological fusion (sandhi), morphological inflection, compounding, and polysemy pose a unique challenge. I assemble a 2.7M-token corpus spanning four canonical periods, recover word boundaries with a neural byte-level sandhi splitter and lemmatizer, and train per-period embeddings across configurations. To evaluate the system, I curate a validation set from historical scholarship and test recovery directionally with anchor displacement. Of 21 testable shifts, 19 move in the philologically attested direction (sign test, p=0.00011). I further show which configuration the language forces and comment on opportunities for improvement.", "url": "https://wpnews.pro/news/a-computational-approach-to-measuring-semantic-change-in-sanskrit-literature", "canonical_source": "https://arxiv.org/abs/2609.25012", "published_at": "2026-09-23 04:00:00+00:00", "updated_at": "2026-09-23 04:25:20.121516+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "ai-research"], "entities": ["arXiv", "Sanskrit"], "alternates": {"html": "https://wpnews.pro/news/a-computational-approach-to-measuring-semantic-change-in-sanskrit-literature", "markdown": "https://wpnews.pro/news/a-computational-approach-to-measuring-semantic-change-in-sanskrit-literature.md", "text": "https://wpnews.pro/news/a-computational-approach-to-measuring-semantic-change-in-sanskrit-literature.txt", "jsonld": "https://wpnews.pro/news/a-computational-approach-to-measuring-semantic-change-in-sanskrit-literature.jsonld"}}