{"slug": "trilingual-topic-modeling-of-sri-lankan-parliamentary-debates", "title": "Trilingual Topic Modeling of Sri Lankan Parliamentary Debates", "summary": "A new arXiv paper (2608.20365v1) presents an end-to-end framework using LLM-based text extraction and multilingual embedding with density-based clustering to topic-model 19,553 trilingual Sri Lankan parliamentary speeches from 2017-2026, recovering 30 macro-topics with a cluster purity of 0.673 and aligning with major national events such as the 2019 Easter Sunday attacks and the 2022 economic crisis. The authors propose a hybrid semantic-lexical extension called BiTopic to improve interpretability and recover discarded speeches, while noting that traditional LDA fails due to cross-lingual fragmentation.", "body_md": "arXiv:2608.20365v1 Announce Type: new\nAbstract: Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs, multilingual scripts, and agglutinative morphology. We present an end-to-end framework that addresses these challenges through LLM-based text extraction followed by a multilingual embedding and density-based clustering pipeline for topic modeling. A hybrid semantic-lexical extension, BiTopic, is further explored to improve interpretability and recover speeches otherwise discarded as noise. Applied to 19,553 speeches spanning 2017-2026, the pipeline recovers 30 macro-topics achieving a cluster purity (BCP) of 0.673, whose temporal trajectories align unsupervised with major national events including the 2019 Easter Sunday attacks and the 2022 economic crisis. Traditional LDA fails on this corpus due to cross-lingual fragmentation, whereas the proposed approach successfully identifies thematic structure across all three languages without supervision.", "url": "https://wpnews.pro/news/trilingual-topic-modeling-of-sri-lankan-parliamentary-debates", "canonical_source": "https://arxiv.org/abs/2608.20365", "published_at": "2026-08-24 04:00:00+00:00", "updated_at": "2026-08-24 04:14:41.945545+00:00", "lang": "en", "topics": ["natural-language-processing", "large-language-models", "machine-learning"], "entities": ["arXiv", "BiTopic", "Sri Lankan Parliament"], "alternates": {"html": "https://wpnews.pro/news/trilingual-topic-modeling-of-sri-lankan-parliamentary-debates", "markdown": "https://wpnews.pro/news/trilingual-topic-modeling-of-sri-lankan-parliamentary-debates.md", "text": "https://wpnews.pro/news/trilingual-topic-modeling-of-sri-lankan-parliamentary-debates.txt", "jsonld": "https://wpnews.pro/news/trilingual-topic-modeling-of-sri-lankan-parliamentary-debates.jsonld"}}