{"slug": "beyond-top-words-monotm-for-topic-modeling-with-interpretable-monosemantic", "title": "Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features", "summary": "Researchers introduced MonoTM, an interpretable topic modeling framework that decouples document-topic mixture estimation from semantic interpretation, according to a paper posted to arXiv as 2609.09575v1. MonoTM estimates topic mixtures from the full sparse autoencoder (SAE) bag-of-features representation and then, with those mixtures fixed, learns topic descriptors over a separate vocabulary of corpus-grounded semantic features. Across three benchmark corpora, the authors report that mixture estimation and semantic interpretation favor different SAE configurations and feature subsets, allowing topics to be represented by semantic units more meaningful than individual words for downstream corpus analysis.", "body_md": "arXiv:2609.09575v1 Announce Type: new \nAbstract: Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics. Sparse autoencoders (SAEs) offer a way to move beyond word-level descriptors by extracting interpretable features from dense representations, yet how feature interpretability relates to topic-inference quality remains unclear. We introduce \\textbf{MonoTM}, an interpretable topic modeling framework that decouples these roles. Across three benchmark corpora, we show that document--topic mixture estimation and semantic interpretation favor different SAE configurations and feature subsets. MonoTM estimates mixtures from the full SAE bag-of-features representation and, with them fixed, learns topic descriptors over a separate vocabulary of corpus-grounded semantic features. This design preserves global topic structure while representing topics with semantic units more meaningful than individual words, making them more useful for downstream corpus analysis.", "url": "https://wpnews.pro/news/beyond-top-words-monotm-for-topic-modeling-with-interpretable-monosemantic", "canonical_source": "https://arxiv.org/abs/2609.09575", "published_at": "2026-09-10 04:00:00+00:00", "updated_at": "2026-09-10 04:19:46.552210+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "ai-research"], "entities": ["MonoTM", "arXiv", "sparse autoencoders"], "alternates": {"html": "https://wpnews.pro/news/beyond-top-words-monotm-for-topic-modeling-with-interpretable-monosemantic", "markdown": "https://wpnews.pro/news/beyond-top-words-monotm-for-topic-modeling-with-interpretable-monosemantic.md", "text": "https://wpnews.pro/news/beyond-top-words-monotm-for-topic-modeling-with-interpretable-monosemantic.txt", "jsonld": "https://wpnews.pro/news/beyond-top-words-monotm-for-topic-modeling-with-interpretable-monosemantic.jsonld"}}