04:00
2026-08-17
arxiv.org
large-language-models
When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics
A new arXiv study (2608.13835v1) evaluating 120 topics from CoNTM and DLDA across NYT, DBLP, and arXiv found that traditional coherence metrics show highly variable agreement with human judgments (rhoโฆ