TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation A new arXiv paper (2608.02975v1) introduces TQLite, a distillation framework that uses a multi-LRM jury to train small language models (SLMs) for MQM-based translation quality evaluation, achieving performance far exceeding off-the-shelf SLMs and offering a scalable, cost-effective alternative to large language models (LLMs) and large reasoning models (LRMs). The study benchmarks SLMs, LLMs, and LRMs across various evaluation setups, establishing best practices. arXiv:2608.02975v1 Announce Type: new Abstract: Large language models LLMs have demonstrated impressive performance in MQM-based translation quality TQ evaluation, and recent advances in large reasoning models LRMs promise even greater improvements. However, both LLMs and LRMs are computationally expensive to deploy at scale, while small language models SLMs ---though much more efficient---struggle with the complex reasoning required for evaluation tasks. In this work, we present an extensive empirical study benchmarking SLMs, LLMs, and LRMs across a wide range of TQ evaluation setups, providing a comprehensive view of the current landscape and establishing best practices. To address the scalability challenge, we introduce TQLite, a novel distillation framework that enables SLMs to approach the MQM evaluation performance of the best LRM-based evaluators. Our approach leverages a multi-LRM jury to generate high-quality synthetic training data via practical data curation techniques and aggregation of evaluation responses across a diverse panel of models. Our results demonstrate that SLMs trained via TQLite achieve strong MQM evaluation performance that far exceeds off-the-shelf evaluation capabilities of standard SLMs, offering a scalable and cost-effective alternative to LLM- and LRM-based evaluators.