{"slug": "judgemoe-distributional-aggregation-for-llm-as-a-judge", "title": "JudgeMoE: Distributional Aggregation for LLM-as-a-Judge", "summary": "JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached LLM judge score distributions before fusing them, improved mean Spearman correlation over uniform log pooling by +0.079 on the original 10-cell benchmark, according to the arXiv paper 2610.07109v1. Applied to six additional cells, the same configuration delivered a +0.0393 mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank p=0.0091. The authors' protocol study also found score-range choice unstable across judge-dataset settings and soft scoring usually outperforming hard decoding, with validation-based analyses showing the preferred aggregation method depends on the task and judge pool.", "body_md": "arXiv:2610.07109v1 Announce Type: new \nAbstract: When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by $+0.079$. Applying the same configuration to six additional cells yields a $+0.0393$ mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank $p=0.0091$. Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.", "url": "https://wpnews.pro/news/judgemoe-distributional-aggregation-for-llm-as-a-judge", "canonical_source": "https://arxiv.org/abs/2610.07109", "published_at": "2026-10-07 04:00:00+00:00", "updated_at": "2026-10-07 04:19:44.355688+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "machine-learning"], "entities": ["JudgeMoE", "arXiv", "2610.07109v1"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/judgemoe-distributional-aggregation-for-llm-as-a-judge", "markdown": "https://wpnews.pro/news/judgemoe-distributional-aggregation-for-llm-as-a-judge.md", "text": "https://wpnews.pro/news/judgemoe-distributional-aggregation-for-llm-as-a-judge.txt", "jsonld": "https://wpnews.pro/news/judgemoe-distributional-aggregation-for-llm-as-a-judge.jsonld"}}