cd/entity/LLM-as-a-judge· home› entities› LLM-as-a-judge
grep -l @llm-as-a-judge /news/*.json | wc -l → 8

LLM-as-a-judge

mentions 8 type Organization feed RSS

// recent coverage 8 mentions

07:19
2026-09-25
infere.com
large-language-models

What is LLM-as-a-Judge, and How It Works?

LLM-as-a-judge has become the default method for evaluating open-ended model output, chatbot conversations, and agent behavior at production volume, according to a guide explaining the technique. The …

12:00
2026-09-23
aiflash.com
large-language-models

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

A study of "jev-as-a-judge" compares a decision-only LLM judge against sixteen generative approaches, testing whether a cheaper first-pass judge can flag when stronger evaluation is needed. The work t…

04:48
2026-06-03
arxiv.org
large-language-models

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

Researchers have introduced LongJudgeBench, a new benchmark designed to evaluate the reliability of large language models (LLMs) when used as judges for long-form outputs. The benchmark reveals a subs…

04:00
2026-05-28
arxiv.org
artificial-intelligence

Cross-Entropy Games and Frost Training

Researchers introduced Frost Training, a method that improves Monte Carlo-based policy optimization for Cross-Entropy Games by exploiting the gradient of the reward function in embedding space. The te…

// co-occurs with top 8 entities
// topics top 6 topics