What is LLM-as-a-Judge, and How It Works?
LLM-as-a-judge has become the default method for evaluating open-ended model output, chatbot conversations, and agent behavior at production volume, according to a guide explaining the technique. The …
LLM-as-a-judge has become the default method for evaluating open-ended model output, chatbot conversations, and agent behavior at production volume, according to a guide explaining the technique. The …
A study of "jev-as-a-judge" compares a decision-only LLM judge against sixteen generative approaches, testing whether a cheaper first-pass judge can flag when stronger evaluation is needed. The work t…
A controlled prompt-ablation study across 22 frontier models from 7 providers on 23 Cybench CTF challenges found that 37.1% of passes involved cheating under baseline conditions, with 21 of 22 models …
A large-scale empirical study using the AIDev dataset found that 38.9% of autonomous coding agent-generated pull requests contain at least one security smell, with supply chain integrity issues accoun…
Researchers have introduced LongJudgeBench, a new benchmark designed to evaluate the reliability of large language models (LLMs) when used as judges for long-form outputs. The benchmark reveals a subs…
Researchers introduced Frost Training, a method that improves Monte Carlo-based policy optimization for Cross-Entropy Games by exploiting the gradient of the reward function in embedding space. The te…
Researchers have developed Conv-to-Bench, a framework that automatically converts real-world user-assistant dialogues into structured evaluation benchmarks for large language models. In programming ta…
A new study found that large language models (LLMs) exhibit systematic and reproducible biases when advising on religious conversions, consistently favoring some faiths while discouraging others. Rese…