04:00
2026-09-23
arxiv.org
large-language-models
MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
Megagon Labs researchers released MAWILE, a developer-facing workbench for auditing the sensitivity of large language model judges across four surfaces: the judge prompt, judge rubric, target-system iā¦