03:35
2026-08-27
vibeleaderboard.ai
artificial-intelligence
Four papers land today against LLM-as-judge, and for executable checks
Four papers published today challenge the reliability of LLM-as-judge, showing that judge annotations can mark working code as failures, leak across rubric dimensions, and misalign with human creativiβ¦