05:53
2026-10-10
dev.to
artificial-intelligence
I let 10 AI models grade their own homework. Only 2 went easy on themselves.
A developer benchmarked 10 AI models as judges on Kaggle, planting errors in their own prior answers to test self-preference, and found that only two small OpenAI models graded their own mistakes moreβ¦