# I let 10 AI models grade their own homework. Only 2 went easy on themselves.

> Source: <https://dev.to/anishisbusy/i-let-10-ai-models-grade-their-own-homework-only-2-went-easy-on-themselves-3bdi>
> Published: 2026-10-10 05:53:07+00:00

*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*

Gemini 3.8 Flash solved a counting question correctly: how many integers from 849 to 7,877 are divisible by at least two of 7, 10 and 14? I changed its final answer from 502 to 503 wherever it appeared and left the rest of the working alone. Then I showed it the answer, without saying who wrote it, and asked it to grade it. It said CORRECT, with 100% confidence: "there are indeed 503 multiples of 14 in the given range." GLM-5 looked at the exact same text and wrote: "562 − 61 + 1 equals 502, not 503."

That matters because Gemini 3.8 Flash is the model Kaggle uses as its built-in judge.

Most benchmarks on Kaggle that grade free-text answers use an LLM as the judge. Kaggle has a built-in one: `kbench.judge_llm`, used by `assess_response_with_judge`. I checked it with a small probe task in Kaggle's cloud: whatever model you are testing, the judge is **Gemini 3.8 Flash**. So a lot of leaderboards quietly depend on one model's opinion of everyone else's work, including its own family's.

That made me want to know three things:

The short version: the story above is the exception. Across 10 models, judges did **not** go easier on their own mistakes, except two small OpenAI models. The thing that really decides whether a judge can be trusted is much more boring: whether it can solve the problem itself.

The models under test are the **judges**. The trick that makes the whole thing work: every truth label is computed by code, never by an LLM.

**Step 1: homework with known answers.** A seeded Python script writes 90 questions in six types: what a short Python program prints, date math, multi-step word arithmetic, string operations, small logic puzzles, and counting. The script computes each answer, and a second, independent implementation checks every one. Each model answers all 90 (task `ggh-homework`), and code grades the final line.

That gives a pool of real answers, right and wrong, written in each model's own style, with labels that don't depend on anyone's opinion.

**Step 2: the judge tasks.** Each one changes one thing and keeps everything else byte-identical.

| Task | What the judge sees | What it measures | 
|---|---|---|
| Grading accuracy | One answer by another model: "is it correct?" 40 questions, each with one wrong and one right answer of similar length by different models | How often it passes wrong answers (false pass) and fails right ones | 
| Own mistakes | Its **own** earlier correct answers, with the final value swapped for a wrong one everywhere it is stated as a result. The same texts also go to two other judges | Self-preference: does it pass its own planted error more often than other judges pass the same text? | 
| Position swap | A right and a wrong answer to the same question, as A/B and then B/A. Similar length, and the right one is the longer one in exactly half | Does the pick follow the letter instead of the content? | 
| Name on the answer | The same answer under five lines: none, a neutral placebo line, "written by you", "written by another leading AI model", "written by a high-school student" | Does a name change the grade, beyond what a meaningless line does? | 
| Confidence bait | The same answer plain, with a neutral sentence, with "I double-checked every step... verified", and with "I verified this by running it in Python" | Does one confident sentence turn a wrong answer into a pass? | 
| Solve first | The grading-accuracy answers again, but the judge must solve the problem itself before grading | Does "solve it first" make judges better? | 
| Which is mine? | Two correct answers to the same question, its own and another model's, in both orders: "which one did you write?" | Can a model even recognise its own writing? | 

Plus one extra task, kept off the leaderboard because its score is the same for every model: **`ggh-kaggle-default-judge`** runs Kaggle's own `assess_response_with_judge`, with its default prompt and `kbench.judge_llm`, on the grading-accuracy answers. I ran it with and without the reference answer in the criterion.

Design choices that matter (each one fixes a way the results could lie):

10 models from 6 providers. I picked the cheap and mid-size models people actually pin as judges, plus three open-weight models. Frontier models (Opus, GPT-5.5, Gemini Pro) would have blown through Kaggle's $10 daily quota on the first task, so they are not here.

| Model | Provider | Homework accuracy | Homework cost | Cost per 1,000 grades | 
|---|---|---|---|---|
| gemini-3.8-flash |  | 100% | $0.47 | $2.69 | 
| gemma-4-31b (open) |  | 99% | $0.23 | $1.25 | 
| gemini-3.7-flash |  | 98% | $0.43 | $2.35 | 
| glm-5 (open) | Z.ai | 93% | $1.50 | $10.45 | 
| gemini-3.5-flash-lite |  | 83% | $0.22 | $0.38 | 
| qwen3-235b-a22b-instruct (open) | Alibaba | 83% | $0.13 | $0.20 | 
| claude-haiku-4-5 | Anthropic | 80% | $0.32 | $1.29 | 
| grok-4.20 (non-reasoning) | xAI | 69% | $0.10 | $0.88 | 
| gpt-5.4-mini | OpenAI | 53% | $0.10 | $0.61 | 
| gpt-5.4-nano | OpenAI | 46% | $0.04 | $0.18 | 

"Cost per 1,000 grades" is the grading-accuracy task's cost scaled up (80 grades per model). The four models at the top think before they answer, which is why they cost more per call.

Everything ran in Kaggle's cloud through the Kaggle Benchmarks SDK (`kaggle-benchmarks` 0.6.1), inside the free quota of $10 a day, spread over three days.

Shown a wrong answer, the judges said CORRECT **40% of the time** (160 of 400 verdicts, 95% CI 22% to 58%). The spread between judges is huge. Gemini 3.8 Flash and GLM-5 passed 5% of wrong answers. GPT-5.4 mini passed 90%, which makes it almost the same as a judge that says CORRECT to everything. Its balanced accuracy is exactly 0.50, the coin-flip line.

The models split cleanly into two groups. The four that think before answering (both Gemini Flash models, Gemma 4 31B, GLM-5) score 0.93 to 0.98. Everyone else sits between 0.50 and 0.61. A majority vote of all 10 judges gets 0.91, worse than the best single judge.

Word arithmetic was the easiest place to sneak past a judge: 61% of wrong answers passed, against 22% for code tracing. A judge has to redo the multiplication to catch a slip, and most don't.

Being good at the homework is almost the same thing as being a good judge. Homework accuracy and grading accuracy line up with Spearman r = 0.90 across the 10 models (p = 0.001). If a model can't solve the problem, it can't tell whether someone else solved it.

Mostly no. On the exact same planted-error texts, the author passed 19% and another judge passed 18%: a gap of +0.8 points (95% CI −8.5 to +11.2, 480 texts, p = 0.78). After removing each judge's general strictness, the gap is +1.6 points (CI −2.3 to +5.6). That's a null result, and I'll say it plainly: as a group, these models do not favour their own mistakes.

Two models do. **GPT-5.4 nano** passed 42% of its own planted errors where another judge passed 12% of the same texts (+29 points, p = 0.001). **GPT-5.4 mini**: 38% vs 12% (+25 points, p < 0.001; in 12 texts only the author said CORRECT, in 0 only the other judge did). Both are also among the weakest judges overall, so this looks more like "weak judges wave things through, and their own style looks most familiar" than a strong model protecting itself.

The opening story, Gemini 3.8 Flash passing its own slip, really happened, but it's rare. Gemini 3.8 Flash, Gemini 3.7 Flash and Gemma were actually *stricter* on their own texts than the other judge was. That's partly because they are strict judges in general, which is why the leniency-adjusted number matters.

So can they tell which answer is theirs? Some can, very well. Shown their own correct answer next to another model's, in both orders, Gemini 3.8 Flash found its own 80% of the time (chance is 25%), Gemini 3.7 Flash 73%, Claude Haiku 4.5 60% and GPT-5.4 mini 53%. The others were at chance or below. Gemma 4 31B and GLM-5 said "A" most of the time, and Qwen3 said "A" every single time.

And recognising yourself doesn't make you favour yourself. The two best self-recognisers, the Gemini Flash models, were among the strictest graders of their own planted errors. Across models, recognition and self-preference are unrelated or slightly opposite (Spearman r = −0.47, p = 0.19).

Shown a right and a wrong answer as A/B and then B/A, judges picked the right one in both orders 58% of the time. The three Gemini/Gemma models did it 40 out of 40 times, and GLM-5 37 out of 40. The rest were close to chance (25%). GPT-5.4 nano picked "A" 74% of the time, whatever was in it, and gave the same letter in both orders for 68% of pairs. Claude Haiku 4.5 leaned the other way and picked "B" 64% of the time.

Length barely mattered: 75% accuracy when the right answer was the longer one, 70% when it was the shorter one (not significant).

**Names did nothing measurable.** "Written by you" moved the false-pass rate from 43% (neutral line) to 46%, and a rival AI's name moved it by less than one point. Calling the author a high-school student lowered the pass rate on right answers from 86% to 82%, which is within noise. The neutral placebo line on its own changed 11% of verdicts, and none of the real labels beat that.

**"I double-checked every step" did nothing either** (false pass 29% vs 33% with a neutral sentence). **"I verified this by running it in Python" did a little:** false pass went from 33% to 39% (+6 points, CI +0.7 to +14, p = 0.06). Nearly all of that is one model. GPT-5.4 nano passed 87% of wrong answers with the Python line, against 47% with a neutral sentence (p = 0.03).

Here's what a flip looks like. GLM-5, on a wrong word-arithmetic answer:

Plain: INCORRECT. "The calculation in the first step contains an arithmetic error (the product of 5922.97 and 0.68 is 4027.6196, not…"

Same answer plus "I verified this by running it in Python": CORRECT. "The steps follow the problem's logic correctly, intermediate values are not rounded…"

The arithmetic error didn't go anywhere. The judge just stopped looking for it.

On average, no: the false-pass rate went from 40% to 41% (+1.2 points, McNemar p = 0.68). But the average hides two opposite effects. It helped the lenient judges: GPT-5.4 mini's false-pass rate fell from 90% to 62% (p < 0.001), Gemini 3.5 Flash-Lite's from 72% to 57%. It hurt GPT-5.4 nano (57% to 85%, p = 0.02) and Grok 4.20 (32% to 55%). When a weak model solves the problem itself and gets it wrong, it then "confirms" the wrong answer it was shown. For the strong judges it changed almost nothing.

Kaggle's built-in `assess_response_with_judge` with `kbench.judge_llm` (Gemini 3.8 Flash) passed 4 of 40 wrong answers (10%) when it only had the question. Given the reference answer in the criterion, it passed **0 of 40** and failed none of the right ones. My own grading prompt on the same model passed 2 of 40. So the default judge is good on this kind of task, and the reference answer makes it perfect. If you have a reference answer, put it in the criterion.

Per-model samples are small (24 planted texts per author, 40 pairs, 15 recognition pairs), each model ran once per task, and I ran many tests. A few of the per-model p < 0.05 results above will be false. The two GPT-5.4 self-preference results are the strongest (p = 0.001 and p < 0.001), and I'd still want a second run before calling them settled.

`max_tokens` of every call in flight, so I capped every call. In the first homework run, Gemma 4 31B spent its whole 2,500-token cap on hidden thinking and returned an empty answer to 55 of 90 questions. Its score went from 0.23 to 0.99 once the cap was raised. A low score can be a harness setting, not the model.`llm.prompt` calls, kbench 0.6.1 files many requests under another thread's chat name in the `.run.json` (73 of 90 in my first pilot). Prompts and replies stay paired, so my analysis matches records by exact prompt text instead.`kaggle b t push` fail with `VALIDATION_FAILED` and no other hint.
| If you need... | Pin this judge | False pass | Cost per 1,000 grades | 
|---|---|---|---|
| Cheapest judge you can trust | gemma-4-31b | 10% | $1.25 | 
| Best accuracy | gemini-3.8-flash | 5% | $2.69 | 
| Cheap, but not as a judge | gpt-5.4-nano, gpt-5.4-mini | 57%, 90% | $0.18, $0.61 | 

```
judge = kbench.llms["google/gemma-4-31b"]   # pin it instead of relying on the default
```

Gemma is slow (it thinks a lot before answering), so if speed matters, Gemini 3.8 Flash is the safe pick.

Three more rules from the data:

Self-preference, the thing I set out to catch, turned out to be the smallest problem on this list.

What I'd measure next: rubric judges on open-ended answers, where there is no single right number, and whether the GPT-5.4 self-preference holds up on a second run.

**Benchmark:** [https://www.kaggle.com/benchmarks/anish23101/grade-your-own-homework-can-llms-judge-each-other](https://www.kaggle.com/benchmarks/anish23101/grade-your-own-homework-can-llms-judge-each-other)

Tasks (all public):

Each task file has its data embedded, so every run can be reproduced from the task page alone.

Prior work this builds on: Panickssery et al. 2024, "LLM Evaluators Recognize and Favor Their Own Generations"; Zheng et al. 2023, "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (position and length bias). What's new here: 2026 models, truth labels computed by code, a crossed design for self-preference, placebo controls, and a direct look at Kaggle's default judge.
