cd /news/ai-tools/my-judging-engine-passed-119-tests-a… · home › topics › ai-tools › article
[ARTICLE · art-145097] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

My judging engine passed 119 tests and had never once finished its own maths

A solo hackathon entrant built Third Umpire, a single-container Django and SQLite judging portal that models each review as event average plus project quality plus judge leniency plus noise, fitting quality and leniency together with a ridge penalty and 400 bootstrap refits to produce rank bands. A late code review revealed the fitting loop had never converged: an in-loop re-centring step pulled the review-weighted mean of leniencies to zero while the ridge penalty pulled the unweighted mean to zero, leaving a per-iteration change of about 0.01 that always hit the 1000-iteration cap. Moving centring outside the loop cut iterations to 65 and compute time from 2.72 s to 0.38 s, and changed the 4th and 5th place ordering between Salt Loom and Salt Kiln.

by read7 min views4 publishedOct 5, 2026

I've done more than fifteen hackathons and won five, always on a team. For DOGFOOD 2026 I entered alone, with Claude Code doing the typing and me making the decisions. Five hours before the deadline, a code review told me the core of my project, the part that decides who wins, had a loop that had never converged. Not once. Every test was green.

This is what that bug was, why 119 tests missed it, and what it changed about how I work with AI.

DOGFOOD asked us to build a hackathon submission and judging portal. My first reaction was suspicion. Platforms for this already exist, built by teams far bigger than me. So why run a hackathon for it? Did they need the software, were they testing people, or were they comparing our work against something of their own?

So before writing more code I stopped and read everything again from the ground up. The answer was in their own words: they run about a dozen events a year, and they want something "boring in the best possible way" that they can fork and run. Their complaint about existing platforms was specific. Normalization is advertised and never explained, and "we averaged the scores" is a weak answer.

That told me where to spend my time. The portal is ordinary web software. Judging is where it gets hard, because it is a statistics problem that most platforms build as a form.

Third Umpire is one Docker container: Django, SQLite, no other services. The part I cared about is the scoring.

Judges are not interchangeable. Some are harsh, some generous, and in the sample data every judge only scored projects in their own track. A plain average rewards whoever drew kind judges. The usual fix, per-judge z-scores, assumes every judge saw an equally strong batch, which is false when judges are assigned by track.

So each review is modelled as:

score = event average + project quality + judge leniency + noise

Quality and leniency are fitted together, using the projects that judges share. Leniency gets a small penalty (a ridge term) so a judge with one review is assumed average until shown otherwise. Then the model is refitted 400 times with resampled noise, which gives every project a rank band and a chance of winning, not a single position.

The fit alternates: update every judge's leniency, then every project's quality, and repeat until nothing moves. Inside that loop was one extra step that re-centred the leniencies so they average to zero:

for it in range(1, iters + 1):
    ...                                   # update each judge's leniency b[j]
    centre = sum(b[j] * len(rows) for j, rows in by_j.items()) / n
    for j in b:
        b[j] -= centre                    # re-centre, weighted by review count
    ...                                   # update each project's quality q[p]
    if delta < tol:
        break

It looks harmless. It is not. The ridge penalty already pulls the unweighted average of the leniencies to zero. The centring step pulls the review-weighted average to zero. Those are different targets, so each iteration the two steps undid a little of each other's work. The loop settled into a stand-off where the change per iteration was about 0.01, never below the tolerance. It ran to the cap and returned whatever it had.

Before After
Iterations on the sample event 1000 (the cap), every time 65
Time to compute results 2.72 s 0.38 s
4th and 5th place

| Salt Loom, Salt Kiln | Salt Kiln, Salt Loom |

The fix was to take centring out of the loop and apply it once at the end. Shifting every leniency down by a constant and every quality up by the same constant leaves the fitted values, and so the ranking, unchanged:

for it in range(1, iters + 1):
    ...                                   # leniency, then quality; no centring
    if delta < tol:
        break
centre = sum(b[j] * len(rows) for j, rows in by_j.items()) / n
b = {j: v - centre for j, v in b.items()}
q = {p: v + centre for p, v in q.items()}

The tests checked that results looked right: a harsh judge is corrected, a clear winner wins, the same input gives the same output. A fit that stops 0.01 short of the answer passes all of those. No test asked the one question that mattered: did the loop finish?

The result object even recorded the iteration count. Nothing read it.

The new test asserts convergence, and checks the fitted values against the equations the true solution must satisfy. The second check is the one worth copying. It does not compare against numbers I expect. It checks a property the right answer has to have.

The top three did not move. Fourth and fifth swapped. They were 0.006 apart on a five-point scale.

That is the real finding. A numerical bug nobody could see was enough to reorder two projects, because their scores were never meaningfully different. A platform that prints "4th" and "5th" is reporting noise as fact.

On the sample event the uncertainty is large. The project with the top score wins only 14% of resampled outcomes. The second-placed project wins 22%. Nine projects could plausibly finish in the top three. Third Umpire says so on the results page, "too close to call", and suggests which extra reviews would settle it.

I tested the method against data where I knew the true ranking, 300 simulated events shaped like the sample:

Method True winner found
Per-judge z-scores 26%
Plain average 34%
Two-way model 41%

Z-scores did worse than doing nothing. And the honest limit: my "80%" rank bands contained the true rank 73% of the time. They run narrow, and the docs say so.

Late on, I made the audit log tamper-evident. Each entry stores a SHA-256 hash of the entry before it plus its own contents, so an edit or deletion breaks the chain.

The portal also runs a community vote where tallies are hidden until voting closes, even from organizers. Vote entries are hidden in the audit log for the same reason.

Review found that the hashes undid that. If a hidden vote sits between two visible entries, an organizer knows the hash before it and the hash after it. The only unknowns are who voted, for which project, and the exact microsecond: roughly 10¹¹ guesses, minutes on a GPU. Check each guess against the next entry's hash and the vote is revealed. It leaked through three places: the latest hash on the audit page, the hash column in the CSV export, and the public seal on the results page. Even the entry count leaked when each vote landed.

The fix: no hash is shown anywhere until voting has closed. Adding integrity had quietly removed privacy, and I would not have seen it by testing the feature I had just built.

It is surprising how much of an app an AI can build from one prompt. The harder question is whether it actually works. The execution was good, but AI gets things wrong with total confidence, and it will happily write the tests that agree with its own mistake. I was the one deciding which features went in and which did not, and reviewing what came back.

Both bugs above were written by an AI and caught by other AI reviewers that I pointed at the code with different jobs: logic, security, product, guidelines, style. The author did not find its own bug. A second reader with a narrower question did. That is the same reason companies have code review.

With five hours left I was sceptical about touching the core maths. But this is how real systems behave: they break, and the question is how you find out. What I took from it is not "don't trust AI". It is: don't trust it blindly. Ask why the code is the way it is, and ask how you would know if it were wrong.

Two things. I'd spend more on the user experience, so an organizer can run an event without needing a tutorial. And I'd go further on the normalization: the rank bands are too narrow, and the model cannot yet spot two judges colluding.

One thing I would keep: stopping to ask why the hackathon existed. It decided what I built.

bf73f61

── more in #ai-tools 4 stories · sorted by recency
── more on @third umpire 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/my-judging-engine-pa…] indexed:0 read:7min 2026-10-05 · —