{"slug": "my-judging-engine-passed-119-tests-and-had-never-once-finished-its-own-maths", "title": "My judging engine passed 119 tests and had never once finished its own maths", "summary": "A solo hackathon entrant built Third Umpire, a single-container Django and SQLite judging portal that models each review as event average plus project quality plus judge leniency plus noise, fitting quality and leniency together with a ridge penalty and 400 bootstrap refits to produce rank bands. A late code review revealed the fitting loop had never converged: an in-loop re-centring step pulled the review-weighted mean of leniencies to zero while the ridge penalty pulled the unweighted mean to zero, leaving a per-iteration change of about 0.01 that always hit the 1000-iteration cap. Moving centring outside the loop cut iterations to 65 and compute time from 2.72 s to 0.38 s, and changed the 4th and 5th place ordering between Salt Loom and Salt Kiln.", "body_md": "I've done more than fifteen hackathons and won five, always on a team. For DOGFOOD 2026 I entered alone, with Claude Code doing the typing and me making the decisions. Five hours before the deadline, a code review told me the core of my project, the part that decides who wins, had a loop that had never converged. Not once. Every test was green.\n\nThis is what that bug was, why 119 tests missed it, and what it changed about how I work with AI.\n\nDOGFOOD asked us to build a hackathon submission and judging portal. My first reaction was suspicion. Platforms for this already exist, built by teams far bigger than me. So why run a hackathon for it? Did they need the software, were they testing people, or were they comparing our work against something of their own?\n\nSo before writing more code I stopped and read everything again from the ground up. The answer was in their own words: they run about a dozen events a year, and they want something \"boring in the best possible way\" that they can fork and run. Their complaint about existing platforms was specific. Normalization is advertised and never explained, and \"we averaged the scores\" is a weak answer.\n\nThat told me where to spend my time. The portal is ordinary web software. Judging is where it gets hard, because it is a statistics problem that most platforms build as a form.\n\n[Third Umpire](https://github.com/vishnu601/Third-umpire) is one Docker container: Django, SQLite, no other services. The part I cared about is the scoring.\n\nJudges are not interchangeable. Some are harsh, some generous, and in the sample data every judge only scored projects in their own track. A plain average rewards whoever drew kind judges. The usual fix, per-judge z-scores, assumes every judge saw an equally strong batch, which is false when judges are assigned by track.\n\nSo each review is modelled as:\n\n```\nscore = event average + project quality + judge leniency + noise\n```\n\nQuality and leniency are fitted together, using the projects that judges share. Leniency gets a small penalty (a ridge term) so a judge with one review is assumed average until shown otherwise. Then the model is refitted 400 times with resampled noise, which gives every project a rank band and a chance of winning, not a single position.\n\nThe fit alternates: update every judge's leniency, then every project's quality, and repeat until nothing moves. Inside that loop was one extra step that re-centred the leniencies so they average to zero:\n\n```\nfor it in range(1, iters + 1):\n    ...                                   # update each judge's leniency b[j]\n    centre = sum(b[j] * len(rows) for j, rows in by_j.items()) / n\n    for j in b:\n        b[j] -= centre                    # re-centre, weighted by review count\n    ...                                   # update each project's quality q[p]\n    if delta < tol:\n        break\n```\n\nIt looks harmless. It is not. The ridge penalty already pulls the *unweighted* average of the leniencies to zero. The centring step pulls the *review-weighted* average to zero. Those are different targets, so each iteration the two steps undid a little of each other's work. The loop settled into a stand-off where the change per iteration was about 0.01, never below the tolerance. It ran to the cap and returned whatever it had.\n\n|  | Before | After | \n|---|---|---|\n| Iterations on the sample event | 1000 (the cap), every time | 65 | \n| Time to compute results | 2.72 s | 0.38 s | \n| 4th and 5th place |  |  | \n\n| Salt Loom, Salt Kiln | Salt Kiln, Salt Loom |\n\nThe fix was to take centring out of the loop and apply it once at the end. Shifting every leniency down by a constant and every quality up by the same constant leaves the fitted values, and so the ranking, unchanged:\n\n```\nfor it in range(1, iters + 1):\n    ...                                   # leniency, then quality; no centring\n    if delta < tol:\n        break\ncentre = sum(b[j] * len(rows) for j, rows in by_j.items()) / n\nb = {j: v - centre for j, v in b.items()}\nq = {p: v + centre for p, v in q.items()}\n```\n\nThe tests checked that results *looked* right: a harsh judge is corrected, a clear winner wins, the same input gives the same output. A fit that stops 0.01 short of the answer passes all of those. No test asked the one question that mattered: did the loop finish?\n\nThe result object even recorded the iteration count. Nothing read it.\n\nThe new test asserts convergence, and checks the fitted values against the equations the true solution must satisfy. The second check is the one worth copying. It does not compare against numbers I expect. It checks a property the right answer has to have.\n\nThe top three did not move. Fourth and fifth swapped. They were 0.006 apart on a five-point scale.\n\nThat is the real finding. A numerical bug nobody could see was enough to reorder two projects, because their scores were never meaningfully different. A platform that prints \"4th\" and \"5th\" is reporting noise as fact.\n\nOn the sample event the uncertainty is large. The project with the top score wins only 14% of resampled outcomes. The second-placed project wins 22%. Nine projects could plausibly finish in the top three. Third Umpire says so on the results page, \"too close to call\", and suggests which extra reviews would settle it.\n\nI tested the method against data where I knew the true ranking, 300 simulated events shaped like the sample:\n\n| Method | True winner found | \n|---|---|\n| Per-judge z-scores | 26% | \n| Plain average | 34% | \n| Two-way model | 41% | \n\nZ-scores did worse than doing nothing. And the honest limit: my \"80%\" rank bands contained the true rank 73% of the time. They run narrow, and the docs say so.\n\nLate on, I made the audit log tamper-evident. Each entry stores a SHA-256 hash of the entry before it plus its own contents, so an edit or deletion breaks the chain.\n\nThe portal also runs a community vote where tallies are hidden until voting closes, even from organizers. Vote entries are hidden in the audit log for the same reason.\n\nReview found that the hashes undid that. If a hidden vote sits between two visible entries, an organizer knows the hash before it and the hash after it. The only unknowns are who voted, for which project, and the exact microsecond: roughly 10¹¹ guesses, minutes on a GPU. Check each guess against the next entry's hash and the vote is revealed. It leaked through three places: the latest hash on the audit page, the hash column in the CSV export, and the public seal on the results page. Even the entry count leaked when each vote landed.\n\nThe fix: no hash is shown anywhere until voting has closed. Adding integrity had quietly removed privacy, and I would not have seen it by testing the feature I had just built.\n\nIt is surprising how much of an app an AI can build from one prompt. The harder question is whether it actually works. The execution was good, but AI gets things wrong with total confidence, and it will happily write the tests that agree with its own mistake. I was the one deciding which features went in and which did not, and reviewing what came back.\n\nBoth bugs above were written by an AI and caught by other AI reviewers that I pointed at the code with different jobs: logic, security, product, guidelines, style. The author did not find its own bug. A second reader with a narrower question did. That is the same reason companies have code review.\n\nWith five hours left I was sceptical about touching the core maths. But this is how real systems behave: they break, and the question is how you find out. What I took from it is not \"don't trust AI\". It is: don't trust it blindly. Ask why the code is the way it is, and ask how you would know if it were wrong.\n\nTwo things. I'd spend more on the user experience, so an organizer can run an event without needing a tutorial. And I'd go further on the normalization: the rank bands are too narrow, and the model cannot yet spot two judges colluding.\n\nOne thing I would keep: stopping to ask why the hackathon existed. It decided what I built.\n\n`bf73f61`", "url": "https://wpnews.pro/news/my-judging-engine-passed-119-tests-and-had-never-once-finished-its-own-maths", "canonical_source": "https://dev.to/vishnu601/my-judging-engine-passed-119-tests-and-had-never-once-finished-its-own-maths-3f1m", "published_at": "2026-10-05 02:07:45+00:00", "updated_at": "2026-10-05 02:12:22.652456+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "mlops"], "entities": ["Third Umpire", "Claude Code", "DOGFOOD 2026", "Django", "SQLite", "Salt Loom", "Salt Kiln"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/my-judging-engine-passed-119-tests-and-had-never-once-finished-its-own-maths", "markdown": "https://wpnews.pro/news/my-judging-engine-passed-119-tests-and-had-never-once-finished-its-own-maths.md", "text": "https://wpnews.pro/news/my-judging-engine-passed-119-tests-and-had-never-once-finished-its-own-maths.txt", "jsonld": "https://wpnews.pro/news/my-judging-engine-passed-119-tests-and-had-never-once-finished-its-own-maths.jsonld"}}