2% bias, 98% noise: what I built, and cut, for a hackathon judging engine A solo developer built a hackathon judging engine for the Hackathon Raptors Dogfood 2026 event, using Claude Code agents to write the code while making the design calls personally. Analyzing 41 projects, 30 judges and 126 reviews, the engine found judge generosity accounted for only about 2% of scoring variance, so it shrinks leniency corrections to at most 0.06 points on a 1-5 scale and instead reports whether each track's top ranking is too close to call. The developer cut three planned features — win-probability percentages, engine-chosen pairwise matchups, and disqualification handling — after tests showed they would invent favorites, miss their accuracy bar, or ship untested. The sample event Hackathon Raptors gave us for Dogfood 2026 has 41 projects, 30 judges and 126 finished reviews. When my judging engine split the judges' disagreement into "this judge is just more generous" and everything else, generosity came out at about 2%. That number ended up deciding most of what I built and most of what I cut. Dogfood's brief was the same for every team: in 72 hours, build the submission and judging platform Raptors will run their own events on. I entered solo. A fleet of AI coding agents Claude Code did the building, and I made the calls. Everything below is in the repo, with the tests behind it. The model is simple. Every review is the project's real level, plus the judge's leniency, plus noise. Noise means everything a correction can't fix, like one judge loving a project another found dull. Leniency is shrunk towards zero until the data supports it, and how much data that takes is re-estimated from the event's own scores on every run. On this event a judge needs about 43 reviews before the engine trusts half of their apparent tilt. Each judge here saw about four projects. So no judge's correction ends up bigger than 0.06 points on a 1 to 5 scale. That looks like the engine doing nothing. On this data, that is the finding. Judges with no tilt at all, scoring the same pairs with this event's noise, still land about 0.37 apart by luck. The real judges are 0.42 apart. A test plants a real tilt, three judges made 0.6 more generous: the spread rises to 0.52 and the engine brings it back to 0.38, so it does correct a tilt that is really there. A permutation test on the same scores finds no project differences beyond chance: in 1,529 of 2,000 shuffles, the shuffled scores spread at least as far as the real ones. So on this event the engine's job is to bound what bias could have done, and disclose it. The usual fix, per-judge z-scores, treats all of a judge's spread as bias. In a planning simulation before kickoff planning was allowed, code wasn't , a shrunk z-score read the harshest judge, who happened to hold the ten strongest projects, as lenient: +0.38 against a true -0.82. With about four reviews per judge, "harsh judge with strong projects" and "lenient judge with average projects" produce the same scores. One judge gave all three of their projects a 4 on every criterion. Leaving that judge out moves 19 of 41 projects, one of them Small Relay from 14th to 30th. So the engine does leave a judge like that out, but never quietly. In simulation the same rule also flags an honest judge in 6 to 8% of panels, so it's a visible flag with its reason, and the organizer can undo it. The first was a chance of winning. It's the obvious feature: "Project X: 63% chance of first place." On simulated events where no project was better than any other, the method named a 50%+ favourite in 46 of 80 tracks. A percentage on a results page reads as fact, and here it would have been inventing favourites. The portal shows no chance of being first anywhere, on any page or in any API answer. The second was letting the engine pick each pairwise question. In pairwise mode a judge just says which of two projects is better. Having the engine choose the next pair from everyone's answers so far sounded smart. The plan fixed the bar before the test ran: it had to pick the right winner at least 3 points more often than the simpler method. The best version managed 1.4, so I said skip it, and it was never built. The third was disqualification. The agents built it on the last day. A review found that reinstating a disqualified project would throw away later pairwise answers, and the feature had never run in a clean container, so it was dropped about two hours before the deadline. Shipping an untested feature in the part that decides who wins would have been the wrong trade. Instead, each track gets a yes or no answer to one question: is the top too close to call? The engine redraws the ranking 4,000 times from its own uncertainty. If the leader comes first in fewer than 3,800 of them, the track is too close to call, and the judges decide it. An exact tie always is. On simulated events with no real differences, it named a winner in 0.13% of 4,000 track-runs. When it did name one, it was the true best 99.2% of the time. The price: where projects really do differ, it names a winner in only 20 to 63% of runs. That's a lot of honest "too close to call". Every score carries a ±, and the ± has to be honest: when the engine says it's at least 95% sure, it should be right at least 95% of the time. The first check pooled all of those claims and counted. A deliberately broken engine, with its ± cut in half, passed: right 98.4% of the time. Easy calls dominate the pool and carry the average. Checking only the claims between 95% and 99% sure dropped the broken engine to 89.7% and failed it, while the honest engine scored 98.6%. The next surprise went the other way. A fixed floor for the worst single run failed the honest engine itself, in one sealed scenario. The floor became "the current engine's own worst run, minus 0.02". After that, every check had to fail on a known-bad input before its pass counted for anything. One Claude Code session Opus 5.5 planned and merged; cheaper models did grind work. Every feature was its own lane in its own git worktree, and nothing reached main until the organizers' checker 7/7 and our own hand checks 35/35 passed in a clean container. On the last day I raised the cap to 16 agents at once. The final suite had 1,670 passing tests. What went wrong: three permission prompts reached my screen within seven minutes during the night run. I was lucky I wasn't asleep yet, since an unanswered prompt stalls the run. So before going to sleep I set up an alarm: a watchdog outside the session, checking every five minutes, that goes off if the run goes quiet for 20 minutes, a tool call hangs for 12, or an agent waits on a prompt for 3. At 4:30 am it actually woke me up. An agent had been stuck for 16 minutes, waiting for permission to delete its own scratch files. The night-run rule since then: nothing gets deleted, the mess waits for the morning. Other things that went wrong: an unquoted heredoc ran a stray docker compose down -v mid-run. A merge silently undid two fixes, so the rule became: rebase onto main at the end, not only at the start. And the organizers' checker makes one 10-second request per check, with no retry. On my loaded laptop the first gallery render took longer, and tier 1 failed once. The portal now warms its own gallery before it says it's ready. Two judges trading favourable scores look like ordinary disagreement to the model; nothing flags that. The audit log's database triggers stop the application, not someone holding the database file, so the chain is only tamper-evident against a hash kept outside the portal. And a judge who strategically calls everything too close to call is caught in only 14 of 120 simulated panels. Code, tests and the full maths: github.com/LippInc/dogfood-2026 JUDGING.md has every number above Live demo: dogfood-portal-demo.onrender.com free instance, give it a minute to wake This is my Write Up Quest entry for Dogfood 2026, run by Hackathon Raptors @partnerships raptors https://dev.to/partnerships raptors .