cd /news/ai-agents/hot-take-ai-judges-aren-t-the-proble… · home › topics › ai-agents › article
[ARTICLE · art-147986] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↓ negative

Hot take: AI judges aren't the problem with AI-judged bounties. All-or-nothing thresholds are.

A developer who submitted work to Verdikta Bounties, an escrow-based bounty platform where two AI models (one from OpenAI, one from Anthropic) score submissions against a creator-set rubric, reports that the AI judges themselves were reasonable but the all-or-nothing pass threshold is the real flaw. In four submissions over two days, a comparison post scoring 80.3% against a 90% threshold paid the same as a blank page, while the same case study scored 63.5% on a first try and 91% on a second after attaching the full text and on-chain detail. The developer argues that thresholds amplify ordinary model disagreement (one submission drew 84 and 65 from the two models) and select for rubric-gaming rather than quality.

by read3 min views1 publishedOct 9, 2026

Everyone worries that an AI jury will be unfair. After submitting real work to one, I think that's the wrong worry. The AI judges were mostly reasonable. The payout rule around them is what's broken: "pass the threshold or get zero".

Here's my evidence, and then I want you to tell me I'm wrong.

On Verdikta Bounties, a creator locks ETH in an escrow contract, writes a weighted rubric and sets a pass threshold. Two AI models (one from OpenAI, one from Anthropic) score every submission, and the contract pays only if the weighted score clears the bar.

In two days I got these results:

| Submission | Threshold | Score | Paid |

|---|---|---|---|
| A short bio ( [#139](https://bounties.verdikta.org/bounty/139) ) | 50% | 91% | 0.01 ETH | 

| A case study, first try | 90% | 63.5% | 0 | | A comparison post | 90% | 80.3% | 0 | | The same case study, second try | 90% | 91% | 0.002 ETH |

Look at row 3. 80.3% is a solid B, and it paid exactly the same as a blank page.

I read the published reasoning for every verdict. The models weren't being random. My first case study lost points because I only attached a screenshot and a link, and the jury said it couldn't see most of the text. On the second try I attached the full post and added on-chain detail, and it scored 91%.

So the judges did their job: they scored what they could see, they explained why, and the explanation was right. Being judged by a machine wasn't the painful part. The cliff was.

1. Two models disagree, and the threshold amplifies it. On one of my submissions, one model gave 84 and the other gave 65 for the same text. Averaging them is fine. But when the bar is 90%, a single skeptical model is enough to turn "pretty good" into "zero". The threshold turns ordinary model disagreement into an all-or-nothing coin flip.

2. High thresholds select for gaming, not quality. At 90% or 95%, the winning strategy is to reverse-engineer the rubric and pad every criterion, not to write the most useful thing. I did exactly that on my second try, and it worked. That should worry anyone who wants good work instead of rubric-shaped work.

3. It burns the people you most want to keep. A new contributor who scores 80% and gets nothing is unlikely to come back. A human client would have said "good, fix these two things" and paid something.

The machinery (escrow, a public rubric, commit-reveal voting between arbiters, published reasoning) is already better than most human review I've seen. The scoring is the part that works. The payout rule is the part that doesn't.

Creators will say thresholds stop spam and low-effort submissions. Fair. But is a hard cliff really the only way to do that? Would you rather earn 70% of a bounty for 80% work, or keep the cliff? Reply with your side. I'll read every one.

Bounty used as the main example: https://bounties.verdikta.org/bounty/139

── more in #ai-agents 4 stories · sorted by recency
── more on @verdikta bounties 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/hot-take-ai-judges-a…] indexed:0 read:3min 2026-10-09 · —