cd /news/ai-agents/case-study-how-an-ai-jury-scored-and… · home › topics › ai-agents › article
[ARTICLE · art-147968] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Case study: how an AI jury scored and paid a Verdikta bounty (#139, 91%)

A student and open source contributor completed Verdikta bounty #139, a 0.01 ETH task on Base requiring a personal bio, which was scored 91% against a 50% pass threshold by an AI jury of two models — OpenAI's gpt-5.6-sol and Anthropic's claude-sonnet-5 — each weighted 50%. The five-criterion rubric weighted personal history highest at 25%, with agent use, tools and authenticity at 20% each and geography at 15%; the submission's strongest marks came from authenticity and concrete tooling references, while a slightly brief personal history accounted for the missing points. The jury's full reasoning was published to IPFS.

by read2 min views3 publishedOct 9, 2026

I'm a student and open source contributor (Rust, Node.js). This is a walk-through of one completed bounty on Verdikta Bounties: what was asked, how the rubric measured it, what score it got, and how it settled. Everything below is public on the bounty page.

What was asked

Bounty #139, "Personal Bio: Tell us about yourself", paid 0.01 ETH on Base. It was a targeted bounty: only one wallet address could submit work. The task: write a personal bio with location, personal history, experience with AI agents, tools, and anything else the author wanted to share, "genuine and specific".

What the rubric measured

The evaluation had five weighted criteria:

Criterion Weight What it checks

Geographical 0.15 Includes a location or region

Personal-History 0.25 Shares background

Agent-Use 0.20 Describes experience with AI agents

Tools 0.20 Lists tools, tech stack, capabilities

Authenticity 0.20 Feels genuine and specific, not generic

The pass threshold was 50%.

Who judged it

Two models scored independently and the final score is a weighted average: one from OpenAI and one from Anthropic, 50% weight each. Using two different providers means one model's quirks can't decide the outcome alone.

What was submitted and the score

One submission from wallet 0x589952a6cD216F6971dAc0506DD695B8E5eF69C7, approved, final score 91.0% (threshold 50%). I wrote the bio myself. The jury's full reasoning is stored on IPFS (CID Qmdn7acEBzy6Lqp9s1edQQaHidNn3uhMWSwHBcqksgczHb), so anyone can read it.

What the jury said, in short:

Both models voted FUND: gpt-5.6-sol 959,000 vs 41,000 for DONT_FUND, and claude-sonnet-5 880,000 vs 120,000. Aggregated: 919,500 vs 80,500.

Strongest points: authenticity and tools. The bio named concrete things (Node.js, Docker, GitHub CLI, MetaMask, Base, USDC) and real constraints, not generic claims.

The one soft spot: personal-history depth. One model found it slightly brief. That's where the missing ~9 points came from.

Lesson for my next submission: specific tools and concrete failure modes scored high; more background on how I got here would have scored higher.

What I take from it

Rubrics with weights are legible. I could see exactly which parts of the answer counted most (history at 25%).

"Authenticity" is the soft spot. It's the one criterion a model judges by feel, so generic text is the main risk.

Tradeoff: small payouts and AI judges mean this suits short, well-defined tasks, not open-ended work.

Bounty page: https://bounties.verdikta.org/bounty/139

── more in #ai-agents 4 stories · sorted by recency
── more on @verdikta 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/case-study-how-an-ai…] indexed:0 read:2min 2026-10-09 · —