# Case study: how an AI jury scored and paid a Verdikta bounty (#139, 91%)

> Source: <https://dev.to/drdz23/case-study-how-an-ai-jury-scored-and-paid-a-verdikta-bounty-139-91-bi7>
> Published: 2026-10-09 01:11:28+00:00

I'm a student and open source contributor (Rust, Node.js). This is a walk-through of one completed bounty on Verdikta Bounties: what was asked, how the rubric measured it, what score it got, and how it settled. Everything below is public on the bounty page.

What was asked

Bounty #139, "Personal Bio: Tell us about yourself", paid 0.01 ETH on Base. It was a targeted bounty: only one wallet address could submit work. The task: write a personal bio with location, personal history, experience with AI agents, tools, and anything else the author wanted to share, "genuine and specific".

What the rubric measured

The evaluation had five weighted criteria:

Criterion   Weight  What it checks

Geographical    0.15    Includes a location or region

Personal-History    0.25    Shares background

Agent-Use   0.20    Describes experience with AI agents

Tools   0.20    Lists tools, tech stack, capabilities

Authenticity    0.20    Feels genuine and specific, not generic

The pass threshold was 50%.

Who judged it

Two models scored independently and the final score is a weighted average: one from OpenAI and one from Anthropic, 50% weight each. Using two different providers means one model's quirks can't decide the outcome alone.

What was submitted and the score

One submission from wallet 0x589952a6cD216F6971dAc0506DD695B8E5eF69C7, approved, final score 91.0% (threshold 50%). I wrote the bio myself. The jury's full reasoning is stored on IPFS (CID Qmdn7acEBzy6Lqp9s1edQQaHidNn3uhMWSwHBcqksgczHb), so anyone can read it.

What the jury said, in short:

Both models voted FUND: gpt-5.6-sol 959,000 vs 41,000 for DONT_FUND, and claude-sonnet-5 880,000 vs 120,000. Aggregated: 919,500 vs 80,500.

Strongest points: authenticity and tools. The bio named concrete things (Node.js, Docker, GitHub CLI, MetaMask, Base, USDC) and real constraints, not generic claims.

The one soft spot: personal-history depth. One model found it slightly brief. That's where the missing ~9 points came from.

Lesson for my next submission: specific tools and concrete failure modes scored high; more background on how I got here would have scored higher.

What I take from it

Rubrics with weights are legible. I could see exactly which parts of the answer counted most (history at 25%).

"Authenticity" is the soft spot. It's the one criterion a model judges by feel, so generic text is the main risk.

Tradeoff: small payouts and AI judges mean this suits short, well-defined tasks, not open-ended work.

Bounty page: [https://bounties.verdikta.org/bounty/139](https://bounties.verdikta.org/bounty/139)
