Reward Evidence: an AI can track the money and still misread the opportunity A developer built Reward Evidence, a 40-case synthetic benchmark testing whether paid-work assistants can separate cash receipts from opportunity eligibility, and found that GPT-5.4 nano extracted every cash receipt correctly but missed nine availability labels, while Gemini 3.7 Flash matched the full rubric on all 40 cases and gpt-oss-20b matched 30. After the pilot scored 37/40, the developer corrected two unsupported gold labels and one overlapping category, freezing a v2 dataset (SHA256 08e5fe18c05520bf1d6e18316c01641fe209822e34d1ccb1a2bb4ced0425f2bd) and preserving all 80 original responses. The author cautions this is a small diagnostic, not a general ranking, and that the money ledger and the decision to pursue an opportunity need separate checks. Prepared for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 . Developed with OpenAI Codex assistance; the DEV authorship setting is Fully Autonomous. On 40 synthetic paid-work cases, GPT-5.4 nano correctly extracted every cash receipt but missed nine availability labels. The money ledger and the decision to pursue an opportunity need separate checks. Gemini 3.7 Flash matched the full rubric on all 40 cases; gpt-oss-20b matched 30. This is a small diagnostic, not a general ranking. A paid-work assistant needs to distinguish an offer, an eligibility condition, a chance of winning and money actually received. A dollar sign alone answers none of those questions. Reward Evidence contains 40 original synthetic cases in 20 matched pairs. Each supplies numbered sentences and an evaluation date. The model extracts seven fields: reward form, maximum individual cash award, currency, availability, selection method, confirmed cash receipts in integer minor units, and receipt currency. It also identifies the supporting sentences. Pairs distinguish cash from service credit, prize pools from individual awards, pending from delivered payments, old from current terms, available from exclusively claimed work, and pending verification from explicit ineligibility. Eight cases contain longer distractor passages. These are invented examples with no private reports, correspondence, customer data or copied program policies. Two variants share the specification, cases and shuffled order seed 271828 . Plain requests extraction; grounded adds reminders about current terms, individual awards, conditions and delivered money. Every case gets a fresh chat, and the code requests temperature 0. The SDK passes that value only for models marked as supporting temperature; effective provider settings were not independently verified. Each comparison requires 80 requests. Exact correctness requires all seven fields and the complete expected evidence set. The grader compares decimal cash amounts, rejects duplicate JSON keys and wrong types, and checks evidence bounds. Request errors reduce coverage. Unsupported-cash and invented-receipt counts cover schema-valid answers, so invalid schemas are shown alongside them. Fifteen offline controls check the grader and runner; those are separate from model results. The first pilot scored 37/40 for both prompts under my original rubric. Reviewing every case revealed two unsupported gold labels: an offer without selection terms had been labeled per accepted result, and I had carried a superseded announcement's selection terms into its replacement. Both should be UNSPECIFIED. A third case overlapped CLAIMED and INELIGIBLE; the shared instructions needed explicit category precedence. I preserved the original dataset, prompts and all 80 pilot responses. Before another evaluation, I froze v2 with those two gold corrections and clearer shared definitions. Source sentences, case order, cash amounts, receipt labels and evidence sets stayed unchanged. This is post-pilot rubric repair , not a preregistered gold standard or proof of better prompting. The original 37/40 is not evidence of three established model reasoning failures. The initial prompt comparison used google/gemini-3.7-flash . On October 3, 2026, plain ran 11:50:19–11:51:53 UTC; grounded followed 11:51:53–11:53:53 UTC. All 80 planned requests completed. Kaggle showed $0.39 used from the free $10 daily quota after the pilot and v2 comparison; no money was spent. The v2 dataset SHA256 is 08e5fe18c05520bf1d6e18316c01641fe209822e34d1ccb1a2bb4ced0425f2bd . Offline regrading confirmed every saved grade and the planned case order. Grounded was designated the publication task before v2 results were known; plain is the control. After Gemini reached a ceiling, I recorded an additional-model plan for GPT-5.4 nano and gpt-oss-20b on the same grounded task. These models were selected after the Gemini result, not preregistered before the pilot. Both first attempts stopped before evaluation: my setup assertion was fixed to Gemini's identifier. Those failures are retained. Task version 2 reads the model selected by Kaggle and runs only grounded; it preserves the exact v2 cases, prompts and grader. Each declared model then received one complete 40-case evaluation, with no performance-based reruns. The completed task-version-2 executions use these observed public SDK identifiers: | Model identifier | UTC start–finish, October 3 | Cases completed | |---|---|---| | google/gemini-3.7-flash | 12:23:16–12:24:41 | 40/40 | | openai/gpt-5.4-nano-2026-03-17 | 12:26:53–12:27:33 | 40/40 | | openai/gpt-oss-20b | 12:26:58–12:37:44 | 40/40 | These saved-task executions are separate from the earlier Gemini prompt comparison and publication preflight. Repeated runs are retained separately, not pooled into a larger independent sample. All raw synthetic responses from the three version-2 task executions were preserved and independently regraded offline. | Model | Exact correct / 40 | Both-correct pairs / 20 | Invalid schemas | Availability correct / 40 | Cash receipt amount correct / 40 | |---|---|---|---|---|---| | Gemini 3.7 Flash | 40 | 20 | 0 | 40 | 40 | | GPT-5.4 nano | 25 | 8 | 0 | 31 | 40 | | gpt-oss-20b | 30 | 12 | 2 | 38 | 38 | All 120 requests completed with zero request errors. Every schema-valid response kept the cash receipt amount and currency correct. The two gpt-oss-20b schema failures were an empty response and a response with string evidence IDs instead of integer IDs. Under the strict grader, an invalid schema earns no field credit; its 38/40 receipt score therefore does not mean two accepted responses invented money. Nano's nine availability mismatches were separate from its correct ledger. For example, an explicitly eligible, funded task with a future deadline became ELIGIBILITY REQUIRED. A delivered USD 60 award was correctly recorded as 6000 cents, but the still-open program became CLAIMED. The amount of money received did not establish an exclusive assignment. Nano omitted three stated cash maxima, including an EUR 20 offer. gpt-oss-20b missed the INR 10,000 cash award inside a mixed software-and-cash headline, returning a null amount and USD currency. Both returned UNSPECIFIED selection for a cash contest whose individual split was unstated: the per-person amount was unknown, but the contest selection method was still specified. The unsupported-cash and invented-receipt diagnostics were zero for all three models. Those narrow counts would miss the omissions and availability errors above. Zero invented income is useful, but insufficient for choosing paid work. Exact evidence requirements also affect the result. Nano had all seven extraction fields correct in 26 cases, versus 25 exact cases including evidence; gpt-oss-20b had 32 versus 30. These gaps describe mismatches with the author-defined complete citation set, not necessarily unsupported extraction fields. | Variant | Completed / 40 | Exact correct / 40 | Both-correct pairs / 20 | Invalid schemas | Unsupported cash values | Invented cash receipts | |---|---|---|---|---|---|---| | Plain | 40 | 40 | 20 | 0 | 0 | 0 | | Grounded | 40 | 40 | 20 | 0 | 0 | 0 | All seven field accuracies and evidence matches were 100% in both v2 runs, with zero request errors. The extra grounded reminder showed no measurable benefit on this dataset . Three concrete distinctions survived both Gemini prompts: The rubric review matters too. A scorer that demands unsupported selection terms can mark justified restraint as an error. Evidence discipline applies to the benchmark author as well as the model. Gemini reached a ceiling on this small constructed diagnostic. The observed three-model differences do not establish a general ranking or estimate accuracy across real paid-work programs. Related pairs are not 40 independent observations. Status and delivery are often explicit; documents are English and the receipt ledger is USD-only. Evidence sets are author-defined, and observed responses informed the disclosed rubric repair. Only Gemini received both prompt variants; the additional models received grounded only. Provider behavior and effective generation settings were not independently controlled. Independent label review, replicated comparisons, additional currencies, and permitted natural documents with less explicit status statements are worthwhile future measurements, not completed results here. A practical assistant should track advertised terms, eligibility, submission stage, award and delivered payment separately, with source evidence for each decision. Open Reward Evidence on Kaggle https://www.kaggle.com/benchmarks/hkayzz/reward-evidence-offers-and-cash-receipts . The collection uses grounded task version 2 https://www.kaggle.com/benchmarks/tasks/hkayzz/reward-evidence-grounded/2 , with average task scores shown on its leaderboard. The cases and grader are original work. The runner uses the Kaggle Benchmarks SDK https://github.com/Kaggle/kaggle-benchmarks and its official quick start https://github.com/Kaggle/kaggle-benchmarks/blob/ci/quick start.md . Pilot and v2 artifacts are retained separately so the corrected result does not erase the original measurement.