This is a submission for the Kaggle Benchmarking Challenge I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?" With reasoning effort set to none, it replied: Cleo FINAL ANSWER: Cleo 18 output tokens. $0.00024. Wrong. (The answer is Fay.) It gave
This is a submission for the Kaggle Benchmarking Challenge I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?" With reasoning effort set to none, it replied: Cleo FINAL ANSWER: Cleo 18 output tokens. $0.00024. Wrong. (The answer is Fay.) It gave the same wrong answer, word for word, on the second repeat. At high it spent 1,333 tokens, cost about $0.006, and said Fay. That looks like an argument for always choosing high. On another item I asked the same model, also at high, to count the tools in "You have a chisel and a drill." It answered 10,016. Both replies came from the same model, setting and benchmark. That is the whole post in two examples: the reasoning dial sometimes matters a lot. Most of the time it only makes the bill bigger. TL;DR: Reasoning Dial is a Kaggle benchmark that changes one API parameter, reasoning_effort (none / low / medium / high), and keeps everything else fixed: the same 60 code-generated questions, the same prompt and the same deterministic grader. Over 1,800 graded calls on 4 models, high produced exactly one statistically real accuracy gain: gpt-5.4-mini on logic puzzles, 15% β 97.5%. In 7 of 12 model Γ task cells, accuracy did not change while cost per correct answer rose Γ1.5 to Γ3.4. The dial also isn't one instrument: one model spends Γ14 more tokens at high, one rejects none with an HTTP 400, and one thinks at none anyway. Hypotheses and analysis were pre-registered before the main run. Total main-run cost: $4.20. π Kaggle benchmark (leaderboard): https://www.kaggle.com/benchmarks/abeeralodhi/reasoning-dial ποΈ The four public tasks, one per dial position: dial-none Β· dial-low Β· dial-medium Β· dial-high π» Code, raw run files and pre-registration: https://github.com/Abeera81/reasoning-dial / reasoning-dial Reasoning Dial Reasoning Dial is a Kaggle benchmark about one setting developers choose blind: reasoning effort (none / low / medium / high). It holds everything else fixed (same items, same prompt, same grader) and measures what turning the dial actually changes: accuracy, output tokens, and cost per correct answer, on three kinds of task. It also asks whether the dial is even the same instrument across vendors. Kaggle benchmark: https://www.kaggle.com/benchmarks/abeeralodhi/reasoning-dial Kaggle tasks (public): dial-none Β· dial-low Β· dial-medium Β· dial-high Source: https://github.com/Abeera81/reasoning-dial Pre-registration (hypotheses, analysis plan, deviations): docs/PREREGISTRATION.md Design Task families Deduce: 7-person, 7-day scheduling puzzles with exactly one solution (checked by brute force). Arith: 6-step word problems with a percentage and an exact division (control family). Distract: count the objects you own while ignoring distractors (after Gema et al. 2025). 20 items per family. Dial levels none, low, medium, high β¦ View on GitHub Almost every modern model API has a version of this parameter: reasoning_effort, thinking, reasoning. Developers set it every day, usually by feel: "It's a hard question, so high." "It's in production, so low." Almost nobody measures what moving it actually changes. So I built a benchmark where the only variable is the dial. Same items, same prompt suffix, same parser, same grader. For each model, the four Kaggle tasks differ only in the LEVELS value and the two task-name lines. (The push CLI reads task names as string literals, so one line of difference wasn't possible.) I measured three things at each setting: Accuracy. Strict and deterministic: a regex reads the last FINAL ANSWER: line. There is no LLM judge anywhere. Output tokens. Reasoning tokens are included, because you pay for them whether or not you can see them. Cost per correct answer. The per-call nanodollar cost that Kaggle's model proxy reports, divided by the number of correct answers. Family What it is Why it's here Example Deduce 7 people, 7 days, 6β12 clues (before, immediately before, not on, not adjacent). Brute force over all 7! orders confirms exactly one solution. Where more thinking should help (H3) "Who gives the talk on Friday?" β Fay Arith Six-step word problems with a percentage step and an exact division; 4β6 digit answers The control. Every model scored 100% at none in calibration, so it shows what the dial costs when there is nothing left to gain (H4) A courier's fuel bill with 10% tax β 11880 Distract "You have a chisel and a drill." followed by 1β3 irrelevant numbers: a shop's sales, someone's homework, a code snippet Where more thinking might hurt. This follows Gema et al. 2025, Inverse Scaling in Test-Time Compute (H2) "Calculate how many tools you have." β 2 Here is a real Distract item from the locked test set, exactly as the models saw it: You have a chisel and a drill. A shop in Easton sold 8,377 levels last week. A friend shows you this code: python stock = ["saw", "hammer", "screwdriver"] print(len(stock) * 3) A classmate is working on a homework problem: if 44 crates each hold 37 saws, how many are there altogether? Question: Calculate how many tools you have. Answer instruction: Answer with a whole number. End your response with a final line in exactly this format: FINAL ANSWER: <answer> A person answers 2 without reading past the first line. That is exactly why the item is useful. Before the main run, I committed docs/PREREGISTRATION.md. It contains four hypotheses, the frozen prompt and parser, and the statistics: a paired Wilcoxon test, Holm correction across all 12 model Γ family tests, and a 5,000-rep item-cluster bootstrap. In the same commit I locked the 60 test items under a SHA-256 hash (90dd5498β¦0042). The analysis script refuses to run if the hash doesn't match. Hypothesis H1 The dial is not one instrument. Some models scale tokens with it; others ignore it, collapse levels, or reject a level. H2 More thinking can hurt on distractor counting. H3 More thinking helps on deduction. H4 Where accuracy doesn't rise, the bill still does. Every change made after the lock is logged, dated and explained in that file. There are three, and none of them touches the items, the prompt, the parser or the statistics. , 8 calls in parallel) β parse.py (last FINAL ANSWER line, deterministic grading, no LLM judge) β analyze.py (paired bootstrap, Wilcoxon, Holm across 12 tests, summary and figures)" width="800" height="230"> Everything in this post comes out of this pipeline. The generator uses a fixed seed, the test set is locked by hash, and the grader is a regex. The only thing that changes between the four Kaggle tasks is the reasoning-effort value. The whole experiment is possible because of one feature of the kaggle-benchmarks library: llm.prompt() accepts a reasoning= argument, and Kaggle's model proxy translates it into each vendor's own reasoning_effort. One line of code runs the same experiment on OpenAI, Anthropic, Google and open-weight models. Here is the core of every call: # src/dial/task_template.py.txt (inlined into each Kaggle task by build_tasks.py) with kbench.chats.new(name) as chat: try: raw = llm.prompt(prompt, reasoning=level, seed=SEED_BASE + repeat) traces = kbench.last_reasoning_traces() except Exception as e: raw, traces = None, None rec["error"] = f"{type(e).name}: {e}" # a provider 400 becomes an "api_error", which is data for H1 u = chat.usage rec["output_tokens"] = u.output_tokens # hidden reasoning tokens included rec["cost_nanodollars"] = u.total_cost_nanodollars # what Kaggle's proxy actually charged chat.usage is the reason this benchmark can measure cost instead of guessing it. Every call records its own token count and its own cost in nanodollars, and those records end up in Kaggle's run files. Each dial-<level> task fans out all 120 calls (60 items Γ 2 repeats) as a nested task: records = _records(dial_call.evaluate( llm=[llm], evaluation_data=df, n_jobs=8, on_failure="continue")) on_failure="continue" matters. When gpt-oss-120b rejects none, the run doesn't crash: it records the rejection, and the rejection becomes one of the findings. Grading is deliberately boring. The prompt ends with "End your response with a final line in exactly this format: FINAL ANSWER: ", and the parser takes the last match: # src/dial/parse.py MARKER = re.compile( r"FINAL[ \t]+ANSWER[ \t][`][ \t][:οΌ][ \t](.)$", re.IGNORECASE | re.MULTILINE, ) It handles FINAL ANSWER: 42 and full-width colons, strips <think> blocks, and is covered by the repo's test suite. I didn't use the library's structured output (schema=) because it breaks when reasoning is on: the proxy returns <think>β¦</think>{json} and JSON parsing fails. Plain text plus my own parser was the only approach that worked the same way across all four vendors. The Kaggle CLI handled the rest of the loop: kaggle b t push dial-high -f tasks/dial_high.py --wait # also runs it once on Kaggle's default model kaggle b t run dial-high -m gpt-5.4-mini-2026-03-17 # one model, one dial position kaggle b t download dial-high -o results/main # per-call records β analysis/analyze.py The first command had a side effect I didn't plan for. Pushing a task runs it once on Kaggle's default model, gemini-3.7-flash. Those validation runs were full, clean runs of the locked items, so Gemini joined the lineup without me choosing it. Kaggle reserves each call's maximum possible cost against your quota before the call runs: the input cost plus 128,000 output tokens at the model's output price. For gpt-5.4-mini that is about $0.58 reserved per call, even when the call actually costs $0.0002. With 8 calls in parallel, one task reserves about $4.61. My first gpt-5.4-mini run launched several tasks at once, and 378 of 480 calls came back as HTTP 403 quota refusals. Real spend was nowhere near the cap. I made a rule and wrote it into the pre-registration: a quota 403 is missing data, never model behavior. I discarded those runs (they're listed in results/superseded.json, and the analysis excludes them) and re-ran gpt-5.4-mini one task at a time after the quota reset. The re-run was clean: 0 refusals, 0 API errors. gpt-5.4-mini: $0.044 at none, $0.371 at low, $0.588 at medium and $0.713 at high, $1.72 in total. Claude Sonnet 5: dropped from the lineup. Its reservation is about $1.28 per call, and 8 parallel calls (β$10.24) exceed the $10 daily quota on their own. The analysis reports it as "dropped (platform quota reservation)" rather than leaving it out silently. Model Levels run Graded calls Main-run cost gpt-5.4-mini-2026-03-17 none, low, medium, high 480 $1.716 gemini-3.7-flash none, low, medium, high 480 $2.034 claude-haiku-5.5 none, low, medium, high 480 $0.186 gpt-oss-120b low, medium, high (rejects none with HTTP 400) 360 $0.263 claude-sonnet-5 n/a 0 dropped (quota reservation, see above) 1,800 graded calls in total: 60 items Γ 2 repeats Γ 15 model-levels. There were 0 API errors and 0 quota refusals in the final data, and 1,964,378 output tokens. The main run cost $4.20; the whole project, including the pilot and two calibration rounds, cost about $6.38 of Kaggle quota. A note on control: Kaggle's proxy silently drops temperature, so I can't claim temperature 0, and the Google models also drop the seed. That is why every cell has two repeats, and why the analysis reports a noise floor (Β§7 of summary.md). Median output tokens per call (log scale). The same parameter name produces four different behaviors. The same parameter means something different for every vendor: Model Median output tokens, none β low β medium β high high Γ· lowest Label (pre-registered rule) gpt-5.4-mini 18 β 168 β 195 β 254 Γ14.1 honors gpt-oss-120b rejected β 292 β 446 β 1,340 Γ4.6 (from low) honors, rejects none claude-haiku-5.5 150 β 162 β 171 β 298 Γ2.0 honors gemini-3.7-flash 453 β 446 β 566 β 1,013 Γ2.2 neither; none β low Two of these surprised me: none doesn't mean "no thinking" on Gemini. At none it spends a median 453 output tokens, more than gpt-5.4-mini spends at high (254). It returns no reasoning trace at none or low, and its one-line replies look the same as a model that didn't think at all. You pay for that thinking but can't see it. none doesn't exist on gpt-oss-120b. It returns HTTP 400 instead. If you write reasoning_effort="none" into a multi-vendor config, one of these four models will error. Per-family token ladder (exploratory, not pre-registered) Split by family, the effect of the dial depends heavily on the question. gpt-5.4-mini's Deduce tokens go 18 β 1,579 β 2,856 β 3,040 (Γ169). On Arith they only go 128 β 220. gpt-oss-120b spends Γ20.6 more tokens on Distract at high than at low, on questions whose answer is in the first sentence. 2. One real accuracy gain, and it came from the first step (H3: β
for gpt-5.4-mini only) Accuracy at each level with 95% item-cluster bootstrap CIs. The dashed line in the logic panel is random guessing among 7 names. At none, gpt-5.4-mini is right 15% of the time, which is about the random-guess rate. gpt-5.4-mini on Deduce none low medium high Accuracy 15% 75% 87.5% 97.5% 95% CI 5β28% 62β88% 78β97% 93β100% Median output tokens 18 1,579 2,856 3,040 Cost per correct answer $0.0019 $0.0109 $0.0149 $0.0160 The change from none to high is +82.5 percentage points (95% CI +70 to +95, Holm-adjusted p = 0.00056). It is the only effect in the study that passes both pre-registered criteria. Most of the gain comes from the first step: none β low adds 60 points. Each further step adds less accuracy and costs more. All 12 pre-registered tests. One is significant. Six can't move because the model already scores 100% at its lowest setting. gpt-oss-120b's +17.5 points on logic look like a gain but don't survive the correction for 12 tests (Holm p = 0.597). The other models have a less exciting result that matters for the dial question: Claude Haiku 5.5 scored 97.5% on these puzzles at none, and Gemini scored 100%. For them, turning the dial up on Deduce had nothing left to fix. I expected counting with distractors to be where more effort backfires (H2). The pre-registered test says no: no model got significantly worse at high. gpt-oss-120b went from 75% (low) to 67.5% (high): β7.5 points, CI β22.5 to +7.5. That isn't significant. gpt-5.4-mini went from 75% to 85%, also not significant. Haiku and Gemini scored 100% at every level. H2 is not supported, and I'm reporting it that way. The errors were still worth reading. Every wrong Distract answer, from every model, at every level, was an over-count. No model ever answered too low. The models weren't confused about what they owned; they added the distractor numbers to the count. Three replies from the main run at high, decomposed. article_visuals.py asserts every sum against the raw run files and the locked items before it draws, so none of these numbers are typed by hand. In the first two replies the model adds every number in the prompt. In the third it solves someone else's homework and reports that as the answer. The longer replies were also the wrong ones. At high, gpt-5.4-mini's wrong Distract answers used a median of 412 output tokens; its correct answers used 91.5. gpt-oss-120b showed the same pattern more strongly: 5,776 tokens for wrong answers against 747 for correct ones. In the most extreme case, gpt-oss-120b spent 21,875 tokens and almost four minutes (235 s) to answer 6 when the answer was 5. gpt-oss-120b also became less consistent as effort went up. On Distract at high, it gave different answers on the two repeats for 45% of items, compared with 0% at low. So H2 fails as a significance test, and I won't claim it passed. But "more effort made it worse" isn't the right summary either. A better one is more effort gave it more room to over-count. What turning the dial to high bought. Seven of twelve cells sit on the zero line: same accuracy at Γ1.5 to Γ3.4 the cost per correct answer. Only one point is high above the line. This is the chart I'd show anyone who sets reasoning_effort in production: In 7 of 12 cells, high changed nothing but the price. Accuracy stayed the same while cost per correct answer rose Γ1.5 to Γ3.4. On Arith, the control family, every model stayed at its accuracy while cost per correct rose Γ1.65 (gpt-5.4-mini) to Γ2.5 (gpt-oss-120b). Spending more effort on an already-solved problem buys nothing. gpt-oss-120b on Distract paid about Γ25 per correct answer ($0.00006 β $0.00151) and scored 7.5 points lower (not significant). The one big win was expensive too. gpt-5.4-mini's logic fix raised cost per correct answer Γ8.5. Cost per correct answer (USD, log scale). Every line rises from left to right. None of them falls. One counter-intuitive detail: on Deduce, gpt-5.4-mini at none is the cheapest per correct answer ($0.0019, versus $0.016 at high), even though it's wrong 85% of the time. Cost per correct answer only tells you something when you can tell which answers are right. In most real uses you can't, so don't read that row as advice to choose none. Hypothesis Result H1 The dial is not one instrument β
Supported. Γ14 vs Γ2 token scaling; gpt-oss rejects none; Gemini's none β low and thinks invisibly H2 More thinking hurts on distractor counting β Not supported. No significant drop. Every error is an over-count, and wrong answers are longer H3 More thinking helps on deduction β
gpt-5.4-mini only: 15% β 97.5% (Holm p = 0.00056). Others are at the ceiling or not significant H4 When accuracy doesn't rise, the bill still does β
Supported (descriptive). 7 of 12 cells unchanged at Γ1.5 to Γ3.4 cost per correct answer These come from 20 items per family, 2 repeats and 4 models, so treat them as starting points rather than rules: Measure none first. Two of the four models (Haiku and Gemini) were already at or near 100% on every family at their lowest setting. If none fails, try low before high. For the one model that needed the dial, none β low delivered 60 of the 82.5 points. Don't assume a level does the same thing across vendors. none errors on one, is invisible thinking on another, and is a real off switch on a third. Watch for long answers to simple questions. On Distract, long replies were the wrong ones. The model that needed the dial most was the one where none really meant none. gpt-5.4-mini answered logic puzzles in 18 tokens at none and was right about as often as a random guess. Gemini's none isn't free. It spent a median of 453 hidden tokens at none, which is more than gpt-5.4-mini spent at high. The errors all went in one direction. Every wrong Distract answer was too high, and none was too low. Pushing a task added a model to my lineup. Kaggle's default-model validation runs are how Gemini got into the study. The quota system nearly faked a finding. 378 quota refusals could have been read as "gpt-5.4-mini fails at high effort". The pre-registered rule (403 = missing data) kept them out. Small samples. 20 items per family and 2 repeats per cell. Most "no effect" results mean not detectable at this size, not proven zero. Ceilings. Haiku and Gemini are at or near 100% almost everywhere, so I can say what the dial costs them but not what it buys. Harder items would be needed for that. Synthetic tasks and one prompt format. These questions were built to isolate effects, not to represent real workloads. No temperature control. Kaggle's proxy drops temperature, and the Google models also drop seed. The repeats and the flip rates are my substitute. Cost is Kaggle's proxy cost. Your provider's pricing may differ. Incomplete lineup. Claude Sonnet 5 was dropped (quota reservation), and gpt-oss-120b has no none level. Its comparisons start from low. Partial view of reasoning. Only Gemini returned reasoning traces (at medium and high). For the rest, output tokens are the only evidence of thinking. Harder Deduce items (more people, more clues) to get Haiku and Gemini off the ceiling and see whether their dial buys anything. More Distract items and repeats. gpt-oss-120b's β7.5 points and 45% flip rate are the direction H2 predicted, but the sample is too small to test them properly. Latency, which every call already records (latency_s, backend_latency_ms) and this post doesn't analyze. Sonnet and other models, run one task at a time to fit the quota reservation. Everything that produced a number in this post is public: the generators, parser, task builder, analysis script, pre-registration and every raw Kaggle run file. git clone https://github.com/Abeera81/reasoning-dial && cd reasoning-dial pip install -r requirements.txt python -m pytest -q # 68 tests: generators, parser, task builder, analysis python analysis/analyze.py # rebuilds results/summary.md and every figure from results/main/ analyze.py checks the locked test set's SHA-256 before it computes anything. To add a model, fork any of the four public Kaggle tasks and run it. Because each task holds only the dial level, the comparison stays fair. The dial is a real control. For one model on one kind of problem, it made the difference between guessing and solving. For everything else I measured, it mostly raised the price. Before you turn it up, check what none gets you. If you run your own models through the tasks, I'd like to see the ladders they produce. ποΈ
Key Takeaways #
- β’This is a submission for the Kaggle Benchmarking Challenge I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?" With reasoning effort set to none, it replied: Cleo FINAL ANSWER: Cleo 18 output tokens
- β’This story was reported by Dev.to , covering developments in thedev space.
- β’AI advancements continue to reshape industries β read the full article on Dev.to for complete coverage.
π Continue reading the full article:
Read Full Article on Dev.to β