I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything. A Kaggle benchmark called Reasoning Dial found that raising the reasoning_effort parameter from none to high produced exactly one statistically real accuracy gain across 1,800 graded calls on 4 models: gpt-5.4-mini on logic puzzles improved from 15% to 97.5%. In 7 of 12 model × task cells, accuracy did not change while cost per correct answer rose ×1.5 to ×3.4, and the same model at high effort answered a tool-counting question with 10,016. The benchmark, run by Abeera Alodhi with pre-registered hypotheses and a total main-run cost of $4.20, also found one model spends ×14 more tokens at high, one rejects none with an HTTP 400, and one reasons at none anyway. I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything. This is a submission for the Kaggle Benchmarking Challenge I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?" With reasoning effort set to none, it replied: Cleo FINAL ANSWER: Cleo 18 output tokens. $0.00024. Wrong. The answer is Fay. It gave This is a submission for the Kaggle Benchmarking Challenge I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?" With reasoning effort set to none, it replied: Cleo FINAL ANSWER: Cleo 18 output tokens. $0.00024. Wrong. The answer is Fay. It gave the same wrong answer, word for word, on the second repeat. At high it spent 1,333 tokens, cost about $0.006, and said Fay. That looks like an argument for always choosing high. On another item I asked the same model, also at high, to count the tools in "You have a chisel and a drill." It answered 10,016. Both replies came from the same model, setting and benchmark. That is the whole post in two examples: the reasoning dial sometimes matters a lot. Most of the time it only makes the bill bigger. TL;DR: Reasoning Dial is a Kaggle benchmark that changes one API parameter, reasoning effort none / low / medium / high , and keeps everything else fixed: the same 60 code-generated questions, the same prompt and the same deterministic grader. Over 1,800 graded calls on 4 models, high produced exactly one statistically real accuracy gain: gpt-5.4-mini on logic puzzles, 15% → 97.5%. In 7 of 12 model × task cells, accuracy did not change while cost per correct answer rose ×1.5 to ×3.4. The dial also isn't one instrument: one model spends ×14 more tokens at high, one rejects none with an HTTP 400, and one thinks at none anyway. Hypotheses and analysis were pre-registered before the main run. Total main-run cost: $4.20. 🏆 Kaggle benchmark leaderboard : https://www.kaggle.com/benchmarks/abeeralodhi/reasoning-dial 🎛️ The four public tasks, one per dial position: dial-none · dial-low · dial-medium · dial-high 💻 Code, raw run files and pre-registration: https://github.com/Abeera81/reasoning-dial / reasoning-dial Reasoning Dial Reasoning Dial is a Kaggle benchmark about one setting developers choose blind: reasoning effort none / low / medium / high . It holds everything else fixed same items, same prompt, same grader and measures what turning the dial actually changes: accuracy, output tokens, and cost per correct answer, on three kinds of task. It also asks whether the dial is even the same instrument across vendors. Kaggle benchmark: https://www.kaggle.com/benchmarks/abeeralodhi/reasoning-dial Kaggle tasks public : dial-none · dial-low · dial-medium · dial-high Source: https://github.com/Abeera81/reasoning-dial Pre-registration hypotheses, analysis plan, deviations : docs/PREREGISTRATION.md Design Task families Deduce: 7-person, 7-day scheduling puzzles with exactly one solution checked by brute force . Arith: 6-step word problems with a percentage and an exact division control family . Distract: count the objects you own while ignoring distractors after Gema et al. 2025 . 20 items per family. Dial levels none, low, medium, high … View on GitHub Almost every modern model API has a version of this parameter: reasoning effort, thinking, reasoning. Developers set it every day, usually by feel: "It's a hard question, so high." "It's in production, so low." Almost nobody measures what moving it actually changes. So I built a benchmark where the only variable is the dial. Same items, same prompt suffix, same parser, same grader. For each model, the four Kaggle tasks differ only in the LEVELS value and the two task-name lines. The push CLI reads task names as string literals, so one line of difference wasn't possible. I measured three things at each setting: Accuracy. Strict and deterministic: a regex reads the last FINAL ANSWER: line. There is no LLM judge anywhere. Output tokens. Reasoning tokens are included, because you pay for them whether or not you can see them. Cost per correct answer. The per-call nanodollar cost that Kaggle's model proxy reports, divided by the number of correct answers. Family What it is Why it's here Example Deduce 7 people, 7 days, 6–12 clues before, immediately before, not on, not adjacent . Brute force over all 7 orders confirms exactly one solution. Where more thinking should help H3 "Who gives the talk on Friday?" → Fay Arith Six-step word problems with a percentage step and an exact division; 4–6 digit answers The control. Every model scored 100% at none in calibration, so it shows what the dial costs when there is nothing left to gain H4 A courier's fuel bill with 10% tax → 11880 Distract "You have a chisel and a drill." followed by 1–3 irrelevant numbers: a shop's sales, someone's homework, a code snippet Where more thinking might hurt. This follows Gema et al. 2025, Inverse Scaling in Test-Time Compute H2 "Calculate how many tools you have." → 2 Here is a real Distract item from the locked test set, exactly as the models saw it: You have a chisel and a drill. A shop in Easton sold 8,377 levels last week. A friend shows you this code: python stock = "saw", "hammer", "screwdriver" print len stock 3 A classmate is working on a homework problem: if 44 crates each hold 37 saws, how many are there altogether? Question: Calculate how many tools you have. Answer instruction: Answer with a whole number. End your response with a final line in exactly this format: FINAL ANSWER: