The persistent instinct in AI engineering is to treat model reasoning as an unalloyed capital asset: buy more compute at inference, produce longer deliberation traces, and harvest superior downstream decisions. In enterprise boards and quantitative desks alike, the working assumption has been linear. If an unprompted model produces a coarse forecast, paying for test-time reasoning ought to act like hiring a more deliberate analyst. The extra token expenditure buys deeper verification, filters out hallucinations, and delivers a portfolio allocation with stronger risk-adjusted survival.
A controlled empirical study by Jiayi Chen and Guiling Wang at the New Jersey Institute of Technology (arXiv:2609.30705) directly tests that calculation across a full year of live market conditions. The authors set up a clean economic experiment. Rather than reshuffling prompts, model architectures, memory modules, and execution rules simultaneously (the standard messy cocktail of agentic benchmarks), they froze everything except reasoning effort. Across 241 trading days in 2024, representative models from the DeepSeek, GPT, and Gemini families evaluated 100 liquid U.S. equities under identical rules. The portfolio engine ranked the assets, bought the top decile, shorted the bottom decile, balanced the sides for market neutrality, and deducted a realistic 10 basis points of one-way transaction friction.
The findings deliver an uncomfortable reality check for teams treating inference tokens as a proxy for alpha: across more than 800,000 asset predictions and nine primary annual baselines, buying extra reasoning produced zero reliable improvement in net portfolio returns.
Paying for inference-time reasoning clearly altered model behavior. It expanded token usage, shifted relative rankings, and produced noticeably more elaborate justifications. Yet across all nine annual low-versus-baseline evaluations, every single 95% confidence interval crossed zero and failed to clear a modest economic hurdle of a quarter basis point per day. For DeepSeek, where the authors mapped the full progression from zero reasoning up to maximum capacity, the response curve was jagged and nonmonotonic. Moving from high reasoning to maximum reasoning helped claw back some performance in identifiable news, yet maximum reasoning still lagged behind the baseline that ran with zero deliberation tokens.
The instability runs deeper than noisy point estimates. When the researchers audited identical decision tasks across repeated stochastic generations on frozen market dates, the treatment effects flipped sign. GPT shifted from positive to negative net returns across repeated calls with identical information. Gemini swung between positive returns and outright losses. In DeepSeek’s audit on masked company filings, paying for reasoning caused active, statistically significant economic harm, eroding returns by 9.0 basis points per day relative to the unreasoned baseline.
The core disconnect stems from what reasoning models actually optimize. On benchmarks like GSM8k, SWE-bench, or competitive programming, the search space contains verifiable ground truth. An extended reasoning trajectory can self-correct logical errors, catch arithmetic blunders, and eliminate invalid solutions. The reward surface is deterministic.
Market pricing is not an arithmetic contest with a single latent answer. It is a noisy, non-stationary clearing mechanism where information is already partially reflected in asset prices. Giving an LLM hundreds of additional reasoning tokens to contemplate a balance sheet or a masked news snippet does not manufacture non-public information. Instead, it frequently leads the model to over-interpret circumstantial noise, construct brittle narrative justifications for random price fluctuations, and abandon sturdy probabilistic priors in favor of elaborate rationalizations. In quantitative terms, test-time compute increases the variance of portfolio selection around the threshold boundaries without increasing the information coefficient.
That operational variance compounds violently once capital meets the physical reality of the order book. When extra reasoning reshuffles rankings near the decile cutoff, it rotates holdings. Rotating holdings generates turnover. And in market microstructure, turnover is an absolute cash drain: execution spread, market impact, and transaction fees subtract capital with certainty, while the speculative edge remains purely hypothetical.
The study also exposes a crucial divergence between structural compliance and economic value. In several conditions, enabling reasoning reduced formatting errors and improved adherence to strict output schemas. But operational reliability is not alpha. An agent can emit a perfectly formatted JSON object that complies flawlessly with every type constraint, yet still select a basket of stocks that systematically bleeds cash. Confusing syntactic adherence with profitable judgment is one of the most expensive errors currently occurring in enterprise agent deployments.
For institutional allocators and engineering leads, the paper establishes a necessary governance rule: inference-time compute is not a universal capability upgrade. It is an operational parameter change with a guaranteed token bill, concrete latency drag, and highly uncertain economic payoff. Until an engineering team can demonstrate that reasoning tokens generate a reproducible shift in the realized distribution after execution friction, baseline compute remains the rational default. Buying unverified deliberation is simply paying a hyperscaler for the privilege of over-analyzing noise.