Nvidia paper finds AI models lose accuracy on long tasks, with steep drops at scale Nvidia researchers reported in the paper "Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability" (arXiv:2609.38712, released September 30, 2026) that average model accuracy fell 62.8% as context length scaled from 4,000 to 128,000 tokens across seven open-weight models. DeepSeek dropped from 0.909 to 0.554 and Nvidia's own Nemotron Super fell from 0.711 to 0.138, while varying input formats caused a 36.5% average decline, removing stable row or item IDs caused relative drops of up to 64.3%, and increasing local step complexity produced a 39.9% average drop. The Long-Transduction benchmark, which used 1,440 documents per model with greedy sampling and exact-match scoring, led the researchers to recommend breaking long tasks into smaller subtasks and labeling each item with a structured identifier. Nvidia paper finds AI models lose accuracy on long tasks, with steep drops at scale A new Nvidia benchmark shows open-weight models stumble as context grows, and missing item labels make things worse AI agents are increasingly asked to handle long, multi-step jobs. New research from Nvidia https://cryptobriefing.com/markets/nvidia/ suggests they get worse the longer those jobs run. In a paper released on September 30, 2026, Nvidia researchers found that average model accuracy fell by 62.8% as context length scaled up. What Nvidia actually tested The paper is titled “Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability” and is listed as arXiv:2609.38712. It introduces a diagnostic benchmark called Long-Transduction. The goal was narrow on purpose. The benchmark tries to isolate one specific failure: keeping track of state while executing a task over a long stretch of input. The researchers evaluated seven open-weight models, meaning models whose trained parameters are publicly available. The tasks were deliberately mundane: arithmetic, sorting UUIDs, looking up variables, and transforming tables. The test was whether models could keep doing them correctly as the input stretched from 4,000 tokens to 128,000 tokens. Each model processed 1,440 documents. The team used greedy sampling, which means the model always picks its single most likely next output, and exact-match scoring, which gives no partial credit. Either the answer is right or it isn’t. The numbers get ugly at length The headline figure is that 62.8% average accuracy drop when moving from 4K to 128K tokens. Individual models told the story more vividly. AI, tech, and the markets they move—in one daily briefing. Daily. Free. Join 34,000+ readers across crypto, finance, and policy. DeepSeek https://cryptobriefing.com/markets/deepseek/ started strong, scoring 0.909 at the shorter length. At the longer end, it slid to 0.554. Nvidia’s own Nemotron Super fared worse. It fell from 0.711 to 0.138. Length was not the only problem. The study identified several other factors that destabilized performance: - Input formats: Varying the format of the input led to a 36.5% average decline in performance. - Missing identifiers: Removing stable row or item IDs caused relative performance drops of up to 64.3%. - Local complexity: Making each individual step more complex produced an average drop of 39.9%. The paper also notes that these weaknesses compound. Greater local complexity and missing IDs had cascading effects on alignment, meaning one slip can throw off everything that comes after it. Why short benchmarks can mislead The central argument of the paper is that doing well on short tasks does not guarantee reliable results in production-scale workflows. The Nvidia study adds to a growing body of research from 2025 and 2026 that has repeatedly documented long-horizon limits in large language models and agents, across both academic and industry settings. This paper’s contribution is isolating the failure cleanly, with tasks simple enough that the only real variable is staying on track. What this means for teams building with AI agents The researchers offer fairly concrete guidance. Break long tasks into smaller subtasks, and label each item with a structured identifier. For developers, the takeaway is that pipeline design may matter as much as model choice. The 64.3% figure for missing identifiers suggests labeling is not a nice-to-have. It is also notable that Nvidia, whose hardware powers much of the AI industry, published results showing its own Nemotron Super falling to 0.138. Model makers have competed on how much text their systems can accept. This research suggests accepting the text and reliably working through it are two different capabilities. Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .