When Should a Fast AI Stop and Think? A new arXiv paper (2610.09683) by the doomLaya developer tests gating a slow reasoning model behind a fast "System One" model in Doom, calling the reasoner only when one of five signals fires: unsure, stuck, blocked, loop, or no new area. Offline, handing over the least confident decisions helped only when the small model's confidence was meaningful, while in live play a fixed rule matched the reasoner's performance and the reasoner's plans mostly bought fewer deaths. The fast models (Laya, Kev, Lev, Julia, Clef, Strands Decider) return per-option probabilities in about 29 ms per decision on a consumer GPU, while the slow side is Qwen3.6-35B-A3B with Gemma 4 26B for comparison, scored against a rule-based referee on 900 questions from unseen seeds. A small decision model plays Doom on its own. A reasoning model is called in only when the small one is unsure or stuck. Offline, handing over the least confident decisions helps, but only when the small model’s confidence means something. In the live game, a fixed rule went as far as the reasoner; the reasoner’s plans mostly bought fewer deaths. I have played Doom for a long time, and I have noticed how I play it. When the action works I don’t think: no inner voice, I just move and shoot. Thinking starts when shooting stops helping. A door that will not open, an enemy that keeps coming back, a corridor I have already walked twice. Then I stop, work out what is going on, and the thinking is what gets me moving again. The game does not pause while I do it. That is the question behind this article, and behind a paper I have just put on arXiv https://arxiv.org/abs/2610.09683 : if a fast model that acts is paired with a slow model that reasons, when should the fast one hand over? Last month I compared two small models playing Doom Laya vs Gemma https://medium.com/towards-artificial-intelligence/laya-vs-gemma-who-wins-the-doom-game-18d3b30f303b . This time a small model plays, and a reasoning model is called only when a gate opens. Agents that pair a fast and a slow model usually do one of two things. In turn-based settings the slow model is called on events, and the world waits for it. In real-time settings the slow model runs all the time, in parallel. I wanted the combination I use as a player: the fast model takes every decision, and the reasoner is called only on events, while the game goes on. The gate watches five signals: unsure low confidence three times in a row , stuck barely moved in 5 seconds with no enemy in sight , blocked half of the recent commands could not be executed , loop the same command 80% of the time over 15 seconds and no new area nothing new reached in 15 seconds . When one fires, the reasoner’s command becomes a plan that overrides the fast model. The game is Doom, with the free Freedoom assets, run through ViZDoom and doomLaya https://github.com/azalio/doomLaya . doomLaya turns the game into questions. Every half second of game time the agent gets a text description of the situation health, ammo, enemies, items, doors and picks one command: attack, collect an item, open a door, go to the exit, retreat, explore or wait. The options always come in the same order, enemies first. An executor turns the command into button presses. The fast side is one of the open “System One” models that appeared in September: Laya, Kev, Lev, Julia, Clef, Strands Decider and others. They don’t generate text. They read the question and the options and return a probability for each option in a single forward pass, about 29 ms per decision on a consumer GPU. llama.cpp now serves all of them through one endpoint, so one bridge works for every model. The slow side is a vision-language model that reasons before answering: Qwen3.6–35B-A3B, with Gemma 4 26B for comparison. Offline it reads the same text as the fast model; in the game it also sees the current frame. It writes a short thought and picks a command. To score decisions, doomLaya has a rule-based referee: fight visible enemies, collect only what you need, open doors in the way, otherwise head for the exit. “Accuracy” below means agreement with that referee, on 900 questions from games played on seeds never used for training. It is one reasonable policy, not the truth, and it never answers “explore”: keep that in mind for the live games. Everything runs on home PCs. On average a question has 7.5 options, and 3.6 of them are items to collect, so many wrong answers would land on an item by chance. The System One models pick an item 1.6 to 1.8 times more often than a random wrong answer would, from 0.15B to 9B parameters and from different makers. The exception is the smallest Kev, whose probabilities are nearly uniform. The reasoner shares the bias without reasoning 1.8 and less with it 1.4 . The order of the options matters too. In doomLaya’s fixed order the correct answer is the first option in 41% of the questions. With the options shuffled, Strands Decider drops from 0.59 to 0.38 and Kev 9B from 0.60 to 0.50, while Clef-Flash is unaffected. The pull toward items stays the same in any order. For reference, a small Gemma fine-tuned on doomLaya’s data doomGemma, from the previous article reaches 0.86 on unseen seeds 0.85 with the options shuffled . A gate needs more than accuracy: the fast model’s confidence must be lower when it is wrong. I measured this as AUROC, where 0.5 means confidence says nothing about being right and 1.0 means it separates right from wrong perfectly. Five models agree with the referee 54–60% of the time, yet their AUROC goes from 0.60 to 0.89. Calibration does not predict it: the two best-calibrated models sit in the middle. But the overall number hides two kinds of signal, which show up when I look within each type of situation fights, items, doors, exits . Laya base is confident on fights and doors, which it gets right, and unsure on exits and items, which it gets wrong. Within one type of situation, its confidence does not tell its right answers from its wrong ones. Strands and Kev 9B do the reverse: weak overall, better within a type. For a gate, the coarse kind is what pays, because accuracy differs so much between situations from almost 0 on exits to almost 1 on fights for Laya base . One more caveat: with shuffled options the spread narrows 0.52 to 0.77 , so “the most accurate models have the least informative confidence” holds only in doomLaya’s fixed order. The first test is offline: take the fast model’s least confident decisions and give them to the reasoner, with all the time it needs, and compare with handing over the same number at random. At 30%, Laya plus the reasoner reaches 0.75, against 0.60 for random handover, with no training. Because Laya and the 30% rate were picked on the same questions, I re-ran the choice 500 times on one half of the games and measured on the other half. The gain over random handover holds: +0.13 0.08, 0.18 . What does not hold is “better than both models”: the pair roughly matches the reasoner alone 0.69 while calling it on 30% of the decisions. Across nine fast models in both option orders, the gain over random tracks AUROC rank correlation 0.87 . With Strands and Kev 9B, handing over by confidence is the same as handing over at random. With shuffled options the gain shrinks to +0.08 0.02, 0.14 . Without reasoning, the reasoner shares the fast model’s bias and the gain halves. The price is time. The reasoner writes a median of 1,300 tokens per decision and the fast model decides every half second, so handing over 30% of decisions is impossible in real time. The reasoner has to be called on events, and each thought has to guide many decisions. I played 33 games on the first map, three seeds per variant, 180 seconds each. The fast model was mostly the typed-decisions Laya. No variant reached the exit; the closest game came within 4 metres of it. Since there are no success rates, I counted how much of the map each game covered and how many doors it opened. With a bare reasoner, it was called about 14 times per game, but half of its thoughts said “collect the nearest item” and its plans lasted about a second. Then I added a few sentences of Doom knowledge to its prompt, removed from its options the command the fast model kept repeating, and let a plan last until done, impossible, a burst of damage, or 45 seconds. Now it prescribed exploring, and the games opened more doors. The control was the same gate and the same commitment, but with a fixed “explore” in place of a thought. The rule opened as many doors as the reasoner and covered as much ground 83 cells against 64–83 ; it also died more often, 2.7 times per game against 1.3. With three games per variant that could be noise, but the same happened with the other fast model. So most of the benefit came from committing to explore, not from what the reasoner said. Whether the gate opened on low confidence or on lack of progress changed when the reasoner was called, not these outcomes. The base Laya makes the point sharper. On its own it stood still for the whole game in all three seeds. With the gate and the reasoner it played: 66 cells, 8.7 doors. With the gate and the fixed rule it played too, and covered even more ground 81 cells, 9 doors , but died 3.7 times per game against 0.7. What unblocked it was committing to explore; what the reasoner added was survival. Thinking is also expensive in game time. With Doom knowledge and commitment, the reasoner was busy for 38–70% of the game, and a plan started on a frame that was a median of 9–29 seconds old. Reading every thought shows where the chain breaks. In the three committed variants, all 25 failed plans ended at a closed door. The executor reported the door and suggested opening it. The reasoner chose to open it once; in 17 of the 25 thoughts it concluded the door needed a key: “Blocked by a closed door. No key. I need to find one.” The Doom knowledge I had added came from the second map, where a door really does need the yellow key. On the first map it does not, and the state cannot settle it: it reports a closed door, not whether it is locked. Without that knowledge, no thought after a door failure mentions a key, but the reasoner goes back to collecting items 45% of its thoughts . One wrong prescription was traded for another. Each fix moved the problem one step: from when to think, to what the thought prescribes, to how long the plan holds, and then back to what it prescribes, now under a state that cannot tell it what it needs to know. The limits are real. The live tests are one map and 33 games, accuracy is measured against one rule-based policy that never rewards exploring, and the option order affects some of the numbers. No game reached the exit. This is a case study of where the chain breaks, not a benchmark. Two questions stay open: whether a single model running in two modes does better than two models, and how faithful each thought is to what is actually on screen. Method note: I designed the experiments. Code, data analysis and drafting were done with Claude Anthropic under my direction and review; all numbers come from the released data. When Should a Fast AI Stop and Think? https://pub.towardsai.net/when-should-a-fast-ai-stop-and-think-56e202a0afdd was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.