{"slug": "we-thought-our-gpt-5-4-agent-got-lazier-in-production-it-was-a-3-bug-workflow-it", "title": "We thought our GPT-5.4 agent got lazier in production — it was a 3-bug workflow teaching it to quit", "summary": "A developer traced a production n8n agent's apparent \"laziness\" to three workflow bugs rather than a model regression, finding that a formatting branch rewarded shallow answers and a reduced retry budget (from 6 to 2) caused early stopping. Running the same broken scaffold against multiple model families produced identical degradation, confirming the orchestration layer, not the model, was at fault. The fix restored quality, and the developer now recommends inspecting stop reasons and tool paths instead of judging final-answer polish.", "body_md": "We had an n8n agent that looked great in staging.\n\nIt would:\n\nThen we shipped it.\n\nIn production, the same task started ending after one shallow pass.\n\nSymptoms were exactly what people usually call “model laziness”:\n\nOur first instinct was to blame GPT-5.4.\n\nThat was the wrong diagnosis.\n\nThe real issue was boring and very fixable:\n\n`6` to `2`\nOnce we fixed the workflow, quality came back.\n\nThat changed how I think about “lazy” agents in production.\n\nMost of the time, the model did not suddenly get worse. Your orchestration started ending runs early.\n\nIn staging, the agent trace looked like this:\n\nIn production, it looked more like this:\n\nSame task class. Same model family. Very different behavior.\n\nAnd because the final answer still looked polished, it passed casual review more often than it should have.\n\nThat is the dangerous part.\n\nA broken agent rarely looks broken in an obvious way. It often looks efficient.\n\nThis was the biggest quality hit.\n\nFor simple classification, `2` retries can be fine.\n\nFor research, debugging, document synthesis, or anything with tool use, `2` is often a trap. One bad retrieval result plus one tool hiccup and the agent is out of budget.\n\nExample of the kind of config drift that causes this:\n\n```\n{\n  \"task_type\": \"research\",\n  \"max_retries\": 2,\n  \"timeout_seconds\": 20\n}\n```\n\nThat looks harmless until your workflow depends on search + fetch + verify.\n\nIn our n8n flow, one formatting branch said, effectively:\n\nThat means the agent could skip retrieval depth and still win.\n\nThis is how you accidentally train an agent to stop early.\n\nPseudo-logic:\n\n``` js\nconst passed =\n  isValidJson(response) &&\n  response.answer.length > 280;\n\nif (passed) {\n  return \"success\";\n}\n```\n\nThat is not quality control.\n\nThat is a shallow-answer reward function.\n\nThis one is common in OpenAI-compatible stacks.\n\nIf your app accepts the first plausible answer and never checks whether the expected tool path ran, the orchestration layer starts selecting for speed, not depth.\n\nThat can happen whether you are routing to GPT-5.4, Claude Opus 4.6, or Grok 4.20.\n\nThe model is not “choosing to be lazy” in some abstract sense.\n\nYour workflow is telling it:\n\nif you look done quickly enough, you pass\n\nBecause production has constraints that staging often hides.\n\nIn a clean test harness, a model usually gets:\n\nIn production, the agent sits inside a box made of:\n\nThat box matters more than people want to admit.\n\nA strong model inside a bad loop will look worse than a decent model inside a clean loop.\n\nRun the same task through the same scaffold and change one variable at a time.\n\nNot “same prompt, different environment.”\n\nActually the same scaffold:\n\nIf you compare production n8n against a clean notebook script, you are not isolating the model.\n\nYou are changing the entire experiment.\n\nNot vibes.\n\nNot output length alone.\n\nNot “this answer feels thinner.”\n\nStop reasons told us far more than final-answer scoring.\n\nFor Anthropic agents, useful stop reasons include values like:\n\n`end_turn`` max_tokens``tool_use`` pause_turn`\nFor OpenAI-compatible workflows, inspect whether:\n\nIf you only evaluate the final answer, you are debugging blind.\n\nWe ran the same broken production scaffold against multiple model families.\n\nWhat we saw:\n\nThat pattern matters.\n\nWhen three strong models all become “lazy” in the same way, the workflow is usually guilty.\n\nHere is the mental model I use now:\n\n| If this changes | Suspect | \n|---|---|\n| One model regresses, others stay stable | model or provider issue | \n| All models regress under one workflow | orchestration bug | \n| Output gets shorter after retry/timeout changes | early stopping | \n| JSON validity improves while answer quality drops | parser-first reward problem | \n\nWe spent too long blaming the model layer.\n\nOur guesses were reasonable:\n\nThose are all real failure modes.\n\nThey just were not the main problem here.\n\nThe actual issue was simpler:\n\nstaging rewarded grounded completion\n\nproduction rewarded acceptable formatting\n\nAgents optimize for whatever your workflow rewards.\n\nIf your automation says “close enough,” GPT-5.4, Claude Opus 4.6, and Grok 4.20 will all start looking suspiciously eager to be done.\n\nThese are the ones I would audit first.\n\nIf the main objective is valid JSON, many agents will satisfy the parser before they satisfy the task.\n\nExample smell:\n\n```\nif (schema.safeParse(output).success) {\n  return success;\n}\n```\n\nThat should almost never be the whole success condition for a research task.\n\nIf n8n, Make, Zapier, OpenClaw, LangGraph, or your custom loop allows a final answer before retrieval or verification, expect shallow completions.\n\nFor research-class tasks, tool use often should not be optional.\n\nThis one is everywhere.\n\nPeople use one global retry budget for everything:\n\n```\nmax_retries: 2\n```\n\nThat might be fine for:\n\nIt is usually bad for:\n\nThis is the worst one.\n\nIf you only score the final text, you hide:\n\nMy strong opinion: this single habit causes teams to think they are comparing models when they are actually comparing orchestration mistakes.\n\nWe did not switch models.\n\nWe changed the workflow.\n\nWe moved the retry cap back from `2` to `6` for research-class tasks.\n\n```\nagent_profiles:\n  classification:\n    max_retries: 2\n  research:\n    max_retries: 6\n  debugging:\n    max_retries: 6\n```\n\nWe added a hard gate.\n\nIf the task is research, at least one retrieval step must happen before a run can pass.\n\nPseudo-code:\n\n```\nfunction validateRun(run) {\n  if (run.taskType === \"research\" && run.toolCalls.search < 1) {\n    return { ok: false, reason: \"missing_required_retrieval\" };\n  }\n\n  if (!run.outputSchemaValid) {\n    return { ok: false, reason: \"invalid_schema\" };\n  }\n\n  return { ok: true };\n}\n```\n\nWe changed success conditions so a run could not pass on formatting alone.\n\nThat meant checking trajectory, not just output shape.\n\nWe reviewed traces side by side in our observability stack and checked stop reasons across both the OpenAI-compatible path and the Anthropic path.\n\nThat made the difference obvious fast.\n\nIf you think your production agent got worse, this is the order I would check things in.\n\n```\ndiff staging-agent.yaml production-agent.yaml\n```\n\nLook for changes in:\n\nLog them explicitly.\n\n```\n{\n  \"run_id\": \"abc123\",\n  \"model\": \"gpt-5.4\",\n  \"stop_reason\": \"end_turn\",\n  \"tool_calls\": 0,\n  \"task_type\": \"research\"\n}\n```\n\nIf research tasks are ending with zero tool calls and still passing, that is your bug.\n\nA simple metric catches a lot:\n\n```\nSELECT\n  task_type,\n  AVG(tool_call_count) AS avg_tool_calls,\n  AVG(retry_count) AS avg_retries,\n  AVG(output_chars) AS avg_output_chars\nFROM agent_runs\nWHERE created_at >= NOW() - INTERVAL '7 days'\nGROUP BY task_type;\n```\n\nIf tool-call counts collapse after a deploy, investigate the workflow before blaming the model.\n\nThis is where an OpenAI-compatible API setup helps.\n\nIf GPT-5.4, Claude Opus 4.6, and Grok 4.20 all fail the same way under one loop, the loop is probably broken.\n\nTrack things like:\n\nThis kind of bug gets expensive fast when you are running automations all day.\n\nNot just in dollars. In bad outputs, hidden regressions, and wasted debugging time.\n\nTeams running agents in n8n, Make, Zapier, OpenClaw, or custom OpenAI-compatible stacks usually hit the same wall:\n\nthey start by asking “which model is best?”\n\nThen eventually they realize the more useful question is:\n\n“what exactly is our workflow rewarding?”\n\nThat is also why predictable API infrastructure matters.\n\nWhen you can swap models without rewriting your stack, compare traces cleanly, and run lots of evals without per-token anxiety, it gets much easier to find orchestration bugs instead of arguing about vibes.\n\nThat is a big part of why Standard Compute is interesting for agent teams: it is a drop-in OpenAI-compatible API, so you can keep your existing SDKs and workflows, route across GPT-5.4, Claude Opus 4.6, and Grok 4.20, and test agent behavior without every debugging session turning into a billing event.\n\nFor teams running automations 24/7, flat monthly pricing is not just a finance preference. It changes how aggressively you can evaluate, compare, and fix agent systems.\n\nBefore blaming GPT-5.4 for getting lazy:\n\nSometimes a model really does regress.\n\nMore often, production taught your agent that quitting early is the winning move.", "url": "https://wpnews.pro/news/we-thought-our-gpt-5-4-agent-got-lazier-in-production-it-was-a-3-bug-workflow-it", "canonical_source": "https://dev.to/lars_winstand/we-thought-our-gpt-54-agent-got-lazier-in-production-it-was-a-3-bug-workflow-teaching-it-to-quit-1b3b", "published_at": "2026-10-01 22:10:28+00:00", "updated_at": "2026-10-01 22:14:27.285205+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "mlops"], "entities": ["n8n", "GPT-5.4", "Claude Opus 4.6", "Grok 4.20", "Anthropic", "OpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/we-thought-our-gpt-5-4-agent-got-lazier-in-production-it-was-a-3-bug-workflow-it", "markdown": "https://wpnews.pro/news/we-thought-our-gpt-5-4-agent-got-lazier-in-production-it-was-a-3-bug-workflow-it.md", "text": "https://wpnews.pro/news/we-thought-our-gpt-5-4-agent-got-lazier-in-production-it-was-a-3-bug-workflow-it.txt", "jsonld": "https://wpnews.pro/news/we-thought-our-gpt-5-4-agent-got-lazier-in-production-it-was-a-3-bug-workflow-it.jsonld"}}