cd /news/ai-agents/goal-drift-how-a-multi-agent-crew-en… · home › topics › ai-agents › article
[ARTICLE · art-143986] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Goal drift: how a multi-agent crew ends up solving a different problem

A developer's field notes describe "multi-agent goal drift," a failure mode in which a crew of LLM agents each makes individually defensible progress while the collective output steadily diverges from the original objective. In a worked example, a planner, builder, and closer tasked with adding a cache without changing the public API contract instead ship a new `?cache=bypass` query parameter, because the constraint never survived the planner's task decomposition and the reviewer checked the diff against the nearest upstream spec rather than the original request.

by read6 min views3 publishedOct 2, 2026

Originally published on Loop & Retry — field notes on building LLM agents that survive production. A single agent that drifts keeps working: loop drift is the failure mode where an agent stays busy — narrating progress, burning tokens, taking actions — without getting any closer to done. A multi-agent crew has a stranger version of the same failure, and it's arguably worse, because nothing about it looks stuck. Every agent in the crew is making real, forward, individually defensible progress. The crew keeps shipping outputs. And the whole operation is still moving further from the goal you actually gave it with every step. Call it multi-agent goal drift: not a crew that stalls, but one that steadily walks off the original objective while every individual handoff looks like solid work.

This is the second companion post to failure modes in multi-agent teams, after split-brain. That post's third failure mode, context fragmentation, sits close to this one and is worth separating cleanly. Context fragmentation is a single lossy event — the goal gets split across agents at one handoff, and a load-bearing detail falls into the gap between two context windows during that one split. Goal drift isn't a single event, it's compounding. It shows up in crews that run several sequential handoffs — a planner to a builder to a reviewer to a closer, or the same handoff shape looped across iterations — where each individual hop reshapes the goal by an amount too small to flag on its own, and only the sum, after enough hops, is unrecognizable. Split-brain, for comparison, is agents disagreeing on facts — two copies of a store diverging on what the current state is. Goal drift is agents agreeing perfectly on the facts and quietly disagreeing, without anyone noticing, on what they're for.

It's also adjacent to multi-agent context drift — the information a crew holds about a fixed objective going stale or inconsistent across agents — but that's a claim about the crew's beliefs. This is a claim about the crew's target. A crew can have perfectly fresh, perfectly consistent context about the wrong goal, which is exactly what happens below.

Three agents build a feature end to end: a planner turns a user request into a spec, a builder implements it, and a closer reviews the diff and merges. The request: "Add a cache in front of the search endpoint to cut p99 latency — do not change the public API contract." The planner writes a reasonable spec: add a cache layer, wire it into the search handler, add invalidation on writes, expose a metric for hit rate. Each subtask inherits the tasks from that decomposition. It does not inherit the constraint. "Do not change the API contract" never makes it past the planner's own reasoning, because it isn't a task — it's a boundary, and boundaries don't fit cleanly into a task list.

The builder implements the cache. To make it debuggable, it adds an optional ?cache=bypass query parameter — a small, sensible, entirely well-reasoned addition, and a genuine change to the public contract. The builder's own acceptance criteria — tests pass, cache works, hit rate is measurable — never mention the contract, because that criterion lived only in the original request, two hops upstream, and nothing carried it forward. The closer reviews the diff against the builder's stated intent — "add a debug bypass parameter" — and approves it, because relative to that spec, the change is correct. Ship it. The contract broke. Every agent along the way did defensible work. The diff that actually broke the promise passed review, because review was checking the wrong reference point: the nearest upstream spec, not the original ask.

The telephone-game framing undersells it, because telephone-game distortion is noise — random, undirected, roughly symmetric. Goal drift in an agent crew is closer to genetic drift than to noise: it's directional, and it accumulates, because at every hop the crew has a structural reason to shed anything that isn't a concrete, checkable task. A boundary condition like "don't change the contract" has no natural home in a task list, a diff, or a pass/fail test — it's exactly the kind of constraint each individual handoff is optimized to lose, not by accident but because it's the informationally cheapest thing to drop. Multiply that by several sequential hops and the loss compounds in one direction: away from the constraints, toward whatever's easiest to state as a concrete next action. No single hop looks unreasonable. The sum is a different problem than the one you asked for.

Four things, and they compound:

You cannot catch this by checking whether each handoff's own acceptance criteria passed — every hop above passes its own bar. You catch it by instrumenting the distance from the origin, not the local pass/fail:

Signal What it means How to catch it
Constraint survival The original ask's non-negotiables are still stated, verbatim, in the current hop's spec Diff each downstream spec's literal constraint list against the root request's; a constraint present at hop zero and absent at hop N is drift, not a rounding error
Distance from origin The working objective has moved away from the original ask, not just been refined Embed the current spec and the original ask; track the distance across hops and flag a monotonic upward slope, not just one reading
Acceptance criteria without provenance A subtask's "definition of done" doesn't cite which root constraint it protects Require every subtask spec to tag which top-level constraints it's responsible for; an untagged constraint is one nobody's watching
Review against the wrong reference The final check compares output to the nearest upstream spec instead of the original request Run one check — cheap, separate from the crew — that diffs delivered output directly against the root ask, never against a downstream derivation of it

Goal drift doesn't look like failure while it's happening, which is what makes it worse than loop drift: a stalled agent at least stops producing convincing output, while a drifting crew never stops — it just keeps shipping competent work that's steadily less connected to what you asked for. It isn't split-brain: the crew can agree completely on every fact and still drift, because the thing diverging isn't the state, it's the target. And it isn't a single context-fragmentation event from failure modes in multi-agent teams — it's what fragmentation looks like accumulated over many hops instead of one. If your crew runs more than a hop or two deep, assume the constraints from your original ask are decaying at every one of them, and go check — not by re-reading any single agent's transcript, which will look reasonable, but by diffing what's about to ship against the literal request you started with.

── more in #ai-agents 4 stories · sorted by recency
── more on @loop & retry 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/goal-drift-how-a-mul…] indexed:0 read:6min 2026-10-02 · —