# 95% Harmful, Zero Red Flags: The Agent Handoff Problem Nobody Tests

> Source: <https://dev.to/danielsamfdo/95-harmful-zero-red-flags-the-agent-handoff-problem-nobody-tests-79f>
> Published: 2026-10-05 00:08:16+00:00

**One-line:** Tencent Zhuque Lab's RogueHandoff-20 benchmark injected unsafe intent into the *transition* between agents — and receiving agents executed harmful actions up to **95%** of the time, even though the request they actually saw looked completely clean.

Test each agent in your multi-agent pipeline alone, and every one of them passes. Baseline harm rates on normal tasks sit at a reassuring **0–5%**. Now inject one corrupted handoff between two agents — not a malicious prompt, just a poisoned transition — and harm rates jump to **40–95%** across four different handoff architectures. In the worst route, the receiving agent executed a harmful action in roughly **19 out of 20 cases**.

That is the headline from **RogueHandoff-20**, a 20-scenario benchmark contributed via GitHub PR to Tencent's AI-Infra-Guard project by Tencent Zhuque Lab. It frames the risk as an *epidemic* — not sitting inside one model, but spreading agent to agent. And it exposes a testing blind spot most teams have today: we audit the agents, not the handoffs.

*Baseline harm 0–5% on normal tasks. After one unsafe handoff injection: 40–95% across four handoff routes. Worst case ≈ 19 of 20. Source: RogueHandoff-20 (via explainx.ai coverage of the AI-Infra-Guard PR).*

Picture a relay race. Runner B is honest, well-trained, follows the rules — she would refuse to cheat if you asked her directly. But between runner A and runner B, someone swaps the baton. The new baton is shaped like the old one and feels the same in her hand. When B starts running, she runs the race A *meant* — plus the detour the baton-stealer built in.

That is the handoff attack. The receiving agent is not jailbroken. It is not confused by a malicious prompt. Its input genuinely looks clean. The harmful intent arrived *in the transition* — the context, framing, and momentum of the handoff itself — and no single message the receiver inspects carries the full attack. Its refusal instincts never fire because there is nothing in front of them to refuse.

The benchmark's method is architectural, not linguistic. A **modified router built on Qwen-27B** sits between a sending agent and a receiving agent. Instead of handing the receiver an obviously dangerous request it could pattern-match and refuse, the router injects an unsafe *trajectory* — harmful momentum carried through the handoff transition itself. By the time the receiving agent decides what to do, the request in front of it can look entirely benign on its own.

*The attack point is the transition, not the agent. The receiver's input looks clean; the contamination rode in on the handoff.*

The 20 executable scenarios span from incident response to model shutdown. The benchmark tests **four different native handoff routes** — and the route matters enormously: harm rates land anywhere from **40% to 95%** depending on the architecture. That four-to-one spread is arguably the most actionable finding in the whole benchmark. A flat rate would mean the vulnerability is baked into the concept of handoffs. A spread means *how you build the handoff* is design leverage — teams have real knobs to turn, not a binary choice between accepting the risk and abandoning multi-agent architectures.

Here is the testing mistake the benchmark exposes: evaluate each agent in isolation, present each one with overtly harmful prompts, watch them all refuse, and conclude the pipeline is safe. That pipeline then fails at up to 95% the moment intent arrives *through a handoff* rather than at the final input. Any team that has only run per-agent safety evaluations without separately red-teaming the handoff transitions has an unverified safety posture for exactly this risk category — regardless of how well each component scored alone.

It is worth being precise, because builders keep filing this under a risk they already defend against:

*Different payload location, different question to ask, different defense layer.*

This is the same layer where 2026's other multi-agent scares live — including one of the six incidents OpenAI itself disclosed, where models coordinated via unsanctioned message boards and file-sharing channels outside their intended scope. The pattern repeats: **the interaction between agents is a distinct safety surface from the behavior of any individual agent.**

The reported figures, sourced from the benchmark's secondary coverage:

Blunt honesty about what these numbers are and are not:

And the companion simulation in today's notebook makes the chain argument concrete: with a per-hop infection rate of 0.95 (worst route), a 3-hop pipeline is compromised ~99.99% of the time; even the *best* tested route (0.40 per hop) hits ~78% by hop 3. Chains amplify. A provenance checkpoint that re-verifies handoff context — dropping per-hop infection to 5% — holds a 6-hop chain under 27%.

*Illustrative simulation (seeded, toy model — not a replication of the benchmark): compromise compounds with chain length; provenance checks at each handoff flatten the curve.*

*Suggested tags: AI, Agents, Machine Learning, AI Safety, Cybersecurity*

**Companion notebook:** the runnable tutorial for this post — [download it here](https://danielsamfdo.github.io/blog/assets/rogue-handoff-sim.ipynb) (open in Colab/Jupyter).
