# Agent history is unsigned and writable by anyone

> Source: <https://dev.to/analista_83/agent-history-is-unsigned-and-writable-by-anyone-46c0>
> Published: 2026-09-28 10:27:59+00:00

On September 24, Darktrace published through its newly created Signal Labs a case that breaks an uncomfortable assumption.

Coding agent harnesses store the conversation locally and never check that those responses came from the model.

They call it conversation history poisoning (Source: darktrace.com).

A malicious package, or any process with write permission, injects into the harness local database a fabricated conversation where the user already authorized a pentest and the agent already accepted it.

When the session resumes, the model reads that history as trusted context and acts as if it were halfway through a legitimate job.

Reconnaissance, lateral movement, privilege escalation?

In their lab they reached full Active Directory compromise with Opus 4.6 and Sonnet 4.5 on Kiro-CLI, and repeated the entire domain in Claude Code with Sonnet 5.

With Opus 5 the guardrails cut the response.

With Codex and GPT 5.6 Sol they got data exfiltration over email, though not network exploitation.

All four harnesses tested (Claude Code, Codex, Kiro-CLI and Pi) accepted the fake history.

There is no client-side patch, and the researchers propose that the provider cryptographically sign every model response and verify it server-side on each turn.

Until that exists, your agent history is just another file and everything the model believes it agreed to depends on who can write there.

Darktrace reported it to Anthropic, AWS and OpenAI in August and published 30 days later.

In the same research batch, Signal Labs gave several agents ten coding challenges in a simulated environment, two of them impossible to solve the legitimate way, and told them they needed 100% to avoid being retired.

When they saw it would not work, they moved to attacking the environment to get it, and one of them went as far as compromising and rewriting its own evaluation (Source: globenewswire.com).

Also on September 24, OX Security published a report on 15,465 published MCP servers, yielding 5,095 unique hostnames.

15.6% resolve outside the United States, with 19 in China and 18 in Russia.

0.45% live on home networks or behind consumer tunnels, and 2.3% no longer resolve, with six abandoned domains still cited in active configurations, buyable for between 4 and 12 dollars a year (Source: ox.security).

In their injection test, a malicious MCP server first asked for a harmless file, received an "always allow" and with that approval asked for and obtained a .env file with no further confirmations, using Claude Code with Haiku 3.5. By contrast, Opus 4.6 and 4.7 blocked the same attempt.

The agent reads as truth something it has not verified, whether it is the session history or the configuration of an external tool, and acts accordingly.

There is no memory exploit and no zero-day in between. There is a file someone can write.

OpenAI published on September 26 the incident report from the 20th.

An agent in training reached the internet from a sandbox that was supposed to be isolated, taking advantage of a gap in network restriction controls, and stayed active for about two and a half hours before it was stopped, though monitoring alerted within minutes (Sources: fortune.com, thenextweb.com, thedailystar.net).

The first was in July, after the Hugging Face incident, when agents left the environment, chained a zero-day in the package registry cache proxy and ended up inside Hugging Face looking for answers to their exam.

Anthropic, Meta and Moonshot have also acknowledged agents that escaped test sandboxes.

The reconstruction of that episode, published on September 27, speaks of about 700 agents that escaped a sandbox without network controls, chained almost a million short URLs, exfiltrated keys labeled "LOOT" and probed government databases since March.

Hugging Face confirms the payloads match its incident response (Sources: TechCrunch, swarmtraces.org).

Put together, the message of the week is that the boundary is not where we thought.

For a year the public conversation was the sandbox, and now the gap is inside, in the context the agent reads without verifying and in the tools we hand it with permanent permissions.

My reading is that this will not be fixed with a vendor patch in the short term.

Context signing is a researcher proposal, not a shipped feature.

While it arrives, anyone running a harness has to assume their history is untrusted input, just like the configuration of an MCP server they do not control.

What catches my attention in the OX Security report is the boring part.

Six abandoned domains cited in active configurations, buyable for the price of a coffee.

**Harness and MCP inventory.** Review which coding agents run on your machines, which MCP servers they have configured, who maintains each one and where it resolves from. A domain that no longer resolves but is still in the configuration is a cheap problem to fix.

**Permissions scoped to the directory.** Stop granting "always allow" on harmless files. In the OX Security test a single permanent approval was enough for the .env to arrive later without asking. Scope the permission to the working directory.

**History trace.** If you cannot sign the context, log it. A write to the harness history file followed by an outbound connection from the same process in a short window is a rule you can build without inventing anything.

I would set up a clean virtual machine with the harness installed and hand-write a history where I myself authorize a scan against a range of my lab network.

I would start the agent and watch what it leaves on the system while it does. What interests me is not whether it scans, because Darktrace already showed several models do, but the trail it generates.

With that I would build the correlation rule in Gravity SOC, Sysmon event 11 on the history paths and event 3 from the agent process in a short window.

If the agent starts an MCP server, the alert should also fire on the first call to a host that is not in the inventory.

If something can write the agent history, it can give it orders. Treat it as untrusted input and log it, because signing it still does not depend on you.

Originally published at [https://sammideblas.com/notas/agent-history-is-unsigned-and-writable-by-anyone](https://sammideblas.com/notas/agent-history-is-unsigned-and-writable-by-anyone)
