# Do coding agents need memory or documentation?

> Source: <https://gethrbr.com/blog/do-agents-need-memory>
> Published: 2026-10-05 00:00:00+00:00

Blog

# Do coding agents need memory or documentation?

Published: October 5, 2026

Neither, as they are usually built. Research from August and September 2026 finds that memory systems which choose what to remember rarely beat no memory at all, and that whether the documentation agents read changes what they edit is still unresolved. What a team's agents need is the team's decisions: admitted through a gate, dated, retired when they stop being true, and counted when they are used.

The question reached the Hacker News front page on 2026-10-03 with Kevin Liao's [Agents don't need memory, they need documentation](https://liao.gg/blog/agents-dont-need-memory) ([369 points and 281 comments](https://news.ycombinator.com/item?id=49945933) by 10-05). His case: memory plugins retrieve snippets by similarity, without context, with stale facts treated as current, and nobody can audit them; a folder of markdown the agent reads before work and updates after is better. The thread pushed back from both sides. Below are the five papers that measure the question, read from their abstracts on 2026-10-05.

## What is the difference between agent memory and documentation?

Both are text an agent reads that it did not get from the code. They differ in who writes it, who checks it, and how it reaches the agent.

|  | Agent memory | Documentation | 
|---|---|---|
| Who writes it | The agent, from what it noticed in a session | People, sometimes with an agent drafting | 
| Who checks it | Usually nobody before the next session reads it | Whoever reviews the pull request | 
| How it reaches the agent | Retrieved by similarity to the current prompt | Loaded whole (AGENTS.md, CLAUDE.md) or opened when the agent decides to | 
| When it goes stale | Often kept and marked invalid, or kept as is | Stays until someone edits it | 
| Typical examples | Claude Code's auto memory, Mem0, Zep, Cognee | AGENTS.md, CLAUDE.md, a docs/ folder, ADRs | 

## Does agent memory help coding agents?

The experience helps. The systems that pick it mostly do not. [VibeMemBench](https://arxiv.org/abs/2609.23570) (arXiv:2609.23570, 2026-09-20) built 111 coding targets from 90 real repositories, each paired with past experience that was verified to help. Injecting that experience directly raised the task resolution of four of five held-out solvers by 1.1 to 4.5 points and cut steps for all five. When four existing memory systems had to build and retrieve the experience from the same history themselves, eleven of twelve solver and system pairings failed to beat the same agent with memory off.

Read that carefully, because both camps on the thread can quote it. It does not say past experience is useless; the direct injection worked. It says the step where a system decides what to keep and what to hand back is where the value is lost.

## What goes wrong when agents share memory across a team?

Two things, both measured in September.

**A false claim sticks.** [arXiv:2609.30813](https://arxiv.org/abs/2609.30813) (2026-09-25) ran multi-agent teams over a shared store. Once an uncontested false belief was in memory, the agent answering from it asserted that belief in 97 to 99 percent of probes, across four agent families. The admission policy decided how often false claims got in: gating on the declared source type held false adoption to 6 to 9 percent, against 22 to 47 percent for the other answering policies. Deduplicating sources rejected true claims along with the false ones.

**A retired fact keeps winning.** [Revoked but Still Authoritative](https://arxiv.org/abs/2609.08258) (arXiv:2609.08258, 2026-09-08) loaded five agent-memory systems with a revoked policy and its replacement. None enforced the revocation by default. The revoked fact came back wherever the retrieval layer could see it, outranked its replacement, and led agents to the unsafe action.

For one developer this is an annoyance. For a team it compounds: one session's mistaken note becomes every session's premise, and the correction sits next to it, losing. [Agent memory poisoning](https://gethrbr.com/blog/agent-memory-poisoning) covers the deliberate version of the same failure.

## Is documentation enough for coding agents?

It is the better half of the argument, and it has its own measured gaps.

**Agents read it. Whether it changes the edit is open.** [arXiv:2608.20195](https://arxiv.org/abs/2608.20195) (2026-08-20) traced 557 agentic sessions and 33,097 agent pull requests. Instruction files and working notes made up 60.5 percent of the agents' documentation interactions, against 10.6 percent for classical technical docs and 1.3 percent for API references. The link from reading to editing was unresolved: the probability that a consultation was followed directly by an edit was 0.002, and only a stage-adjusted model put the association above chance. And docs trail code: in multi-commit pull requests that changed both, code was touched first 4.7 times more often.

**A written rule is not an enforced one.** [When “Do Not” Is Not Deny](https://arxiv.org/abs/2608.23550) (arXiv:2608.23550, 2026-08-24) took the security rules from 481 public `CLAUDE.md` files and checked each for a built-in Claude Code control that would enforce it. Depending on how close a match had to be, only about 4 to 16 percent had one; 4.4 percent under the strictest standard. The author calls `CLAUDE.md` a write-only channel: you write a rule and never learn whether anything enforces it.

Two objections in the thread line up with these numbers. One: a markdown folder has the unknown-unknowns problem Liao charges memory with, since the agent still has to guess which page matters before it starts. Two: whatever is written needs enforcing, because a rule such as using jq rather than an ad-hoc Python script to parse JSON gets broken anyway. Whether [AGENTS.md helps at all](https://gethrbr.com/blog/is-agents-md-useful) is its own debate, with its own studies.

## What should a team give its coding agents instead?

Put the five papers side by side and they describe the same object from different angles. It is neither a memory store nor a docs folder. It is the set of decisions the team has made, held to five properties.

| Property | Why | Evidence | 
|---|---|---|
| A gate on what gets in | An uncontested false claim in shared memory is repeated in 97 to 99 percent of probes | 2609.30813 | 
| Retirement enforced when it is read | A revoked fact that is still retrievable outranks its replacement | 2609.08258 | 
| Delivered with the task | The link from reading to editing is unresolved, and an agent cannot search for what it does not know exists | 2608.20195 | 
| Hard rules as controls, not prose | Only 4 to 16 percent of written security rules map to a control that enforces them | 2608.23550 | 
| A count of what was used | Eleven of twelve pairings failed to beat memory off. Without a count per item, you cannot see which items earn their place | 2609.23570 | 

Take one decision through each option. In `#eng-payments` the team agrees that webhook handlers acknowledge with a 200 and enqueue, and never do the work inline. As memory, it exists if an agent happened to be in a session where it came up, and it comes back if a later prompt happens to resemble that one. As documentation, it exists if someone copies it into `AGENTS.md`, and it stays there after the queue is replaced. As a decision, it is captured from the thread, approved, dated to the day it was said, served to the session that is editing a webhook handler, and retired when the queue it assumed is gone.

None of this argues against writing things down. Liao's consult-then-update loop is a good habit for one developer. It stops being enough when ten people and their agents write to the same folder and nobody owns what is true.

## Where Harbor fits

Harbor is built on that table. It reads where the team decides: Slack threads, pull request reviews and docs. A person approves what it finds, or a policy you set approves the routine ones, and a rule that contradicts a live one always waits for a person. Each fact is dated by the message it came from, a rule can carry an end date, and when a contradiction is settled for the newer rule, the old one is superseded and the sessions that used it are told what replaced it. Hooks give Claude Code, Codex and Cursor the rules that apply to the task at session start, and [guardrails](https://gethrbr.com/docs/guardrails) turn the hard ones into blocks.

It also keeps the fifth column. Every rule shows how often it was served and how often an answer cited it, so a rule nobody uses is visible instead of assumed. A cite is evidence that a rule was used, not proof that it was obeyed; the method is in [served vs cited](https://gethrbr.com/blog/served-vs-cited).

## Questions

### Do coding agents need memory?

They benefit from past experience, but most memory systems lose the benefit. In VibeMemBench (arXiv:2609.23570), injecting verified experience directly helped four of five solvers, while eleven of twelve pairings of a solver with an existing memory system failed to beat memory off.

### Is documentation better than memory for AI agents?

It is reviewable and auditable, which memory usually is not. But agents spend most of their documentation reading on instruction files, the link from reading to editing is unresolved (arXiv:2608.20195), and only 4 to 16 percent of security rules written in CLAUDE.md files map to a control that enforces them (arXiv:2608.23550).

### What goes wrong with shared agent memory on a team?

A false claim that enters shared memory uncontested is repeated in 97 to 99 percent of probes (arXiv:2609.30813), and in five memory systems a revoked fact still outranked its replacement by default (arXiv:2609.08258). Without a gate and enforced retirement, one session's mistake becomes every session's premise.

### What should a team give its coding agents instead of memory?

The decisions the team made, admitted through a gate, dated, delivered with the task, retired when they stop being true, with the hard ones enforced as controls, and counted when they are used.
