cd /news/ai-agents/do-coding-agents-need-memory-or-docu… · home › topics › ai-agents › article
[ARTICLE · art-145658] src=gethrbr.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Do coding agents need memory or documentation?

Research published in August and September 2026 found that agent memory systems which choose what to remember rarely beat no memory at all, according to five papers reviewed on 2026-10-05, including VibeMemBench (arXiv:2609.23570, 2026-09-20), in which eleven of twelve solver-and-system pairings failed to beat the same agent with memory off. A separate study (arXiv:2609.30813, 2026-09-25) found agents asserted an uncontested false belief in 97 to 99 percent of probes across four agent families, while gating on declared source type held false adoption to 6 to 9 percent versus 22 to 47 percent for other answering policies. The review followed Kevin Liao's post "Agents don't need memory, they need documentation," which reached the Hacker News front page on 2026-10-03 with 369 points and 281 comments.

by read8 min views1 publishedOct 5, 2026
Do coding agents need memory or documentation?
Image: Gethrbr (auto-discovered)

Blog

Published: October 5, 2026

Neither, as they are usually built. Research from August and September 2026 finds that memory systems which choose what to remember rarely beat no memory at all, and that whether the documentation agents read changes what they edit is still unresolved. What a team's agents need is the team's decisions: admitted through a gate, dated, retired when they stop being true, and counted when they are used.

The question reached the Hacker News front page on 2026-10-03 with Kevin Liao's Agents don't need memory, they need documentation (369 points and 281 comments by 10-05). His case: memory plugins retrieve snippets by similarity, without context, with stale facts treated as current, and nobody can audit them; a folder of markdown the agent reads before work and updates after is better. The thread pushed back from both sides. Below are the five papers that measure the question, read from their abstracts on 2026-10-05.

What is the difference between agent memory and documentation? #

Both are text an agent reads that it did not get from the code. They differ in who writes it, who checks it, and how it reaches the agent.

Agent memory Documentation
Who writes it The agent, from what it noticed in a session People, sometimes with an agent drafting
Who checks it Usually nobody before the next session reads it Whoever reviews the pull request
How it reaches the agent Retrieved by similarity to the current prompt Loaded whole (AGENTS.md, CLAUDE.md) or opened when the agent decides to
When it goes stale Often kept and marked invalid, or kept as is Stays until someone edits it
Typical examples Claude Code's auto memory, Mem0, Zep, Cognee AGENTS.md, CLAUDE.md, a docs/ folder, ADRs

Does agent memory help coding agents? #

The experience helps. The systems that pick it mostly do not. VibeMemBench (arXiv:2609.23570, 2026-09-20) built 111 coding targets from 90 real repositories, each paired with past experience that was verified to help. Injecting that experience directly raised the task resolution of four of five held-out solvers by 1.1 to 4.5 points and cut steps for all five. When four existing memory systems had to build and retrieve the experience from the same history themselves, eleven of twelve solver and system pairings failed to beat the same agent with memory off.

Read that carefully, because both camps on the thread can quote it. It does not say past experience is useless; the direct injection worked. It says the step where a system decides what to keep and what to hand back is where the value is lost.

What goes wrong when agents share memory across a team? #

Two things, both measured in September.

A false claim sticks. arXiv:2609.30813 (2026-09-25) ran multi-agent teams over a shared store. Once an uncontested false belief was in memory, the agent answering from it asserted that belief in 97 to 99 percent of probes, across four agent families. The admission policy decided how often false claims got in: gating on the declared source type held false adoption to 6 to 9 percent, against 22 to 47 percent for the other answering policies. Deduplicating sources rejected true claims along with the false ones.

A retired fact keeps winning. Revoked but Still Authoritative (arXiv:2609.08258, 2026-09-08) loaded five agent-memory systems with a revoked policy and its replacement. None enforced the revocation by default. The revoked fact came back wherever the retrieval layer could see it, outranked its replacement, and led agents to the unsafe action.

For one developer this is an annoyance. For a team it compounds: one session's mistaken note becomes every session's premise, and the correction sits next to it, losing. Agent memory poisoning covers the deliberate version of the same failure.

Is documentation enough for coding agents? #

It is the better half of the argument, and it has its own measured gaps.

Agents read it. Whether it changes the edit is open. arXiv:2608.20195 (2026-08-20) traced 557 agentic sessions and 33,097 agent pull requests. Instruction files and working notes made up 60.5 percent of the agents' documentation interactions, against 10.6 percent for classical technical docs and 1.3 percent for API references. The link from reading to editing was unresolved: the probability that a consultation was followed directly by an edit was 0.002, and only a stage-adjusted model put the association above chance. And docs trail code: in multi-commit pull requests that changed both, code was touched first 4.7 times more often.

A written rule is not an enforced one. When “Do Not” Is Not Deny (arXiv:2608.23550, 2026-08-24) took the security rules from 481 public CLAUDE.md files and checked each for a built-in Claude Code control that would enforce it. Depending on how close a match had to be, only about 4 to 16 percent had one; 4.4 percent under the strictest standard. The author calls CLAUDE.md a write-only channel: you write a rule and never learn whether anything enforces it.

Two objections in the thread line up with these numbers. One: a markdown folder has the unknown-unknowns problem Liao charges memory with, since the agent still has to guess which page matters before it starts. Two: whatever is written needs enforcing, because a rule such as using jq rather than an ad-hoc Python script to parse JSON gets broken anyway. Whether AGENTS.md helps at all is its own debate, with its own studies.

What should a team give its coding agents instead? #

Put the five papers side by side and they describe the same object from different angles. It is neither a memory store nor a docs folder. It is the set of decisions the team has made, held to five properties.

Property Why Evidence
A gate on what gets in An uncontested false claim in shared memory is repeated in 97 to 99 percent of probes 2609.30813
Retirement enforced when it is read A revoked fact that is still retrievable outranks its replacement 2609.08258
Delivered with the task The link from reading to editing is unresolved, and an agent cannot search for what it does not know exists 2608.20195
Hard rules as controls, not prose Only 4 to 16 percent of written security rules map to a control that enforces them 2608.23550
A count of what was used Eleven of twelve pairings failed to beat memory off. Without a count per item, you cannot see which items earn their place 2609.23570

Take one decision through each option. In #eng-payments the team agrees that webhook handlers acknowledge with a 200 and enqueue, and never do the work inline. As memory, it exists if an agent happened to be in a session where it came up, and it comes back if a later prompt happens to resemble that one. As documentation, it exists if someone copies it into AGENTS.md, and it stays there after the queue is replaced. As a decision, it is captured from the thread, approved, dated to the day it was said, served to the session that is editing a webhook handler, and retired when the queue it assumed is gone.

None of this argues against writing things down. Liao's consult-then-update loop is a good habit for one developer. It stops being enough when ten people and their agents write to the same folder and nobody owns what is true.

Where Harbor fits #

Harbor is built on that table. It reads where the team decides: Slack threads, pull request reviews and docs. A person approves what it finds, or a policy you set approves the routine ones, and a rule that contradicts a live one always waits for a person. Each fact is dated by the message it came from, a rule can carry an end date, and when a contradiction is settled for the newer rule, the old one is superseded and the sessions that used it are told what replaced it. Hooks give Claude Code, Codex and Cursor the rules that apply to the task at session start, and guardrails turn the hard ones into blocks.

It also keeps the fifth column. Every rule shows how often it was served and how often an answer cited it, so a rule nobody uses is visible instead of assumed. A cite is evidence that a rule was used, not proof that it was obeyed; the method is in served vs cited.

Questions #

Do coding agents need memory?

They benefit from past experience, but most memory systems lose the benefit. In VibeMemBench (arXiv:2609.23570), injecting verified experience directly helped four of five solvers, while eleven of twelve pairings of a solver with an existing memory system failed to beat memory off.

Is documentation better than memory for AI agents?

It is reviewable and auditable, which memory usually is not. But agents spend most of their documentation reading on instruction files, the link from reading to editing is unresolved (arXiv:2608.20195), and only 4 to 16 percent of security rules written in CLAUDE.md files map to a control that enforces them (arXiv:2608.23550).

What goes wrong with shared agent memory on a team?

A false claim that enters shared memory uncontested is repeated in 97 to 99 percent of probes (arXiv:2609.30813), and in five memory systems a revoked fact still outranked its replacement by default (arXiv:2609.08258). Without a gate and enforced retirement, one session's mistake becomes every session's premise.

What should a team give its coding agents instead of memory?

The decisions the team made, admitted through a gate, dated, delivered with the task, retired when they stop being true, with the hard ones enforced as controls, and counted when they are used.

── more in #ai-agents 4 stories · sorted by recency
── more on @kevin liao 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/do-coding-agents-nee…] indexed:0 read:8min 2026-10-05 · —