cd /news/ai-research/open-call-test-your-agent-memory-lay… · home topics ai-research article
[ARTICLE · art-119040] src=discuss.huggingface.co ↗ pub= topic=ai-research verified=true sentiment=· neutral

Open call: test your agent memory layer on an adversarial coding benchmark

A developer has released AMB, an open and preregistered benchmark that tests whether memory layers in coding agents improve real coding work, using executable tests to grade artifacts in a sandboxed repository. The public feed contains 195 session transcripts, and a hard corpus scales to 4,900 documents with roughly 143,000 chunks. The benchmark is adversarially constructed with absent, stale, superseded, contradictory, adjacent, and irrelevant memories, and small memory vendors are invited to test their systems at no cost.

read1 min views1 publishedSep 2, 2026

I built AMB, an open and preregistered benchmark for memory layers in coding agents.

Most memory benchmarks test whether a system retrieves a relevant chunk. AMB tests whether that memory actually helps an agent complete real coding work, or causes a worse change.

Claude Code runs inside a sandboxed repository, and executable tests grade the resulting artifact. The benchmark is adversarially constructed with absent, stale, superseded, contradictory, adjacent, and irrelevant memories.

The current public feed contains 195 pre-authored session transcripts. A separate hard corpus record scales to 4,900 documents, including 196 real documents and 4,704 synthetic distractors, with roughly 143,000 chunks.

I am inviting small memory vendors to test their systems. There is no participation fee. Vendors can review the adapter contract, run with their own credentials, and publish their configuration and results. Negative results are welcome.

The current suite mainly measures the read path. It does not yet measure whether a memory layer learns from the agent’s own work across sessions, and no multi product ranking has been published.

Repository:

What would you want this benchmark to measure before trusting its results?

── more in #ai-research 4 stories · sorted by recency
── more on @amb 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/open-call-test-your-…] indexed:0 read:1min 2026-09-02 ·