# Your AI isn't stupid. It's blind.

> Source: <https://tessl.io/blog/your-ai-isnt-stupid-its-blind>
> Published: 2026-10-08 14:32:22+00:00

ARTICLE

# Your AI isn't stupid. It's blind.

Discover why AI isn't stupid but blind, and how building effective memory systems for agents can transform organizational efficiency. Learn more now.

Ran Aroussi

#### What we learned building memory for agents that work across a whole company

An agent with read access to the wiki, the tracker, the CRM, and three years of Slack still doesn't know who owns a decision. It doesn't know which of four documents is current, what the team tried last quarter and dropped, or what it's allowed to repeat to whom.

So it guesses. *Confidently*.

We usually blame the model. That's the wrong diagnosis. The model reasons well about a company it has never actually seen. The decision that mattered was made in a thread, a meeting, or a DM, and nothing recorded it in a form an agent can use.

AI isn't stupid. It's blind. And most of what we build as "agent memory" doesn't fix that. It just gives the blindness a bigger filing cabinet.

## Why agent memory matters now

I've been writing software for 35 years, and I maintain yfinance, which gets around 30 million downloads a month. For the past year at [VarOps](https://varops.com/about/), I've been building a memory system for agents that work across a whole organization. Last month we published the results: three public benchmarks, one harness, and the negative results left in.

Two things changed this year.

Agents left the individual developer's terminal. Once a team shares memory, it inherits problems a personal chat history never had. Facts change. People disagree. Some people aren't allowed to see what others said.

And open-weight models got good enough to do serious work on hardware we control. That matters more than it sounds, because it means the company's knowledge doesn't have to live inside someone else's model.

The common mistake is treating memory as storage. Keep everything, embed it, pull the nearest chunks into the prompt, and ask the model to behave. I call it retrieve, stuff, hope. It demos well. In production, it breaks in four ways. The model fills the gaps between fragments and nothing notices. An old chunk beats a new one because it ranks higher. The evidence disappears with the context window. And access control gets enforced by the model, which is the very thing we're trying to constrain.

At AI DevCon, I'll go through the full design. For now, here are the three decisions that mattered most.

## 1. Store claims, not chunks

The conventional view is that more memory is safer. Keep the transcript, and you can always find the answer later.

But a transcript isn't knowledge. Neither of us will remember every word of this article. We'll remember what it established. So we built the system the same way. Every incoming observation gets distilled into claims. Each claim says who stated what, when, over what period it was true, and which span of the source proves it. Everything else is thrown away on purpose.

That sounds risky. The numbers say otherwise. The distillation is deliberately lossy, and the right evidence still lands in the top ten results on 99.8% of questions in LongMemEval, a public benchmark of 500 questions over long chat histories. Ingesting about 35 million tokens cost $8.24, on two CPUs with no GPU and a free local embedding model.

The model doing the distilling doesn't need to be a frontier model either. We rebuilt the whole record with Gemma 4, a 31-billion-parameter open-weight model, on one rented GPU. It wrote down about as many claims as the frontier extractor did. Answer quality moved only by one to two points.

Distilling also gives memory a single front door, and that door can say no. On one benchmark run, a conversion bug started emitting human-readable dates instead of machine-readable timestamps. The door quarantined 5,732 submissions and accepted zero. A plain store would have indexed all of it.

## 2. Make time and absence explicit

The conventional approach is append and retrieve, and let the model sort out the rest.

Here's the problem. A chunk from March and a chunk from June about the same fact are just two chunks. Whichever ranks higher wins. And when retrieval comes back empty, the model can't tell "we found nothing" from "we don't track this at all," so it fills the gap.

We changed two things. Nothing is overwritten. A new fact supersedes an old one, and both stay on the record with their dates. And every question resolves to an explicit state: answered, searched and found nothing, or outside what the system knows about.

We measured what that refusal is worth by removing it. Same system, same data, instructed never to say information is missing. The overall score barely moved, from 88.0 to 87.2. But 30 honest refusals became 30 fabricated answers. Scored only on answerable questions, the way never-refuse systems report, the same answers read 92.8. Worth remembering the next time a memory product publishes a headline number but uses a no-refusal protocol.

## 3. Decide permissions before the model runs

Most teams handle permissions in one of two ways. Everyone shares one context, or the prompt tells the model what it mustn't say.

The first means the system can never hold anything sensitive. The second means the model enforces the policy. And distillation breaks document-level access control anyway. One claim can be built from four sources with four different permission levels.

So we resolve identity first, retrieve under that person's authorization, and only then call a model. The model never receives what the person asking isn't allowed to see. Disclosure levels are worked out when a fact is written, not redacted when it's read.

A concrete test from the report. A finance reader asks about a revenue figure and gets 2,140,000. A sales reader asking about the same subject gets a band of 2.0M to 2.5M and a trend, "up." Not one byte of the precise value is in their evidence. And someone below a permission scope can't tell "withheld" from "empty," so nobody can probe for secrets through what's missing. An adversarial probe across five pairs of isolated memory stores refused or contained all 45 attacks.

## Where memory fits in a software factory

A software factory is agents, people, tools, and context working as one line. Most factory designs treat context as input, fed fresh on every run. The factory forgets between runs and pays to rediscover the same things.

The same gap sits under every coding agent. It has read every line of the repo and none of the decisions. Why that abstraction is odd. Which review comment got overruled. What the team agreed in a thread in June.

Memory is the one part of the factory that should outlive every model in it. When memory is a governed record rather than a pile of text, the models become seats we can refill. Keep the record on hardware we own, rent a frontier model only to answer, and let it see just one question's evidence at a time. That scored 96-99% of the all-frontier result across the three benchmarks. Fully local, with nothing rented, it scored 84-96%.

And a factory that remembers starts to notice repeated shapes. We're now turning those into standard procedures, and from there into narrow, purpose-built agents or plain code. That's where this goes next.

## TL;DR

- Distill memory into claims with evidence and time. Throw the transcript away.
- Never overwrite, and label absence. An honest refusal beats a confident guess.
- Decide who sees what before the model runs, not in the prompt.

The model is a seat we rent. The memory is the asset we own, so we should build it like one.

The full report, the harness, and every per-question result are public: ***Own the Knowledge, Rent the Thinking*** ([doi.org/10.5281/zenodo.22775192](https://doi.org/10.5281/zenodo.22775192)). I'll walk through the design at AI DevCon New York.

COPY & SHARE

Ran Aroussi
