cd /news/ai-agents/mits-minimalist-jaz-agent-beats-lett… · home › topics › ai-agents › article
[ARTICLE · art-146788] src=cryptobriefing.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

MIT’s minimalist JAZ agent beats Letta and ACE on memory tasks

Researchers at MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) introduced JAZ, a minimalist agent framework that scored 70% on the StuLife long-term recall benchmark versus 62% for Letta at roughly half the cost, and 74% on the AppWorld self-improvement benchmark, 4 percentage points ahead of ACE. The paper, titled "Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity" and posted to arXiv as arXiv:2609.26891, exposes an agent's history and prompt as variables in a code environment so the model can write executable code to inspect and manipulate its own past, relying on a single LLM-based primitive called invoke. The framework and evaluation code are published on GitHub at jaz-lang/jaz and jaz-lang/jaz-evals, though the results cover only two benchmarks and the StuLife comparison ran on the single small model GPT-5.4 nano.

by read3 min views2 publishedOct 7, 2026
MIT’s minimalist JAZ agent beats Letta and ACE on memory tasks
Image: Cryptobriefing (auto-discovered)

Photo: Tara Winstead / Pexels

A new MIT CSAIL paper finds that letting an agent treat its own history as code variables can outperform dedicated memory systems at lower cost

Researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) have introduced JAZ, a stripped-down agent framework. In benchmark tests it beat Letta, formerly known as MemGPT, on recall-heavy tasks. It also topped ACE, a specialized self-improvement harness, while spending less money to do it.

Instead of giving a language model a filing cabinet, give it the ability to write code that reaches into its own past.

One primitive to rule them all #

The paper is titled “Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity.” It was posted to arXiv under the identifier arXiv:2609.26891.

At its core, JAZ relies on a single LLM-based primitive called invoke. That’s the whole toolkit, more or less.

JAZ takes a different route. The agent’s history and its prompt are exposed as variables inside a code environment. The model can then write executable code to inspect, slice, or manipulate that history directly.

The benchmark numbers #

The researchers tested JAZ on two fronts: long-term recall and self-improvement.

For recall, they used the StuLife benchmark and ran JAZ on the GPT-5.4 nano model. JAZ scored 70%. Letta scored 62%.

AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

That eight-point gap is notable on its own. The cost side makes it more interesting: JAZ reportedly got there at approximately half the cost of Letta.

Letta’s design is built around a stateful memory hierarchy for managing context. In plainer terms, it organizes what the model remembers into tiers, deciding what stays close at hand and what gets archived.

The second test focused on self-improvement, using the AppWorld benchmark. Here JAZ went up against ACE, a harness designed specifically for iterative self-improvement in CodeAct environments.

JAZ scored 74%, finishing 4 percentage points ahead of ACE. It did so at lower cost.

What this means for agent builders #

For developers building AI agents, the most practical takeaway is about cost. Getting a higher score at roughly half the price, as JAZ reportedly did against Letta, is the kind of result that gets attention in engineering budget meetings. When agents run thousands of tasks, per-task savings compound quickly. There are reasons for caution. The results cover two benchmarks, StuLife and AppWorld, and the StuLife comparison was run on a single small model, GPT-5.4 nano.

Giving a model the ability to write and run arbitrary code against its own history also raises its own engineering questions. Sandboxing, error handling, and predictability matter more when the agent is effectively programming its own memory access.

The team has published the framework and its evaluation code on GitHub. The repositories are jaz-lang/jaz and jaz-lang/jaz-evals, so other researchers can poke at the results themselves.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our

Editorial Policy.

── more in #ai-agents 4 stories · sorted by recency
── more on @mit csail 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mits-minimalist-jaz-…] indexed:0 read:3min 2026-10-07 · —