# Glasshouse v0.1 Is Out: A Memory Benchmark for AI Systems

> Source: <https://dev.to/woochan/glasshouse-v01-is-out-a-memory-benchmark-for-ai-systems-51h4>
> Published: 2026-09-22 07:34:24+00:00

Glasshouse v0.1 is out. It's a long-term memory benchmark for AI systems, and it's what my last two benchmark posts here were about. If you haven't read those, here's the short version.

Going through developer communities, I kept running into people raising the same problems with memory benchmarks. The numbers a vendor publishes don't match the numbers someone else measures, and swapping the model that does the grading moves the results more than the gap between the systems being compared.

If you read my earlier post, some of the numbers have moved since. That post said they would.

`submissions/` is empty. We haven't submitted either.

If you're an individual, run it, and open a pull request or an issue when something is wrong. There's no threshold for individuals, on purpose, and an objection that names a specific error gets answered in public. There's a file in the repo listing what people suggested on Reddit while v0.1 was being built, and what each suggestion became. The stale fact axis, the contradiction axis and the false memory probes all started as someone's comment. One suggestion wasn't used, and it's listed anyway, because a record that only shows what was taken can't be checked.

If you're a company, you can submit a result or add your company, and how that works is in the repo.

Everything, including how to run it, is here: [github.com/wontopos/glasshouse](https://github.com/wontopos/glasshouse)

If you run it and a number looks wrong to you, that's exactly what I want to hear.
