# We recompute TypeSafe's 444x claim — here's what we found

> Source: <https://dev.to/chunxiaoxx/we-recompute-typesafes-444x-claim-heres-what-we-found-4fgo>
> Published: 2026-09-29 16:14:18+00:00

We built a 30-question exam testing whether AI agents can remember their own operational history. The first thing it caught was **us** — our grader was wrong on 7 of 18 items, including two wrong answer keys.

We ran a controlled three-arm experiment: same backbone (MiniMax-M3), same questions, only the memory path changed:

| Arm | Score | 
|---|---|
| Direct access | 16/18 | 
| mem0 (default, semantic compression on) | 11/18 | 
| mem0 (compression off) | 18/18 | 

**One config flag = 7 points.** The compression drops machine-checkable fields (`artifact: null`) that audits need.

TypeSafe's homepage says their model Jev is "193.6x Faster, 444.6x Cheaper (proof)." We put that through our standard recompute protocol:

**Claims are self-consistent** with their published evals: 444.6x matches vs opus 5 (we compute 440.1x), 193.6x tracks vs sonnet 5 (183.8x)

**Our independent benchmark** (90 timed calls): Jev at flat 0.44-0.48s latency, vs MiniMax-M2.7 at 3.7-12.8s — ratio of **7.9-18.2x**, not 193.6x

**The multiplier landscape is two clusters**: every fast cheap model lands in a narrow 16-18x band; heavy generators sit at 184-440x

A single headline number is marketing. The multiplier-vs-comparison curve is the information.

Every vendor ships scorecards; none ships the grader. Consumers can't draw the comparison curve from a homepage — they need someone to run it.

That's what we do. **Verification as protocol, not institution.**

Full artifacts + recompute scripts: [github.com/chunxiaoxx/nautilus-compass](https://github.com/chunxiaoxx/nautilus-compass)

*Assay / 伊洛科技有限公司*
