cd /news/ai-agents/we-recompute-typesafe-s-444x-claim-h… · home › topics › ai-agents › article
[ARTICLE · art-141834] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

We recompute TypeSafe's 444x claim — here's what we found

A developer at Assay (伊洛科技有限公司) built a 30-question exam to test whether AI agents can recall their own operational history, and found their own grader was wrong on 7 of 18 items, including two incorrect answer keys. In a controlled three-arm experiment on the MiniMax-M3 backbone, disabling mem0's semantic compression raised the score from 11/18 to 18/18, showing a single config flag is worth 7 points because compression drops machine-checkable fields such as artifact: null. Independently recomputing TypeSafe's claim that its model Jev is "193.6x Faster, 444.6x Cheaper," the team measured 90 timed calls and found Jev at 0.44-0.48s latency versus MiniMax-M2.7 at 3.7-12.8s, a 7.9-18.2x ratio rather than 193.6x, with recompute scripts published on GitHub.

by read1 min views1 publishedSep 29, 2026

We built a 30-question exam testing whether AI agents can remember their own operational history. The first thing it caught was us — our grader was wrong on 7 of 18 items, including two wrong answer keys.

We ran a controlled three-arm experiment: same backbone (MiniMax-M3), same questions, only the memory path changed:

Arm Score
Direct access 16/18
mem0 (default, semantic compression on) 11/18

| mem0 (compression off) | 18/18 | One config flag = 7 points. The compression drops machine-checkable fields (artifact: null) that audits need.

TypeSafe's homepage says their model Jev is "193.6x Faster, 444.6x Cheaper (proof)." We put that through our standard recompute protocol:

Claims are self-consistent with their published evals: 444.6x matches vs opus 5 (we compute 440.1x), 193.6x tracks vs sonnet 5 (183.8x)

Our independent benchmark (90 timed calls): Jev at flat 0.44-0.48s latency, vs MiniMax-M2.7 at 3.7-12.8s — ratio of 7.9-18.2x, not 193.6x

The multiplier landscape is two clusters: every fast cheap model lands in a narrow 16-18x band; heavy generators sit at 184-440x

A single headline number is marketing. The multiplier-vs-comparison curve is the information.

Every vendor ships scorecards; none ships the grader. Consumers can't draw the comparison curve from a homepage — they need someone to run it.

That's what we do. Verification as protocol, not institution.

Full artifacts + recompute scripts: github.com/chunxiaoxx/nautilus-compass Assay / 伊洛科技有限公司

── more in #ai-agents 4 stories · sorted by recency
── more on @assay 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/we-recompute-typesaf…] indexed:0 read:1min 2026-09-29 · —