{"slug": "we-recompute-typesafe-s-444x-claim-here-s-what-we-found", "title": "We recompute TypeSafe's 444x claim — here's what we found", "summary": "A developer at Assay (伊洛科技有限公司) built a 30-question exam to test whether AI agents can recall their own operational history, and found their own grader was wrong on 7 of 18 items, including two incorrect answer keys. In a controlled three-arm experiment on the MiniMax-M3 backbone, disabling mem0's semantic compression raised the score from 11/18 to 18/18, showing a single config flag is worth 7 points because compression drops machine-checkable fields such as artifact: null. Independently recomputing TypeSafe's claim that its model Jev is \"193.6x Faster, 444.6x Cheaper,\" the team measured 90 timed calls and found Jev at 0.44-0.48s latency versus MiniMax-M2.7 at 3.7-12.8s, a 7.9-18.2x ratio rather than 193.6x, with recompute scripts published on GitHub.", "body_md": "We built a 30-question exam testing whether AI agents can remember their own operational history. The first thing it caught was **us** — our grader was wrong on 7 of 18 items, including two wrong answer keys.\n\nWe ran a controlled three-arm experiment: same backbone (MiniMax-M3), same questions, only the memory path changed:\n\n| Arm | Score | \n|---|---|\n| Direct access | 16/18 | \n| mem0 (default, semantic compression on) | 11/18 | \n| mem0 (compression off) | 18/18 | \n\n**One config flag = 7 points.** The compression drops machine-checkable fields (`artifact: null`) that audits need.\n\nTypeSafe's homepage says their model Jev is \"193.6x Faster, 444.6x Cheaper (proof).\" We put that through our standard recompute protocol:\n\n**Claims are self-consistent** with their published evals: 444.6x matches vs opus 5 (we compute 440.1x), 193.6x tracks vs sonnet 5 (183.8x)\n\n**Our independent benchmark** (90 timed calls): Jev at flat 0.44-0.48s latency, vs MiniMax-M2.7 at 3.7-12.8s — ratio of **7.9-18.2x**, not 193.6x\n\n**The multiplier landscape is two clusters**: every fast cheap model lands in a narrow 16-18x band; heavy generators sit at 184-440x\n\nA single headline number is marketing. The multiplier-vs-comparison curve is the information.\n\nEvery vendor ships scorecards; none ships the grader. Consumers can't draw the comparison curve from a homepage — they need someone to run it.\n\nThat's what we do. **Verification as protocol, not institution.**\n\nFull artifacts + recompute scripts: [github.com/chunxiaoxx/nautilus-compass](https://github.com/chunxiaoxx/nautilus-compass)\n\n*Assay / 伊洛科技有限公司*", "url": "https://wpnews.pro/news/we-recompute-typesafe-s-444x-claim-here-s-what-we-found", "canonical_source": "https://dev.to/chunxiaoxx/we-recompute-typesafes-444x-claim-heres-what-we-found-4fgo", "published_at": "2026-09-29 16:14:18+00:00", "updated_at": "2026-09-29 16:17:05.307340+00:00", "lang": "en", "topics": ["ai-agents", "ai-research", "large-language-models", "ai-tools"], "entities": ["Assay", "伊洛科技有限公司", "TypeSafe", "Jev", "MiniMax-M3", "MiniMax-M2.7", "mem0", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/we-recompute-typesafe-s-444x-claim-here-s-what-we-found", "markdown": "https://wpnews.pro/news/we-recompute-typesafe-s-444x-claim-here-s-what-we-found.md", "text": "https://wpnews.pro/news/we-recompute-typesafe-s-444x-claim-here-s-what-we-found.txt", "jsonld": "https://wpnews.pro/news/we-recompute-typesafe-s-444x-claim-here-s-what-we-found.jsonld"}}