We recompute TypeSafe's 444x claim — here's what we found A developer at Assay (伊洛科技有限公司) built a 30-question exam to test whether AI agents can recall their own operational history, and found their own grader was wrong on 7 of 18 items, including two incorrect answer keys. In a controlled three-arm experiment on the MiniMax-M3 backbone, disabling mem0's semantic compression raised the score from 11/18 to 18/18, showing a single config flag is worth 7 points because compression drops machine-checkable fields such as artifact: null. Independently recomputing TypeSafe's claim that its model Jev is "193.6x Faster, 444.6x Cheaper," the team measured 90 timed calls and found Jev at 0.44-0.48s latency versus MiniMax-M2.7 at 3.7-12.8s, a 7.9-18.2x ratio rather than 193.6x, with recompute scripts published on GitHub. We built a 30-question exam testing whether AI agents can remember their own operational history. The first thing it caught was us — our grader was wrong on 7 of 18 items, including two wrong answer keys. We ran a controlled three-arm experiment: same backbone MiniMax-M3 , same questions, only the memory path changed: | Arm | Score | |---|---| | Direct access | 16/18 | | mem0 default, semantic compression on | 11/18 | | mem0 compression off | 18/18 | One config flag = 7 points. The compression drops machine-checkable fields artifact: null that audits need. TypeSafe's homepage says their model Jev is "193.6x Faster, 444.6x Cheaper proof ." We put that through our standard recompute protocol: Claims are self-consistent with their published evals: 444.6x matches vs opus 5 we compute 440.1x , 193.6x tracks vs sonnet 5 183.8x Our independent benchmark 90 timed calls : Jev at flat 0.44-0.48s latency, vs MiniMax-M2.7 at 3.7-12.8s — ratio of 7.9-18.2x , not 193.6x The multiplier landscape is two clusters : every fast cheap model lands in a narrow 16-18x band; heavy generators sit at 184-440x A single headline number is marketing. The multiplier-vs-comparison curve is the information. Every vendor ships scorecards; none ships the grader. Consumers can't draw the comparison curve from a homepage — they need someone to run it. That's what we do. Verification as protocol, not institution. Full artifacts + recompute scripts: github.com/chunxiaoxx/nautilus-compass https://github.com/chunxiaoxx/nautilus-compass Assay / 伊洛科技有限公司