JevBench, a reproducible benchmark for typed decision models TypeSafe AI's Jev 1.13.0 topped the JevBench leaderboard for typed decision models with a score of 74.4, ahead of Theodore Lee's SemIf (73.1) and Maisa's djev (73.0), according to the published benchmark table. Jev 1.13.0 posted 85.7 on one subscore, 82.7 on another, 83.3 on a third, 52.0 on a fourth, a cost of $0.040, and 0.65 s raw latency with 0.72 s raw p95 on a production API. OpenAI's GPT-5.6 Luna, run at low reasoning effort, ranked 14th at 65.9 with a $0.242 cost, the highest listed price in the table. | 1 | by TypeSafe AI Jev 1.13.0 https://docs.typesafe.ai | 74.4 | 85.7 | 82.7 | 83.3 | 52.0 | $0.040 | 100.0% | 99.0% | 94.5% | 74.1% | 0.65 s rawp95 0.72 s raw | production API | |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | 2 | by Theodore Lee TheoLeeCJ SemIf https://github.com/TheoLeeCJ/openjev formerly OpenJev Qwen3.5-4B, TheoLeeCJ | 73.1 | 79.0 | 72.6 | 83.7 | 59.5 | ~$0.022 est. | 100.0% | 97.9% | 95.2% | 59.5% | 0.20 s raw→ 0.55 s adjustedp95 0.32 s raw → 0.78 s | our RunPod GPU | | 3 | by Maisa David Villalón djev https://github.com/Davipar/djev-dev