cd /news/ai-research/jevbench-a-reproducible-benchmark-fo… · home topics ai-research article
[ARTICLE · art-137035] src=benchmarkheaven.com ↗ pub= topic=ai-research verified=true sentiment=· neutral

JevBench, a reproducible benchmark for typed decision models

TypeSafe AI's Jev 1.13.0 topped the JevBench leaderboard for typed decision models with a score of 74.4, ahead of Theodore Lee's SemIf (73.1) and Maisa's djev (73.0), according to the published benchmark table. Jev 1.13.0 posted 85.7 on one subscore, 82.7 on another, 83.3 on a third, 52.0 on a fourth, a cost of $0.040, and 0.65 s raw latency with 0.72 s raw p95 on a production API. OpenAI's GPT-5.6 Luna, run at low reasoning effort, ranked 14th at 65.9 with a $0.242 cost, the highest listed price in the table.

by read12 min views2 publishedSep 22, 2026
JevBench, a reproducible benchmark for typed decision models
Image: source
1 by TypeSafe AI Jev 1.13.0 74.4 85.7 82.7 83.3 52.0 $0.040 100.0% 99.0% 94.5% 74.1% 0.65 s rawp95 0.72 s raw production API
2 by Theodore Lee (TheoLeeCJ) SemIf formerly OpenJev (Qwen3.5-4B, TheoLeeCJ 73.1 79.0 72.6 83.7 59.5 ~$0.022 est. 100.0% 97.9% 95.2% 59.5% 0.20 s raw→ 0.55 s adjustedp95 0.32 s raw → 0.78 s our RunPod GPU
3 by Maisa (David Villalón) djev<sup></sup> Maisa, diffusion-gemma 73.0 82.7 65.4 91.4 57.6 $0.026 announced 100.0% 97.9% 93.2% 69.5% 0.24 s rawp95 0.31 s raw production API
4 by Eldan Ring Winnow-12B Q8<sup></sup> 71.2 82.0 72.0 82.3 52.9 ~$0.037 est. 100.0% 96.9% 91.1% 70.9% 0.23 s raw→ 0.60 s adjustedp95 0.41 s raw → 0.98 s our RunPod GPU
5 by kshetrajna12 reflex 4B<sup></sup> 70.3 80.1 75.2 68.0 59.7 ~$0.022 est. 100.0% 94.8% 97.3% 63.2% 1.80 s raw→ 3.75 s adjustedp95 2.05 s raw → 4.26 s our RunPod GPU
6 by hjmurmur (Octalab) jqv<sup></sup> Qwen3-32B zero-shot 68.6 79.3 79.0 74.6 47.5 ~$0.056 est. 100.0% 95.8% 92.5% 64.5% 0.75 s raw→ 1.64 s adjustedp95 0.97 s raw → 2.10 s our RunPod GPU
7 by milliseconds.ai (Baptiste Laget) decision-machine-1<sup></sup> milliseconds.ai 68.3 62.1 70.4 92.9 53.7 $0.035 100.0% 76.0% 89.7% 46.8% 0.17 s rawp95 0.30 s raw production API
8 by Mapika decider-35b-a3b<sup></sup> 67.6 79.6 71.5 80.8 45.3 ~$0.067 est. 100.0% 96.9% 91.1% 65.5% 0.29 s raw→ 0.73 s adjustedp95 0.49 s raw → 1.14 s our RunPod GPU
9 by IkerMoel open-alternative-jev<sup></sup> Qwen3.5-4B, IkerMoel 67.0 64.0 63.2 83.5 59.6 ~$0.022 est. 100.0% 84.4% 74.7% 56.8% 0.21 s raw→ 0.56 s adjustedp95 0.32 s raw → 0.80 s our RunPod GPU
10 by mithalouni system-one-open Gemma 4 E2B LoRA on an L4 66.6 69.5 56.7 77.0 64.8 ~$0.015 est. 100.0% 93.8% 87.7% 49.1% 0.65 s raw→ 1.30 s adjustedp95 0.77 s raw → 1.54 s author's demo server
11 by razorback16 / Codiv OpenJev DiffusionGemma 26B-A4B NVFP4, razorback16 66.4 79.2 64.8 83.2 45.5 ~$0.066 est. 100.0% 95.8% 91.1% 65.5% 0.24 s raw→ 0.63 s adjustedp95 0.31 s raw → 0.76 s our RunPod GPU
12 by Featherless AI SimpleJev Qwen3.8-27B<sup></sup> 66.3 84.7 81.1 71.2 39.5 ~$0.104 est. 100.0% 96.9% 93.2% 75.0% 1.01 s raw→ 2.03 s adjustedp95 1.88 s raw → 3.76 s author's demo server
13 by ZeroEntropy ZeroEntropy zerank-2<sup></sup> 66.0 63.0 76.5 79.0 49.8 $0.047 100.0% 79.2% 88.4% 47.3% 0.13 s raw→ 0.40 s adjustedp95 1.50 s raw → 3.15 s our RunPod GPU
14 by OpenAI GPT-5.6 Luna low reasoning effort 65.9 95.3 89.8 77.5 28.5 $0.242 100.0% 97.9% 96.6% 94.5% 0.97 s rawp95 1.82 s raw production API
15 by ekzhang openjev-sglang Qwen3.6-35B-A3B on SGLang 65.3 83.4 77.4 77.1 36.5 ~$0.131 est. 100.0% 95.8% 95.2% 71.4% 0.68 s raw→ 1.36 s adjustedp95 0.73 s raw → 1.45 s author's demo server
16 by Qwen Qwen3-Reranker-4B<sup></sup> 63.8 64.0 67.0 78.7 49.2 $0.050 100.0% 79.2% 87.7% 50.0% 0.13 s raw→ 0.41 s adjustedp95 1.56 s raw → 3.27 s our RunPod GPU
17 by kshetrajna12 reflex-27b<sup></sup> Qwen3.8-27B 63.3 85.8 86.2 67.5 32.3 ~$0.181 est. 100.0% 95.8% 95.9% 75.9% 1.89 s raw→ 3.93 s adjustedp95 2.21 s raw → 4.57 s our RunPod GPU
18 by Zhengxu Yu LitJev<sup></sup> Qwen3.8-27B 62.7 82.4 83.5 66.7 33.6 ~$0.163 est. 100.0% 97.9% 88.4% 73.2% 2.03 s raw→ 4.20 s adjustedp95 2.46 s raw → 5.06 s our RunPod GPU
19 by Jared Palmer kev 0.6B<sup></sup> research preview 62.5 51.9 51.1 75.6 76.1 ~$0.0063 est. 100.0% 81.3% 66.4% 40.0% 0.59 s raw→ 1.33 s adjustedp95 0.97 s raw → 2.09 s our RunPod GPU
20 by Featherless AI SimpleJev Qwen3.6-35B-A3B<sup></sup> 62.5 79.5 67.1 75.0 38.1 ~$0.116 est. 100.0% 93.8% 93.2% 66.4% 0.85 s raw→ 1.70 s adjustedp95 0.93 s raw → 1.86 s author's demo server
21 by David Villalon / Maisa djev<sup></sup> thinking 62.4 80.8 92.7 75.2 26.9 ~$0.274 est. 95.8% 99.0% 80.1% 77.7% 0.43 s raw→ 1.00 s adjustedp95 1.45 s raw → 3.05 s our RunPod GPU
22 by us (GitHub) jev-local<sup></sup> Qwen3.5-9B 61.8 70.8 68.7 69.2 43.3 ~$0.077 est. 100.0% 84.4% 89.0% 59.1% 1.05 s raw→ 2.24 s adjustedp95 2.62 s raw → 5.38 s our RunPod GPU
23 by Mapika decider-2b<sup></sup> 61.7 61.2 46.6 83.2 61.0 ~$0.020 est. 100.0% 85.4% 77.4% 47.3% 0.26 s raw→ 0.67 s adjustedp95 0.28 s raw → 0.72 s our RunPod GPU
24 by Bespoke Labs Bespoke Nimble 9B<sup></sup> 60.5 77.9 65.3 78.7 33.4 ~$0.166 est. 100.0% 94.8% 89.0% 65.5% 0.39 s raw→ 0.93 s adjustedp95 0.65 s raw → 1.46 s our RunPod GPU
25 by Google Gemini 3.1 Flash-Lite 60.1 85.6 68.1 81.8 27.4 $0.264 100.0% 99.0% 93.2% 75.0% 0.76 s rawp95 0.88 s raw production API
26 by razorback16 OpenJev<sup></sup> thinking, BF16 60.0 88.0 69.6 76.1 27.8 ~$0.255 est. 100.0% 100.0% 94.5% 78.2% 0.46 s raw→ 1.08 s adjustedp95 1.08 s raw → 2.31 s our RunPod GPU
27 by Jared Palmer kev 4B<sup></sup> research preview 59.7 64.8 42.0 75.7 61.8 ~$0.019 est. 100.0% 91.7% 85.6% 42.3% 0.55 s raw→ 1.25 s adjustedp95 0.99 s raw → 2.13 s our RunPod GPU
28 by DeepSeek DeepSeek V4.1 Flash thinking default 57.5 94.3 96.7 71.6 16.8 $0.594 98.6% 99.0% 93.2% 95.0% 1.42 s rawp95 4.89 s raw production API
29 by Jared Palmer kev 8B<sup></sup> research preview 56.4 69.4 44.2 74.9 44.0 ~$0.073 est. 100.0% 92.7% 90.4% 47.3% 0.59 s raw→ 1.33 s adjustedp95 1.15 s raw → 2.45 s our RunPod GPU
30 by Zefan Cai (@Zefan_Cai) Open-Jev 9B<sup></sup> Zefan Cai 55.0 71.2 63.3 72.0 28.1 ~$0.249 est. 100.0% 90.6% 81.5% 60.9% 0.75 s raw→ 1.66 s adjustedp95 1.81 s raw → 3.77 s our RunPod GPU
31 by Sean Goedecke system-one Qwen3-8B, Sean Goedecke 54.8 70.3 36.8 84.4 41.5 ~$0.089 est. 100.0% 90.6% 91.8% 50.0% 0.17 s raw→ 0.48 s adjustedp95 0.30 s raw → 0.76 s our RunPod GPU
32 by Logan Markewich jeff<sup></sup> Logan Markewich, GLiFormer 400M 54.4 46.9 64.6 63.5 76.6 ~$0.0060 est. 100.0% 76.0% 61.6% 37.7% 0.94 s raw→ 2.03 s adjustedp95 10.97 s raw → 22.09 s our CPU
33 by Convai Innovations Laya<sup></sup> Convai Innovations, ModernBERT-large 421M 54.4 45.8 62.5 71.1 86.2 ~$0.0029 est. 94.4% 72.9% 69.2% 34.1% 0.79 s raw→ 1.72 s adjustedp95 2.20 s raw → 4.54 s our CPU
34 by Zefan Cai (@Zefan_Cai) Open-Jev 2B<sup></sup> Zefan Cai 51.3 61.0 55.1 73.5 28.1 ~$0.249 est. 100.0% 79.2% 88.4% 42.7% 0.66 s raw→ 1.48 s adjustedp95 1.45 s raw → 3.05 s our RunPod GPU
35 by Deepan Wadhwa OpenDecision<sup></sup> ModernBERT-large zero-shot 40.6 40.8 56.1 79.9 75.3 ~$0.0066 est. 87.5% 62.5% 71.2% 33.2% 0.34 s raw→ 0.83 s adjustedp95 0.54 s raw → 1.24 s our RunPod GPU
36 by Hemant (heman10x) openJev Verdict 1.4<sup></sup> 38.9 38.6 74.1 78.1 82.4 ~$0.0039 est. 86.1% 67.7% 56.2% 37.7% 0.31 s raw→ 0.78 s adjustedp95 0.92 s raw → 2.00 s our CPU
37 by Hemant (heman10x) openJev Verdict<sup></sup> heman10x, ModernBERT-base 151M 38.1 39.8 51.3 76.7 83.1 ~$0.0037 est. 86.1% 65.6% 61.0% 38.2% 0.28 s raw→ 0.71 s adjustedp95 1.45 s raw → 3.04 s our CPU
38 by Jared Palmer kev 0.5B<sup></sup> 33.2 38.2 47.4 77.0 76.1 ~$0.0063 est. 95.8% 52.1% 71.2% 30.9% 0.43 s raw→ 1.01 s adjustedp95 0.92 s raw → 1.99 s our RunPod GPU
39 by Fastino GLiNER2 large<sup></sup> 29.6 40.1 24.3 61.7 73.3 ~$0.0077 est. 98.6% 62.5% 61.0% 36.4% 1.10 s raw→ 2.34 s adjustedp95 14.49 s raw → 29.13 s our CPU
40 by Aditya (isHeSatoshi) smalljev semantic-v9<sup></sup> 27.4 35.1 58.9 79.8 57.9 ~$0.025 est. 97.2% 68.8% 40.4% 38.2% 0.41 s raw→ 0.98 s adjustedp95 0.46 s raw → 1.07 s our RunPod GPU
41 by Fastino GLiNER2<sup></sup> Fastino, gliner2.5-base 24.0 35.6 23.7 71.8 83.1 ~$0.0037 est. 97.2% 66.7% 45.9% 36.4% 0.31 s raw→ 0.78 s adjustedp95 4.15 s raw → 8.46 s our CPU
42 by Kotoba Labs open-jev-deberta-v3-large local CPU 23.1 31.9 66.4 66.0 74.0 ~$0.0073 est. 100.0% 49.0% 53.4% 36.4% 1.77 s raw→ 3.69 s adjustedp95 3.35 s raw → 6.85 s our CPU
43 by Fastino GLiNER2.5 multi<sup></sup> Fastino, 287M 16.6 27.7 56.1 67.8 82.4 ~$0.0039 est. 90.3% 51.0% 43.8% 37.7% 0.43 s raw→ 1.01 s adjustedp95 8.18 s raw → 16.50 s our CPU
44 by Fastino GLiNER2.5 small<sup></sup> Fastino, 74M 13.8 25.6 47.2 77.8 82.4 ~$0.0039 est. 83.3% 47.9% 50.0% 33.2% 0.11 s raw→ 0.38 s adjustedp95 2.10 s raw → 4.35 s our CPU
45 by Mixedbread Mixedbread mxbai-rerank-base-v2<sup></sup> 0.8 6.7 83.1 87.5 67.9 $0.012 44.4% 33.3% 26.7% 40.0% 0.07 s raw→ 0.29 s adjustedp95 0.23 s raw → 0.62 s our RunPod GPU
46 by BAAI BAAI bge-reranker-v2-m3<sup></sup> 0.7 6.3 83.8 89.5 73.4 $0.0077 43.1% 36.5% 8.9% 36.8% 0.03 s raw→ 0.22 s adjustedp95 0.18 s raw → 0.51 s our RunPod GPU
47 by Alibaba-NLP Alibaba GTE Reranker ModernBERT-base<sup></sup> 0.3 4.6 76.8 90.6 69.6 $0.010 33.3% 39.6% 30.1% 33.6% 0.05 s raw→ 0.25 s adjustedp95 0.10 s raw → 0.35 s our RunPod GPU
48 by AltSlate Labs Certo v1<sup></sup> 0.0 0.0 82.0 94.0 100.0 ~$0.0010 est. 27.8% 30.2% 21.9% 31.8% 0.02 s raw→ 0.19 s adjustedp95 0.03 s raw → 0.21 s our RunPod GPU
Honorable mentions — services built on another entrant's model — shown, not ranked: A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.
by mrmps (@michael_chomsky) classifier.dev<sup></sup> fast tierhonorable mention · not ranked 83.6 85.1 77.9 87.6 84.3 ~$0.0033 est. 100.0% 99.0% 97.3% 70.5% 0.39 s rawp95 0.45 s raw production API
Partial runs — shown, not ranked: a tier attempted for fewer than 95 % of its decisions.
by Qwen / Chutes Qwen3.8 27B Chutes TEEpartial run · not ranked 24.8 67.4 92.1 61.3 0.0 ~$2.669 est. 98.6% 99.0% 95.3% 21.4% 5.75 s rawp95 12.97 s raw production API
by Cactus Compute Needle 3, options as tools post-hoc adapter modepartial run · not ranked 1.1 13.5 none (label only) 52.8 65.3 ~$0.014 est. 66.7% 31.3% 34.2% 3.78 s raw→ 7.71 s adjustedp95 33.64 s raw → 67.42 s our CPU
by Cactus Compute Needle 3 Cactus, 2-bit, local CPUpartial run · not ranked 0.1 4.6 none (label only) 59.9 58.7 ~$0.024 est. 47.2% 16.7% 31.5% 7.7% 1.69 s raw→ 3.52 s adjustedp95 14.36 s raw → 28.88 s our CPU
── more in #ai-research 4 stories · sorted by recency
── more on @typesafe ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/jevbench-a-reproduci…] indexed:0 read:12min 2026-09-22 ·