13:49
2026-09-25
github.com
artificial-intelligence
Benchmarking Jev, Laya, and five open models under three stresses
TypeSafe's Jev decision model answered 60% of items correctly at 128 candidate answers, versus 39% for Laya and 41% for the best open model, in a benchmark of seven systems run on a frozen dataset of …