I built an independent benchmark to test TypeSafe's Jev model against GPT-4, Claude, and Gemini on classification tasks.
Jev is a different kind of model — instead of generating text, it outputs probabilities for given answer choices. This makes it particularly interesting for:
The benchmark covers spam detection, sentiment analysis, and topic classification tasks with full reproducible code.
Read the full article + code on Medium:
https://github.com/PavelRavvich/jev-bench Connect with me on LinkedIn:
https://www.linkedin.com/in/pavel-ravvich/