cd /news/ai-agents/a-typed-classifier-out-judges-llms-o… · home topics ai-agents article
[ARTICLE · art-135211] src=vibeleaderboard.ai ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

A typed classifier out-judges LLMs on agent scoring

LangChain tested TypeSafe AI's Jev, a typed classifier rather than a text-generating LLM, as an agent-eval judge, and Jev matched a human reviewer on all 500 repeated decisions while costing a fraction of a cent per call, far less than Claude's LLM-judge cost. The result points to typed classifiers as a cheaper alternative to LLM judges for scoring agent behavior.

read1 min views1 publishedSep 20, 2026

LangChain tested TypeSafe AI's Jev, a typed classifier rather than a text-generating LLM, as an agent-eval judge. It matched a human reviewer on every one of 500 repeated decisions while costing a fraction of a cent per call, far less than Claude's LLM-judge cost. Read: LangChain tested TypeSafe AI's Jev, a typed classifier rather than a text-generating LLM, as an agent-eval judge. It matched a human reviewer on every one of 500 repeated decisions while costing a fraction of a cent per call, far less than Claude's LLM-judge cost. Watch: Filip Makraduli's FlashNorm folds a transformer's norm layer into its projection weights and overlaps the remaining divide on a separate CUDA stream, cutting norm-plus-projection cost by a third with no retraining needed. Watch: A rare vLLM bug corrupted about one in a thousand prompts with no error. The cause: a scheduler race let decode run before prefill for Jamba's Mamba layers, computing a fresh request over a stale prior state. Read: Ben Swerdlow ran 171 real-time StarCraft matches between Codex, Claude, and Grok models, surfacing concrete failure modes in continuous, multi-unit agent control that single-turn benchmarks miss. Read: PlanetScale released Tin, a GA Postgres extension with boolean, phrase, fuzzy, and BM25-ranked search that keeps correct transactional visibility, cutting a common reason teams bolt on Elasticsearch. Read: Hacktron's fuller HEIF Heist disclosure shows the libheif bugs behind Tuesday's OpenAI account takeover also reach Slack, Meta, GitHub Enterprise, Rails, and Next.js through indirect dependencies.

── more in #ai-agents 4 stories · sorted by recency
── more on @langchain 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-typed-classifier-o…] indexed:0 read:1min 2026-09-20 ·