{"slug": "a-typed-classifier-out-judges-llms-on-agent-scoring", "title": "A typed classifier out-judges LLMs on agent scoring", "summary": "LangChain tested TypeSafe AI's Jev, a typed classifier rather than a text-generating LLM, as an agent-eval judge, and Jev matched a human reviewer on all 500 repeated decisions while costing a fraction of a cent per call, far less than Claude's LLM-judge cost. The result points to typed classifiers as a cheaper alternative to LLM judges for scoring agent behavior.", "body_md": "LangChain tested TypeSafe AI's Jev, a typed classifier rather than a text-generating LLM, as an agent-eval judge. It matched a human reviewer on every one of 500 repeated decisions while costing a fraction of a cent per call, far less than Claude's LLM-judge cost.\nRead: LangChain tested TypeSafe AI's Jev, a typed classifier rather than a text-generating LLM, as an agent-eval judge. It matched a human reviewer on every one of 500 repeated decisions while costing a fraction of a cent per call, far less than Claude's LLM-judge cost.\nWatch: Filip Makraduli's FlashNorm folds a transformer's norm layer into its projection weights and overlaps the remaining divide on a separate CUDA stream, cutting norm-plus-projection cost by a third with no retraining needed.\nWatch: A rare vLLM bug corrupted about one in a thousand prompts with no error. The cause: a scheduler race let decode run before prefill for Jamba's Mamba layers, computing a fresh request over a stale prior state.\nRead: Ben Swerdlow ran 171 real-time StarCraft matches between Codex, Claude, and Grok models, surfacing concrete failure modes in continuous, multi-unit agent control that single-turn benchmarks miss.\nRead: PlanetScale released Tin, a GA Postgres extension with boolean, phrase, fuzzy, and BM25-ranked search that keeps correct transactional visibility, cutting a common reason teams bolt on Elasticsearch.\nRead: Hacktron's fuller HEIF Heist disclosure shows the libheif bugs behind Tuesday's OpenAI account takeover also reach Slack, Meta, GitHub Enterprise, Rails, and Next.js through indirect dependencies.", "url": "https://wpnews.pro/news/a-typed-classifier-out-judges-llms-on-agent-scoring", "canonical_source": "https://www.vibeleaderboard.ai/intel/brief/2026-09-20", "published_at": "2026-09-20 16:05:16+00:00", "updated_at": "2026-09-20 16:55:00.340763+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models"], "entities": ["LangChain", "TypeSafe AI", "Jev", "Claude"], "alternates": {"html": "https://wpnews.pro/news/a-typed-classifier-out-judges-llms-on-agent-scoring", "markdown": "https://wpnews.pro/news/a-typed-classifier-out-judges-llms-on-agent-scoring.md", "text": "https://wpnews.pro/news/a-typed-classifier-out-judges-llms-on-agent-scoring.txt", "jsonld": "https://wpnews.pro/news/a-typed-classifier-out-judges-llms-on-agent-scoring.jsonld"}}