System 1 (Jev) Models: Faster and Cheaper Proxy to Frontier Models — With a Working Router Four "System 1" decision models — Jev, Laya, Kev and Nimble — are now available as faster, cheaper alternatives to full frontier LLM calls for routing and typed decisions, according to a Towards AI article by Muhammad Soliman published October 6, 2026. The article presents a proof-of-concept tier router implemented as a LiteLLM callback that uses a classifier to pick among backend models such as haiku, sonnet and opus before a request leaves the machine, rewriting only an opt-in alias and failing open while logging decisions. Soliman argues these models answer typed questions in a single forward pass, returning probabilities with no generated tokens, and should be treated as bases to specialize via small fine-tuning sets with small option sets and calibrated confidence. Last Updated on October 6, 2026 by Editorial Team Author s : Muhammad Soliman Originally published on Towards AI. Jev, Laya, Kev and Nimble in one place — plus a short, open proof of concept that puts one in front of a LiteLLM proxy and picks the model for every request. Ask your coding agent to rename a variable, and it will load a few hundred billion parameters to do it. Is it worth to pass everything to LLM as is and waste hundreds of thousands of tokens for simple questions/routing/decisions — The Most Expensive Question You Ask an AI Is “Yes or No” and System 1 model was built to save alot of respond faster. The article explains why “System 1” decision models can replace expensive full LLM calls for routing and other typed decisions, contrasting it with System 2 token-generating models. It introduces how these models work typed questions answered in one forward pass, returning probabilities with no generated tokens , covers the new availability of four options Jev, Laya, Kev, and Nimble , and discusses how to read benchmarks carefully—these models are meant as bases to specialize via small fine-tuning sets, keep option sets small, calibrate confidence, and match hardware to the latency/claims. It then shows practical integration: using a classifier to choose among backend models e.g., haiku/sonnet/opus before the request leaves the machine, demonstrating a proof-of-concept tier router implemented as a LiteLLM callback that rewrites only an opt-in alias and fails open while logging decisions. Finally, it outlines where to take it next measuring on real traffic, fine-tuning the classifier, adding guardrails in the same pass, quota awareness, and other proxy/gateway hooks and concludes with the key takeaway about the strong economics and relatively small integration effort. Read the full blog for free on Medium. Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor. Published via Towards AI