{"slug": "we-swapped-our-llms-for-jev-it-s-39-cheaper", "title": "We swapped our LLMs for Jev. It's 39% cheaper", "summary": "TypeSafe AI's decision model Jev cut routing costs by almost half and made P90 latency 3x faster — from 1.5 seconds to under 500 ms — for an always-on on-call agent that had previously used LLMs, according to a first-person account of a week-long production trial. Jev, which does not generate text but returns typed values and calibrated probabilities via Noul, Choice and Score question types, is pitched for \"smart if-statements\" at 70 to 500 ms per response with free output tokens. The team put Jev in production for routing and three classification tasks, each replacing a different LLM.", "body_md": "# We swapped our LLMs for Jev. It's 39% cheaper.\n\n## Explore with AI\n\nWe’re building an always-on on-call agent, which constantly makes decisions, for example:\n\n- how severe is this issue?\n- should we be pro-active on this slack thread?\n- have we seen this incident before?\n\nUp until last week, we’ve been using LLMs for these tasks. As soon as we got access to [Jev](https://docs.typesafe.ai/models) this week we started experimenting with it.\n\n## What is Jev?\n\nJev is a decision model from [TypeSafe AI](https://typesafe.ai/blog/introducing-system-one-models-and-jev). It doesn’t generate text: it answers questions about your data with typed values and calibrated probabilities. TypeSafe pitches it for “smart if-statements”, the classify, route and score steps where hand-written logic is too brittle, and quotes 70 to 500 ms per response with free output tokens. Our agents make those decisions all day, so we put it in production.\n\n### How it works\n\nA request has two parts: the `state`, which is the text or JSON you want a decision about, and the `questions` you want answered. Each question has one of three types:\n\n- **Noul** returns the probability of`yes` .\n- **Choice** selects one of the provided options.\n- **Score** returns a probability-weighted value across an ordered rubric.\n\n## Use-case 1: Routing, or when should the agent respond?\n\nOur agents are pro-active on Slack and Github. They will respond to Slack messages or Github comments if they have a meaningful insight to provide to the user. The agents need to act when a user requests it, without overreacting to every event. This is a perfect use-case for Jev’s `Noul` questions.\n\n- **Comments on PRs our agents submitted:** a reviewer asking for a change wakes the agent up. A CI status update or a “thanks” doesn’t.\n- **Slack channels:** the agent jumps in uninvited only when it can clearly help, like a direct infrastructure question. It stays out of humans coordinating with each other.\n- **Slack threads the agent is in:** it answers messages meant for it, ignores humans talking to each other, and leaves when someone asks it to.\n\nHere is an example of a Jev request to determine if an agent should follow up on a Slack message.\n\n```\n{\n  \"model\": \"jev-1.13.0\",\n  \"state\": {\n    \"slack_channel\": \"#Deployment\",\n    \"user_message\": \"Watch this PR until fully deployed\"\n  },\n  \"questions\": {\n    \"respond\": {\n      \"type\": \"noul\",\n      \"instructions\": {\n        \"question\": \"Should the agent follow up on this user message?\"\n      }\n    }\n  }\n}\n```\n\nWe ran Jev for about a week and compared with the LLM we previously used for this task.\n\nRouting got 3x faster at P90, from 1.5 seconds to under 500 ms, and the cost dropped by almost half.\n\n## Use case 2: Classification, or what does the evidence mean?\n\nWe also run a few tasks where the agent needs to classify things into different buckets. For example, based on the provided evidence, should the agent start investigating the issue, should it fold it as related to an existing issue, or is it a duplicate of a previously resolved issue?\n\nA Jev Choice question turns that evidence into one label from criteria we define, and the next step is decided by the label. It looks like this:\n\nWe run three classifications this way. Each one uses a Choice question with explicit criteria, and each one replaced a different LLM, so we measured them separately.\n\n### Are these incidents related?\n\nWhen a new incident opens, the agent compares it with existing incidents: do they share one root cause, are they related, or are they independent? This illustrative example shows one candidate; a production call compares several against the same new incident.\n\n```\n{\n  \"model\": \"jev-1.13.0\",\n  \"state\": {\n    \"incident\": \"Checkout cannot authenticate to the database.\",\n    \"candidate_0\": \"Billing cannot authenticate to the same database.\",\n    \"evidence\": \"Both services use a credential revoked at 14:00.\"\n  },\n  \"questions\": {\n    \"candidate_0\": {\n      \"type\": \"choice\",\n      \"instructions\": \"How is candidate_0 connected to the new incident?\",\n      \"criteria\": {\n        \"duplicate_same_root_cause\": \"One underlying problem explains both\",\n        \"related\": \"Distinct problems share a trigger or blast radius\",\n        \"independent\": \"No evidenced connection\"\n      }\n    }\n  }\n}\n```\n\nBy switching to Jev for this use-case, we made it almost 8x faster at P90, from 2.9 seconds to under 400 ms, and 27% cheaper.\n\n### Why did this PR we submitted get closed?\n\nOur agents submit pull requests to developers, and one of our key success metrics is the merge rate. We need to understand why pull requests get closed such that we can improve the product.\n\nWhen one of our PRs is closed without merging, the agent reads the reviews, the discussion, and references to other work, then picks the reason: the fix was wrong, a human fixed it another way, the issue was a false positive, the PR went stale, or the behaviour was intended.\n\nBy switching to Jev for this use-case, we improved latency to 6x faster at P90, but only 17% cheaper.\n\n### How pressing is this pull request?\n\nBefore submitting a pull request to developers, our agents need to rank it such that higher severity things rise to the top. For this classification, the agent grades the underlying problem on a ladder from critical to info. It grades current impact, not hypothetical risk.\n\nSwitching to Jev here substantially reduced cost: 59% cheaper than DeepSeek V4.1 Flash, and nearly 5x faster at P90.\n\nJev is faster on every classification. Where it replaced GPT-OSS 120B, the gain is mostly latency. Where it replaced DeepSeek V4.1 Flash, it also cut the bill by more than half.\n\n## Use case 3: Ranking, or how important is this cloud resource?\n\nTo get the best out of Polylane, teams connect their cloud accounts. We create a context graph of all the cloud resources, such that the agents can quickly understand the relationship between compute nodes, databases, queues, etc.\n\nWe have teams on the platform with extremely busy cloud accounts, with 10s of thousands of nodes. Each server, sandbox, database, and queue is a node in our context graph. It’s necessary to rank each of these nodes such that agents know what is critical to your application, and what is essentially “fine” to fail.\n\nWe assign one of four priority tiers to each resource: Critical, Standard, Low, or Minimal.\n\nJev evaluates a Choice question for each resource, with context based on configuration, environment, recent metrics, and dependencies.\n\n```\n{\n  \"model\": \"jev-1.13.0\",\n  \"state\": {\n    \"instructions\": \"Assign importance relative to the other resources in this cohort.\",\n    \"cohort\": [\n      {\n        \"id\": \"database-a\",\n        \"environment\": \"production\",\n        \"daily_queries\": 80000,\n        \"dependents\": 6\n      },\n      {\n        \"id\": \"database-b\",\n        \"environment\": \"preview\",\n        \"daily_queries\": 0,\n        \"dependents\": 0\n      }\n    ]\n  },\n  \"questions\": {\n    \"resource_0\": {\n      \"type\": \"choice\",\n      \"instructions\": \"Assign the importance tier for database-a.\",\n      \"criteria\": {\n        \"1\": \"Critical: substantial production traffic or blast radius\",\n        \"2\": \"Standard: active and operationally relevant\",\n        \"3\": \"Low: limited activity or importance\",\n        \"4\": \"Minimal: idle or disposable, without meaningful dependents\"\n      }\n    }\n  }\n}\n```\n\nRanking is where Jev shines: more than 10x faster at P90, from 5.4 seconds to about half a second, and 34% cheaper. It’s also our highest-volume decision, so it drives most of the overall savings.\n\n## Summary\n\nOverall, Jev delivered substantial overall reductions in latency and estimated cost per 1,000 calls:\n\n- P90 latency reduction: 4,752 ms —> 508 ms.\n- Cost per 1,000 calls reduction: $0.76199 —> $0.46369.\n\nPer model, Jev is both the fastest and the cheapest: slightly cheaper than DeepSeek V4.1 Flash, and well below both GPT-OSS models.\n\nWherever our agents pick from a fixed set of answers, Jev is now the default: it is faster and cheaper on every decision we moved.", "url": "https://wpnews.pro/news/we-swapped-our-llms-for-jev-it-s-39-cheaper", "canonical_source": "https://polylane.com/blog/we-swapped-our-llms-for-jev/", "published_at": "2026-09-28 17:37:04+00:00", "updated_at": "2026-09-28 17:49:40.557958+00:00", "lang": "en", "topics": ["ai-agents", "ai-products", "large-language-models", "ai-tools"], "entities": ["TypeSafe AI", "Jev", "Slack", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/we-swapped-our-llms-for-jev-it-s-39-cheaper", "markdown": "https://wpnews.pro/news/we-swapped-our-llms-for-jev-it-s-39-cheaper.md", "text": "https://wpnews.pro/news/we-swapped-our-llms-for-jev-it-s-39-cheaper.txt", "jsonld": "https://wpnews.pro/news/we-swapped-our-llms-for-jev-it-s-39-cheaper.jsonld"}}