# This is how Jev makes your AI assistant faster, and Judge Jev

> Source: <https://texposit.com/blog/using-jev-for-ai-decisions>
> Published: 2026-10-03 13:46:18+00:00

An AI writing assistants usually spend the first few turns getting ready. Browsing skills, loading them, reaading project files, tool schemas and so on. You can optimise by introducing some level of parallelism, using faster models for the context gathering or having the harness recommend/pre-load things based on the user's request, but these add complexity and you reach a time floor quite fast.

We started experimenting with [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) for this specific part of AI assistant sessions. We pass a short description of the task and a batch of narrow questions to Jev: does this request need this skill, this file, or this tool? Jev answers them together and TeXposit loads everything Jev asked for all at once.

Below is a real recording from a single, slighly extreme case: "What are complex concepts in this paper that need to be visualised with TikZ?" The assistant answered after 73 seconds, and the first 59 of them were setup: it loaded the TikZ skill, read the project's comments file in two passes, and read two ranges of main.tex. Only the last 14 seconds were the answer itself.

The same prompt, using the new Jev-based preflight step, included 42 yes-or-no questions about the project's 23 skills, 14 deferred tools, and 5 candidate files. Jev picked the TikZ skill, main.tex, and the comments file, which are what the assistant had gone and fetched itself as well. It also picked a second skill, a project notes file, and a PDF download tool that were never used. A minor waste of tokens, but insignificant in comparison to the time savings and kept in check with the thresholds and caps. We read up to three files, and unlock up to four tools for the main model's first turn.

We have not measured the aggregate time savings or benchmarkes, but the difference is perceptible, especially when working with slower main models. It is also important to point out that this setup does not give Jev authority to act on the project. The main model still decides what to do, and any edit goes through the same validation and approval as before, so the final output quality does not suffer. In fact, having Jev also act as a judge improves the output quality.

## Judge Jev

In a recent update, we introduced [Rigour Mode](https://texposit.com/updates/update-260907) that runs the model's final answers against a series of customisable requirements automatically, to avoid the user having to go through their requirements one by one or trusting the model not to violate any of its rules. While this helped with the output quality, it came at a significant time and token cost. Further, the judge does not produce a long explanation. It reports which criteria failed, making Jev a good candidate to replace an LLM to power this feature.

Jev tests each active requirement item in parallel against the conversation and the LLM's proposed final answer or edit. A failing check sends feedback back to the main model, which will then revise its answer before the user sees it or an edit reaches approval.

In the example comparison shown below, five examples took about 28 seconds to evaluate with the old LLM-based judge and 2.4 seconds with Jev. Jev agreed with the old judge on four, while in the fifth example, Jev incorrectly rejected a good answer for not disclosing limitations. Further testing and benchmarking might be justified here, but we see that Rigour Mode is useless if it is not enabled, and users do not enable it if it slows down their workflow, so perhaps occasional false positives are acceptable.

We also cap the correction loop: after three failed attempts, the task proceeds while giving the user an explicit warning that the answer may not be rigorous enough, avoiding a strict judge or an impossible request block the task forever. This should help the user see if, for example, their question cannot be answered with the available evidence, rather than having the LLM come back with a dodgy answer with questionable logical connections to the sources.

## Jev as a turn-level router

Some turns are mechanical, such as running a named tool or fixing an error the compiler already identified. Others need ordinary writing judgment, and some involve key claims or complex instructions. TeXposit uses a three-tiered setup where a fast, a general and a smart model (Think e.g. Haiku, Sonnet and Opus) are flexibly available for each session. With Jev, we can decide for every turn which tier to use for the next turn. When Jev is uncertain, it moves up a tier and if routing fails, it defaults to the smart model.

Jev also handles the turn-level automated skill discovery. Until now, based on very simple NLP, we recommended skills to the LLM based on what the user asked about. Now, we replaced a hand-written phrase list with Jev's yes-or-no questions as a request may point to a tool or skill even when the user does not use the words in its description.

As a sidenote, for now, the router does not apply to users who bring their own model OpenRouter key. Switching them to a different model without asking would be a poor trade, even if it saved time.
