# Understanding System One Models Through Their Architecture

> Source: <https://pub.towardsai.net/understanding-system-one-models-through-their-architecture-4d57f90b2624?source=rss----98111c9905da---4>
> Published: 2026-10-03 12:37:24+00:00

On September 15, 2026, TypeSafe AI launched Jev and introduced a new category of models called **System One models**. Instead of generating text, they take a state, such as a support ticket, and a set of possible options, and return a probability for each option in a single forward pass. Unlike a traditional classifier, whose labels are fixed during training, they receive the options as input, so the same model can make different decisions without retraining.

Jev is closed, but within two weeks open-source alternatives appeared, including Laya, Lev and CLM. Despite the new interface, all three are built on pretrained language models: Lev and CLM on **Qwen**, Laya on **ModernBERT**. What changes is where the decision is read from the network, and that choice determines how each model behaves.

This post derives expected behaviors from the architecture of the three models, tests them through controlled experiments, and then applies the same tests to Jev.

All three models take the same inputs: a **state**, which contains the information to be classified, and a set of **options**, each with a name and, optionally, a description.

**Laya** (421M, ModernBERT) concatenates the options and the state into a single sequence, options first, each one preceded by a [MASK] token that marks where it begins. The encoder reads the whole sequence in one pass, so the embedding at each [MASK] absorbs information from the state and ends up reflecting how well its option matches it. Instead of generating anything, Laya passes these embeddings through a small MLP that scores each option.

**Lev** (Qwen3.5–4B + LoRA) turns the task into a multiple-choice question: the state first, then the options labeled A, B, C…, then “Answer:”. After one pass, Qwen computes the probability of every token coming next, as it would before writing an answer. Lev keeps only the probabilities of the option letters and rescales them to sum to one. A LoRA adapter, trained on this format, pushes the probability toward the correct letter, and the result is averaged over two orders of the options. I call this mode A. Lev also has a mode B, in which each option is embedded separately and compared with the state.

**CLM** (Qwen3–8B + two MLPs) never puts the state and the options in the same input. Qwen reads the state, and then each option, in separate passes, keeping one embedding per text; for options with a description, only the description is encoded. These embeddings capture the subject, not the fit: a state about a duplicate charge is close to every payment-related option. Two small MLPs, the only trained part, map states and options into a shared space where closeness means “correct option for this state”, and the cosine similarity decides. Since each option is encoded on its own, its embedding can be computed once and cached.

All experiments ran on NVIDIA T4 GPUs, with the same states from the banking77 test split, the same options and the same seeds for every model. One caveat applies throughout: Lev was trained on banking77, so its absolute accuracy is inflated, and the comparisons focus on how each model’s behavior changes between conditions.

Models that write the options into the input should get slower as options are added, while models that encode them separately should stay flat, as long as the option embeddings are cached. I ran 100 states with 2 to 284 options, measuring CLM and Lev’s mode B with an empty cache (cold) and with the options already cached (warm), and Laya both with its default budget of 512 tokens and with the budget raised so that every option fits whole.

With a warm cache, CLM and mode B stayed flat; with an empty one, both grew about 10×. The flat curve comes from the cache, not from the architecture by itself. Lev’s mode A grew 18.6×, since it rereads every option twice per question. Laya’s default configuration also looked flat, but only because it truncates the options to fit its budget. At 150 options, it rejected 99 of 100 questions, and the one it accepted had lost all 32 tokens of the state, yet it still returned probabilities without any warning.

Accuracy also diverged. Laya fell from 0.93 to 0.20, and CLM fell close to zero from 50 options on, because its training on long texts does not separate short intent names. Experiments 2 and 5 explain both.

The options have no natural order, so a well-behaved model should not care how they are arranged. Each state was answered with the same options in five different orders, measuring how often the answer changed and how much probability moved between options.

CLM never changed its answer, and its probabilities did not move at all: each option embedding is computed once and reused in every order. Lev’s mode A changed its answer for only 6% of the states with 77 options. Laya changed for 79%, and its accuracy depended on where the correct option sat: 32% when it was in the first quarter of the list, 62% when it was in the last.

Laya places the options before the state, so the last options are also the closest to it. To separate the place in the list from the distance to the state, every option was made longer with the same meaningless words (“card arrival: banking intent”), pushing each position farther from the state. The fifth option was found 55% of the time in the short list and only 33% in the long one. A logistic regression confirms that distance explains the errors (p = 0.0008) and the place in the list does not (p = 0.57).

The cause is ModernBERT’s attention: in two out of every three layers, each token sees only the 128 tokens around it, so a distant option can read the state in only about 9 of its 28 layers. Lev showed no such effect, because it reads the answer at the end of the prompt, from a position that sees every token.

Often a single word decides the answer: charged once or twice, arrived or never arrived. Pairs of hand-written sentences that differ by one edit, in counts, negations, numbers, roles, time and quantifiers, tested whether each model follows the change. A pair only counts as correct when both sentences are answered correctly.

Lev followed almost every edit, helped by its size and by training on similar examples. Laya followed edits that add or remove a word, such as “twice” or “never”, but failed when the edit changed how existing words relate: swapped amounts, swapped flight times, or “I owe my brother money” against “My brother owes me money”. CLM failed almost every pair. It compresses the state into a single embedding before seeing any option, and a one-word edit barely moves that embedding.

Each option can carry a name and a description. To see which one each model relies on, I replaced the names with meaningless codes and, in one condition, gave every option a misleading name while keeping its correct description.

CLM produced the same accuracy, to the last digit, in every condition with descriptions, because it never encodes the name. Laya followed the misleading name in 72% of the cases, since the [MASK] that scores each option sits right next to the name. Lev followed the description 92% of the time, reading each option as a whole line, although its training may also contribute. In practice, the same set of options can lead to different decisions: for Laya the names decide, while for CLM only the descriptions count.

The last experiment compares the embeddings at the point where each model reads its decision, with and without the part of the model that was trained.

In the untrained ModernBERT, Laya’s embeddings group by option and barely distinguish correct options from incorrect ones. After fine-tuning, the separation is almost perfect (AUROC 0.995), already inside the encoder.

In Lev, the LoRA is not what makes the model answer with letters: the base Qwen already places 99.97% of its probability on them and is right 98% of the time. What the adapter changes is the internal organization, from the letter the model is about to say to the intent of the state.

In CLM’s raw Qwen space, almost every embedding points in the same direction (average cosine 0.98). The MLPs spread the states apart and group them by intent, but the short options stay bundled together, which is why the correct option is often not the closest one.

Jev’s weights are closed, but most of the experiments only need inputs and output probabilities, so they can run through its API. Since its internals are not accessible, the token counts reported for each call took the place of Experiment 5.

The results consistently placed Jev next to Lev. Its latency was flat, dominated by a fixed cost of about 250 ms, which alone does not distinguish the designs. Its answers changed with the order of the options 2.5 to 4.3 times more than between identical calls, which rules out CLM’s separate encoding, and neither the place nor the distance of the correct option mattered, which rules out Laya’s layout. It also solved almost every minimal edit and weighed the name and the description of each option together.

The token counts complete the picture:

The most likely design is therefore Lev’s, a decoder that reads the state followed by coded options, with one addition: a state shared by several questions evaluated in parallel. This is an inference from behavior, not a confirmed architecture, but it shows that tests derived from open models are enough to place a closed one among them.

A benchmark score shows how a model performed on one dataset, not how it behaves when the input changes. Every difference observed here traced back to the architecture:

A System One model is not a new kind of network, but a pretrained language model read at a different point. That leaves room for new variants: CLM’s cached embeddings could narrow hundreds of options down to a few, and Lev’s joint reading could decide among them. Understanding why each design behaves as it does is what makes such combinations possible.

The full version of this post, with all tables, statistics and interactive charts, is available on my portfolio. [[Full post](https://www.silasdata.com/system-one-models/)]

**Repositories:** [Laya](https://github.com/NandhaKishorM/laya) (Convai Innovations) · [Lev](https://github.com/InterfazeAI/lev) (Interfaze) · [CLM](https://github.com/Contrastive-LM/CLM) (Contrastive-LM)

[Understanding System One Models Through Their Architecture](https://pub.towardsai.net/understanding-system-one-models-through-their-architecture-4d57f90b2624) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.
