# An Independent Evaluation of TypeSafe's Jev

> Source: <https://www.vals.ai/blogs/independent-evaluation-of-jev>
> Published: 2026-10-06 21:29:38+00:00

[Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev), a TypeSafe model, does not generate text. Instead, you give it a state, which can be any text or structured data, and questions, each with a list of allowed answers. It returns one answer per question, as a probability distribution over the possible options. Input costs $0.042 per million tokens, and output is free.

Our proprietary benchmarks evaluate written answers, so they are not compatible with Jev. We therefore built two benchmarks that do: one has 400 claim-verification items made from SEC filings, and the other is a preregistered 12-subtask slice of LegalBench, a public set of legal-reasoning tasks. We ran twelve systems on both benchmarks, and Jev matched frontier models on claim verification at about 1/500th the cost of GPT-6 Astra.

## Key Takeaways

- On claim verification, all twelve systems score 0.953 to 0.995. Jev, the cheapest model by far, scores 0.975 at $0.02 per 1,000 cases and **ties with GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna** .
- On LegalBench, Jev finishes last of twelve on raw accuracy and on class-balanced accuracy. On raw accuracy, every other system’s lead is statistically reliable.
- Jev answers in about a tenth of a second whether asked for one judgment about a document or thirty-two. At thirty-two, its median latency is 194x lower than GPT-6 Astra’s and 241x lower than Gemini 4 Argon’s.
- With a 1% error budget, six of twelve systems exceeded the limit on held-out items. Jev was one of them, automating 95% of cases at 1.6% error.

## How we evaluated Jev

### Claim verification

A system reads a claim and a source excerpt. It classifies the claim as `supported`, `contradicted`, or `insufficient_evidence`. We built 400 items from the latest 10-Ks of 48 large US companies.

The dataset has five categories with 80 items each:

| Category | How it was built | Correct answer | 
|---|---|---|
| Direct support | The claim paraphrases a sentence in the excerpt | `supported` | 
| Subtle contradiction | One meaning-inverting edit to that sentence | `contradicted` | 
| Numeric discrepancy | A figure altered to a plausible wrong value | `contradicted` | 
| Missing evidence | The claim comes from elsewhere in the same filing | `insufficient_evidence` | 
| Non-supporting citation | The excerpt comes from a different company on the same topic | `insufficient_evidence` | 

Every item goes through the same steps:

1. **Pull a document.** A script downloads each company’s latest 10-K from EDGAR and keeps the paragraphs that read as prose: two or more sentences, mostly letters, no tables of figures. That gave 1,960 passages.
2. **Build the claim.** One passage becomes the source excerpt. The claim is built from one of its sentences (or, for the two`insufficient_evidence` categories, from a different passage). The construction rule decides the correct answer. No model labels anything, and we fix the answer before any system sees the item.
3. **Ask every system the same question.** Each system gets the claim, the excerpt, three defined labels with one-line definitions, and an instruction to judge only from the excerpt. It does not receive the company name, the filing, or a hint about the category. LLMs receive a prompt with a structured-output schema; Jev receives a single`choice` question with the same labels as its options and the same definitions as its rubric.
4. **Record the answer.** We keep the label, the probabilities, the wall-clock time, and the token cost of every call. A system is right when its label matches the fixed answer.

**Gemini 3.8 Flash** wrote the 160 paraphrases and contradictions. It only rewrote sentences; automatic checks rejected rewrites that changed too much or too little, and the labels come from the construction rule, not from the model.

Every item and fixed answer was audited by a human.

#### Example

Here’s one passage from a McDonald’s 10-K, along with the two items we built from it. Switch between the items to see how we edited the sentence and how each system responded.

This is the exact prompt the LLMs saw for the first item, with the excerpt shortened:

```
Decide how the source excerpt relates to the claim. Judge only from the excerpt; do not use outside knowledge about the company.

Allowed labels:
- supported: The excerpt states or directly entails the claim.
- contradicted: The excerpt states something that conflicts with the claim.
- insufficient_evidence: The excerpt neither supports nor contradicts the claim.

Claim: The Company expects 2026 capital expenditures to be between $3.7 and $3.9 billion, with the minority directed towards new restaurant unit expansion across the U.S. and International Operated Markets.

Source excerpt:
• The Company expects 2026 capital expenditures to be between $3.7 and $3.9 billion, with the majority directed towards new restaurant unit expansion across the U.S. and International Operated Markets. Globally, the Company expects to open approximately 2,600 restaurants in 2026 [...]

Answer with exactly one label from the allowed set, plus your probability for every allowed label. Probabilities must be between 0 and 1 and sum to 1. Do not explain.
```

All twelve systems got both items right, consistent with the broader results. Accuracy on this set ranged from 0.953 to 0.995, and no system missed more than 32 of the 400. The differences are in those items, as well as in time and cost. Here, Jev answered in about a tenth of a second for a few thousandths of a cent, but the LLMs took up to fifteen seconds and up to a cent and a half. These cost and latency figures end up having an impact after thousands of claim verification tasks.

### LegalBench

LegalBench is a public benchmark of legal-reasoning tasks. We chose it because most of its subtasks have a fixed set of answers, which Jev can answer directly. Before calling any model, we fixed 12 distinct subtasks and 33 questions from each in writing: 396 questions, sampled from the 1,808 available in those subtasks. **Because this is a small slice of LegalBench, it is not representative of performance on the full benchmark.** LegalBench is public and probably sits in the LLMs’ training data.

## Results

### Claim verification

Here, Jev performs like the frontier models: all twelve systems land within seventeen items out of 400 items, and Jev ties GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna at 390 correct answers. Claude Sonnet 5.5 and Opus 5.5 score a few items higher, but the gap is small enough that we can’t call it real.

The difference is price: Jev costs 1/498th as much as GPT-6 Astra, 1/295th as much as Opus 5.5, and 1/138th as much as Sonnet 5.5. GPT-6 Luna, the cheapest model that matches Jev’s score, costs about three times as much ($0.08 per 1,000 cases against $0.02). Sonnet 5.5, the top scorer, gets eight more items right (out of four hundred) at 138 times the price.

All twelve systems score 1.00 on the 80 numeric-discrepancy items, so that category does not separate them.

### LegalBench

The answers for some subtasks are not evenly distributed between options, so our protocol focuses on class-balanced accuracy. On that metric, GPT-5.6 Terra leads at 0.932, and Jev ranks last at 0.730, a hair behind Haiku 4.5 (0.732). On raw accuracy, every other system’s lead over Jev remains statistically significant after correcting for eleven comparisons.

Jev struggled most on the two contract entailment subtasks, the closest LegalBench match to claim verification. On `contract_nli_permissible_copy`, it answered “Yes” to all 33 items. The gold answers split 16 yes and 17 no, so a constant “Yes” scores 0.485. On `contract_nli_confidentiality_of_agreement`, it answered yes to 30 of 33. Every LLM scored 0.91 to 1.00 on these items, but each subtask has only 33 items, so trust the pattern across both over any single number.

Calibration varied across the two benchmarks. Expected calibration error (ECE) measures the gap between a model’s stated confidence and its actual accuracy. Jev’s ECE was 0.011 on claim verification, the lowest of the twelve, but 0.108 on LegalBench, the highest. Some of the rise is mechanical, since error grows as accuracy falls, but a calibration measured on one task did not hold on the other.

### How much work can be automated

A calibrated model can accept the answers it is confident about and send the rest to a person. We asked how much work each system can take before its error rate passes a budget. To avoid tuning on the data that measures it, we set each threshold on the first 30% of items to meet a 1% error budget, froze it, and applied it to the other 70%.

**Coverage and error at a 1% error budget**

| Router |  |  | 
|---|---|---|
| Claude Sonnet 5.5 | 100% | 0.4% | 
| Gemini 4 Argon | 100% | 0.8% | 
| Claude Opus 5.5 | 99% | 0.4% | 
| GPT-6.1 Sol | 99% | 1.9% | 
| GPT-6 Astra | 99% | 1.5% | 
| GPT-6 Luna | 97% | 1.9% | 
| GPT-5.6 Sol | 97% | 1.6% | 
| DeepSeek V4.1 Flash | 96% | 0.8% | 
| Jev 1.13.0 | 95% | 1.6% | 
| GPT-5.6 Terra | 94% | 1.2% | 
| GPT-5.4 mini | 84% | 0.9% | 
| Claude Haiku 4.5 | 22% | 0.0% | 

Six of twelve systems exceeded the budget. Jev was among them, automating 95% of cases at 1.6% error. Haiku 4.5 stayed within budget only by automating 22% of cases.

In hindsight, sweeping the threshold over all 400 items would have allowed Jev to automate 91% of cases within the budget, compared with 100% for Opus 5.5, Sonnet 5.5, and Gemini 4 Argon. That is the upper limit, since no one could pick the threshold in advance. Jev’s `confidence` score tracks its probability distribution almost exactly and adds no additional signal.

To match Astra’s accuracy on the held-out items, Jev sends 1% of cases to Astra, at 27 cents per thousand cases against Astra’s $12, though GPT-6 Luna clears that bar alone at eight cents. On LegalBench, Jev must defer 64% of cases and saves 18%; routing through GPT-6.1 Sol saves 72%.

### Latency and cost

We asked each system for 1 to 64 judgments about one long document. Jev’s latency barely rises, from 0.10s to 0.12s, but the LLMs slow sharply: at sixty-four judgments GPT-6 Astra takes 38s and Gemini 4 Argon 65s. On single judgments under load, Jev’s median is still 4x to 74x lower across the roster.

Asking an LLM for probabilities instead of a bare label raises its cost by 29% (GPT-6 Luna) to 101% (GPT-5.6 Sol). Gemini 4 Argon is the exception: its probability calls cost 3% less than its label calls. Jev’s output is the distribution, so it adds nothing.

No system logged a malformed response, and Jev logged no service failures. All calls on LegalBench that timed out after 180 seconds were retried. After it, GPT-5.6 Terra was still missing 10 answers, Claude Opus 5.5 6, Sonnet 5.5 2, and GPT-5.4 mini 1; 16 of those 19 were on `international_citizenship_questions`. Google blocked Gemini 4 Argon on one `learned_hands_crime` item. Systems are scored on the items they answered. Jev gave the same answer on all five repeats of 25 items.

## Takeaways

For bounded verification over a document, Jev sits within noise of every system we tested. It gives calibrated probabilities at no extra cost, it answers in about a tenth of a second under load, and it costs thousandths of a frontier model’s price. Perhaps most importantly, more questions about the same document add almost no latency.

TypeSafe’s [launch post](https://typesafe.ai/blog/introducing-system-one-models-and-jev) claims “193.6x faster, 444.6x cheaper” than the average of GPT-6 Astra and Claude Fable 5.1. It calls the figures “on the higher end of real world gains.” TypeSafe’s own model capabilities team built the workflows.

This headline holds up: we measure 1/498th of Astra’s cost, against their 444.6x. Their latency figure covers a multi-step workflow. Against Astra we measure 194x when thirty-two judgments are requested, in line with their figure, and 47x on single judgments under load.

Even cheap LLMs can often match Jev’s performance but are still several times more costly. GPT-6 Luna scored 0.975 on claim verification at eight cents per thousand cases, equivalent to Jev’s score at about three times its price.

To test your own workload:

1. Collect 200 to 400 real cases with known answers, and audit them manually before showing them to any model.
2. Ask each candidate for a label and a probability, pricing them separately.
3. Split the cases 30/70 and set a confidence threshold on the first 30% that meets your error budget.
4. Apply it to the other 70%, and read the realized error.

Step 4 is easy to skip, and it showed six of our twelve systems exceeding the budget they were tuned for.

## Methodology

We called every system through its provider’s API with native structured output. We never prompted for JSON and did not retry schema failures. Comparisons pair systems on identical items. Intervals come from bootstrap resampling of whole filings, Bonferroni-corrected for eleven comparisons against Jev. We planned McNemar’s exact test as well and ran only the bootstrap.

We measured latency under concurrent load from one machine on a residential connection, with 8 Jev requests in flight against 4 per LLM. The judgments sweep ran sequentially, three repeats per cell. We added the 64-judgment cell after the preregistered design. Jev’s isolated single-call median is 0.10s, against 0.11s under load.
