cd /news/ai-agents/jigsaw-judgement-atomic-questions-as… · home › topics › ai-agents › article
[ARTICLE · art-144068] src=interestingengineering.substack.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Jigsaw Judgement: Atomic Questions, Assembled Decisions

A developer has published a detailed pattern book for combining the Jev fast decision model with Python rule engines and Claude for explanation generation, applying the architecture to eleven high-value decision workflows including corporate credit approval, anti-money-laundering alerts, trade finance, insurance claims, prior authorisation, venture screening, covenant monitoring, semiconductor assembly and test, M&A contract review, and university credit for prior learning. The writeup argues that model API costs are a rounding error — Jev is priced at $0.042 per million input tokens with no charge for output, putting a year of the anchor case at roughly eleven dollars — and that automatable volume, not model quality, determines whether automation pays back. It cites a Carnegie Mellon study dated 29 September 2026 that tested Jev against sixteen other judges, finding it within three points of the strongest reasoning model where a verdict can be read off the text, at 0.36% of that model's fee, and behind it where the verdict must be derived.

read35 min views2 publishedOct 2, 2026
Jigsaw Judgement: Atomic Questions, Assembled Decisions
Image: Interestingengineering (auto-discovered)

That 5% problem #

Yesterday, I gave a simple example of an Enterprise Banking Loan approval process (part of a far wider integrated chain, but taking a sliver of that chain was “fit for purpose” of showcasing how these two models - Jev and Claude - might come together) here:

In this article, I will now be adding a more thorough review and expanding the scope to showcase possibilities within high-value industries or well-defined segments of highly regulated sectors.

To remind of the scenario, Alpha Global Holdings wants a credit facility. Its ratios are the kind a lender frames and hangs on the wall: debt service covered 2.8 times, net leverage at 1.5, a current ratio of 2.1. Sixty per cent of revenue comes from cloud software. Looking good! Then there is Division 4. 5% of revenue contribution, coming from a sector the lender may hesitate to take exposure to. Say as a result of concentration risk, or limitation of risk exposure. The way Alpha Global Holdings classfied it: “** consulting and transport infrastructure for offshore deepwater oil extraction**”.

A keyword filter looking for “oil” or “gas” would catch it (to facilitate segmentation analysis), but what of a business that classifies a segment slightly differently? Writing “subsurface hydrocarbon logistics”. A seasoned analyst likely catches both, at a cost of hours per file. My note to a banking friend, “Enterprise Banking: Jev, Python, Claude Implementation Example”, showed a different route: a fast decision model (Jev) reads the sentence, Python applies the bank’s rules, and Claude writes the explanation only when the file needs one.

This guide takes that example apart (again, but much refined), and extends it to 10 other examples (including the lending limit to oil and gas). Each rebuild is a different industry with money on the line. Numbers are based on assumptions that can vary by case: anti-money-laundering alerts, trade finance, insurance claims, prior authorisation, venture screening, covenant monitoring, a semiconductor assembly and test plant, M&A contract review, and university credit for prior learning. In all cases, every one carries an assumption-based costed build, a maintenance bill and a payback worked line by line, so a relevant officer or committee can check the arithmetic before running a line of code. These are provided to illustrate how decisions might be made. They in no way represent circumstances of any particular business. “Automatable-volume” has a big role in determining potential for application.

Consider this a more practical look at our Pattern Book for Agent Design (with Jev):

Two findings shape everything that follows. First, the model bill will typically look like a rounding error. Meaning, with Jev, if a big part of the process flow is “automatable”, then the bill you pay will likely be very low. Jev charges $0.042 per million input tokens and nothing for output [3]; a whole year of the anchor case costs about eleven dollars in API calls. The real money lives in analyst minutes (or hours). Second, volume decides. The example that started this guide, corporate credit at 2,000 files a year, is the weakest business case of the ten. So volume is critical. Analyzing 100,000s+ or Millions of files keeps many analysts busy. Too few, however, and it’s hard to justify automation, in any form. Prior authorisation and insurance claims, at hundreds of thousands of cases, pay back within a quarter on my assumptions.

Independent evidence now backs the design. References are found further below. A Carnegie Mellon study published on 29 September 2026 tested Jev against sixteen other judges and found it within three points of the strongest reasoning model wherever a verdict can be read off the text, at 0.36% of its fee, and behind wherever the verdict must be derived [25]. Its recipe for setting thresholds on about 100 local labels is now the calibration loop in Part III.

How to read the numbers

Known: Jev and Claude prices, Jev’s documented limits and failure modes, the independent JEV-as-a-Judge results [25], and the industry baselines cited in brackets, which point to the numbered references at the end. Some baselines come from vendors selling automation; those are flagged in the text.

Assumed: volumes, path shares, minutes saved, build weeks and loaded hourly rates, in US dollars. Two cost settings recur: a high-cost market and a lower-cost market. Every assumption sits in the economics table for its deep dive, so you can swap in your own and redo the sum.

Unknown: how Jev performs on your documents. Nothing here replaces a shadow-mode pilot on your own labelled cases.

Tested: Appendix A sets out each assumption, its basis and its strength, what is left out, and how the verdicts move under ten scenarios. An organisation building its own case may classify the same industry differently.

Part I · The jigsaw #

Strip the example to its parts, and I have eight pieces to consider. Some are models, some are code, one is a person with a pen. The skill lies in the order.

Figure 1. The eight pieces. The colours repeat in every decision tree that follows.

Terms, and why each matters #

Jev is good at what a knowledgeable person could judge in a few seconds, and TypeSafe’s own documentation says so: decompose broad judgments into atomic questions and combine them in code [1]. Its published list of weak spots is candid. It reads instructions literally, does not count reliably, reads dates as text, loses accuracy when the state fills with irrelevant material and can be steered by adversarial text [4]. Each of those weaknesses maps to a piece of the jigsaw. Arithmetic, dates and counts go to Piece 1. Filtering happens before the call. Steering it is caught by the loop.

Part II · The anchor: “Enterprise Banking: Jev, Python, Claude Implementation Example” #

The original note gets the architecture right. A cheap, fast gate reads every file; hard rules decide the obvious cases; Claude only runs on the contested ones; a committee votes. That shape survives every rebuild in this guide. The details need work before the note goes near a production system (hence why i initially issue it as a note), and most of the fixes come straight from TypeSafe’s documentation. So the original structure has been further fine-tuned.

Twelve changes before production #

Item 7 is the one to explain to a loan officer. The original pipeline asks Jev for “the probability of violating concentration limits” and gets 72%. That number mixes two questions. Is Division 4 in a restricted sector? That is a judgment about words, and Jev is built for it. Does a 5% revenue slice breach a 10% sector cap? That is arithmetic against a policy the bank already had in place. Split them and each piece does the job it is good at. The committee then sees an exposure figure it can audit, instead of a probability it has to trust.

Figure 2. Alpha Global Holdings through the rebuilt cascade. Chips 1 to 4 follow one file; the dashed frame is the harness; the gold loop retunes the router monthly.

The rebuilt code #

About forty lines. Field names follow TypeSafe’s quick start [2]; check them against the SDK reference before you run anything, because a young SDK moves.

Run Alpha Global through it. The ratios pass in step 1. Jev flags Division 4 alone, so exposure is 5%. Five per cent is above zero and below the 10% cap, so the file goes to committee with a memo. If the Choice comes back unsure whether Division 4 owns rigs or sells software to people who do, the file goes to a credit officer instead. The committee’s question, “is 5% of consulting revenue a problem?”, remains their responsibility!

What my own bench already showed

In “The Judgment Line” I ran Jev against a generative model on eight trap announcements and three decision rules, twice. Jev scored 24/24 both times and its cost matched to the last digit across runs ($0.000261); the generative model scored 22/24 both times, failing the same two rule-application texts, and its cost and latency moved between runs. Determinism, more than price, is the property a regulated lender is buying. Eight texts is a small bench, and your documents are not my documents.

What an independent study found #

Eight texts is my bench. JEV-as-a-Judge, from Yubo Li, Yidi Miao, Ramayya Krishnan and Rema Padman at Carnegie Mellon, is the first large outside test [25]. They ran Jev against sixteen generative and reward-model judges on 5,172 base judgments, audited disputed labels with blinded human reviewers, and froze their routing rule before looking at held-out results. Then they ran a live test on two workloads nobody had seen.

Two cautions. The authors tested public benchmarks, and say specialised professional domains are untested; every deep dive here is a professional domain, so the shadow pilot is not optional. And they close on a line any risk committee can adopt: confident errors argue against using any automated judge as the sole arbiter of consequential decisions [25]. For that reason, where relevant I have included chains of responsibility which must still remain with a person of responsibility.

Part III · The money, and the maintenance #

Per case: the model bill #

Work the anchor case one file at a time, at the prices published (in my case) on 30 September 2026 [3][5].

  1. Jev on every file. State plus four questions comes to about 2,500 input tokens. 2,500 × $0.042 ÷ 1,000,000 = $0.000105. Output is free.

  2. Claude on the escalated 20%. A memo reads about 6,000 tokens and writes about 1,500. On Sonnet 5.5 ($2 in, $10 out per million): 6,000 × 2 ÷ 1,000,000 = $0.012, plus 1,500 × 10 ÷ 1,000,000 = $0.015, so $0.027 a memo.

  3. Blended. $0.000105 + 20% × $0.027 = $0.0055 a file. Claude on every file would cost $0.027. Routing saves $0.0215 a file, or $21.50 per thousand.

  4. Analyst time. On my assumptions the cascade saves 3 hours on the 70% of files it fast-tracks, 1.5 hours on the 20% where Claude drafts the memo, and 1 hour on the 10% that fail ratios before anyone opens them: 0.7 × 3 + 0.2 × 1.5 + 0.1 × 1 = 2.5 hours a file. At a $90 loaded rate that is $225 a file.

Figure 3. Two scales, on purpose. Routing trims the model bill by dollars; moving analyst hours is worth hundreds of thousands.

Every deep dive repeats this shape. Model spend lands between $11 and about $1,800 a year even at half a million cases. Design effort belongs on the human minutes, and on making sure the cases that skip people deserve to.

Set-up: what a build costs #

Three lines, the same in every deep dive. Worked here for the anchor case.

Maintenance #

Freed hours are not cash until someone stops hiring, reassigns a team or takes on more volume with the same staff. I count half of freed time as cash in year one. For the anchor case: 2,000 files × 2.5 hours = 5,000 hours; × $90 = $450,000; half is $225,000; minus $136,611 of running costs leaves $88,389 a year. Payback: $246,400 ÷ ($88,389 ÷ 12) = 33.5 months. The volume at which the same build pays back inside a year: ($246,400 + $136,600) ÷ (2.5 h × $90 × 50%) = about 3,400 files. The reading for my banking friend

At 2,000 files a year the anchor case is a good idea and a slow payback. It works as a shared service across business lines, as an add-on to a spreading platform the bank already runs, or as a pilot that proves the pattern before the bank points it at AML alerts, where the volume is fifty to a hundred times higher.

Is 80% really 80%? The calibration loop #

The original note says that if Jev flags 100 cases at 80%, exactly 80 must be real. Close, but sampling noise also applies. The standard error of a proportion is the square root of p × (1 − p) ÷ n. For p = 0.8 and n = 100: 0.8 × 0.2 = 0.16; 0.16 ÷ 100 = 0.0016; the square root is 0.04. Double it for a 95% band: ±0.08. So anything from 72 to 88 real cases is consistent with a well-calibrated model. Seventy real cases is a warning. Forty, the note’s example, means the line is wrong for your documents.

Choosing the line: the pessimistic rule

The JEV-as-a-Judge authors give a recipe for setting the line on about 100 labelled cases from your own workload [25]. Run the cascade and the full human path on the same cases. For each case write d = −1 if the fast path got it wrong where the full path got it right, and 0 otherwise. Then take the loosest line, the one that fast-tracks the most cases, whose pessimistic loss stays within two points.

  1. Line at 0.9. Of 100 labelled cases, two fast-tracked cases were wrong. Average loss d̄ = −2 ÷ 100 = −0.02. The point estimate just passes a two-point tolerance.

  2. Spread. Sum of squared deviations = 2 × (−1 + 0.02)² + 98 × (0.02)² = 1.9208 + 0.0392 = 1.96. Divide by n − 1 = 99: 0.0198. Square root: sd = 0.141.

  3. Pessimistic loss. d̄ − 1.645 × sd ÷ √n = −0.02 − 1.645 × 0.141 ÷ 10 = −0.02 − 0.023 = −0.043. A possible loss of 4.3 points fails the two-point test.

  4. Line at 0.95. No fast-tracked case was wrong, so d̄ = 0 and the bound is 0. It passes. More cases go to people, and the bill rises a little.

The study found the same pattern on real data: a 0.9 line lost exactly two points and passed on the point estimate, the bound came out at −4.3, and the rule moved to 0.95 [25]. With about 100 labels the pessimistic rule cut the chance of losing more than two points from about 45% to about 5%. Set the line for one pinned model version and one question type; a line tuned on a Noul does not carry over to a Choice [4][25].

Figure 4. Illustrative reliability check (left) and the five-step monthly loop (right). The red point is the pattern to hunt: high probabilities that come true less often than they claim.

The loop in Figure 4 costs between about $7,000 and $110,000 a year across the ten deep dives, counting audits and re-validation. It is also the part regulators ask about. Banking supervisors’ model-risk guidance expects ongoing monitoring and outcomes analysis for any model used in decisions, and SR 11-7 is the most widely cited example [22]; this loop is that, written down. The most developed AI statute so far classes credit scoring of individuals, life and health insurance pricing and education admissions as high-risk, with obligations applying from 2 December 2027 [20][21]. Other jurisdictions are moving the same way, so check your own. Alpha Global is a company, so its facility sits outside the credit-scoring category for individuals.

Three rules that hold in every industry

Automate the yes, not the no. Where a wrong refusal harms a person (a claim, a treatment, a student’s credit), the model may approve or route, and a human signs every refusal.

Treat applicant text as hostile. TypeSafe lists adversarial content as a known weakness [4], and the independent study found confidence routing weakest when a wrong answer is more elaborately written [25]. Anyone who learns that “we are a software company” moves a Noul will write it. Test with planted texts; watch for drift in the loop.

Log the version with every answer. An auditor who cannot replay a decision will not accept it.

Part IV · 10 deep dives #

Each deep dive follows: with a sourced baseline, the decision tree, the pieces used, the economics worked line by line, where it breaks etc. Shares of cases on each path and minutes saved are my assumptions, stated in the tables. You can use these as starter templates and then later adapt them to your own circumstances. After all, every business is unique.

Deep dive 1 · Corporate credit #

Banks spend real money to put a commercial loan on the books. One correspondent-banking analysis put the average origination cost at about $11,319 for a commercial real estate loan averaging $614,000 [6]. A commercial file can absorb 8 to 20 staff hours across collection, spreading, analysis, memo writing and review (vendor estimate) [7]. Sector screening is a small slice of that work, and the slice most exposed to phrasing tricks.

Figure 2 is this deep dive’s tree. The build extends naturally: the same Nouls can screen against a bank’s full restricted-sector list, pinned to a standard industry classification such as ISIC [24] or NAICS [23], and the same exposure arithmetic can run for every sector cap the bank holds.

PIECES Code (ratios, exposure) · Jev (one Noul per division, one Choice) · Router · Claude Sonnet 5.5 (committee memo) · Credit committee · Harness · Monthly loop · Ledger

The economics, worked

Where it breaks

• Volume. At 2,000 files a year payback runs 33 months, and 116 months if the fast-track share halves. Below about 3,400 files a year, share the build.

• Phrasing games. Applicants learn what flags. Plant paraphrased restricted activities in the audit set every month.

• Division boundaries. If the application does not split revenue by division, code cannot compute exposure and every file needs a person.

First move

Run the pipeline in shadow mode on the last 300 closed files with known sector outcomes. Measure how many Jev would have fast-tracked and how many of those an analyst later flagged. Simple enough to understand.

Deep dive 2 · AML alert triage #

Rule-based transaction monitoring is famous for noise. Between 90% and 95% of alerts turn out to be false positives, a figure traced to PwC analysis and repeated across the industry [8][9]. Manual review costs roughly $25 to $50 an alert at mid-size institutions, and fewer than 5% of alerts become suspicious activity reports (practitioner estimates) [8].

The design keeps every human decision. Jev reads the free text a rule cannot (wire memos, counterparty names, the customer’s stated business) and code bands the alert. Claude drafts. The analyst signs.

Figure 5. An alert from firing to a signed decision. Code owns counts and velocity; Jev reads whether the story fits the customer.

PIECES Code (lists, counts, velocity) · Jev (three Nouls, one Choice, one Score) · Router (bands) · Claude Haiku 4.5 (closure rationale, context) · Analyst · Harness · Below-the-line loop · Ledger

The economics, worked

Where it breaks

• Tuning debt. Examiners treat an undocumented threshold as a finding. The loop’s paperwork is part of the product.

• False negatives. A quick-close band that swallows a real case is the costly error. Sample low-band closures hardest, not escalations.

• Language. Jev is strongest in English [3]; cross-border wire memos often are not. Route non-English memos to the medium band until the audit says otherwise.

First move

Shadow-band six months of closed alerts with known outcomes. The number to report is the SAR rate inside the low band. It should be near zero, with an error band to prove it.

Deep dive 3 · Trade finance and dual-use goods #

Documentary credits remain paperwork-heavy. First-presentation discrepancy rates of 60% to 70% are widely cited [11], and one vendor puts manual examination at two to three hours for a standard credit (vendor estimate) [10]. UCP 600 gives the examining bank five banking days to decide [11].

Most discrepancies are clerical and belong to code: amounts, dates, tolerances, ports. Jev takes the part that needs a reader: whether the invoice describes the same goods as the credit, whether a bill of lading notation declares damage, and whether the goods might sit in a dual-use category.

Figure 6. A presentation inside the five-day window. Code checks the numbers; Jev reads the descriptions; compliance owns any dual-use hit.

PIECES Code (UCP 600 arithmetic) · Jev (three Nouls, one Choice) · Router · Claude Sonnet 5.5 (discrepancy notice) · Examiner and trade compliance · Harness · Refusal-audit loop · Ledger

The economics, worked

Where it breaks

• OCR quality. Jev takes text only [3]. A smudged bill of lading becomes a bad state before any model sees it.

• Dual-use is a regulatory call. A Noul can raise the flag. Classification under export-control lists stays with trade compliance.

• Examiner trust. Discrepancy practice rests on examiner experience. If drafted notices get rewritten every time, the saving disappears; track the edit rate.

First move

Take 500 historical presentations with the examiner’s final discrepancy list. Score the drafted notices field by field against the examiner’s calls before any live use.

Deep dive 4 · Insurance claims at first notice #

Claims handling is the biggest operating cost an insurer controls. One industry analysis puts loss adjustment expense at 10% to 12% of earned premium for personal auto, and contrasts a traditional carrier’s $185 to $240 handling cost on a simple auto claim with far lower marginal costs at automated insurers (vendor analysis) [12].

The fast path is small, clear claims: policy in force, estimate below a line, liability clear, no injury, no fraud markers. Code owns the money. Jev reads the claimant’s own account, where injuries and inconsistencies hide.

Figure 7. A claim from the first call to the right desk. The injury flag carries the lowest threshold because a missed injury is the expensive error.

PIECES Code (coverage, estimate, counts) · Jev (three Nouls, one Choice) · Router · Claude Haiku 4.5 (file summary) · Adjusters and SIU · Harness · Leakage and fairness loop · Ledger

The economics, worked

Where it breaks

• Unfair-practice exposure. A fraud flag must open an investigation, never a denial. Monitor outcomes by customer group in the loop.

• Sideways injuries. Claimants mention pain late and casually. Re-run the injury Noul on every later contact, not just first notice.

• Photos. Jev reads text only [3]. Damage photos need a separate vision step that writes a text description first.

First move

Shadow six months of closed simple claims. Report how many would have gone straight through, the paid-amount difference against the adjuster’s actual settlement, and the reopen rate.

Deep dive 5 · Prior authorisation, payer side #

Prior authorisation, where an insurer approves a treatment before it happens, is one of the most expensive routine transactions in healthcare administration. The largest published benchmark put a manual request at $10.97 to the provider and $3.52 to the payer, against $0.05 for a fully electronic one on the payer side [13]. A manual request takes about 24 minutes of provider staff time, and in 2024 only about 40% of requests in that benchmark ran fully electronically [14].

Payer-side clinical review is where Jev fits: matching clinical notes to written coverage criteria, one Noul per criterion. The rule that governs the design is simple. Automation may approve. Only clinicians may deny.

Figure 8. A request from intake to decision. No path leads to an automated denial.

PIECES Code (code lists, eligibility, clocks) · Jev (one Noul per criterion) · Router (approve-only) · Claude Haiku 4.5 (criteria map for the clinician) · Nurses and physicians · Harness · Approval-accuracy loop · Ledger

The economics, worked

Where it breaks

• Policy drift. Plans rewrite coverage criteria. Each rewrite resets calibration; the loop has to fire on policy changes as well as monthly.

• Long records. Clinical notes can exceed Jev’s 32k-token state budget [3]. Retrieve the relevant notes in code first.

• Regulation. Where individuals’ health cover is priced or decided, sector rules apply, and AI statutes increasingly class health-cover decisions as high-risk [20].

First move

Pick the three highest-volume procedure codes. Shadow 2,000 historical requests and report auto-approval precision against clinician decisions, with its error band.

Deep dive 6 · Private markets deal screening #

Venture firms wade through volume to make a handful of bets. In the largest survey of institutional VCs, firms considered roughly 100 opportunities for every deal they closed, and about one in four reached a meeting with management [15].

Labour savings here are modest because volumes are low. The case rests on speed of response and coverage: every deck read against the thesis within a day, and a courteous, specific reply to every founder.

Figure 9. An inbound deck from inbox to partner. The loop audits declines, because a missed winner costs more than a wasted meeting.

PIECES Code (mandate rules, weights) · Jev (two Scores, one Noul, one Choice) · Router (composite score) · Claude Haiku 4.5 (declines) and Sonnet 5.5 (memos) · Partners · Harness · Quarterly miss audit · Ledger

The economics, worked

Where it breaks

• Pattern-matching bias. Scores trained on what partners liked before will reproduce what partners liked before. The quarterly miss audit is the counterweight.

• Deck theatre. Founders write to be scored. Treat claims of recurring revenue as unverified until data-room evidence arrives.

• Scale. At 3,000 decks a year this is a four-week configuration job, not a platform.

First move

Replay last year’s inbound decks. Check where this year’s portfolio companies would have landed, and how many declined companies later raised from strong funds.

Deep dive 7 · Covenant and project-finance monitoring #

Lenders to data centres, power and infrastructure receive a steady stream of compliance certificates and construction reports. Ratios arrive late; warning signs arrive early, in prose. A slipped energisation date or a renegotiated offtake shows up in a monitor’s report months before debt service coverage moves.

This is the thread from my datacenter-credit study: capex carry depends on dates and counterparties that live in narrative, not in spreadsheets.

Figure 10. A reporting pack read for what the ratios have not shown yet. Code recomputes every covenant; Jev reads the narrative.

PIECES Code (covenant arithmetic, headroom) · Jev (four Nouls) · Router (weights by exposure and stage) · Claude Sonnet 5.5 (watchlist note) · Credit officer and committee · Harness · Quarterly hindsight loop · Ledger

The economics, worked

Where it breaks

• Long packs. Construction reports run to dozens of pages. Chunk them and ask the Nouls per section to stay inside the state budget [3].

• Rare events. Defaults are few, so calibration data is thin. Pool the hindsight audit across portfolios or across lenders.

• Borrower drafting. Sponsors learn which phrases trip flags. Keep criteria private and refresh them.

First move

Take ten facilities that went on watch in the past three years. Run their earlier packs and record how many quarters of warning the flags would have given.

Deep dive 8 · Semiconductor assembly and test #

Semiconductor assembly and test plants run on tickets: shift logs, operator notes, customer complaints and failure-analysis requests. At leading-edge fabs, a single wafer can cost $16,000 to $22,000 to process, and root-cause analysis for a yield excursion commonly takes two to five days (industry case estimate) [16]. Back-end economics differ, but the lesson carries: the cost of an excursion grows every hour it runs.

Numbers belong to statistical process control, already in code. Jev reads the words: which failure mode a note describes, whether an automotive customer reports a field failure, and whether a note matches the symptoms of an open excursion.

Figure 11. A floor ticket from note to containment. Language is itself a question, because Jev is strongest in English.

PIECES Code (SPC, genealogy) · Jev (one Choice, three Nouls) · Router · Claude Sonnet 5.5 (8D first draft) · Process and quality engineers · Harness · Taxonomy loop · Ledger

The economics, worked

Where it breaks

• Mixed-language notes. Several languages in one line is normal on many floors. Measure accuracy by language and route non-English notes to an engineer until the loop says otherwise.

• Taxonomy churn. Failure classes change with each new package. Criteria text has to move with them.

• The real prize is outside the table. The economics count ticket minutes only. One excursion caught a shift earlier can exceed the whole year’s benefit, but I have not counted it.

First move

Relabel 1,500 historical tickets with the plant’s current taxonomy, run Jev in shadow, and report class accuracy by language and the recall of automotive field-failure flags.

Deep dive 9 · M&A contract triage #

A mid-sized acquisition hands buyers a data room of 5,000 to 20,000 documents, often with two to three weeks to review them, and contract reviewers bill $45 to $100 an hour (vendor estimates) [17]. Teams sample because they cannot read everything.

Jev asks the same literal questions of every contract: change of control, assignment consent, exclusivity or most-favoured-nation rights, non-competes. Code parses and chunks. Claude quotes the flagged clause with a page reference.

Figure 12. A data room from upload to a complete clause index. The change-of-control line sits lowest because a miss costs more than a false alarm.

PIECES Code (parsing, chunking, dates) · Jev (one Noul per clause type) · Router · Claude Sonnet 5.5 (clause extract) · Associates and partners · Harness · Seeded-contract loop · Ledger

The economics, worked

Where it breaks

• Recall is the only metric that matters. A missed change-of-control clause can cost more than the whole review. Seed every data room with known clauses and report recall per type.

• Scans. Poor OCR turns into poor state. The 5% unreadable path is an estimate; old targets run higher.

• Privilege and confidentiality. Enterprise data terms matter here. TypeSafe offers zero data retention to enterprise customers [3]; confirm the same for every model in the chain.

First move

Run the pipeline on one closed deal’s data room where the final disclosure schedule is known, and compare the index with the schedule clause by clause.

Deep dive 10 · Higher education: credit for prior learning #

Universities in many countries let working adults earn course credit for what they already know, under schemes usually called recognition of prior learning. UNESCO’s guidelines describe the common shape: the learner documents evidence, an assessor judges it against stated standards or learning outcomes, and successful claims are certified [18][19]. Each portfolio means an assessor matching work samples to outcomes, item by item.

The fit is exact: one Noul per evidence item and outcome, the institution’s pass rule in code, Claude drafting gap feedback, the assessor signing. The economics are thin where assessor rates are low, which is why this deep dive carries the verdict it does.

Figure 13. A prior-learning portfolio from submission to the assessor’s signature. The harness prepares; it never awards.

PIECES Code (eligibility, credit caps, the pass rule) · Jev (one Noul per item and outcome) · Router · Claude Sonnet 5.5 (gap feedback) · Assessors and credit panel · Harness · Moderation loop · Ledger

The economics, worked

Where it breaks

• Money. At $30 an hour and 10,000 portfolios the build takes eight years to pay back. It needs a shared service across institutions, or the value has to be counted in turnaround and consistency.

• Language. Portfolios often arrive in more than one language; Jev is strongest in English [3]. Items in other languages go to a full read until the loop measures otherwise.

• Authenticity. Whether evidence is the learner’s own is a judgment with consequences. Flag, never decide.

First move

Pilot one high-volume programme across one intake, with moderation comparing pre-filled maps to assessor outcomes. For an open and distance learning institution, pitch it as a shared service to peer institutions from the start.

Figure 14. Net annual benefit after running costs, half of freed time counted as cash. Ticks show the stress case where the fast path halves.

Read the table from the right. Build means payback inside a year on the base case and inside eighteen months if the fast path halves. Pilot first means the case holds but depends on assumptions worth testing in shadow mode. Share or wait means the volume cannot carry a single-institution build. The twelve- and twenty-four-month lines are my hurdle; Appendix A shows how the verdicts shift under other hurdles, build costs and captured shares.

The pattern is pretty simple, really. Volume and minutes per case decide the outcome; model prices barely register. The four builds each run more than 10,000 cases a year and save between 9 and 73 minutes a case. The anchor case saves the most time per file of any deep dive, two and a half hours, and still ranks seventh, because 2,000 files cannot carry $383,000 of build and first-year maintenance.

Adapting the jigsaw to a new industry #

Eight questions, to start. Answer them before anyone writes code.

  1. What is the case? One alert, one claim, one contract. Count them per year.

  2. What is the fast path worth? Minutes saved per case on the path that skips people. If it is under five, stop.

  3. What can code decide alone? Arithmetic, dates, lists, counts. Move all of it to Piece 1 before thinking about models.

  4. What judgments remain? Write each one as a single statement a knowledgeable person could judge in seconds. Those are your Nouls.

  5. Which way is the costly error? A missed injury, a missed clause, a missed winner. Set that line lowest and audit that path hardest.

  6. Who must sign? Name the role for every path that ends in a refusal. Automation approves or routes.

  7. What will the loop sample? Cases per month, minutes each, who labels them. Put the cost in the business case.

  8. What is the volume for a 12-month payback? (Build + maintenance) ÷ (minutes saved × rate × 50%). If your volume is below it, share the build or wait.

Appendix A · How the numbers were built, and how far to trust them #

Every figure in this guide comes from one small model with stated inputs. This appendix sets out what each input is, where it came from, how strong it is and what was left out, then shows how the verdicts move when the inputs move. The numbers are a worked method for a mid-sized operator. They are not a forecast for any particular firm. An organisation that builds its own case with its own volumes, rates, platform costs and hurdle rate may reach a different verdict for the same industry, and should trust its own numbers over these.

A.1 The method #

  1. Hours freed = cases a year × minutes saved per case ÷ 60. Minutes saved per case = Σ (share of cases on a path × minutes saved on that path).

  2. Cash captured = hours freed × loaded hourly rate × 50%.

  3. Running cost = Jev bill + Claude bill + maintenance (engineering share + monthly calibration audit + yearly re-validation).

  4. Net annual benefit = cash captured − running cost.

  5. Payback = build cost ÷ (net annual benefit ÷ 12). Simple payback: one steady-state year, no discounting, no tax, no ramp-up.

A loaded hourly rate is salary plus benefits, space and overhead, divided by working hours. Worked: a $120,000 salary × 1.35 overhead ÷ 1,800 working hours = $90, the rate used for the credit analyst. The other rates ($30 to $120) follow the same formula with role-appropriate salaries.

A.2 Each assumption, its basis and its strength #

A.3 Minutes saved, checked against the baselines #

Minutes saved are the softest input, so each one is checked against the sourced baseline where one exists. The test is simple: the saving should be a part of the measured handling time, not more than all of it.

Several baselines come from vendors selling automation, who have reason to report high manual costs. Where the guide’s saving sits inside a vendor baseline, the bias works against the case rather than for it.

A.4 Why the build cost sets the verdict #

Payback within twelve months means build cost ÷ (net ÷ 12) ≤ 12, which reduces to build cost ≤ one year’s net benefit. So the Build verdict is a direct test of the build bill against one year of net savings. Low-volume cases fail it because the build is a fixed cost: a calibrated harness, a labelled gold set and a validation review cost roughly the same at 2,000 cases as at 200,000.

Worked for the anchor case: net benefit $88,389 a year against a build of $246,400. To pass, the build must fall to $88,389, a cut of 1 − 88,389 ÷ 246,400 = 64%. Or volume must rise to about 3,400 files a year (Part III). Or the hurdle must change: at 33.5 months the case clears a three-year payback, and its three-year net is 3 × $88,389 − $246,400 = $18,767.

The same arithmetic shows where the money can come from. Halving both build and maintenance gives almost exactly the same payback as doubling the captured share, because both halve the ratio of cost to benefit. A firm that already runs a document pipeline, a model-risk function and an audit process is in the first position. It should expect a lower build and a better verdict than the tables show.

Table A1. The build test for every deep dive. Three-year net = 3 × net annual benefit − build. Under a three-year hurdle, nine of ten cases clear.

A.5 How the verdicts move #

Table A2 changes one input at a time, then two combinations: a shared platform that halves build and maintenance, and a downside in which only a quarter of freed time is captured and minutes saved come in 30% lower. Figure A1 draws the same table as ranges.

Table A2. Payback in months under each scenario. “Both down” is the downside: a quarter captured and minutes 30% lower. “Never” means running costs exceed captured savings; “240+” means more than twenty years.

Figure A1. Range of payback across the ten scenarios. Green: under 12 months. Gold: 12 to 24 months.

Three readings follow. The model bill barely matters: a tenfold rise changes no verdict, and outside the education case, already past eight years, it moves no payback by more than half a month. The captured share matters most for every case, which is why the guide halves it by default. And the downside breaks seven of the ten cases outright, so a firm that cannot redeploy freed time, or whose shadow pilot shows smaller savings, should expect only the highest-volume cases to pay.

A.6 What the model leaves out #

A.7 How to build your own case #

  1. Replace the volumes. Use last year’s actual case counts. This year’s management accounts work too.

  2. Replace the rates. Use your own loaded rates by role.

  3. Measure the paths. Run the pipeline in shadow mode on a few hundred closed cases. Count how many would have taken each path, and time the work saved.

  4. Price your build. Start from what you already own. A firm with document pipelines and model validation in place pays less.

  5. Choose your hurdle. Twelve months, three years or a net present value at your cost of capital. The verdict labels in this guide use twelve and twenty-four months; yours may differ.

  6. Price your errors. Estimate the cost of one wrong fast-track and multiply by the fast-path error rate the shadow pilot measures.

The qualifications that apply for all calculations or numbers:

The figures in this guide show a method and an order of magnitude. Prices and baselines are sourced; volumes, path shares, minutes saved, build and maintenance are stated assumptions for an illustrative mid-sized operator. The verdicts depend on a twelve-month hurdle that I chose. A different organisation, with different volumes, platforms or hurdles, can reasonably classify the same industry differently, and its own shadow-mode numbers should replace mine.

[1] TypeSafe. Introduction. [https://docs.typesafe.ai/introduction](https://docs.typesafe.ai/introduction)

[2] TypeSafe. Quick start. [https://docs.typesafe.ai/introduction/quickstart](https://docs.typesafe.ai/introduction/quickstart)

[3] TypeSafe. Models: pricing, rate limits, context, aliases, data handling. [https://docs.typesafe.ai/models](https://docs.typesafe.ai/models)

[4] TypeSafe. Jev 1.13 jaggedness. [https://docs.typesafe.ai/model-jaggedness/jev-1.13](https://docs.typesafe.ai/model-jaggedness/jev-1.13)

[5] Anthropic. Claude API pricing. [https://platform.claude.com/docs/en/about-claude/pricing](https://platform.claude.com/docs/en/about-claude/pricing)

[6] SouthState Correspondent. The Cost of a Commercial Real Estate Loan. https://southstatecorrespondent.com/banker-to-banker/commercial/cost_of_a_commercial_real_estate_loan/

[7] Crediflow. Credit Underwriting Automation: The Real Cost of Manual Review (vendor). https://www.crediflow.ai/blog/cost-of-manual-underwriting

[8] Fluxforce. False Positive Rates in Transaction Monitoring: 2024 Data (vendor). https://www.fluxforce.ai/statistics/false-positive-rates-transaction-monitoring

[9] DataVisor, citing PwC. End the False Positive Alerts Plague in AML Systems. https://www.datavisor.com/blog/guest-post-end-the-false-positive-alerts-plague-in-anti-money-laundering-aml-systems

[10] Finantrix. What Is Trade Finance Automation? (vendor). https://www.finantrix.com/articles/what-is-trade-finance-automation-letters-of-credit-bank-guarantees

[11] FloridAI. Letters of Credit and the 70% Discrepancy Rate. https://floridai.agency/blog/letter-of-credit-discrepancy-rate-ai [12] Finantrix. Claims Automation: FNOL to Settlement (vendor). https://www.finantrix.com/in-focus/next-gen-digital-insurer-pc-transformation/claims-automation-fnol-to-settlement

[13] 4sight Health. The Costly Lever of Prior Authorization (CAQH 2023 Index figures). https://www.4sighthealth.com/the-costly-lever-of-prior-authorization/

[14] Vstorm. The hidden cost of manual prior authorization (CAQH 2023 to 2025 figures). https://vstorm.co/agentic-ai/the-hidden-cost-of-manual-prior-authorization-a-framework-for-calculating-your-automation-roi/

[15] Gompers, Gornall, Kaplan and Strebulaev. How Do Venture Capitalists Make Decisions? NBER Working Paper 22587. https://www.nber.org/papers/w22587

[16] Case-studies.ai. Semiconductor fab yield optimization. https://case-studies.ai/use-cases/production-and-quality/PQ-005-semiconductor-yield-optimization/

[17] Groath.ai. The Real Cost of Manual Document Review in 2026 (vendor). https://groath.ai/strategy/ai-document-review [18] UNESCO Institute for Lifelong Learning. UNESCO Guidelines on the Recognition, Validation and Accreditation of the Outcomes of Non-formal and Informal Learning. https://uil.unesco.org/lifelong-learning/recognition-validation-accreditation/unesco-guidelines-recognition-validation-and

[19] UNESCO Institute for Lifelong Learning. Recognition, Validation and Accreditation of Non-formal and Informal Learning in UNESCO Member States. https://uil.unesco.org/lifelong-learning/recognition-validation-accreditation/recognition-validation-and-accreditation-non

[20] Gibson Dunn. EU AI Act Omnibus Agreement: Postponed High-Risk Deadlines. https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/

[21] Usercentrics. EU AI Act Deal: Digital Omnibus Now in Force. https://usercentrics.com/knowledge-hub/eu-ai-act-high-risk-delay-article-50-transparency-consent/

[22] Board of Governors of the Federal Reserve System. SR 11-7: Guidance on Model Risk Management. https://www.federalreserve.gov/supervisionreg/srletters/sr1107.htm

[23] US Census Bureau. North American Industry Classification System. https://www.census.gov/naics/ [24] United Nations Statistics Division. International Standard Industrial Classification of All Economic Activities (ISIC). https://unstats.un.org/unsd/classifications/ISIC/revision

[25] Li, Miao, Krishnan and Padman. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure. arXiv:2609.26550v3, 29 September 2026. https://arxiv.org/abs/2609.26550

── more in #ai-agents 4 stories · sorted by recency
── more on @jev 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/jigsaw-judgement-ato…] indexed:0 read:35min 2026-10-02 · —