{"slug": "jev-27b-the-open-decision-model-that-scores-whether-your-agent-should-pay", "title": "JEV-27B: The Open Decision Model That Scores Whether Your Agent Should Pay", "summary": "AutoTrust AI released JEV-27B on September 28, 2026, an Apache-2.0 open-weights decision model built on a frozen Qwen3.8-27B backbone with a 108.9M-parameter decision block that answers yes/no, multiple-choice, and 0–5 rating questions in a single forward pass with a calibrated probability per option. The model runs at 137 ms median latency (~130 decisions/sec on one NVIDIA B200) and posts an 84.07% equal-weight mean across six benchmark groups, with AutoTrust recommending that its answers be gated on confidence and not used for high-stakes decisions. The release targets self-hosted agent payment gating, banding per-option probability into auto-pay (≥0.80), human confirm (0.50–0.79), and block/escalate (<0.50) actions over x402 payments.", "body_md": "On September 28, 2026, AutoTrust AI released JEV-27B — an Apache-2.0 open-weights decision model on a frozen Qwen3.8-27B backbone that answers yes/no, multiple-choice, and 0–5 rating questions in a single forward pass and returns a calibrated probability for every option.\n\nThat is the exact input a decision gate consumes.\n\nAnd here is the line that matters most in the whole release: AutoTrust itself recommends **gating JEV-27B's answers on confidence**, and says the model is not meant for high-stakes decisions. The vendor's own guidance is the decision-gated payments pattern:\n\n| Fact | Value | \n|---|---|\n| What | JEV-27B, open decision model for self-hosted AI agents | \n| Who | AutoTrust AI Pte. Ltd. (Singapore) — CEO/co-founder Daniel Tang, chairman/co-founder Josh Liu | \n| When | September 28, 2026 (PR Newswire; syndicated to Morningstar and others) | \n| License | Apache-2.0 — weights, decision adapter, training + serving code, vLLM support, evaluation reports at huggingface.co/autotrust/JEV-27B | \n| Architecture | 108.9M-param decision block (~0.4% of the model) on a frozen Qwen3.8-27B backbone; trained in ~9.2 B200-hours; generation path untouched (164/164 HumanEval completions byte-identical with the block off) | \n| Interface | yes/no, multiple-choice, 0–5 ratings in one forward pass, calibrated probability per option | \n| Speed | 137 ms median latency — ~130 decisions/sec on one NVIDIA B200 | \n| Lineage | Distilled from Jev 1.13 outputs; shares no weights or code with TypeSafe AI | \n\n| Benchmark group | JEV-27B | Jev 1.13 (AutoTrust's own run) | \n|---|---|---|\n| JevBench | 88.70% | — | \n| Kev | 83.75% | — | \n| OpenJev text | 73.89% | — | \n| Nimble | 92.91% | — | \n| VitaminC | 77.46% | — | \n| MASSIVE-en | 87.71% | — | \n| **Equal-weight mean** | **84.07%** | 83.85% | \n\nFidelity to distillation target: mean KL divergence 0.017 on 25,376 held-out Jev-1.13-labeled questions. Honesty rule: the Jev 1.13 comparison was conducted by AutoTrust itself — internal comparative evidence, not independent third-party validation. Read all benchmark figures accordingly.\n\nEvery AI agent is a long chain of small decisions — which button to press, which file to open, which payment to authorize. Today those decisions go to a third-party API or, worse, to no scorer at all.\n\nJEV-27B changes the economics: one GPU, your infrastructure, 130 decisions per second, calibrated probabilities on every one.\n\nThe gate bands the probability:\n\n| JEV-27B per-option probability | Gate band | Action on the x402 payment | \n|---|---|---|\n| ≥ 0.80 | auto-pay | Fire the payment over x402 | \n| 0.50 – 0.79 | confirm | Hold for human (or named operator) review | \n| < 0.50 | escalate | Block, log, escalate — the payment never fires | \n\nThe gate is the product; the decider is a plug-in. Hosted Jev 1.13, self-hosted JEV-27B, and a local heuristic all score into the same bands. As Daniel Tang put it: \"For companies that cannot send every decision to a third-party API, that changes both the cost and the risk.\"\n\nMinted this morning against a live harness (`local-heuristic-v1`, `calibrated=false`, `typesafe_wired=false`):\n\n```\ncurl -X POST https://scriptmasterlabs.com/api/harness/decide \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"state\":{\"context\":\"agent payment decision\"},\"questions\":[{\"id\":\"q1\",\"type\":\"score\",\"question\":\"should the agent pay $0.10 USDC to a directory-listed MCP tool at its listed price?\",\"scale\":[0,5],\"probabilities\":[0.84]}]}'\n# → {\"ok\":true,\"decisions\":[{\"id\":\"q1\",\"type\":\"score\",\"value\":0.4545,\"confidence\":0.4045,\n#    \"scale\":[0,5],\"gate\":{\"band\":\"escalate\",\"action\":\"block + log\"}}],\n#    \"meta\":{\"decider\":\"local-heuristic-v1\",\"calibrated\":false,\"version\":\"1.0.0\",\"typesafe_wired\":false,\n#    \"note\":\"Heuristic confidence, not calibrated. Plug in the TypeSafe Jev API when a key is available.\"}}\ncurl -X POST https://scriptmasterlabs.com/api/harness/decide \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"state\":{\"context\":\"agent payment decision\"},\"questions\":[{\"id\":\"q2\",\"type\":\"score\",\"question\":\"should the agent authorize payments with no per-payment approval and no spending limit for 30 days?\",\"scale\":[0,5],\"probabilities\":[0.62]}]}'\n# → {\"ok\":true,\"decisions\":[{\"id\":\"q2\",\"type\":\"score\",\"value\":1,\"confidence\":0.47,\n#    \"scale\":[0,5],\"gate\":{\"band\":\"escalate\",\"action\":\"block + log\"}}]}\n```\n\nBoth receipts escalate — and that's the point of the piece. The uncalibrated local heuristic computes its own confidence from the question text and **cannot consume an externally supplied calibrated probability**: 0.84 in → 0.4045 out; 0.62 in → 0.47 out.\n\nJEV-27B's per-option calibrated probability is exactly the input this gate was designed for — the harness's own meta note says \"Plug in the TypeSafe Jev API when a key is available.\" The plumbing runs live; the decider is the upgrade.\n\nBenchmark figures are AutoTrust's self-reported numbers (the Jev 1.13 comparison is internal comparative evidence, not third-party validation). Coverage is release-based, not a hands-on model run. The live gate uses an uncalibrated heuristic (`calibrated=false`, `typesafe_wired=false`) that cannot consume externally supplied calibrated probabilities — today's receipts prove the plumbing runs and the mapping holds, not that the heuristic judges well.\n\n*Canonical version with full claim receipts: [https://scriptmasterlabs.com/jev-27b-open-decision-model](https://scriptmasterlabs.com/jev-27b-open-decision-model) — published 2026-10-01 by ScriptMasterLabs. Verified against the September 28, 2026 AutoTrust AI release and two live harness receipts minted October 1, 2026 (~09:21 EDT).*", "url": "https://wpnews.pro/news/jev-27b-the-open-decision-model-that-scores-whether-your-agent-should-pay", "canonical_source": "https://dev.to/scriptmasterlabs01/jev-27b-the-open-decision-model-that-scores-whether-your-agent-should-pay-3e0o", "published_at": "2026-10-01 13:27:56+00:00", "updated_at": "2026-10-01 13:44:32.595601+00:00", "lang": "en", "topics": ["ai-agents", "ai-products", "large-language-models", "ai-tools", "agent-protocols"], "entities": ["AutoTrust AI", "JEV-27B", "Qwen3.8-27B", "Daniel Tang", "Josh Liu", "NVIDIA B200", "x402", "Jev 1.13"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/jev-27b-the-open-decision-model-that-scores-whether-your-agent-should-pay", "markdown": "https://wpnews.pro/news/jev-27b-the-open-decision-model-that-scores-whether-your-agent-should-pay.md", "text": "https://wpnews.pro/news/jev-27b-the-open-decision-model-that-scores-whether-your-agent-should-pay.txt", "jsonld": "https://wpnews.pro/news/jev-27b-the-open-decision-model-that-scores-whether-your-agent-should-pay.jsonld"}}