{"slug": "show-hn-at0m-a-jev-like-system-one-decision-model-in-a-single-rust-binary", "title": "Show HN: AT0M – a Jev Like System One decision model in a single Rust binary", "summary": "AT0M, a decision model that returns one probability per option in a single forward pass, scored 0.789 on 2,000 typed decisions, ahead of Laya's 0.766 by 2.3 points and marginally outside the ±1.8-point sampling error, according to the project's Show HN post. The model ships as a single self-contained Rust binary with no Python at inference, and its v0 evaluation endpoint measured 0.056 expected calibration error versus Laya's 0.081 and Jev's 0.121–0.161 range. The author states calibration decays off-distribution, from 0.037 on seen spam to 0.171 on unseen phishing, and that AT0M cannot compare one field of a record against another, so those predicates must be computed in code.", "body_md": "Give it a **state** and a set of **questions** — choice, score, or yes-no. It returns **one probability per option**, in a single forward pass.\n\nNo tokens and no free text, so it **cannot answer outside your candidate set**. It runs as a single local executable, which keeps inference inside your perimeter.\n\n**Accuracy against what a decision costs.** A workflow is five typed decisions about one record. Competitor accuracies and prices are as published; ours is measured over all 2,000 test decisions and priced from a laptop’s measured throughput plus its amortised system cost. The full basis is in the report.\n\nAT0M will ship as a single self-contained Rust binary — no Python at inference, no service in the path, nothing to rent. That is the deployment, and it will be the only thing that ends up in your perimeter.\n\nThe hosted endpoint exists for one reason: so you can check every number on this page before you download anything.\n\n**v0 is the evaluation endpoint; v1 is the executable release** — the hosted service is a way to try the model, not a way to run it. Every figure on this page was measured against v0, the evaluation build, so they are the floor rather than the ceiling: v1 is expected to be more accurate, and we will republish every number on this page when it ships.\n\nOne configuration across every public test set — no per-dataset tuning or prompt tweaks. 38,862 decisions, scored in one sweep.\n\n| Benchmark | AT0M | Laya | Jev | Condition | \n|---|---|---|---|---|\n| typed-decisions (2,000) | 0.789 | 0.766 | 0.740 | ahead of Laya by 2.3pt, marginally outside the ±1.8pt sampling error | \n| calibration (ECE) | 0.056 | 0.081* | 0.121–0.161 | *Laya after fitting a temperature, and from its specialist checkpoint; Jev is the spread across four independent studies — TypeSafe publishes none. Jev and AT0M are general models, one configuration for every row; AT0M fits no temperature | \n| DAIR Emotion | 0.921 | 0.595 | 0.480 | bare labels | \n| Banking77 (77 intents) | 0.853 | 0.425 | 0.870 | Jev on 72 labels — not the same label space. Ours moves between 85% and 88% with how the intents are worded, so read the margin here as a band, not a point | \n| AG News | 0.896 | 0.950 | 0.910 | the one clean multi-way we trail | \n| CLINC150 + out-of-scope | 0.795 | — | — | 151 candidates, scored in full | \n| Enron spam / held-out phishing | 0.955 / 0.805 | — | — | phishing never trained on — the honest transfer test | \n| toxicity guardrail (F1) | 0.626 | — | — | 7.1% positive — answering \"no\" to everything scores 92.9% accuracy and catches nothing | \n\nLaya figures, and Jev's on the classification sets, are as reported in the Laya write-up. Jev's typed-decisions figure is DecisionEval's independent rerun. On calibration, TypeSafe publishes no figure at all, so Jev's range is assembled from four independent studies, none temperature-fitted.\n\nThe accuracy number is near the top but sits inside admitted label noise, so we don't hang the pitch on it. The claims we'd stand behind under scrutiny are architectural:\n\nAnd the parts we wouldn't: it compares a value to a threshold well, but not one field of a record against another, so compute those predicates in code. Calibration decays off-distribution — 0.037 on spam it has seen, 0.171 on phishing it hasn't. Both are measured below.\n\nEvery figure on this page came from a public test split scored through this endpoint. Point the same benchmark at it and check. One POST carries a record and as many questions as you like, all answered in the same pass — no SDK, no client library.\n\nWe'll help you scope a local-training or private-deployment evaluation against your own decisions — and share the full technical report.\n\nA comparative evidence inventory across AT0M, Jev 1.13.0 and Laya, carried from the technical paper. It is not a harmonized three-model benchmark: figures were obtained under different conditions, and each is tagged with its source and how it was measured.\n\n| Parameter | AT0M | Laya | Jev 1.13.0 | \n|---|---|---|---|\n| Developer | Pienomial | Convai Innovations | TypeSafe AI | \n| Delivery form | Standalone Rust binary; HTTP / CLI / JSON-file | Open weights; Python SDK / Jev-compatible HTTP server | Managed API | \n| Backbone | Not disclosed in report | ModernBERT-large 421M; mmBERT-base 322M | Not disclosed | \n| Primitives | choice · score · noul | choice · score · noul | choice · score · noul | \n| Inference | Non-autoregressive; no free text | Non-autoregressive; no free text | Non-autoregressive; no free text | \n| Multi-question call | Yes | Yes | Yes | \n| Max context | 32k request — the verification server caps requests at 4,096 tokens | 512 EN / 1,024 TD, multi 1,024 | 32k state + longest Q / 64k request | \n| Model selection | One configuration across all tests | Router: EN / multilingual / specialist | One managed version | \n\n| Measure | AT0M [S1] | Laya [L2] | Jev [I1] | \n|---|---|---|---|\n| TD overall | 0.789 | 0.766 specialist | 0.740 (0.727 older [B1]) | \n| TD choice | 0.755 | 0.733 | 0.738 | \n| TD score | 0.763 | 0.723 | 0.705 | \n| TD noul | 0.857 | 0.857 | 0.787 | \n| Invoice / service / obs / security | .846 / .786 / .764 / .758 | .804 / .764 / .730 / .766 | .786 / .788 / .640 / .744 | \n| AG News | 0.896 described | 0.950 routed | 0.910 (secondary) | \n| DAIR Emotion | 0.921 bare | 0.595 routed | 0.480 (secondary) | \n| Banking77 | 0.853 / 77 labels | 0.425 / 77 labels | 0.870 / 72 labels ≠ | \n| Held-out phishing | 0.802 (excluded from training) | — | — | \n| CLINC150 + out-of-scope | 0.794 / 151 choices | — | — | \n| Base vs specialist | One config throughout | General EN checkpoint 0.362 on TD | Zero-shot general weights | \n\n| Measure | AT0M | Laya | Jev | \n|---|---|---|---|\n| TD ECE (10-bin) | 0.057 no temp fitted [S1] | 0.213 raw specialist [L2] | 0.121–0.161 across four independent studies | \n| Per-type ECE | choice .052 · score .086 · noul .057 | 0.081 after temperature [L1] | TypeSafe publishes none | \n| Brier | Not provided | 0.062 specialist | 0.148 independent | \n| Confidence gating | Full coverage curve not published | No matched curve | ≥0.7 → 57.6% kept @ 86.4%; ≥0.9 → 24.5% @ 93.7% | \n| Option-order | 98% stable | Not established | Not established | \n| Counterfactual | 13/15 directional; identifier-repeat pair fails | — | — | \n| Empty record | Confidence falls 0.723→0.616; still no abstention | Not measured | Not measured | \n\n| Metric | AT0M | Laya | Jev | \n|---|---|---|---|\n| 5-question, T4 | 82 ms server time | 84.5 ms EN / 40.1 ms multi | 687 ms p50, network incl. [I1] | \n| 1-question, T4 | 16 ms T4 server; 22.1 ms M1 Max | 39.5 ms EN / 32.8 ms multi | older 236–276 ms hosted | \n| Throughput | M1 Max ~41→54 dps (1→32 clients) | 103–332 q/s batched, T4 | Provider-managed | \n| 77 / 151 candidates | 92.7 ms / 165.5 ms, M1 Max | Budget competition on options | Not matched | \n| Hardware | CPU · Apple Metal · CUDA | CPU / GPU self-host | Provider-managed | \n| Network | Server figure excludes round-trip | Inference; network not included | 687 ms explicitly end-to-end | \n\n| Parameter | AT0M | Laya | Jev | \n|---|---|---|---|\n| Weights | Training emits own weights; distribution terms per engagement | Apache-2.0 downloadable | Proprietary; no download | \n| Self-hosting | Standalone executable | Python / HTTP server | Not offered publicly | \n| Data transmission | Local inference stays in perimeter | Local inference stays in perimeter | Sent to provider unless contracted otherwise | \n| Version management | Export validation; customer handles rollout | Pin model hashes / runtime | Versioned IDs / alias pinning | \n| Audit trail | Raw probabilities; app owns the log | Opt-in prediction hooks | Version returned; app logs decisions | \n\n| Item | AT0M | Laya | Jev | \n|---|---|---|---|\n| Published price | Commercials will be released with v1. What do you think they should be? Please let us know at [info@pienomial.com](mailto:info@pienomial.com?subject=AT0M%20pricing) | Apache-2.0; hosting not free | $0.042 / 1M input tokens; output free | \n| Recurring cost | Local hardware; no per-decision token fee | Self-host compute + ops | Per-token input charges | \n| Correct unit | Normalized cost per verified useful decision = (inference HW + service ops + adaptation amortization + review + provider fees) ÷ verified useful decisions. No party publishes a hardware-normalized total — do not infer a universal lowest cost. |  |  | \n\nSources: [S1] Pienomial AT0M benchmark report v1.0, 27 Sep 2026 · [L1/L2] Convai Laya model cards · [I1] DecisionEval Jev 1.13.0 independent rerun, 20 Sep 2026 · [J1/J2] TypeSafe documentation · [B1] public typed-decisions leaderboard. Latencies are not hardware- or network-normalized; pricing and availability may change after the report date. \"Not disclosed\" means the cited sources don't establish the item — not that it's unavailable.\n\nV1 is the single binary you run on your own hardware. Leave an address and we'll tell you when it ships — nothing else, and no one else gets it.", "url": "https://wpnews.pro/news/show-hn-at0m-a-jev-like-system-one-decision-model-in-a-single-rust-binary", "canonical_source": "https://at0m.pienomial.com/", "published_at": "2026-09-30 12:23:19+00:00", "updated_at": "2026-09-30 12:50:04.026320+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "developer-tools"], "entities": ["AT0M", "Rust", "Laya", "Jev 1.13.0", "TypeSafe", "DecisionEval", "DAIR Emotion", "Banking77"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-at0m-a-jev-like-system-one-decision-model-in-a-single-rust-binary", "markdown": "https://wpnews.pro/news/show-hn-at0m-a-jev-like-system-one-decision-model-in-a-single-rust-binary.md", "text": "https://wpnews.pro/news/show-hn-at0m-a-jev-like-system-one-decision-model-in-a-single-rust-binary.txt", "jsonld": "https://wpnews.pro/news/show-hn-at0m-a-jev-like-system-one-decision-model-in-a-single-rust-binary.jsonld"}}