cd /news/machine-learning/show-hn-at0m-a-jev-like-system-one-d… · home › topics › machine-learning › article
[ARTICLE · art-142502] src=at0m.pienomial.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Show HN: AT0M – a Jev Like System One decision model in a single Rust binary

AT0M, a decision model that returns one probability per option in a single forward pass, scored 0.789 on 2,000 typed decisions, ahead of Laya's 0.766 by 2.3 points and marginally outside the ±1.8-point sampling error, according to the project's Show HN post. The model ships as a single self-contained Rust binary with no Python at inference, and its v0 evaluation endpoint measured 0.056 expected calibration error versus Laya's 0.081 and Jev's 0.121–0.161 range. The author states calibration decays off-distribution, from 0.037 on seen spam to 0.171 on unseen phishing, and that AT0M cannot compare one field of a record against another, so those predicates must be computed in code.

read8 min views1 publishedSep 30, 2026
Show HN: AT0M – a Jev Like System One decision model in a single Rust binary
Image: source

Give it a state and a set of questions — choice, score, or yes-no. It returns one probability per option, in a single forward pass.

No tokens and no free text, so it cannot answer outside your candidate set. It runs as a single local executable, which keeps inference inside your perimeter.

Accuracy against what a decision costs. A workflow is five typed decisions about one record. Competitor accuracies and prices are as published; ours is measured over all 2,000 test decisions and priced from a laptop’s measured throughput plus its amortised system cost. The full basis is in the report.

AT0M will ship as a single self-contained Rust binary — no Python at inference, no service in the path, nothing to rent. That is the deployment, and it will be the only thing that ends up in your perimeter.

The hosted endpoint exists for one reason: so you can check every number on this page before you download anything.

v0 is the evaluation endpoint; v1 is the executable release — the hosted service is a way to try the model, not a way to run it. Every figure on this page was measured against v0, the evaluation build, so they are the floor rather than the ceiling: v1 is expected to be more accurate, and we will republish every number on this page when it ships.

One configuration across every public test set — no per-dataset tuning or prompt tweaks. 38,862 decisions, scored in one sweep.

Benchmark AT0M Laya Jev Condition
typed-decisions (2,000) 0.789 0.766 0.740 ahead of Laya by 2.3pt, marginally outside the ±1.8pt sampling error
calibration (ECE) 0.056 0.081* 0.121–0.161 *Laya after fitting a temperature, and from its specialist checkpoint; Jev is the spread across four independent studies — TypeSafe publishes none. Jev and AT0M are general models, one configuration for every row; AT0M fits no temperature
DAIR Emotion 0.921 0.595 0.480 bare labels
Banking77 (77 intents) 0.853 0.425 0.870 Jev on 72 labels — not the same label space. Ours moves between 85% and 88% with how the intents are worded, so read the margin here as a band, not a point
AG News 0.896 0.950 0.910 the one clean multi-way we trail
CLINC150 + out-of-scope 0.795 — — 151 candidates, scored in full
Enron spam / held-out phishing 0.955 / 0.805 — — phishing never trained on — the honest transfer test
toxicity guardrail (F1) 0.626 — — 7.1% positive — answering "no" to everything scores 92.9% accuracy and catches nothing

Laya figures, and Jev's on the classification sets, are as reported in the Laya write-up. Jev's typed-decisions figure is DecisionEval's independent rerun. On calibration, TypeSafe publishes no figure at all, so Jev's range is assembled from four independent studies, none temperature-fitted.

The accuracy number is near the top but sits inside admitted label noise, so we don't hang the pitch on it. The claims we'd stand behind under scrutiny are architectural:

And the parts we wouldn't: it compares a value to a threshold well, but not one field of a record against another, so compute those predicates in code. Calibration decays off-distribution — 0.037 on spam it has seen, 0.171 on phishing it hasn't. Both are measured below.

Every figure on this page came from a public test split scored through this endpoint. Point the same benchmark at it and check. One POST carries a record and as many questions as you like, all answered in the same pass — no SDK, no client library.

We'll help you scope a local-training or private-deployment evaluation against your own decisions — and share the full technical report.

A comparative evidence inventory across AT0M, Jev 1.13.0 and Laya, carried from the technical paper. It is not a harmonized three-model benchmark: figures were obtained under different conditions, and each is tagged with its source and how it was measured.

Parameter AT0M Laya Jev 1.13.0
Developer Pienomial Convai Innovations TypeSafe AI
Delivery form Standalone Rust binary; HTTP / CLI / JSON-file Open weights; Python SDK / Jev-compatible HTTP server Managed API
Backbone Not disclosed in report ModernBERT-large 421M; mmBERT-base 322M Not disclosed
Primitives choice · score · noul choice · score · noul choice · score · noul
Inference Non-autoregressive; no free text Non-autoregressive; no free text Non-autoregressive; no free text
Multi-question call Yes Yes Yes
Max context 32k request — the verification server caps requests at 4,096 tokens 512 EN / 1,024 TD, multi 1,024 32k state + longest Q / 64k request
Model selection One configuration across all tests Router: EN / multilingual / specialist One managed version
| Measure | AT0M [S1] | Laya [L2] | Jev [I1] | 
|---|---|---|---|
| TD overall | 0.789 | 0.766 specialist | 0.740 (0.727 older [B1]) | 

| TD choice | 0.755 | 0.733 | 0.738 | | TD score | 0.763 | 0.723 | 0.705 | | TD noul | 0.857 | 0.857 | 0.787 | | Invoice / service / obs / security | .846 / .786 / .764 / .758 | .804 / .764 / .730 / .766 | .786 / .788 / .640 / .744 | | AG News | 0.896 described | 0.950 routed | 0.910 (secondary) | | DAIR Emotion | 0.921 bare | 0.595 routed | 0.480 (secondary) | | Banking77 | 0.853 / 77 labels | 0.425 / 77 labels | 0.870 / 72 labels ≠ | | Held-out phishing | 0.802 (excluded from training) | — | — | | CLINC150 + out-of-scope | 0.794 / 151 choices | — | — | | Base vs specialist | One config throughout | General EN checkpoint 0.362 on TD | Zero-shot general weights |

| Measure | AT0M | Laya | Jev |

|---|---|---|---|
| TD ECE (10-bin) | 0.057 no temp fitted [S1] | 0.213 raw specialist [L2] | 0.121–0.161 across four independent studies | 

| Per-type ECE | choice .052 · score .086 · noul .057 | 0.081 after temperature [L1] | TypeSafe publishes none | | Brier | Not provided | 0.062 specialist | 0.148 independent | | Confidence gating | Full coverage curve not published | No matched curve | ≥0.7 → 57.6% kept @ 86.4%; ≥0.9 → 24.5% @ 93.7% | | Option-order | 98% stable | Not established | Not established | | Counterfactual | 13/15 directional; identifier-repeat pair fails | — | — | | Empty record | Confidence falls 0.723→0.616; still no abstention | Not measured | Not measured |

Metric AT0M Laya Jev
5-question, T4 82 ms server time 84.5 ms EN / 40.1 ms multi 687 ms p50, network incl. [I1]
1-question, T4 16 ms T4 server; 22.1 ms M1 Max 39.5 ms EN / 32.8 ms multi older 236–276 ms hosted
Throughput M1 Max ~41→54 dps (1→32 clients) 103–332 q/s batched, T4 Provider-managed
77 / 151 candidates 92.7 ms / 165.5 ms, M1 Max Budget competition on options Not matched
Hardware CPU · Apple Metal · CUDA CPU / GPU self-host Provider-managed
Network Server figure excludes round-trip Inference; network not included 687 ms explicitly end-to-end
Parameter AT0M Laya Jev
Weights Training emits own weights; distribution terms per engagement Apache-2.0 downloadable Proprietary; no download
Self-hosting Standalone executable Python / HTTP server Not offered publicly
Data transmission Local inference stays in perimeter Local inference stays in perimeter Sent to provider unless contracted otherwise
Version management Export validation; customer handles rollout Pin model hashes / runtime Versioned IDs / alias pinning
Audit trail Raw probabilities; app owns the log Opt-in prediction hooks Version returned; app logs decisions
Item AT0M Laya Jev
Published price Commercials will be released with v1. What do you think they should be? Please let us know at info@pienomial.com Apache-2.0; hosting not free $0.042 / 1M input tokens; output free
Recurring cost Local hardware; no per-decision token fee Self-host compute + ops Per-token input charges
Correct unit Normalized cost per verified useful decision = (inference HW + service ops + adaptation amortization + review + provider fees) ÷ verified useful decisions. No party publishes a hardware-normalized total — do not infer a universal lowest cost.

Sources: [S1] Pienomial AT0M benchmark report v1.0, 27 Sep 2026 · [L1/L2] Convai Laya model cards · [I1] DecisionEval Jev 1.13.0 independent rerun, 20 Sep 2026 · [J1/J2] TypeSafe documentation · [B1] public typed-decisions leaderboard. Latencies are not hardware- or network-normalized; pricing and availability may change after the report date. "Not disclosed" means the cited sources don't establish the item — not that it's unavailable. V1 is the single binary you run on your own hardware. Leave an address and we'll tell you when it ships — nothing else, and no one else gets it.

── more in #machine-learning 4 stories · sorted by recency
── more on @at0m 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-at0m-a-jev-l…] indexed:0 read:8min 2026-09-30 · —