Attributed compile + one small real CPU run
Primary sources (Liquid AI, 2026-10-07): Open d1: Edge decision models for text, vision, and audio (Liquid AI blog), Multimodal open d1 decision models for the edge (Hugging Face blog), and the model cards for d1-3B and d1-omni-600M
Background: Introducing d1: The most capable decision model, now with vision (Liquid AI, 2026-10-05)
All benchmark and latency figures for d1-3B are Liquid AI's. The numbers I report for d1-omni-600M come from one run on a shared 8-vCPU Linux box with 12 tickets I wrote myself. That is a smoke test, not an evaluation.
This week Liquid AI put two of its d1 "decision models" on Hugging Face as open weights: d1-3B (text + images) and an experimental d1-omni-600M (text + images, or text + audio). It landed the same week OpenAI put its Decisions API into public beta, and the conversation on X quickly turned to running these decisions locally instead of paying per call.
The idea is simple and, for backend people, more interesting than another chat model. A decision model does not write text. You hand it a state (a string, JSON, an image) and a set of named questions, and it returns typed, calibrated answers in one forward pass, with zero output tokens. Three question types cover most pipeline glue:
| type | you give | you get back |
|---|---|---|
noul |
a yes/no question | P(yes) |
choice |
named options with descriptions | choice ,confidence ,probabilities |
score |
2 to 10 ordered levels | expected level, confidence ,probabilities |
That is exactly the shape of the ticket routers, intent classifiers, and moderation checks that a lot of us currently build with an LLM call, a prompt that says "answer only with JSON", and a parser that hopes for the best.
From the blog and model cards, in short:
transformers support via trust_remote_code, and Liquid mentions llama.cpp support for the family.
The 3B model is the one Liquid is pushing. I ran the 600M one, for a boring reason: d1-3B wants about 12 GB in float32 on CPU, and the box I used had about 6 GB free. If you have a GPU or an M-series Mac, start with the 3B.
The 600M model card asks for transformers>=5.15 (the 3B card says >=5.14). so pin it explicitly:
python3 -m venv .venv && . .venv/bin/activate
pip install --index-url https://download.pytorch.org/whl/cpu torch torchvision
pip install "transformers>=5.15" pillow soundfile
Then one call answers three questions over the same ticket:
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained(
"LiquidAI/d1-omni-600M", trust_remote_code=True, dtype=torch.float32
)
questions = {
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "App or site faults",
"fraud": "Suspected unauthorised use",
},
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"],
},
}
out = model.system_one(
"I was charged twice for my October plan. Please refund the duplicate.", questions
)
print(out["answers"])
On my run, that ticket came back as billing with confidence 0.994, P(refund) = 0.999, and an urgency score of 1.69 on the 0 to 2 scale, using 157 input tokens and 0 output tokens. Two notes from the card: use float32 on CPU (the model was trained in it), and avoid bfloat16 on GPU, which Liquid says flipped the top answer on a small share of rows. trust_remote_code also downloads Python files from the repo, so pin a revision before you ship this anywhere.
I wrote 12 short support tickets and their expected team and refund labels before running anything: four billing, four technical, four fraud-ish, including a phishing email question and a hijacked account.
system_one_batch took 3.63 s, about 13 tickets per second.
Those are CPU numbers on a shared Xeon box, so treat them as "fast enough for a background queue", not as a latency benchmark. Twelve easy tickets also say nothing about your data; the point is that the API shape works and the probabilities are usable.
The more useful part was where confidence dropped. The phishing question ("I got an email asking me to confirm my card number on a weird link. Is that you?") still went to fraud, but at 0.477. The hijacked-account ticket went to fraud at 0.633. Those are exactly the tickets I'd want a human to see, and the confidence field flags them without any prompt engineering.
A choice question always picks one of your options. I sent two inputs that are not tickets at all:
billing at 0.496. The low confidence gives it away.technical at So a confidence threshold alone does not catch junk. The obvious fix is a catch-all option, so I added "other": "None of the above, or not a real support request" and a separate noul gate, "Is this a genuine customer support request?". That fixed the junk: the gibberish moved to other (0.663) with P(request) = 0.004, and the address question moved to other (0.896).
It also moved two real tickets. With other available, the phishing question went to other (0.603) and "Do you have a nonprofit discount?" went to other (0.699) instead of billing. The clear tickets (double charge, crashing app) stayed put at 0.90 or higher. Adding a catch-all moves the decision boundary for every borderline input, not just the junk.
noul "is this a real request?" question and drop or queue anything with very low P(yes).
For years the pattern has been to call an LLM, beg for JSON, and parse it. A small model that returns probabilities in one pass, runs on a CPU, and costs nothing per call is a real option for the boring 80% of classification work. It doesn't replace judgment on the hard 20%, but it tells you which tickets those are.
The scripts and raw output from this run (triage.py, other_option.py) need only the pip install above and no API key.
YongBo Yu is an AI engineer in Toronto building LLM workflows and agent systems. More at yongbo-yu.vercel.app and GitHub.