{"slug": "multimodal-open-d1-decision-models-for-the-edge", "title": "Multimodal open d1 decision models for the edge", "summary": "Liquid AI released two open decision models, d1-3B and d1-omni-600M, with d1-3B scoring 48.57 on the Decision Index 0.2.1 — the best result for any model under 10B, ahead of Decider 35B-A3B's 47.11. d1-3B answers a single question in 16 ms on an NVIDIA Jetson AGX Thor and 50 ms on a Jetson Orin Nano, while the experimental d1-omni-600M handles text with either images or audio. On seven public benchmarks, d1-3B posted a mean score of 82.9 and d1-omni-600M scored 78.4, surpassing Decider 2B's 77.1 with a quarter of the parameters.", "body_md": "[Fill-Mask •  0.4B • Updated   •  33.1k  •  149](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M)  \n\n# \n\t\tMultimodal open d1 decision models for the edge\n\t\n\n [Team Article](https://huggingface.co/blog)\n\n[d1 decision model family](https://www.liquid.ai/blog/d1-decision-model):\n\n[d1-3B](https://huggingface.co/LiquidAI/d1-3b)and\n\n[d1-omni-600M](https://huggingface.co/LiquidAI/d1-omni-600M)(experimental).\n\n- **Best decision model under 10B on the Decision Index 0.2.1:** d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).\n- **Multimodal:** d1-3B supports text and images, while d1-omni-600M supports text and images or text and audio\n- **Fast:** d1-3B answers a question in 16 ms on an NVIDIA Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50ms on a Jetson Orin Nano\n\n## \n\t\tHow we built decision models for the edge\n\t\n\nThese open d1 decision models are built on our Liquid Foundation Models (LFMs). Unlike our generative models, decision models don’t produce tokens but answer in a single forward pass.\n\nd1-3B and d1-omni-600M are trained from two very different backbones:\n\n- **d1-3B** is trained from[LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B) , our latest VLM, which is decoder-only. It accepts text and images as inputs.\n- **d1-omni-600M** is trained from[LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) , a bidirectional encoder. It adds vision and audio encoders to handle all three modalities. It accepts either text and image, or text and audio as inputs. This model is currently in an early research release and is undergoing further development.\n\n## \n\t\tBenchmark results\n\t\n\nWe benchmarked d1-3B and d1-omni-600M on seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding. d1-3B achieves a mean score of 82.9, the highest in the table and above Decider 4B. d1-omni-600M scores 78.4, surpassing Decider 2B (77.1) with only a quarter of the parameters.\n\n| Benchmark | d1-omni-600M | d1-3B | Decider 2B | Decider 4B | \n|---|---|---|---|---|\n| SQuAD 2.0 | 74.0 | **83.3** | 67.7 | 76.0 | \n| Civil Comments | **95.8** | 93.3 | 93.6 | 92.8 | \n| MASSIVE intent | 86.1 | 86.9 | 81.1 | **88.3** | \n| PubMedQA | 61.3 | **68.3** | 65.7 | 63.3 | \n| BoolQ | 77.7 | 86.3 | 87.3 | **89.0** | \n| XNLI | 74.7 | 85.6 | 85.0 | **88.6** | \n| PAWS-X | **79.5** | **76.4** | 59.5 | 69.8 | \n| **Mean** | 78.4 | **82.9** | 77.1 | 81.1 | \n\nWe validated that d1-3B retains the vision capabilities of its LFM2.5-VL-3B backbone on standard vision benchmarks, and that d1-omni-600M handles all three modalities. We do not report any vision or audio benchmarks, as the Decision Index v0.3 includes only a private vision split and audio decision benchmarks are currently an open problem.\n\n## \n\t\tSpeed\n\t\n\nIn collaboration with NVIDIA, we evaluated d1-3B on the NVIDIA stack across NVIDIA GeForce RTX 4090, NVIDIA Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. Since d1-omni-600M is an early research release, we don’t report any speed numbers for it in this release.\n\n**Edge inference.** d1-3B answers a single question in under 50 ms on every measured device. Three questions take only 1.3x the time of one, with the AGX Thor going from 16 ms to 20 ms.\n\n|  | One question | 3 questions | 3.4K-token state | 384px image | 64 states, packed | \n|---|---|---|---|---|---|\n| Apple M5 Pro | 30 ms | 41 ms | 640 ms | 62 ms | 78 / s | \n| Jetson AGX Thor | 16 ms | 20 ms | 220 ms | 35 ms | 262 / s | \n| Jetson AGX Orin 64 GB | 26 ms | 35 ms | 560 ms | 83 ms | 110 / s | \n| Jetson Orin Nano | 50 ms | 73 ms | 1,640 ms | 202 ms | 38 / s | \n\n**GPU inference.** On GPU, d1-3B answers a question in under 10 ms and processes a 384px image in under 18 ms on both platforms. \n\n|  | One question | 3 questions | 3.4K-token state | 384px image | 64 states, packed | \n|---|---|---|---|---|---|\n| NVIDIA RTX 4090 | 8 ms | 21 ms | 102 ms | 17 ms | 475 / s | \n| AMD MI325X | 9 ms | 14 ms | 44 ms | 18 ms | 1,106 / s | \n\n## \n\t\tHow to use open d1 decision models\n\t\n\nReach for d1 decision models when you need fast, structured decisions, including multimodal inputs. d1-3B delivers the highest decision quality at its size, while d1-omni-600M fits where footprint matters.\n\nInstall the dependencies (requires `transformers>=5.14`):\n\n```\npip install \"transformers>=5.14\" torch torchvision pillow\n```\n\nThese model ship their own code, so load it with `trust_remote_code=True`:\n\n``` python\nimport io\nimport urllib.request\n\nimport torch\nfrom PIL import Image\nfrom transformers import AutoModel\n\ndevice = \"cuda\" if torch.cuda.is_available() else \"mps\" if torch.backends.mps.is_available() else \"cpu\"\nmodel = AutoModel.from_pretrained(\"LiquidAI/d1-3B\", trust_remote_code=True,\n                                  dtype=torch.float32 if device == \"cpu\" else torch.bfloat16).to(device)\n\n# Several named questions over one text state, answered in one pass\nquestions = {\n    \"refund\": {\"type\": \"noul\", \"instructions\": \"Is the customer asking for a refund?\"},\n    \"team\": {\"type\": \"choice\", \"instructions\": \"Which team should handle this?\",\n             \"criteria\": {\"billing\": \"Charges, refunds, invoices\", \"technical\": \"App or site faults\",\n                          \"fraud\": \"Suspected unauthorised use\"}},\n    \"urgency\": {\"type\": \"score\", \"instructions\": \"How urgent is this?\",\n                \"criteria\": [\"Can wait\", \"Today\", \"Blocking the customer now\"]},\n}\nprint(model.system_one(\"I was charged twice this month, please refund one of them.\", questions))\n\n# An image as the whole state\nurl = \"http://images.cocodataset.org/val2017/000000039769.jpg\"  # two cats on a sofa\nphoto = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))\nprint(model.system_one(None, {\"cats\": {\"type\": \"choice\", \"instructions\": \"How many cats are there?\",\n                                       \"criteria\": {\"one\": \"One\", \"two\": \"Two\", \"more\": \"Three or more\"}}},\n                       images=[photo]))\n\n# Many requests, packed together with no padding\ntickets = [\"Where is my parcel? It was due Monday.\", \"The app crashes when I open settings.\"]\nprint(model.system_one_batch([(t, {\"team\": questions[\"team\"]}) for t in tickets]))\n```\n\nFor brevity, we only include the example for d1-3B. See the [d1-omni-600M model card](https://huggingface.co/LiquidAI/d1-omni-600M) for instructions on how to run it.\n\n## \n\t\tGet Started with open d1 decision models\n\t\n\nBoth decision models are open-weight and available on Hugging Face today:\n\n- **Download:**[d1-3B](https://huggingface.co/LiquidAI/d1-3b) and[d1-omni-600M](https://huggingface.co/LiquidAI/d1-omni-600M) on Hugging Face.\n- **Try:** run the demos in our[System One Arcade](https://huggingface.co/spaces/LiquidAI/system-one-arcade) Hugging Face Space.\n\nWe can't wait to see what you build.\n\n## \n\t\tCitation\n\t\n\nIf you use this work, please cite the release blog:\n\n```\n@article{liquidAI2026opend1,\n  author  = {Liquid AI},\n  title   = {Open d1: Edge decision models for text, vision, and audio},\n  journal = {Liquid AI Blog},\n  year    = {2026},\n  note    = {www.liquid.ai/blog/open-d1},\n}\n```\n\n", "url": "https://wpnews.pro/news/multimodal-open-d1-decision-models-for-the-edge", "canonical_source": "https://huggingface.co/blog/LiquidAI/open-d1", "published_at": "2026-10-07 16:54:33+00:00", "updated_at": "2026-10-07 17:18:53.090501+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-products", "ai-infrastructure"], "entities": ["Liquid AI", "d1-3B", "d1-omni-600M", "Liquid Foundation Models", "LFM2.5-VL-3B", "LFM2.5-Encoder-350M", "NVIDIA", "Jetson AGX Thor"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/multimodal-open-d1-decision-models-for-the-edge", "markdown": "https://wpnews.pro/news/multimodal-open-d1-decision-models-for-the-edge.md", "text": "https://wpnews.pro/news/multimodal-open-d1-decision-models-for-the-edge.txt", "jsonld": "https://wpnews.pro/news/multimodal-open-d1-decision-models-for-the-edge.jsonld"}}