{"slug": "d1-the-most-capable-decision-model-now-with-vision", "title": "d1: The most capable decision model, now with vision", "summary": "Liquid AI released d1, its first decision model, adding image support to a system it says became the first to rival Jev on text decisions with last week's experimental release. Liquid AI reported that d1 matches or beats GPT-6.1 Sol on four of six real applications, costs 19x to 200x less than GPT-6.1 Sol and Claude Opus 5.5, and answers a text decision in 200 to 300 ms without generating tokens. The model is available at console.liquid.ai and the d1 Playground, and Liquid AI said it sorts good and defective parts at 85-97% accuracy on the public VisA dataset despite never being trained for that task.", "body_md": "[News](https://www.liquid.ai/news)\n\n[Models](https://www.liquid.ai/news/models)\n\n# Introducing d1: The most capable decision model, now with vision\n\nToday, we introduce **d1, our first decision model**, now supporting both text and images. With [last week's experimental release](https://x.com/liquidai/status/2105003472332693869), d1 became the first model to rival Jev on text decisions, and it is now the first to extend these capabilities to images. You can try it today at [console.liquid.ai](http://console.liquid.ai/) and in the [d1 Playground](http://d1.liquid.ai/).\n\nWe tested d1 against GPT-6.1 Sol and Claude Opus 5.5 on six real applications, from filtering support tickets to inspecting circuit boards. d1 matches or beats GPT-6.1 Sol on four of them. It costs 19x to 200x less than both models and answers significantly faster on every task.\n\n## How the d1 decision model works\n\nDecision models answer questions about a situation with a probability for each possible answer. The d1 decision model takes unstructured data (e.g., text, images, or both) and one or more questions as input, reads them in one forward pass, and returns the probabilities, without generating any tokens. A text decision takes 200 to 300 ms, fast enough for real-time applications.\n\nd1 answers three types of questions:\n\n- **Noul** : a yes/no question, answered with a probability between 0 and 1.\n- **Choice** : pick one label among many, answered with a probability per label.\n- **Score** : a position on a scale, weighted by the probability of each level.\n\nOne request can ask several questions about the same state, saving input tokens.\n\n## d1 in action\n\nDecision models can replace expensive calls to language models when the answer is a structured decision. They also enable a variety of new use cases, some shown in this section. Every demo below runs live in the [d1 Playground](http://d1.liquid.ai/).\n\n**Visual inspection.** This industrial application reviews parts from four production lines that pass under a camera: circuit boards, candles, cashews, and chewing gum (public VisA dataset). d1 sorts good and defective parts with 85-97% accuracy. The most interesting part is that the model was never trained for this. Thanks to its excellent generalizability, it understands the task from a short description.\n\n**Text applications.** Five applications use d1 for every decision:\n\n- **Smart Filter:** d1 becomes a function in a SQL query,`WHERE d1(ticket, 'the customer wants to cancel)` . It answers yes or no for each of 150 support tickets.\n- **Code Search:** d1 browses the Hugging Face transformers repository (6,511 files) one folder at a time, down to the function that answers a question.\n- **Smart Folders:** d1 files each new document into a folder, then a subfolder. It files search questions the same way, so keyword search only looks in one subfolder.\n- **Web Agent:** d1 operates a flight-search website from a one-sentence goal. At each step, it picks the next action among everything the page allows.\n- **Context Compaction:** d1 reads each tool output in a coding agent's session and keeps, trims, or drops it for the next task. It removes 52% of the tokens and keeps every output the task needs.\n\nWe adapted four of these applications from the following open-source projects: [pg-jev](https://github.com/realZachi/pg-jev), [jevgrep](https://github.com/dzhng/jevgrep), [jev-ultrafast](https://github.com/browser-use/jev-ultrafast), and [fast-jev-compaction](https://github.com/tamaratran/fast-jev-compaction).\n\n**Games.** d1 plays eight classic games live and picks every move. Vision helps it in two ways:\n\n- **Better decisions** : Tetris can be fully described in text, but adding the screen raises d1's score from 70 to 81 cleared lines.\n- **Simpler integration** : In Wordle, d1 reads the board directly from a screenshot. Developers don't need to write a text version of the game, and d1 still solved 12 of 12 games in 3.8 guesses on average.\n\nd1 also handles purely visual tasks. In Quick, Draw!, it guesses what a player's doodle shows among 62 words, and recognizes 5.2 of 6 drawings (random guessing gets 0.6).\n\n## Availability and pricing\n\nStart building today with the d1 decision model, available on the  [Liquid AI API](http://console.liquid.ai/) as `d1`. Create an API key at [console.liquid.ai](https://console.liquid.ai/) (Dashboard > API Keys).\n\nTo use vision, you can simply send images as base64 data URLs in `images`:\n\n``` python\nimport base64, os, requests\n\nimage = base64.b64encode(open(\"board.jpg\", \"rb\").read()).decode()\n\nresponse = requests.post(\n    \"https://api.liquid.ai/decisions/v1/systemone\",\n    headers={\"Authorization\": f\"Bearer {os.environ['LIQUID_API_KEY']}\"},\n    json={\n        \"model\": \"d1\",\n        \"images\": [f\"data:image/jpeg;base64,{image}\"],\n        \"state\": \"Camera image of a circuit board on the production line.\",\n        \"questions\": \n            {\"defect\": {\n                \"type\": \"noul\", \n                \"instructions\": \"Does this circuit board have a defect?\"\n             }\n         },\n    },\n)\n\nprint(response.json()[\"answers\"][\"defect\"][\"noul\"])\n```\n\nThe full API reference is in the [decision models documentation](https://docs.liquid.ai/lfm/models/decision-models).\n\nd1 is billed on input tokens only, with no output tokens. Images are counted as input tokens at the same rate as text: 1.5 tokens per 32×32-pixel patch, so a 1024×1024 image costs 1,536 tokens. Each question is billed as its own prompt, including its text and all images.\n\nd1 is also available through [Vercel](https://vercel.com/ai-gateway/models/d1) and [OpenRouter](https://openrouter.ai/liquid/d1), with text only for now. Vision is coming to both soon.\n\nThis is an exciting time for creativity and exploration in AI, and d1 is only the beginning. We're continuing our work on decision models across a range of sizes, with new features, higher decision quality, and lower latency, and we plan to release open weights for upcoming models on Hugging Face soon.\n\n[Try on Liquid Console](https://console.liquid.ai/)\n\n[Read our docs](https://docs.liquid.ai/lfm/models/decision-models)\n\n## Citation\n\nFor citations, please use the following reference or BibTeX:\n\n*Methodology. We ran each application once per model on October 5, 2026, with the d1 Playground's comparison script. GPT-6.1 Sol and Claude Opus 5.5 get each request as one chat message and answer in JSON, at their default reasoning setting. Where many questions share one input, such as the filter's tickets or the folders' passages, the chat models answer them in batches. Costs use list prices, without prompt-cache discounts. d1's costs use $0.04 per million input tokens. Time is per run, with up to 8 requests in flight. A Smart Filter run is one query over 150 tickets. A Smart Folders run files 105 passages, and its cost is per 1,000 passages. The filter's quality is its F1 score against hand labels, which counts both missed and wrong matches. The other text applications count the goals, questions, passages or needed outputs handled correctly. We wrote six of the 15 code questions and two of the four compaction sessions after d1's pipeline was set. In Visual Inspection, every model sees a good part from the same line next to the part to inspect.*", "url": "https://wpnews.pro/news/d1-the-most-capable-decision-model-now-with-vision", "canonical_source": "https://www.liquid.ai/blog/d1-decision-model", "published_at": "2026-10-05 17:05:57+00:00", "updated_at": "2026-10-05 17:20:29.121239+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-products", "ai-tools", "computer-vision"], "entities": ["Liquid AI", "d1", "GPT-6.1 Sol", "Claude Opus 5.5", "Jev", "Hugging Face transformers", "VisA dataset", "d1 Playground"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/d1-the-most-capable-decision-model-now-with-vision", "markdown": "https://wpnews.pro/news/d1-the-most-capable-decision-model-now-with-vision.md", "text": "https://wpnews.pro/news/d1-the-most-capable-decision-model-now-with-vision.txt", "jsonld": "https://wpnews.pro/news/d1-the-most-capable-decision-model-now-with-vision.jsonld"}}