# D1-Omni-600M Locally: Testing Liquid AI's Audio-Vision-Text Decision Model

> Source: <https://www.mindstudio.ai/blog/d1-omni-600m-locally/>
> Published: 2026-10-08 00:00:00+00:00

# D1-Omni-600M Locally: Testing Liquid AI's Audio-Vision-Text Decision Model

Hands-on test of Liquid AI's D1-Omni-600M decision model running locally, covering text, vision, and audio accuracy plus VRAM usage.

## What is D1-Omni-600M?

D1-Omni-600M is a small decision model from Liquid AI that takes text, images, or audio as input and returns probabilities for a set of possible answers instead of generating written text. It is built on Liquid AI’s LFM2.5-Encoder-350M base, adds a SigLip2 vision tower for images and a 17-layer fast conformer for audio, and routes every modality through a shared 16-layer trunk so the model can compare a question against the evidence in a single pass. At roughly 587 million parameters, it’s small enough to run on a single consumer GPU or even a CPU, and it ships under Liquid AI’s own license on Hugging Face.

Unlike a chatbot, D1-Omni doesn’t write sentences. You give it a “state” (text, an image, or an audio clip), a question type, and a list of answer options, and it scores each option with a probability. That makes it closer to a classifier with multimodal perception than to a conversational model, which is the whole point: it’s meant for fast, structured decisions in pipelines, not open-ended generation.

## TL;DR

- **D1-Omni-600M** is a ~587M parameter decision model from Liquid AI that answers questions with probability scores instead of generated text, across text, image, and audio inputs.
- It’s built on **LFM2.5-Encoder-350M** , paired with a**SigLip2 vision tower** and a**17-layer fast conformer** for audio, all feeding into a shared 16-layer encoder trunk.
- In hands-on text testing, the model correctly identified a refund request, routed it to the right team, and scored urgency accurately from a single customer complaint message.
- Vision testing was strong on **object-level facts** (holding two phones, looking distressed, both over 99% confidence) but weaker on**inferred states** like the primary cause of stress, which landed close to a coin flip at 51%.
- Audio testing across 13 languages returned answers with no errors, but the model mostly classified **how speech was delivered** rather than**what it meant** , suggesting it tracks tone and register more than semantic content.
- On GPU, the model stayed **under 2GB of VRAM** fully loaded, making it viable for edge devices and low-resource servers.
- Liquid AI also offers a **3-billion-parameter sibling** with text and vision support, but no audio modality.

## 
Plans first.
*Then code.*

Remy writes the spec, manages the build, and ships the app.

## How does D1-Omni-600M actually work?

The architecture follows a fairly simple recipe. Everything you feed the model, whether it’s plain text, an image, or an audio clip, gets converted into tokens. Text goes through a standard tokenizer. Images pass through the SigLip2 vision encoder, and audio passes through the fast conformer, a 17-layer architecture designed for speech and sound. All of these tokenized representations land in a single sequence and get processed by one shared 16-layer encoder, where every token can attend to every other token at once.

That shared-trunk design is what lets the model reason jointly across a question, a set of answer options, and the underlying evidence (text, image, or audio) in one forward pass. On the output side, a decision head converts the final representations into probabilities. A yes/no question gets a single confidence score, like 0.93. A multiple-choice question gets a probability distribution across the offered options. There’s no decoding step, no sampling, no generated tokens: just a probability readout.

This is the same general approach used by a wave of recent “decision models” that skip free-form generation in favor of calibrated scoring, but D1-Omni distinguishes itself by extending that approach to audio, not just text and images.

## How well does D1-Omni-600M handle text classification?

In a direct test using a customer-service scenario (“I was charged twice and nobody answers the phone”), the model was asked three questions simultaneously: whether the customer wants a refund, which team should handle it, and how urgent the issue is. It returned a refund request as the top-scoring answer, billing as the correct team, and an urgency score of around 1.90 out of 2, correctly flagging the issue as near-critical. All three answers were correct in a single pass, which is the kind of structured, multi-question output this model class is designed to produce. For straightforward text classification and triage tasks, the model performed reliably.

## How accurate is D1-Omni-600M on vision tasks?

For vision, the test used an AI-generated image of a woman at a desk holding two phones, visibly stressed. Three questions were asked: what’s stressing her most, is she holding two phones, and does she look overwhelmed.

The model nailed the two objective, visually grounded questions: it confirmed she was holding two phones with 99.7% confidence and flagged her as distressed with 99.9% confidence. But when asked to infer the *cause* of her stress (work, in this case), the model’s confidence dropped to about 51%, essentially a coin flip. This points to a pattern worth remembering: D1-Omni is strong at describing what’s visibly present in an image, but much shakier when asked to infer intent, cause, or context that isn’t directly depicted.

On the hardware side, the vision test showed the model staying under 2GB of VRAM when fully loaded onto a GPU, a figure consistent across modalities and well within range for most consumer and edge GPUs.

## Does D1-Omni-600M actually understand audio content?

This is the most interesting and most mixed result. The audio test fed the model 13 short speech clips across 13 languages, including English, Spanish, Yoruba, Persian, Russian, and others, each reading aloud news or encyclopedia-style text. The model had to classify each clip as information, action, complaint, or statement.

All 13 languages returned an answer with no errors. Twelve were classified as “information,” with confidence highest on Portuguese and Spanish (over 99%) and lowest on Persian (65%). Czech was the outlier, scored as “action” at just 51%, another coin-flip result.

The catch: these clips were all read-aloud informational text, so classifying everything as “information” isn’t necessarily a sign of deep comprehension, it may just reflect tone and delivery style. A follow-up test using a short spoken English question (“What is happiness?”) showed the model correctly identifying it as a request for information. That’s a reasonable result, but it still looks like the model is picking up on *how* something is said (register, cadence, delivery) rather than parsing the actual semantic content of the speech. In other words, D1-Omni’s audio modality seems to track delivery style more reliably than meaning.

## Is D1-Omni-600M worth running locally?

For a 600-million-parameter model, D1-Omni is a solid citizen on modest hardware: full GPU load stayed under 2GB of VRAM, and Liquid AI’s documentation indicates it can also run on CPU, which matters for edge deployment or low-cost inference servers. For structured text classification and basic visual fact-checking (is this object present, does this look distressed), it performs well and confidently.

Where it gets shakier is in tasks that require inference beyond what’s directly stated or shown, like guessing the cause of someone’s stress from an image, or correctly parsing meaning from speech content rather than its tonal qualities. If your use case is clear-cut classification (urgency scoring, routing, binary visual checks), it’s a reasonable, lightweight option. If you need nuanced semantic understanding of audio content, or reliable causal inference from images, it falls short, and you should validate its confidence scores rather than trusting them outright, especially anything landing near 50%, which tends to signal genuine uncertainty rather than a wrong-but-confident answer.

Liquid AI also offers a 3-billion-parameter variant with text and vision support, though notably without the audio modality found in the 600M model.

## Frequently Asked Questions

### What is a “decision model” and how is it different from a chatbot?

A decision model takes an input (text, image, or audio) plus a question and a list of possible answers, then returns calibrated probabilities for each answer instead of generating free-form text. It’s designed for classification, routing, and scoring tasks rather than conversation.

### How much VRAM does D1-Omni-600M need?

In hands-on testing on a GPU, the model stayed under 2GB of VRAM when fully loaded, making it runnable on most consumer GPUs and even CPU-only setups.

### Can D1-Omni-600M understand spoken language content, not just tone?

Testing suggests it’s more reliable at detecting how speech is delivered (its register or style) than at parsing the actual semantic content of what’s said. It correctly classified a simple spoken English question as an information request, but broader multilingual tests suggested it leans on delivery cues more than deep comprehension.

### What’s the difference between D1-Omni-600M and Liquid AI’s 3B model?

### Built like a system. Not vibe-coded.

Remy manages the project — every layer architected, not stitched together at the last second.

The 3B model supports text and vision but drops the audio modality that D1-Omni-600M includes. The smaller 600M model is the only one of the two with audio support.

### Is D1-Omni-600M good for vision tasks?

It handles directly observable visual facts well (like confirming an object is present or that someone looks distressed), often with confidence above 99%. It performs noticeably worse on questions requiring inference about causes or context not directly visible in the image.
