{"slug": "run-decision-models-on-vllm-and-red-hat-ai-using-diffusiongemma", "title": "Run Decision Models on vLLM and Red Hat AI Using DiffusionGemma", "summary": "The vLLM community added a structured-read mode to DiffusionGemma 26B-A4B via vLLM PR #57250, letting the open model return typed probabilistic decisions in a single denoising step instead of token-by-token generation. The approach targets the same high-volume decision role as Jev, the first model from San Francisco startup TypeSafe AI, which reports latency of 70–500 ms but remains a hosted API in early access with no published weights, parameter count, or self-hosting option. DiffusionGemma 26B-A4B is Google's block-diffusion language model built on Gemma 4's mixture-of-experts backbone with 26B total and 4B active parameters.", "body_md": "Why Jev went viral, what \"System One\" models are good for, and how the vLLM community turned DiffusionGemma, already a [Red Hat AI](https://developers.redhat.com/products/red-hat-ai) validated model, into an open, self-hostable decision engine.\n\n## A new shape of model\n\nOver the last couple of weeks, a new kind of model took over developer feeds. Jev is the first model from TypeSafe AI, a San Francisco startup founded by former OpenAI researcher Diogo Almeida. Jev doesn't chat or write. It takes unstructured state in and returns typed, probabilistic decisions out, like a function call backed by frontier intelligence.\n\nThe API has 3 question types. A Choice question picks 1 option from a set, a Score question places input on an ordered scale, and a Noul question is a yes-or-no. The name comes from the Bernoulli distribution. Each answer comes back with a probability that your code can branch on.\n\nTypeSafe calls this category \"System One,\" after Daniel Kahneman's fast, intuitive System 1 thinking: The model picks immediately from defined options instead of deliberating token by token.\n\n## Why this resonated\n\nEnterprises have been making these kinds of decisions with AI for a while. Search and e-commerce teams routinely run small generative models with structured output as zero-shot classifiers. As one analysis put it, composing AI into software is not new; Jev's contribution is a model and API built only for that role, with low latency, low cost, typed answers, and probabilities as the normal output.\n\nSo why did it go viral? Three reasons stand out:\n\n- **Guaranteed structure.** The answer is always one of the options you defined, so there's no JSON to repair and no free text to parse.\n- **Probabilities by default.** Confidence scores let you set thresholds, escalate uncertain cases to a person or a larger model, and act automatically on confident ones.\n- **Speed and cost.** One structured pass is much cheaper than a full decode loop. TypeSafe reports latency of 70–500 ms.\n\nThe use cases are the unglamorous, high-volume decisions inside every application: ticket routing, content moderation, risk scoring, agent branching, and guardrails. A person opens a chatbot a few times a day, but software could make thousands of tiny decisions in the background.\n\nThere's a catch for many enterprises. Jev is a hosted API in early access, and TypeSafe hasn't published weights, a parameter count, or a self-hosting option as of this writing. Regulated industries, air-gapped environments, and sovereign or public sector deployments can't send every routing decision to a third-party endpoint. They need the pattern, not the dependency.\n\n## The open path: Decisions on DiffusionGemma in vLLM\n\nThe vLLM community found that 1 open model already contains most of what a decision engine needs.\n\nDiffusionGemma 26B-A4B is Google's block-diffusion language model, built on Gemma 4's mixture of experts (MoE) backbone with 26B total and 4B active parameters. An autoregressive large language model (LLM) writes left to right, 1 token at a time. DiffusionGemma instead works on a fixed-length \"canvas\" of tokens and fills in every position in parallel through iterative denoising.\n\nThat parallelism is what makes parallel decisions possible. [vLLM PR #57250](https://github.com/vllm-project/vllm/pull/57250) adds a *structured-read mode*, which works like this:\n\n1. **Seed the canvas.** The client prefills the canvas with the answer template, such as`urgent: @` /`category: @` /`severity: @` . Only the answer slots are left as noise.\n2. **Run 1 denoising step.** The request is capped at a single step and marked read-only, so vLLM skips the extra work that full generation would do.\n3. **Read the probability distribution at each slot.** Every answer slot returns calibrated logprobs from a single forward pass. The top choice is the decision, and the entropy of the distribution is the confidence.\n4. **Reread only when uncertain.** If a slot's entropy is above a threshold, the client samples a few more reads and measures agreement. Confident answers return after a single read.\n\nAnswers must be single tokens so the canvas layout stays fixed. That's easy to handle on the client: map `moderation_spam` to `B`, for example. The same mechanism covers yes-or-no, multiple-choice, and scored questions, and several questions can be asked in 1 request.\n\nEarly numbers are encouraging. On a single DGX Spark, the PR author measured 8.7 req/s at 0.12 s with 1 request at a time. With 32 concurrent requests they measured 54 req/s at 0.58 s; each request answered 3 questions, for roughly 162 decisions/s. Google has also published a Cloud Run deployment of this approach running on vLLM, reporting about 35–60 ms single-step latency and 100–123 req/s at batch 32; at three questions per request this translates to 300+ decisions/s.\n\n## Try it: The vLLM recipe\n\nDecisions currently need a nightly build. The [vLLM recipe](https://recipes.vllm.ai/Google/diffusiongemma-26B-A4B-it#jev-style-structured-decisions-nightly) walks through 3 steps.\n\n1. Start vLLM with a diffusion canvas sized for decisions. A 64-token canvas is enough room for a template with several questions.\n\n```\npodman run -d --name dgemma --gpus '\"device=0\"' --ipc=host \\\n  -p 8000:8000 -p 8011:8011 \\\n  -v \"$HOME/.cache/huggingface:/root/.cache/huggingface\" \\\n  vllm/vllm-openai:nightly-e9757321527ca1ecd514c07c1418dd2c53da3d19 \\\n  google/diffusiongemma-26B-A4B-it \\\n  --served-model-name dgemma \\\n  --diffusion-config '{\"canvas_length\":64}' \\\n  --max-logprobs 32 \\\n```\n\n2. Start the example decision server. It converts a schema into a seeded canvas, so clients never have to build one by hand.\n\n```\nuntil curl -fsS localhost:8000/health >/dev/null; do sleep 2; done\n\npodman exec -d dgemma python \\\n  /vllm-workspace/examples/features/structured_diffusion/structured_server.py \\\n  --upstream http://127.0.0.1:8000 \\\n  --tokenizer google/diffusiongemma-26B-A4B-it --canvas 64\n```\n\n3. Ask a question. The example server exposes a Jev-compatible `/v1/systemone` endpoint, so existing Jev client code can point at it.\n\n```\nuntil curl -fsS localhost:8011/health >/dev/null; do sleep 2; done\n\ncurl -sS localhost:8011/v1/systemone -H 'content-type: application/json' -d '{\n  \"model\": \"jev-latest\",\n  \"state\": {\"ticket\": \"Everything is down, demo at noon.\"},\n  \"questions\": {\n    \"urgent\": {\"type\": \"noul\", \"instructions\": \"Needs a reply within the hour?\"}\n  }\n}'\n```\n\nKeep a few operational details in mind:\n\n- The `/v1/systemone` endpoint is an*example server* , not a standard vLLM API. The engine-level machinery (seeded canvases, step caps, read-only requests) is in vLLM core. The request format is left to the community, which is still settling on one and will release it as an experimental endpoint.\n- For standard generation, the recipe keeps `--max-num-seqs` low (4 at a 256-token canvas) because the diffusion state buffers scale with batch size, canvas length, and Gemma's 262K vocabulary. Decisions use a much smaller canvas, which is why the PR's 32-way concurrency numbers are possible. Size your deployment against the variant and canvas length you actually plan to use.\n- You can also try the quantized variants of DiffusionGemma for similar capability with a lower memory footprint.\n\n## Running decision models on Red Hat AI and Red Hat AI Inference\n\nDiffusionGemma is already a Red Hat AI validated model for multimodal workloads (handling text, image, and video inputs). Red Hat has tested it for its existing use cases on the Red Hat AI platform and published optimized checkpoints in the RedHatAI Hugging Face collection, including [FP8-dynamic](https://huggingface.co/RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic) and [NVFP4](https://huggingface.co/RedHatAI/diffusiongemma-26B-A4B-it-NVFP4) variants. The NVFP4 variant cuts the memory footprint to roughly a third of BF16.\n\nCustomers can evaluate decision mode today on Red Hat AI Inference preview builds with confidence that the model architecture is validated on Red Hat infrastructure. It will land in our supported version shortly after.\n\nStructured decisions follow a clear rollout path across the platform:\n\n- **Prototyping today:** Decisions are available in the upstream vLLM nightly, which teams can run on Red Hat AI Inference and Red Hat OpenShift AI through a custom serving runtime, which is unsupported, for prototyping against their own data.\n- **Try it now** via the[unsupported Red Hat AI Inference preview image](http://registry.redhat.io/rhaii-preview/vllm-cuda-rhel9:diffusiongemma-jev) .\n- **Next:**  When structured-read support ships in a stable vLLM release, Red Hat AI Inference Server will pick it up in stages, starting with a preview and hardening toward a supported endpoint based on customer feedback.\n  - Example server Developer Preview in 3.6 general availability (GA), assuming vLLM >= 0.31.0 is picked up.\n  - Hardened endpoint: timing and support level to be determined, based on feedback.\n\nFor customers who can't send data to a hosted API, such as regulated industries, air-gapped sites, and sovereign or public sector deployments, this is the key point. The model is already validated, the weights are open, and the decision engine runs on hardware you control.\n\n## The bigger picture\n\nDecision models aren't a replacement for LLMs. They're a new building block that sits beside them. The pattern many teams will adopt is a fast decision model in the request path for routing, gating, and classification, with a generative model behind it for work that needs language and reasoning. Whatever interface the community settles on, the aim is for vLLM to run it in the open. As new models are developed, we will likely see many more decision-style models and one of our goals is to ensure these models run great on vLLM.\n\nStart running decision models on infrastructure you control. Test the [__vLLM recipe__](https://recipes.vllm.ai/Google/diffusiongemma-26B-A4B-it) on Red Hat OpenShift AI today, explore [__PR #57250__](https://github.com/vllm-project/vllm/pull/57250), download the [__FP8-quantized DiffusionGemma checkpoints__](https://huggingface.co/RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic), or read the [__Red Hat AI Inference__](https://www.redhat.com/en/engage/get-started-with-ai-inference-ebook?sc_cid=RHCTN0250000439163&gad_source=1&gad_campaignid=22196380703&gbraid=0AAAAADsbVMQ_Z9TftAX1OtCmicywoFBzF&gclid=Cj0KCQjwt9jVBhDXARIsAFSP-6c6SZPGH-lDgGC9rys9EL-HNvO0K91KNTIUjHQvKhgIvGN7PWd6R6saAjkCEALw_wcB) guide to plan your deployment roadmap.", "url": "https://wpnews.pro/news/run-decision-models-on-vllm-and-red-hat-ai-using-diffusiongemma", "canonical_source": "https://developers.redhat.com/articles/2026/09/28/run-decision-model-vllm-and-red-hat-ai", "published_at": "2026-09-29 03:13:02+00:00", "updated_at": "2026-09-29 03:47:33.387096+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "machine-learning", "ai-products"], "entities": ["vLLM", "DiffusionGemma 26B-A4B", "TypeSafe AI", "Jev", "Diogo Almeida", "Google", "Gemma 4", "Red Hat AI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/run-decision-models-on-vllm-and-red-hat-ai-using-diffusiongemma", "markdown": "https://wpnews.pro/news/run-decision-models-on-vllm-and-red-hat-ai-using-diffusiongemma.md", "text": "https://wpnews.pro/news/run-decision-models-on-vllm-and-red-hat-ai-using-diffusiongemma.txt", "jsonld": "https://wpnews.pro/news/run-decision-models-on-vllm-and-red-hat-ai-using-diffusiongemma.jsonld"}}