Xuan-Son Nguyen's change brings yes-or-no, classification and scoring outputs to llama-server, with development builds available ahead of an expected v0.6.0 release.
By [RuntimeWire Staff](https://runtimewire.com/author/runtimewire-staff)
· Published
Primary source: [Aligned News - AI Intelligence](https://x.com/ggerganov/status/2106029758350032937)
Why it matters #
A local API for typed decisions lets software use specialized open models for routing, classification and scoring. Its value depends on model coverage and whether the returned probabilities hold up against real workloads.
llama.cpp, the open-source inference runtime associated with Georgi Gerganov (@ggerganov), merged a local typed-decision API on October 2nd. Gerganov announced the endpoint on X; the change, authored by Xuan-Son Nguyen, adds a /v1/systemone endpoint for five model families.
A support system can send a ticket and ask whether it needs urgent attention, which team should receive it, and how frustrated the customer sounds. The endpoint returns typed answers and probability distributions for those decisions. It is designed for software that needs a structured result it can act on, such as a label or score, instead of a paragraph it must parse.
A local route for typed decisions
The pull request names Laya, Julia-1, Lev, OpenJev and Kev as supported model families. Its example sends a shared state containing a customer's message and order details, then asks several questions at once: whether the customer wants a refund, which team should handle the case, and how urgent or frustrating it is. The answer format can be yes-or-no (noul), a choice among labels, or a score on an ordered scale.
TypeSafe AI introduced hosted model Jev on September 15th, describing its /v1/systemone format as a way to return typed decisions for software workflows. The llama.cpp implementation gives developers a local option for a subset of specialized models. Operators can run inference on their own hardware or avoid routing every classification request through a hosted service. It does not make every open model compatible: the pull request describes model-specific conversion, metadata and server handling.
A community discussion had proposed a native endpoint and described a wrapper built from existing llama-server features. Nguyen's merged change adds that capability to the runtime, where conversion and server support can be maintained alongside the project. Gerganov reviewed and approved the pull request, while Nguyen authored it.
The contributor behind the change
Nguyen's path into llama.cpp started as an open-source contributor. In his own account of joining Hugging Face, he says he began contributing while finishing his studies at INSA Centre Val de Loire and joined Hugging Face in August 2024. His biography describes earlier work in full-stack development and cybersecurity, including at Snowpack. Since then, he has worked on the project's server and model integration.
Gerganov created llama.cpp to run models locally through a lightweight C/C++ runtime. In February 2026, he and the GGML team joined Hugging Face, which said it would provide long-term resources while leaving the project open source and its technical direction with the team. The new API extends local inference beyond text generation for developers running the software themselves.
In the pull request's author disclosure, Nguyen wrote that most of the code was AI-written, while he owned the design. The disclosure describes how this feature was built; the change went through the project's review process before merging.
What the probabilities do not prove
The endpoint returns probabilities, but a probability field alone does not establish that a model's confidence is calibrated or reliable for a particular workflow. The pull request includes comparisons between reference outputs and llama.cpp outputs for example inputs. Those checks show implementation agreement on the tested cases; they are not a broad evaluation of decision quality across real customer data.
In applications such as moderation, routing or escalation, teams may use a score to decide whether a human should intervene. Developers will still need to test supported models against their own cases and verify that reported confidence matches observed accuracy. The current support list and model-specific requirements also limit the feature's reach.
Nguyen wrote in a follow-up on X that the feature was expected to ship with llama.cpp v0.6.0 "early next week" and that developers could use development builds in the meantime. The merged code is in the project; the stable release timing remains an expectation, not a completed release.