llama.cpp adds a typed-decision API for five specialized open models Llama.cpp merged a local typed-decision API on October 2nd that adds a /v1/systemone endpoint supporting five model families — Laya, Julia-1, Lev, OpenJev and Kev — returning yes-or-no, label and ordered-scale answers with probability distributions, according to the pull request authored by Xuan-Son Nguyen and announced by Georgi Gerganov. The change, approved by Gerganov and available in development builds ahead of an expected v0.6.0 release, lets operators run classification, routing and scoring inference on their own hardware instead of routing every request through a hosted service such as TypeSafe AI's Jev, which introduced the /v1/systemone format on September 15th. The pull request notes the implementation requires model-specific conversion, metadata and server handling, so it does not make every open model compatible. llama.cpp adds a typed-decision API for five specialized open models Xuan-Son Nguyen's change brings yes-or-no, classification and scoring outputs to llama-server, with development builds available ahead of an expected v0.6.0 release. By RuntimeWire Staff https://runtimewire.com/author/runtimewire-staff · Published Primary source: Aligned News - AI Intelligence https://x.com/ggerganov/status/2106029758350032937 Why it matters A local API for typed decisions lets software use specialized open models for routing, classification and scoring. Its value depends on model coverage and whether the returned probabilities hold up against real workloads. llama.cpp https://github.com/ggml-org/llama.cpp?ref=runtimewire , the open-source inference runtime associated with Georgi Gerganov @ggerganov https://x.com/ggerganov?ref=runtimewire , merged a local typed-decision API on October 2nd. Gerganov announced the endpoint on X https://x.com/ggerganov/status/2106029758350032937?ref=runtimewire ; the change https://github.com/ggml-org/llama.cpp/pull/29818?ref=runtimewire , authored by Xuan-Son Nguyen https://github.com/ngxson?ref=runtimewire , adds a /v1/systemone endpoint for five model families. A support system can send a ticket and ask whether it needs urgent attention, which team should receive it, and how frustrated the customer sounds. The endpoint returns typed answers and probability distributions for those decisions. It is designed for software that needs a structured result it can act on, such as a label or score, instead of a paragraph it must parse. A local route for typed decisions The pull request names Laya, Julia-1 https://runtimewire.com/models/huggingface/supersoniclabs-julia-1-ef56314765bda517 , Lev, OpenJev https://runtimewire.com/models/huggingface/openjev-openjev-4c2f2f313f2c2b1e and Kev as supported model families. Its example sends a shared state containing a customer's message and order details, then asks several questions at once: whether the customer wants a refund, which team should handle the case, and how urgent or frustrating it is. The answer format can be yes-or-no noul , a choice among labels, or a score on an ordered scale. TypeSafe AI introduced hosted model Jev on September 15th, describing its /v1/systemone format as a way to return typed decisions for software workflows. The llama.cpp implementation gives developers a local option for a subset of specialized models. Operators can run inference on their own hardware or avoid routing every classification request through a hosted service. It does not make every open model compatible: the pull request describes model-specific conversion, metadata and server handling. A community discussion had proposed a native endpoint and described a wrapper built from existing llama-server features. Nguyen's merged change adds that capability to the runtime, where conversion and server support can be maintained alongside the project. Gerganov reviewed and approved the pull request, while Nguyen authored it. The contributor behind the change Nguyen's path into llama.cpp started as an open-source contributor. In his own account of joining Hugging Face https://huggingface.co/blog/ngxson/ggml-and-llama-cpp-join-hugging-face?ref=runtimewire , he says he began contributing while finishing his studies at INSA Centre Val de Loire and joined Hugging Face in August 2024. His biography https://ngxson.com/about/?ref=runtimewire describes earlier work in full-stack development and cybersecurity, including at Snowpack. Since then, he has worked on the project's server and model integration. Gerganov created llama.cpp to run models locally through a lightweight C/C++ runtime. In February 2026, he and the GGML team joined Hugging Face https://huggingface.co/blog/ggml-joins-hf?ref=runtimewire , which said it would provide long-term resources while leaving the project open source and its technical direction with the team. The new API extends local inference beyond text generation for developers running the software themselves. In the pull request's author disclosure, Nguyen wrote that most of the code was AI-written, while he owned the design. The disclosure describes how this feature was built; the change went through the project's review process before merging. What the probabilities do not prove The endpoint returns probabilities, but a probability field alone does not establish that a model's confidence is calibrated or reliable for a particular workflow. The pull request includes comparisons between reference outputs and llama.cpp outputs for example inputs. Those checks show implementation agreement on the tested cases; they are not a broad evaluation of decision quality across real customer data. In applications such as moderation, routing or escalation, teams may use a score to decide whether a human should intervene. Developers will still need to test supported models against their own cases and verify that reported confidence matches observed accuracy. The current support list and model-specific requirements also limit the feature's reach. Nguyen wrote in a follow-up on X https://x.com/heave448/status/2106032986042810517?ref=runtimewire that the feature was expected to ship with llama.cpp v0.6.0 "early next week" and that developers could use development builds in the meantime. The merged code is in the project; the stable release timing remains an expectation, not a completed release.