{"slug": "frink-inference-the-rust-alternative-to-llama-cpp", "title": "Frink Inference: the Rust alternative to llama.cpp", "summary": "Frink, a pure-Rust inference engine for GGUF models created by developer Antonello F, now runs 100 architectures on its generic path with evidence, each verified against llama.cpp's own logits via libllama at a typical KL divergence around 1e-13, according to the project's v0.24 to v0.49 release notes. The engine loads the same checkpoints as llama.cpp and runs them on CPU, Apple Metal or CUDA, ships an OpenAI-compatible server with Anthropic Messages and the Responses API on the same port, and includes a CLI, a Studio web UI, and tools such as download, bench, quantize, imatrix, gguf-split, verify and parity in the same binary. Frink's stated goal is to be used instead of llama.cpp: it refuses to load a model rather than compute something different, matches llama.cpp tokens at temperature 0, and aims to be no slower on the same host, file and backend.", "body_md": "# Frink Inference: the Rust alternative to llama.cpp\n\n[Frink](https://github.com/antonellof/frink) is a pure-Rust inference engine for GGUF models. It loads the same checkpoints as llama.cpp and runs them on CPU, Apple Metal or CUDA, dense and mixture-of-experts alike. There are no llama.cpp bindings and no ggml wrapper: the loader, the quantized kernels, attention and expert routing are all written in Rust.\n\nWhat you get:\n\n- **A CLI with llama.cpp’s flags** :`-m` ,`-p` ,`-n` ,`-ngl` ,`-c` ,`-hf` ,`--jinja` . A command copied from a model card just works.\n- **An OpenAI-compatible server** , with Anthropic Messages and the Responses API on the same port, continuous batching and a prefix cache over paged KV.\n- **Studio** , a small web UI that talks to the server: chat, the model inventory and live request activity.\n- **The tools, in the same binary** :`download` ,`bench` ,`quantize` ,`imatrix` ,`gguf-split` ,`verify` and`parity` , which checks Frink against llama.cpp on your own machine.\n\nThis post covers the goal Frink is built around, what changed in the last twenty-five releases (v0.24 to v0.49), and where it already beats llama.cpp. The [first post](https://www.fratepietro.com/2026/frink-rust-gguf-inference-engine/) covers the design, the [second](https://www.fratepietro.com/2026/frink-metal-parity-llama-cpp/) the Metal backend, and the [Bonsai post](https://www.fratepietro.com/2026/frink-bonsai-ternary-27b-local/) running a 27B ternary model on a 16 GB Mac.\n\n## The goal, stated once\n\nThe project now has a written north star, and every other plan is ranked against it:\n\nFrink should be what somebody reaches for **instead of** llama.cpp. Same models, same command shapes, same or better performance, on the hardware people actually own.\n\nThat is a bigger claim than “a fast Rust inference engine”, so it comes with a bar you can check. For any GGUF you can run under llama.cpp:\n\n1. It **loads** , or refuses with a sentence naming exactly what is missing. It never loads and then computes something else.\n2. It produces the **same tokens** at temperature 0, checked against**llama.cpp’s own logits** , not against a golden file this project wrote.\n3. It is **not slower** on the same host, file and backend.\n4. The **command you already know** works, or the difference is documented.\n\nThe first point is where Frink differs from llama.cpp on purpose, and I think it is the most important one. llama.cpp will often run something approximately right. Frink refuses instead. A refusal is a gap you can see, while a model that loads and quietly computes the wrong thing is a bug you only find after trusting it.\n\nThe route there is **recent, big models on consumer hardware**: MoE checkpoints, including ones too large for the machine’s memory, because that is where llama.cpp is weakest and where a new engine can win on merit rather than on being written in Rust.\n\n## What changed: v0.24 to v0.49\n\nIt has been a dense month: twenty-five releases, grouped here by theme.\n\n### Architecture coverage: 100 models with evidence\n\nThe main number is **100 architectures running on the generic path, each with evidence**, plus four dedicated engines (MLA/DeepSeek, GLM, Kimi, Gemma 4). “Evidence” has a specific meaning here: a tiny synthetic GGUF of that architecture is run through **libllama** to produce reference logits, and Frink has to match them, typically to a KL divergence around 1e-13. An architecture without that evidence does not run on a guess. It refuses, and the refusal says which line of which llama.cpp graph it still needs.\n\nSome of what closed this month:\n\n- **The Mamba and state-space families** : Mamba-1 and Mamba-2, Jamba, Granite 4.0 hybrids (H-Micro / H-Tiny / H-Small), Nemotron-H including the 30B-A3B MoE, Falcon-H1, and PLaMo-2 with its own tokenizer.\n- **Gated delta-net hybrids** : Qwen3.5 dense and MoE, and Qwen3-Next-80B-A3B.\n- **LFM2 and LFM2-MoE** , with their short convolutions running on the ordinary layer rather than on a separate engine.\n- **Llama 4 Scout and Maverick** , with chunked attention, the attention temperature on the no-RoPE layers, and the expert weight applied to the*input* of the expert. Llama 4 is the only graph that does that last one, and on the test fixture it moves the logits by 0.86.\n- **MiniMax-Text-01** (456B-A45B) with lightning attention,**Spark-2.5** ,**Maple-20B** ,**Granite 4.1** ,**DFM Mimir** and**muse-glimmer** : six of the eight architectures that arrived when I moved the llama.cpp reference pin forward 792 commits.\n- **DeepSeek-V2/V3 attention** , finally with a libllama golden, in both the legacy and the current tensor layout, and with YaRN read the way llama.cpp reads it across three separate files.\n- **The 2023 tail** : GPT-2, StarCoder, BLOOM, MPT, Falcon, GPT-NeoX/Pythia, Phi-2, Command-R, Cohere2, Arctic, Grok, DBRX and many more.\n- **Embeddings, reranking and scoring** : nomic-bert,`llama-embed` , and a new`/v1/score` endpoint.\n\nThe refusal count briefly went *up* when I moved the pin, from 2 to 10. That is what keeping pace with a moving target looks like. A count that only ever fell would mean nobody was reading upstream.\n\n### A serving stack with the full sampling surface\n\nMost of the version numbers went to the server. Frink now accepts the whole OpenAI and vLLM sampling surface instead of a subset:\n\n- `n` and`best_of` from**one shared prefill** , using copy-on-write on the paged KV store, and interleaved when streaming.\n- `logprobs` ,`prompt_logprobs` ,`logit_bias` ,`allowed_token_ids` ,`bad_words` ,`echo` ,`truncate_prompt_tokens` ,`skip_special_tokens: false` ,`return_tokens_as_token_ids` .\n- `cache_salt` , so each caller gets isolated prefix-cache pages. Identical prompt text under a different salt reuses nothing.\n- **Sleep mode** :`POST /sleep` ,`POST /wake_up` ,`GET /is_sleeping` .\n- **Speculative decoding in the server** , lossless at any temperature.\n- **Cache-aware admission** : a queued request whose system prompt is already cached can be admitted first, with hard bounds so a large request cannot be starved.\n\nOne fix mattered more than any single feature. v0.30.0 found **eleven request fields that changed the answer but were being accepted and ignored**. Now any field the server does not implement is **refused by name**, because a dropped field looks exactly like one that was honoured.\n\nMeasured on an M2 Pro with Llama-3.2-1B, one build against itself: continuous batching takes aggregate throughput from **37.8 tok/s at one client to 67.0 at sixteen**, and a shared 757-token system prompt reuses **736 tokens** from the prefix cache, cutting time to first answer from 938 ms to 410 ms.\n\n### The engine\n\n- **4-bit KV cache with a Hadamard rotation** (`--ctk q4_0` ). K is rotated when written and Q goes through the same rotation when read, so`q·k` is unchanged. Against an f16 cache, next-token agreement over sixty long-context windows went from 45/60 to 52/60. The more useful finding came by accident: two builds that differed only in an arbitrary sign pattern scored 48 and 52, so a four-window difference is the*noise of the metric* . A KV comparison is only worth believing when the difference is wider than that.\n- **Bonsai 27B got faster, and its decode speed no longer drops with context length** : 11.17 / 11.19 / 11.16 tok/s at 32 / 300 / 600 tokens of context on an M2 Pro, against the reference’s 11.46. The per-token command-buffer waits went from 192 to 69 to 20. A four-layer group on a hybrid model is now one GPU submission.\n- **`--ctk` takes llama.cpp’s value names** and refuses the ones it does not implement.\n- **The CLI and server logs follow llama.cpp’s format** : the same launcher, the same colours on the prompt echo and the timing line, and startup lines in the same timestamped, levelled format, so the output you are used to reading reads the same way.\n\n## Where Frink is already better than llama.cpp\n\nI want to be precise here, because “better” is easy to say and hard to check.\n\n**It refuses instead of guessing.** This is the design difference everything else rests on. When an architecture, a GGUF key or a request field is not implemented, Frink stops and names it. Along the way, the libllama comparisons turned up several places where llama.cpp itself runs something wrong or does not run at all:\n\n- libllama runs every current **PLaMo-2** export**without RoPE** , because it takes the rotary dimension from layer 0, which is a state-space layer. On the one public PLaMo-2 GGUF, libllama returns all-NaN logits. Frink answers “Paris. Paris is the capital of France.”\n- libllama **segfaults** on a Granite-4.0 hybrid file that omits an optional convolution bias, because it adds that bias unconditionally. Frink treats it as required.\n- No interleaved **ERNIE-4.5 MoE** checkpoint can load in llama.cpp: the tensor loader and the graph disagree about the interleave step. Frink refuses the case by name.\n- A **DeciLM** layer with attention but no FFN has its attention output silently discarded by llama.cpp. Frink refuses rather than copy that behaviour.\n\n**It runs models mainline llama.cpp does not.** PrismML’s **Ternary-Bonsai-2-27B** (`PTQ1_0`, 1.75 bits per weight, 5.95 GB) needs PrismML’s fork of llama.cpp. Frink runs it as released, matched against that fork to a first-token KL of 2e-5.\n\n**It decodes faster on Apple Silicon.** On the M2 Pro, Frink decodes faster than llama.cpp on **every one of the 15 comparable Metal rows**, from 4% ahead on OLMoE to 1.67x on the smallest models: Qwen3-0.6B at 191 tok/s against 116, SmolLM2-135M at 363 against 217, Gemma-3-1B at 122 against 83. Llama-3.1-8B is 32.0 against 30.4. Prefill is within 9% of llama.cpp on every row, level on most. (Those rows were measured on v0.20; the ledger marks them as such.)\n\n**Its server does more.** llama-server is a good server. Frink’s has the vLLM-style surface on top of it: `n`/` best_of` from one prefill, prompt log-probabilities, per-caller cache isolation, sleep/wake, runtime model swap, resumable streams, slot save/restore that refuses a mismatched checkpoint, and Anthropic Messages plus the Responses API on the same port. Structured output is enforced per token, whether from a GBNF grammar, a forced `tool_choice`, or a tool’s own JSON schema, so an invalid answer cannot be generated at all. Tool calls are parsed in the eleven formats real checkpoints emit.\n\n**It is one Rust binary.** 23 MB on macOS arm64, with the server, `download`, `bench`, `quantize`, `imatrix`, `gguf-split` and `verify` all included. No Python, no CUDA userspace to match against your driver. The engine is published as ordinary crates too: `frink-inference` as a facade, or `frink-gguf`, `frink-quant`, `frink-core` and `frink-models` individually if you want to embed it.\n\n**It checks itself against llama.cpp.** `frink parity` runs both engines on the same token ids and compares the full logit distribution. `frink bench --compare` runs `llama-bench` beside it on the same file. Every benchmark row now records which build measured it, and the tables mark outdated rows instead of mixing them in. The tools that would catch Frink being wrong are part of Frink.\n\n**And where it matches, it matches exactly.** `frink quantize` writes Q8_0 and the K-quants **byte-identically** to `llama-quantize`, with or without an importance matrix. Tokenization is checked against libllama on twenty checkpoints. The flags are llama.cpp’s flags: `-m`, `-p`, `-n`, `-ngl`, `-c`, `-hf`, `--jinja`, `--alias`, `--api-key`.\n\n## Where llama.cpp is still ahead\n\nA post comparing itself to llama.cpp without this section would not be worth reading.\n\n- **CUDA is far behind.** On an RTX 3090, Frink’s prefill is 25x to 43x slower and decode 2.7x to 9x slower, from rows measured on v0.21. Newer work (a resident prefill and a tensor-core GEMM) has not been re-measured on that box yet. Today CUDA is correct but not fast.\n- **CPU is behind.** On a Ryzen 9 3900X, decode is 1.04x to 1.34x slower and prefill up to 4.45x. The gap is largest on the smallest models, which points at fixed per-matmul overhead rather than slow kernels.\n- **No Vulkan.** That means no AMD or Intel GPUs, which llama.cpp covers with one backend. This is the largest hardware gap.\n- **Coverage is not complete.** llama.cpp has 155 hand-written graphs. Frink runs 100 architectures with evidence plus four on dedicated engines, and the rest refuse with a reason.\n- **K-quant logits drift slightly** because llama.cpp quantizes activations to`Q8_K` before the dot product and Frink keeps them in f32. That is a documented difference, not a bug, and on`Q8_0` and`IQ4_NL` the logits match.\n\nThose are the next priorities, in roughly that order, along with the item that would do most to justify choosing Frink: **running an MoE larger than the machine’s memory** by keeping the experts that actually fire resident and streaming the rest. The residency policy is written. Wiring it into the engine is next.\n\n## Try it\n\n```\ncurl -fsSL https://raw.githubusercontent.com/antonellof/frink/main/scripts/install.sh | bash\n\n# Fetch and serve in one step, llama-server style\nfrink serve -hf bartowski/Llama-3.2-3B-Instruct-GGUF:Q4_K_M -c 8192 --alias local\n\n# Or check it against llama.cpp on your own machine\nfrink bench -m model.gguf -p 512 -n 128 -r 3 --compare\n```\n\nThat installs `frink` and `frink-server` into `~/.local/bin`. The prebuilt binaries are macOS arm64 with Metal and Linux x86_64 for CPU. `cargo install frink-cli --features metal` (or `--features cuda`) builds the same thing from source.\n\nIf you run a model that Frink refuses, the error message is the bug report. Open an issue with it and the GGUF’s architecture string: that is usually everything needed to add the model.\n\n*None of this would exist without [llama.cpp](https://github.com/ggml-org/llama.cpp) and GGML. Frink does not link against them, but their kernels, formats and years of openly shared engineering were the reference for every line of it.*", "url": "https://wpnews.pro/news/frink-inference-the-rust-alternative-to-llama-cpp", "canonical_source": "https://www.fratepietro.com/2026/frink-rust-alternative-to-llama-cpp/", "published_at": "2026-09-27 22:00:00+00:00", "updated_at": "2026-09-28 17:50:56.382134+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "developer-tools", "ai-tools"], "entities": ["Frink", "llama.cpp", "Antonello F", "GGUF", "libllama", "Apple Metal", "CUDA", "Qwen3.5"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/frink-inference-the-rust-alternative-to-llama-cpp", "markdown": "https://wpnews.pro/news/frink-inference-the-rust-alternative-to-llama-cpp.md", "text": "https://wpnews.pro/news/frink-inference-the-rust-alternative-to-llama-cpp.txt", "jsonld": "https://wpnews.pro/news/frink-inference-the-rust-alternative-to-llama-cpp.jsonld"}}