{"slug": "bonsai-a-27b-reasoning-model-on-a-16-gb-m2-mac-with-frink", "title": "Bonsai: a 27B reasoning model on a 16 GB M2 Mac, with Frink", "summary": "PrismML's Ternary-Bonsai-2-27B, a 27-billion-parameter reasoning model quantized to 1.75 bits per weight and 5.95 GB on disk, runs on a 16 GB M2 Mac through Frink v0.24.0, a pure-Rust GGUF inference engine that added the PTQ1_0 ternary packing and Walsh-Hadamard rotation fold. PrismML reports the model retains 98.2% of FP16 intelligence with an 84.78 average across 14 thinking-mode benchmarks, versus 72.59 for a conventional IQ2_XXS build at more than half again the footprint, and Frink's parity check against PrismML's llama.cpp fork measured a KL divergence of 2e-5 on CPU and Metal decode and 2e-6 through the Metal prefill GEMM.", "body_md": "# Bonsai: a 27B reasoning model on a 16 GB M2 Mac, with Frink\n\n[Frink](https://github.com/antonellof/frink) is a pure-Rust inference engine for GGUF models, with a llama.cpp-shaped CLI, an OpenAI-compatible server, and a small web UI called Studio. The [first post](https://www.fratepietro.com/2026/frink-rust-gguf-inference-engine/) covers the design, the [second](https://www.fratepietro.com/2026/frink-metal-parity-llama-cpp/) the Metal backend.\n\nThis one is about a single model: [Ternary-Bonsai-2-27B](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf) from PrismML. 27 billion parameters at 1.75 bits each, 5.95 GB on disk, running on a 16 GB laptop with room to spare.\n\n## What Bonsai is\n\nEvery language weight is one of three values, -1, 0 or +1, with one 16-bit scale per group of 128. Five trits pack into a byte in base 3, which is where 1.75 bits per weight comes from; that packing is the GGUF type `PTQ1_0`.\n\nTernary alone would wreck a 27B model. What makes it work is a rotation: each weight matrix is transformed blockwise by a Walsh-Hadamard matrix with fixed signs, which spreads outliers across a block so a three-level grid can hold them. The rotation is folded into the stored weights, so the runtime has to apply the matching transform to the activations before every matmul, and undo it on the embedding table after each lookup. Get that wrong and the model loads and talks nonsense.\n\nIt is ternary end to end — embeddings, attention projections, MLP projections and the LM head, a true 1.72 bits per weight, with no high-precision escape hatch behind a low-bit label. The vision tower ships separately as a Q8_0 `mmproj` pack. PrismML publish two packings: `PTQ1_0` packs trits densely (1.75 bits/weight, 5.95 GB) and `PQ2_0` gives each trit its own 2-bit slot (2.13 bits/weight, 7.21 GB). Either way the packed weights are consumed directly and never expanded back to FP16.\n\nThe numbers PrismML report for it are the reason it is interesting rather than a curiosity: **98.2% of FP16 intelligence retained**, 84.78 average across 14 thinking-mode benchmarks, against 72.59 for a conventional IQ2_XXS build at more than half again the footprint — and within 0.4 points of UD-Q4_K_XL at three times the size. Reasoning survives well below 4 bits where conventional low-bit formats collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92. The backbone is the Qwen3.8-27B hybrid attention (roughly 75% linear), which is what keeps its 262K context practical on-device. There is an MLX build too, `Ternary-Bonsai-2-27B-mlx-2bit`, for native Apple Silicon.\n\nFrom ~54 GB in FP16 to ~5.9 GB. That is the whole pitch: 27B-class reasoning on a laptop or one GPU.\n\nFrink v0.24.0 adds the packing (a CPU dot, a Metal matvec, a Metal GEMM) and the fold, read from the checkpoint’s own `prism.hadamard.*` metadata. Anything outside the configuration that has been verified is refused by name rather than guessed at.\n\nPrismML’s [llama.cpp fork](https://github.com/PrismML-Eng/llama.cpp) is the reference implementation for the format, so that is what Frink is checked against. `frink parity` feeds both engines identical token ids and compares the full first-token logit distribution: a KL divergence of 2e-5 on CPU, 2e-5 on the Metal decode kernel and 2e-6 through the Metal prefill GEMM, with the same ten top tokens in the same order. The tokenizer matches on 1198 tokens across 21 test strings.\n\n## Download and run\n\nOne line, no toolchain:\n\n```\ncurl -fsSL https://raw.githubusercontent.com/antonellof/frink/main/scripts/install.sh | bash\n```\n\nThat drops `frink` and `frink-server` into `~/.local/bin`. The macOS\nbuild is arm64 with Metal already on; Linux x86_64 is CPU. Then:\n\n```\n# Same argument shape as `hf download`, no Python.\nfrink download prism-ml/Ternary-Bonsai-2-27B-gguf \\\n  Ternary-Bonsai-2-27B-PTQ1_0.gguf --local-dir models\n\n# Chat. Frink applies the checkpoint's own template.\nfrink -m models/Ternary-Bonsai-2-27B-PTQ1_0.gguf \\\n  -p \"Explain ternary quantization in three sentences\" -n 400 -ngl 99\n```\n\nIf you would rather build it, `cargo install frink-cli --features metal`\n(or `--features cuda`) gives you the same binary.\n\nOn an M2 Pro it decodes at about 10.5 tokens per second and prefills at about 43: reading speed rather than skimming speed, with most of the machine’s memory still free. That decode figure was 7.1 when the format first landed; most of the difference is a recurrent layer now running as a single Metal submission instead of three.\n\n## The server and Studio\n\n```\nfrink serve -m models/Ternary-Bonsai-2-27B-PTQ1_0.gguf \\\n  -ngl 99 --alias bonsai-2-27b --port 8383\n```\n\nThat is an OpenAI-compatible API on `http://127.0.0.1:8383/v1`: chat completions, completions, embeddings and models, with Anthropic Messages and Responses on the same port. `--alias` names the model in `/v1/models` and in every response; without it you get the checkpoint’s own `general.name`, which for this file is the unhelpful string `Hf`.\n\nNote what is missing from that command: `-c`. With no context flag the server prices the checkpoint against the device and derives both the per-request ceiling and the KV block budget from the same arithmetic, so the context it advertises is one it can actually serve.\n\nStudio is a separate app that talks to that API over HTTP. From a checkout:\n\n```\ncd ui && npm install && npm run dev   # http://localhost:5173/ui/\n```\n\nIts Models page lists every GGUF in your models directory with its quant, architecture, context and size, and loads one without restarting the server:\n\nAnd a chat, at the real speed:\n\nThe finished turn carries its own numbers underneath: time to first token, prefill and decode rates.\n\nThe model thinks before it answers. Studio folds the thinking into a collapsible block; over the API it arrives in `reasoning_content`, separated from the answer, and `reasoning_effort: \"medium\"` shortens it while `\"none\"` turns it off.\n\n## Point a coding agent at it\n\nAny tool that speaks the OpenAI API can use the server. Here is [Pi](https://github.com/earendil-works/pi), the minimal coding agent I covered in [an earlier post](https://www.fratepietro.com/2026/running-glm-5-2-locally-rondine-pi/): four tools, one loop, one provider file.\n\n```\nnpm install -g --ignore-scripts @earendil-works/pi-coding-agent\n```\n\nCheck the server answers first, using the id you gave `--alias`:\n\n```\ncurl http://127.0.0.1:8383/v1/models\n```\n\nThen add Frink as a provider in `~/.pi/agent/models.json`:\n\n```\n{\n  \"providers\": {\n    \"frink\": {\n      \"baseUrl\": \"http://127.0.0.1:8383/v1\",\n      \"api\": \"openai-completions\",\n      \"apiKey\": \"frink\",\n      \"models\": [\n        {\n          \"id\": \"bonsai-2-27b\",\n          \"name\": \"Bonsai 2 27B ternary, local via Frink\",\n          \"contextWindow\": 16384\n        }\n      ]\n    }\n  }\n}\n```\n\nFrink does not validate the key, but Pi wants a non-empty one. Then `cd` into a repository, run `pi`, pick the Frink entry with `/model`, and start with something small and checkable. Two practical notes: keep `max_tokens` generous, because a thinking model spends part of the budget before it writes any code, and a few tokens per second suits reviewing each step rather than firing and forgetting.\n\n## Links\n\n## AI full disclosure\n\nThis software is developed with strong assistance from Cursor, Grok 4.5, GPT 5.6, and Claude Fable 5, with humans leading the ideas, testing, and debugging. We say this openly because it shaped how the project was built. If you are not happy with AI-developed code, this software is not for you.\n\n## Acknowledgements\n\nFrink does not link against GGML, but exists thanks to the path opened by the llama.cpp project and the kernels, quantization formats, GGUF ecosystem, and hard-won engineering knowledge developed there. The ternary format, the Hadamard fold and the model itself are PrismML’s, and the Metal kernel in this release is a port of the design in their fork. We keep the GGML authors’ copyright notice in [docs/THIRD_PARTY_NOTICES.md](https://github.com/antonellof/frink/blob/main/docs/THIRD_PARTY_NOTICES.md).", "url": "https://wpnews.pro/news/bonsai-a-27b-reasoning-model-on-a-16-gb-m2-mac-with-frink", "canonical_source": "https://www.fratepietro.com/2026/frink-bonsai-ternary-27b-local/", "published_at": "2026-09-17 22:00:00+00:00", "updated_at": "2026-10-08 10:48:10.818483+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "machine-learning", "ai-tools", "ai-research"], "entities": ["PrismML", "Ternary-Bonsai-2-27B", "Frink", "Qwen3.8-27B", "PTQ1_0", "PQ2_0", "llama.cpp", "Apple M2"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/bonsai-a-27b-reasoning-model-on-a-16-gb-m2-mac-with-frink", "markdown": "https://wpnews.pro/news/bonsai-a-27b-reasoning-model-on-a-16-gb-m2-mac-with-frink.md", "text": "https://wpnews.pro/news/bonsai-a-27b-reasoning-model-on-a-16-gb-m2-mac-with-frink.txt", "jsonld": "https://wpnews.pro/news/bonsai-a-27b-reasoning-model-on-a-16-gb-m2-mac-with-frink.jsonld"}}