Ternary Bonsai 2 27B Ternary Bonsai 2 27B, a ternary-weight multimodal model built on a Qwen3.8 27B base, ships at roughly 6GB with weights restricted to {-1, 0, +1} at 1.76 bits per weight with FP16 group scales, according to the model's provider. The 27.4B-parameter model offers hybrid attention, a 262k context window, and vision input, and retains 98.2 percent of the FP16 baseline on a 20-benchmark thinking suite while running on llama.cpp and MLX. The provider reports the model adds no embedded watermarks or provenance metadata to generated output, and no per-token API provider pricing is tracked for it yet. Ternary Bonsai 2 27B consumer Ternary-weight multimodal model on a Qwen3.8 27B base: weights are values in {-1, 0, +1} at 1.76 bits per weight with FP16 group scales, so the full 27B model ships at roughly 6GB. Hybrid attention, 262k context, vision input. Keeps 98.2 percent of the FP16 baseline on a 20-benchmark thinking suite. Runs on llama.cpp and MLX. AI-generated content marks The provider reports that this model does not add embedded watermarks or provenance metadata to generated output. - 27.4B - 262k - apache 2.0 - Undisclosed - Sep 2026 Save your hardware and every model page answers the real question: will it run on your machine, and how fast? Join free - save your rig → https://tokenstead.ai/login?return to=%2Fonboarding Or run it in the cloud No per-token API provider pricing tracked for Ternary Bonsai 2 27B yet. For flagship list prices, see the calculator https://tokenstead.ai/calculator . Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.