Testing PrismML Bonsai 2
PrismML released Bonsai 2 27B, a ternary-weight compression of Qwen3.8 27B that cuts model size from about 56 GB to 5.9 GB while retaining roughly 98% of the original model's benchmark score. In a han…
PrismML released Bonsai 2 27B, a ternary-weight compression of Qwen3.8 27B that cuts model size from about 56 GB to 5.9 GB while retaining roughly 98% of the original model's benchmark score. In a han…
A tester ran Bonsai 2 27B, a ternary model that stores each weight as one of three values (1, 0, -1), locally on a 12GB Nvidia GPU using a special build of llama-cpp from PrismML's GitHub. The 7.2 GB …
PrismML released a ternary-weight version of Qwen3.8-27B under its Bonsai family on Tuesday, cutting the model from more than 50 GB of weights at full precision to just under 6 GB, which lets it run o…
PrismML released Bonsai 2 27B on September 17, 2026, a ternary-quantized version of Alibaba's open Qwen3.8-27B that retains 98.2% of the full-precision model's benchmark average while shrinking from 5…
PrismML released Ternary Bonsai 2 27B on September 17, an open-weight model built on Alibaba's Qwen3.8 27B that compresses 27.8 billion parameters to roughly 1.76 bits per weight and a 5.9-gigabyte fi…
PrismML released Bonsai 2 27B on Thursday, a compressed version of Alibaba's Qwen3.8 27B open-source model shrunk to 5.9 GB — a 9x to 10x memory reduction that matches 98% of Qwen's aggregate benchmar…
Ferrox v0.23.0, released today, adds support for PrismML's Ternary-Bonsai-2-27B, a 27-billion-parameter model quantized to 1.75 bits per weight that occupies 5.5 GB on disk and runs on a 16 GB laptop,…
PrismML launched Bonsai 2 27B on September 17th, an Apache 2.0-licensed ternary-weight version of Qwen3.8 27B whose language model fits in 5.95 GB and, according to the company, preserves 98.2% of the…
Edge0, an open-source framework posted to arXiv on September 16, 2026, reports running a 4-bit Qwen3.6-35B-A3B mixture-of-experts model at 20.4 tokens per second inside 2.9 GiB of peak active memory o…
PrismML's 1-bit quantized Bonsai 27B model, weighing 3.9 GB, runs on a consumer GPU with an NVIDIA GeForce RTX 5060, achieving 10-20 tokens per second on average, but slower than smaller models like Q…
Mozilla's llamafile v0.10.5 release adds support for two large local models: the 6GB Ternary Bonsai 27B and Poolside's 118B Laguna-S-2.1 coding MoE, by syncing with a more recent llama.cpp core. The u…
DeepGrove's Maple-Preview, a 20B-parameter mixture-of-experts reasoning model with ternary weights, achieves 120 tokens per second on an iPhone and 218 tok/s on a base M4 Mac mini, according to the co…
PrismML released the 1-bit Bonsai-27B language model, deployable via a specialized fork of llama.cpp with CUDA kernels for the Q1_0_g128 GGUF format. The model requires only ~5.2 GB peak memory at 4K …
A new architecture called CaSA (Charge-Sharing Architecture) performs LLM inference directly inside commodity DRAM using processing-in-memory, bypassing the memory bus to solve the memory wall bottlen…
A developer known as pcdeni has created CaSA, an architecture that runs PrismML's ternary Bonsai LLM models directly inside commodity DRAM by breaking DDR4 timing rules and using charge-sharing, bypas…
PrismML CEO told CNBC that Apple is interested in the startup's technology, coinciding with the release of its Bonsai 27B model designed to run on iPhones, iPads, and Macs. The AI startup's press rele…
RightNow AI open-sourced bonsai-turbo, a batch-1 decode engine that runs PrismML's Bonsai 27B ternary LLM 1.76x faster than the official llama.cpp fork on an H100, achieving 151 tok/s (ternary) and 15…
PrismML, a Caltech-spinout AI company backed by Khosla Ventures, Cerberus, Google, and Samsung, announced Bonsai 27B on July 14, the first 27.8-billion-parameter AI model that runs locally on mobile d…
PrismML released Bonsai 27B, a 27-billion-parameter AI model compressed to 3.9 GB that runs on an iPhone 17 Pro Max at 11 tokens per second, the first model at that capability tier to fit on a smartph…
Alibaba's U.S.-listed shares rose up to 4% in premarket trading after the company confirmed its Qwen AI model will power Apple Intelligence features for users in China across iOS, iPadOS, macOS, and v…