cd/entity/PrismML· home entities PrismML
grep -l @prismml /news/*.json | wc -l → 39

PrismML

mentions 39 type Organization page 1/2 feed RSS

// recent coverage 39 mentions

12:00
2026-09-18
julin.ai
large-language-models

Testing PrismML Bonsai 2

PrismML released Bonsai 2 27B, a ternary-weight compression of Qwen3.8 27B that cuts model size from about 56 GB to 5.9 GB while retaining roughly 98% of the original model's benchmark score. In a han…

07:00
2026-09-18
dotnetperls.com
large-language-models

Tested Bonsai 2 27B Locally (Ternary Model)

A tester ran Bonsai 2 27B, a ternary model that stores each weight as one of three values (1, 0, -1), locally on a 12GB Nvidia GPU using a special build of llama-cpp from PrismML's GitHub. The 7.2 GB …

06:22
2026-09-18
thewatershed.markpesce.com
large-language-models

Did the Device Watershed Just Arrive?

PrismML released a ternary-weight version of Qwen3.8-27B under its Bonsai family on Tuesday, cutting the model from more than 50 GB of weights at full precision to just under 6 GB, which lets it run o…

22:34
2026-09-17
techcrunch.com
large-language-models

PrismML hopes its tiny LLM will change how we all use AI

PrismML released Bonsai 2 27B on Thursday, a compressed version of Alibaba's Qwen3.8 27B open-source model shrunk to 5.9 GB — a 9x to 10x memory reduction that matches 98% of Qwen's aggregate benchmar…

22:00
2026-09-17
fratepietro.com
large-language-models

A 27B model in 5.5 GB: Bonsai ternary on Ferrox, locally

Ferrox v0.23.0, released today, adds support for PrismML's Ternary-Bonsai-2-27B, a 27-billion-parameter model quantized to 1.75 bits per weight that occupies 5.5 GB on disk and runs on a 16 GB laptop,…

20:58
2026-09-17
runtimewire.com
large-language-models

PrismML launches a 5.9 GB Qwen3.8 model for local AI

PrismML launched Bonsai 2 27B on September 17th, an Apache 2.0-licensed ternary-weight version of Qwen3.8 27B whose language model fits in 5.95 GB and, according to the company, preserves 98.2% of the…

00:00
2026-09-17
digitalapplied.com
large-language-models

Run a 35B AI Model on a 24GB Mac: Two New Ways Compared

Edge0, an open-source framework posted to arXiv on September 16, 2026, reports running a 4-bit Qwen3.6-35B-A3B mixture-of-experts model at 20.4 tokens per second inside 2.9 GiB of peak active memory o…

10:01
2026-08-05
blog.mozilla.ai
artificial-intelligence

llamafile v0.10.5

Mozilla's llamafile v0.10.5 release adds support for two large local models: the 6GB Ternary Bonsai 27B and Poolside's 118B Laguna-S-2.1 coding MoE, by syncing with a more recent llama.cpp core. The u…

09:10
2026-08-05
sourcefeed.dev
artificial-intelligence

How a 20B Model Hits 120 tok/s on an iPhone

DeepGrove's Maple-Preview, a 20B-parameter mixture-of-experts reasoning model with ternary weights, achieves 120 tokens per second on an iPhone and 218 tok/s on a base M4 Mac mini, according to the co…

22:03
2026-07-23
promptcube3.com
artificial-intelligence

CaSA: Computing LLM Inference Directly in RAM

A new architecture called CaSA (Charge-Sharing Architecture) performs LLM inference directly inside commodity DRAM using processing-in-memory, bypassing the memory bus to solve the memory wall bottlen…

10:08
2026-07-16
machinebrief.com
artificial-intelligence

PrismML Dangles Tech Carrot Before Apple: Bold or Just Bold-faced?

PrismML CEO told CNBC that Apple is interested in the startup's technology, coinciding with the release of its Bonsai 27B model designed to run on iPhones, iPads, and Macs. The AI startup's press rele…

21:15
2026-07-15
cryptobriefing.com
artificial-intelligence

Bonsai debuts as the first 27B AI model for mobile devices

PrismML, a Caltech-spinout AI company backed by Khosla Ventures, Cerberus, Google, and Samsung, announced Bonsai 27B on July 14, the first 27.8-billion-parameter AI model that runs locally on mobile d…

19:55
2026-07-15
decrypt.co
artificial-intelligence

Meet Bonsai: The First 27B AI Model That Fits on Your Phone

PrismML released Bonsai 27B, a 27-billion-parameter AI model compressed to 3.9 GB that runs on an iPhone 17 Pro Max at 11 tokens per second, the first model at that capability tier to fit on a smartph…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics