# PrismML launches Bonsai 2 27B, a high-intelligence AI model so small it fits on consumer hardware

> Source: <https://siliconangle.com/2026/09/18/prismml-launches-bonsai-2-27b-a-high-intelligence-ai-model-so-small-it-fits-on-consumer-hardware/>
> Published: 2026-09-18 16:10:08+00:00

### PrismML launches Bonsai 2 27B, a high-intelligence AI model so small it fits on consumer hardware

[Prism ML Inc.](https://prismml.com/news/bonsai-2-27b) announced Thursday the launch of [Bonsai 2 27B](https://prismml.com/news/bonsai-2-27b), the second generation of its ultra-compact multimodal generative artificial intelligence small enough to fit on PCs and some high-end mobile devices.

The company said it used [ternary](https://en.wikipedia.org/wiki/Ternary_numeral_system), which uses three parts, to scale down its [Qwen3.8 27B-based](https://huggingface.co/Qwen/Qwen3.8-27B) model. Qwen3.8 weighs around 56 gigabytes at its full 16-bit uncompressed size, and Bonsai 2 reduces it to around 5.9 gigabytes while retaining around 98.2% of its capabilities.

Although it is possible to shrink AI models using other compression techniques called quantization, these methods usually strip away accuracy, knowledge and other systematic capabilities. Qwen3.8’s minimal memory footprint is around 9.4 gigabytes.

Ternary provides an interesting compression method when shrinking the “weights,” or parameters that make up the model. Weights are the model’s numerical dials that control how it processes information and generates outputs. In full-size models, these are represented by 16 bits; with PrismML’s approach, these are simplified down to ternary, or three bits, represented by +1, 0, and -1. This lets the company store information in a much smaller memory footprint while still holding onto reasonably high intelligence.

In essence, this allows Bonsai 2 to punch well above its weight class at a very small size.

On benchmarks, Bonsai 2 showed close performance on agentic and tool calling compared to Qwen3.8 within 3 points, at 77.6 and 79.8 respectively; with aggregate scores of 81.6 and 82.2 for coding across HumanEval+, LiveCodeBench v6, MBPP+ and BigCodeBench; and 82.7 and 81.3 for knowledge and reasoning across MMLU-Redux, GPQA Diamond and AA-LCR.

The model can run on an Nvidia GeForce GTX 5090 card without quantization, reaching 143 tokens per second and 46.8 tokens per second on Apple Inc.’s M5 Max chip. The company said the model consumes extremely low power per token at 0.714 megawatt-hours, making it 40% more energy-efficient than other 8B models running at full precision, meaning uncompressed.

Ultra-small models let users run AI locally on their own machines without sending inference to the cloud. Any time data is sent across the internet, there can be a delay in receiving a response, or sensitive information might be sent to a third party. Bringing intelligence onto a local machine eliminates third-party data sharing, keeps prompts and responses local, helps meet strict privacy regulations, and can improve security.

For example, simple translation, summarization, and search organization could run on device, while long-horizon task comprehension and research might need to be sent to an expensive cloud model. For an everyday user, or even an enterprise use case, running a local model that is far less expensive and respects privacy when a task is simple and involves sensitive information, while scaling up to highly intelligent, cloud-based models to handle complex, high-touch, goal-oriented work.

The new model runs on Nvidia graphics processing units via [CUDA](https://en.wikipedia.org/wiki/CUDA) and on Apple devices, including Mac, iPhone and iPad, via [MLX](<https://en.wikipedia.org/wiki/MLX_(machine_learning_framework)>), through low-bit kernels. Model weights [are available today](https://huggingface.co/collections/prism-ml/bonsai-2) under Apache 2.0 licenses.

##### Image: [Pixabay](https://pixabay.com/illustrations/ai-generated-global-communication-8942972/)

# A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.

- **15M+ viewers of theCUBE videos** , powering conversations across AI, cloud, cybersecurity and more
- **11.4k+ theCUBE alumni** — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

### Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: [https://siliconangle.com/aws-marketplace/](https://siliconangle.com/aws-marketplace/)

##### **About SiliconANGLE Media**

[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),

[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),

[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),

[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),

[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.
