llamafile v0.10.5
Mozilla's llamafile v0.10.5 release adds support for two large local models: the 6GB Ternary Bonsai 27B and Poolside's 118B Laguna-S-2.1 coding MoE, by syncing with a more recent llama.cpp core. The u…
Mozilla's llamafile v0.10.5 release adds support for two large local models: the 6GB Ternary Bonsai 27B and Poolside's 118B Laguna-S-2.1 coding MoE, by syncing with a more recent llama.cpp core. The u…
DeepGrove's Maple-Preview, a 20B-parameter mixture-of-experts reasoning model with ternary weights, achieves 120 tokens per second on an iPhone and 218 tok/s on a base M4 Mac mini, according to the co…
PrismML released the 1-bit Bonsai-27B language model, deployable via a specialized fork of llama.cpp with CUDA kernels for the Q1_0_g128 GGUF format. The model requires only ~5.2 GB peak memory at 4K …
A new architecture called CaSA (Charge-Sharing Architecture) performs LLM inference directly inside commodity DRAM using processing-in-memory, bypassing the memory bus to solve the memory wall bottlen…
A developer known as pcdeni has created CaSA, an architecture that runs PrismML's ternary Bonsai LLM models directly inside commodity DRAM by breaking DDR4 timing rules and using charge-sharing, bypas…
PrismML CEO told CNBC that Apple is interested in the startup's technology, coinciding with the release of its Bonsai 27B model designed to run on iPhones, iPads, and Macs. The AI startup's press rele…
Apple is in talks to acquire PrismML, a startup that shrinks AI models up to 15x to run on iPhones, according to a report from CNBC.…
RightNow AI open-sourced bonsai-turbo, a batch-1 decode engine that runs PrismML's Bonsai 27B ternary LLM 1.76x faster than the official llama.cpp fork on an H100, achieving 151 tok/s (ternary) and 15…
PrismML, a Caltech-spinout AI company backed by Khosla Ventures, Cerberus, Google, and Samsung, announced Bonsai 27B on July 14, the first 27.8-billion-parameter AI model that runs locally on mobile d…
PrismML released Bonsai 27B, a 27-billion-parameter AI model compressed to 3.9 GB that runs on an iPhone 17 Pro Max at 11 tokens per second, the first model at that capability tier to fit on a smartph…
Alibaba's U.S.-listed shares rose up to 4% in premarket trading after the company confirmed its Qwen AI model will power Apple Intelligence features for users in China across iOS, iPadOS, macOS, and v…
PrismML has compressed a 27-billion-parameter AI model to under 4 GB, enabling it to run on an iPhone while retaining 90 percent of the original model's performance in company benchmarks, with math an…
PrismML has compressed a 27-billion-parameter AI model to under 4 GB, small enough to run on an iPhone, with the smallest version retaining 90 percent of original performance in benchmarks. Apple is r…
PrismML, a Caltech spinout, compressed Alibaba's Qwen3.6-27B AI model from 54 gigabytes to 3.9 gigabytes using 1-bit and ternary quantization, enabling it to run natively on an iPhone 17 Pro at 11 tok…
Apple has been in talks with bankers and semiconductor startups about potential acquisitions to strengthen its AI server chip capabilities, according to The Information. The company's next-generation …
Apple is in early talks with PrismML, a startup that compresses large AI models to run on a phone, as the company seeks to run more of Siri's work locally. PrismML CEO Babak Hassibi told CNBC that App…
Alibaba's US-listed shares rose 3.7% on Wednesday after China's Cyberspace Administration approved Alibaba's Qwen AI model to power Apple Intelligence across iOS, iPadOS, macOS, and visionOS for users…
PrismML has unveiled a compressed version of Alibaba's Qwen model that slashes memory requirements by up to 15 times, a technical leap that could accelerate Apple's AI ambitions by enabling more sophi…
PrismML has developed a technique to run a 27-billion-parameter AI model in just 4GB of RAM, potentially enabling on-device AI on iPhones and giving Apple a competitive edge in mobile AI.…
PrismML released Bonsai 27B, a 27-billion parameter AI model compressed to 3.9GB via native 1-bit training that runs on an iPhone 17 Pro at 11 tokens per second, but developers report tool-calling reg…