# Samsung's LittleBit squeezes AI models to a tenth of a bit per weight

> Source: <https://startupfortune.com/samsungs-littlebit-squeezes-ai-models-to-a-tenth-of-a-bit-per-weight/>
> Published: 2026-10-08 15:00:37+00:00

*Samsung Research says its LittleBit technique can shrink a 13-billion-parameter model to under a gigabyte, at a compression level most quantization methods can't survive without breaking the model.*

Run a large language model today and you're almost certainly asking it to store its knowledge in 16-bit or 8-bit numbers, one for every parameter. Samsung Research built something that gets the same model down to roughly 0.1 bits per weight, a nearly 31x memory cut, according to the team's paper presented as a NeurIPS 2025 poster in December. Llama2-13B, which normally needs tens of gigabytes, fits under 0.9GB. Llama2-70B comes in under 2GB. That's phone territory.

The method is called LittleBit, and it works by splitting each weight matrix into smaller latent factors through low-rank factorization, then binarizing those factors down to single bits. Pushed that far, quantized models usually fall apart. So Samsung's researchers, Banseok Lee, Dongkyu Kim, Youngcheon You and Youngmin Kim, added what they call a multi-scale compensation mechanism, learned importance weights across rows, columns and the latent dimension that claw back some of the accuracy the binarization would otherwise destroy. Two techniques do the heavy lifting: Dual Sign-Value-Independent Decomposition for initializing the quantization-aware training, and a residual compensation step that keeps chipping away at the approximation error.

The headline number from the paper is a strange one to brag about, until you see the comparison. At 0.1 bits per weight, LittleBit on Llama2-7B outperforms competing quantization methods running at 0.7 bits per weight, seven times the storage, according to the team's benchmarks. Samsung also says the compressed format unlocks up to an 11.6x inference speedup over FP16. The code is public on GitHub under SamsungLabs/LittleBit, and Samsung Research later published a LittleBit-2 post tied to ICML 2026, describing a successor that uses latent geometry alignment to push accuracy further across the same sub-1-bit regime.

Here's why this lands differently in October 2026 than it would have a year ago. The AI industry has spent most of this year fixated on a hardware story: not enough HBM, not enough advanced packaging, soaring memory prices as every hyperscaler fights for the same supply. Compression research like LittleBit is the other half of that equation, the argument that you don't always need a bigger chip if you can make the model smaller.

[PrismML Squeezes a 27-Billion-Parameter AI Model Into 5.9 Gigabytes](https://startupfortune.com/prismml-squeezes-a-27-billion-parameter-ai-model-into-59-gigabytes/)

PrismML released Ternary Bonsai 2 27B, a 27.8-billion-parameter open-weight model that fits in just 5.9GB while keeping 98.2% of its full-precision benchmark score. Free under Apache 2.0, it runs on laptops and phones and challenges the idea that capable AI needs expensive hardware. - [how to compress large AI models for laptops](https://startupfortune.com/prismml-squeezes-a-27-billion-parameter-ai-model-into-59-gigabytes/) - [billion parameter model file size reduction techniques](https://startupfortune.com/prismml-squeezes-a-27-billion-parameter-ai-model-into-59-gigabytes/)

It isn't alone in that race, and the company it's up against for headlines right now isn't a chipmaker. The Information reported in July that Apple had held meetings with PrismML, and CNBC later reported that Apple was in early talks with the Caltech-rooted startup. PrismML's Bonsai 27B, based on Qwen3.6-27B, compresses the model from about 54GB to as little as 3.9GB, small enough for recent high-end iPhones. The company offers binary and ternary versions and says Bonsai 27B delivers a 10x to 15x memory reduction, a 6x to 8x speed boost, and 3x to 6x lower energy use. PrismML emerged from stealth in March 2026 and raised $16.25 million from Khosla Ventures, Cerberus Ventures and Caltech, while Google is listed by the company as a supporter. Its CEO, Babak Hassibi, told CNBC the Apple discussions were early but "progressing nicely."

None of this comes free. PrismML's own numbers show the tradeoff tightening as compression gets more aggressive: its ternary Bonsai 27B retains 95% of the full-precision baseline across a 15-benchmark suite, while the 1-bit version retains 90%. That's the tradeoff every compression method on the market is negotiating in its own way, and Samsung's compensation mechanisms and PrismML's calibrated value reduction are both, in effect, bets on where to spend the error budget.

Google is working the problem from a different angle entirely. Its TurboQuant algorithm, unveiled in late March and headed for ICLR 2026, doesn't touch model weights at all. It compresses the key-value cache, the memory that balloons as a model holds a longer conversation, to roughly 3 bits per value with what Google says is no accuracy compromise, while 4-bit TurboQuant showed up to an 8x speedup in attention-logit computation on Nvidia H100 hardware. The announcement was big enough to knock memory stocks around: SK Hynix fell 6%, Samsung dropped 5%, Micron slid 3%, all on the logic that less memory demand per AI workload means less urgency to buy more chips. Microsoft's BitNet framework, pursuing native 1-bit training rather than post-hoc compression, rounds out a field that now has at least four serious players attacking the same bottleneck from different directions.

Samsung hasn't said whether LittleBit is headed into its own phones or Galaxy Watch lineup, and the company's public materials describe it as a research contribution rather than a shipping product. But the timing isn't subtle. Samsung is simultaneously one of the two companies whose stock got hit by Google's memory-saving algorithm and one of the labs publishing techniques that make memory matter less. Whichever side of that tension wins out, the direction is the same: the industry no longer treats "the model needs more hardware" as the only available answer.

**Also read:** [Researchers find letting AI coding agents write their own tests backfires](https://startupfortune.com/researchers-find-letting-ai-coding-agents-write-their-own-tests-backfires/) • [Independent mathematicians are now Lean-checking OpenAI's proof claims line by line](https://startupfortune.com/independent-mathematicians-are-now-lean-checking-openais-proof-claims-line-by-line/) • [Marvell raises its AI chip forecast and Wall Street likes what it hears](https://startupfortune.com/marvell-raises-its-ai-chip-forecast-and-wall-street-likes-what-it-hears/)

*This article is posted in [AI News](https://startupfortune.com/category/ai/), check it out for more related stories.*

[A Caltech Startup Shrank a 27 Billion Parameter AI Model to Fit on an iPhone](https://startupfortune.com/a-caltech-startup-shrank-a-27-billion-parameter-ai-model-to-fit-on-an-iphone/)

PrismML, a Caltech spinout backed by Khosla Ventures, released Bonsai 27B on July 14, 2026, a 1-bit and ternary compressed version of Qwen3.6-27B that runs natively on an iPhone 17 Pro. CNBC reports Apple is already evaluating the technology for on-device AI. - [how to run AI models on](https://startupfortune.com/a-caltech-startup-shrank-a-27-billion-parameter-ai-model-to-fit-on-an-iphone/) - [bit quantization for mobile deployment](https://startupfortune.com/a-caltech-startup-shrank-a-27-billion-parameter-ai-model-to-fit-on-an-iphone/)

## Join the discussion

[Open in the community →](https://startupfortune.com/community/)

Almost there. Sign in and your reply posts straight away.
