21:29
2026-08-10
promptcube3.com
artificial-intelligence
Needle 2 fits a functional LLM into just 14MB
Cactus Compute's Needle 2 model fits a functional LLM into just 14MB, using Simple Attention Networks to reduce per-token computation to 70 MFLOPs, compared to 164 MFLOPs for a standard transformer ofโฆ