A Caltech Startup Shrank a 27 Billion Parameter AI Model to Fit on an iPhone
PrismML, a Caltech spinout, compressed Alibaba's Qwen3.6-27B AI model from 54 gigabytes to 3.9 gigabytes using 1-bit and ternary quantization, enabling it to run natively on an iPhone 17 Pro at 11 tok…