ESP32 BitNet Cluster Runs a 0.4B LLM for $28 Developer Low-Zi-Hong published a working distributed language model inference engine on GitHub that runs a 0.4-billion-parameter model across seven ESP32-S3 microcontrollers in a SPI daisy-chain pipeline for roughly $28 in total hardware, with each chip costing about $4. The project, which relies on Microsoft BitNet's 1.58-bit quantization technique, reached Hacker News's front page and drew attention from the embedded AI community. A developer named Low-Zi-Hong published a working distributed language model inference engine on GitHub over the weekend, and it is exactly as surprising as it sounds. Seven ESP32-S3 microcontrollers — chips you can buy for about $4 apiece — are running a 0.4-billion-parameter language model in a SPI daisy-chain pipeline. Total hardware cost: around $28. The project landed on Hacker News’s front page today and has the embedded AI community paying close attention. The reason this is possible in 2026, when it would have been absurd in 2024, is a technique called 1.58-bit quantization. Microsoft’s BitNet research team published a … The post