00:00
2026-08-09
mikeayles.com
artificial-intelligence
On-Chip LLM: War Stories and the Method
A $250 FPGA-based on-chip LLM project achieved a record 59,965.5 tokens per second with a split-brain N=16 configuration at 200 MHz, according to the developer's blog post. The project, built with hanβ¦