11:52
2026-08-10
mikeayles.com
artificial-intelligence
Show HN: Taalas-style on-chip LLM weights on a $250 AMD FPGA (60k tok/s)
A developer achieved 59,965 tokens per second running a 3.16M-parameter INT4 transformer entirely in the on-chip memory of a $250 AMD Xilinx Kria KV260 FPGA, with zero DRAM in the token loop. The bit-โฆ