LLMs on Consumer Hardware — Part 2: Prefill and the Failure of the AI PC
A developer's benchmark of LLM inference on consumer hardware reveals that prefill speed, not generation, is the bottleneck for large prompts, with an 18-fold variation across machines. The test also …