Running a 35B LLM at 128K Context, Full Speed, on €870 of Used Hardware A developer demonstrated running a 35B-parameter large language model with 128K context at full speed using only €870 of used hardware, eliminating the need for cloud services. The setup leverages cost-effective second-hand components to achieve high-performance inference, highlighting the feasibility of local AI deployment. Article URL: https://medium.com/ai-advances/running-a-35b-llm-at-128k-context-full-speed-on-870-of-used-hardware-no-cloud-required-c4f7629810b8 Comments URL: https://news.ycombinator.com/item?id=49142301 Points: 1 Comments: 0