cd /news/large-language-models/running-a-35b-llm-at-128k-context-fu… · home topics large-language-models article
[ARTICLE · art-83569] src=medium.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Running a 35B LLM at 128K Context, Full Speed, on €870 of Used Hardware

A developer demonstrated running a 35B-parameter large language model with 128K context at full speed using only €870 of used hardware, eliminating the need for cloud services. The setup leverages cost-effective second-hand components to achieve high-performance inference, highlighting the feasibility of local AI deployment.

read1 min views1 publishedAug 2, 2026
Article URL: https://medium.com/ai-advances/running-a-35b-llm-at-128k-context-full-speed-on-870-of-used-hardware-no-cloud-required-c4f7629810b8

Comments URL: https://news.ycombinator.com/item?id=49142301

Points: 1

── more in #large-language-models 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/running-a-35b-llm-at…] indexed:0 read:1min 2026-08-02 ·