{"slug": "occulink-and-llama-cpp", "title": "Occulink and llama.cpp", "summary": "A hobbyist reports successfully running local AI inference on a Dell E7470 laptop with an OCuLink eGPU adapter and a Radeon Vega 64, achieving 25–40 tokens per second. The setup, which uses 8GB Vega 64 cards costing $60–70 each, can provide 24–32GB VRAM with 3–4 cards, enough for a 30B dense or 35B MoE model, and the user claims it codes well with Qwen 3.6 35B MoE with CPU offload.", "body_md": "This might be late ( I am new to this forum). But I have done it.\n\nI used a very old E7470 Dell laptop.\n\nBroke off a tab on the back cover and have m.2 to occulink adapter, occulink cable, a nice 750TX Corsair PSU and an egpu to occulink board.\n\nThe GPU used was an 8GB radeon vega 64.\n\nI am so surprised that I actually have it on sale and not more people are running this setup on cheap XEONs.\n\n3 to 4 vega64s will give you 24GB to 32GB VRAM at 60/70$/card. That will fit a 30B dense or 35B MoE model easily with tensor-split.\n\nThe next meaningful jump is 64GB cards and that doesn’t give you proportional benefit compared to $/gb of vram.\n\nMy setup runs local inference at a very decent 25 to 40T/s depending on models loaded.\n\nIt codes fairly complex code well with Qwen 3.6 35B moe with some cpu offload.\n\nIts meaningfully fast at 20+ TPS consistently, and up to 40tps (not 5 to 8 t/s that people live with).", "url": "https://wpnews.pro/news/occulink-and-llama-cpp", "canonical_source": "https://forum.level1techs.com/t/occulink-and-llama-cpp/253245#post_3", "published_at": "2026-08-31 17:23:37+00:00", "updated_at": "2026-08-31 17:51:47.903767+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-tools"], "entities": ["Dell E7470", "OCuLink", "Radeon Vega 64", "Corsair 750TX", "Qwen 3.6 35B MoE"], "alternates": {"html": "https://wpnews.pro/news/occulink-and-llama-cpp", "markdown": "https://wpnews.pro/news/occulink-and-llama-cpp.md", "text": "https://wpnews.pro/news/occulink-and-llama-cpp.txt", "jsonld": "https://wpnews.pro/news/occulink-and-llama-cpp.jsonld"}}