Occulink and llama.cpp A hobbyist reports successfully running local AI inference on a Dell E7470 laptop with an OCuLink eGPU adapter and a Radeon Vega 64, achieving 25–40 tokens per second. The setup, which uses 8GB Vega 64 cards costing $60–70 each, can provide 24–32GB VRAM with 3–4 cards, enough for a 30B dense or 35B MoE model, and the user claims it codes well with Qwen 3.6 35B MoE with CPU offload. This might be late I am new to this forum . But I have done it. I used a very old E7470 Dell laptop. Broke off a tab on the back cover and have m.2 to occulink adapter, occulink cable, a nice 750TX Corsair PSU and an egpu to occulink board. The GPU used was an 8GB radeon vega 64. I am so surprised that I actually have it on sale and not more people are running this setup on cheap XEONs. 3 to 4 vega64s will give you 24GB to 32GB VRAM at 60/70$/card. That will fit a 30B dense or 35B MoE model easily with tensor-split. The next meaningful jump is 64GB cards and that doesn’t give you proportional benefit compared to $/gb of vram. My setup runs local inference at a very decent 25 to 40T/s depending on models loaded. It codes fairly complex code well with Qwen 3.6 35B moe with some cpu offload. Its meaningfully fast at 20+ TPS consistently, and up to 40tps not 5 to 8 t/s that people live with .