This might be late ( I am new to this forum). But I have done it.
I used a very old E7470 Dell laptop.
Broke off a tab on the back cover and have m.2 to occulink adapter, occulink cable, a nice 750TX Corsair PSU and an egpu to occulink board.
The GPU used was an 8GB radeon vega 64.
I am so surprised that I actually have it on sale and not more people are running this setup on cheap XEONs.
3 to 4 vega64s will give you 24GB to 32GB VRAM at 60/70$/card. That will fit a 30B dense or 35B MoE model easily with tensor-split.
The next meaningful jump is 64GB cards and that doesn’t give you proportional benefit compared to $/gb of vram.
My setup runs local inference at a very decent 25 to 40T/s depending on models loaded.
It codes fairly complex code well with Qwen 3.6 35B moe with some cpu offload.
Its meaningfully fast at 20+ TPS consistently, and up to 40tps (not 5 to 8 t/s that people live with).