# Occulink and llama.cpp

> Source: <https://forum.level1techs.com/t/occulink-and-llama-cpp/253245#post_3>
> Published: 2026-08-31 17:23:37+00:00

This might be late ( I am new to this forum). But I have done it.

I used a very old E7470 Dell laptop.

Broke off a tab on the back cover and have m.2 to occulink adapter, occulink cable, a nice 750TX Corsair PSU and an egpu to occulink board.

The GPU used was an 8GB radeon vega 64.

I am so surprised that I actually have it on sale and not more people are running this setup on cheap XEONs.

3 to 4 vega64s will give you 24GB to 32GB VRAM at 60/70$/card. That will fit a 30B dense or 35B MoE model easily with tensor-split.

The next meaningful jump is 64GB cards and that doesn’t give you proportional benefit compared to $/gb of vram.

My setup runs local inference at a very decent 25 to 40T/s depending on models loaded.

It codes fairly complex code well with Qwen 3.6 35B moe with some cpu offload.

Its meaningfully fast at 20+ TPS consistently, and up to 40tps (not 5 to 8 t/s that people live with).
