JBinFla 1 So I’m late to the party guys. A couple months back finally started paying for Codex and wow. But. I don’t want to get API fees so I’d like to stay local. I need a GPU as I don’t wanna use my laptop 5080 and my server is an RX5700 I use for Plex transcoding. I’m looking at the Strix Halo stuff but am wondering if I get two AMD R9700 32gb cards (card vs card roughly equivalent perf to 9070 XT). This is way faster and will have 64gb VRAM for about $3k, and since my server is a threadripper (I know) I have enough slots and lanes to go to 4 cards at 128gb should the need arise.
Question: Can I use 2x 32gb cards roughly equivalent to one 64gb card just twice as fast or does the model have to fit on one GPU?
You can get two cards and run them once as fast. This is still pretty great.
the idea would be to up the vram total to be able to load a larger model / larger context / overall less quantized llm - at speed. with 2x r9700 you will be having a good time with Qwen3.8-27B. 4x r9700 then Qwen3.8-Flash-Next, Deepseek v4 Flash 0731 / GLM 5.3 flash (with some ram overflow) is probably where you will be aiming at - depending on use case (eg how many users need access to the ai?)