cd /news/ai-infrastructure/nvidia-v100-vs-amd-395 · home › topics › ai-infrastructure › article
[ARTICLE · art-149059] src=forum.level1techs.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Nvidia v100 vs AMD 395

A Level 1 Techs forum thread weighing used Nvidia V100 GPUs against AMD Strix Halo machines for local AI found V100-16GB cards selling for $800-$1,000 each, making a two-card build cost at least $3,000, while Minisforum's MS-S1 MAX Strix Halo system with over 100GB unified memory runs $3,000-$4,000. Forum user drbsg noted the Strix Halo offers more slower VRAM with current drivers and low power draw, whereas V100s require three or four cards for equivalent VRAM plus older CUDA drivers and custom llama.cpp builds. Another poster reported 400 tokens per second on four V100-16GB cards running single-stream 27B-NVFP4-DFlash and 800 tokens per second across eight cards, and suggested running Qwen3.8 Flash (170B parameters) with MoE offload on a 24GB card at 25-35 tokens per second with 64GB of dual-channel DDR5.

read3 min views1 publishedOct 11, 2026
Nvidia v100 vs AMD 395
Image: Forum (auto-discovered)

DS_DV 1

Hi Level 1 Techs (: im not quite sure if this is the best category since its hardware question regarding AI performance. (so if not im sorry)

in a recent Video Wendel mentioned that used Nvidia V100s are somewhat affordable to get into local AI.

But when i checked they are 800-1000 Bucks.

So building a System with two of them like showcased in the Video will come to at least 3k+ with the other nessesarry components.

While the AMD 395 comes with over 100GB unifed Memmory. So my Question is when it comes to AI is it worth building a V100 Setup over for example the MINISFORUM MS-S1 MAX which is only 3k-4k

with kind regards

I tend to think of them like cars:

V100 = GT-R R34

GB300 = GT-R NISMO

Strix Halo = Riced out prius

Spark = Riced out civic with “we made it big” energy

drbsg 3

This very much depends on what you want to do (and is a question I have been pondering).

The Strix Halo machines offer a lot of slower VRAM and simplicity. All drivers are recent and everything (IIRC) is now supported. The machine draws comparatively little power.

The V100s use a lot of power, you will need three or four of them to get the same amount of VRAM, and you need to mess around with older CUDA drivers and custom builds of e.g. llama.cpp.

So do you want a machine to play with LLMs on, get some work done, or a project to tinker with?

DS_DV 4

this is a very good question.

Currently AI is not in any of my workflows.

but ofc i follow the topic.

So far i could not see any workflow where it would benefit me.

But i recently needed to help a friend with creative work.

And since AI stole learned all the creative work i consider taking a peak into Image generation

Or assisted editing.

Ofc if i really buy into it i would be tempted to test out the other options too ^^

but

We managed to get 400tps on 4X V100-16GB running single-stream 27B-NVFP4-DFlash 2

And 8way at 800tps

DS_DV 6

Wow that is a super expensive Setup o.o

tbh i think if i try to tinker with this i rather stay with my allready expensive 7900xtx.

Sure it has only 24GB VRAM but its already there.

A 3-4k Minisforum pc is more expensive then most gpus a few years back

I guess it wont print me money so i wont spend a car or house on it

You can experiment with the latest hotness - Qwen3.8 Flash (170B params) with MoE offload on that card and get around a nice 25-35t/s if it’s DDR5 dual channel using 60k context out of 90k you’ll have available @ Q4 quant - if you have 64GB DRAM.

Only have 32GB DRAM? Some are playing with Q1 sizes by off the n-gran table to SSD (it’s fine) and still have 262k context window available within 24GB VRAM (off ~39 layers to CPU/32GB DRAM).

The key components to speed with MoE like this is both PCIe and DDR4/DDR5 bandwidth. My 9yr old quad channel DDR4 ThreadRipper 2950X clocks in around 80GB/s, which is faster than dual channel DDR5 - though only 15GB/s PCIe 3.0 speed. An $800 Rome motherboard will yield you eight channels of DDR4 and PCIe 4.0 speeds.

This Qwen3.8 Flash architecture is a preview of the format they will be using for Qwen4.0.

DS_DV 8

thank you for the hint (: i currently have 64GB 6000Mhz dual channel

with that recommendation i think ill definitely skip any investment and try it on the gaming system (:

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nvidia-v100-vs-amd-3…] indexed:0 read:3min 2026-10-11 · —