DS_DV 1
Hi Level 1 Techs (: im not quite sure if this is the best category since its hardware question regarding AI performance. (so if not im sorry)
in a recent Video Wendel mentioned that used Nvidia V100s are somewhat affordable to get into local AI.
But when i checked they are 800-1000 Bucks.
So building a System with two of them like showcased in the Video will come to at least 3k+ with the other nessesarry components.
While the AMD 395 comes with over 100GB unifed Memmory. So my Question is when it comes to AI is it worth building a V100 Setup over for example the MINISFORUM MS-S1 MAX which is only 3k-4k
with kind regards
I tend to think of them like cars:
V100 = GT-R R34
GB300 = GT-R NISMO
Strix Halo = Riced out prius
Spark = Riced out civic with “we made it big” energy
drbsg 3
This very much depends on what you want to do (and is a question I have been pondering).
The Strix Halo machines offer a lot of slower VRAM and simplicity. All drivers are recent and everything (IIRC) is now supported. The machine draws comparatively little power.
The V100s use a lot of power, you will need three or four of them to get the same amount of VRAM, and you need to mess around with older CUDA drivers and custom builds of e.g. llama.cpp.
So do you want a machine to play with LLMs on, get some work done, or a project to tinker with?
DS_DV 4
this is a very good question.
Currently AI is not in any of my workflows.
but ofc i follow the topic.
So far i could not see any workflow where it would benefit me.
But i recently needed to help a friend with creative work.
And since AI stole learned all the creative work i consider taking a peak into Image generation
Or assisted editing.
Ofc if i really buy into it i would be tempted to test out the other options too ^^
but
We managed to get 400tps on 4X V100-16GB running single-stream 27B-NVFP4-DFlash 2
And 8way at 800tps
DS_DV 6
Wow that is a super expensive Setup o.o
tbh i think if i try to tinker with this i rather stay with my allready expensive 7900xtx.
Sure it has only 24GB VRAM but its already there.
A 3-4k Minisforum pc is more expensive then most gpus a few years back
I guess it wont print me money so i wont spend a car or house on it
You can experiment with the latest hotness - Qwen3.8 Flash (170B params) with MoE offload on that card and get around a nice 25-35t/s if it’s DDR5 dual channel using 60k context out of 90k you’ll have available @ Q4 quant - if you have 64GB DRAM.
Only have 32GB DRAM? Some are playing with Q1 sizes by off the n-gran table to SSD (it’s fine) and still have 262k context window available within 24GB VRAM (off ~39 layers to CPU/32GB DRAM).
The key components to speed with MoE like this is both PCIe and DDR4/DDR5 bandwidth. My 9yr old quad channel DDR4 ThreadRipper 2950X clocks in around 80GB/s, which is faster than dual channel DDR5 - though only 15GB/s PCIe 3.0 speed. An $800 Rome motherboard will yield you eight channels of DDR4 and PCIe 4.0 speeds.
This Qwen3.8 Flash architecture is a preview of the format they will be using for Qwen4.0.
DS_DV 8
thank you for the hint (: i currently have 64GB 6000Mhz dual channel
with that recommendation i think ill definitely skip any investment and try it on the gaming system (: