{"slug": "nvidia-v100-vs-amd-395", "title": "Nvidia v100 vs AMD 395", "summary": "A Level 1 Techs forum thread weighing used Nvidia V100 GPUs against AMD Strix Halo machines for local AI found V100-16GB cards selling for $800-$1,000 each, making a two-card build cost at least $3,000, while Minisforum's MS-S1 MAX Strix Halo system with over 100GB unified memory runs $3,000-$4,000. Forum user drbsg noted the Strix Halo offers more slower VRAM with current drivers and low power draw, whereas V100s require three or four cards for equivalent VRAM plus older CUDA drivers and custom llama.cpp builds. Another poster reported 400 tokens per second on four V100-16GB cards running single-stream 27B-NVFP4-DFlash and 800 tokens per second across eight cards, and suggested running Qwen3.8 Flash (170B parameters) with MoE offload on a 24GB card at 25-35 tokens per second with 64GB of dual-channel DDR5.", "body_md": "[DS_DV](https://forum.level1techs.com/u/DS_DV)\n1\n \nHi Level 1 Techs (:\n\nim not quite sure if this is the best category since its hardware question regarding AI performance. (so if not im sorry)\n\nin a recent Video Wendel mentioned that used Nvidia V100s are somewhat affordable to get into local AI.\n\nBut when i checked they are 800-1000 Bucks.\n\nSo building a System with two of them like showcased in the Video will come to at least 3k+ with the other nessesarry components.\n\nWhile the AMD 395 comes with over 100GB unifed Memmory.\n\nSo my Question is when it comes to AI is it worth building a V100 Setup over for example the MINISFORUM MS-S1 MAX which is only 3k-4k\n\nwith kind regards\n\n \n\n \nI tend to think of them like cars:\n\nV100 = GT-R R34\n\nGB300 = GT-R NISMO\n\nStrix Halo = Riced out prius\n\nSpark = Riced out civic with “we made it big” energy\n\n \n\n \n[drbsg](https://forum.level1techs.com/u/drbsg)\n3\n \nThis very much depends on what you want to do (and is a question I have been pondering).\n\nThe Strix Halo machines offer a lot of slower VRAM and simplicity. All drivers are recent and everything (IIRC) is now supported. The machine draws comparatively little power.\n\nThe V100s use a lot of power, you will need three or four of them to get the same amount of VRAM, and you need to mess around with older CUDA drivers and custom builds of e.g. llama.cpp.\n\nSo do you want a machine to play with LLMs on, get some work done, or a project to tinker with?\n\n \n\n \n[DS_DV](https://forum.level1techs.com/u/DS_DV)\n4\n \nthis is a very good question.\n\nCurrently AI is not in any of my workflows.\n\nbut ofc i follow the topic.\n\nSo far i could not see any workflow where it would benefit me.\n\nBut i recently needed to help a friend with creative work.\n\nAnd since AI ~~stole~~ learned all the creative work i consider taking a peak into Image generation\n\nOr assisted editing.\n\nOfc if i really buy into it i would be tempted to test out the other options too ^^\n\nbut\n\n \n\n \nWe managed to get 400tps on 4X V100-16GB running single-stream 27B-NVFP4-DFlash 2\n\nAnd 8way at 800tps\n\n \n\n \n[DS_DV](https://forum.level1techs.com/u/DS_DV)\n6\n \nWow that is a super expensive Setup o.o\n\ntbh i think if i try to tinker with this i rather stay with my allready expensive 7900xtx.\n\nSure it has only 24GB VRAM but its already there.\n\nA 3-4k Minisforum pc is more expensive then most gpus a few years back \n\nI guess it wont print me money so i wont spend a car or house on it \n\n \n\n \nYou can experiment with the latest hotness - Qwen3.8 Flash (170B params) with MoE offload on that card and get around a nice 25-35t/s if it’s DDR5 dual channel using 60k context out of 90k you’ll have available @ Q4 quant - if you have 64GB DRAM.\n\nOnly have 32GB DRAM?  Some are playing with Q1 sizes by offloading the n-gran table to SSD (it’s fine) and still have 262k context window available within 24GB VRAM (offloading ~39 layers to CPU/32GB DRAM).\n\nThe key components to speed with MoE like this is both PCIe and DDR4/DDR5 bandwidth.  My 9yr old quad channel DDR4 ThreadRipper 2950X clocks in around 80GB/s, which is faster than dual channel DDR5 - though only 15GB/s PCIe 3.0 speed.  An $800 Rome motherboard will yield you eight channels of DDR4 and PCIe 4.0 speeds.\n\nThis Qwen3.8 Flash architecture is a preview of the format they will be using for Qwen4.0.\n\n \n\n \n[DS_DV](https://forum.level1techs.com/u/DS_DV)\n8\n \nthank you for the hint (:\n\ni currently have 64GB 6000Mhz dual channel \n\nwith that recommendation i think ill definitely skip any investment and try it on the gaming system (:", "url": "https://wpnews.pro/news/nvidia-v100-vs-amd-395", "canonical_source": "https://forum.level1techs.com/t/nvidia-v100-vs-amd-395/258040#post_8", "published_at": "2026-10-11 07:18:45+00:00", "updated_at": "2026-10-11 07:21:55.485034+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips", "large-language-models", "ai-tools"], "entities": ["Nvidia", "Nvidia V100", "AMD", "AMD Strix Halo", "Minisforum MS-S1 MAX", "Level 1 Techs", "Qwen3.8 Flash", "llama.cpp"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/nvidia-v100-vs-amd-395", "markdown": "https://wpnews.pro/news/nvidia-v100-vs-amd-395.md", "text": "https://wpnews.pro/news/nvidia-v100-vs-amd-395.txt", "jsonld": "https://wpnews.pro/news/nvidia-v100-vs-amd-395.jsonld"}}