{"slug": "post-show-off-your-ai-rig-whatchya-got-in-there", "title": "Post / Show off your Ai Rig - Whatchya got in there?", "summary": "A user detailed a home AI inference rig built on an AMD Epyc Rome 7532 CPU with 128GB DDR4, roughly 5TB of NVMe storage, and 4x AMD V620 GPUs on an open-air crypto-mining bench. The rig runs a LiteLLM frontend to vLLM serving Qwen3.8-flash-next at int4 quantization with the n-gram table offloaded to system memory, reaching about 70 tokens per second for single-request inference and roughly 90TPS peak across 2-3 simultaneous streams. The user reported that multi-token prediction (MTP) is not reliably performant with multiple simultaneous requests and said they are still investigating the issue.", "body_md": "Oh, you wanna see my m~~e~~ath lab?\n\nEpyc Rome 7532\n\n128GB DDR4\n\nVarious NVMe totaling ~5TB of storage\n\n4x AMD V620\n\nAll in an open air crypto-mining bench. No, this does not trip breakers, but it comes close.\n\nCurrent stack is a LiteLLM frontend to vLLM (I’ll do a write-up on how I optimized vLLM soon), serving Qwen3.8-flash-next at int4 quant, with the n-gram table offloaded to system memory. Gets me ~70TPS of single-request inference and ~90TPS peak for 2-3 simultaneous streams. This is without MTP, because MTP is not reliably performant with multiple simultaneous requests being served. Still trying to figure that one out.", "url": "https://wpnews.pro/news/post-show-off-your-ai-rig-whatchya-got-in-there", "canonical_source": "https://forum.level1techs.com/t/post-show-off-your-ai-rig-whatchya-got-in-there/255709#post_3", "published_at": "2026-09-10 23:44:48+00:00", "updated_at": "2026-09-10 23:47:21.442388+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-tools"], "entities": ["AMD Epyc Rome 7532", "AMD V620", "LiteLLM", "vLLM", "Qwen3.8-flash-next"], "alternates": {"html": "https://wpnews.pro/news/post-show-off-your-ai-rig-whatchya-got-in-there", "markdown": "https://wpnews.pro/news/post-show-off-your-ai-rig-whatchya-got-in-there.md", "text": "https://wpnews.pro/news/post-show-off-your-ai-rig-whatchya-got-in-there.txt", "jsonld": "https://wpnews.pro/news/post-show-off-your-ai-rig-whatchya-got-in-there.jsonld"}}