{"slug": "trying-hybrid-inference-gpu-cpu-with-200-400b-models", "title": "Trying hybrid inference (GPU+CPU) with 200-400B models", "summary": "A user running hybrid GPU+CPU inference on 4x Nvidia V100 GPUs reported beating an Nvidia RTX 5090 in token generation (TG) for Qwen 3.8 27B and outperforming two DGX Spark systems with Qwen 3.8 Flash Next, according to a post on X by jkyamog. The report covers hybrid inference experiments with 200-400B models.", "body_md": "Some updates, been running on 4x V100 so far its been beating in TG the 5090 for Qwen 3.8 27B\n\nI have been using it more for Qwen 3.8 Flash Next, beats x2 DGX Spark.\n\nhttps://x.com/jkyamog/status/2096954517984346459?s=20", "url": "https://wpnews.pro/news/trying-hybrid-inference-gpu-cpu-with-200-400b-models", "canonical_source": "https://forum.level1techs.com/t/trying-hybrid-inference-gpu-cpu-with-200-400b-models/247625#post_11", "published_at": "2026-09-15 00:52:37+00:00", "updated_at": "2026-09-15 01:00:39.707953+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-chips"], "entities": ["Nvidia V100", "Nvidia RTX 5090", "Nvidia DGX Spark", "Qwen 3.8 27B", "Qwen 3.8 Flash Next", "jkyamog"], "alternates": {"html": "https://wpnews.pro/news/trying-hybrid-inference-gpu-cpu-with-200-400b-models", "markdown": "https://wpnews.pro/news/trying-hybrid-inference-gpu-cpu-with-200-400b-models.md", "text": "https://wpnews.pro/news/trying-hybrid-inference-gpu-cpu-with-200-400b-models.txt", "jsonld": "https://wpnews.pro/news/trying-hybrid-inference-gpu-cpu-with-200-400b-models.jsonld"}}