{"slug": "hermes-getting-so-good", "title": "Hermes getting so good", "summary": "A Hugging Face forum user reported that the Hermes model, version V0.21, achieves 68 tokens per second on an older NVIDIA GeForce RTX 3090 GPU, with speeds occasionally reaching 84 tokens per second, and mentioned plans to upgrade to an NVIDIA DGX Spark next year. The post highlights the model's performance on consumer hardware, suggesting strong optimization for local inference.", "body_md": "Hugging Face Forums\nHermes getting so good\nBeginners\nMartinPilarski\nSeptember 4, 2026, 9:50pm\n1\n19480\n1920×885 272 KB\nV0.21..\n19481\n1920×885 345 KB\nMy old 3090 68t/s … amazing\n19483\n1844×4000 1.2 MB\nMartinPilarski\nSeptember 6, 2026, 9:24pm\n2\ngoes up to 84 sometime, I already promised new body for my agent\nDGX spark in next year pipeline\nRelated topics\nTopic\nReplies\nViews\nActivity\nPractical match for 128Gb Strix Halo with 2x3090s? (inference for coding)\nBeginners\n4\n891\nMay 21, 2026\nHow many tokens will an old 3090 produce?\nBeginners\n1\n190\nAugust 26, 2026\nPalit Gaming Pro 24GB OC 3090 + Palit X3060 8GB? is it worth bothering even?\nBeginners\n1\n107\nAugust 26, 2026\nHow to add GB10 into NVIDIA hardware list in https://huggingface.co/settings/local-apps\nSite Feedback\n2\n217\nJanuary 27, 2026\nRtx 5090 35b nvfp4\n🤗Transformers\n3\n210\nJuly 18, 2026", "url": "https://wpnews.pro/news/hermes-getting-so-good", "canonical_source": "https://discuss.huggingface.co/t/hermes-getting-so-good/179906#post_2", "published_at": "2026-09-06 21:24:13+00:00", "updated_at": "2026-09-07 02:03:00.617732+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure"], "entities": ["Hugging Face", "Hermes", "NVIDIA GeForce RTX 3090", "NVIDIA DGX Spark"], "alternates": {"html": "https://wpnews.pro/news/hermes-getting-so-good", "markdown": "https://wpnews.pro/news/hermes-getting-so-good.md", "text": "https://wpnews.pro/news/hermes-getting-so-good.txt", "jsonld": "https://wpnews.pro/news/hermes-getting-so-good.jsonld"}}