{"slug": "so-what-local-inference-models-are-we-using", "title": "So, what local inference models are we using?", "summary": "A user reports running Gemma4:26b on a GPU at 2000+ tokens/s pre-fill and 80+ tokens/s evaluation, and Laguna S 2.1 118B on a CPU server at 13-14 tokens/s, expressing frustration at the lack of models between 30B and 120B parameters that fit in 64GB VRAM. The user calls for a modern MoE model in the 55B-75B range to optimize their hardware.", "body_md": "So, I am not ready to trust Chinese models yet, and have thus been keeping things western.\n\nI have Gemma4:26b running super fast on my GPU (2000+ tokens/s pre-fill, and 80+ tokens/s evaluation output. I generally use it for web assisted research and problem solving.\n\nThen I have a secondary model (Laguna S 2.1 118B) running on my CPU on my server. Normally this would be extremely frustrating, but with 8 channels of DDR4, I get a total memory bandwidth pretty close to the likes of an Nvidia Spark or a Strix 395+.\n\nCPU’s can’t keep up with the matrix math of even a low end GPU, so pre-fill can be a little slow, but eval output actually isn’t bad on the CPU, with some 13-14 tokens/s\n\nI generally use the heavier Laguna logic model to do things like code review or general purpose reasoning where Gemma4 can’t keep up, and I have spare time to run it in the background.\n\nI’d be curious to hear what others may be using.\n\nI am looking for better models for my use case, but right now I feel stuck.\n\nI am finding that the modern 30B-class models - while *excellent* for their size still don’t perform anywhere near as well as a good 120B-class model like Laguna S 2.1\n\nMy main gripe is that there is very little in-between the 30B-class and the 120B-class. The 30B is kind of small, using only ~27GB of my 64GB GPU RAM with 8bit quantization.120B-class is too large to fit in my 64GB of VRAM even at 4 bit quantization, and thus I can’t benefit from GPU acceleration.\n\nI would really love a modern MoE model in the 55B-75B range which I could run at 4-6bit quantization and make the most out of my VRAM with, but there just isnt much there right now.\n\nI did briefly run the old llama 3.3B 70B model, but it is actually out-performed logic wise by the newer generation of 30B-class models. I just can’t help but wonder how good this new generation of models could be if they could have weights in that size class.\n\nI’m hoping someone comes up with a model like that in the not too distant future.", "url": "https://wpnews.pro/news/so-what-local-inference-models-are-we-using", "canonical_source": "https://forum.level1techs.com/t/so-what-local-inference-models-are-we-using/254516#post_1", "published_at": "2026-08-27 16:09:55+00:00", "updated_at": "2026-08-27 16:21:15.389472+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools"], "entities": ["Gemma4:26b", "Laguna S 2.1 118B", "Nvidia Spark", "Strix 395+", "llama 3.3B 70B"], "alternates": {"html": "https://wpnews.pro/news/so-what-local-inference-models-are-we-using", "markdown": "https://wpnews.pro/news/so-what-local-inference-models-are-we-using.md", "text": "https://wpnews.pro/news/so-what-local-inference-models-are-we-using.txt", "jsonld": "https://wpnews.pro/news/so-what-local-inference-models-are-we-using.jsonld"}}