Show HN: Gainz.fast – Local Inference, Faster Carsen Klock launched Gainz.fast, a new site for benchmarking local AI inference speed, reporting that the Laguna XS 2.1 model on an AMD R9700 with llama.cpp HIP achieved 143.3 tokens per second, a 31.14% improvement over baseline. The site currently supports Laguna XS and S 2.1 models, with plans to add Qwen3.8 upon release, and Klock is seeking contributors and funding via X (formerly Twitter). Come help push the frontier of token speed across local models and hardware with your agents Current frontier Laguna XS 2.1 · AMD R9700 llama.cpp HIP +31.14% 143.3 tok/s Laguna XS 2.1 · DGX Spark GB10 vLLM NVFP4 +5.28% 37.3 tok/s Laguna XS 2.1 · DGX Spark GB10 llama.cpp CUDA +0.82% 92.6 tok/s Laguna S 2.1 · DGX Spark GB10 vLLM NVFP4 +0.13% 14.1 tok/s Laguna S 2.1 · DGX Spark GB10 llama.cpp CUDA baseline 23.6 tok/s The current Laguna XS and S 2.1 models are supported, I will be adding Qwen3.8 once it is released and possibly some others. If you are interested in helping out with runners or funding, DM me on x.com/carsenklock The site is new, so I am working on further improvements, it is the worst it will ever be, let's make the future faster Comments URL: https://news.ycombinator.com/item?id=49165219 https://news.ycombinator.com/item?id=49165219 Points: 1 Comments: 0