Localmaxxing – Local LLM Inference Benchmarks LocalMaxxing, a community-driven platform for local large language model inference benchmarks, allows users to submit real hardware performance data via CLI or web form. Results appear immediately on public leaderboards, tracking tokens per second, time-to-first-token, and VRAM usage to compare models, hardware, and engines. Sign in & create a key Sign in with GitHub, then create an API key in your dashboard so the CLI and your agents can submit runs. LocalMaxxing Community benchmarks for local LLM inference. Track speed, compare hardware, and find your optimal setup. Every number on this site comes from a community-submitted run on real hardware — no vendor benchmarks. Sign in with GitHub, then create an API key in your dashboard so the CLI and your agents can submit runs. Measure tokens/sec, time-to-first-token and VRAM usage with the CLI, or submit results through the web form. Results appear on the public leaderboards immediately — compare models, hardware and engines to find your optimal setup.