{"slug": "crucible-llm-bench-your-llm-and-server-is-it-fast-enough-for-everyone", "title": "Crucible LLM :: Bench your LLM and server. Is it fast enough… for everyone?", "summary": "Developer MadGoatHaz released Crucible LLM, a free, GPL-3.0 terminal-based benchmarking suite written in Rust that stress-tests any OpenAI-compatible inference server and reports per-user throughput rather than aggregate tokens per second. Crucible LLM ships as a roughly 14 MB static binary with zero runtime dependencies and SQLite compiled in, and it illustrates the gap it targets with the example of a server serving 64 users at 2.6 tokens per second each. The tool is available as source via git clone or as pre-compiled binaries from its GitHub Releases page.", "body_md": "I built **Crucible LLM** — a terminal-based benchmarking suite that stress-tests any OpenAI-compatible inference server and measures what a *real user* actually experiences. One static Rust binary. No Python. No venv. No dependencies.\n\nMost benchmarks shout “xx.x tokens/sec!” and stop. But when 32 users hit that server at once, each one gets 2 tokens/sec. A number that looks fast in the aggregate can feel painfully slow per user. Crucible measures **that gap** — the difference between a server’s peak throughput and what each individual connection actually gets as well as how competent the LLM is at tasks.\n\n`/v1` endpoint and go.\n*Live Monitor screenshot — real-time throughput graph with auto-scaling y-axis, engine-transition markers, and live event log.*\n\nBuild from source (requires a stable Rust toolchain):\n\n```\ngit clone https://github.com/MadGoatHaz/crucible-llm\ncd crucible-llm\ncargo build --release\n./target/release/crucible-llm\n```\n\nOr grab a pre-compiled static binary from the [Releases](https://github.com/MadGoatHaz/crucible-llm/releases) page. On first run, the TUI walks you through entering your server URL and picking a model.\n\nA few examples of the kind of insight a run produces:\n\nA server that serves 64 users at 2.6 t/s each is not “fast.” It is slow for everyone. Crucible reports the difference.\n\nCrucible is **free and open source** (GPL-3.0). If you run inference servers, benchmark models, or just want to know what your hardware can actually do, I’d love your feedback, issues, and PRs.\n\n **[github.com/MadGoatHaz/crucible-llm](https://github.com/MadGoatHaz/crucible-llm)** — star the repo, open an issue, or fork and contribute.\n\n*Technical: Rust · GPL-3.0 · ~14 MB static binary · zero runtime dependencies (SQLite compiled in) · GPU telemetry feature-gated (NVIDIA NVML built-in)*", "url": "https://wpnews.pro/news/crucible-llm-bench-your-llm-and-server-is-it-fast-enough-for-everyone", "canonical_source": "https://forum.level1techs.com/t/crucible-llm-bench-your-llm-and-server-is-it-fast-enough-for-everyone/257531#post_1", "published_at": "2026-10-01 19:19:04+00:00", "updated_at": "2026-10-01 21:31:36.028119+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "mlops", "developer-tools", "ai-tools"], "entities": ["Crucible LLM", "MadGoatHaz", "GitHub", "Rust", "SQLite", "NVIDIA NVML", "OpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/crucible-llm-bench-your-llm-and-server-is-it-fast-enough-for-everyone", "markdown": "https://wpnews.pro/news/crucible-llm-bench-your-llm-and-server-is-it-fast-enough-for-everyone.md", "text": "https://wpnews.pro/news/crucible-llm-bench-your-llm-and-server-is-it-fast-enough-for-everyone.txt", "jsonld": "https://wpnews.pro/news/crucible-llm-bench-your-llm-and-server-is-it-fast-enough-for-everyone.jsonld"}}