I built Crucible LLM — a terminal-based benchmarking suite that stress-tests any OpenAI-compatible inference server and measures what a real user actually experiences. One static Rust binary. No Python. No venv. No dependencies.
Most benchmarks shout “xx.x tokens/sec!” and stop. But when 32 users hit that server at once, each one gets 2 tokens/sec. A number that looks fast in the aggregate can feel painfully slow per user. Crucible measures that gap — the difference between a server’s peak throughput and what each individual connection actually gets as well as how competent the LLM is at tasks.
/v1 endpoint and go.
Live Monitor screenshot — real-time throughput graph with auto-scaling y-axis, engine-transition markers, and live event log.
Build from source (requires a stable Rust toolchain):
git clone https://github.com/MadGoatHaz/crucible-llm
cd crucible-llm
cargo build --release
./target/release/crucible-llm
Or grab a pre-compiled static binary from the Releases page. On first run, the TUI walks you through entering your server URL and picking a model.
A few examples of the kind of insight a run produces:
A server that serves 64 users at 2.6 t/s each is not “fast.” It is slow for everyone. Crucible reports the difference.
Crucible is free and open source (GPL-3.0). If you run inference servers, benchmark models, or just want to know what your hardware can actually do, I’d love your feedback, issues, and PRs.
github.com/MadGoatHaz/crucible-llm — star the repo, open an issue, or fork and contribute.
Technical: Rust · GPL-3.0 · ~14 MB static binary · zero runtime dependencies (SQLite compiled in) · GPU telemetry feature-gated (NVIDIA NVML built-in)