cd /news/ai-infrastructure/crucible-llm-bench-your-llm-and-serv… · home › topics › ai-infrastructure › article
[ARTICLE · art-143439] src=forum.level1techs.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Crucible LLM :: Bench your LLM and server. Is it fast enough… for everyone?

Developer MadGoatHaz released Crucible LLM, a free, GPL-3.0 terminal-based benchmarking suite written in Rust that stress-tests any OpenAI-compatible inference server and reports per-user throughput rather than aggregate tokens per second. Crucible LLM ships as a roughly 14 MB static binary with zero runtime dependencies and SQLite compiled in, and it illustrates the gap it targets with the example of a server serving 64 users at 2.6 tokens per second each. The tool is available as source via git clone or as pre-compiled binaries from its GitHub Releases page.

read1 min views1 publishedOct 1, 2026

I built Crucible LLM — a terminal-based benchmarking suite that stress-tests any OpenAI-compatible inference server and measures what a real user actually experiences. One static Rust binary. No Python. No venv. No dependencies.

Most benchmarks shout “xx.x tokens/sec!” and stop. But when 32 users hit that server at once, each one gets 2 tokens/sec. A number that looks fast in the aggregate can feel painfully slow per user. Crucible measures that gap — the difference between a server’s peak throughput and what each individual connection actually gets as well as how competent the LLM is at tasks.

/v1 endpoint and go. Live Monitor screenshot — real-time throughput graph with auto-scaling y-axis, engine-transition markers, and live event log.

Build from source (requires a stable Rust toolchain):

git clone https://github.com/MadGoatHaz/crucible-llm
cd crucible-llm
cargo build --release
./target/release/crucible-llm

Or grab a pre-compiled static binary from the Releases page. On first run, the TUI walks you through entering your server URL and picking a model.

A few examples of the kind of insight a run produces:

A server that serves 64 users at 2.6 t/s each is not “fast.” It is slow for everyone. Crucible reports the difference.

Crucible is free and open source (GPL-3.0). If you run inference servers, benchmark models, or just want to know what your hardware can actually do, I’d love your feedback, issues, and PRs.

github.com/MadGoatHaz/crucible-llm — star the repo, open an issue, or fork and contribute.

Technical: Rust · GPL-3.0 · ~14 MB static binary · zero runtime dependencies (SQLite compiled in) · GPU telemetry feature-gated (NVIDIA NVML built-in)

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @crucible llm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/crucible-llm-bench-y…] indexed:0 read:1min 2026-10-01 · —