{"slug": "show-hn-rigmark-benchmarks-local-ai-the-way-coding-agents-use-it", "title": "Show HN: RigMark benchmarks local AI the way coding agents use it", "summary": "RigMark released a reproducible serving benchmark that measures local AI stacks on agent-shaped code and prose workloads, structured output, cache-busting prefill and concurrent requests rather than a single headline number. On Qwen3.8-27B-FP8-vllm, protocol 1.3 recorded single-stream medians of 84.2 tok/s for prose, 129.5 tok/s for code and 136.0 tok/s for structured JSON, with prefill first-token throughput falling from 8,173 tok/s at 8K depth to 5,808 tok/s at 64K. The benchmark is model-agnostic across Qwen, GLM, DeepSeek, vLLM, SGLang and multi-node DGX Spark, and its author notes identical weights and software are required before claiming a topology-only speed-up.", "body_md": "**RigMark benchmarks local AI how coding agents actually use it.**\n\nMost AI benchmarks reduce a serving stack to one flattering number. RigMark is a reproducible stress test for the whole OpenAI-compatible appliance: agent-shaped code and prose, an explicitly labelled structured-output ceiling, exact cache-busting/immediate-replay prefill, and concurrent request behaviour. It is a serving benchmark, not a claim that three prompts reproduce a complete multi-turn coding session.\n\nIt is deliberately model-agnostic: use it with Qwen, GLM, DeepSeek, vLLM, SGLang, local GPU servers, or multi-node DGX Spark recipes. A different model or serving stack is an appliance comparison; identical weights and software are required before claiming a topology-only speed-up.\n\n[Superterm](https://superterm.dev/), from RigMark's author, gives your coding\nagents a home in the browser. Use OpenCode with your local models, keep track\nof which sessions need attention, and check progress or send a follow-up from\nyour phone.\n\n- **Community Edition:** free for personal, non-commercial use on one device,\nwith session attention, logbook, and mobile access.\n- **Pro:** for daily and professional use, with access across your devices,\nmulti-line input, image paste, voice dictation, and email support.\n\n[Get Community Edition or explore Pro](https://superterm.dev/pricing/).\n\nThe current protocol (1.3) allows 8,192 generated tokens per decode sample, including reasoning, with five samples per workload. Shorter answers stop normally. Prefill and the 256-token concurrency cap are unchanged. The archived receipt below retains its original protocol and 4,096-token default; old and new caps must not be presented as a matched comparison.\n\n```\n╭────────────────────────────────────────────────────────────────────────╮\n│                                                                        │\n│  R I G M A R K   //   AGENT WORKLOAD RECEIPT                           │\n│  BENCHMARKS LOCAL AI HOW CODING AGENTS ACTUALLY USE IT                 │\n│  ●  15/15 BASIC OUTPUT GATES PASSED                                    │\n│                                                                        │\n├─ MODEL ────────────────────────────────────────────────────────────────┤\n│                                                                        │\n│  Qwen3.8-27B-FP8-vllm                                                  │\n│                                                                        │\n├─ SINGLE STREAM ────────────────────────────────────────────────────────┤\n│                                                                        │\n│  WORKLOAD         tok/s       Range tok/s   Last (s)    Checks         │\n│  PROSE             84.2        81.9–100.1       19.9     ✓ 5/5         │\n│  CODE             129.5       124.0–130.4       25.2     ✓ 5/5         │\n│  STRUCTURED*      136.0       135.0–136.1        7.9     ✓ 5/5         │\n│                                                                        │\n│  * Predictable JSON ceiling; not general agent performance.            │\n│  Decode medians are estimates; rates include streamed reasoning.       │\n│                                                                        │\n├─ PREFILL ──────────────────────────────────────────────────────────────┤\n│                                                                        │\n│  DEPTH      First tok/s     First TTFT (s)    Replay TTFT (s)          │\n│  8K               8,173               1.00               0.25          │\n│  32K              6,968               4.70               0.44          │\n│  64K              5,808              11.28               0.77          │\n│  Cold cache UNVERIFIED: cache-hit usage unavailable.                   │\n│  Medians; replay is an immediate repeat, not a proven hit.             │\n│                                                                        │\n├─ CAPPED CONCURRENT GENERATION ─────────────────────────────────────────┤\n│                                                                        │\n│  Aggregate tok/s: C1 108.7 | C2 206.0 | C4 385.4                       │\n│  C4: 0/12 normal stops; 2/12 with visible output.                      │\n│  Capped throughput includes reasoning; not completed agent tasks.      │\n│                                                                        │\n├─ APPLIANCE ────────────────────────────────────────────────────────────┤\n│                                                                        │\n│  Hardware: 1x NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB │\n│  Topology: TP1, local PCIe GPU                                         │\n│  Checkpoint: Qwen/Qwen3.8-27B-FP8 @ local checkp                       │\n│  Quantisation: FP8 E4M3, dynamic activations, 128x128 blocks           │\n│  KV cache: FP8                                                         │\n│  Engine: vLLM 0.27.1                                                   │\n│                                                                        │\n├─ SETTINGS ─────────────────────────────────────────────────────────────┤\n│                                                                        │\n│  SUITE: DEFAULT SETTINGS                                               │\n│  REQUEST: {\"chat_template_kwargs\": {\"reasoning_effort\": \"low\"}}        │\n│  Temperature 0.0 | top_p 1.0 | seed 20260905 | protocol 1.0.0          │\n│  Decode: 5 runs, 4096-token cap | Prefill: 8K/32K/64K, 3 pairs         │\n│  Concurrency: C1/C2/C4, 3 rounds, 256-token cap; workload code         │\n│  Comparison ID: 2026-09-05-rigmark-standard-v2                         │\n│                                                                        │\n├─ RECEIPT ──────────────────────────────────────────────────────────────┤\n│                                                                        │\n│  SOURCE     git:046e92cbe941  •  clean                                 │\n│  JSON sha256:604ea2c48107a69f…                                         │\n│  SHARE THE CARD • LINK THE JSON RECEIPT • #RIGMARK                     │\n│  github.com/alexellis/rigmark                                          │\n│                                                                        │\n╰────────────────────────────────────────────────────────────────────────╯\n```\n\nThis card is rendered from a published protocol 1.0 receipt. The current runner emits protocol 1.2 receipts with per-pair cache isolation and cache-usage evidence. This older receipt has no cache-hit evidence, so its prefill is explicitly labelled unverified by the current renderer.\n\nThe card is the shareable headline. It always shows the benchmark Git revision\nand whether that worktree was clean or dirty. GitHub source archives embed the\noriginating commit too, so downloading a tarball does not lose the benchmark\nrevision. The content-hashed\n[result JSON](https://github.com/alexellis/rigmark/blob/master/results/reference/qwen38-27b-fp8-rtxpro6000-low.json) is the\nreceipt: complete outputs, ranges, TTFT, settings, and appliance metadata.\nThe checksum identifies those exact bytes; it is not a signature or independent\nattestation. A linked Git commit or release supplies the public anchor.\n\nPython 3.10 or newer is required; there are no third-party packages.\n\n```\ngit clone https://github.com/alexellis/rigmark\ncd rigmark\n./rigmark configure\n\n./rigmark run \\\n  --base-url http://SERVER:8000 \\\n  --model auto \\\n  --label my-appliance \\\n  --comparison-id weekend-sweep-1 \\\n  --metadata metadata.json\n```\n\nThe configurator explains every metadata field and writes the ignored local\n`metadata.json`. See [` METADATA.md`](https://github.com/alexellis/rigmark/blob/master/METADATA.md) if an agent is filling it in\nfor you. The benchmark refuses unchanged placeholders or missing required\nfields. Use the same comparison ID for every appliance in one A/B sweep. This\nmakes the corresponding generated prompts byte-for-byte identical. Keep the ID\nwhen repeating the same inputs; changing it changes the prompt nonces. Fresh\nper-pair cache salts isolate prefill independently of that ID. If the model supports graded effort or a thinking toggle, set it\nexplicitly with `--extra-body` and use the identical value throughout the\nsweep; model defaults are not assumed equivalent.\n\nThe repository also includes clearly labelled\n[first-party serving records](https://github.com/alexellis/rigmark/blob/master/examples/serving-records) for the Qwen 27B,\nQwen Flash Next, DeepSeek V4 Flash 0731, and GLM-5.3-Flash appliances operated\nby Alex Ellis. These are\nversioned metadata inputs, not benchmark results or defaults for other users.\nThey pin details that are easy to misreport—especially the drafter, image,\nrecipe revision, scheduler, context limit, and co-resident services. Copy one\nonly when reproducing that appliance, then change every field that differs.\n\nThe default suite performs:\n\n- five runs of code, prose, and structured JSON with a 4,096-token ceiling;\n- three cold/immediate-warm pairs at 8,192, 32,768, and 65,536 tokens; and\n- three rounds of code at concurrency 1, 2, and 4 with 256-token outputs.\n\nThe concurrency phase is deliberately a capped generation-capacity test. It reports normal-stop and visible-output counts and must not be described as completed multi-agent coding work; reasoning models may spend the entire cap on reasoning.\n\nResults are written to `results/LABEL-TIMESTAMP.json`. Neither the endpoint URL\nnor API credentials are written to the result. Credentials are read from\n`OPENAI_API_KEY` by default; use `--api-key-env NAME` to select another\nenvironment variable. Visible generated output is retained for auditability;\nreasoning text is not retained, although its size and hash are recorded.\nEvery decode workload has a basic output gate. Code and prose must emit a\nvisible answer and finish normally; structured JSON must additionally match\nevery requested value. This is not a code-correctness score. Throughput from a\nfailed gate remains diagnostic but must not be cited as a successful workload\nresult. RigMark also reports time to the last generated output, because tok/s\nalone can conceal how long a verbose or reasoning-heavy answer takes.\n\nThe card retains medians and min/max ranges. Protocol 1.2 JSON also includes\n`pooled_decode_tokens_per_second`, calculated from summed decode-token\nnumerators and decode windows. It supplements the median; longer answers carry\nmore weight. Neither ranges nor pooled rates are confidence intervals.\n\nPrefill sends a fresh `cache_salt` for every cold/replay pair. Known cache hits\non a cold sample fail the run. If the server omits cache usage, the card marks\nthat measurement unverified. Skipped phases remain visible as an incomplete\nsuite. Custom suite flags and the full extra request body appear below the\nresults, alongside wrapped appliance details.\n\nAt the end, the runner prints a terminal result card designed to be\nscreenshotted and saves a stable `RESULT.card.txt` beside the JSON. Reprint or\nregenerate it at any time with:\n\n```\n./rigmark report --save results/YOUR-RESULT.json\n```\n\nShare the card with the complete JSON—the card is the headline, and the JSON is the receipt.\n\nThe code answer deliberately includes an implementation and its own test suite. Replay those model-supplied tests in a locked-down Docker container with:\n\n```\ndocker pull golang:1.25\n./rigmark audit-code results/YOUR-RESULT.json\n```\n\nGenerated code is untrusted. The command disables networking, drops capabilities, makes the container root filesystem and source mount read-only, and applies CPU, memory, and process limits. A disposable VM remains the stronger isolation boundary. Passing model-written tests is useful evidence of self-consistency, not proof of correctness.\n\nRun the standard suite unchanged against a quiet endpoint and publish its card and JSON receipt. That is the whole challenge: agent-shaped code and prose, structured output labelled as a ceiling, cache-busting and immediate-replay prefill, and concurrent service load under one disclosed protocol. A recipe may optimise anything on the server side; RigMark keeps the client-side work and claim boundaries fixed.\n\nThere is deliberately no universal `thinking=off` profile. Some templates\nimplement that switch, some ignore it, and others expose different controls.\nPublish such a run as a separately labelled, model-specific ceiling and compare\nit only with receipts carrying the identical request body.\n\nFor an engine without vLLM's `/tokenize` extension or token-ID completion\ninput, add `--skip-prefill`. To omit load testing, add `--skip-concurrency`.\n\n```\n./rigmark compare results/recipe-a.json results/recipe-b.json\n```\n\nFor a matched, screenshot-ready A/B card:\n\n```\n./rigmark compare --card results/recipe-a.json results/recipe-b.json\n```\n\nThe comparison validates raw rows against reported summaries and gates, then\nstops when the benchmark revision, protocol, prompt corpus, seed,\nthinking/request body, run counts, output lengths, prefill depths/runs, or\nconcurrency settings differ. `--allow-mismatch` exists for exploratory\ncomparisons and prints the mismatches. A matched card means the client requests\nmatch; token/s ratios still require matching server-side token definitions.\n\n- Run against a quiet endpoint and disclose any competing traffic.\n- Pin and publish model, drafter, engine/image, and benchmark revisions.\n- Publish the complete result JSON and command, not only a screenshot.\n- Check the retained outputs and code-test audit; throughput alone is not a quality score.\n- Headline code and prose medians with their ranges—not a best run.\n- Label structured output as a speculative-decoding ceiling.\n- Report cache-busting and immediate-replay prefill separately.\n- Verify the negotiated fabric rate rather than copying a product headline.\n\nThe exact measurement definitions and claim boundaries are in\n[`PROTOCOL.md`](https://github.com/alexellis/rigmark/blob/master/PROTOCOL.md).\n\nReference results include their complete generated outputs and appliance\nmetadata, not just headline numbers. See [`RESULTS.md`](https://github.com/alexellis/rigmark/blob/master/RESULTS.md) for the\ncurrent table and exact commands.\n\nIf RigMark helps you tune your setup or publish reproducible recipes,\n[consider sponsoring Alex Ellis](https://github.com/sponsors/alexellis) to\nsupport its maintenance and testing.\n\n```\npython3 -m unittest discover -s tests -v\npython3 -m py_compile audit_code.py bench.py compare.py configure.py receipt.py report.py rigmark\n```\n\nMIT", "url": "https://wpnews.pro/news/show-hn-rigmark-benchmarks-local-ai-the-way-coding-agents-use-it", "canonical_source": "https://github.com/alexellis/rigmark", "published_at": "2026-10-06 14:25:25+00:00", "updated_at": "2026-10-06 14:50:31.769899+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-agents", "mlops", "developer-tools"], "entities": ["RigMark", "Qwen3.8-27B-FP8-vllm", "vLLM", "SGLang", "Superterm", "OpenCode", "Qwen", "DeepSeek"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-rigmark-benchmarks-local-ai-the-way-coding-agents-use-it", "markdown": "https://wpnews.pro/news/show-hn-rigmark-benchmarks-local-ai-the-way-coding-agents-use-it.md", "text": "https://wpnews.pro/news/show-hn-rigmark-benchmarks-local-ai-the-way-coding-agents-use-it.txt", "jsonld": "https://wpnews.pro/news/show-hn-rigmark-benchmarks-local-ai-the-way-coding-agents-use-it.jsonld"}}