Practical tools for evaluating, observing, pruning, and quantizing language models. Each tool also includes a portable SKILL.md for coding agents. Pick a folder, install that tool's dependencies, and run it. There is no server or orchestration framework to deploy.
This repository collects the workflows used in 0xSero's REAP and EXL3 experiments, including the GLM-5.3 fidelity campaign and Harbor coding evaluations. It contains reusable Python tools plus small launchers for the original upstream implementations. Model weights, benchmark questions, private session logs, and credentials are not bundled.
| I want to… | Tool | What runs |
|---|---|---|
| Test an agent on terminal tasks | Terminal-Bench | Harbor; explicit 2.1 default, configurable dataset |
| Test long software engineering tasks | DeepSWE | Official DeepSWE task set through Pier |
| Measure graduate-level science reasoning | GPQA Diamond | Included API runner, deterministic choice shuffle, retained responses |
| Evaluate scientific code | SciCode | Official SciCode Inspect evaluator |
| Run reasoning/knowledge/ instruction benchmarks | LM evaluation | EleutherAI task registry: MMLU, GSM8K, IFEval, and more |
| Measure fidelity lost through pruning/quantization | KL divergence | Included full-logit capture and aligned forward-KL comparison |
| Collect expert activation statistics | REAP observation | Pinned Cerebras REAP observer; observation-only by default |
| Select which experts to keep | REAP pruning | Global, domain-balanced, and max-over-domain keep plans |
| Convert a supported model to EXL3 | EXL3 quantization | Pinned ExLlamaV3 converter, recipes and resume support |
| Check calibration inputs and expert coverage | Calibration | Corpus hashes, counts, per-layer/expert coverage gate |
| Check that the intended model really works | Runtime checks | Model-list check, real completion, JSON output, graph-log audit |
| Compare inference throughput fairly | Throughput | Common-window analysis of exact token-emission traces |
| Package reproducible evidence | Evidence | SHA-256 sealing and tamper/missing-file verification |
“Intelligence” is measured through specific tasks, not a single invented composite score. DeepSWE is the concrete software benchmark included here; it is not labeled as a separate benchmark called “Deep Suite.”
Python 3.10+ is required. External tools may require newer Python; the pinned Cerebras REAP checkout requires Python 3.12+.
git clone https://github.com/0xSero/model-toolkit.git
cd model-toolkit
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python toolkit.py list
Most launchers and GPQA use only the Python standard library. NumPy is needed for KL, pruning plans, and route coverage. Install large GPU libraries and benchmark harnesses in separate environments using each folder's README; their dependencies may conflict.
export OPENAI_API_KEY='your-local-or-provider-key'
python toolkit.py run runtime-checks \
--base-url http://localhost:8000/v1 --model your-served-model \
--output runs/smoke.json
After accepting GPQA's dataset terms and down gpqa_diamond.csv:
python toolkit.py run gpqa-diamond \
--csv data/gpqa_diamond.csv --model your-served-model \
--base-url http://localhost:8000/v1 --limit 2 \
--output runs/gpqa-smoke
Remove --limit for the full 198-question set. Configure sampling, reasoning effort, and output budget explicitly for your comparison. Request failures withhold the aggregate accuracy; output truncations are reported separately.
python toolkit.py run terminal-bench \
--model openai/your-served-model --jobs-dir runs/tb21 \
--limit 1 --dry-run
python toolkit.py run exl3-quantization \
--input models/source --output models/exl3-4bpw \
--work runs/exl3-work --bits 4 --dry-run
The previews do not require the external harness or a GPU. Actual runs do. Terminal-Bench and DeepSWE require a container environment. SciCode executes generated code: run its harness inside a disposable evaluation environment.
Pinned model + calibration corpus
↓
REAP observations → coverage checks → expert keep plan
↓ ↓
upstream pruning / architecture-specific weight surgery
↓
EXL3 conversion → runtime checks → KL + task benchmarks
↓
sealed evidence
Observation formats, expert tensor layouts, and quantization kernels vary by architecture. The tool READMEs state the supported boundary. A keep plan does not modify weights. A completed conversion does not prove runtime support or retained model quality.
For KL, compare the quantized model against pruned BF16 to isolate quantization loss, and separately against original BF16 to measure total pruning-plus-quantization loss. Keep sampling, prompts, tokenization, and task sets fixed when comparing task scores.
- GLM-5.3 REAP fidelity study : criterion sweeps, keep plans, KL reports, and GPQA artifacts.
- REAP calibration data : calibration inputs and collection methodology. Audit constituent sources before claiming a benchmark-clean evaluation.
- Qwen layerwise observations andGLM layerwise observations : prior observation artifacts.
- GLM-5.3 observations : referenced by the fidelity study, butnot anonymously accessible when checked on 2026-09-10 . No tool here requires access to it.
- Cerebras REAP andPR #17 : underlying observation and pruning work.
See REAP observation for the executable path and provenance for the connection to earlier sessions and skills.
upstreams.json records exact upstream commits. python toolkit.py fetch NAME checks out one of them under .external/; it does not install dependencies. Record harness/package versions, model and dataset revisions, hardware, context, sampling, timeouts, and all failed attempts with your run.
The included tools are tested with synthetic data, mocked API responses, and CLI checks. This release does not claim a fresh GPU conversion, full benchmark, or architecture compatibility test. See validation. Historical model scores are deliberately not presented as results of this release.
Run the local checks:
python -m unittest discover -s tests -v
MIT for this toolkit's code; imported MIT notices are retained in licenses/. External harnesses, datasets, models, and model-specific code retain their own licenses. REAP belongs to Cerebras Research; ExLlamaV3/EXL3 to turboderp and contributors; Terminal-Bench/Harbor to their authors; DeepSWE/Pier to Datacurve; SciCode to its authors; lm-evaluation-harness to EleutherAI and contributors. See provenance.