# Judge cheap, audit confidence: a CI gate for LLM evals (open source)

> Source: <https://dev.to/gjusev/judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source-269l>
> Published: 2026-10-01 06:00:00+00:00

Your LLM passes the demo every time. Then a customer sends a prompt you didn't test, and the answer is garbage. The gap between "it worked in the eval" and "it works in production" is almost always the same: you measured accuracy, but you never measured whether the model's confidence was trustworthy.

[laya-evals](https://github.com/Gjusev/laya-evals) closes that gap. It is the fifth tool in my open-source series on the laya decision engine, and the one that keeps the other four honest.

```
Judge accuracy: 0.8968 (Cohen's kappa 0.7938)
Calibration: ECE 0.0253, Brier 0.0779
Cost per 1,000 judgments: $0 (local inference)
CI gate: fails the build when accuracy or calibration regresses
```

Three things, in one pipeline:

**1. Rubric-based LLM-as-judge.** Score your eval sets with a local laya decision model instead of paying for GPT-4 as a judge. Choice questions, score questions, or no-ground-truth mode. 6.4 decisions per second on CPU.

**2. Calibration audit.** Accuracy tells you how often the model is right. Calibration tells you whether you can trust its confidence. laya-evals computes Expected Calibration Error, Brier score, and reliability bins, then tells you which confidence threshold to use for each question shape.

Why per shape? Because a 2-option question needs a different gate than a 20-option question. The SST-2 benchmark (2 options) needs a threshold of 0.865 for 95% accuracy. The MASSIVE intent benchmark (20 options) needs 0.994. One global `min_confidence` is not a policy, and laya-evals proves it with your own numbers.

**3. CI regression gate.** Wire it into GitHub Actions. Exit code 0 means pass, exit code 1 means a measured regression, exit code 2 means the run itself was invalid. Your eval suite becomes a build gate, not a suggestion.

872 SST-2 sentences, 2-option rubric, local laya judge:

| Measurement | Result | 
|---|---|
| Judge accuracy vs gold | 0.8968 | 
| Cohen's kappa | 0.7938 | 
| ECE (15 bins) | 0.0253 | 
| Brier score | 0.0779 | 
| Throughput | 6.4 decisions/s | 
| API cost | $0 | 

Cross-checked against Z.ai's GLM-5.3-flash as reference judge: 95.1% accuracy, 90.7% agreement, kappa 0.814. The local judge is cheaper and nearly as reliable.

I also re-measured four public laya benchmark claims with laya-evals and got values within 0.0003 of the published numbers. The reproduction pack is committed.

```
pip install laya-evals
python
from laya_evals import Judge, ece, brier_score, advise_thresholds

judge = Judge()
results = judge.run(eval_set, rubric)
print(ece(results.confidences, results.correct))
print(advise_thresholds(results))
```

This is the fifth and last tool in the laya series, built on the [laya decision engine](https://github.com/NandhaKishorM/laya) by NandhaKishorM. The same engine that TypeSafe shipped as Jev and OpenAI shipped as the Decisions API on Luna this week. The difference: this one runs on your machine, for free, and the benchmarks are committed.

Every tool ships with its benchmarks, including the numbers that hurt.
