cd /news/ai-tools/judge-cheap-audit-confidence-a-ci-ga… · home › topics › ai-tools › article
[ARTICLE · art-143039] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

Judge cheap, audit confidence: a CI gate for LLM evals (open source)

A developer released laya-evals, an open-source CI gate for LLM evaluation that combines rubric-based LLM-as-judge scoring, calibration auditing (Expected Calibration Error, Brier score, reliability bins), and a GitHub Actions regression gate that fails builds on measured accuracy or calibration regressions. On 872 SST-2 sentences with a local laya judge, it reported 0.8968 judge accuracy, Cohen's kappa 0.7938, ECE 0.0253, Brier 0.0779, and 6.4 decisions per second at $0 API cost, with per-question-shape confidence thresholds (0.865 for 2-option SST-2, 0.994 for 20-option MASSIVE).

by read2 min views1 publishedOct 1, 2026

Your LLM passes the demo every time. Then a customer sends a prompt you didn't test, and the answer is garbage. The gap between "it worked in the eval" and "it works in production" is almost always the same: you measured accuracy, but you never measured whether the model's confidence was trustworthy.

laya-evals closes that gap. It is the fifth tool in my open-source series on the laya decision engine, and the one that keeps the other four honest.

Judge accuracy: 0.8968 (Cohen's kappa 0.7938)
Calibration: ECE 0.0253, Brier 0.0779
Cost per 1,000 judgments: $0 (local inference)
CI gate: fails the build when accuracy or calibration regresses

Three things, in one pipeline:

1. Rubric-based LLM-as-judge. Score your eval sets with a local laya decision model instead of paying for GPT-4 as a judge. Choice questions, score questions, or no-ground-truth mode. 6.4 decisions per second on CPU.

2. Calibration audit. Accuracy tells you how often the model is right. Calibration tells you whether you can trust its confidence. laya-evals computes Expected Calibration Error, Brier score, and reliability bins, then tells you which confidence threshold to use for each question shape.

Why per shape? Because a 2-option question needs a different gate than a 20-option question. The SST-2 benchmark (2 options) needs a threshold of 0.865 for 95% accuracy. The MASSIVE intent benchmark (20 options) needs 0.994. One global min_confidence is not a policy, and laya-evals proves it with your own numbers.

3. CI regression gate. Wire it into GitHub Actions. Exit code 0 means pass, exit code 1 means a measured regression, exit code 2 means the run itself was invalid. Your eval suite becomes a build gate, not a suggestion.

872 SST-2 sentences, 2-option rubric, local laya judge:

Measurement Result
Judge accuracy vs gold 0.8968
Cohen's kappa 0.7938
ECE (15 bins) 0.0253
Brier score 0.0779
Throughput 6.4 decisions/s
API cost $0

Cross-checked against Z.ai's GLM-5.3-flash as reference judge: 95.1% accuracy, 90.7% agreement, kappa 0.814. The local judge is cheaper and nearly as reliable.

I also re-measured four public laya benchmark claims with laya-evals and got values within 0.0003 of the published numbers. The reproduction pack is committed.

pip install laya-evals
python
from laya_evals import Judge, ece, brier_score, advise_thresholds

judge = Judge()
results = judge.run(eval_set, rubric)
print(ece(results.confidences, results.correct))
print(advise_thresholds(results))

This is the fifth and last tool in the laya series, built on the laya decision engine by NandhaKishorM. The same engine that TypeSafe shipped as Jev and OpenAI shipped as the Decisions API on Luna this week. The difference: this one runs on your machine, for free, and the benchmarks are committed.

Every tool ships with its benchmarks, including the numbers that hurt.

── more in #ai-tools 4 stories · sorted by recency
── more on @laya-evals 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/judge-cheap-audit-co…] indexed:0 read:2min 2026-10-01 · —