Judge cheap, audit confidence: a CI gate for LLM evals (open source) A developer released laya-evals, an open-source CI gate for LLM evaluation that combines rubric-based LLM-as-judge scoring, calibration auditing (Expected Calibration Error, Brier score, reliability bins), and a GitHub Actions regression gate that fails builds on measured accuracy or calibration regressions. On 872 SST-2 sentences with a local laya judge, it reported 0.8968 judge accuracy, Cohen's kappa 0.7938, ECE 0.0253, Brier 0.0779, and 6.4 decisions per second at $0 API cost, with per-question-shape confidence thresholds (0.865 for 2-option SST-2, 0.994 for 20-option MASSIVE). Your LLM passes the demo every time. Then a customer sends a prompt you didn't test, and the answer is garbage. The gap between "it worked in the eval" and "it works in production" is almost always the same: you measured accuracy, but you never measured whether the model's confidence was trustworthy. laya-evals https://github.com/Gjusev/laya-evals closes that gap. It is the fifth tool in my open-source series on the laya decision engine, and the one that keeps the other four honest. Judge accuracy: 0.8968 Cohen's kappa 0.7938 Calibration: ECE 0.0253, Brier 0.0779 Cost per 1,000 judgments: $0 local inference CI gate: fails the build when accuracy or calibration regresses Three things, in one pipeline: 1. Rubric-based LLM-as-judge. Score your eval sets with a local laya decision model instead of paying for GPT-4 as a judge. Choice questions, score questions, or no-ground-truth mode. 6.4 decisions per second on CPU. 2. Calibration audit. Accuracy tells you how often the model is right. Calibration tells you whether you can trust its confidence. laya-evals computes Expected Calibration Error, Brier score, and reliability bins, then tells you which confidence threshold to use for each question shape. Why per shape? Because a 2-option question needs a different gate than a 20-option question. The SST-2 benchmark 2 options needs a threshold of 0.865 for 95% accuracy. The MASSIVE intent benchmark 20 options needs 0.994. One global min confidence is not a policy, and laya-evals proves it with your own numbers. 3. CI regression gate. Wire it into GitHub Actions. Exit code 0 means pass, exit code 1 means a measured regression, exit code 2 means the run itself was invalid. Your eval suite becomes a build gate, not a suggestion. 872 SST-2 sentences, 2-option rubric, local laya judge: | Measurement | Result | |---|---| | Judge accuracy vs gold | 0.8968 | | Cohen's kappa | 0.7938 | | ECE 15 bins | 0.0253 | | Brier score | 0.0779 | | Throughput | 6.4 decisions/s | | API cost | $0 | Cross-checked against Z.ai's GLM-5.3-flash as reference judge: 95.1% accuracy, 90.7% agreement, kappa 0.814. The local judge is cheaper and nearly as reliable. I also re-measured four public laya benchmark claims with laya-evals and got values within 0.0003 of the published numbers. The reproduction pack is committed. pip install laya-evals python from laya evals import Judge, ece, brier score, advise thresholds judge = Judge results = judge.run eval set, rubric print ece results.confidences, results.correct print advise thresholds results This is the fifth and last tool in the laya series, built on the laya decision engine https://github.com/NandhaKishorM/laya by NandhaKishorM. The same engine that TypeSafe shipped as Jev and OpenAI shipped as the Decisions API on Luna this week. The difference: this one runs on your machine, for free, and the benchmarks are committed. Every tool ships with its benchmarks, including the numbers that hurt.