{"slug": "judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source", "title": "Judge cheap, audit confidence: a CI gate for LLM evals (open source)", "summary": "A developer released laya-evals, an open-source CI gate for LLM evaluation that combines rubric-based LLM-as-judge scoring, calibration auditing (Expected Calibration Error, Brier score, reliability bins), and a GitHub Actions regression gate that fails builds on measured accuracy or calibration regressions. On 872 SST-2 sentences with a local laya judge, it reported 0.8968 judge accuracy, Cohen's kappa 0.7938, ECE 0.0253, Brier 0.0779, and 6.4 decisions per second at $0 API cost, with per-question-shape confidence thresholds (0.865 for 2-option SST-2, 0.994 for 20-option MASSIVE).", "body_md": "Your LLM passes the demo every time. Then a customer sends a prompt you didn't test, and the answer is garbage. The gap between \"it worked in the eval\" and \"it works in production\" is almost always the same: you measured accuracy, but you never measured whether the model's confidence was trustworthy.\n\n[laya-evals](https://github.com/Gjusev/laya-evals) closes that gap. It is the fifth tool in my open-source series on the laya decision engine, and the one that keeps the other four honest.\n\n```\nJudge accuracy: 0.8968 (Cohen's kappa 0.7938)\nCalibration: ECE 0.0253, Brier 0.0779\nCost per 1,000 judgments: $0 (local inference)\nCI gate: fails the build when accuracy or calibration regresses\n```\n\nThree things, in one pipeline:\n\n**1. Rubric-based LLM-as-judge.** Score your eval sets with a local laya decision model instead of paying for GPT-4 as a judge. Choice questions, score questions, or no-ground-truth mode. 6.4 decisions per second on CPU.\n\n**2. Calibration audit.** Accuracy tells you how often the model is right. Calibration tells you whether you can trust its confidence. laya-evals computes Expected Calibration Error, Brier score, and reliability bins, then tells you which confidence threshold to use for each question shape.\n\nWhy per shape? Because a 2-option question needs a different gate than a 20-option question. The SST-2 benchmark (2 options) needs a threshold of 0.865 for 95% accuracy. The MASSIVE intent benchmark (20 options) needs 0.994. One global `min_confidence` is not a policy, and laya-evals proves it with your own numbers.\n\n**3. CI regression gate.** Wire it into GitHub Actions. Exit code 0 means pass, exit code 1 means a measured regression, exit code 2 means the run itself was invalid. Your eval suite becomes a build gate, not a suggestion.\n\n872 SST-2 sentences, 2-option rubric, local laya judge:\n\n| Measurement | Result | \n|---|---|\n| Judge accuracy vs gold | 0.8968 | \n| Cohen's kappa | 0.7938 | \n| ECE (15 bins) | 0.0253 | \n| Brier score | 0.0779 | \n| Throughput | 6.4 decisions/s | \n| API cost | $0 | \n\nCross-checked against Z.ai's GLM-5.3-flash as reference judge: 95.1% accuracy, 90.7% agreement, kappa 0.814. The local judge is cheaper and nearly as reliable.\n\nI also re-measured four public laya benchmark claims with laya-evals and got values within 0.0003 of the published numbers. The reproduction pack is committed.\n\n```\npip install laya-evals\npython\nfrom laya_evals import Judge, ece, brier_score, advise_thresholds\n\njudge = Judge()\nresults = judge.run(eval_set, rubric)\nprint(ece(results.confidences, results.correct))\nprint(advise_thresholds(results))\n```\n\nThis is the fifth and last tool in the laya series, built on the [laya decision engine](https://github.com/NandhaKishorM/laya) by NandhaKishorM. The same engine that TypeSafe shipped as Jev and OpenAI shipped as the Decisions API on Luna this week. The difference: this one runs on your machine, for free, and the benchmarks are committed.\n\nEvery tool ships with its benchmarks, including the numbers that hurt.", "url": "https://wpnews.pro/news/judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source", "canonical_source": "https://dev.to/gjusev/judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source-269l", "published_at": "2026-10-01 06:00:00+00:00", "updated_at": "2026-10-01 06:16:39.515880+00:00", "lang": "en", "topics": ["ai-tools", "mlops", "large-language-models", "ai-safety", "developer-tools"], "entities": ["laya-evals", "laya decision engine", "NandhaKishorM", "GitHub Actions", "SST-2", "MASSIVE", "GLM-5.3-flash", "Z.ai"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source", "markdown": "https://wpnews.pro/news/judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source.md", "text": "https://wpnews.pro/news/judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source.txt", "jsonld": "https://wpnews.pro/news/judge-cheap-audit-confidence-a-ci-gate-for-llm-evals-open-source.jsonld"}}