A nightly prompt eval that collapses every failure into one boolean will blame the model for a cut-off response. The useful split is smaller than a full scoring platform and stricter than a single green check. You should treat transport loss, truncation, refusal, malformed JSON, and a wrong answer as five different verdicts. Only a wrong answer, among those verdicts, should be allowed to move a pinned content baseline.
When a completion stops at a token limit, the tail of a JSON object disappears and the parser throws. A boolean grader records that throw as a miss, then a diff against last week looks like a reasoning regression. The same distortion appears when the connection drops, the server returns an empty body, or the model refuses a case it previously answered. Those events are operational facts, and they should not wear the label of a content failure.
This workflow assumes you already store each raw completion beside its case id for a later replay. Free model access and a free server option matter because the nightly job is repetitive, not because a free lane is an oracle. Disclosure: This article was prepared as part of MonkeyCode's product outreach. If those two options are available, use them for the scheduled completion calls, and do not treat that availability as a measured quality bar.
The following grader is a proposal, not a log from a production fleet, and this draft has not executed it. It reads a manifest of golden cases, classifies each saved completion, and writes one typed verdict per case. Expected answers stay in the case file, while the grader returns a class before any numeric score exists. That order is the whole point, because a score without a class invites the wrong comparison later.
from __future__ import annotations
import hashlib
import json
from dataclasses import dataclass
from enum import Enum
from typing import Any
class Verdict(str, Enum):
OK = "ok"
WRONG = "wrong"
MALFORMED = "malformed"
REFUSED = "refused"
TRUNCATED = "truncated"
TRANSPORT = "transport"
UNKNOWN = "unknown"
CONTENT = {Verdict.OK, Verdict.WRONG}
@dataclass(frozen=True)
class Case:
case_id: str
prompt: str
expect: dict[str, Any]
@dataclass(frozen=True)
class Completion:
text: str
finish: str
error: str | None = None
def classify(case: Case, done: Completion) -> Verdict:
if done.error or done.finish == "error":
return Verdict.TRANSPORT
if done.finish == "length":
return Verdict.TRUNCATED
if done.finish == "refusal":
return Verdict.REFUSED
if done.finish != "stop":
return Verdict.UNKNOWN
try:
payload = json.loads(done.text)
except json.JSONDecodeError:
return Verdict.MALFORMED
if not isinstance(payload, dict):
return Verdict.MALFORMED
for key, expected in case.expect.items():
if payload.get(key) != expected:
return Verdict.WRONG
return Verdict.OK
def manifest_sha256(cases: list[Case]) -> str:
blob = json.dumps(
[{"id": c.case_id, "prompt": c.prompt, "expect": c.expect} for c in cases],
sort_keys=True,
separators=(",", ":"),
)
return hashlib.sha256(blob.encode()).hexdigest()
def content_rate(rows: list[Verdict]) -> float | None:
judged = [row for row in rows if row in CONTENT]
if not judged:
return None
hits = sum(row is Verdict.OK for row in judged)
return hits / len(judged)
def content_diff_allowed(rows: list[Verdict], min_judged: int = 8) -> bool:
judged = [row for row in rows if row in CONTENT]
if len(judged) < min_judged:
return False
blocked = len(rows) - len(judged)
return blocked * 5 <= len(rows)
def summarize(cases: list[Case], rows: list[Verdict]) -> dict[str, Any]:
counts = {item.value: sum(row is item for row in rows) for item in Verdict}
return {
"manifest_sha256": manifest_sha256(cases),
"counts": counts,
"content_rate": content_rate(rows),
"content_diff_allowed": content_diff_allowed(rows),
}
def diff_content(previous: dict[str, Any], current: dict[str, Any]) -> str:
if previous["manifest_sha256"] != current["manifest_sha256"]:
raise SystemExit("manifest hash changed; refuse the content diff")
if not current["content_diff_allowed"]:
raise SystemExit("operational mix too high; content diff suppressed")
prev = previous["content_rate"]
cur = current["content_rate"]
if prev is None or cur is None:
raise SystemExit("no judged content rows; nothing to compare")
return f"content_rate {prev:.3f} -> {cur:.3f}"
def demo() -> None:
expect = {"priority": "p2", "queue": "billing"}
prompt = "Return JSON with priority and queue."
cases = [
Case("ticket-14", prompt, expect),
Case("ticket-15", prompt, expect),
Case("ticket-16", prompt, expect),
]
saved = {
"ticket-14": Completion('{"priority": "p2"', "length"),
"ticket-15": Completion(
'{"priority": "p1", "queue": "billing"}', "stop"
),
"ticket-16": Completion("", "refusal"),
}
rows = [classify(case, saved[case.case_id]) for case in cases]
print(json.dumps(summarize(cases, rows), indent=2))
if __name__ == "__main__":
demo()
A content rate that ignores truncated rows will look healthier than the raw pass rate on the same run. That healthier figure is acceptable only when the same record also publishes the blocked count beside it. The helper refuses a content diff when fewer than eight cases were judged, or when non-content verdicts exceed one fifth of the run. Those thresholds are policy choices in this sketch, not measured optima, so retune them against your own manifest size.
The demo at the bottom of the module is also unexecuted in this draft, and its expected result follows from reading the branches. By inspection, one row is truncated, one is refused, and one is wrong, so the judged set contains only the wrong answer. The content rate on that judged set is zero, and the diff flag stays false because three rows cannot satisfy a minimum of eight. That closed flag is the desired outcome of the demo, and it is not a measurement of any hosted model.
A postal scale that records a missing reading as zero kilograms will announce a famine every time the truck is late. The typed verdict is the practical difference between an empty tray and a tray that never reached the scale. A support-ticket case that expects priority p2 and queue billing is enough structure to show that split without a large fixture. When finish equals length and the object is only half written, the classifier returns truncated and that row stays outside the content rate.
A later clean completion that returns priority p1 for the same queue is wrong, and that row does enter the rate. Those two outcomes must not share a cell, or the weekly chart will narrate a quality drop that was a cut-off. Keep the raw text, the finish flag, and the case id in the incident file so a reviewer can confirm the class. Rerunning immediately would mix a new sample into the evidence, which is a different question from how this sample was graded.
Before any nightly model call, lock the verdict taxonomy with saved completions that never leave the test process. A three-line fixture for truncated, refused, and wrong is enough to prove the classes do not collapse under a parser refactor. The proposed test below is unexecuted here, and it should fail if a later edit maps length onto wrong. Once that test is green, the nightly runner may call a model, but it must reuse the same classify function the test already pinned.
from verdicts import Case, Completion, Verdict, classify
def test_length_is_not_a_content_miss() -> None:
case = Case(
"ticket-14",
"Return JSON with priority and queue.",
{"priority": "p2", "queue": "billing"},
)
assert classify(case, Completion('{"priority": "p2"', "length")) is Verdict.TRUNCATED
assert (
classify(case, Completion('{"priority": "p1", "queue": "billing"}', "stop"))
is Verdict.WRONG
)
assert classify(case, Completion("", "refusal")) is Verdict.REFUSED
assert classify(case, Completion("not-json", "stop")) is Verdict.MALFORMED
assert classify(case, Completion("", "stop", error="reset")) is Verdict.TRANSPORT
[
{
"id": "ticket-14",
"prompt": "Return JSON with priority and queue.",
"expect": {"priority": "p2", "queue": "billing"}
}
]
python -m venv .venv
. .venv/bin/activate
python -m pip install pytest
python -m pytest tests/test_verdict_class.py -q
python verdicts.py
From a clean virtual environment, install only the test runner, then execute the taxonomy tests before you wire any client. The commands below match the two files in this draft, and they do not call a hosted model. A later wrapper around summarize and diff_content should refuse to start when the manifest hash cannot be computed. That wrapper should also exit non-zero when the hashes differ or the newer content-diff flag is false.
The run record is one JSON object with a manifest hash, class counts, a content rate, and a flag for the content diff. The hash binds expected JSON to the run, so a quiet edit to a golden case cannot look like a model change. Class counts sit beside the rate because a high rate can still hide a night when one fifth of the rows never finished. If the flag is false, keep the previous content series and store the new rows as an incident beside that series.
A refusal is closer to a policy outcome than to a wrong field, and folding it into the content rate punishes a declined answer. Transport errors sit on the path between processes, and they should not be narrated as text the model actually produced. Malformed JSON after a clean stop is the ambiguous middle, because the model finished and still broke the contract you requested. This sketch files that middle as malformed rather than wrong, so a format break stays visible when you later tighten the prompt.
This sketch assumes the provider tells you why the completion ended, and many clients simply never send that flag. If finish is missing, you cannot honestly separate truncation from malformed text, and you should label the row unknown rather than guess. Exact JSON equality is a narrow grader, because free-text drift will look wrong even when a human would accept it. Do not use this harness for open-ended essays, legal advice, or any case where partial credit is the real product requirement.
A free model lane can change how often responses are cut off without changing the answers that finish. That shift in mix is why the blocked share must gate the content diff rather than the incident log. That gate is not a benchmark of the vendor, and this draft records no latency, cost, or accuracy numbers. If your client cannot return a finish reason, or your cases are prose rather than a small JSON contract, skip this pattern.
If free model access and a free server are already on your MonkeyCode account, that pair can carry this nightly job. Keep the grading script in your own repo, and read the finish flag before you read any score. Nothing in this note claims a quota, a model name, or a permanent offer, because those facts were not supplied here as primary sources. The comparison worth keeping is the content rate beside the class counts, for one pinned manifest, on one pinned endpoint.