Study OW: The open-weight backfill: the cheapest model passed everything; two frontiers failed the same gate A study backfill found that open-weight model gpt-oss-120b passed all 13 guardrail gates for $0.33 in API spend, while frontier models moonshotai/kimi-k3 and anthropic/claude-fable-5 both failed the same gate (goal-safe eviction at the 20-note memo cap edge). The total spend was $53.39, and the results license a map for future tier swaps, not immediate ship decisions. Study OW · Confirmations & machinery The open-weight backfill: the cheapest model passed everything; two frontiers failed the same gate Open-weight and frontier tier maps Study Overview Why a Backfill The series had measured three closed-weight models throughout, while open-weight frontier models became genuinely competitive. This backfill maps moonshotai/kimi-k3 open-weight frontier , openai/gpt-oss-120b open-weight non-frontier , and anthropic/claude-fable-5 the closed frontier tier above the shipped one against two standing instruments: the 13-gate regression suite and the tag-steering family. Registered as a tier-map extension: no new hypotheses, no new gates, prompts byte-identical, and all registered verdicts pinned to the original trio. One disclosure carried in the brief: the analyst model is among those measured, so every grader used is mechanical and predates the runs; the judge-graded Track 2 studies are excluded. The Cheapest Model Passed Everything gpt-oss-120b cleared all thirteen shipped-guardrail gates, anchored patches through memo mechanics, for 33 cents of API spend. The guardrail stack context handed to the model, tools over rules, server-owned invariants was designed not to depend on tier; the cheapest model in the series' history just delivered the strongest confirmation of that design yet. Two Frontiers, One Red Gate kimi-k3 8/10 and fable-5 7/10 both went red on a single slice, goal-safe eviction at the 20-note memo cap edge, with the exact client-prune anatomy Study AK registered as the app's honest boundary: each sends a list already trimmed to the cap, so the app's goal-preserving eviction never engages, and the victim is an old goal, the one thing only the memo carries. The failure is cap-obedience. The aging shipped tier passes this slice by over-sending and letting the app decide; the two newest frontier-class models politely self-edit and choose victims badly, while gpt-oss-120b passes it 10/10. Per the standing procedure, neither red-gated model ships onto the chat surfaces as-is, and the filed Study AL powered re-measurement the client-prune prompt fence, verdict "unproven, not disproven" now has the motivation it lacked: the likely next tiers exhibit exactly the pathway it targets. The Tag Family Tier Map All the family's universals replicated: bare-arm drift 0/36 on every model, unprompted tags list reads in every shipped-arm cell, zero invented tags anywhere. kimi-k3 posted the best adjacent-class numbers measured to date, 12/12 in all four AP arms under the AP′ independent key. fable-5 profiles as the shipped tier's successor, with the same refusal to tag off-catalog topics and the same one-sentence fix, but it mints freely where opus abstains. gpt-oss-120b is the cautionary shape: perfect catalog obedience, yet it never returns an empty list unaided, the guidance sentences only partially rescue it, and its warnings-only arm collapsed to 1/36, the starkest the-tool-is-the-mechanism datapoint in the family. What It Cost and What It Licenses Total spend was $53.39 against a registered fence of $25 to $45; the overage came from underestimating fable-5's regression input volume and is disclosed in the report. Phase 2 deeper family re-runs on the new models stays unregistered and unspent. What this backfill licenses is a map, not a ship decision: the guardrails hold at every price point measured, the two newest frontier models share one specific, well-understood exposure at the memo cap edge, and any future tier swap starts from these gate summaries.