{"slug": "study-ow-the-open-weight-backfill-the-cheapest-model-passed-everything-two-the", "title": "Study OW: The open-weight backfill: the cheapest model passed everything; two frontiers failed the same gate", "summary": "A study backfill found that open-weight model gpt-oss-120b passed all 13 guardrail gates for $0.33 in API spend, while frontier models moonshotai/kimi-k3 and anthropic/claude-fable-5 both failed the same gate (goal-safe eviction at the 20-note memo cap edge). The total spend was $53.39, and the results license a map for future tier swaps, not immediate ship decisions.", "body_md": "Study OW · Confirmations & machinery\n\n# The open-weight backfill: the cheapest model passed everything; two frontiers failed the same gate\n\nOpen-weight and frontier tier maps\n\n## Study Overview\n\n### Why a Backfill\n\nThe series had measured three closed-weight models throughout, while open-weight frontier models became genuinely competitive. This backfill maps moonshotai/kimi-k3 (open-weight frontier), openai/gpt-oss-120b (open-weight non-frontier), and anthropic/claude-fable-5 (the closed frontier tier above the shipped one) against two standing instruments: the 13-gate regression suite and the tag-steering family. Registered as a tier-map extension: no new hypotheses, no new gates, prompts byte-identical, and all registered verdicts pinned to the original trio. One disclosure carried in the brief: the analyst model is among those measured, so every grader used is mechanical and predates the runs; the judge-graded Track 2 studies are excluded.\n\n### The Cheapest Model Passed Everything\n\ngpt-oss-120b cleared all thirteen shipped-guardrail gates, anchored patches through memo mechanics, for 33 cents of API spend. The guardrail stack (context handed to the model, tools over rules, server-owned invariants) was designed not to depend on tier; the cheapest model in the series' history just delivered the strongest confirmation of that design yet.\n\n### Two Frontiers, One Red Gate\n\nkimi-k3 (8/10) and fable-5 (7/10) both went red on a single slice, goal-safe eviction at the 20-note memo cap edge, with the exact client-prune anatomy Study AK registered as the app's honest boundary: each sends a list already trimmed to the cap, so the app's goal-preserving eviction never engages, and the victim is an old goal, the one thing only the memo carries. The failure is cap-obedience. The aging shipped tier passes this slice by over-sending and letting the app decide; the two newest frontier-class models politely self-edit and choose victims badly, while gpt-oss-120b passes it 10/10. Per the standing procedure, neither red-gated model ships onto the chat surfaces as-is, and the filed Study AL powered re-measurement (the client-prune prompt fence, verdict \"unproven, not disproven\") now has the motivation it lacked: the likely next tiers exhibit exactly the pathway it targets.\n\n### The Tag Family Tier Map\n\nAll the family's universals replicated: bare-arm drift 0/36 on every model, unprompted tags_list reads in every shipped-arm cell, zero invented tags anywhere. kimi-k3 posted the best adjacent-class numbers measured to date, 12/12 in all four AP arms under the AP′ independent key. fable-5 profiles as the shipped tier's successor, with the same refusal to tag off-catalog topics and the same one-sentence fix, but it mints freely where opus abstains. gpt-oss-120b is the cautionary shape: perfect catalog obedience, yet it never returns an empty list unaided, the guidance sentences only partially rescue it, and its warnings-only arm collapsed to 1/36, the starkest the-tool-is-the-mechanism datapoint in the family.\n\n### What It Cost and What It Licenses\n\nTotal spend was $53.39 against a registered fence of $25 to $45; the overage came from underestimating fable-5's regression input volume and is disclosed in the report. Phase 2 (deeper family re-runs on the new models) stays unregistered and unspent. What this backfill licenses is a map, not a ship decision: the guardrails hold at every price point measured, the two newest frontier models share one specific, well-understood exposure at the memo cap edge, and any future tier swap starts from these gate summaries.", "url": "https://wpnews.pro/news/study-ow-the-open-weight-backfill-the-cheapest-model-passed-everything-two-the", "canonical_source": "https://www.lightningjar.com/research/barkup-bench/ow", "published_at": "2026-07-29 12:00:00+00:00", "updated_at": "2026-07-29 17:55:03.373804+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "ai-products"], "entities": ["moonshotai/kimi-k3", "openai/gpt-oss-120b", "anthropic/claude-fable-5"], "alternates": {"html": "https://wpnews.pro/news/study-ow-the-open-weight-backfill-the-cheapest-model-passed-everything-two-the", "markdown": "https://wpnews.pro/news/study-ow-the-open-weight-backfill-the-cheapest-model-passed-everything-two-the.md", "text": "https://wpnews.pro/news/study-ow-the-open-weight-backfill-the-cheapest-model-passed-everything-two-the.txt", "jsonld": "https://wpnews.pro/news/study-ow-the-open-weight-backfill-the-cheapest-model-passed-everything-two-the.jsonld"}}