# Study OW: The open-weight backfill: the cheapest model passed everything; two frontiers failed the same gate

> Source: <https://www.lightningjar.com/research/barkup-bench/ow>
> Published: 2026-07-29 12:00:00+00:00

Study OW · Confirmations & machinery

# The open-weight backfill: the cheapest model passed everything; two frontiers failed the same gate

Open-weight and frontier tier maps

## Study Overview

### Why a Backfill

The series had measured three closed-weight models throughout, while open-weight frontier models became genuinely competitive. This backfill maps moonshotai/kimi-k3 (open-weight frontier), openai/gpt-oss-120b (open-weight non-frontier), and anthropic/claude-fable-5 (the closed frontier tier above the shipped one) against two standing instruments: the 13-gate regression suite and the tag-steering family. Registered as a tier-map extension: no new hypotheses, no new gates, prompts byte-identical, and all registered verdicts pinned to the original trio. One disclosure carried in the brief: the analyst model is among those measured, so every grader used is mechanical and predates the runs; the judge-graded Track 2 studies are excluded.

### The Cheapest Model Passed Everything

gpt-oss-120b cleared all thirteen shipped-guardrail gates, anchored patches through memo mechanics, for 33 cents of API spend. The guardrail stack (context handed to the model, tools over rules, server-owned invariants) was designed not to depend on tier; the cheapest model in the series' history just delivered the strongest confirmation of that design yet.

### Two Frontiers, One Red Gate

kimi-k3 (8/10) and fable-5 (7/10) both went red on a single slice, goal-safe eviction at the 20-note memo cap edge, with the exact client-prune anatomy Study AK registered as the app's honest boundary: each sends a list already trimmed to the cap, so the app's goal-preserving eviction never engages, and the victim is an old goal, the one thing only the memo carries. The failure is cap-obedience. The aging shipped tier passes this slice by over-sending and letting the app decide; the two newest frontier-class models politely self-edit and choose victims badly, while gpt-oss-120b passes it 10/10. Per the standing procedure, neither red-gated model ships onto the chat surfaces as-is, and the filed Study AL powered re-measurement (the client-prune prompt fence, verdict "unproven, not disproven") now has the motivation it lacked: the likely next tiers exhibit exactly the pathway it targets.

### The Tag Family Tier Map

All the family's universals replicated: bare-arm drift 0/36 on every model, unprompted tags_list reads in every shipped-arm cell, zero invented tags anywhere. kimi-k3 posted the best adjacent-class numbers measured to date, 12/12 in all four AP arms under the AP′ independent key. fable-5 profiles as the shipped tier's successor, with the same refusal to tag off-catalog topics and the same one-sentence fix, but it mints freely where opus abstains. gpt-oss-120b is the cautionary shape: perfect catalog obedience, yet it never returns an empty list unaided, the guidance sentences only partially rescue it, and its warnings-only arm collapsed to 1/36, the starkest the-tool-is-the-mechanism datapoint in the family.

### What It Cost and What It Licenses

Total spend was $53.39 against a registered fence of $25 to $45; the overage came from underestimating fable-5's regression input volume and is disclosed in the report. Phase 2 (deeper family re-runs on the new models) stays unregistered and unspent. What this backfill licenses is a map, not a ship decision: the guardrails hold at every price point measured, the two newest frontier models share one specific, well-understood exposure at the memo cap edge, and any future tier swap starts from these gate summaries.
