Study AP: The off-catalog fork: one sentence settles the foreign branch; the registration loses the adjacent one A study by AP found that a guidance sentence instructing models to return an empty tag list for wholly-foreign topics achieved 12/12 accuracy across all three tested models (Opus, Gemini, Sonnet) with no side effects on adjacent topics. The pre-registered gate for adjacent topics failed because Opus cleared only 8/12 cells against a 10-cell bar, with misses involving editorially arguable tags. The study filed a re-registered definition AP′ and deferred the guidance-text decision, while the foreign branch guidance is ready to ship. Study AP · Agent-surface steering The off-catalog fork: one sentence settles the foreign branch; the registration loses the adjacent one The off-catalog fork Study Overview The Fork AO Filed Study AO ended with a protocol lesson: its non-empty "clean" definition mis-scored opus declining to tag off-catalog topics. AP registered that fork explicitly, splitting the old uncovered class into adjacent a general canonical tag genuinely fits and foreign nothing does , and measured two candidate guidance sentences plus a guarded tag create tool: 48 tasks, four arms, three models, 576 cells. The Foreign Branch Settles Unaided, the frontier tier already refuses: opus returned empty tag lists on 12/12 wholly-foreign topics, replicating AO exactly, while gemini and sonnet stretched general canonical tags onto 5 of 12. Either guidance sentence closes that gap. The minimal empty text "if no canonical tag genuinely fits, return an empty tags list" hit 12/12 on all three models, and its feared side effect, suppressing legitimate tags on adjacent topics, never appeared: zero empty adjacent cells in any arm on any model. The Tool Is Discipline-Safe, With a Tier Map With a guarded tag create in hand, no model produced a single invented tag in 144 tool-arm cells, and the server-side guard never fired: all 40 mints were well-formed, Title Case, and free of near-duplicates. Willingness split sharply by tier: gemini minted on every foreign topic 26 mints , sonnet 10, opus 4; the frontier tier prefers empty over minting. A sub-frontier agent grows the catalog by a dozen well-formed, unwanted tags per batch, so if the tool ships it ships tier-gated or with human approval on new tags. The Gate Failed on the Registration The pre-registered gate failed by the letter: opus cleared 8/12 adjacent cells under the fork text against a ≥10 bar. The audit shows every failing cell was all-canonical; the misses are one or two extra, editorially arguable tags beyond the solo-authored acceptable sets UX on a checkout-conversion piece, CSS on web components, Security on email deliverability . Under an anchored reading all tags canonical, at least one from the registered set adjacent scores 144/144 across every model and arm; that number is published as exploratory, not the registered definition. AO's lesson recurred one class over: acceptance criteria authored alone under-cover defensible editorial judgment. What Ships and What Waits Nothing ships from a failed gate. Filed: AP′, which re-registers the adjacent conformance definition and re-scores the same frozen raw results with no new API spend; the guidance-text decision waits for it. The foreign branch does not wait: on registered gates, either sentence stops sub-frontier stretching onto wholly-foreign topics, and the frontier tier never needed the help.