# Study AP: The off-catalog fork: one sentence settles the foreign branch; the registration loses the adjacent one

> Source: <https://www.lightningjar.com/research/barkup-bench/ap>
> Published: 2026-07-28 12:00:00+00:00

Study AP · Agent-surface steering

# The off-catalog fork: one sentence settles the foreign branch; the registration loses the adjacent one

The off-catalog fork

## Study Overview

### The Fork AO Filed

Study AO ended with a protocol lesson: its non-empty "clean" definition mis-scored opus declining to tag off-catalog topics. AP registered that fork explicitly, splitting the old uncovered class into *adjacent* (a general canonical tag genuinely fits) and *foreign* (nothing does), and measured two candidate guidance sentences plus a guarded tag_create tool: 48 tasks, four arms, three models, 576 cells.

### The Foreign Branch Settles

Unaided, the frontier tier already refuses: opus returned empty tag lists on 12/12 wholly-foreign topics, replicating AO exactly, while gemini and sonnet stretched general canonical tags onto 5 of 12. Either guidance sentence closes that gap. The minimal empty text ("if no canonical tag genuinely fits, return an empty tags list") hit 12/12 on all three models, and its feared side effect, suppressing legitimate tags on adjacent topics, never appeared: zero empty adjacent cells in any arm on any model.

### The Tool Is Discipline-Safe, With a Tier Map

With a guarded tag_create in hand, no model produced a single invented tag in 144 tool-arm cells, and the server-side guard never fired: all 40 mints were well-formed, Title Case, and free of near-duplicates. Willingness split sharply by tier: gemini minted on every foreign topic (26 mints), sonnet 10, opus 4; the frontier tier prefers empty over minting. A sub-frontier agent grows the catalog by a dozen well-formed, unwanted tags per batch, so if the tool ships it ships tier-gated or with human approval on new tags.

### The Gate Failed on the Registration

The pre-registered gate failed by the letter: opus cleared 8/12 adjacent cells under the fork text against a ≥10 bar. The audit shows every failing cell was all-canonical; the misses are one or two extra, editorially arguable tags beyond the solo-authored acceptable sets (UX on a checkout-conversion piece, CSS on web components, Security on email deliverability). Under an anchored reading (all tags canonical, at least one from the registered set) adjacent scores 144/144 across every model and arm; that number is published as exploratory, not the registered definition. AO's lesson recurred one class over: acceptance criteria authored alone under-cover defensible editorial judgment.

### What Ships and What Waits

Nothing ships from a failed gate. Filed: AP′, which re-registers the adjacent conformance definition and re-scores the same frozen raw results with no new API spend; the guidance-text decision waits for it. The foreign branch does not wait: on registered gates, either sentence stops sub-frontier stretching onto wholly-foreign topics, and the frontier tier never needed the help.
