{"slug": "study-al-the-gate-fails-on-a-moved-baseline-and-unproven-does-not-ship", "title": "Study AL: The gate fails on a moved baseline; and unproven does not ship", "summary": "Study AL found that a prompt-level fence against self-eviction reduced client-side pruning from 5 to 1 pooled across models, but the pre-registered significance gate failed at p = .219 because the control arm's baseline drifted: sonnet's prune rate fell from 4 of 10 to 2, and gemini's from 6 of 10 to 3. The study's authors conclude the pathway is unproven, not disproven, and will not ship the change.", "body_md": "Study AL · Sessions & memory\n\nThe prompt-side fence\n\nStudy AK left one pathway open: the app-side eviction cannot restore a note the model pruned before sending, and the mid tiers pruned goals client-side in 4 to 6 cells of ten. AK's interpretation table filed the obvious candidate; one prompt sentence forbidding self-eviction (\"never drop or trim an existing note to make room; the app decides evictions\"); for measurement before shipping, and Study AL is that measurement, with the prompt rule as the only variable against the AK-eviction arm as its contemporaneous control.\n\nThe fence looked right everywhere the gates could not see: pooled client prunes fell 5 to 1, goal survival rose to 10 of 10 on sonnet and 9 of 10 on gemini, all sixty under-cap updates stayed untouched, and the sentence cost nothing; the models were already sending complete lists; the fence changed which notes survived, not how many tokens moved.\n\nBut the pre-registered significance gate failed at p = .219, for a reason nobody registered: the control arm did not replicate its own three-day-old baseline. Sonnet's prune rate halved (4 of 10 to 2), gemini's too (6 of 10 to 3); same prompts, same corpus, same handler, temperature 0, the largest drift the series has caught between two measurements of the same construction. Against a five-prune pooled baseline, ten cells per tier cannot clear the bar.\n\nSo the published verdict is deliberately double-edged: the registered reading (\"the prune pathway is not instruction-closable\") overstates what this run measured, and the honest reading; unproven, not disproven; still ships nothing, because unproven does not ship.\n\nWhat the study bought instead is sharper than a pass: the client-prune baseline itself is unstable week to week, which is the strongest argument yet for the filed alternative; an update shape that cannot express omission at all, so the guarantee stops depending on tier behavior holding still. Two smaller facts round it out: opus pruned for the first time in the series (a fact, not a goal; its goal safety stayed perfect), and consolidation-on-notice recurred, two sightings in two studies.", "url": "https://wpnews.pro/news/study-al-the-gate-fails-on-a-moved-baseline-and-unproven-does-not-ship", "canonical_source": "https://www.lightningjar.com/research/barkup-bench/al", "published_at": "2026-07-18 12:00:00+00:00", "updated_at": "2026-07-24 15:27:21.413761+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "large-language-models"], "entities": ["Study AL", "Study AK", "sonnet", "gemini", "opus"], "alternates": {"html": "https://wpnews.pro/news/study-al-the-gate-fails-on-a-moved-baseline-and-unproven-does-not-ship", "markdown": "https://wpnews.pro/news/study-al-the-gate-fails-on-a-moved-baseline-and-unproven-does-not-ship.md", "text": "https://wpnews.pro/news/study-al-the-gate-fails-on-a-moved-baseline-and-unproven-does-not-ship.txt", "jsonld": "https://wpnews.pro/news/study-al-the-gate-fails-on-a-moved-baseline-and-unproven-does-not-ship.jsonld"}}