{"slug": "when-do-attention-head-ablations-support-causal-claims-projection-level-floor", "title": "When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls", "summary": "A new arXiv paper (2610.00373v1) reports that attention-head ablation in GPT-2 small can produce misleading causal claims unless intervention placement, evaluation metric, and controls are validated. The authors found a post-projection implementation of \"zeroing a head\" correlates only weakly with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads, while binary accuracy hides effects at behavioral floors and ceilings that gold-token log-probability still captures. Using a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the corrected per-head ranking was stable across splits (Spearman rho = 0.974) and the top-5 heads exceeded both control distributions (Monte Carlo p = 0.001), though task specificity did not replicate robustly on GPT-2; DistilGPT2 preserved the intervention-semantic and matched-control findings.", "body_md": "arXiv:2610.00373v1 Announce Type: new \nAbstract: Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a common method for inferring which components of a language model are causally responsible for a behavior. We show using GPT-2 small that this inference can be fragile unless the intervention semantics, evaluation metric, and controls are carefully validated. A natural post-projection implementation of \"zeroing a head\" is nearly uncorrelated with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads. We also show that binary accuracy can hide effects at behavioral floors and near ceilings, whereas gold-token log-probability remains graded. Using a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the corrected per-head effect ranking is highly stable across splits (Spearman rho = 0.974), and the top-5 selected heads significantly exceed both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is not robust on GPT-2. Replication on DistilGPT2 preserves the intervention-semantic and matched-control findings. These results show that single-head ablation does not by itself justify a causal claim; defensible interpretation requires correct intervention placement, a non-saturated continuous metric, and matched held-out controls.", "url": "https://wpnews.pro/news/when-do-attention-head-ablations-support-causal-claims-projection-level-floor", "canonical_source": "https://www.machinebrief.com/news/when-do-attention-head-ablations-support-causal-claims-proje-9z6i", "published_at": "2026-10-03 04:00:00+00:00", "updated_at": "2026-10-03 04:38:12.709201+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-safety"], "entities": ["GPT-2 small", "DistilGPT2", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-do-attention-head-ablations-support-causal-claims-projection-level-floor", "markdown": "https://wpnews.pro/news/when-do-attention-head-ablations-support-causal-claims-projection-level-floor.md", "text": "https://wpnews.pro/news/when-do-attention-head-ablations-support-causal-claims-projection-level-floor.txt", "jsonld": "https://wpnews.pro/news/when-do-attention-head-ablations-support-causal-claims-projection-level-floor.jsonld"}}