{"slug": "the-curse-of-multiple-mediators-hidden-interaction-effects-in-activation", "title": "The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching", "summary": "Researchers at arXiv have discovered that activation patching, a key tool in mechanistic interpretability, contains hidden interaction effects (INT) that can distort causal attributions. These effects, which measure how a component's causal influence depends on other components, can cause some mechanisms to be invisible or artificially inflated, as demonstrated in GPT-2's IOI circuit. The findings suggest INT should be used as a diagnostic rather than eliminated, as it reveals prompt-dependent conclusions and limitations of greedy component ranking.", "body_md": "arXiv:2606.27510v1 Announce Type: new\nAbstract: Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its natural indirect effect (NIE). Re-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component. It also contains interaction effects (INT) that measure how much the component's causal effect itself depends on the state of other components in the model. A natural response may be to try to eliminate INT by adjusting the estimator or unit of analysis, but each of these potential remedies has predictable failure modes. We demonstrate these failure modes in the GPT-2 IOI circuit; components whose causal importance is conditional on the state of other components are either invisible or artificially inflated, and INT variance explains the previously documented instability of faithfulness scores. We prove that INT scales with the distance between clean and patched component activations, is negligible when the model is locally affine, and decomposes combinatorially into pairwise and higher-order group interactions. Despite its inevitability, INT is not a nuisance to be eliminated, but rather a diagnostic for interpretability studies. Its individual and group-level magnitude and sign signal when causal conclusions are prompt-dependent, and when greedy NIE-based component ranking will miss mechanisms only discoverable through combinatorial search.", "url": "https://wpnews.pro/news/the-curse-of-multiple-mediators-hidden-interaction-effects-in-activation", "canonical_source": "https://arxiv.org/abs/2606.27510", "published_at": "2026-06-29 04:00:00+00:00", "updated_at": "2026-06-29 04:09:33.116752+00:00", "lang": "en", "topics": ["machine-learning", "neural-networks", "ai-research"], "entities": ["arXiv", "GPT-2"], "alternates": {"html": "https://wpnews.pro/news/the-curse-of-multiple-mediators-hidden-interaction-effects-in-activation", "markdown": "https://wpnews.pro/news/the-curse-of-multiple-mediators-hidden-interaction-effects-in-activation.md", "text": "https://wpnews.pro/news/the-curse-of-multiple-mediators-hidden-interaction-effects-in-activation.txt", "jsonld": "https://wpnews.pro/news/the-curse-of-multiple-mediators-hidden-interaction-effects-in-activation.jsonld"}}