cd /news/machine-learning/the-curse-of-multiple-mediators-hidd… · home topics machine-learning article
[ARTICLE · art-42923] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

Researchers at arXiv have discovered that activation patching, a key tool in mechanistic interpretability, contains hidden interaction effects (INT) that can distort causal attributions. These effects, which measure how a component's causal influence depends on other components, can cause some mechanisms to be invisible or artificially inflated, as demonstrated in GPT-2's IOI circuit. The findings suggest INT should be used as a diagnostic rather than eliminated, as it reveals prompt-dependent conclusions and limitations of greedy component ranking.

read1 min views1 publishedJun 29, 2026

arXiv:2606.27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its natural indirect effect (NIE). Re-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component. It also contains interaction effects (INT) that measure how much the component's causal effect itself depends on the state of other components in the model. A natural response may be to try to eliminate INT by adjusting the estimator or unit of analysis, but each of these potential remedies has predictable failure modes. We demonstrate these failure modes in the GPT-2 IOI circuit; components whose causal importance is conditional on the state of other components are either invisible or artificially inflated, and INT variance explains the previously documented instability of faithfulness scores. We prove that INT scales with the distance between clean and patched component activations, is negligible when the model is locally affine, and decomposes combinatorially into pairwise and higher-order group interactions. Despite its inevitability, INT is not a nuisance to be eliminated, but rather a diagnostic for interpretability studies. Its individual and group-level magnitude and sign signal when causal conclusions are prompt-dependent, and when greedy NIE-based component ranking will miss mechanisms only discoverable through combinatorial search.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-curse-of-multipl…] indexed:0 read:1min 2026-06-29 ·