The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

wpnews.pro

cd /news/machine-learning/the-curse-of-multiple-mediators-hidd… · home › topics › machine-learning › article

[ARTICLE · art-42923] src=arxiv.org ↗ pub=2026-06-29T04:00Z topic=machine-learning verified=true sentiment=· neutral

The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

Researchers at arXiv have discovered that activation patching, a key tool in mechanistic interpretability, contains hidden interaction effects (INT) that can distort causal attributions. These effects, which measure how a component's causal influence depends on other components, can cause some mechanisms to be invisible or artificially inflated, as demonstrated in GPT-2's IOI circuit. The findings suggest INT should be used as a diagnostic rather than eliminated, as it reveals prompt-dependent conclusions and limitations of greedy component ranking.

read1 min views1 publishedJun 29, 2026

arXiv:2606.27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its natural indirect effect (NIE). Re-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component. It also contains interaction effects (INT) that measure how much the component's causal effect itself depends on the state of other components in the model. A natural response may be to try to eliminate INT by adjusting the estimator or unit of analysis, but each of these potential remedies has predictable failure modes. We demonstrate these failure modes in the GPT-2 IOI circuit; components whose causal importance is conditional on the state of other components are either invisible or artificially inflated, and INT variance explains the previously documented instability of faithfulness scores. We prove that INT scales with the distance between clean and patched component activations, is negligible when the model is locally affine, and decomposes combinatorially into pairwise and higher-order group interactions. Despite its inevitability, INT is not a nuisance to be eliminated, but rather a diagnostic for interpretability studies. Its individual and group-level magnitude and sign signal when causal conclusions are prompt-dependent, and when greedy NIE-based component ranking will miss mechanisms only discoverable through combinatorial search.

source & further reading

arxiv.org — original article

~/api · this article 200

$curl api.wpnews.pro/v1/news/the-curse-of-multiple-me…

Read original on arxiv.org → arxiv.org/abs/2606.27510

mentioned entities

arXiv

GPT-2

metadata

slugthe-curse-of-multiple-mediators-hidden-interaction-effects-in-activation

topic#machine-learning

secondary2 topics

sentimentneutral

canonicalarxiv.org

navigation

← prevv0.5.6

next →Media Buying Briefing: The holdc…

── more in #machine-learning 4 stories · sorted by recency

arxiv.org · 29 Jun · #machine-learning

Prism Transformer: Progressive Head Schedules for Hierarchical Attention Processing

arxiv.org · 29 Jun · #machine-learning

The Context-Ready Transformer

arxiv.org · 29 Jun · #machine-learning

Tessellating The Earth

arxiv.org · 29 Jun · #machine-learning

Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment

── more on @arxiv 3 stories trending now

wpnews · 28 May · #ai-startups

[AINews] Cognition raises $1B in $26B Series D

wpnews · 5 Jun · #ai-agents

Miasma Worm Targets AI Coding Agents via GitHub Repos

wpnews · 28 Jun · #ai-agents

OpenCode v1.17: Session Snapshots Undo Your AI Agent

sponsored brought to you by zahid.host 4,200+ EU-deployed projects

reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main

→ Live at https://your-agent.zahid.host ✓

Get free account → Pricing

from €0/mo · no card required