One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them

wpnews.pro

cd /news/machine-learning/one-mask-to-rule-them-all-on-hidden-… · home › topics › machine-learning › article

[ARTICLE · art-17125] src=arxiv.org ↗ pub=2026-05-29T04:00Z topic=machine-learning verified=true sentiment=· neutral

One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them

Researchers at arXiv have discovered that knowledge editing methods ROME and MEMIT, which modify factual associations in transformer models, rely on a shared functional mechanism rather than fact-specific weight changes. By training a binary mask over edited weights, the team reversed over 70% of edits on test sets and found that edits suppress rather than overwrite knowledge, explaining why these methods fail to propagate changes to related facts. The findings reveal a common functional subspace that could inform detection and defense against unwanted edits in AI systems.

read1 min views12 publishedMay 29, 2026

arXiv:2605.28839v1 Announce Type: new Abstract: Knowledge editing methods such as ROME and MEMIT update factual associations in transformer models by modifying MLP weights. While evaluated mainly by output behavior, their internal mechanism remains underexplored. We investigate whether edits rely on a common mechanism, regardless of which fact is modified. Despite fact-specific weight changes, we argue that ROME and MEMIT target the same subset of weights critical for maintaining edits. To isolate this subset, we train a compact binary mask over the edited weights. The mask reverses 80% of edits on the training set and over 70% on the test set, confirming that diverse edits share a common functional structure. Our analysis reveals that the mask reverses edits by eliminating overattention in later layers. Additionally, we show that injecting the mask during editing drops editing success from 98% to 38%, demonstrating that this mechanism is necessary for edits to succeed. Our finding that edits suppress rather than overwrite knowledge explains why ROME and MEMIT fail to propagate changes to related facts. The identified common functional subspace informs detection and defense against unwanted edits.

source & further reading

arxiv.org — original article

~/api · this article 200

$curl api.wpnews.pro/v1/news/one-mask-to-rule-them-al…

Read original on arxiv.org → arxiv.org/abs/2605.28839

mentioned entities

ROME

MEMIT

metadata

slugone-mask-to-rule-them-all-on-hidden-facts-after-editing-and-how-to-find-them

topic#machine-learning

secondary4 topics

sentimentneutral

canonicalarxiv.org

navigation

← prevChatGPT glitch is leaking OpenAI…

next →New infosec products of the mont…

── more in #machine-learning 4 stories · sorted by recency

machinebrief.com · 16 Jul · #machine-learning

Precision Unlearning: The Revolutionary Approach to Data Deletion

machinebrief.com · 16 Jul · #machine-learning

Reinforcement Learning: The Future of Cyber-Defense

machinebrief.com · 16 Jul · #machine-learning

Model Evaluation: Beyond Reward Prediction

machinebrief.com · 16 Jul · #machine-learning

Operator Approximation: A New Theorem Challenges the Norm

── more on @rome 3 stories trending now

wpnews · 27 May · #artificial-intelligence

How I Run Two Claude Accounts as One

wpnews · 8 Jul · #ai-chips

D-Matrix launches Corsair AI inference platform, challenging Nvidia’s GPU dominance

wpnews · 8 Jul · #artificial-intelligence

What Is Vibe Coding? How AI Builds Games From Scratch

sponsored brought to you by zahid.host 4,200+ EU-deployed projects

reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main

→ Live at https://your-agent.zahid.host ✓

Get free account → Pricing

from €0/mo · no card required