{"slug": "measuring-the-depth-of-llm-unlearning-via-activation-patching", "title": "Measuring the Depth of LLM Unlearning via Activation Patching", "summary": "Researchers from Sungkyunkwan University, Samsung Research, and KAIST presented the Unlearning Depth Score (UDS), a mechanistic metric that quantifies how deeply a model has unlearned target knowledge by measuring how much of it remains recoverable through two-stage activation patching, at EMNLP 2026 as an Oral presentation. The team benchmarked UDS against 20 unlearning metrics on the TOFU forget10 dataset using Llama-3.2-1B-Instruct and the Open-Unlearning framework, evaluating faithfulness via AUC-ROC over 30 knowledge-present and 30 knowledge-absent models and robustness under 4-bit NF4 quantization and 1-epoch relearning. UDS was also evaluated across 152 models spanning 8 unlearning methods, where it is combined with membership inference attacks as Privacy = HM(MIA, UDS).", "body_md": "<sup>1</sup>Sungkyunkwan University    <sup>2</sup>Samsung Research    <sup>3</sup>KAIST\n\nEMNLP 2026 Oral\n\n    We present the **Unlearning Depth Score (UDS)**, a mechanistic metric that quantifies the *depth of unlearning* by measuring how much target knowledge is recoverable through two-stage activation patching.\n    This page provides: **(i) an interactive walkthrough of the UDS pipeline**, **(ii) a meta-evaluation comparing 20 metrics on faithfulness and robustness**, and **(iii) per-method benchmark results across 150 unlearned models**. The results below benchmark UDS on the TOFU [\\[12\\]](#ref12) dataset (forget10) using Llama-3.2-1B-Instruct and the Open-Unlearning [\\[18\\]](#ref18) framework.\n  \n\nConstruct a prompt from question and answer prefix, and mark the entity span to evaluate.\n\nPatch retain model's hidden states into the full model at each layer.\n\nLog-probs are compared token-by-token under teacher forcing.\n\nLarge drops reveal layers encoding knowledge the retain model lacks.\n\nRepeat with the unlearned model's hidden states.\n\nDrops matching Stage 1 = erased; near-zero = knowledge intact.\n\n① Filter layers that significantly encode the target knowledge.\n\n② Compute per-layer erasure ratio from the two stages.\n\n③ Aggregate into a final score: 0 (intact) → 1 (erased).\n\n      We evaluate **20 metrics** to measure **how reliable each unlearning metric is**, using the Open-Unlearning benchmark framework.\n    \n\n**Faithfulness**: How well each metric separates knowledge-present (P, 30 models) vs knowledge-absent (N, 30 models), measured by AUC-ROC.\n\n**Robustness**: Stability of each metric under post-hoc perturbations (4-bit quantization, 1-epoch relearning).\n    \n\n`Faithfulness` = AUC-ROC over P/N pools (30 knowledge-present vs 30 knowledge-absent models)\n      `Robustness` = HM(Q, R) — `Quantization (Q)` $= 1 - \\dfrac{|m_{\\text{after}} - m_{\\text{before}}|}{|m_{\\text{before}}| + |m_{\\text{after}}|}$ — penalizes both recovery and destruction after 4-bit NF4 quantization`Relearning (R)` $= 1 - \\dfrac{|\\Delta_{\\text{unl}} - \\Delta_{\\text{ret}}|}{|\\Delta_{\\text{unl}}| + |\\Delta_{\\text{ret}}|}$,   $\\Delta = m_{\\text{after}} - m_{\\text{before}}$ — penalizes both over- and under-recovery relative to retain`Overall` = HM(Faithfulness, Robustness)\n      `ES` `EM` `Prob` — geometric mean of per-token probabilities. $\\exp\\!\\bigl(-(1/T)\\sum \\text{CE}_t\\bigr)$`ParaProb` — geometric mean of Prob across paraphrased answer variants.`Truth Ratio` — normalized correct vs. incorrect probability. $p_c / (p_c + p_w)$,   $p = \\exp(-\\text{avg loss})$\n`ROUGE` — standard prompt.`Para. ROUGE` — paraphrased prompt.`Jailbreak ROUGE` — adversarial prompt (prefixed with “Sure, here is the answer:”).\n      `MIA-LOSS` `MIA-ZLib` `MIA-Min-K` `MIA-Min-K++` `s`<sub>LOSS</sub> · `s`<sub>ZLib</sub> · `s`<sub>Min-K</sub> · `s`<sub>Min-K++</sub>\n`CKA` `Logit Lens` `Fisher Masked` `UDS (Ours)` — Measures whether knowledge remains recoverable via activation patching.\n| Metric | Overall ↑ | Faithfulness ↑ | Robustness |  |  | \n|---|---|---|---|---|---|\n|  |  |  | Aggregate ↑ | Quantization ↑ | Relearning ↑ | \n| Loading... |  |  |  |  |  | \n\n      We evaluate **152 models** (8 methods × varying hyperparameters × 2 epochs + full + retain) across three axes.\n      All unlearned model checkpoints are from the [Open-Unlearning](https://github.com/locuslab/open-unlearning) framework.\n    \n\n**Memorization**: How much target knowledge was forgotten. (↑ higher = more forgotten)\n\n**Privacy**: How well sensitive information from the forget set is protected from being extracted. (↑ higher = better protected)\n\n**Utility**: How well the model retains general capabilities on non-target knowledge. (↑ higher = better retention)\n\n**Overall**: Harmonic mean of all three axes. (↑ higher = better)\n\n`Privacy = HM(MIA, UDS)`, capturing both statistical (MIA) and mechanistic (UDS) aspects.\n    `Mem.` $= \\text{HM}(1{-}\\text{ES},\\; 1{-}\\text{EM},\\; 1{-}\\text{ParaProb},\\; 1{-}\\text{TruthRatio})$`MIA` $= \\text{HM}(s_{\\text{LOSS}},\\; s_{\\text{ZLib}},\\; s_{\\text{Min-K}},\\; s_{\\text{Min-K++}})$`Privacy` = HM(MIA, UDS)\n      `ModelUtility (MU)` = HM(retain_Prob, retain_ROUGE, retain_TruthRatio, ra_Prob, ra_ROUGE, ra_TruthRatio, wf_Prob, wf_ROUGE, wf_TruthRatio)`Fluency` = generation fluency score`Utility` = HM(MU, Fluency), then normalized: $\\text{Utility}_{\\text{rel}} = \\text{Utility}\\, /\\, \\text{Utility}_{\\text{full(epoch)}}$\n      | Method | Learning Rate | Swept Hyperparameters | Fixed | Epochs | Models | \n|---|---|---|---|---|---|\n| GradDiff [\\[12\\]](#ref12) , IdkNLL[\\[12\\]](#ref12) , IdkDPO[\\[12\\]](#ref12) , NPO[\\[13\\]](#ref13) , AltPO[\\[14\\]](#ref14) | {1e-5, 2e-5, 5e-5} | $\\alpha$ ∈ {1, 2, 5} | β = 0.1 | {5, 10} | 5 × 3 × 3 × 2 = 90 | \n| SimNPO [\\[15\\]](#ref15) | {1e-5, 2e-5, 5e-5} | $\\beta$ ∈ {3.5, 4.5}, $\\gamma$ ∈ {0.125, 0.25} | δ = 1, α = 1 | {5, 10} | 1 × 3 × 4 × 2 = 24 | \n| RMU [\\[16\\]](#ref16) | {1e-5, 2e-5, 5e-5} | layer ∈ {5, 10, 15} | steering coeff = 10 | {5, 10} | 1 × 3 × 3 × 2 = 18 | \n| UNDIAL [\\[17\\]](#ref17) | {1e-5, 1e-4, 3e-4} | $\\alpha$ ∈ {1, 2, 5} | β = 10 | {5, 10} | 1 × 3 × 3 × 2 = 18 | \n| Total: 150 unlearned + full + retain |  |  |  |  | 152 | \n\n| Model | Overall <sub>w/o UDS</sub> ↑ | Overall <sub>w/ UDS</sub> ↑ | Mem. ↑ | Privacy ↑ | Utility ↑ | LL ↑ | UDS ↑ |  | \n|---|---|---|---|---|---|---|---|---|\n| Loading... |  |  |  |  |  |  |  |  |", "url": "https://wpnews.pro/news/measuring-the-depth-of-llm-unlearning-via-activation-patching", "canonical_source": "https://gnueaj.github.io/unlearning-depth-score/", "published_at": "2026-10-11 16:42:16+00:00", "updated_at": "2026-10-11 17:00:33.450282+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "ai-safety"], "entities": ["Sungkyunkwan University", "Samsung Research", "KAIST", "Unlearning Depth Score", "TOFU", "Llama-3.2-1B-Instruct", "Open-Unlearning", "EMNLP 2026"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/measuring-the-depth-of-llm-unlearning-via-activation-patching", "markdown": "https://wpnews.pro/news/measuring-the-depth-of-llm-unlearning-via-activation-patching.md", "text": "https://wpnews.pro/news/measuring-the-depth-of-llm-unlearning-via-activation-patching.txt", "jsonld": "https://wpnews.pro/news/measuring-the-depth-of-llm-unlearning-via-activation-patching.jsonld"}}