cd /news/machine-learning/measuring-the-depth-of-llm-unlearnin… Β· home β€Ί topics β€Ί machine-learning β€Ί article
[ARTICLE Β· art-149221] src=gnueaj.github.io β†— pub= topic=machine-learning verified=true sentiment=Β· neutral

Measuring the Depth of LLM Unlearning via Activation Patching

Researchers from Sungkyunkwan University, Samsung Research, and KAIST presented the Unlearning Depth Score (UDS), a mechanistic metric that quantifies how deeply a model has unlearned target knowledge by measuring how much of it remains recoverable through two-stage activation patching, at EMNLP 2026 as an Oral presentation. The team benchmarked UDS against 20 unlearning metrics on the TOFU forget10 dataset using Llama-3.2-1B-Instruct and the Open-Unlearning framework, evaluating faithfulness via AUC-ROC over 30 knowledge-present and 30 knowledge-absent models and robustness under 4-bit NF4 quantization and 1-epoch relearning. UDS was also evaluated across 152 models spanning 8 unlearning methods, where it is combined with membership inference attacks as Privacy = HM(MIA, UDS).

read4 min views1 publishedOct 11, 2026

<sup>1</sup>Sungkyunkwan University Β Β  <sup>2</sup>Samsung Research Β Β  <sup>3</sup>KAIST EMNLP 2026 Oral

    We present the **Unlearning Depth Score (UDS)**, a mechanistic metric that quantifies the *depth of unlearning* by measuring how much target knowledge is recoverable through two-stage activation patching.
    This page provides: **(i) an interactive walkthrough of the UDS pipeline**, **(ii) a meta-evaluation comparing 20 metrics on faithfulness and robustness**, and **(iii) per-method benchmark results across 150 unlearned models**. The results below benchmark UDS on the TOFU [\[12\]](#ref12) dataset (forget10) using Llama-3.2-1B-Instruct and the Open-Unlearning [\[18\]](#ref18) framework.

Construct a prompt from question and answer prefix, and mark the entity span to evaluate.

Patch retain model's hidden states into the full model at each layer.

Log-probs are compared token-by-token under teacher forcing.

Large drops reveal layers encoding knowledge the retain model lacks.

Repeat with the unlearned model's hidden states.

Drops matching Stage 1 = erased; near-zero = knowledge intact. β‘  Filter layers that significantly encode the target knowledge.

β‘‘ Compute per-layer erasure ratio from the two stages.

β‘’ Aggregate into a final score: 0 (intact) β†’ 1 (erased).

      We evaluate **20 metrics** to measure **how reliable each unlearning metric is**, using the Open-Unlearning benchmark framework.
    

**Faithfulness**: How well each metric separates knowledge-present (P, 30 models) vs knowledge-absent (N, 30 models), measured by AUC-ROC.

**Robustness**: Stability of each metric under post-hoc perturbations (4-bit quantization, 1-epoch relearning).
    

`Faithfulness` = AUC-ROC over P/N pools (30 knowledge-present vs 30 knowledge-absent models)
      `Robustness` = HM(Q, R) β€” `Quantization (Q)` $= 1 - \dfrac{|m_{\text{after}} - m_{\text{before}}|}{|m_{\text{before}}| + |m_{\text{after}}|}$ β€” penalizes both recovery and destruction after 4-bit NF4 quantization`Relearning (R)` $= 1 - \dfrac{|\Delta_{\text{unl}} - \Delta_{\text{ret}}|}{|\Delta_{\text{unl}}| + |\Delta_{\text{ret}}|}$,   $\Delta = m_{\text{after}} - m_{\text{before}}$ β€” penalizes both over- and under-recovery relative to retain`Overall` = HM(Faithfulness, Robustness)
      `ES` `EM` `Prob` β€” geometric mean of per-token probabilities. $\exp\!\bigl(-(1/T)\sum \text{CE}_t\bigr)$`ParaProb` β€” geometric mean of Prob across paraphrased answer variants.`Truth Ratio` β€” normalized correct vs. incorrect probability. $p_c / (p_c + p_w)$,   $p = \exp(-\text{avg loss})$

ROUGE β€” standard prompt.Para. ROUGE β€” paraphrased prompt.Jailbreak ROUGE β€” adversarial prompt (prefixed with β€œSure, here is the answer:”). MIA-LOSS MIA-ZLib MIA-Min-K MIA-Min-K++ s<sub>LOSS</sub> Β· s<sub>ZLib</sub> Β· s<sub>Min-K</sub> Β· s<sub>Min-K++</sub> CKA Logit Lens Fisher Masked UDS (Ours) β€” Measures whether knowledge remains recoverable via activation patching.

Metric Overall ↑ Faithfulness ↑ Robustness
Aggregate ↑ Quantization ↑ Relearning ↑
...
      We evaluate **152 models** (8 methods Γ— varying hyperparameters Γ— 2 epochs + full + retain) across three axes.
      All unlearned model checkpoints are from the [Open-Unlearning](https://github.com/locuslab/open-unlearning) framework.

Memorization: How much target knowledge was forgotten. (↑ higher = more forgotten)

Privacy: How well sensitive information from the forget set is protected from being extracted. (↑ higher = better protected)

Utility: How well the model retains general capabilities on non-target knowledge. (↑ higher = better retention)

**Overall**: Harmonic mean of all three axes. (↑ higher = better)

`Privacy = HM(MIA, UDS)`, capturing both statistical (MIA) and mechanistic (UDS) aspects.
    `Mem.` $= \text{HM}(1{-}\text{ES},\; 1{-}\text{EM},\; 1{-}\text{ParaProb},\; 1{-}\text{TruthRatio})$`MIA` $= \text{HM}(s_{\text{LOSS}},\; s_{\text{ZLib}},\; s_{\text{Min-K}},\; s_{\text{Min-K++}})$`Privacy` = HM(MIA, UDS)
      `ModelUtility (MU)` = HM(retain_Prob, retain_ROUGE, retain_TruthRatio, ra_Prob, ra_ROUGE, ra_TruthRatio, wf_Prob, wf_ROUGE, wf_TruthRatio)`Fluency` = generation fluency score`Utility` = HM(MU, Fluency), then normalized: $\text{Utility}_{\text{rel}} = \text{Utility}\, /\, \text{Utility}_{\text{full(epoch)}}$
      | Method | Learning Rate | Swept Hyperparameters | Fixed | Epochs | Models | 
|---|---|---|---|---|---|

| GradDiff [12] , IdkNLL[12] , IdkDPO[12] , NPO[13] , AltPO[14] | {1e-5, 2e-5, 5e-5} | $\alpha$ ∈ {1, 2, 5} | Ξ² = 0.1 | {5, 10} | 5 Γ— 3 Γ— 3 Γ— 2 = 90 |

| SimNPO [\[15\]](#ref15) | {1e-5, 2e-5, 5e-5} | $\beta$ ∈ {3.5, 4.5}, $\gamma$ ∈ {0.125, 0.25} | Ξ΄ = 1, Ξ± = 1 | {5, 10} | 1 Γ— 3 Γ— 4 Γ— 2 = 24 | 
| RMU [\[16\]](#ref16) | {1e-5, 2e-5, 5e-5} | layer ∈ {5, 10, 15} | steering coeff = 10 | {5, 10} | 1 Γ— 3 Γ— 3 Γ— 2 = 18 | 
| UNDIAL [\[17\]](#ref17) | {1e-5, 1e-4, 3e-4} | $\alpha$ ∈ {1, 2, 5} | Ξ² = 10 | {5, 10} | 1 Γ— 3 Γ— 3 Γ— 2 = 18 | 

| Total: 150 unlearned + full + retain | | | | | 152 |

Model Overall <sub>w/o UDS</sub> ↑ Overall <sub>w/ UDS</sub> ↑ Mem. ↑ Privacy ↑ Utility ↑ LL ↑ UDS ↑
...
── more in #machine-learning 4 stories Β· sorted by recency
── more on @sungkyunkwan university 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/measuring-the-depth-…] indexed:0 read:4min 2026-10-11 Β· β€”