Measuring the Depth of LLM Unlearning via Activation Patching Researchers from Sungkyunkwan University, Samsung Research, and KAIST presented the Unlearning Depth Score (UDS), a mechanistic metric that quantifies how deeply a model has unlearned target knowledge by measuring how much of it remains recoverable through two-stage activation patching, at EMNLP 2026 as an Oral presentation. The team benchmarked UDS against 20 unlearning metrics on the TOFU forget10 dataset using Llama-3.2-1B-Instruct and the Open-Unlearning framework, evaluating faithfulness via AUC-ROC over 30 knowledge-present and 30 knowledge-absent models and robustness under 4-bit NF4 quantization and 1-epoch relearning. UDS was also evaluated across 152 models spanning 8 unlearning methods, where it is combined with membership inference attacks as Privacy = HM(MIA, UDS).