{"slug": "olmo-detect-a-multi-stage-confounder-controlled-benchmark-for-membership-on", "title": "OLMo-Detect: A Multi-Stage, Confounder-Controlled Benchmark for Membership Inference on Large Language Models", "summary": "A new benchmark called OLMo-Detect, built on the fully open OLMo 2 pipeline, finds that membership inference attacks on large language models remain weak, with the best unsupervised and supervised attacks both reaching only 0.68 AUC across the OLMo 2 family. The benchmark spans pre-training, mid-training, and post-training, aligns members and non-members on three axes, and filters non-members via infini-gram; testing 15 unsupervised and 3 supervised membership inference attacks showed performance peaks at mid-training, driven by curated math data rather than a stage effect, improves from 1B to 13B parameters but plateaus at 32B, and no unsupervised attack is robust to distribution shifts, with AUCs shifting by up to 0.42. The authors report the findings generalize to OLMo 3 and non-OLMo models.", "body_md": "arXiv:2610.02986v1 Announce Type: new \nAbstract: Membership inference on large language models (LLMs) aims to determine whether a given text sample was included in an LLM's training data, without access to its training corpus. Despite recent progress, existing benchmarks suffer from three limitations: limited coverage of training stages, insufficient distributional alignment between members and non-members, and lack of rigorous filtering of non-members against the training corpus. To address these limitations, we propose OLMo-Detect, a multi-stage, confounder-controlled benchmark built upon the fully open OLMo 2 pipeline. OLMo-Detect spans pre-training, mid-training, and post-training, explicitly aligns members and non-members on three key axes, and rigorously filters non-members via infini-gram. To assess robustness to distribution shifts, we further introduce OLMo-Detect (Shifted), a variant where members are misaligned with non-members. We evaluate 15 unsupervised and 3 supervised membership inference attacks (MIAs) across the OLMo 2 family, finding that: (i) overall performance is limited: the best unsupervised and supervised MIAs both reach an AUC of only 0.68, and supervised MIAs degrade under cross-domain evaluation; (ii) MIA performance peaks at mid-training and is lower at pre-training and post-training, a pattern driven by data type rather than a stage effect: curated math data is far more detectable than other types; (iii) overall scores improve from 1B to 13B but plateau at 32B; and (iv) no unsupervised MIA is robust to distribution shifts, with AUCs shifting by up to 0.42. Finally, we find that our findings on OLMo 2 generalize to OLMo 3 and non-OLMo models.", "url": "https://wpnews.pro/news/olmo-detect-a-multi-stage-confounder-controlled-benchmark-for-membership-on", "canonical_source": "https://www.machinebrief.com/news/olmo-detect-a-multi-stage-confounder-controlled-benchmark-fo-2n3i", "published_at": "2026-10-05 04:00:00+00:00", "updated_at": "2026-10-05 05:12:42.657419+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "machine-learning", "ai-research", "ai-safety"], "entities": ["OLMo-Detect", "OLMo 2", "OLMo 3", "infini-gram"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/olmo-detect-a-multi-stage-confounder-controlled-benchmark-for-membership-on", "markdown": "https://wpnews.pro/news/olmo-detect-a-multi-stage-confounder-controlled-benchmark-for-membership-on.md", "text": "https://wpnews.pro/news/olmo-detect-a-multi-stage-confounder-controlled-benchmark-for-membership-on.txt", "jsonld": "https://wpnews.pro/news/olmo-detect-a-multi-stage-confounder-controlled-benchmark-for-membership-on.jsonld"}}