{"slug": "frozen-bert-neurons-detect-ai-text-via-sparse-probing-patching-raid", "title": "Frozen BERT neurons detect AI text via sparse probing patching RAID", "summary": "A sparse-probing study on a frozen BERT-base-uncased encoder identified fewer than 1% of the model's 9,216 CLS hidden-state units (12 layers × 768) as stable AI-text detectors across six generators on the RAID benchmark, with the selected neurons retaining 86-94% of full-feature detection performance on unseen generator families in a leave-one-family-out test. Bidirectional activation patching flipped predictions an order of magnitude more often than size-matched random neuron sets, confirming causal relevance, while mean-ablating the same neurons barely affected accuracy, indicating the signal is redundantly encoded. Instruction-tuned generators placed 30-36% of their stable neurons in BERT's final layer versus under 14% for pure-base generators, and the work is published as arXiv:2609.30287v1.", "body_md": "# Frozen BERT neurons detect AI text via sparse probing patching RAID\n\nThe study recovered fewer than 1 % of BERT’s CLS hidden dimensions as stable detectors for six different generators spanning pure‑base and instruction‑tuned models. I ran the L1‑to‑L2 sparse‑probing protocol from Gurnee et al. (2023) on a frozen BERT‑base‑uncased encoder, probing all 9,216 CLS hidden‑state units (12 layers × 768). The probing was performed on the RAID benchmark, and the procedure consistently isolated a small, reproducible set of neurons per generator across folds and random seeds.\n\n**Running the sparse‑probe**\n\n1. Load the frozen BERT‑base‑uncased checkpoint; keep all weights static.  \n\n2. Extract CLS hidden states for each generated text (both base and instruction‑tuned).  \n\n3. Apply L1‑to‑L2 regularization to a linear probe across the full 9,216‑dimensional space.  \n\n4. Evaluate probe accuracy on held‑out folds; record the subset of neurons that contributed the most weight.  \n\n5. Repeat with different random seeds to confirm stability.\n\nThe result was a compact subspace—under 1 % of the total dimensions—for each generator. Restricting a fresh probe to only those selected neurons retained most of the original detection performance, showing that the rest of the space adds little signal.\n\n**Causal verification with activation patching**\n\nNext I performed bidirectional activation patching to see whether the identified neurons actually drive predictions.  \n\n- In the forward direction, I replaced the CLS activations of the candidate neurons with random values while leaving the rest of the layer intact.\n- In the backward direction, I restored random CLS activations and left the candidate neurons fixed.\n\nBoth directions flipped predictions an order of magnitude more often than size‑matched random neuron sets, confirming causal relevance. By contrast, mean‑ablating the same neurons barely impacted accuracy, indicating that the detection signal is redundantly encoded across the network.\n\n**Neuron distribution across layers**\n\nCross‑generator analysis revealed a clear split: instruction‑tuned generators packed 30‑36 % of their stable neurons into BERT’s final layer (layer 12), while pure‑base generators kept that figure below 14 %. This matches the intuition that post‑training alignment leaves a stronger footprint in the topmost layer for instruction‑following models.\n\n**Generalizing to unseen families**\n\nA leave‑one‑family‑out test showed that the selected neurons kept 86‑94 % of the full‑feature ceiling when applied to a completely new set of generators. In practice, this means a detector can be built once on a small fixed subspace and reused without re‑identifying new neurons for each model that comes out.\n\n**What you can take away**\n\nIf you need a lightweight AI‑text detector that works across multiple generators, the key is to focus on that tiny, stable subset of CLS neurons rather than the whole hidden state. The protocol is straightforward: freeze BERT, run a sparse probe on the CLS dimensions, and validate causality with patching. The redundancy observed also suggests you won’t lose much performance if you accidentally drop a few of those neurons in a production pipeline.\n\nThe paper is available as arXiv:2609.30287v1, and the numbers (12 × 768 dimensions, <1 % stable neurons, 86‑94 % cross‑family retention) give you concrete checkpoints when you replicate the experiment.\n\n[Next Xpeng G9L Launches from 231,800 Yuan, Targets Global Segment Leadership →](https://promptcube3.com/en/news/9497/)\n\n## All Replies （1）\n\nWant a live back-and-forth? [Join the global AI chat room](https://promptcube3.com/en/chat/) — login to talk.\n\nDid the sparse probing keep the same neurons across the six generators, or did each one get its own distinct set?", "url": "https://wpnews.pro/news/frozen-bert-neurons-detect-ai-text-via-sparse-probing-patching-raid", "canonical_source": "https://promptcube3.com/en/news/9649/", "published_at": "2026-09-28 04:09:06+00:00", "updated_at": "2026-09-28 04:17:48.646341+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-research", "large-language-models"], "entities": ["BERT", "BERT-base-uncased", "RAID benchmark", "Gurnee et al. (2023)", "arXiv:2609.30287v1"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/frozen-bert-neurons-detect-ai-text-via-sparse-probing-patching-raid", "markdown": "https://wpnews.pro/news/frozen-bert-neurons-detect-ai-text-via-sparse-probing-patching-raid.md", "text": "https://wpnews.pro/news/frozen-bert-neurons-detect-ai-text-via-sparse-probing-patching-raid.txt", "jsonld": "https://wpnews.pro/news/frozen-bert-neurons-detect-ai-text-via-sparse-probing-patching-raid.jsonld"}}