cd /news/large-language-models/decodable-but-not-detachable-trainin… · home topics large-language-models article
[ARTICLE · art-93062] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models

A new arXiv study (2608.10214v1) finds that large language models contain domain-specific parametric shells only when training data was modular at the token level, with 0.65–1.14% of neurons exceeding 60% selectivity for language and modality domains, while academic subject domains show no such structure. The study, spanning three model families (1.5B to 7B parameters) and eight domains, reports that masking code-selective neurons reduces mathematical reasoning accuracy by 16–24 percentage points, and shell strength increases with scale.

read1 min views1 publishedAug 12, 2026

arXiv:2608.10214v1 Announce Type: new Abstract: Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selectively degrades a target domain while sparing others? We apply a uniform causal methodology across two domain granularities, three model families (1.5B to 7B parameters), and eight domains. At the academic subject level, zero neurons exceed 60% domain selectivity across 939,008 combined FFN neurons and causal damage matrices are flat, despite domain identity being linearly decodable above 85% accuracy. At the language and modality level, 0.65--1.14% of neurons exceed 60% selectivity, damage matrices are near-perfectly diagonal (ratios up to 595:1), and shell neuron sets are essentially disjoint (IoU $< 0.003$). Masking code-selective neurons reduces mathematical reasoning accuracy by 16--24 percentage points across all models; masking Spanish or Chinese neurons leaves it at or below random. Shell strength increases monotonically with scale and shells are spatially interleaved in a pattern that precludes group-level selective quantization. Parametric shells form where and only where training data was modular at the token level.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/decodable-but-not-de…] indexed:0 read:1min 2026-08-12 ·