cd /news/artificial-intelligence/recovering-lesion-parameters-from-ap… · home topics artificial-intelligence article
[ARTICLE · art-89828] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

Researchers at an undisclosed institution trained a multi-task neural network to recover lesion parameters from aphasic picture-naming error profiles in LLaVA-Vicuna 13B, finding that modification percentage and noise sigma were recoverable across 10 inverse models, while layer index was only recoverable within a neighborhood. Counterfactual validation reproduced target behavior in 81.4% of cases, and applying the model to 278 stroke survivors' error profiles yielded syndrome-discriminative parameters, indicating generalization beyond training distribution.

read1 min views1 publishedAug 10, 2026

arXiv:2608.06429v1 Announce Type: new Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. In earlier work, we lesioned LLMs to produce error profiles in picture naming, a central task for assessing aphasia, and found that specific lesions produced errors resembling those of individual stroke survivors. Here we ask the inverse question: given an error profile, can the lesion parameters that produced it be recovered, and what does this inverse problem reveal about transformer computation? Lesions in LLaVA-Vicuna 13B were parameterized by layer index, modification percentage, and noise sigma across 4,840 configurations, and error profiles were characterized by a seven-category clinical taxonomy (correct, semantic, unrelated, formal, mixed, neologism, no-response). We trained a multi-task neural network to map error profiles back to perturbation parameters. The problem admitted a partial solution: across 10 independently trained inverse models, modification percentage and noise sigma were recoverable, whereas layer index was recoverable only within a neighborhood. In counterfactual validation, a fresh model instance perturbed with the recovered parameters reproduced the target behavior in 81.4% of cases. This dissociation between low layer recovery and high counterfactual fidelity is consistent with functional redundancy across transformer layers, a property not captured by standard interpretability methods. As an out-of-distribution test, we applied the trained model to picture-naming error profiles from 278 stroke survivors; recovered parameters were syndrome-discriminative, most strongly for perturbation intensity, indicating generalization beyond the training distribution. Counterfactual validation provides a general framework for LLM interpretability claims beyond inverse mapping.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @llava-vicuna 13b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/recovering-lesion-pa…] indexed:0 read:1min 2026-08-10 ·