cd /news/artificial-intelligence/a-few-neurons-reveal-when-llms-misus… · home topics artificial-intelligence article
[ARTICLE · art-85568] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use

Researchers introduced PRISMS, a closed-loop framework that detects and steers tool-use failures in agentic LLMs using sparse MLP neuron readouts, achieving ROC-AUC 0.90-1.00 for over-calling and missing detection and 0.86-0.90 for validity across six models from Qwen3, Llama, and Gemma families. PRISMS reduces pooled over-calling rate by 80% (from 0.131 to 0.026) and increases tool-required accuracy by 14.2 percentage points (from 0.689 to 0.831), using only 1-128 neurons per failure type.

read1 min views2 publishedAug 4, 2026

arXiv:2608.00218v1 Announce Type: new Abstract: Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools are needed (missing). We find that a small, failure-specific set of MLP neurons could distinguish such failures with linearly separable decision boundaries. Building on this observation, we introduce PRISMS (Probing Representations In Support of Monitoring and Steering), a closed-loop framework that shares a failure-specific neuron basis between sparse detection and activation steering. PRISMS selects contribution-critical MLP neurons and fits an L1-regularized detector on their activations. Across six models from the Qwen3, Llama, and Gemma families, over-calling and missing are detected at the pre-generation prompt boundary with ROC-AUC 0.90-1.00, while validity is detected from the generated tool-call span with ROC-AUC 0.86-0.90. These results are achieved with highly sparse readouts: only 1-2 MLP neurons for missing, 2-16 for over-calling, and approximately 128 for validity. These sparse detectors match or outperform dense residual-stream baselines using 23-627 times fewer features. The shared neuron basis also supports bidirectional control over tool-calling behavior, suppressing unnecessary calls and eliciting omitted ones. PRISMS therefore gates intervention on predicted failure risk to mitigate the collateral effects of unconditional steering. Across all six models, PRISMS reduces pooled over-calling rate by 80% (from 0.131 to 0.026) while increasing tool-required accuracy by 14.2 percentage points (from 0.689 to 0.831). PRISMS thus provides lightweight failure detection and selective intervention across model families.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @prisms 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-few-neurons-reveal…] indexed:0 read:1min 2026-08-04 ·