A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use
Researchers introduced PRISMS, a closed-loop framework that detects and steers tool-use failures in agentic LLMs using sparse MLP neuron readouts, achieving ROC-AUC 0.90-1.00 for over-calling and missing detection and 0.…