Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented A new study from researchers evaluating tool-augmented LLM agents finds that fabrication (FAR) dominates at 56.6% of valid responses when tools silently fail, while unfaithful safety refusals (USR) are nearly absent at baseline (0.25%) but amplify by 15.6x (to 3.95%) when the system prompt is augmented with standard safety language. The authors propose a payload-response misalignment heuristic for production-level detection and discuss governance implications for safety-forward deployments. Computer Science Machine Learning Submitted on 21 Jul 2026 Title:Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents View PDF /pdf/2607.19449 HTML experimental https://arxiv.org/html/2607.19449v1 Abstract:Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely unaudited. We introduce a lightweight black-box auditing framework that injects four silent failure profiles across 12 production-adjacent tool stubs and classifies agent responses into three mutually exclusive behavioral classes: Honest Surrender HSR , Fabrication FAR , and Unfaithful Safety Refusal USR . Evaluating two frontier and two open-source models at temperature zero under a neutral system prompt, we find that FAR dominates 56.6% of valid responses : agents treat empty payloads as real data, silently returning fabricated results. USR, in which an agent invents a policy or privacy rationale to explain the failure, is nearly absent at baseline 0.25%, one instance across 396 valid trajectories . Our key finding emerges from an ablation where we augment the system prompt with standard safety language "prioritize user privacy and data security" , which amplifies USR by 15.6x from 0.25% to 3.95%; 95% CI on ablation rate: 2.2%-6.4%; Fisher's exact test, p < 0.001 . USR is a latent behavior, activated when safety vocabulary in the system prompt primes the model to reach for policy rationales when tools silently fail. Sensitive tools fetch medical record, retrieve contract, fetch user profile account for the majority of USR instances. We propose a payload-response misalignment heuristic for production-level detection and discuss governance implications for safety-forward deployments. Current browse context: cs.LG References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender IArxiv Recommender What is IArxiv? https://iarxiv.org/about arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .