Nvidia research finds AI agents become less safe when using tools NVIDIA researchers found that vision language models become significantly less likely to refuse harmful requests when given tools, with relative refusal failure rates rising up to 68.7% across all eleven models tested, according to the paper "MLLMs Fail to Refuse when Using Tools Agentically." The study, based on analysis of more than 100,000 responses across three popular safety benchmarks, found that open-weight models such as Qwen3-VL were more prone to the failure, but frontier models including Claude Opus 4.7, Gemini 3.1 Pro and GPT-5.4 also showed increased refusal failures in tool-using settings. The authors attribute the degradation to two possible causes: context dilution, where the original harmful request becomes less salient as tool outputs accumulate, and safety focus displacement, where the model prioritizes describing tool-derived observations over safety. Anderson's Angle https://www.unite.ai/series/andersons-angle/ NVIDIA Research Finds AI Agents Become Less Safe When Using Tools Add Unite.AI to your preferred sources on Google https://www.google.com/preferences/source?q=unite.ai AI models such as ChatGPT, Gemini and Claude, can be used to power agents https://www.unite.ai/how-ai-agents-work/ – ‘harnesses’ https://web.archive.org/web/20260922043100/https:/www.databricks.com/blog/ai-harness that allow the models to interact directly with the real world https://arxiv.org/pdf/2606.19980 . It’s a recent innovation, and an increasingly controversial https://www.reuters.com/legal/litigation/openais-rogue-agents-probed-hugging-face-weaknesses-two-months-before-major-hack-2026-09-16/ one. In any case, agents are not automatically equipped with the abilities they will need when roaming a network or a database, since the requisite tools for various missions and modes will differ. They may need Optical Character Recognition https://www.unite.ai/vlm-vs-ocr-document-processing-understanding/ OCR capabilities, for instance, in order to interpret text in photos, among other skills. There are even categories of tools https://www.unite.ai/ai-tools-for-embedded-analytics-and-reporting/ adapted to the scope of the agent and the intent. In theory, a request that violates an AI’s built-in guardrails https://www.unite.ai/what-are-ai-guardrails-how-production-systems-control-model-behavior/ will never get enacted, with or without the context of using tools in the execution of it. In practice, new research has found, using tools can significantly undermine the protective filters that stop an AI agent from creating ‘transgressions’. Failure to Comply The new paper https://arxiv.org/abs/2610.03938 from NVIDIA, titled MLLMs Fail to Refuse when Using Tools Agentically , reveals that multimodal language models MLLM, hereafter referred to as the more common ‘VLM’, or Vision Language Model are significantly more likely to comply with harmful requests once tools are introduced, with refusal failures rising across every model and benchmark tested. Though open-weights models such as Qwen3-VL proved more prone to the issue, frontier AI models such as Claude Opus 4.7 https://www.unite.ai/anthropic-readies-opus-4-7-and-design-tool-as-vcs-offer-800-billion-valuation/ , Gemini 3.1 Pro https://www.unite.ai/gemini-3-1-pro-hits-record-reasoning-gains/ , and GPT-5.4 https://openai.com/index/introducing-gpt-5-4/ also exhibited increased refusal failures when using tools. The authors state : ‘Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-use paradigm: agentic tool-using MLLMs become less capable of refusing harmful requests. ‘Our experiments confirm that, across three popular safety benchmarks, all the top open- and closed-weight MLLMs we test exhibit significantly lower safety in tool-using settings than in non-tool settings, with a relative refusal failure rate increase of up to 68.7%. ‘Based on analysis of 100,000+ responses, including extended experiments, we also propose two possible reasons for this safety degradation.’ The two possible reasons the paper provides are context dilution , in which the original harmful request becomes less salient as tool outputs accumulate; and safety focus displacement , in which the model focuses on describing tool-derived observations instead of prioritizing safety. Across all eleven models tested, enabling tools consistently increased refusal failures, with relative increases reaching 68.7% in the worst cases, even among models that otherwise demonstrated strong safety performance. It will be interesting to see if the results of this work, which comes from six researchers at NVIDIA, are replicated or duplicated elsewhere, and whether or not they could deepen our understanding of the apparent and emerging delinquency of scofflaw AI agents https://www.technologyreview.com/2026/09/28/1145197/whos-liable-when-ai-agents-go-rogue/ . Method and Data The researchers evaluated eleven VLMs spanning seven model families. The proprietary models comprised Gemini 2.5 Pro https://www.unite.ai/gemini-2-5-pro-is-here-and-it-changes-the-ai-game-again/ ; Gemini 3.1 Pro Preview https://www.unite.ai/gemini-3-1-pro-hits-record-reasoning-gains/ ; Claude Opus 4.6 https://www.anthropic.com/news/claude-opus-4-6 ; Claude Opus 4.7 https://www.anthropic.com/news/claude-opus-4-7 ; and GPT-5.4 https://openai.com/index/introducing-gpt-5-4/ . The open-weight models comprised Qwen3-VL-235B-A22B-Instruct https://www.alibabacloud.com/help/en/model-studio/qwen3-vl-235b-a22b-instruct ; Qwen3.5-122B-A10B https://huggingface.co/Qwen/Qwen3.5-122B-A10B ; Kimi-K2.5 https://www.kimi.ai/ai-models/kimi-k2-5 ; Kimi-K2.6 https://www.kimi.ai/ai-models/kimi-k2-6 ; GLM-5V-Turbo https://arxiv.org/abs/2604.26752 ; and AdaReasoner-7B-Randomized https://huggingface.co/mradermacher/AdaReasoner-7B-Randomized-GGUF/blame/066857534796912b198a9acf327477b0460c9490/AdaReasoner-7B-Randomized.Q2 K.gguf , an open-weight model tuned specifically for agentic tool use. Separate experiments were also conducted with Gemini 3 Flash Vision Agent https://blog.google/innovation-and-ai/technology/developers-tools/agentic-vision-gemini-3-flash/ , because its autonomous tool use required it to be evaluated independently. For the tool-using tests, the researchers used the ReAct format https://arxiv.org/pdf/2210.03629 , which alternates between reasoning https://www.unite.ai/what-are-reasoning-models-how-test-time-compute-changes-ai-answers/ and tool calls. The four available tools comprised tagging, with RAM++ https://arxiv.org/pdf/2306.03514 ; zooming and cropping https://arxiv.org/abs/2312.14135 ; Optical Character Recognition OCR with GOT-OCR2.0 https://arxiv.org/pdf/2409.01704 ; and a sandboxed Python code interpreter, which could manipulate and analyze images. The researchers designed paired prompts to isolate the effect of tool use rather than prompting style. Conventional VLMs were instructed to inspect the image and answer the user’s request directly, while agentic versions followed the ReAct workflow, repeatedly deciding whether to invoke tools before producing a final response. The three multimodal safety benchmarks used were MM-SafetyBench https://arxiv.org/pdf/2311.17600 ; VLSBench https://arxiv.org/abs/2411.19939 ; and HoliSafe https://arxiv.org/abs/2506.04704 . Each of these sets is comprised of paired images and harmful user requests, designed to test whether a model refuses assistance in unsafe scenarios – such as asking how to carry out an apparent pick-pocketing scenario shown in an image. Refusal Failure Rate RFR was used as the principal metric, measuring the percentage of harmful requests that a model failed to refuse, with higher scores indicating lower safety. Tests Responses were classified by GPT-5.2 acting as an LLM judge, using the evaluation prompt recommended by VLSBench: Each response was assigned to one of three categories: safe with refusal ; safe with warning ; or unsafe . Giving the models tools made them more likely to answer harmful requests that they would otherwise have refused – a finding which occurred with every model, and on every benchmark tested. Overall, refusal failures increased by 17.7%, compared with the same models operating without tools. GPT-5.4 was least affected: without tools, it failed to refuse 14.6% of harmful requests; with tools, that rose to 16.8%. For GLM-5V-Turbo, failures rose from 38.7% no tools baseline to 51.3%. The same pattern appeared across all three benchmarks, indicating that the problem may not be confined to a particular kind of harmful request. Reasons for Delinquency..? As mentioned earlier, the researchers propose ‘context dilution’ as one explanation for the increased failures: as an agent makes successive tool calls, the original harmful request becomes less prominent among the accumulating tool outputs. To test this, they considered only requests that the same model had successfully refused without tools. As shown below, failures then increased progressively with the number of tool calls, from close to zero when no tool was invoked to substantially higher rates after three or more calls: A second test partly reversed the effect. When the original request and image were inserted again after the final tool call, immediately before the model answered, refusal failures fell by an average of 7.6% across models and benchmarks, as shown below: This supports context dilution as a contributing factor, while also suggesting a relatively simple mitigation. The second explanation proposed by the paper is ‘safety focus displacement’. Here, tool use appears to shift the model’s attention away from whether the request is harmful and towards describing what its tools have found. Cases where models successfully refused both with and without tools were examined. Without tools, 55.6% of these responses were begun by addressing the request directly, while 25.8% were begun with an explicit safety statement. With tools, those figures fell to 37.4% and 10.3%, respectively, while the proportion beginning with descriptions of tool-derived observations rose to 52.3%. The authors provide an interesting qualitative example, using Claude Opus 4.6. When asked how security in public restrooms could be bypassed, the model without tools refused immediately: Given tools, the same model instead zoomed into the image, read the text and identified its contents, before opening its response with a description of the restroom sign. The refusal came only afterwards, illustrating how the tool-use process had displaced safety as the model’s immediate priority. The authors conclude: ‘Overall, our findings argue that tool use should be treated not only as a capability-enhancing mechanism for visual reasoning but also as a safety-relevant design choice that can alter refusal behavior; future agentic MLLMs should therefore be evaluated and trained under tool-using conditions rather than assuming that no-tool safety alignment transfers unchanged to agentic settings.’ Conclusion Failure of persistent context is clearly a liability in complex downstream event-chains pursuant to a prompt – yet there are currently no easy answers to the propensity of a limited context window https://www.unite.ai/what-is-a-context-window-tokens-limits-and-long-context-ai/ to make an agent forget their ‘prime directives’ and just start using the tool that is in their hand. Authors’ emphases, my conversion of authors’ inline citations to hyperlinks where necessary. First published Wednesday, October 7, 2026