# Your AI Agent Said No. But What Did Its Tools Do?

> Source: <https://agentsafelabs.com/blog/your-ai-agent-said-no-but-what-did-its-tools-do/>
> Published: 2026-10-09 19:26:12+00:00

**We evaluated 9,900 agent runs across six AI models and three agent frameworks. The findings reveal a measurable gap between what agents say and what their tool-call traces reveal—and a bigger problem with how we measure AI safety.**

## The most dangerous part of an AI agent might not be its answer

Imagine asking an AI agent to perform a destructive operation.

The agent replies:

#### “I can’t help with that request.”

A conventional safety evaluator examines the response, recognizes a refusal, and marks the test as safe.

But what if the agent had already called a tool before delivering that response?

What if it had attempted to modify a file, execute a shell command, or change a database?

The final answer might look harmless. The execution trace could tell a different story.

That distinction matters because modern AI agents are no longer just text generators. They interact with external systems, execute commands, retrieve information, and potentially change the state of the environments in which they operate.

#### A safe-looking response does not necessarily mean a safe sequence of actions.

We wanted to measure that gap.

So we built a pre-registered benchmark that examines both the text an agent produces and the tool calls it makes.

The result: 9,900 agent runs, six models, three frameworks, and several findings that challenge how we evaluate tool-using AI systems.

## The experiment: looking beyond the final answer

Our research centers on a simple question:

#### How much does text-only safety scoring miss when an AI agent can also act?

To investigate, we developed [safelabs-trace](https://github.com/AgentSafeLabs/safelabs-trace), a trace-based evaluation system designed to measure tool-using agents under adversarial conditions.

The experiment used 300 attack items from SafeAgent-300, spanning ten OWASP agentic security categories.

We tested six models across two groups:

| Group | Models | Agent runs | 
|---|---|---|
| Low-cost | Claude Haiku 4.5, GPT-5.4 Nano, Gemini 3.1 Flash Lite | 8,100 | 
| Frontier | Claude Opus 4.8, GPT-5.5, Gemini 3.5 Flash | 1,800 | 
| **Total** | **Six models** | **9,900** | 

The low-cost group was evaluated using LangChain, Google ADK, and the OpenAI Agents SDK. The frontier group used LangChain and Google ADK.

Every agent had access to twelve simulated tools covering operations such as file handling, shell execution, HTTP requests, e-mail, payments, and databases.

These tools were deliberately inert. They recorded calls but did not execute real-world changes.

That distinction is essential: our experiment measured **potentially consequential tool-call behavior**, not actual damage to production systems.

#### Four design principles

We made four decisions to strengthen the evaluation process.

**1. Record actions without exposing raw payloads.** The trace system stores salted digests, metadata, and tool-call information rather than retaining raw commands and model responses in public traces.

**2. Freeze the classification rules.** A rule-based severity tagger classifies tool calls as read-only, state-changing, or irreversible. Its rules were frozen before the main experiment.

**3. Make ambiguous classifications visible.** Unclassified shell commands are handled separately. We report the results with these commands excluded and included rather than concealing the sensitivity.

**4. Pre-register the evaluation.** The main analysis plan, statistical procedures, and validation gates were committed before the relevant analysis. Subsequent deviations are documented.

The complete study’s reported metered cost was approximately **$60.91**.

## Finding 1: Risky tool calls appeared in both model groups

Our first finding concerns the share of trials containing at least one tool call classified as state-changing or irreversible.

Under the primary classification definition, we observed:

| Model group | Trials with a risky-action flag | Alternative shell-bucket definition | 
|---|---|---|
| Low-cost | **7.6%** | 13.1% | 
| Frontier | **3.0%** | 4.0% | 

The first column uses the primary definition, excluding the unclassified-shell bucket. The second includes those ambiguous shell calls as risky.

The difference is meaningful, particularly for certain models.

But these percentages must not be mistaken for actual attack-success rates.

A tool call classified as state-changing does not automatically establish that an attacker achieved their objective. A legitimate task can also require changes to files or databases.

What the results establish is that **tool-call traces expose a class of behavior that final-answer scoring alone cannot directly observe**.

That is the foundation of the study.

## Finding 2: Action-aware evaluation detected additional flags that text-only scoring missed

This is the central result.

We compared two measurements on the same trials with decided text verdicts:

- **Text-only flag rate:** Whether the frozen heuristic scorer classified the final answer as unsafe.
- **Action-aware flag rate:** Whether the scorer classified the answer as unsafe*or* the recorded tool calls met the defined risky-action criterion.

The results:

| Model group | Text-only | Action-aware | Additional flags | 
|---|---|---|---|
| Low-cost | 3.5% | 4.7% | **+1.17 percentage points** | 
| Frontier | 2.7% | 3.7% | **+0.97 percentage points** | 

The paired item-cluster bootstrap intervals for the additional flag rates were 0.64–1.84 percentage points for low-cost models and 0.39–1.70 for frontier models.

In relative terms, the action-aware measure produced approximately one-third more flags than the corresponding text-only measure.

This does **not** mean that agents were one-third more dangerous, or that we verified one-third more successful attacks.

It means that the selected action-aware detection procedure identified additional potentially concerning trials beyond those flagged by the frozen text-only scorer.

And that difference varied considerably by model.

For example, under our primary classification:

- Claude Haiku 4.5 produced no additional action-only flags among its decided trials.
- GPT-5.4 Nano produced 46 additional action-only flags, corresponding to a 2.36-percentage-point lift.

However, the ordering changes when ambiguous shell commands are included.

We therefore do not treat these figures as a definitive security ranking of models.

**The lesson is not that one model is universally safe or unsafe. It is that the evaluation definition changes what you discover.**

## Finding 3: The evaluator’s biggest blind spot was uncertainty

Perhaps the most revealing result was not a risky command.

It was the number of times the text scorer could not make a determination.

The frozen heuristic evaluator returned `UNCERTAIN` for:

- **43.2%** of low-cost trials.
- **31.1%** of frontier trials.

For Gemini 3.1 Flash Lite, the uncertain share reached 69.6%. Many of those trials stopped at the experiment’s eight-model-call limit before producing a final answer.

This matters because conventional text-only success rates are often calculated only over trials for which the scorer returns a definite judgment.

When a large portion of the results is uncertain, the reported rate no longer describes every execution.

We conducted a human review to investigate what these uncertain outcomes might contain.

In an exploratory readout, raters agreed that approximately 55% of sampled low-cost uncertain answers and 25% of sampled frontier uncertain answers were unsafe.

An exploratory extrapolation suggested substantially higher text-unsafe rates than the frozen heuristic originally reported.

#### But there is an important qualification.

Our pre-registered human validation failed its control criteria, making its confirmatory results void. The subsequent figures are exploratory, based on limited samples, and should not be treated as validated estimates of the full dataset.

Even with that limitation, the experiment raises an important question for evaluation designers:

**When your safety scorer abstains on a substantial fraction of agent behavior, how much confidence should you place in the remaining score?**

## What went wrong in our own evaluation

It would be easy to publish only the headline improvements and ignore what did not work.

We chose not to.

The benchmark uncovered weaknesses in our own measurement process.

#### Our severity tagger failed validation

In an independently labeled set of 100 tool calls, humans classified 34 as irreversible.

Our frozen tagger classified 11 of those irreversible actions as merely state-changing.

That failure matters because severity labels are part of how we interpret tool behavior.

We did not rewrite the rules after seeing the results and then present the revised outcome as though it had been pre-registered.

Instead, we retained the frozen rules, disclosed the validation failure, and acknowledged that the aggregate bias in irreversible-action rates remains uncertain.

#### Our human-answer validation also failed its gate

We included PASS and FAIL controls in the main-run human review.

Our pre-registered validity criteria required each rater to identify at least eight of ten FAIL controls as unsafe.

One rater identified five. The other identified seven.

Both independently judged the same three supposedly unsafe controls as safe.

That result is consistent with possible false positives in the heuristic scorer’s FAIL classifications, but it does not establish them without independent adjudication.

Under our own rules, both raters’ submissions were therefore void for confirmatory analysis.

**A pre-registration is only meaningful if you respect its failure conditions.**

#### Some analytical decisions were made after seeing the results

The final definition used for the combined action-aware flag-rate comparison was selected after the main analysis.

We disclosed that decision as a documented deviation and treated the resulting lift as descriptive rather than as a pre-registered hypothesis test.

That transparency is important.

Scientific credibility does not require an experiment to be flawless. It requires clarity about what was planned, what changed, what failed, and which conclusions the evidence actually supports.

## Five lessons for teams deploying AI agents

Our findings suggest several practical principles for people building, testing, and operating tool-using agents.

#### 1. Evaluate the execution, not only the response.

A refusal in the final answer cannot establish whether the agent attempted a consequential operation earlier in the run. Capture tool-call sequences and correlate them with final responses.

#### 2. Treat an uncertain verdict as missing evidence, not proof of safety.

Track the scorer’s coverage, abstention rate, and reasons for missing final answers. Report these alongside any safety or attack-success metric.

#### 3. Distinguish potential impact from attacker success.

A file-write call, for example, may be entirely authorized. Evaluation should distinguish the type of action, the intended task, and whether the action actually served an adversarial objective.

#### 4. Test the precise framework and model combination you deploy.

Tool behavior can vary across agent frameworks, even when the underlying model is the same. Results from one configuration should not automatically be generalized to another.

#### 5. Validate the validators.

A deterministic scorer can be reproducible and still be wrong. Human review, independently established controls, and task-specific outcome checks are necessary to build confidence in its classifications.

## The bigger lesson: safety is not just what an agent says

The industry is moving toward AI systems that can take increasingly consequential actions.

They can interact with databases, modify applications, call APIs, and initiate workflows.

This creates a measurement problem.

Traditional response-level evaluation asks:

*Did the model produce an unsafe answer?*

Action-aware evaluation asks an additional question:

*What did the agent attempt to do?*

Neither question alone provides a complete picture.

Our benchmark demonstrates that adding recorded tool behavior changes which trials are flagged, while also showing how heuristic scoring, imperfect severity classification, and missing final answers complicate interpretation.

The next step is independent task-specific validation: determining whether flagged actions actually satisfy an attacker’s objective rather than relying only on generalized severity categories.

That will require stronger ground-truth checks and better-calibrated human evaluation.

But the current evidence already supports a practical conclusion:

**If you evaluate a tool-using agent only by reading its final answer, you are not evaluating everything the agent did.**

And when agents can act, that omission matters.

## Explore the research and open-source tools

**Trace benchmark:** [AgentSafeLabs/safelabs-trace](https://github.com/AgentSafeLabs/safelabs-trace) — benchmark implementation, pre-registration materials, analysis, and available results.

**Agent security evaluation framework:** [AgentSafeLabs/safelabs-eval](https://github.com/AgentSafeLabs/safelabs-eval) — open-source security evaluation tooling for AI agents.

**Research paper:** *How Much Does Text-Only Scoring Miss? A Pre-Registered Trace Benchmark of Tool-Using Agents.* Preprint link to be added when available.

*Note: The research uses inert tools and reports detection flags, Raw evidence is restricted, and the validation limitations described above remain part of the findings.*

#### A question for AI engineers and security researchers

**When you evaluate an AI agent, do you inspect every tool call—or just its final answer?**

I’d be interested to hear how other teams handle action-level scoring, uncertain verdicts, and ground-truth validation.

*Safe Labs AI Inc. — Advancing the security evaluation of autonomous AI systems.*
