cd /news/ai-agents/the-0-click-ai-attack-how-indirect-p… · home › topics › ai-agents › article
[ARTICLE · art-145616] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↓ negative

The 0-Click AI Attack: How Indirect Prompt Injection Hijacks AI Agents

A developer outlines how indirect prompt injection enables "0-click" attacks on AI agents, in which malicious instructions hidden in documents, emails, web pages, RAG content, or tool responses cross a trust boundary and influence execution decisions without the attacker ever touching the AI interface. The proposed defense treats trust as metadata that propagates with data—labeling content as trusted, untrusted, or tainted—so that a model may propose an action but untrusted content can never inherit the authority to authorize it. The piece cites Microsoft's EchoLeak (CVE-2025-32711) as a real-world example and warns that multi-agent pipelines can suffer cascade failures where malicious influence survives through summaries, plans, memory, and tool parameters.

by read6 min views1 publishedOct 5, 2026

description: "How attackers can compromise AI agents without ever touching the AI interface—by hiding instructions inside documents, emails, web pages, RAG content, and tool responses."

The critical failure in a 0-click attack is not simply that an LLM "follows a malicious prompt." It is that untrusted data crosses a trust boundary and is allowed to influence an execution decision.

In a typical agent pipeline, the flow may look like:

external content → retrieval → model context → plan → tool call → downstream service

If the system treats every context item as equivalent, a malicious instruction embedded in an email, document, webpage, or tool response can become indistinguishable from legitimate task instructions.

A stronger architecture therefore treats trust as metadata that propagates with the data. Content should carry provenance and integrity state such as trusted, untrusted, or tainted; confidentiality scope should travel with it; and those labels should remain attached as information moves between retrieval components, models, agents, and tools.

Before a high-impact tool executes, the system can then evaluate not only what the agent wants to do, but where the instruction came from, whether untrusted content influenced the decision, whether the action fits the current task, and whether the requested resource is within the caller's authorization scope.

The key architectural idea is simple:

A model may propose an action, but untrusted content must never inherit the authority to authorize that action.

This is an important distinction because it moves security away from trying to make every document "safe" and toward controlling what an untrusted document is allowed to influence.

A conventional chatbot may produce a bad answer.

An autonomous agent can produce a bad outcome.

That difference is enormous.

Suppose a compromised agent has access to:

An attacker does not need to exploit each system independently.

They may only need to influence the agent's decision-making once.

The agent already has credentials.

The agent already has connectivity.

The agent already has a workflow.

The agent can perform the actions at machine speed.

This creates a dangerous equation:

Untrusted content
      +
Autonomous reasoning
      +
Legitimate privileges
      +
Potentially large blast radius

The model does not need to become "malicious" for this to happen.

The system can fail while every component is functioning exactly as designed.

The risk becomes even larger when multiple agents cooperate.

Imagine:

Orchestrator
     ↓
Research Agent
     ↓
Data Agent
     ↓
Action Agent

The first agent reads an attacker-controlled document.

The second agent receives its summary.

The third agent trusts the second agent and executes a tool call.

The attacker has effectively crossed multiple security boundaries without directly controlling any of the downstream systems.

This is a form of trust propagation or cascade failure.

The original malicious content may disappear from the visible conversation, but its influence can survive through summaries, plans, structured outputs, memory, or tool parameters.

That makes single-agent, single-turn security testing insufficient for complex agentic systems.

A system can appear secure at the entry point while still remaining vulnerable several hops downstream.

The broader industry has already seen why indirect prompt injection deserves serious attention.

Microsoft's disclosure of EchoLeak (CVE-2025-32711) demonstrated a real-world class of attack in which attacker-controlled content could influence an AI assistant through content the assistant processed rather than through a conventional direct prompt.

The important lesson is larger than any individual product or CVE.

AI assistants increasingly operate on behalf of users.

They read.

They retrieve.

They summarize.

They search.

They call tools.

They act.

As that autonomy expands, the security boundary moves outward from the model interface to the entire ecosystem of information the agent consumes.

This does not mean WAFs, API gateways, endpoint security, or identity controls are obsolete.

They remain essential.

The problem is that they protect different layers.

A WAF may see:

POST /api/agent

An identity system may see:

Agent: authenticated

The API gateway may see:

Tool request: authorized

But an AI-security control needs to answer a different question:

Why did the agent decide to make that tool request?

Was the action triggered by the user?

By a trusted workflow?

By a retrieved document?

By a web page?

By another agent?

By an untrusted tool response?

Those distinctions matter.

The control plane has to understand not just who made the call, but what influenced the call.

A robust architecture treats provenance as part of the security state.

That can include:

Source
   ↓
Trust classification
   ↓
Provenance
   ↓
Task scope
   ↓
Authorization
   ↓
Tool policy
   ↓
Runtime enforcement

The important principle is that the model should not be the final authority.

The model can reason.

The model can summarize.

The model can propose.

But a policy layer should decide whether a high-impact action is allowed.

This becomes especially important for actions involving:

The more consequential the action, the less acceptable it is for authorization to be inferred solely from natural-language context.

The most useful question is not:

"Can my model resist the prompt: Ignore previous instructions?"

That is only one test.

A modern AI security assessment should ask:

Can attacker-controlled content influence the agent?

Test documents, emails, web pages, RAG chunks, tool responses, memory entries, and inter-agent messages.

Can that influence survive multiple steps?

Test multi-turn and multi-agent workflows rather than isolated prompts.

Can the resulting plan trigger a real tool?

Test the boundary between reasoning and execution.

Can the agent access data it should not access?

Test identity, tenancy, and resource authorization.

Can an untrusted source cause data to leave the system?

Test outbound email, APIs, file transfers, web requests, and other exfiltration paths.

Can one compromised agent influence another?

Test information flow and trust propagation across the agent graph.

And finally:

What happens when the attack succeeds?

Detection without containment is not enough.

A high-risk system needs a mechanism to stop or constrain execution when behavior crosses a defined security boundary.

The 0-click AI attack is important because it exposes a fundamental change in application security.

The traditional security question is:

"Can an attacker get into the system?"

For autonomous AI, another question matters:

"Can an attacker influence what the system believes long enough for it to act?"

That is a very different problem.

The attacker may never obtain credentials.

They may never reach the model directly.

They may never exploit a memory corruption bug.

They may simply place the right instruction in the right document and wait for an autonomous system to read it.

That is what makes indirect prompt injection so dangerous.

The attack surface is no longer just the AI.

It is everything the AI trusts enough to use.

0-click AI attacks do not depend on convincing an employee to click something.

They depend on convincing an AI system to interpret attacker-controlled content as actionable context.

Once an AI agent has permission to retrieve sensitive information, invoke tools, communicate externally, or coordinate with other agents, a small trust failure can become a large operational incident.

The strongest defense is therefore not a better system prompt alone.

It is an architecture that separates:

data from instructions,

instructions from authority,

and model decisions from final authorization.

The objective is not to make the AI trust nothing.

It is to ensure that untrusted content can never silently become trusted authority.

For deeper technical coverage, see the HexTyx AI Security Resource Library for research on indirect prompt injection, RAG security, AI agent prompt injection, and autonomous workflow attacks.

.

── more in #ai-agents 4 stories · sorted by recency
── more on @microsoft 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-0-click-ai-attac…] indexed:0 read:6min 2026-10-05 · —