cd /news/ai-agents/agentleak-cloning-stronger-llm-agent… · home topics ai-agents article
[ARTICLE · art-125488] src=machinebrief.com ↗ pub= topic=ai-agents verified=true sentiment=↓ negative

AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing

A black-box attack called AgentLeak can clone a stronger proprietary LLM agent's task-solving capabilities onto a weaker attacker-controlled agent, improving task pass rates by over 40% versus direct skill reuse and recovering more than 80% of the victim-attacker capability gap, according to an arXiv paper (2609.07131v1). The attack exploits the skill execution gap as a leakage surface, identifying capability-critical behaviors from differences between successful victim executions and failed attacker executions and incorporating them into attacker-side skills while leaving the attacker's model, harness, and tools unchanged. Tested across 20 task scenarios comprising 600 instances, diverse agent systems, and multiple backbone models, the findings indicate that protecting explicit skill artifacts alone is insufficient because observable execution behavior leaks the procedural knowledge needed to reconstruct proprietary capabilities.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.07131v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly achieve long-horizon tasks by combining foundation models with explicit skills and implicit procedural knowledge acquired through execution. The resulting task-solving capabilities have become valuable proprietary assets, raising a new security question: can a substantially weaker attacker-controlled agent acquire the capabilities of a stronger proprietary agent through limited black-box interaction? Existing skill-stealing attacks recover explicit skill artifacts, yet we show that artifact leakage does not necessarily transfer capability: a weaker agent may possess the same skills but still fail because it lacks procedural behaviors implicitly realized by the stronger agent. Our key insight is that the skill execution gap itself forms a leakage surface, where missing behaviors are exposed through observable differences between successful victim executions and failed attacker executions. Based on this, we present AgentLeak, a black-box capability-cloning attack that identifies capability-critical behaviors from these execution differences and incorporates them into attacker-side skills, while keeping the attacker's model, harness, and tools unchanged. Across 20 task scenarios comprising 600 instances, diverse agent systems, and multiple backbone models, AgentLeak improves task pass rates by over 40% compared with direct skill reuse and recovers more than 80% of the victim--attacker capability gap. Our findings reveal a confidentiality risk in LLM agents: protecting explicit artifacts alone is insufficient, as observable execution behavior can leak the procedural knowledge required to reconstruct proprietary task-solving capabilities in low-capability and attacker-controlled agents.

── more in #ai-agents 4 stories · sorted by recency
── more on @agentleak 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agentleak-cloning-st…] indexed:0 read:1min 2026-09-10 ·