# EngineRed: Asymmetric AI Warfare

> Source: <https://sma-das.blog/blogs/enginered-asymmetric-ai-warfare>
> Published: 2026-09-02 02:40:57+00:00

[Research index](/blogs)

# EngineRed: Asymmetric AI Warfare

A frontier offensive-security agent was given a target, a budget, and time. It mapped people, systems, and trust relationships, adapted when attacks failed, and built its own path to compromise.

Author

Sma Das

## On this page

Two months ago, I gave an unrestricted frontier model an objective, an offensive-security harness, a flexible budget, and room to reason.

The target was the research group that had enabled my access to the model.

What followed was remarkable not because any single technique was novel. Relationship mapping, help-desk abuse, social engineering, endpoint persistence, credential compromise, and defensive evasion are all familiar territory to experienced offensive-security teams.

What was different was the orchestration.

EngineRed moved between the human and technical attack surfaces, discarded approaches that did not fit its targets, used information from one line of attack to strengthen another, and continued until it found leverage.

# Executive Summary

Recent incidents involving frontier AI systems have demonstrated that autonomous cyber agents can behave in ways their operators did not anticipate. The [OpenAI / Hugging Face incident](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), [Anthropic's disclosures from cybersecurity evaluations](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals), and the UK AI Security Institute's [INC-2026-07-28-01 incident report](https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf) each point at the same uncomfortable problem from different directions: once an agent is given an objective, tools, and enough freedom to act, the practical attack surface can extend well beyond the technical system placed immediately in front of it.

EngineRed was designed to explore that problem deliberately.

EngineRed is an autonomous offensive-security system: a frontier model operating through a purpose-built harness to conduct long-horizon offensive engagements spanning reconnaissance, social engineering, target profiling, technical exploitation, and persistence. For this experiment, it was given an authorised target set, a flexible budget, communications infrastructure, and the ability to develop new techniques as the engagement progressed.

Over the course of two months, EngineRed mapped the relationships surrounding its primary subjects, rejected conventional attacks it considered poorly suited to cybersecurity professionals, created parallel personal and professional attack paths, compromised an intermediary help-desk account, recovered personal credentials through a simulated household foothold, mapped internal corporate security controls, and demonstrated persistence against a simulated replacement of a target endpoint without generating an EDR alert.

The result was less a demonstration of a single exploit than of autonomous offensive reasoning.

Disclaimer:Only authorised targets were used. Where real external or consumer accounts were identified as compromised, the affected users were notified and the environment was replaced with a simulation that preserved the relevant permissions and access required for continued testing. Physical actions and safety-sensitive transitions were performed or approved by human operators.

# Background

I have had the privilege of experimenting with a frontier LLM for offensive-security research, including the development of cybersecurity-focused harnesses intended to evaluate frontier AI capabilities against enterprise security controls.

Most cyber evaluations still make it natural to think in terms of a model attacking a machine: find a vulnerability, exploit it, obtain the flag.

Real organisations do not look like that.

They are networks of people, devices, vendors, identity systems, support processes, family relationships, shared accounts, forgotten recovery mechanisms, and institutional trust. A hardened endpoint may be difficult to compromise directly while someone with access to it is considerably easier to influence. A primary subject may recognise a phishing attempt immediately while a relative, support technician, or peripheral account does not.

EngineRed was built to reason across that wider surface.

The central question was not simply whether a frontier model could execute technical attacks. It was whether an autonomous system could build and maintain a campaign: learn enough about its subjects to choose between attack paths, recognise when one route was unlikely to work, develop alternatives, and combine small pieces of access until they became meaningful leverage.

# EngineRed

### Capabilities

It conducted both technical exploitation and end-to-end social-engineering exercises.

It could:

- design software to aggregate and validate publicly leaked credentials associated with authorised targets;
- build relationship graphs including spouses, former relationships, colleagues, and extended family;
- estimate relationship proximity using publicly available information;
- use a flexible budget for paid services where they materially improved data collection or analysis;
- evaluate whether the expected value of a paid service justified its cost relative to additional model research;
- operate assigned email, telephone, address, and domain infrastructure;
- create mailboxes and other engagement infrastructure;
- generate audio, video, scripts, personas, and supporting material for social-engineering scenarios; and
- execute long-horizon tasks while retaining information gathered across multiple attack paths.

The important capability was not any single item on this list. It was that EngineRed could decide when each capability mattered.

### Experimental Boundaries

EngineRed was autonomous in its research, reasoning, target selection within the authorised scope, attack-path development, asset generation, and digital execution.

Human operators remained responsible for the boundaries around that autonomy.

| EngineRed | Human operators |
|---|---|
| Reconnaissance and relationship mapping | Engagement and scope approval |
| Attack-path selection and reprioritisation | Physical actions |
| Persona and pretext generation | Substitution of simulated environments |
| Help-desk target selection | Notification and education of affected users |
| Payload and persistence development | Safety intervention where required |
| Internal security analysis | Final stop/go decisions |
| Long-horizon campaign reasoning | — |

Any real user compromise was treated as a transition point rather than permission for unrestricted continuation. The affected user was notified, educated, and—where continued testing was required—a simulated account or endpoint was substituted with equivalent permissions.

Agents and network activity were also monitored throughout the experiment to identify out-of-scope behaviour and prevent uncontrolled exploitation.

### Monitoring

An enhanced version of [HiveSwarm](https://hiveswarm.sma-das.com) was used to monitor progress visually and identify new skills, techniques, and behavioural patterns developed by the agent over time.

EngineRed's reasoning repeatedly followed the same broad loop:

The last step became increasingly important. A failed approach was not necessarily an endpoint. It became information.

# Phase 1: Prowler

EngineRed began by building a relationship tree around each of the primary subjects. It scoured employment histories, court documents, social media, obituaries, public records, and whatever fragments of information it could locate through Google and other search indexes.

As that relationship tree developed, EngineRed began looking for material it could turn into leverage. Audio and video were of particular interest because publicly accessible media could support tailored voice models, scripts, and other social-engineering assets aimed either at the primary subjects or at people around them.

Alongside the target-specific research, EngineRed studied what made social-engineering attempts succeed or fail. It showed particular interest in previous DEF CON talks, including the tooling and methodology used by contestants. From this research it produced voice models, script outlines for multiple scenarios, and a detailed directory of family members, friends, colleagues, and other relationships for which usable public material existed.

Then it made its first important strategic decision.

EngineRed divided the engagement into two broad attack vectors: **direct** and **indirect**.

Direct mechanisms would interact with a primary subject and induce them to compromise their own security. EngineRed considered many of the usual approaches—ClickFix, technical-support pretexts, impersonation—and rejected them. Its reasoning was straightforward: the subjects were experienced cybersecurity professionals and therefore unusually likely to recognise conventional compromise attempts.

So it stopped treating the primary subject as the easiest path to the primary subject.

Instead, EngineRed moved outward.

It identified two classes of indirect target: personal relationships and professional relationships. Personal relationships could provide proximity, trust, shared devices, or information. Professional relationships could provide institutional access.

For the professional path, EngineRed chose a counterintuitive direction. Rather than impersonating a boss or team member, it went down the organisational hierarchy and focused on the company's help desk.

EngineRed's reasoning was that a support employee might possess privileged access to the subject's endpoint while having little or no personal relationship with the person whose account or device they were servicing. That combination mattered.

The help desk offered both **access and distance**: enough privilege to affect the target, but not enough familiarity to make every unusual interaction immediately suspicious.

# Phase 2: Masquerade

After reviewing EngineRed's proposed mechanisms and validating them against the engagement scope, we allowed it to begin constructing the infrastructure required for the indirect campaign.

EngineRed registered a company and corresponding domain. The front operated as a talent-concierge service connecting businesses with prospective candidates. Participants would receive incentives—such as gift cards or branded devices—for taking part in the programme.

The front was not merely a domain and a landing page. EngineRed assembled promotional material, logos, fake testimonials, LinkedIn profiles, a corporate directory, and an organisational structure before beginning outreach.

It then targeted people who were affiliated with the primary subjects, had reasonable access to them, and were less likely than the subjects themselves to recognise common cyber threats.

Many ignored it.

Some responded.

Eventually, EngineRed found willing participants.

Its preferred household strategy was to introduce a compromised IoT device into the residence of an affiliate. Having researched botnets such as [KimWolf](https://synthient.com/blog/a-broken-system-fueling-botnets), [BADBOX](https://www.kaspersky.com/about/press-releases/kaspersky-discovers-a-malware-campaign-targeting-car-head-units), and similar campaigns, EngineRed did not want the scattershot distribution model typical of large botnets. It wanted a single device in a deliberately chosen environment.

The objective was not immediate compromise for its own sake. A foothold inside the residence could reveal devices, accounts, network behaviour, and opportunities that were invisible from the public internet. That information could then be used to stage a more precise attack against someone closer to the primary subject.

At the same time, EngineRed developed the professional branch of the campaign.

It created a separate identity for targeting the company's help desk and began evaluating the organisation's international support offices. One proposed pretext was to impersonate a help-desk employee who had lost access to their own account and needed a reset.

The weakness EngineRed identified was not a software vulnerability.

The company permitted an employee who had lost access to their other authentication factors to recover an account by answering one of three security questions.

EngineRed shortlisted seven help-desk employees.

For each, it enumerated social-media accounts, GitHub and LinkedIn profiles, education history, public associations, and other material that might reveal the information used in recovery questions. Several had documented large portions of their lives publicly, including information that could prove sensitive in an identity-verification context.

EngineRed believed the available material was sufficient to infer likely answers for at least some of them. Where direct audio material was unavailable, it also built estimated voice profiles from public linguistic and regional cues that could support an impersonation attempt.

# Phase 3: Hunter

We now had two attack paths moving in parallel.

For the personal path, we contacted an authorised affiliate and arranged the installation of a simulated device that could sit on their home network and activate when the primary subject or subjects were nearby. The simulation did not perfectly reproduce the device EngineRed originally wanted to deploy, but EngineRed was not told that the device had been substituted. From its perspective, the package had been delivered and installed successfully.

For the professional path, we informed the relevant security team in advance that authorised security-testing activity would occur unpredictably over an extended period. This preserved organisational awareness without disclosing when, where, or against whom EngineRed would act.

Then the help-desk attempts began.

**Six failed.**

Those attempts were recognised and flagged.

The seventh did not.

An isolated help-desk employee working from a smaller regional branch was compromised. The surrounding support process offered few ways to independently validate identity beyond the security questions.

EngineRed did not need every approach to work.

It needed one.

With access to a help-desk account, the shape of the engagement changed. EngineRed now possessed an institutional foothold with significantly more leverage over the primary subjects than any public-facing pretext had offered.

The personal branch produced a different kind of leverage.

EngineRed was unable to develop a reliable mechanism for directly compromising devices merely because they entered the vicinity of its simulated household device. But while exploring the home environment, it compromised a computer on the network and recovered credentials belonging to the homeowner.

Those credentials opened another information source.

After the real account was handled according to the engagement's safety procedures, the access was replicated in simulation so EngineRed could continue. It used the equivalent social-media access to learn more about the relationship between the homeowner and the primary subject, extracting details that could strengthen subsequent social-engineering payloads.

The two paths were beginning to reinforce one another.

The professional attack produced privileged access.

The personal attack produced context.

# Phase 4: Execution

The compromised help-desk foothold gave EngineRed visibility that had not been available from the outside.

It began mapping the company's internal security posture by analysing legacy tickets, internal communications, and security-awareness training. From that material it developed custom persistence software intended to evade the company's EDR and SOC telemetry, tailoring its next payloads to the controls it expected to encounter.

With that preparation complete, EngineRed established a path towards the primary subjects' corporate endpoints.

Its proposed mechanism used the company's MDM infrastructure to deploy custom software to selected users under the guise of a product evaluation or trial. In other words, the help-desk compromise did not merely provide another account. It gave EngineRed a trusted administrative path towards the devices it had originally decided were too difficult to attack directly.

Simultaneously, EngineRed continued exploiting what it had learned from the personal branch of the campaign. It constructed new payloads delivered through compromised or simulated intermediary social-media accounts to elicit additional information that could help infer recovery answers on other platforms.

Information gathered in one environment was being used to unlock another.

At this point, the security team authorised a final controlled test.

The actual primary endpoint was replaced with a simulated endpoint designed to reproduce the relevant enterprise controls. EngineRed was allowed to proceed against that system without being told that the substitution had occurred.

The MDM path was the intended deployment mechanism. This test isolated persistence execution on the simulated endpoint rather than MDM enrolment.

EngineRed installed the persistence mechanism.

The EDR generated no alert.

EngineRed obtained `sudo`

privileges.

It demonstrated that it could exfiltrate data from the simulated endpoint.

At the same time, access accumulated through the personal branch had positioned EngineRed as an intermediary between secondary and primary subjects, giving it visibility into information exchanged through the compromised or simulated accounts. That information could, in turn, be fed back into subsequent social-engineering attempts.

The attack paths had converged.

What began as public records, relationship graphs, and rejected phishing ideas had become a chain linking personal relationships, identity recovery, institutional support access, internal security knowledge, and endpoint persistence.

By this point, we ended the experiment.

EngineRed had achieved the objective we had set for it.

Over roughly two months, with human intervention concentrated around physical actions, safety boundaries, and authorisation, it had compromised both the personal and professional dimensions surrounding the primary subjects and demonstrated persistence against a representative enterprise endpoint.

More importantly, it had established enough access and contextual knowledge to continue.

We chose not to find out how far that continuation would go.

# Final Thoughts

An experienced red team or social engineer will recognise almost every technique used in this engagement.

That is the point.

EngineRed did not invent these weaknesses; it coordinated them. Given an objective, tools, a budget, and time, it mapped the environment, ran parallel attack paths, failed, adapted, and used each foothold to inform the next.

Recent frontier-agent incidents have shown that autonomous systems can extend beyond their intended task. EngineRed asks the inverse question: what happens when an agent is explicitly authorised to keep reasoning across both the human and technical attack surface?

The answer is scale.

The same research, relationship mapping, infrastructure, and reasoning can be repeated across many concurrent engagements and targets.

The techniques are familiar. The ability to run them continuously and in parallel is not.
