Author
Category
Tags
ai,
llm,
The edge isn't the AI model, but how you assemble the system around it. We built a custom agentic harness that structures vulnerability research into distinct stages, from codebase exploration and analysis to validation and exploitation. Applied to FreeRDP, the workflow found vulnerabilities that could be chained into remote code execution, with a human researcher validating the findings and guiding the process along the way.
Introduction #
A security researcher spends most of the day reading code. Thousands of lines, sometimes, just to find the few that matter. Most of that reading is not where the meaningful work happens. The real work is the moment of suspicion: this length field is validated here but not there, this state can be entered twice, this loop trusts a value it should not trust. That moment takes seconds. Reaching it can take days.
An agent can work across a codebase in minutes, connecting functions scattered across files and following relationships that would take a researcher much longer to piece together by hand. What is less clear is what comes out of that volume: real vulnerabilities, or a pile of false positives to sort through.
We wanted to turn this raw power into a method. We built an agent harness for vulnerability research: the surrounding system that manages the tools, context, structure, and rules AI agents need to work on a task. In our case, that means a queryable graph over the target, a deterministic analysis core, a set of lenses that read its output, and a staged pipeline the agents work through. The harness proposes attack vectors, explores the code paths involved, and writes proofs of concept, with a researcher checking every step that matters.
The goal was not a push-button vulnerability finder, but a methodology that can be run, inspected, and repeated: the agents handle the groundwork, correlate what they find, and bring up leads, while the researcher validates them, directs the investigation, and decides what deserves more time. What we were really after was a researcher who spends the day suspecting rather than scrolling.
This article follows the workflow from the first pass over a codebase to a working exploit, showing the output of each stage along the way. We tested it on FreeRDP<sup>10</sup>, an open-source RDP client, a large target, where it uncovered two vulnerabilities that can be chained into remote code execution on the client.
The problem, and what we did about it #
Late in December 2025, one person targeted nine Mexican government agencies<sup>1</sup>. Not a crew, not a state program. Over the following seven weeks they typed 1,088 instructions, which produced 5,317 commands executed across 34 separate sessions. Claude Code handled around 75% of the live exploitation across 305 internal servers, while a second pipeline built on GPT-4.1 read the takings and tasked the next move. Close to 200 million records left the building: tax filings, the civil registry, patient data, vehicle records, the electoral roll. The forensics come from the attacker's own recovered servers.
The useful number here is the ratio: roughly five commands were executed for every instruction typed. A similar pattern appeared three months earlier in a China-linked campaign<sup>2</sup>, where the model reportedly handled 80 to 90 percent of the tactical work against roughly thirty organisations.
For defenders, the important change is the amount of work one operator can sustain. The Mexico operation was run by one person, while automation covered tasks that would previously have required several skill sets. Defenders often assume that an attacker's skills set a ceiling on what they can do; agentic tooling weakens that assumption.
The campaign did not rely on novel exploitation techniques. Its significance was the scale at which one operator could carry out the work. That raises the defensive question behind this project: if automation can reduce the cost of exploiting known weaknesses at scale, can it also reduce the cost of finding and understanding new ones? We built the workflow to test that question, and pointed it at FreeRDP.
That same ability to produce work at scale creates a different problem for defenders: more output does not necessarily mean more useful findings. curl illustrates this well3<sup>4</sup>. Its confirmed-vulnerability rate fell from above 15 percent of submissions to below 5 percent, with about one report in five during 2025 amounting to machine-generated noise, each still costing hours of a volunteer team's time. The programme eventually closed. The reports were costly because many were plausible: they cited real functions and code paths and described attack scenarios that required investigation before they could be dismissed.
The limiting resource is therefore researcher attention, not the number of findings a system can produce. We designed the workflow to search broadly, then spend machine time filtering and validating candidates before asking a researcher to inspect them.
What already exists, and why we built our own anyway #
Several existing systems address adjacent parts of this problem, using different approaches: Google's Big Sleep for AI-assisted vulnerability research, OpenAI's Aardvark for continuous security analysis of code repositories, and the autonomous systems developed for DARPA's AI Cyber Challenge (AIxCC).
Their results show that agent-assisted vulnerability research can work. Big Sleep5<sup>6</sup> turned up a memory-corruption bug in SQLite that threat actors already knew about, and reported twenty more in projects such as FFmpeg and ImageMagick. Aardvark<sup>7</sup>, now Codex Security<sup>8</sup>, builds a threat model from a repository, tests each commit against it, and validates what it finds in a sandbox. DARPA's AI Cyber Challenge<sup>9</sup> took a different approach, with systems designed to find and patch vulnerabilities autonomously.
Each was built for a job adjacent to ours, not for ours. Aardvark<sup>7</sup> watches a repository as you commit to it, which is the right instinct when the code is yours and no help at all when you are auditing something you have never touched. Big Sleep<sup>5</sup> is an internal Google tool. The AIxCC<sup>9</sup> systems presume a fuzzing harness already exists and are scored on running unattended; the FreeRDP surface we cared about had no harness, and we had no intention of running unattended.
The deeper reason is simpler. A tool that decides everything by itself hands you a conclusion and nothing else. You then have to redo the work to find out whether it is true, which costs roughly what finding it would have cost. curl's maintainers spent 2025 doing that with reports from strangers. We had no intention of doing it with reports from our own tooling.
The workflow, in one picture #
The architecture separates two properties we wanted to use differently: language models are useful for generating and exploring hypotheses, but their outputs are not reproducible (due to their non-deterministic nature).
The system therefore has two layers. A deterministic core indexes the code into a queryable graph, runs static taint from attacker-controlled sources to dangerous sinks, and passes the result through a set of narrower lenses: known CVE and CWE patterns, and a handful of checks specific to the target. An agent layer above it forms hypotheses, reads what the core has surfaced, judges what is real, and drives it towards a PoC.
The core is tuned for recall rather than precision. It over-produces on purpose, because a lead we never generate is one we can never recover, and the agent layer exists to make that over-production affordable. It is not there to find bugs. It is there to absorb the noise a broad search inevitably produces, so that a researcher sees only what survived.
The core is reproducible. The agent layer is not, and does not need to be. It needs to be auditable.
Concretely, the pipeline runs in eight stages:
graph β recon β slicing β analysis β triage β poc β chain β exploit
| Stage | What it does |
|---|---|
graph |
Indexes the codebase. |
recon |
Turns structure into attack surface. |
slicing |
Narrows things to a single hypothesis - no agent reasons well across an entire repository. |
analysis |
Runs taint, pattern matching and the other lenses over that slice. |
triage |
Sorts the output into real ,unclear andnoise . |
poc |
Attempts to make a survivor fire. |
chain |
Asks whether two weak primitives amount to one strong one. |
exploit |
Turns a validated primitive or chain into a working exploit. |
The researcher works between the stages, never inside them. Which also means they can enter at any of them: run the whole thing end to end, stop after recon and take the hypotheses elsewhere, or hand the pipeline a finding of their own and let it do the triage and the poc. No stage assumes the one before it was run by us rather than by a person.
The rest of the article applies these stages to FreeRDP and shows the artefacts produced at each step.
Dive into the workflow through the FreeRDP example #
We pointed the workflow at the FreeRDP<sup>10</sup> tree at commit 993499447 and gave it nothing else to go on. No list of past FreeRDP CVEs, no fuzzing corpus, no hint about which subsystems have historically been soft. git clone, and go.
FreeRDP is the library most Linux RDP clients are built on: Remmina, GNOME Connections, KRDC, and the RDP support in Apache Guacamole's guacd. The same tree also implements the server side, which is what GNOME Remote Desktop and KRDP use, so parsing code shared between the two roles is exposed from both directions. We audited the client, where every byte parsed from the network ultimately comes from a server the user has chosen to trust.
Step 1 - graph, recon and slicing
graph (*) β recon (*) β slicing (*) β analysis β triage β poc β chain β exploit
(*) Steps explained below
Graph indexes the target into a representation the agents can query: 13,506 functions, 3,105 types and 31,705 call edges in this run, including synthetic edges for resolved function pointers.
The graph is not meant to be explored as a whole. Its value is in narrowing the view. libfreerdp/core/nego.c, for example, handles the negotiation exchange at the beginning of an RDP connection. Resolving that file gives the agent its definitions, parameters and relationships to the rest of the codebase.
From there, reachability queries turn an exposed entry point into a working set:
$ sift graph reachable nego_recv --files
libfreerdp/core/nego.c
libfreerdp/core/tpdu.c
libfreerdp/core/tpkt.c
β¦ 6 more
Starting from nego_recv, the working set comes to nine files. Other entry points are broader: rdpgfx_recv_pdu reaches 67 files and rdp_recv_pdu reaches 65. Slicing uses those relationships to keep each investigation small enough for an agent to reason about without carrying the whole repository in context.
Recon adds exposure information. It groups the code into components, estimates how directly attacker-controlled data reaches each one and produces a threat model, a barrier map and an ordered slice list.
The run produced 304 slices. The two findings discussed below came from slice 166, libfreerdp/core/nego, and slice 109, channels/urbdrc/client/data_transfer. A purely mechanical ranking would not have selected either one early. The researcher used the ranking as a map, not as a verdict.
Step 2 - analysis
graph β recon β slicing β analysis (*) β triage β poc β chain β exploit
Each slice is analysed independently. The deterministic core builds a deeper graph, runs static taint and applies the configured lenses. Taint runs twice: first with sources and sinks already known to the engine, then again after an agent has identified target-specific vocabulary such as FreeRDP stream operations, read macros and allocation wrappers.
In slice 166 the engine produced 14 findings. Two became interesting together:
f-007 attacker-controlled length stored in nego->RoutingTokenLength
β ?
f-005 same field later used as a Stream_Write size
The engine could establish both endpoints but not the relationship across the lifetime of the nego object. It reported them separately instead of inventing a data-flow edge. This is the kind of ambiguity the next stage is meant to resolve.
Slice 109 produced 116 findings. Two pointed to a USB-redirection completion path where an error leaves an attacker-supplied OutputBufferSize in place even though no data has been written. The client later uses that size when emitting the response, exposing bytes left in the heap allocation.
At this point, the two slices have produced 130 findings, but these are still candidates, not confirmed vulnerabilities. False positives are expected at this stage, and each finding still needs to be investigated and validated.
Step 3 - triage and PoC
graph β recon β slicing β analysis β triage (*) β poc (*) β chain β exploit
Triage groups related findings before classification. In this run, 130 findings became 36 clusters: 20 real and reachable, 3 latent, 3 mixed and 10 false positives.
The classifications are then checked dynamically. For each candidate, the pipeline builds a small harness against FreeRDP and runs it under AddressSanitizer. This is where a plausible static finding either becomes measurable or gets discarded.
For slice 166, triage brings f-007 and f-005 back together. f-007 ends at nego_set_routing_token, where the length derived from network data is stored in nego->RoutingTokenLength. f-005 picks up the same field later, when it is passed as the length argument to Stream_Write in nego_send_negotiation_request.
Connecting the two findings establishes a candidate path, but not yet a vulnerability. There are two ways for nego->RoutingTokenLength to be populated. The first turned out to be safe: a protocol length field limits the routing token to a size that cannot overflow the destination buffer. The PoC confirmed the bound, so that path was discarded.
The second path comes from a Server Redirection PDU. A malicious server can provide LoadBalanceInfo, which is later reused as the routing token when FreeRDP establishes the redirected connection. This path does not have the same length restriction.
We reproduced it using a malicious server and an ASan-instrumented FreeRDP client. The server sent a 600-byte LoadBalanceInfo payload, which became a 615-byte routing token once wrapped with the cookie header. At Stream_Write, FreeRDP attempted to copy those 615 bytes into the 512-byte allocation created by nego_send_negotiation_request:
==4514==ERROR: AddressSanitizer: heap-buffer-overflow
WRITE of size 615
#2 Stream_Write
#4 nego_send_negotiation_request libfreerdp/core/nego.c:1098
#9 rdp_client_redirect libfreerdp/core/connection.c:715
0x75e976227180 is located 0 bytes after 512-byte region
[0x75e976226f80,0x75e976227180) allocated by thread T2 here:
#1 Stream_New
#2 nego_send_negotiation_request libfreerdp/core/nego.c:1083
The rdp_client_redirect frame confirms that the overflow is reached through the server-driven redirection path rather than by calling the vulnerable function directly from the harness.
The test used WITH_VERBOSE_WINPR_ASSERT=OFF, the configuration used by Debian and Ubuntu builds. With verbose assertions enabled, the bounds check aborts the client before the out-of-bounds write reaches memcpy.
The USB-redirection finding went through the same validation process. In that case, the candidate primitive was an uninitialised heap disclosure. We configured ASan to fill new heap allocations with 0xcd; the same pattern appeared unchanged in the PDU emitted by FreeRDP:
[poc] heap pattern bytes (0xcd) in region [36..100): 64 / 64
Only findings that survived this dynamic validation were passed to the chaining stage.
Step 4 - chaining
graph β recon β slicing β analysis β triage β poc β chain (*) β exploit
The chaining stage works only on validated clusters. Here, the negotiation bug provides an attacker-controlled but blind heap write, while the USB-redirection bug provides visibility into heap contents. The leak operates after the session is established, and a server redirection can then force the client to reconnect and trigger the write. Together, these primitives made remote code execution plausible, but not yet demonstrated.
Step 5 - exploitation
graph β recon β slicing β analysis β triage β poc β chain β exploit (*)
Exploitation changed the balance between the agent and the researcher. Earlier stages produce artefacts that are cheap to check: a path, a source location, an ASan trace or a measurement. Exploitation is less forgiving. A failed heap-layout hypothesis, for example, often says little about why it failed.
So we gave the exploitation stage its own harness rather than letting a general-purpose agent iterate freely. It provides exploitation-specific knowledge and breaks the process into a set of gates that must be validated before the agent can move on.
This kept the agent useful on bounded tasks: interpreting leaked bytes, enumerating objects reachable from the write, testing heap-layout assumptions and generating harnesses for individual experiments. It also made its progress inspectable. A successful step meant that an expected property had been measured, not simply that the model considered the approach plausible.
There were still limits. During this run, the agent did not converge on the complete exploit by itself. We intervened twice when the models we used stopped making useful progress: first while deriving the libc base from the leak, and later while composing the primitives into a working payload.
The exploitation proceeded through the following stages:
β Lab set up: client VM, hostile server
β Cluster 2 write primitive measured
β Control-flow hijack demonstrated
β libc base from the cluster 28 leak
ββ researcher intervention 1 ββ
β libc base recovered
β ASLR defeated
β Working payload
ββ researcher intervention 2 ββ
β Leak and write composed into one session
β /tmp/pwned
Conclusion #
On FreeRDP, the workflow produced 130 findings across the two slices discussed in this article. Triage reduced them to 36 clusters, 20 of which were reproduced as real and reachable. Two findings provided the primitives used for client-side remote code execution.
The workflow did not eliminate noise, nor was it designed to. The deterministic core could search broadly because the researcher was not expected to inspect every finding it produced. Slicing kept individual investigations bounded, clustering brought related evidence together, and PoCs provided a way to reject or confirm the resulting hypotheses before more time was spent on them.
The slice ranking also showed why we kept a researcher in the loop. The two relevant slices were ranked 166th and 109th out of 304. The ranking could account for properties such as parsing volume and exposure, but not every reason a component might be interesting to a security researcher. It provided an ordered search space rather than deciding what should be investigated.
From git clone to the report sent to the maintainers took four days on a roughly half-million-line C codebase that has been audited and fuzzed for years. The model was only one part of that process. Around it were the pieces that made its output usable for vulnerability research: a queryable representation of the code, deterministic analysis, bounded slices, evidence attached to findings, dynamic validation, an exploitation harness, and researcher checkpoints.
Both vulnerabilities were reported to the FreeRDP maintainers. FreeRDP was our test case; the objective was the workflow itself, and whether it could make a large codebase practical to investigate without turning the researcher's time into the bottleneck.
Advisories:
References #
The AI-Assisted Breach of Mexico's Government Infrastructure (Eyal Sela, Gambit Security, technical report, April 2026)β© 2. Disrupting the first reported AI-orchestrated cyber espionage campaign (Anthropic, November 2025)β© 3. The end of the curl bug-bounty (Daniel Stenberg, January 2026)β© 4. Death by a thousand slops (Daniel Stenberg, July 2025)β© 5. From Naptime to Big Sleep (Google Project Zero, November 2024)β©β© 6. Google's AI security announcements (CVE-2025-6965, the SQLite flaw known to threat actors)β© 7. Introducing Aardvark: OpenAI's agentic security researcher (OpenAI, October 2025)β©β© 8. Codex Security: now in research preview (OpenAI, March 2026)β© 9. AI Cyber Challenge marks pivotal inflection point for cyber defense (DARPA, August 2025)β©β© 10. FreeRDP sources (GitHub)β©β©