From AI Agents to RCE - Building a Vulnerability Research Workflow A custom agentic harness built for vulnerability research uncovered two vulnerabilities in the open-source FreeRDP RDP client that can be chained into remote code execution on the client, with a human researcher validating findings and guiding the process. The harness structures vulnerability research into distinct stages — codebase exploration, analysis, validation and exploitation — using a queryable graph over the target, a deterministic analysis core, and lenses that read its output. The article cites a December 2025 campaign in which one person typed 1,088 instructions that produced 5,317 commands across 34 sessions, with Claude Code handling around 75% of live exploitation across 305 internal servers and roughly 200 million records exfiltrated. Author Julien Lair ./author/julien-lair.html Category AI ./category/ai.html Tags ai ./tag/ai.html , agentic ./tag/agentic.html , llm ./tag/llm.html , vulnerability ./tag/vulnerability.html , 2026 ./tag/2026.html The edge isn't the AI model, but how you assemble the system around it. We built a custom agentic harness that structures vulnerability research into distinct stages, from codebase exploration and analysis to validation and exploitation. Applied to FreeRDP, the workflow found vulnerabilities that could be chained into remote code execution, with a human researcher validating the findings and guiding the process along the way. Introduction A security researcher spends most of the day reading code. Thousands of lines, sometimes, just to find the few that matter. Most of that reading is not where the meaningful work happens. The real work is the moment of suspicion : this length field is validated here but not there, this state can be entered twice, this loop trusts a value it should not trust. That moment takes seconds. Reaching it can take days. An agent can work across a codebase in minutes, connecting functions scattered across files and following relationships that would take a researcher much longer to piece together by hand. What is less clear is what comes out of that volume: real vulnerabilities, or a pile of false positives to sort through . We wanted to turn this raw power into a method. We built an agent harness for vulnerability research: the surrounding system that manages the tools, context, structure, and rules AI agents need to work on a task. In our case, that means a queryable graph over the target, a deterministic analysis core, a set of lenses that read its output, and a staged pipeline the agents work through. The harness proposes attack vectors, explores the code paths involved, and writes proofs of concept, with a researcher checking every step that matters. The goal was not a push-button vulnerability finder, but a methodology that can be run, inspected, and repeated : the agents handle the groundwork, correlate what they find, and bring up leads, while the researcher validates them, directs the investigation, and decides what deserves more time. What we were really after was a researcher who spends the day suspecting rather than scrolling . This article follows the workflow from the first pass over a codebase to a working exploit, showing the output of each stage along the way. We tested it on FreeRDP