# Decepticon: Autonomous Red Team Orchestration with LangGraph

> Source: <https://dev.to/mech_app_ai/decepticon-autonomous-red-team-orchestration-with-langgraph-569b>
> Published: 2026-10-10 20:06:58+00:00

Decepticon is an autonomous red-team agent built on LangChain and LangGraph. It chains reconnaissance, exploitation, and privilege escalation tools without human intervention. The project hit GitHub Trending at #13 for Python with 5,825 stars and 1,100 forks, making it a real-world case study of agentic orchestration in a domain where tool execution must be both autonomous and auditable.

The red-team use case exposes hard problems. How do you let an agent run nmap, parse results, decide on next steps, and recover when a tool fails, all while maintaining an audit trail? Decepticon's architecture reveals how LangGraph state machines handle multi-phase adversarial workflows where tool boundaries, failure recovery, and context preservation matter.

Traditional penetration testing is a linear, human-driven process. An operator runs nmap, reads the output, picks an exploit, runs Metasploit, checks for success, and repeats. Each step depends on the previous one. The operator holds state in their head.

Autonomous agents flip this model. The agent must:

Decepticon uses LangGraph to model this as a state machine. Each node represents a phase (reconnaissance, exploitation, post-exploitation). Edges define transitions based on tool output. The graph holds state across phases, so the agent does not duplicate work or lose context when a tool fails.

LangGraph is a stateful orchestration layer on top of LangChain. It models workflows as directed graphs where nodes are functions and edges are conditional transitions. State flows through the graph, accumulating context as the agent progresses.

Decepticon's graph has three primary phases:

Each phase is a node. The agent transitions to the next phase only if the current phase succeeds. If a tool fails, the graph can backtrack, retry with different parameters, or escalate to a human operator.

The state object carries information between nodes. For Decepticon, this includes:

LangGraph persists this state in memory or a database. If the agent crashes, it can resume from the last checkpoint without re-running reconnaissance.

Edges in LangGraph can be conditional. After reconnaissance, the agent checks if any exploitable services were found. If yes, it transitions to exploitation. If no, it either retries with different nmap flags or exits.

This is where orchestration gets interesting. The agent does not blindly execute a fixed sequence. It adapts based on tool output. If Metasploit fails, the graph can route to a fallback exploit or a manual review node.

Offensive security tools are dangerous. An agent that runs arbitrary commands can cause unintended lateral movement, data loss, or compliance violations. Decepticon needs tool boundaries that let the agent operate autonomously while preventing runaway execution.

Each tool runs in a sandboxed subprocess. The agent passes input via stdin and reads output from stdout. The subprocess has a timeout. If the tool hangs, the agent kills it and logs the failure.

This prevents a single tool from blocking the entire workflow. It also isolates tool crashes. If nmap segfaults, the agent logs the error and moves to the next phase.

The agent does not execute arbitrary shell commands. It has an allowlist of approved tools (nmap, Metasploit, custom scripts). Each tool has a schema that defines valid arguments.

For example, nmap can only scan a predefined IP range. The agent cannot run `nmap -A 0.0.0.0/0` and scan the entire internet. The schema enforces constraints at runtime.

Tool output is untrusted. The agent parses it with a strict schema. If nmap returns malformed XML, the agent logs the error and retries. If Metasploit reports a shell but the session is not reachable, the agent marks the exploit as failed.

This prevents the agent from making decisions based on garbage data. It also creates an audit trail. Every tool invocation, input, and output is logged.

Red-team workflows are full of dead ends. An exploit might fail because the target is patched. A shell might die because the network is unstable. The agent needs a recovery strategy.

If a tool fails, the agent retries with exponential backoff. The first retry happens after 1 second, the second after 2 seconds, the third after 4 seconds. After three retries, the agent gives up and transitions to a fallback node.

This handles transient failures (network glitches, rate limits) without wasting time on permanent failures (patched targets, wrong exploits).

If the primary exploit fails, the agent tries a fallback. For example, if a remote code execution exploit fails, the agent tries a privilege escalation exploit. The graph defines fallback edges for each exploit.

This is where LangGraph shines. The agent does not need a hardcoded decision tree. The graph defines the fallback logic declaratively. If the primary edge fails, the graph follows the fallback edge.

Some failures require human judgment. If all exploits fail, the agent escalates to a human operator. The graph pauses, logs the state, and sends a notification. The operator reviews the logs, decides on a next step, and resumes the graph.

This hybrid model balances autonomy and safety. The agent handles routine tasks. Humans handle edge cases.

Offensive security is a regulated activity. Every action must be logged for compliance and post-mortem analysis. Decepticon's observability stack includes:

The logs are queryable. An operator can search for all nmap scans against a specific IP, or all Metasploit exploits that failed with a specific error.

LangGraph supports trace propagation. Each tool invocation gets a trace ID. The trace ID flows through the graph, linking all related actions. This makes it easy to reconstruct the attack path.

For example, if the agent escalates privileges on a target, the trace shows:

The trace is a directed acyclic graph (DAG) of tool invocations. It is the audit trail.

Decepticon runs as a long-lived service. The agent polls a queue for new targets, executes the workflow, and writes results to a database. The deployment has three components:

| Component | Role | Technology | 
|---|---|---|
| **Agent runtime** | Executes LangGraph workflows, manages state | Python, LangChain, LangGraph | 
| **Tool sandbox** | Runs offensive security tools in isolated subprocesses | Docker, systemd-nspawn | 
| **Observability backend** | Stores logs, traces, and metrics | Elasticsearch, Prometheus, Grafana | 

The agent runtime is stateless. It reads state from the database, executes the workflow, and writes state back. This makes it easy to scale horizontally. Multiple agent instances can process targets in parallel.

The tool sandbox is ephemeral. Each tool invocation gets a fresh container. After the tool exits, the container is destroyed. This prevents state leakage between invocations.

Autonomous red-team agents have sharp edges. Here are the failure modes to watch for:

If the agent's allowlist is too permissive, it can execute unintended commands. For example, if the agent can run arbitrary Python scripts, it can exfiltrate data or pivot to unintended hosts.

**Mitigation**: Strict allowlists, schema validation, and human-in-the-loop escalation for high-risk actions.

If the agent crashes mid-workflow, the state might be inconsistent. For example, the agent might log a successful exploit but fail to save the session handle.

**Mitigation**: Atomic state updates, checkpointing, and idempotent tool execution.

Offensive security tools are brittle. They crash, hang, or return malformed output. If the agent does not handle these failures gracefully, the workflow stalls.

**Mitigation**: Timeouts, retries, fallback exploits, and structured error logging.

If the agent scans or exploits unintended targets, it can violate compliance policies or legal boundaries.

**Mitigation**: IP allowlists, pre-flight validation, and audit trails.

Here is a simplified LangGraph workflow for Decepticon. It shows how the agent transitions between reconnaissance, exploitation, and post-exploitation phases.

``` python
from langgraph.graph import StateGraph, END
from typing import TypedDict

class AgentState(TypedDict):
    target_ip: str
    open_ports: list[int]
    exploit_results: dict
    session_handle: str | None

def reconnaissance(state: AgentState) -> AgentState:
    # Run nmap, parse open ports
    ports = run_nmap(state["target_ip"])
    return {**state, "open_ports": ports}

def exploitation(state: AgentState) -> AgentState:
    # Match ports to exploits, run Metasploit
    results = run_metasploit(state["open_ports"])
    return {**state, "exploit_results": results}

def post_exploitation(state: AgentState) -> AgentState:
    # Escalate privileges, pivot
    session = escalate_privileges(state["exploit_results"])
    return {**state, "session_handle": session}

def should_exploit(state: AgentState) -> str:
    return "exploitation" if state["open_ports"] else END

def should_post_exploit(state: AgentState) -> str:
    return "post_exploitation" if state["exploit_results"].get("success") else END

workflow = StateGraph(AgentState)
workflow.add_node("reconnaissance", reconnaissance)
workflow.add_node("exploitation", exploitation)
workflow.add_node("post_exploitation", post_exploitation)

workflow.set_entry_point("reconnaissance")
workflow.add_conditional_edges("reconnaissance", should_exploit)
workflow.add_conditional_edges("exploitation", should_post_exploit)
workflow.add_edge("post_exploitation", END)

app = workflow.compile()
```

This workflow runs reconnaissance, checks if any exploitable ports were found, runs exploitation if yes, and escalates privileges if the exploit succeeded. The conditional edges handle failures gracefully.

**Use Decepticon when:**

**Avoid Decepticon when:**

Decepticon is not a drop-in replacement for human penetration testers. It is an orchestration layer that automates routine tasks and escalates edge cases to humans. The LangGraph architecture is sound, but the deployment requires careful boundary enforcement, observability, and failure recovery.
