{"slug": "decepticon-autonomous-red-team-orchestration-with-langgraph", "title": "Decepticon: Autonomous Red Team Orchestration with LangGraph", "summary": "A developer built Decepticon, an autonomous red-team agent on LangChain and LangGraph that chains reconnaissance, exploitation, and privilege escalation tools without human intervention, reaching #13 on GitHub Trending for Python with 5,825 stars and 1,100 forks. The project models penetration testing as a LangGraph state machine with conditional edges, checkpointed state, sandboxed subprocess tool execution, and an allowlist of approved tools with argument schemas. Tool output is parsed against strict schemas so the agent can backtrack, retry, or escalate to a human when a phase fails.", "body_md": "Decepticon is an autonomous red-team agent built on LangChain and LangGraph. It chains reconnaissance, exploitation, and privilege escalation tools without human intervention. The project hit GitHub Trending at #13 for Python with 5,825 stars and 1,100 forks, making it a real-world case study of agentic orchestration in a domain where tool execution must be both autonomous and auditable.\n\nThe red-team use case exposes hard problems. How do you let an agent run nmap, parse results, decide on next steps, and recover when a tool fails, all while maintaining an audit trail? Decepticon's architecture reveals how LangGraph state machines handle multi-phase adversarial workflows where tool boundaries, failure recovery, and context preservation matter.\n\nTraditional penetration testing is a linear, human-driven process. An operator runs nmap, reads the output, picks an exploit, runs Metasploit, checks for success, and repeats. Each step depends on the previous one. The operator holds state in their head.\n\nAutonomous agents flip this model. The agent must:\n\nDecepticon uses LangGraph to model this as a state machine. Each node represents a phase (reconnaissance, exploitation, post-exploitation). Edges define transitions based on tool output. The graph holds state across phases, so the agent does not duplicate work or lose context when a tool fails.\n\nLangGraph is a stateful orchestration layer on top of LangChain. It models workflows as directed graphs where nodes are functions and edges are conditional transitions. State flows through the graph, accumulating context as the agent progresses.\n\nDecepticon's graph has three primary phases:\n\nEach phase is a node. The agent transitions to the next phase only if the current phase succeeds. If a tool fails, the graph can backtrack, retry with different parameters, or escalate to a human operator.\n\nThe state object carries information between nodes. For Decepticon, this includes:\n\nLangGraph persists this state in memory or a database. If the agent crashes, it can resume from the last checkpoint without re-running reconnaissance.\n\nEdges in LangGraph can be conditional. After reconnaissance, the agent checks if any exploitable services were found. If yes, it transitions to exploitation. If no, it either retries with different nmap flags or exits.\n\nThis is where orchestration gets interesting. The agent does not blindly execute a fixed sequence. It adapts based on tool output. If Metasploit fails, the graph can route to a fallback exploit or a manual review node.\n\nOffensive security tools are dangerous. An agent that runs arbitrary commands can cause unintended lateral movement, data loss, or compliance violations. Decepticon needs tool boundaries that let the agent operate autonomously while preventing runaway execution.\n\nEach tool runs in a sandboxed subprocess. The agent passes input via stdin and reads output from stdout. The subprocess has a timeout. If the tool hangs, the agent kills it and logs the failure.\n\nThis prevents a single tool from blocking the entire workflow. It also isolates tool crashes. If nmap segfaults, the agent logs the error and moves to the next phase.\n\nThe agent does not execute arbitrary shell commands. It has an allowlist of approved tools (nmap, Metasploit, custom scripts). Each tool has a schema that defines valid arguments.\n\nFor example, nmap can only scan a predefined IP range. The agent cannot run `nmap -A 0.0.0.0/0` and scan the entire internet. The schema enforces constraints at runtime.\n\nTool output is untrusted. The agent parses it with a strict schema. If nmap returns malformed XML, the agent logs the error and retries. If Metasploit reports a shell but the session is not reachable, the agent marks the exploit as failed.\n\nThis prevents the agent from making decisions based on garbage data. It also creates an audit trail. Every tool invocation, input, and output is logged.\n\nRed-team workflows are full of dead ends. An exploit might fail because the target is patched. A shell might die because the network is unstable. The agent needs a recovery strategy.\n\nIf a tool fails, the agent retries with exponential backoff. The first retry happens after 1 second, the second after 2 seconds, the third after 4 seconds. After three retries, the agent gives up and transitions to a fallback node.\n\nThis handles transient failures (network glitches, rate limits) without wasting time on permanent failures (patched targets, wrong exploits).\n\nIf the primary exploit fails, the agent tries a fallback. For example, if a remote code execution exploit fails, the agent tries a privilege escalation exploit. The graph defines fallback edges for each exploit.\n\nThis is where LangGraph shines. The agent does not need a hardcoded decision tree. The graph defines the fallback logic declaratively. If the primary edge fails, the graph follows the fallback edge.\n\nSome failures require human judgment. If all exploits fail, the agent escalates to a human operator. The graph pauses, logs the state, and sends a notification. The operator reviews the logs, decides on a next step, and resumes the graph.\n\nThis hybrid model balances autonomy and safety. The agent handles routine tasks. Humans handle edge cases.\n\nOffensive security is a regulated activity. Every action must be logged for compliance and post-mortem analysis. Decepticon's observability stack includes:\n\nThe logs are queryable. An operator can search for all nmap scans against a specific IP, or all Metasploit exploits that failed with a specific error.\n\nLangGraph supports trace propagation. Each tool invocation gets a trace ID. The trace ID flows through the graph, linking all related actions. This makes it easy to reconstruct the attack path.\n\nFor example, if the agent escalates privileges on a target, the trace shows:\n\nThe trace is a directed acyclic graph (DAG) of tool invocations. It is the audit trail.\n\nDecepticon runs as a long-lived service. The agent polls a queue for new targets, executes the workflow, and writes results to a database. The deployment has three components:\n\n| Component | Role | Technology | \n|---|---|---|\n| **Agent runtime** | Executes LangGraph workflows, manages state | Python, LangChain, LangGraph | \n| **Tool sandbox** | Runs offensive security tools in isolated subprocesses | Docker, systemd-nspawn | \n| **Observability backend** | Stores logs, traces, and metrics | Elasticsearch, Prometheus, Grafana | \n\nThe agent runtime is stateless. It reads state from the database, executes the workflow, and writes state back. This makes it easy to scale horizontally. Multiple agent instances can process targets in parallel.\n\nThe tool sandbox is ephemeral. Each tool invocation gets a fresh container. After the tool exits, the container is destroyed. This prevents state leakage between invocations.\n\nAutonomous red-team agents have sharp edges. Here are the failure modes to watch for:\n\nIf the agent's allowlist is too permissive, it can execute unintended commands. For example, if the agent can run arbitrary Python scripts, it can exfiltrate data or pivot to unintended hosts.\n\n**Mitigation**: Strict allowlists, schema validation, and human-in-the-loop escalation for high-risk actions.\n\nIf the agent crashes mid-workflow, the state might be inconsistent. For example, the agent might log a successful exploit but fail to save the session handle.\n\n**Mitigation**: Atomic state updates, checkpointing, and idempotent tool execution.\n\nOffensive security tools are brittle. They crash, hang, or return malformed output. If the agent does not handle these failures gracefully, the workflow stalls.\n\n**Mitigation**: Timeouts, retries, fallback exploits, and structured error logging.\n\nIf the agent scans or exploits unintended targets, it can violate compliance policies or legal boundaries.\n\n**Mitigation**: IP allowlists, pre-flight validation, and audit trails.\n\nHere is a simplified LangGraph workflow for Decepticon. It shows how the agent transitions between reconnaissance, exploitation, and post-exploitation phases.\n\n``` python\nfrom langgraph.graph import StateGraph, END\nfrom typing import TypedDict\n\nclass AgentState(TypedDict):\n    target_ip: str\n    open_ports: list[int]\n    exploit_results: dict\n    session_handle: str | None\n\ndef reconnaissance(state: AgentState) -> AgentState:\n    # Run nmap, parse open ports\n    ports = run_nmap(state[\"target_ip\"])\n    return {**state, \"open_ports\": ports}\n\ndef exploitation(state: AgentState) -> AgentState:\n    # Match ports to exploits, run Metasploit\n    results = run_metasploit(state[\"open_ports\"])\n    return {**state, \"exploit_results\": results}\n\ndef post_exploitation(state: AgentState) -> AgentState:\n    # Escalate privileges, pivot\n    session = escalate_privileges(state[\"exploit_results\"])\n    return {**state, \"session_handle\": session}\n\ndef should_exploit(state: AgentState) -> str:\n    return \"exploitation\" if state[\"open_ports\"] else END\n\ndef should_post_exploit(state: AgentState) -> str:\n    return \"post_exploitation\" if state[\"exploit_results\"].get(\"success\") else END\n\nworkflow = StateGraph(AgentState)\nworkflow.add_node(\"reconnaissance\", reconnaissance)\nworkflow.add_node(\"exploitation\", exploitation)\nworkflow.add_node(\"post_exploitation\", post_exploitation)\n\nworkflow.set_entry_point(\"reconnaissance\")\nworkflow.add_conditional_edges(\"reconnaissance\", should_exploit)\nworkflow.add_conditional_edges(\"exploitation\", should_post_exploit)\nworkflow.add_edge(\"post_exploitation\", END)\n\napp = workflow.compile()\n```\n\nThis workflow runs reconnaissance, checks if any exploitable ports were found, runs exploitation if yes, and escalates privileges if the exploit succeeded. The conditional edges handle failures gracefully.\n\n**Use Decepticon when:**\n\n**Avoid Decepticon when:**\n\nDecepticon is not a drop-in replacement for human penetration testers. It is an orchestration layer that automates routine tasks and escalates edge cases to humans. The LangGraph architecture is sound, but the deployment requires careful boundary enforcement, observability, and failure recovery.", "url": "https://wpnews.pro/news/decepticon-autonomous-red-team-orchestration-with-langgraph", "canonical_source": "https://dev.to/mech_app_ai/decepticon-autonomous-red-team-orchestration-with-langgraph-569b", "published_at": "2026-10-10 20:06:58+00:00", "updated_at": "2026-10-10 20:16:17.109485+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "developer-tools"], "entities": ["Decepticon", "LangGraph", "LangChain", "GitHub", "nmap", "Metasploit"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/decepticon-autonomous-red-team-orchestration-with-langgraph", "markdown": "https://wpnews.pro/news/decepticon-autonomous-red-team-orchestration-with-langgraph.md", "text": "https://wpnews.pro/news/decepticon-autonomous-red-team-orchestration-with-langgraph.txt", "jsonld": "https://wpnews.pro/news/decepticon-autonomous-red-team-orchestration-with-langgraph.jsonld"}}