Beyond CI/CD: Building Agentic Validation and Inspection Layers with LangGraph, MCP, and A2A A developer proposes an agentic validation and inspection layer built with LangGraph, MCP, and A2A to complement traditional CI/CD pipelines. The architecture combines deterministic tools like SQL linters and SAST scanners with stateful reasoning agents that evaluate deployment context and risk before production releases. Modern CI/CD pipelines are exceptionally good at automation. They compile code, run tests, build artifacts, scan dependencies, deploy infrastructure, and promote releases. But there is an important difference between automating a deployment and reasoning about whether a change should be deployed. Consider a pull request containing: DROP TABLE customer transactions; Or: DELETE FROM orders; Or an application change that accidentally: removes authorization middleware, introduces an unsafe database operation, exposes a secret, changes an infrastructure security boundary, bypasses an architectural convention, deploys a breaking schema migration, or modifies production resources outside the expected scope. Traditional CI/CD can detect many of these problems when explicit rules already exist. The harder problem is determining contextual risk. Is this DROP statement part of an approved migration? Does a DELETE contain an appropriate predicate? Does the changed service still satisfy the system's architectural constraints? Does the migration preserve compatibility with applications currently running in production? Does a seemingly harmless configuration change expand the attack surface? This is where an agentic inspection layer can complement—not replace—traditional CI/CD security controls. The idea is simple: Before production deployment, introduce a stateful reasoning layer that collects evidence from deterministic tools, evaluates the change from multiple perspectives, and produces an auditable deployment decision. A practical architecture can combine: deterministic linters and security scanners, LangGraph and StateGraph, specialized validation agents, Model Context Protocol MCP , Agent2Agent A2A communication, policy engines, CI/CD status checks, and human approval for high-risk operations. This article develops that architecture from first principles. A conventional pipeline often resembles: Developer ↓ Pull Request ↓ Lint ↓ Unit Tests ↓ SAST / Dependency Scan ↓ Build ↓ Deploy Test ↓ Deploy Production This architecture is essential. NIST's Secure Software Development Framework recommends integrating secure development practices throughout the software development lifecycle rather than treating security as a final-stage activity. OWASP similarly emphasizes that CI/CD infrastructure is itself security-sensitive because pipelines frequently have access to source code, credentials, artifacts, infrastructure, and deployment environments. The problem is therefore not that traditional CI/CD is obsolete. The problem is that most controls are intentionally specialized. A SQL linter understands SQL. A SAST engine understands known code patterns and data flows. A dependency scanner understands packages and vulnerabilities. A unit test validates expected behavior. None necessarily understands the complete deployment context. That distinction motivates a second layer. An important architectural rule is: Never ask an LLM to replace a deterministic control when a deterministic control can reliably perform the job. For example, SQL syntax should normally be validated using a SQL parser or linter. SQLFluff provides dialect-aware SQL parsing and linting and can also work with templated SQL. Similarly, source-code security analysis should continue using tools such as CodeQL, Semgrep, dependency scanners, secret scanners, and language-native test frameworks. The agentic layer sits above those tools. Instead of asking: LLM → Is this SQL valid? prefer: SQL Parser ↓ SQL Linter ↓ Policy Rules ↓ Agentic Risk Analysis The first three stages generate evidence. The agent interprets that evidence within the broader deployment context. This gives us a hybrid architecture: Deterministic Controls + Contextual Agentic Reasoning + Human Governance That combination is considerably safer than treating an LLM as a universal security scanner. A useful high-level architecture is: ┌───────────────────┐ │ Pull Request │ └─────────┬─────────┘ │ ▼ ┌───────────────────┐ │ Change Detection │ └─────────┬─────────┘ │ ▼ ┌───────────────────┐ │ Gatekeeper Agent │ └─────────┬─────────┘ │ ┌────────────────┼────────────────┐ ▼ ▼ ▼ Validation Agent Security Agent Review Agent │ │ │ └────────────────┼────────────────┘ ▼ ┌───────────────────┐ │ Policy Evaluation │ └─────────┬─────────┘ │ ┌─────────┴─────────┐ ▼ ▼ APPROVE BLOCK │ │ ▼ ▼ Deployment Human Review I generally divide the system into four logical responsibilities. Gatekeeper The Gatekeeper determines what changed and what inspection path is required. Examples: Python changed → Python validation + security inspection SQL changed → SQL parser + SQL policy inspection Terraform changed → IaC security inspection Dockerfile changed → container security inspection The Gatekeeper should not blindly analyze the entire repository for every commit. Its first responsibility is scope reduction. The Validation Agent answers: Is the proposed change structurally valid? For Python, this might include: ruff check . pytest mypy . For SQL: sqlfluff lint migrations/ For infrastructure: terraform validate For containers: docker build . The important point is that the agent does not invent validation. It orchestrates established tools and normalizes their results. Conceptually: validation result = { "syntax": "PASS", "tests": "PASS", "lint": "PASS", "violations": } That becomes evidence for subsequent agents. The Security Agent focuses on exploitable or operationally dangerous changes. Its inputs may include results from: CodeQL Semgrep Secret Scanner Dependency Scanner IaC Scanner SQL Policy Engine Container Scanner For example: security result = { "critical": 0, "high": 1, "medium": 3, "secrets detected": False } The agent can then evaluate those findings alongside the actual code diff. This distinction matters. A scanner detects findings. The agent helps reason about their deployment significance. Database changes are particularly interesting because perfectly valid SQL can still be operationally catastrophic. Consider: DELETE FROM customer orders; This may be syntactically valid. But the absence of a WHERE clause can make it extremely dangerous. Likewise: DROP DATABASE production; may be syntactically valid while obviously requiring exceptional authorization. Other operations worth policy inspection include: DROP TABLE DROP SCHEMA TRUNCATE TABLE DELETE without WHERE UPDATE without WHERE ALTER TABLE DROP COLUMN GRANT REVOKE CREATE OR REPLACE But naive keyword blocking is insufficient. For example: DROP TABLE temp integration test; inside an isolated ephemeral environment may be perfectly acceptable. Therefore the inspection decision should consider: Statement + Environment + Object classification + Migration context + Permissions + Dependencies + Policy This is where agentic reasoning becomes useful. One mistake I would avoid in production is implementing SQL security entirely with expressions such as: if "DROP" in sql.upper : reject This is fragile. SQL has: comments, aliases, nested statements, stored procedures, dialect differences, templating, dynamic SQL, quoted identifiers. A stronger architecture parses SQL into an Abstract Syntax Tree. Conceptually: SQL ↓ Lexer ↓ Parser ↓ AST ↓ Policy Engine ↓ Risk Evaluation For example: DELETE ├── target: customer orders └── where: null Now the policy engine can reason structurally: if statement.type == "DELETE" and statement.where is None: risk = "CRITICAL" This is far more reliable than substring matching. A deployment gate is not naturally a single prompt. It is a workflow. LangGraph models workflows using: state, nodes, edges, conditional edges. That maps remarkably well to CI/CD inspection. Our shared state could look conceptually like: class PipelineState TypedDict : changed files: list str validation results: dict security results: dict sql results: dict architecture results: dict risk score: int decision: str Then define nodes: detect changes validate inspect security inspect sql review architecture calculate risk deployment gate And connect them: START ↓ detect changes ↓ validate ↓ inspect security ↓ inspect sql ↓ review architecture ↓ calculate risk ↓ deployment gate ↓ END But real pipelines need branching. That is where StateGraph becomes especially useful. Suppose only documentation changed. Running database inspection makes little sense. We can route dynamically: def route change state : files = state "changed files" if any f.endswith ".sql" for f in files : return "sql inspection" if any f.endswith ".py" for f in files : return "code validation" return "lightweight review" The graph then represents deployment policy explicitly. Conceptually: Change Detector │ ┌──────────┼───────────┐ ▼ ▼ ▼ SQL Python Docs │ │ │ ▼ ▼ ▼ SQL Agent Code Agent Lightweight │ │ Review └──────────┼───────────┘ ▼ Review ↓ Decision This is substantially easier to audit than hiding the entire decision process inside one enormous system prompt. The Review Agent asks a different question: Even if the change is technically valid, does it make architectural sense? Imagine a service that previously used: API ↓ Service Layer ↓ Repository ↓ Database A developer introduces: API ↓ Direct Database Query The code may: compile, pass tests, pass SQL linting, contain no obvious vulnerability. But it may violate an established architecture. A Review Agent can compare the diff against architecture policies stored in repository documentation or structured policy resources. For example: architecture/ principles.md service-boundaries.yaml database-policy.yaml security-policy.yaml The output might be: { "status": "WARN", "rule": "ARCH-DB-004", "reason": "API layer directly accesses persistence layer", "confidence": 0.94 } Critically, this should usually be treated as decision-support evidence, not unquestionable truth. Our agents need tools. The SQL agent may need: SQL parser Schema metadata Dependency graph Migration history The Security Agent may need: SAST results Dependency scanner Secret scanner Repository metadata The Review Agent may need: Architecture documents Policy repository Pull-request diff Service catalog Hard-coding every integration directly into every agent quickly becomes difficult to maintain. This is where Model Context Protocol becomes useful. MCP defines a client-server architecture through which servers can expose primitives including: tools, resources, prompts. A conceptual deployment architecture becomes: Agent │ MCP Client │ ├── Repository MCP Server ├── SQL MCP Server ├── Security MCP Server ├── Policy MCP Server └── Deployment MCP Server For example: SQL MCP Server tools: parse sql inspect schema calculate dependency impact resources: schema://production migrations://history This creates a clean separation between: Reasoning and: System Access That separation is extremely valuable in production agent architectures. This point deserves emphasis. MCP standardizes interaction. It does not magically make every exposed operation trustworthy. The MCP specification explicitly treats tool execution as security-sensitive and recommends strong user control, authorization, and data protection. Therefore a production CI/CD MCP server should enforce: Authentication Authorization Least privilege Input validation Audit logging Rate limiting Environment boundaries Credential isolation A SQL inspection agent, for example, should normally receive read-only metadata access. It should not automatically receive credentials capable of executing: DROP DATABASE production; Tool permissions must follow the principle of least privilege. MCP and A2A solve different architectural problems. A useful mental model is: MCP Agent ↔ Tools / Resources A2A Agent ↔ Agent A2A is designed for interoperability between independent agent systems. That becomes useful when validation responsibilities are separated into independently deployed services. For example: Orchestrator │ ┌──────────┼───────────┐ │ │ │ ▼ ▼ ▼ SQL Agent Security Architecture Agent Agent Each agent could be independently: deployed, versioned, secured, scaled, owned by another team, or implemented using a different agent framework. A2A provides a standardized interaction model between such agents. Now the architecture becomes more interesting. CI/CD Pipeline │ ▼ LangGraph Orchestrator │ A2A Coordination Layer ┌───────────┼───────────┐ ▼ ▼ ▼ Validation Security Review Agent Agent Agent │ │ │ └──── MCP Tool Layer ───┘ │ ┌──────────────────┼─────────────────┐ ▼ ▼ ▼ Repository MCP Security MCP SQL MCP │ │ │ ▼ ▼ ▼ Git Provider Code Scanners Database LangGraph manages workflow state. A2A handles communication between independently operating agents. MCP provides controlled access to tools and contextual resources. Traditional security tooling produces deterministic evidence. The CI/CD platform enforces the final deployment gate. Each technology therefore has a distinct responsibility. This may be the most important production rule in the article. Do not implement: if llm response == "SAFE": deploy production Instead: Deterministic Evidence ↓ Agentic Analysis ↓ Policy Engine ↓ Risk Classification ↓ Human Approval if Required ↓ CI/CD Gate The final decision should be constrained by explicit policy. For example: if critical security findings 0: decision = "BLOCK" elif destructive sql detected: decision = "HUMAN APPROVAL" elif test coverage failed: decision = "BLOCK" elif architecture risk == "HIGH": decision = "HUMAN APPROVAL" else: decision = "PASS" The LLM contributes evidence and contextual reasoning. It does not become an unrestricted production administrator. Binary pass/fail is sometimes insufficient. A better system can classify deployment risk. Example: 0–20 LOW 21–50 MEDIUM 51–75 HIGH 76–100 CRITICAL Signals might include: Critical vulnerability +50 Secret detected +50 DROP/TRUNCATE +40 DELETE without WHERE +40 Breaking schema change +30 Architecture violation +20 Insufficient tests +15 Large dependency impact +15 Then: LOW → automatic continuation MEDIUM → continue + warning HIGH → mandatory reviewer approval CRITICAL → deployment blocked The exact weights should be calibrated using organizational incidents, false-positive analysis, and risk appetite rather than copied blindly from an example. An agentic deployment gate should never simply return: DEPLOYMENT BLOCKED A useful result looks more like: { "decision": "BLOCK", "risk": "CRITICAL", "score": 92, "findings": { "file": "migrations/042 cleanup.sql", "line": 18, "type": "UNBOUNDED DELETE", "evidence": "DELETE statement contains no WHERE clause" } , "required action": "Add predicate or obtain destructive-operation approval" } This provides: evidence, location, policy, severity, remediation. That transforms an AI gate from a mysterious reviewer into an auditable engineering control. Some operations should intentionally stop automation. For example: DROP production table Major IAM change Destructive migration High-confidence architecture violation Production network change Critical scanner finding The graph can interrupt: Agent Analysis ↓ HIGH RISK ↓ Human Approval ↙ ↘ Reject Approve ↓ ↓ STOP Continue LangGraph's persistence and human-in-the-loop capabilities are particularly useful for workflows that must pause and later resume. This makes the system appropriate for controlled production environments where automation must coexist with governance. Putting everything together: Developer Push ↓ Pull Request ↓ Changed File Detection ↓ Traditional CI ┌────┼─────┐ │ │ │ Lint Test SAST └────┼─────┘ ↓ Agentic Inspection Layer ↓ LangGraph StateGraph ↓ Gatekeeper ↓ ┌───────────────┐ │ Parallel │ │ Inspection │ └───────────────┘ ↓ ↓ ↓ Validation Security SQL ↓ ↓ ↓ └──────┼──────┘ ↓ Architecture Review ↓ Policy Engine ↓ Risk Score ↓ ┌────────┼────────┐ ▼ ▼ ▼ PASS REVIEW BLOCK │ │ │ Human │ Approval │ │ └────────┘ ↓ Artifact Build ↓ Test Deployment ↓ Smoke Tests ↓ Production Gate ↓ Production A good rule of thumb: Problem Preferred mechanism Syntax Parser/compiler Formatting Linter Unit behavior Tests Known vulnerabilities SAST/SCA Secrets Secret scanner SQL structure SQL parser Policy invariants Policy engine Contextual impact Agent Architecture reasoning Agent + policy Cross-system investigation Agent Final critical approval Human/policy This separation is essential. Agentic systems become more trustworthy when they are surrounded by deterministic boundaries. If I were implementing this architecture for an enterprise pipeline, I would consider these controls non-negotiable: Agents receive minimum required permissions. Production inspection should prefer read-only access. LLM-generated SQL must never execute automatically merely because the model considers it safe. Scanner findings remain authoritative evidence rather than being silently overridden by an LLM. Critical operations require deterministic policy gates. Every agent decision should be logged with evidence and policy identifiers. Prompt injection must be considered because repository files, PR descriptions, comments, and documentation can all contain untrusted text. Tool outputs must be validated before they enter downstream workflows. MCP tools should expose narrowly scoped capabilities rather than generic shell/database access. Credentials should be short-lived and environment-scoped wherever possible. Agent outputs should use structured schemas rather than unconstrained natural language. High-risk deployment decisions should support human review. The interesting evolution is not: CI/CD → AI It is: Automation ↓ Evidence ↓ Reasoning ↓ Policy ↓ Governance ↓ Deployment Traditional CI/CD answers: Can this software be built and deployed? An agentic inspection layer adds another question: Given the available evidence, organizational policy, architecture, environment, and potential blast radius, should this particular change proceed automatically? That is a much richer engineering problem. And it is exactly why technologies such as LangGraph, MCP, and A2A become interesting in DevSecOps—not because we need an LLM inside every pipeline stage, but because complex delivery environments increasingly require stateful coordination across tools, policies, agents, and humans. Conclusion Agentic CI/CD should not replace deterministic CI/CD. It should sit above it as an inspection and decision-support layer. The architecture I would advocate is: LangGraph → stateful orchestration StateGraph → explicit nodes, state, and conditional control flow MCP → standardized access to tools and contextual resources A2A → interoperability between independently deployed agents SAST / linters / parsers / tests → deterministic technical evidence Policy engine → enforceable organizational rules Human approval → governance for high-risk operations The result is not merely a pipeline that executes faster. It is a pipeline capable of collecting evidence, understanding deployment context, escalating uncertainty, and protecting production boundaries before a risky change reaches them. That is the direction I expect serious agentic DevSecOps architectures to move toward: deterministic where possible, agentic where useful, and human-controlled where consequences matter. References NIST — Secure Software Development Framework SSDF , SP 800-218 https://csrc.nist.gov/pubs/sp/800/218/final https://csrc.nist.gov/pubs/sp/800/218/final OWASP — CI/CD Security Cheat Sheet https://cheatsheetseries.owasp.org/cheatsheets/CI CD Security Cheat Sheet.html https://cheatsheetseries.owasp.org/cheatsheets/CI CD Security Cheat Sheet.html LangChain — LangGraph Graph API / StateGraph https://docs.langchain.com/oss/python/langgraph/graph-api https://docs.langchain.com/oss/python/langgraph/graph-api LangChain — LangGraph v1 https://docs.langchain.com/oss/python/releases/langgraph-v1 https://docs.langchain.com/oss/python/releases/langgraph-v1 Model Context Protocol — Architecture https://modelcontextprotocol.io/docs/learn/architecture https://modelcontextprotocol.io/docs/learn/architecture Model Context Protocol — Specification https://modelcontextprotocol.io/specification/2025-11-25 https://modelcontextprotocol.io/specification/2025-11-25 Agent2Agent Protocol — Specification https://a2a-protocol.org/dev/specification/ https://a2a-protocol.org/dev/specification/ SQLFluff — SQL Linter Documentation https://docs.sqlfluff.com/en/stable/ https://docs.sqlfluff.com/en/stable/ GitHub — CodeQL Code Scanning Documentation https://docs.github.com/en/code-security/reference/code-scanning/codeql https://docs.github.com/en/code-security/reference/code-scanning/codeql OpenSSF — Scorecard https://openssf.org/projects/scorecard/ https://openssf.org/projects/scorecard/