Originally published on tamiz.pro.
Autonomous AI coding agents β from code generation copilots to full-pipeline automation tools β are moving from demo to production. They can generate pull requests, modify infrastructure-as-code, update dependency manifests, and trigger deployments. The promise is radical velocity. The danger is radical blast radius. Without guardrails, a single hallucinated rm -rf / or a misconfigured IAM policy can cascade through an entire deployment pipeline. The engineering question is no longer can we automate this? but how do we automate this safely?
This article dissects the architecture, policy models, and implementation patterns for Human-in-the-Loop (HITL) Gates β the checkpoint system that sits between an autonomous AI agent and the real world, ensuring that high-risk actions require human approval while low-risk actions flow through unimpeded.
Before designing gates, you need a precise model of how autonomous agents fail. These aren't hypotheticals β they're drawn from documented incidents in AI-assisted development environments.
| Category | Description | Example | Blast Radius |
|---|---|---|---|
| Semantic Hallucination | Agent generates plausible but incorrect code/config | Wrong S3 bucket ACL, incorrect Terraform resource | MediumβHigh |
| Scope Creep | Agent modifies files outside its intended scope | Updates main.tf when asked to fix a test |
High |
| Dependency Injection | Agent introduces vulnerable or malicious dependencies | npm install of a typosquatted package |
Critical |
| Secret Leakage | Agent commits credentials or keys | Hardcoded API key in a new file | Critical |
| Infrastructure Mutation | Agent deploys config that changes production behavior | Changes max_connections in RDS |
Critical |
| Cascade Failure | A sequence of individually safe actions creates an unsafe state | 5 small config changes that together break auth | High |
The key insight: risk is contextual and cumulative. A single change to a test file is trivial. The same change applied 50 times across 50 microservices, each subtly altering behavior, is a systemic risk. Gates must account for both individual action risk and accumulated pipeline risk.
LOW IMPACT HIGH IMPACT
ββββββββββββββββββββββββββββββββββββββ
LOW PROB β LOG (no gate) β REVIEW (async gate) β
ABILITY β e.g., fix typo β e.g., config tweak β
ββββββββββββββββββββββββββββββββββββββ
HIGH PROB β LOG + MONITOR β BLOCK (sync gate) β
ABILITY β e.g., test update β e.g., prod deploy β
ββββββββββββββββββββββββββββββββββββββ
The gate system maps every action to a cell in this matrix and applies the corresponding policy: LOG (pass through, record), REVIEW (async human approval with timeout), or BLOCK (synchronous human approval required before proceeding).
The gate system is a sidecar service that intercepts actions between an AI agent and the execution environment. It operates as a middleware layer in the CI/CD pipeline, not as a replacement for the pipeline itself.
ββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ
β β β β β β
β AI Agent ββββββΆβ Gate Service ββββββΆβ Execution Env β
β (Copilot, β β (Policy Engine, β β (CI/CD Runner, β
β Codex, β β Risk Scorer, β β Terraform, β
β Custom) β β Approval Flow) β β K8s, Cloud) β
β β β β β β
ββββββββ¬ββββββββ ββββββββββ¬ββββββββββ ββββββββββ¬ββββββββββ
β β β
β βββββββββΌβββββββββ β
β β Human Approverβ β
β β (Slack, Email,β β
β β Web Console) β β
β ββββββββββββββββββ β
β β β
ββββββββββββββββββββββββ΄ββββββββββββββββββββββββββ
(Audit Log)
Action Interceptor β Captures every action the agent wants to perform. In CI/CD, this hooks into pipeline steps. In infrastructure, it wraps Terraform/CloudFormation calls.
Risk Scorer β Evaluates each action against a policy engine, producing a risk score (0β100) and a gate decision (PASS, REVIEW, BLOCK).
Approval Orchestrator β Manages the human-in-the-loop workflow: sends approval requests, handles timeouts, manages escalations.
Audit Ledger β Immutable record of every action, decision, and approval. Critical for compliance and forensics.
Rollback Coordinator β If a gated action is rejected post-execution (e.g., during async review), orchestrates automatic rollback.
Every intercepted action produces a GateDecision:
// types/gate.ts
export type GateAction = 'PASS' | 'REVIEW' | 'BLOCK' | 'ROLLBACK';
export interface GateDecision {
actionId: string;
riskScore: number; // 0-100
gateAction: GateAction;
policyViolations: PolicyViolation[];
reviewerRequired: boolean;
timeoutMs: number; // For REVIEW actions
createdAt: Date;
metadata: Record<string, unknown>;
}
export interface PolicyViolation {
ruleId: string;
severity: 'info' | 'warning' | 'critical';
description: string;
autoRemediable: boolean;
}
The distinction between REVIEW and BLOCK is critical:
The risk scorer is the brain of the gate system. It must be fast (sub-second latency for most decisions), accurate, and configurable.
The scorer evaluates each action across multiple weighted dimensions:
// risk/scorer.ts
export interface RiskProfile {
fileScope: FileScopeRisk; // Which files are touched?
changeMagnitude: ChangeMagnitude; // How large is the diff?
environment: EnvironmentRisk; // Which environment?
dependencyImpact: DependencyRisk; // New packages, version changes?
secretExposure: SecretRisk; // Potential credential leaks?
infrastructureImpact: InfraRisk; // Terraform, K8s, IAM changes?
cumulativeRisk: CumulativeRisk; // Pipeline-level accumulation
}
export interface FileScopeRisk {
touchedFiles: string[];
outOfScopeFiles: string[]; // Files the agent shouldn't touch
sensitiveFiles: string[]; // e.g., prod configs, IAM policies
}
export interface ChangeMagnitude {
linesAdded: number;
linesDeleted: number;
filesModified: number;
isBinaryChange: boolean;
}
export interface EnvironmentRisk {
targetEnvironment: 'dev' | 'staging' | 'production';
deploymentFrequency: number; // deploys per day historically
criticality: 'low' | 'medium' | 'high' | 'critical';
}
js
// risk/scorer.ts (continued)
const WEIGHTS = {
fileScope: 0.20,
changeMagnitude: 0.15,
environment: 0.25,
dependencyImpact: 0.15,
secretExposure: 0.15,
infrastructureImpact: 0.10,
cumulativeRisk: 0.00, // applied as multiplier
} as const;
export function computeRiskScore(profile: RiskProfile): number {
let score = 0;
// File scope: penalize out-of-scope and sensitive files
score += WEIGHTS.fileScope * (
profile.fileScope.outOfScopeFiles.length * 15 +
profile.fileScope.sensitiveFiles.length * 25
);
// Change magnitude: exponential scaling for large diffs
const magnitudeRaw =
profile.changeMagnitude.linesAdded +
profile.changeMagnitude.linesDeleted +
profile.changeMagnitude.filesModified * 10;
score += WEIGHTS.changeMagnitude * Math.min(100, magnitudeRaw * 0.5);
// Environment: production is inherently riskier
const envMultipliers = { dev: 1, staging: 3, production: 10 };
score += WEIGHTS.environment * envMultipliers[profile.environment.targetEnvironment] * 8;
// Dependency impact
const depChanges = profile.dependencyImpact.newPackages.length;
const versionBumps = profile.dependencyImpact.majorVersionBumps;
score += WEIGHTS.dependencyImpact * (depChanges * 20 + versionBumps * 30);
// Secret exposure
if (profile.secretExposure.detected) {
score += WEIGHTS.secretExposure * 100; // Near-instant fail
}
// Infrastructure impact
const infraChanges = profile.infrastructureImpact.terraformChanges;
score += WEIGHTS.infrastructureImpact * infraChanges * 20;
// Cumulative risk multiplier
const cumulativeMultiplier = 1 + (profile.cumulativeRisk.actionsThisPipeline / 10);
score = Math.min(100, score * cumulativeMultiplier);
return Math.round(score);
}
js
// risk/thresholds.ts
export const GATE_THRESHOLDS = {
PASS: { maxScore: 25, description: 'Auto-approve, log only' },
REVIEW: { maxScore: 60, description: 'Async human review with rollback option' },
BLOCK: { maxScore: 100, description: 'Synchronous human approval required' },
} as const;
export function decideGate(score: number): GateAction {
if (score <= GATE_THRESHOLDS.PASS.maxScore) return 'PASS';
if (score <= GATE_THRESHOLDS.REVIEW.maxScore) return 'REVIEW';
return 'BLOCK';
}
The thresholds are configurable per team, per repository, and per environment. A mature DevOps team might start strict (everything BLOCKed) and progressively relax as trust in the agent grows β a pattern we'll call Gate Decay.
The gate service is a lightweight HTTP server that acts as a proxy. The AI agent submits actions to the gate service instead of directly to the execution environment.
// server/gate-service.ts
import { FastifyInstance } from 'fastify';
import { computeRiskScore, decideGate } from '../risk/scorer';
import { PolicyEngine } from '../policy/engine';
import { ApprovalOrchestrator } from '../approval/orchestrator';
import { AuditLogger } from '../audit/logger';
import { RollbackCoordinator } from '../rollback/coordinator';
export interface GateRequest {
actionId: string;
agentId: string;
action: {
type: 'code-change' | 'deploy' | 'config-change' | 'dependency-update';
target: string; // e.g., 'https://github.com/org/repo'
description: string;
diff?: string; // Unified diff for code changes
metadata: Record<string, unknown>;
};
pipelineContext: {
pipelineId: string;
actionsSoFar: number;
previousRiskScores: number[];
environment: 'dev' | 'staging' | 'production';
};
}
export function buildGateService(fastify: FastifyInstance) {
const policyEngine = new PolicyEngine(loadDefaultPolicies());
const orchestrator = new ApprovalOrchestrator();
const auditLogger = new AuditLogger();
const rollbackCoordinator = new RollbackCoordinator();
fastify.post('/gate/evaluate', async (request, reply) => {
const req = request.body as GateRequest;
// Step 1: Build risk profile
const profile = policyEngine.buildRiskProfile(req);
// Step 2: Compute score
const score = computeRiskScore(profile);
// Step 3: Check policy violations
const violations = policyEngine.checkViolations(req, profile);
// Step 4: Decide gate action
let gateAction = decideGate(score);
// Override: any critical violation forces BLOCK
if (violations.some(v => v.severity === 'critical')) {
gateAction = 'BLOCK';
}
// Step 5: Audit
await auditLogger.log({
actionId: req.actionId,
agentId: req.agentId,
score,
gateAction,
violations,
timestamp: new Date(),
});
// Step 6: Execute gate decision
switch (gateAction) {
case 'PASS':
return reply.send({
actionId: req.actionId,
decision: 'APPROVED',
riskScore: score,
message: 'Action auto-approved. Logged for audit.',
});
case 'REVIEW':
const reviewId = await orchestrator.requestAsyncReview({
actionId: req.actionId,
agentId: req.agentId,
action: req.action,
riskScore: score,
violations,
timeoutMs: 30 * 60 * 1000, // 30 minutes
});
return reply.send({
actionId: req.actionId,
decision: 'SUBMITTED_FOR_REVIEW',
riskScore: score,
reviewId,
message: 'Action submitted for async review. Will execute if not rejected within timeout.',
});
case 'BLOCK':
const approvalId = await orchestrator.requestSyncApproval({
actionId: req.actionId,
agentId: req.agentId,
action: req.action,
riskScore: score,
violations,
timeoutMs: 24 * 60 * 60 * 1000, // 24 hours
});
return reply.send({
actionId: req.actionId,
decision: 'PENDING_APPROVAL',
riskScore: score,
approvalId,
message: 'Action blocked. Awaiting human approval.',
});
}
});
// Webhook endpoint for human approvals
fastify.post('/gate/approval/callback', async (request, reply) => {
const { approvalId, decision, reviewerId } = request.body;
const result = await orchestrator.processApproval({
approvalId,
decision, // 'approved' | 'rejected'
reviewerId,
});
if (decision === 'rejected' && result.originalAction.type === 'code-change') {
await rollbackCoordinator.initiateRollback(result.originalAction);
}
return reply.send({ status: 'processed' });
});
// Health and metrics
fastify.get('/gate/health', async () => ({
status: 'healthy',
pendingApprovals: orchestrator.pendingCount(),
pendingReviews: orchestrator.reviewCount(),
uptime: process.uptime(),
}));
}
On the AI agent side, the integration is a simple wrapper that routes all actions through the gate service:
// agent/safe-agent.ts
import axios from 'axios';
export class GatedAgent {
constructor(
private gateUrl: string,
private agentId: string,
private pipelineId: string,
private environment: 'dev' | 'staging' | 'production'
) {}
private riskHistory: number[] = [];
async submitAction(action: {
type: string;
target: string;
description: string;
diff?: string;
metadata?: Record<string, unknown>;
}): Promise<GateResponse> {
const actionId = crypto.randomUUID();
const response = await axios.post(`${this.gateUrl}/gate/evaluate`, {
actionId,
agentId: this.agentId,
action,
pipelineContext: {
pipelineId: this.pipelineId,
actionsSoFar: this.riskHistory.length,
previousRiskScores: this.riskHistory,
environment: this.environment,
},
});
const result = response.data;
this.riskHistory.push(result.riskScore);
switch (result.decision) {
case 'APPROVED':
// Execute the action directly
return this.executeAction(action);
case 'SUBMITTED_FOR_REVIEW':
// Execute now, but monitor for rejection
const executePromise = this.executeAction(action);
this.monitorForRejection(result.reviewId, action);
return executePromise;
case 'PENDING_APPROVAL':
// Wait for human approval
return this.waitForApproval(result.approvalId, action);
}
}
private async waitForApproval(
approvalId: string,
action: any
): Promise<GateResponse> {
while (true) {
await sleep(5000); // Poll every 5 seconds
const status = await axios.get(
`${this.gateUrl}/gate/approval/${approvalId}/status`
);
if (status.data.decision === 'approved') {
return this.executeAction(action);
}
if (status.data.decision === 'rejected') {
return {
decision: 'REJECTED',
message: `Action rejected by ${status.data.reviewerId}`,
};
}
}
}
private async monitorForRejection(
reviewId: string,
action: any
): Promise<void> {
const start = Date.now();
const timeout = 30 * 60 * 1000; // 30 min
while (Date.now() - start < timeout) {
await sleep(10000);
const status = await axios.get(
`${this.gateUrl}/gate/review/${reviewId}/status`
);
if (status.data.decision === 'rejected') {
await this.rollback(action);
break;
}
}
}
private async executeAction(action: any): Promise<GateResponse> {
// Dispatch to actual execution: git push, terraform apply, etc.
// Implementation depends on the execution environment
return { decision: 'EXECUTED', actionId: action.id };
}
private async rollback(action: any): Promise<void> {
// Execute rollback logic
console.log(`Rolling back action: ${action.description}`);
}
}
function sleep(ms: number) {
return new Promise(resolve => setTimeout(resolve, ms));
}
Hardcoding gate rules in the scorer is brittle. The production approach is Policy-as-Code β declarative rules stored in version control, loaded at startup, and hot-reloadable.
version: "1.0"
environment: production
policies:
- id: no-prod-config-changes
description: "AI agents cannot modify production config files"
condition:
file_patterns:
- "config/prod/**"
- "*.prod.yml"
- "*.prod.yaml"
action: BLOCK
severity: critical
reviewer_group: "platform-team"
- id: dependency-major-version-bump
description: "Major version bumps require human review"
condition:
dependency_change:
type: "major"
action: BLOCK
severity: warning
reviewer_group: "security-team"
- id: large-diff
description: "Diffs over 500 lines require review"
condition:
change_magnitude:
lines_total: { gt: 500 }
action: REVIEW
severity: info
reviewer_group: "team-leads"
- id: iam-policy-change
description: "Any IAM policy change is blocked"
condition:
file_patterns:
- "iam/**"
- "*.iam.json"
action: BLOCK
severity: critical
reviewer_group: "security-team"
- id: terraform-destroy
description: "Terraform destroy operations are always blocked"
condition:
command_contains:
- "terraform destroy"
- "terraform apply -destroy"
action: BLOCK
severity: critical
reviewer_group: "infra-team"
- id: cumulative-action-limit
description: "More than 10 actions in one pipeline triggers review"
condition:
pipeline:
max_actions: 10
action: REVIEW
severity: warning
reviewer_group: "team-leads"
- id: test-file-auto-pass
description: "Changes to test files are auto-approved"
condition:
file_patterns:
- "**/*.test.*"
- "**/*.spec.*"
- "**/__tests__/**"
action: PASS
severity: info
reviewer_group: null
js
// policy/engine.ts
import { Policy, PolicyCondition } from './types';
export class PolicyEngine {
private policies: Policy[] = [];
constructor(policies: Policy[]) {
this.policies = policies;
}
checkViolations(
request: GateRequest,
profile: RiskProfile
): PolicyViolation[] {
const violations: PolicyViolation[] = [];
for (const policy of this.policies) {
if (this.evaluateCondition(policy.condition, request, profile)) {
violations.push({
ruleId: policy.id,
severity: policy.severity,
description: policy.description,
autoRemediable: policy.autoRemediable ?? false,
});
}
}
return violations;
}
private evaluateCondition(
condition: PolicyCondition,
request: GateRequest,
profile: RiskProfile
): boolean {
// File pattern matching
if (condition.file_patterns) {
const patterns = condition.file_patterns;
const touched = profile.fileScope.touchedFiles;
return touched.some(file =>
patterns.some(p => matchGlob(p, file))
);
}
// Dependency change detection
if (condition.dependency_change) {
const change = condition.dependency_change;
if (change.type === 'major') {
return profile.dependencyImpact.majorVersionBumps > 0;
}
}
// Change magnitude
if (condition.change_magnitude) {
const mag = condition.change_magnitude;
const total =
profile.changeMagnitude.linesAdded +
profile.changeMagnitude.linesDeleted;
if (mag.lines_total?.gt && total > mag.lines_total.gt) return true;
}
// Command matching
if (condition.command_contains) {
const commands = condition.command_contains;
const actionStr = JSON.stringify(request.action);
return commands.some(cmd => actionStr.includes(cmd));
}
// Pipeline-level conditions
if (condition.pipeline) {
const pipe = condition.pipeline;
if (pipe.max_actions && request.pipelineContext.actionsSoFar >= pipe.max_actions) {
return true;
}
}
return false;
}
}
// Simple glob matcher (use minimatch in production)
function matchGlob(pattern: string, file: string): boolean {
// Simplified β use a proper library in production
if (pattern.endsWith('**')) {
return file.startsWith(pattern.slice(0, -2));
}
if (pattern.includes('*')) {
const regex = new RegExp(
'^' + pattern.replace(/\*/g, '.*').replace(/\?/g, '.') + '$'
);
return regex.test(file);
}
return pattern === file;
}
In production, policies are stored in a Git repository. A file watcher or webhook detects changes and reloads the policy engine without restarting the service:
// policy/re.ts
import { watch } from 'chokidar';
import { PolicyEngine } from './engine';
export class PolicyRe {
private engine: PolicyEngine;
private watcher: import('chokidar').FSWatcher | null = null;
constructor(engine: PolicyEngine, policyPath: string) {
this.engine = engine;
}
async start(policyDir: string) {
this.watcher = watch(`${policyDir}/*.yaml`, {
persistent: true,
});
this.watcher.on('change', async (filePath) => {
console.log(`[PolicyRe] Detected change: ${filePath}`);
const newPolicies = await loadPoliciesFromDir(policyDir);
this.engine.reload(newPolicies);
console.log(
`[PolicyRe] Reloaded ${newPolicies.length} policies`
);
});
}
stop() {
this.watcher?.close();
}
}
The gate is only as good as the human interaction. A BLOCK that takes 48 hours to get approved is worse than no gate at all β it creates bottlenecks that engineers will work around.
The orchestrator must support multiple channels:
| Channel | Latency | Best For | Implementation |
|---|---|---|---|
| Slack/Teams | Minutes | REVIEW, low-risk BLOCK | Bot integration with buttons |
| Hours | Escalation, BLOCK for critical | SMTP + tracking | |
| Web Console | Real-time | Dashboard view, batch approval | React + WebSocket |
| PagerDuty/Opsgenie | Minutes | Critical BLOCK during incidents | API integration |
When a human receives an approval request, they need enough context to make a decision quickly:
// approval/request.ts
export interface ApprovalRequest {
approvalId: string;
agentId: string;
agentName: string; // e.g., "CodeGen-Agent v2.1"
actionDescription: string; // Human-readable summary
riskScore: number;
riskLevel: 'low' | 'medium' | 'high' | 'critical';
violations: PolicyViolation[];
diffSummary: DiffSummary;
context: {
pipelineId: string;
environment: string;
branch: string;
commitHash: string;
relatedActions: ActionSummary[]; // Previous actions in this pipeline
};
recommendedAction: 'approve' | 'reject' | 'modify';
expiresAt: Date;
}
export interface DiffSummary {
filesChanged: number;
linesAdded: number;
linesDeleted: number;
keyChanges: string[]; // Top 5 most important changes
riskyPatterns: string[]; // Detected risky code patterns
}
js
// approval/slack-integration.ts
import { WebClient } from '@slack/web-api';
export class SlackApprovalChannel {
constructor(
private slackToken: string,
private defaultChannel: string
) {}
private client = new WebClient(this.slackToken);
async sendApprovalRequest(req: ApprovalRequest): Promise<string> {
const riskEmoji = {
low: 'π’',
medium: 'π‘',
high: 'π ',
critical: 'π΄',
}[req.riskLevel];
const blocks = [
{
type: 'header',
text: {
type: 'plain_text',
text: `${riskEmoji} Approval Required: ${req.actionDescription}`,
},
},
{
type: 'section',
text: {
type: 'mrkdwn',
text: [
`*Agent:* ${req.agentName}`,
`*Risk Score:* ${req.riskScore}/100 (${req.riskLevel})`,
`*Environment:* ${req.context.environment}`,
`*Pipeline:* ${req.context.pipelineId}`,
`*Expires:* ${req.expiresAt.toISOString()}`,
].join('\n'),
},
},
{
type: 'section',
text: {
type: 'mrkdwn',
text: `*Key Changes:*
${req.diffSummary.keyChanges.map(c => `β’ ${c}`).join('\n')}`,
},
},
req.violations.length > 0 && {
type: 'section',
text: {
type: 'mrkdwn',
text: `*Policy Violations:*
${req.violations.map(v => `β οΈ ${v.description}`).join('\n')}`,
},
},
{
type: 'actions',
elements: [
{
type: 'button',
text: { type: 'plain_text', text: 'β
Approve' },
style: 'primary',
value: `approve:${req.approvalId}`,
},
{
type: 'button',
text: { type: 'plain_text', text: 'β Reject' },
style: 'danger',
value: `reject:${req.approvalId}`,
},
{
type: 'button',
text: { type: 'plain_text', text: 'π View Full Diff' },
url: req.context.diffUrl,
},
],
},
].filter(Boolean);
const result = await this.client.chat.postMessage({
channel: this.defaultChannel,
blocks,
thread_ts: req.threadTs, // Thread related approvals together
});
return result.ts;
}
}
T+0min β Slack notification to primary reviewer
T+15min β Slack notification to backup reviewer
T+30min β Email notification to team lead
T+1h β PagerDuty alert (for BLOCK actions on critical resources)
T+4h β Auto-reject (configurable) + incident ticket created
This ensures that BLOCK actions don't silently stall pipelines indefinitely.
Every gate decision must be logged immutably. This is your forensic record.
// audit/logger.ts
export interface AuditEntry {
entryId: string;
timestamp: Date;
actionId: string;
agentId: string;
pipelineId: string;
action: {
type: string;
target: string;
description: string;
diffHash: string; // SHA-256 of the diff for integrity
};
riskAssessment: {
score: number;
level: 'low' | 'medium' | 'high' | 'critical';
dimensions: Record<string, number>;
};
policyCheck: {
policiesEvaluated: number;
violations: PolicyViolation[];
};
gateDecision: GateAction;
humanInteraction?: {
approvalId: string;
reviewerId: string;
reviewerName: string;
decision: 'approved' | 'rejected';
decisionTimestamp: Date;
comments: string;
};
executionResult?: {
status: 'executed' | 'failed' | 'rolled_back';
timestamp: Date;
rollbackId?: string;
};
}
Expose these metrics via Prometheus/OpenTelemetry:
agent_gate_decisions_total{gate_action="PASS"} 142
agent_gate_decisions_total{gate_action="REVIEW"} 37
agent_gate_decisions_total{gate_action="BLOCK"} 12
agent_gate_risk_score_bucket{le="25"} 142
agent_gate_risk_score_bucket{le="60"} 179
agent_gate_risk_score_bucket{le="100"} 191
agent_gate_approval_duration_seconds_bucket{le="60"} 3
agent_gate_approval_duration_seconds_bucket{le="300"} 8
agent_gate_approval_duration_seconds_bucket{le="3600"} 11
agent_gate_rejections_total 4
agent_gate_rollbacks_total 2
agent_gate_pending_approvals 3
agent_gate_policy_violations_total{rule_id="no-prod-config-changes"} 7
agent_gate_policy_violations_total{rule_id="large-diff"} 23
Multiple AI agents may submit actions simultaneously. The gate service must handle this:
// gate/concurrency.ts
import { Mutex } from 'async-mutex';
export class ConcurrencyAwareGate {
private mutexes = new Map<string, Mutex>();
async withPipelineLock(
pipelineId: string,
fn: () => Promise<any>
): Promise<any> {
let mutex = this.mutexes.get(pipelineId);
if (!mutex) {
mutex = new Mutex();
this.mutexes.set(pipelineId, mutex);
}
return mutex.runExclusive(async () => {
try {
return await fn();
} finally {
// Clean up mutex if no more actions in pipeline
if (this.isPipelineComplete(pipelineId)) {
this.mutexes.delete(pipelineId);
}
}
});
}
}
Actions must be idempotent. If the gate service crashes after logging but before executing, a retry should not double-execute:
// gate/idempotency.ts
export class IdempotencyStore {
private store: Map<string, IdempotencyRecord> = new Map();
async check(actionId: string): Promise<'new' | 'in-progress' | 'completed'> {
const record = this.store.get(actionId);
if (!record) return 'new';
if (record.status === 'completed') return 'completed';
return 'in-progress';
}
async markInProgress(actionId: string): Promise<void> {
this.store.set(actionId, {
status: 'in-progress',
startedAt: new Date(),
});
}
async markCompleted(
actionId: string,
result: any
): Promise<void> {
this.store.set(actionId, {
status: 'completed',
result,
completedAt: new Date(),
});
}
}
As agents prove themselves reliable, teams should gradually relax gates. This is Gate Decay β a deliberate, policy-driven reduction of gate strictness over time:
agent_id: "codegen-agent-v2.1"
trust_level: "building" # building β trusted β autonomous
gate_config:
code_changes: REVIEW
config_changes: BLOCK
deployments: BLOCK
max_daily_actions: 50
gate_config:
code_changes: PASS
config_changes: REVIEW
deployments: BLOCK
max_daily_actions: 200
gate_config:
code_changes: PASS
config_changes: PASS
deployments: REVIEW
max_daily_actions: 1000
always_block:
- iam-policy-change
- terraform-destroy
- secret-exposure
The gate service must be treated as a critical infrastructure component:
PASS-eligible actions, fail-open (log the gap). For BLOCK-eligible actions, fail-closed (never execute without a gate).
// gate/failover.ts
export class GateFailover {
async evaluateWithFailover(
request: GateRequest,
gateAction: GateAction
): Promise<GateDecision> {
try {
return await this.primary.evaluate(request);
} catch (error) {
console.error('[GateFailover] Primary gate failed:', error);
if (gateAction === 'BLOCK') {
// CRITICAL: Never bypass BLOCK gates
throw new Error(
`Gate service unavailable for BLOCK action. ` +
`Action ${request.actionId} will NOT be executed. ` +
`Retry when service is restored.`
);
}
// For PASS/REVIEW actions, fail-open with audit trail
return {
actionId: request.actionId,
riskScore: 0,
gateAction: 'PASS',
policyViolations: [],
reviewerRequired: false,
timeoutMs: 0,
createdAt: new Date(),
metadata: {
failover: true,
originalError: error.message,
bypassedGate: true,
},
};
}
}
}
BLOCK gates add latency by definition. The mitigation is action batching: group multiple low-risk actions and submit them as a single BLOCK request with a consolidated diff. For truly critical actions (IAM, Terraform destroy), the latency is acceptable β these should never be automated without review anyway. Consider also implementing a pre-flight mode where the agent submits a plan, gets approval for the plan, and then executes individual steps without per-step gating.
Use a shared gate service with per-agent policy profiles. A shared service gives you centralized auditing, consistent metrics, and a single deployment to maintain. Per-agent profiles (loaded via policy-as-code) handle the fact that different agents have different trust levels and risk profiles. The agentId field in every request enables per-agent metrics and policy routing.
Three strategies: (1) Gate Decay β progressively reduce gates as trust builds, so reviewers only see genuinely risky actions. (2) Randomized audit sampling β randomly select 5% of PASS decisions for post-hoc human review, catching drift without blocking the pipeline. (3) Reviewer rotation β prevent the same person from always approving, which reduces both fatigue and single-point-of-failure risk. Track approval rates per reviewer; a 99% approval rate is a signal to investigate, not celebrate.
Building safe autonomous pipelines isn't about preventing automation β it's about making automation accountable. Every gate you build is a contract between your organization and your AI agents: you'll give you freedom, you'll keep us safe. The architecture described here gives you that contract, enforced in code, audited in logs, and reviewed by humans who know the difference between a typo fix and a production outage. For more on production AI safety patterns, see Tamiz's Insights.
The conceptual model is sound, but DevOps lives or dies on implementation. Let's build a production-grade gate engine that sits between your AI agent's proposed changes and your deployment pipeline. The architecture we're targeting is straightforward: every agent-generated artifact passes through a deterministic gate that evaluates it against policy, risk scoring, and human review thresholds before anything reaches a merge queue or production environment.
from dataclasses import dataclass, field
from enum import Enum
from typing import Optional
import hashlib
import json
import time
import uuid
class GateDecision(Enum):
AUTO_APPROVE = "auto_approve"
HUMAN_REVIEW = "human_review"
BLOCK = "block"
ESCALATE = "escalate"
@dataclass
class RiskProfile:
"""Quantified risk assessment for a proposed change."""
scope_score: float = 0.0 # 0-1: how broad is the change?
blast_radius: float = 0.0 # 0-1: how many systems affected?
reversibility: float = 0.0 # 0-1: how easily can we undo?
novelty: float = 0.0 # 0-1: how different from prior changes?
confidence: float = 0.0 # 0-1: how sure is the agent?
test_coverage_delta: float = 0.0 # change in test coverage
@property
def composite_risk(self) -> float:
"""Weighted composite risk score (0-1, higher = riskier)."""
return (
0.25 * self.scope_score +
0.30 * self.blast_radius +
0.20 * (1.0 - self.reversibility) +
0.15 * self.novelty +
0.10 * (1.0 - self.confidence)
)
@dataclass
class GateResult:
decision: GateDecision
risk_profile: RiskProfile
gate_id: str
timestamp: float
metadata: dict = field(default_factory=dict)
reviewer_id: Optional[str] = None
reviewer_decision: Optional[str] = None
reviewer_notes: Optional[str] = None
class GateEngine:
"""
Deterministic gate engine that evaluates AI agent proposals
against configurable policy thresholds.
"""
def __init__(self, config: dict):
self.config = config
self.audit_log: list[dict] = []
self._review_queue: list[GateResult] = []
def evaluate(self, proposal: dict) -> GateResult:
"""
Evaluate a single agent proposal against gate policy.
Args:
proposal: Dict containing:
- change_id: Unique identifier
- diff_stats: {files_changed, lines_added, lines_removed}
- affected_services: List of service names
- test_results: {passed, failed, coverage_before, coverage_after}
- agent_confidence: 0-1 confidence score
- change_type: 'refactor' | 'feature' | 'bugfix' | 'hotfix'
- dependencies: List of dependency changes
- is_security_related: bool
"""
gate_id = str(uuid.uuid4())
risk = self._compute_risk(proposal)
decision = self._apply_policy(risk, proposal)
result = GateResult(
decision=decision,
risk_profile=risk,
gate_id=gate_id,
timestamp=time.time(),
metadata={
"change_id": proposal.get("change_id"),
"change_type": proposal.get("change_type"),
"diff_stats": proposal.get("diff_stats"),
}
)
self._audit(result)
if decision == GateDecision.HUMAN_REVIEW:
self._review_queue.append(result)
return result
def _compute_risk(self, proposal: dict) -> RiskProfile:
"""Compute risk profile from proposal metadata."""
diff_stats = proposal.get("diff_stats", {})
test_results = proposal.get("test_results", {})
files_changed = diff_stats.get("files_changed", 0)
lines_changed = diff_stats.get("lines_added", 0) + diff_stats.get("lines_removed", 0)
affected = len(proposal.get("affected_services", []))
deps_changed = len(proposal.get("dependencies", []))
scope_score = min(1.0, (files_changed / 20) + (lines_changed / 500))
blast_radius = min(1.0, (affected / 5) + (deps_changed / 3))
reversibility = 1.0
if proposal.get("is_security_related"):
reversibility -= 0.3
if deps_changed > 0:
reversibility -= 0.2
if test_results.get("failed", 0) > 0:
reversibility -= 0.2
reversibility = max(0.0, reversibility)
novelty = min(1.0, files_changed / 15 + deps_changed / 5)
cov_before = test_results.get("coverage_before", 0.0)
cov_after = test_results.get("coverage_after", 0.0)
cov_delta = cov_after - cov_before
return RiskProfile(
scope_score=scope_score,
blast_radius=blast_radius,
reversibility=reversibility,
novelty=novelty,
confidence=proposal.get("agent_confidence", 0.5),
test_coverage_delta=cov_delta,
)
def _apply_policy(self, risk: RiskProfile, proposal: dict) -> GateDecision:
"""
Apply deterministic policy rules. These are NOT learned β
they are explicit, auditable, and versioned.
"""
if proposal.get("is_security_related") and risk.blast_radius > 0.5:
return GateDecision.BLOCK
if risk.composite_risk > self.config.get("hard_block_threshold", 0.8):
return GateDecision.BLOCK
if proposal.get("test_results", {}).get("failed", 0) > 0:
return GateDecision.BLOCK
if risk.composite_risk > self.config.get("escalation_threshold", 0.6):
return GateDecision.ESCALATE
if risk.composite_risk > self.config.get("review_threshold", 0.4):
return GateDecision.HUMAN_REVIEW
if risk.test_coverage_delta < self.config.get("min_coverage_delta", -0.02):
return GateDecision.HUMAN_REVIEW
if risk.composite_risk <= self.config.get("auto_approve_threshold", 0.3):
return GateDecision.AUTO_APPROVE
return GateDecision.HUMAN_REVIEW
def _audit(self, result: GateResult):
"""Append to immutable audit log."""
entry = {
"gate_id": result.gate_id,
"decision": result.decision.value,
"composite_risk": result.risk_profile.composite_risk,
"timestamp": result.timestamp,
"metadata": result.metadata,
}
self.audit_log.append(entry)
The gate engine's power comes from its policy being explicit and versioned. Store it alongside your infrastructure code:
version: "2.3.1"
last_updated: "2025-01-15"
approved_by: "platform-security-team"
thresholds:
auto_approve: 0.30 # Below this: machine approves
review: 0.40 # Between review and escalation: human reviews
escalation: 0.60 # Between escalation and hard_block: senior review
hard_block: 0.80 # Above this: never auto-approve
coverage:
min_coverage_delta: -0.02 # Allow 2% regression before requiring review
require_green_tests: true
overrides:
hotfix:
review_threshold: 0.55
max_files_changed: 10
require: ["oncall-approval"]
security:
always_review: true
require: ["security-team-approval"]
migrations:
review_threshold: 0.25
require: ["db-team-approval"]
max_tables_affected: 3
review_routing:
default_reviewers: ["platform-team"]
max_review_time_hours: 4
escalation_after_hours: 2
backup_reviewers: ["senior-engineers-pool"]
audit:
retention_days: 365
export_format: "jsonl"
destination: "s3://org-audit-logs/gate-decisions/"
The gate engine doesn't live in isolation β it hooks into your existing pipeline. Here's how to wire it into a GitHub Actions workflow:
name: Agent Gate Evaluation
on:
pull_request:
types: [opened, synchronize, reopened]
jobs:
gate-evaluation:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v4
- name: Run Gate Evaluation
id: gate
run: |
python scripts/run_gate.py \
--proposal "${{ github.event.pull_request.number }}" \
--policy gates/policy.yaml \
--output gate-result.json
- name: Process Gate Decision
run: |
DECISION=$(jq -r '.decision' gate-result.json)
RISK=$(jq -r '.composite_risk' gate-result.json)
GATE_ID=$(jq -r '.gate_id' gate-result.json)
case "$DECISION" in
"auto_approve")
echo "β
Auto-approved (risk: $RISK)"
echo "GATE_STATUS=approved" >> $GITHUB_ENV
;;
"human_review")
echo "π Requires human review (risk: $RISK)"
echo "GATE_STATUS=needs_review" >> $GITHUB_ENV
python scripts/request_review.py \
--gate-id "$GATE_ID" \
--pr "${{ github.event.pull_request.number }}" \
--risk "$RISK"
exit 1
;;
"escalate")
echo "β οΈ Escalated for senior review (risk: $RISK)"
python scripts/escalate.py \
--gate-id "$GATE_ID" \
--pr "${{ github.event.pull_request.number }}"
exit 1
;;
"block")
echo "π« Blocked by policy (risk: $RISK)"
python scripts/block_pr.py \
--gate-id "$GATE_ID" \
--pr "${{ github.event.pull_request.number }}" \
--reason "Exceeds hard block threshold"
exit 1
;;
esac
- name: Record Audit Entry
if: always()
run: |
python scripts/audit.py \
--gate-result gate-result.json \
--destination "s3://org-audit-logs/gate-decisions/"
A gate is only as effective as the quality of human review it enables. The review interface must be purpose-built: showing reviewers exactly what they need to decide, in the minimum time required.
@dataclass
class ReviewPayload:
"""
What a human reviewer sees. Optimized for 2-minute decisions
on low-risk items and 15-minute decisions on complex items.
"""
change_title: str
change_type: str
risk_score: float
risk_factors: list[str] # "3 services affected", "new dependency"
recommended_action: str # "Approve", "Request changes", "Reject"
diff_summary: str # "12 files, +87 lines, -23 lines"
test_results: str # "All 342 tests passing, coverage: 87% β 88%"
affected_services: list[str]
agent_rationale: str # Why the agent made this change
full_diff: str
agent_confidence_breakdown: dict
similar_past_changes: list[dict] # "3 similar changes approved in last 30d"
approve: callable
request_changes: callable
reject: callable
escalate: callable
Not all human reviews are equal. The interface should adapt to the risk level:
class ReviewWorkflow:
"""
Adapts the review experience based on risk tier.
"""
def get_review_mode(self, risk: RiskProfile) -> str:
if risk.composite_risk < 0.4:
return "quick_approve" # 1-click with diff preview
elif risk.composite_risk < 0.6:
return "standard_review" # Full diff + test results
elif risk.composite_risk < 0.8:
return "deep_review" # Context, history, architecture impact
else:
return "adversarial_review" # Second reviewer + architecture review
def build_quick_approve_payload(self, result: GateResult) -> dict:
"""
For low-risk changes: show the reviewer everything they need
to approve in under 60 seconds.
"""
return {
"mode": "quick_approve",
"summary": self._summarize_change(result),
"risk_factors": self._extract_risk_factors(result.risk_profile),
"diff_preview": self._get_diff_preview(result, max_lines=50),
"test_status": "β
All passing",
"quick_actions": ["Approve", "Flag for deeper review"],
"confidence_note": (
f"Agent confidence: {result.risk_profile.confidence:.0%} | "
f"Risk score: {result.risk_profile.composite_risk:.2f}"
),
}
def build_adversarial_review_payload(self, result: GateResult) -> dict:
"""
For high-risk changes: structured adversarial review.
Reviewer is prompted to actively try to find problems.
"""
return {
"mode": "adversarial_review",
"summary": self._summarize_change(result),
"risk_factors": self._extract_risk_factors(result.risk_profile),
"questions": [
"Could this change cause a production outage?",
"Is there a safer way to achieve the same goal?",
"Are the tests actually testing the right behavior?",
"What happens if this runs during peak traffic?",
"Can we roll this back in under 5 minutes?",
],
"required_findings": [
"Rollback procedure verified",
"Monitoring alerts identified",
"Blast radius confirmed",
],
"second_reviewer_required": True,
"architecture_review_required": True,
"full_diff": self._get_full_diff(result),
}
A gate system without observability is a black box. You need to answer: Is the gate blocking too much? Too little? Are reviewers rubber-stamping? Is the risk model calibrated?
class GateMetrics:
"""
Metrics that answer the critical questions about gate effectiveness.
"""
def __init__(self, audit_log: list[dict]):
self.audit_log = audit_log
def get_health_metrics(self) -> dict:
return {
"decision_distribution": self._decision_distribution(),
"review_latency": self._review_latency_stats(),
"override_rate": self._override_rate(),
"false_positive_rate": self._estimate_false_positives(),
"auto_approve_pass_rate": self._auto_approve_quality(),
"risk_calibration": self._calibration_check(),
}
def _decision_distribution(self) -> dict:
"""What percentage of changes get each decision?"""
total = len(self.audit_log)
if total == 0:
return {}
counts = {}
for entry in self.audit_log:
decision = entry["decision"]
counts[decision] = counts.get(decision, 0) + 1
return {k: v / total for k, v in counts.items()}
def _review_latency_stats(self) -> dict:
"""How long do human reviews take? Are we blocking velocity?"""
review_times = [
entry["review_time_seconds"]
for entry in self.audit_log
if entry["decision"] in ("human_review", "escalate")
and "review_time_seconds" in entry
]
if not review_times:
return {"count": 0}
review_times.sort()
return {
"count": len(review_times),
"p50_seconds": review_times[len(review_times) // 2],
"p95_seconds": review_times[int(len(review_times) * 0.95)],
"p99_seconds": review_times[int(len(review_times) * 0.99)],
"max_seconds": review_times[-1],
}
def _override_rate(self) -> dict:
"""
How often do humans override the gate's recommendation?
High override rate = model needs recalibration.
"""
overrides = [
entry for entry in self.audit_log
if entry.get("gate_decision") != entry.get("final_decision")
]
total = len(self.audit_log)
return {
"total_overrides": len(overrides),
"override_rate": len(overrides) / total if total > 0 else 0,
"by_gate_decision": self._overrides_by_decision(overrides, total),
}
def _auto_approve_quality(self) -> dict:
"""
Of changes that were auto-approved, how many were later
reverted or caused incidents?
"""
auto_approved = [
entry for entry in self.audit_log
if entry["decision"] == "auto_approve"
]
incidents = [
entry for entry in auto_approved
if entry.get("caused_incident", False)
]
reverts = [
entry for entry in auto_approved
if entry.get("was_reverted", False)
]
total = len(auto_approved)
return {
"total_auto_approved": total,
"incident_rate": len(incidents) / total if total > 0 else 0,
"revert_rate": len(reverts) / total if total > 0 else 0,
"target_incident_rate": 0.001, # < 0.1%
}
def _calibration_check(self) -> dict:
"""
Is the risk score actually predictive of outcomes?
If high-risk changes rarely cause problems, thresholds are too conservative.
If low-risk changes cause incidents, thresholds are too permissive.
"""
risk_buckets = {"low": [], "medium": [], "high": []}
for entry in self.audit_log:
risk = entry.get("composite_risk", 0)
if risk < 0.4:
risk_buckets["low"].append(entry)
elif risk < 0.7:
risk_buckets["medium"].append(entry)
else:
risk_buckets["high"].append(entry)
return {
bucket: {
"count": len(entries),
"incident_rate": sum(
1 for e in entries if e.get("caused_incident", False)
) / len(entries) if entries else 0,
"revert_rate": sum(
1 for e in entries if e.get("was_reverted", False)
) / len(entries) if entries else 0,
}
for bucket, entries in risk_buckets.items()
}
The gate's thresholds shouldn't be static. They should adapt based on what actually happens:
class GateCalibrator:
"""
Periodically recalibrates gate thresholds based on outcomes.
Runs as a scheduled job, not in the hot path.
"""
def __init__(self, config: dict):
self.config = config
self.min_samples_per_bucket = 100
self.lookback_days = 30
def propose_threshold_adjustments(self, metrics: dict) -> list[dict]:
"""
Generate proposed threshold changes with justification.
These proposals go through their own gate (human approval).
"""
proposals = []
calibration = metrics.get("risk_calibration", {})
override_data = metrics.get("override_rate", {})
auto_quality = metrics.get("auto_approve_quality", {})
if auto_quality.get("incident_rate", 0) > auto_quality.get("target_incident_rate", 0.001):
proposals.append({
"parameter": "auto_approve_threshold",
"current": self.config.get("thresholds", {}).get("auto_approve", 0.3),
"proposed": self.config.get("thresholds", {}).get("auto_approve", 0.3) - 0.05,
"reason": (
f"Auto-approved changes caused incidents at rate "
f"{auto_quality['incident_rate']:.4f} (target: "
f"{auto_quality['target_incident_rate']:.4f}). "
f"Lowering threshold to reduce auto-approve volume."
),
"confidence": "high" if auto_quality.get("total_auto_approved", 0) > 500 else "medium",
})
if override_data.get("override_rate", 0) > 0.3:
proposals.append({
"parameter": "review_threshold",
"current": self.config.get("thresholds", {}).get("review", 0.4),
"proposed": self.config.get("thresholds", {}).get("review", 0.4) + 0.05,
"reason": (
f"Reviewers override gate decisions {override_data['override_rate']:.1%} "
f"of the time. Gate may be too conservative β consider raising threshold."
),
"confidence": "medium",
})
for bucket, stats in calibration.items():
if stats.get("count", 0) < self.min_samples_per_bucket:
continue
if bucket == "high" and stats.get("incident_rate", 0) < 0.01:
proposals.append({
"parameter": "hard_block_threshold",
"current": self.config.get("thresholds", {}).get("hard_block", 0.8),
"proposed": self.config.get("thresholds", {}).get("hard_block", 0.8) + 0.05,
"reason": (
f"High-risk bucket (risk > 0.7) has incident rate "
f"{stats['incident_rate']:.4f}, well below expected. "
f"Hard block threshold may be too aggressive."
),
"confidence": "medium",
})
return proposals
No system is perfect. The gate engine itself can fail, and the failure modes matter:
class ResilientGateEngine:
"""
Wraps the gate engine with failure handling.
Key principle: the gate's failure mode must be safe.
"""
def __init__(self, engine: GateEngine, config: dict):
self.engine = engine
self.config = config
self.failure_mode = config.get("failure_mode", "fail_closed")
def evaluate_safe(self, proposal: dict) -> GateResult:
"""
Evaluate with guaranteed safe failure behavior.
"""
try:
result = self.engine.evaluate(proposal)
self._record_success(result)
return result
except Exception as e:
self._record_failure(e)
return self._handle_failure(proposal, e)
def _handle_failure(self, proposal: dict, error: Exception) -> GateResult:
"""
When the gate can't evaluate, what do we do?
"""
gate_id = str(uuid.uuid4())
timestamp = time.time()
if self.failure_mode == "fail_closed":
decision = GateDecision.BLOCK
metadata = {"failure_reason": str(error), "mode": "fail_closed"}
elif self.failure_mode == "fail_open":
decision = GateDecision.AUTO_APPROVE
metadata = {"failure_reason": str(error), "mode": "fail_open"}
else: # fail_review
decision = GateDecision.HUMAN_REVIEW
metadata = {"failure_reason": str(error), "mode": "fail_review"}
result = GateResult(
decision=decision,
risk_profile=RiskProfile(), # Empty β couldn't compute
gate_id=gate_id,
timestamp=timestamp,
metadata=metadata,
)
self._audit_failure(result, error)
self._alert_on_failure(error)
return result
def _alert_on_failure(self, error: Exception):
"""
Gate failures are operational incidents.
Alert the team immediately.
"""
alert = {
"severity": "warning",
"source": "gate_engine",
"error": str(error),
"failure_mode": self.failure_mode,
"timestamp": time.time(),
"action": "Gate engine unavailable β operating in degraded mode",
}
print(f"π¨ GATE ENGINE FAILURE: {alert}")
class GateCircuitBreaker:
"""
Prevents cascading failures when the gate engine is repeatedly failing.
After N consecutive failures, opens the circuit and applies
a pre-defined emergency policy.
"""
def __init__(self, max_failures: int = 5, recovery_seconds: int = 300):
self.max_failures = max_failures
self.recovery_seconds = recovery_seconds
self.consecutive_failures = 0
self.last_failure_time = 0
self.state = "closed" # closed, open, half_open
def should_use_fallback(self) -> bool:
if self.state == "open":
if time.time() - self.last_failure_time > self.recovery_seconds:
self.state = "half_open"
return True # Try once to see if it recovered
return True # Still open, use fallback
if self.state == "half_open":
return True # Let one through to test
return False # Closed β use normal path
def record_failure(self):
self.consecutive_failures += 1
self.last_failure_time = time.time()
if self.consecutive_failures >= self.max_failures:
self.state = "open"
def record_success(self):
self.consecutive_failures = 0
self.state = "closed"
def get_fallback_policy(self) -> dict:
"""
Emergency policy when circuit is open.
Conservative by default β requires human review for everything.
"""
return {
"all_changes_require_review": True,
"auto_approve_threshold": 0.0, # Effectively disabled
"escalation_threshold": 0.0, # Everything escalates
"alert_team": True,
"message": (
"Gate engine circuit breaker OPEN. All changes require "
"manual review until engine recovers."
),
}
Different environments warrant different gate strictness:
environments:
development:
thresholds:
auto_approve: 0.70
review: 0.70
escalation: 1.0
hard_block: 1.0
review_required: false
failure_mode: "fail_open"
rationale: "Developer velocity matters more than safety in dev"
staging:
thresholds:
auto_approve: 0.40
review: 0.55
escalation: 0.75
hard_block: 0.90
review_required: true
failure_mode: "fail_review"
rationale: "Balance safety and velocity; staging is for catching issues"
production:
thresholds:
auto_approve: 0.20
review: 0.35
escalation: 0.55
hard_block: 0.70
review_required: true
failure_mode: "fail_closed"
rationale: "Safety first β production changes affect real users"
critical_production:
thresholds:
auto_approve: 0.0
review: 0.25
escalation: 0.40
hard_block: 0.55
review_required: true
double_review: true
failure_mode: "fail_closed"
rationale: "Zero tolerance for automated changes in critical systems"
class AgentRateLimiter:
"""
Prevents agent-driven change flooding.
Even if individual changes pass the gate, too many changes
in rapid succession create systemic risk.
"""
def __init__(self, config: dict):
self.window_seconds = config.get("window_seconds", 3600)
self.max_changes_per_window = config.get("max_changes_per_window", 10)
self.max_concurrent_reviews = config.get("max_concurrent_reviews", 3)
self._change_timestamps: list[float] = []
self._active_reviews: int = 0
def can_proceed(self, proposal: dict) -> tuple[bool, str]:
"""Check if this change can proceed given rate limits."""
now = time.time()
self._change_timestamps = [
t for t in self._change_timestamps
if now - t < self.window_seconds
]
if len(self._change_timestamps) >= self.max_changes_per_window:
return False, (
f"Rate limit exceeded: {self.max_changes_per_window} changes "
f"per {self.window_seconds}s window. "
f"Try again in {self.window_seconds - (now - self._change_timestamps[0]):.0f}s"
)
if self._active_reviews >= self.max_concurrent_reviews:
return False, (
f"Too many concurrent reviews ({self.max_concurrent_reviews} max). "
f"Please wait for existing reviews to complete."
)
self._change_timestamps.append(now)
return True, "OK"
def record_review_start(self):
self._active_reviews += 1
def record_review_complete(self):
self._active_reviews = max(0, self._active_reviews - 1)
Here's the complete integration showing how all pieces connect:
class AgentSafePipeline:
"""
Complete pipeline orchestrating agent proposals through
rate limiting, gate evaluation, human review, and deployment.
"""
def __init__(self, config: dict):
self.config = config
self.gate_engine = GateEngine(config)
self.resilient_gate = ResilientGateEngine(
self.gate_engine,
config.get("failure_handling", {})
)
self.circuit_breaker = GateCircuitBreaker(
max_failures=config.get("circuit_breaker", {}).get("max_failures", 5),
recovery_seconds=config.get("circuit_breaker", {}).get("recovery_seconds", 300),
)
self.rate_limiter = AgentRateLimiter(
config.get("rate_limits", {})
)
self.metrics = GateMetrics(self.gate_engine.audit_log)
def process_proposal(self, proposal: dict) -> GateResult:
"""
Full pipeline for processing an agent proposal.
"""
can_proceed, reason = self.rate_limiter.can_proceed(proposal)
if not can_proceed:
return GateResult(
decision=GateDecision.BLOCK,
risk_profile=RiskProfile(),
gate_id=str(uuid.uuid4()),
timestamp=time.time(),
metadata={"reason": reason, "stage": "rate_limit"}
)
if self.circuit_breaker.should_use_fallback():
fallback = self.circuit_breaker.get_fallback_policy()
result = GateResult(
decision=GateDecision.HUMAN_REVIEW,
risk_profile=RiskProfile(),
gate_id=str(uuid.uuid4()),
timestamp=time.time(),
metadata={"reason": "circuit_breaker_open", "fallback": fallback}
)
return result
try:
result = self.resilient_gate.evaluate_safe(proposal)
self.circuit_breaker.record_success()
if result.decision in (GateDecision.HUMAN_REVIEW, GateDecision.ESCALATE):
self.rate_limiter.record_review_start()
result = self._initiate_review(result, proposal)
return result
except Exception as e:
self.circuit_breaker.record_failure()
return