Most Kubernetes AI agents are a rootkit waiting to happen. They demand cluster-admin rights and stream raw stdout to external LLM APIs. SREK3S is a zero-trust, read-only incident response agent. It intercepts pod crashes, scrubs secrets in-memory before network egress, and sandboxes LLM triage in a POSIX-jailed worker. It generates verified GitOps patches with strictly zero cluster write authority and deterministically fails closed to human review.
get/ list/ watch only. No ClusterRole, no write field on either wire contract, no mounted token on the Agent.git apply --check against the target's own bytes.
Most "AI SRE agents" are catastrophic control failures dressed as features. By binding service accounts to cluster-admin and mounting raw credentials, they turn container stdout—which routinely leaks AWS keys and Postgres passwords—into an exfiltration pipeline to third-party LLMs. Worse, granting the model write authority turns a log-line prompt injection into a full cluster mutation requiring zero vulnerabilities in the model itself.
I built SREK3S to take the opposite approach: remove the capability entirely. SREK3S is a zero-trust, read-only AI incident responder. Because it holds no cluster write authority anywhere in its architecture, the question of whether the model would do something dangerous never arises—it physically can't.
Instead of relying on prompt hardening, SREK3S enforces strict, structural security boundaries:
Role limited strictly to ClusterRole or binding.automountServiceAccountToken: false as UID 10001 with a read-only root filesystem and all capabilities dropped.
Live execution: The Sentinel intercepts an AWS Secret Access Key in a crashing pod and scrubs it in-memory before it ever hits the network.
Instead of blindly piping model hallucinations to kubectl apply, SREK3S treats LLM output with extreme suspicion.
verify_yaml_ast function parses original and patched documents into Abstract Syntax Trees to prove exactly one semantic field changed, catching type errors simple diffs miss.CrashLoopBackOff), SREK3S refuses to guess, defaulting to a Tier-2 architectural review that dispatches a Markdown war room report without proposing a patch.
I didn't just write happy-path tests; I tried to break this system in every way imaginable. The codebase passes every gate I could throw at it (go vet, gofmt, -race, black, flake8, mypy, pytest), currently sitting at 1087 passing tests (179 in Go).
negative_controls group of 6 cases that test_scrubber_manifest_spec.py) that parses the normative rule table straight out of CONTRIBUTING.md and fails the build if the code's rule IDs, order, or patterns drift from the documentation.TestSentinelRoleGrantsNoMutatingVerb parses the deployment YAML and fails the build if any verb other than lessons-learned.md. Exhaustive testing caught edge cases like a CrashLoop fixture rendered undetectable by restartPolicy: Never, a Service selector that routed nowhere despite 16 green manifest tests, and a terminationGracePeriodSeconds block mistakenly placed at the container level—which text-matching tests missed, but the actual API server rejected. SREK3S extracts the genuinely useful capabilities of LLMs—reading crash evidence and forming hypotheses—without handing over the keys to the cluster.
SREK3S is open-source and MIT-licensed. Because it relies on standard client-go informers, it deploys cleanly to any conformant cluster (k8s, k3s, Minikube, EKS).
You do not need to build Go binaries or compile Python to test it. We publish multi-arch images directly to GHCR.
Clone the repo, apply the quickstart overlay, and detonate the provided memory-leak fixture to watch the in-memory redaction happen live: