Agentic AI Infrastructure: What It Takes to Do It Safely Red Hat shipped an OpenShift diagnostic MCP server, mcp-sre-tools, as read-only by design, with nine diagnostic tools and no write tools, after finding that RBAC alone cannot safely enable agentic remediation because meaningful fixes often require access to Secrets. The team designed a maturity-gated approval architecture for write access but has not deployed it, citing the challenge of trusting an LLM's reasoning in production infrastructure. ⚡ Byte Size Summary - See why we shipped an OpenShift diagnostic MCP server as read-only by design , and the RBAC wall that made write access harder than it looks- Walk through a real failed remediation test where an agent recommended a correct-looking fix built on stale, deprecated config — and what that failure mode actually is - Get the maturity-gated approval architecture we designed for write access — and why it’s still sitting on paper, not in production The Story the-story In Article 06 https://pipelineandprompts.com/posts/ai-in-the-stack-06-n8n-workflows/ we wired an n8n workflow to MCP and RAG for automated incident triage. That article ended with a question: what happens when the agent gets a longer leash? We built mcp-sre-tools — an MCP server that exposes OpenShift and Kubernetes diagnostics to an LLM, wired into Claude Desktop and n8n, covering ARO, ROSA HCP, OSD-GCP, and generic clusters. Nine diagnostic tools: get cluster health , diagnose crashloop , get failing pods , and others in that family. READ ONLY MODE is on by default, and there are no write tools in the codebase at all. That part shipped clean. The friction started when we scoped what came next: a remediation mode, where the agent wouldn’t just diagnose a broken deployment — it would patch it. That’s where the story stopped being a build story and became an organizational one. The Problem the-problem Platform engineers and developers landed on opposite sides of the same question almost immediately, and for reasons that turned out to be more substantial than the usual risk-aversion reflex. Developers were comfortable trusting agent-proposed changes roughly the way they’d trust a colleague’s pull request — read the diff, sanity-check it, merge it. Platform engineers pushed back hard, and their objection wasn’t reflexive. It was specific: a PR from a colleague comes with inspectable reasoning. You can ask them why. An LLM’s proposed patch doesn’t carry that same trail — the “why” is buried in a forward pass, not a code review comment. Business stakeholders, meanwhile, were worried about something simpler and more immediate: an autonomous agent breaking a critical application in production. Three legitimate concerns, three different vocabularies for the same underlying question — how much do we trust a system whose reasoning we can’t fully inspect, applied to infrastructure we can’t afford to break? Why RBAC alone doesn’t solve it why-rbac-alone-doesnt-solve-it The instinct is to reach for RBAC and call it solved. Scope the agent’s service account to a namespace, give it patch permissions on Deployments and nothing else, and let it operate inside a fence. That fence has a hole in it. Meaningful remediation almost always eventually needs to touch Secrets or environment variables — a misconfigured database connection string, an expired credential reference, a missing env var causing a crash loop. The moment your remediation scope includes Secrets, “namespace-scoped RBAC” stops being clean sandboxing and starts being a much bigger trust surface than the phrase implies. We didn’t have a way around that with RBAC alone. So we fell back to a narrower, honest justification for read-only: even without write access, a diagnostic agent cuts human mean-time-to-resolution. It’s a smaller value proposition than full self-healing, but it’s a real one — and it’s the one we could actually defend without hand-waving. The Architecture the-architecture Diagram 1 — as built: the shipped, read-only MCP architecture. The controls we designed the server to work with — note the repo intentionally ships without a default rbac.yaml , to stay adaptable across cluster types and org policies. Deployment teams are expected to write their own scoped ClusterRole / RoleBinding https://kubernetes.io/docs/reference/access-authn-authz/rbac/ tailored to their access model; the sample below shows the shape we recommend, not a default that ships: - Namespace-scoped RBAC recommended, not shipped — bind the MCP server’s service account to Role / RoleBinding resources scoped per-namespace, not a cluster-wide ClusterRole - Service-account-based access — no static kubeconfig or personal credentials in the agent’s execution path - NetworkPolicy egress/ingress restriction — the MCP server’s pod network is fenced to only the cluster API and the LLM endpoint it needs to reach - Logging and observability as non-functional requirements — every tool call is logged, not bolted on after the fact Recommended shape, not a shipped default — deployment teams write their own apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: mcp-sre-tools-reader namespace: