{"slug": "your-ai-policy-doesn-t-run-in-production-your-gateway-does", "title": "Your AI Policy Doesn't Run in Production. Your Gateway Does.", "summary": "A developer-facing analysis argues that AI governance fails in practice because policies live in documents while enforcement must run on every model call, and proposes an AI gateway as the single choke point for inbound policy, routing, identity and budgets. The piece cites OWASP's 2025 Top 10 for LLM Applications ranking prompt injection first and shows a Python snippet pointing an OpenAI client at a gateway base URL so provider credentials and routing stay out of each codebase.", "body_md": "Here is a quick test for your AI governance setup. An auditor asks: \"Show me every prompt sent last quarter that contained customer personal data, which model received it, and what policy was applied.\"\n\nHow long would that take you?\n\nFor most teams, the honest answer involves grepping logs across a handful of repos, discovering that two services log prompts in different formats, one logs nothing, and one team is calling a provider directly with a key nobody in security knows about. The policy document says all of this is controlled. The infrastructure says otherwise.\n\nThat gap is the real problem. Governance is not a PDF. It is a set of controls that run on every model call, plus the ability to prove they ran. NeuralTrust makes this case in detail in a recent post on AI gateways and enterprise governance. This is the condensed, developer-facing version.\n\nThe default pattern is that every team bolts governance onto its own app. Someone writes a PII regex, someone else adds a moderation call, a third team logs prompts to their APM tool. It works for the first LLM feature. By the tenth it has turned into drift:\n\nAgents make this worse. An agent does not just send a prompt. It retrieves context, calls tools and hands work to other agents, often on behalf of a user whose identity is lost after the first hop. If you are working on that side of the problem, [Agent Security](https://agentsecurity.com/) is a useful knowledge hub covering agent threat models, tool governance and runtime enforcement.\n\nAn AI gateway sits between your applications and agents on one side and model providers (and increasingly MCP servers and tools) on the other. Every request passes through it, which makes it the one place where policy can be enforced consistently and where a complete record can be produced, regardless of which team or SDK generated the call.\n\nFrom the app's perspective, adoption is usually a config change. With an SDK that supports a custom base URL, it looks something like this:\n\n``` python\nimport os\nfrom openai import OpenAI\n\n# Point the client at the gateway instead of the provider.\n# The key is scoped to this app by the gateway, not a raw provider key.\nclient = OpenAI(\n    base_url=os.environ[\"LLM_GATEWAY_URL\"],\n    api_key=os.environ[\"LLM_GATEWAY_APP_KEY\"],\n)\n\nresp = client.chat.completions.create(\n    model=\"support-assistant\",  # a logical route, resolved by gateway policy\n    messages=[{\"role\": \"user\", \"content\": \"Where is my order?\"}],\n)\n```\n\nThis snippet is illustrative. The exact endpoint format and routing model depend on the gateway you use. The important part is that provider credentials, routing decisions and policy live in the gateway, not in each codebase.\n\nGovernance breaks down into four jobs. A gateway can handle all of them in one place.\n\n**Inbound policy.** Inspect the prompt before the model sees it. That means detecting and redacting personal data, credentials and payment data, flagging prompt injection attempts, and keeping the app within its intended scope (a support bot should not quietly become a general research assistant). Prompt injection is ranked first in the 2025 edition of the [OWASP Top 10 for LLM Applications](https://genai.owasp.org/llmrisk/llm01-prompt-injection/), and enforcing detection at the gateway means every app gets the same defence instead of whatever each team got around to building.\n\n**Routing policy.** Decide where each request is allowed to go. Requests tagged as carrying regulated data can be forced to a private or self-hosted model, or to a provider in an approved region. Developers do not have to reimplement this logic in every service.\n\n**Identity and budgets.** Replace shared keys with scoped credentials and attach real identity: which user, which app, which agent. From there you can enforce per-user and per-app token quotas, restrict which models a given role can reach, and cap how many tool calls an agent can make in a session. Budgets double as a scope signal. An app that keeps blowing through its quota is often handling requests it was never meant to handle.\n\n**Evidence.** Emit a structured record for every request. This is the part that turns \"we have controls\" into something you can show an auditor. A useful record looks roughly like this:\n\n```\n{\n  \"timestamp\": \"2026-09-14T10:32:07Z\",\n  \"request_id\": \"req_8f2c1a\",\n  \"app\": \"support-assistant\",\n  \"user\": \"u_19384\",\n  \"route\": \"support-assistant -> eu-private-model\",\n  \"tokens\": { \"input\": 412, \"output\": 188 },\n  \"policy\": {\n    \"pii_detected\": [\"email\"],\n    \"pii_action\": \"redacted\",\n    \"injection_score\": 0.03,\n    \"decision\": \"allowed\"\n  },\n  \"latency_ms\": 940\n}\n```\n\nOnce every request produces a record like this, the auditor's question from the start of this post becomes a query rather than a project. The same data feeds alerting, SIEM pipelines and cost attribution. NeuralTrust has a deeper guide on this layer in [LLM Observability with an AI Gateway](https://neuraltrust.ai/blog/ai-gateway-llm-observability), including which latency, token and fallback metrics are worth tracking.\n\nBe careful here, because vendors (and blog posts) tend to oversell this part. A gateway does not make you compliant. It gives you the controls and the evidence that compliance work depends on.\n\nTake the EU AI Act. [Article 12](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-12) requires high-risk AI systems to technically allow the automatic recording of events over the lifetime of the system, so that risks can be identified and operation can be monitored. Deployers of high-risk systems are also required under Article 26 to keep those automatically generated logs for at least six months, unless other Union or national law says otherwise. These obligations apply to high-risk use cases (hiring, credit scoring and similar categories listed in Annex III), not to every chatbot. But if you are building in those categories, manual documentation will not cover automatic logging, and per-app logs with inconsistent schemas will be painful to defend.\n\nGDPR is the other obvious one. The data minimisation principle means that sending more personal data to an external model than the task requires is a problem in itself. Redacting at the gateway, before the request leaves your perimeter, is a practical way to enforce that across every app at once.\n\nYou do not need to turn everything on at day one. A sequence that tends to work:\n\n[TrustGate](https://neuraltrust.ai/ai-gateway) is NeuralTrust's gateway for this layer. It sits in front of models, MCP servers, tools and agent-to-agent calls, and applies policy such as prompt inspection and data masking centrally, so individual developers are not each responsible for securing their own AI traffic. It forwards end-user identity through each hop with per-agent and per-tool access control, and records every call for audit. Deployment options include SaaS, a hybrid model with the data plane hosted in your own environment, and fully on-premises or air-gapped installs on Kubernetes.\n\nWhatever you use, the principle is the same. If your governance lives in documents and per-app code, it will drift. If it lives in the one layer every request has to pass through, you can enforce it and prove it.\n\nFor the full breakdown, including a mapping of gateway controls to the EU AI Act, GDPR and UK NCSC guidance, read the original article on the [NeuralTrust blog](https://neuraltrust.ai/blog/ai-gateway-enterprise-governance).", "url": "https://wpnews.pro/news/your-ai-policy-doesn-t-run-in-production-your-gateway-does", "canonical_source": "https://dev.to/alessandro_pignati/your-ai-policy-doesnt-run-in-production-your-gateway-does-jgj", "published_at": "2026-09-28 14:40:07+00:00", "updated_at": "2026-09-28 14:51:00.811081+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-infrastructure", "ai-policy", "developer-tools"], "entities": ["NeuralTrust", "OWASP", "OpenAI", "Agent Security", "MCP"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-ai-policy-doesn-t-run-in-production-your-gateway-does", "markdown": "https://wpnews.pro/news/your-ai-policy-doesn-t-run-in-production-your-gateway-does.md", "text": "https://wpnews.pro/news/your-ai-policy-doesn-t-run-in-production-your-gateway-does.txt", "jsonld": "https://wpnews.pro/news/your-ai-policy-doesn-t-run-in-production-your-gateway-does.jsonld"}}