{"slug": "give-muse-code-an-on-call-mcp-toolkit-with-kong-ai-gateway", "title": "Give Muse Code an On-Call MCP Toolkit with Kong AI Gateway", "summary": "A developer built an on-call MCP toolkit that lets Meta's Muse Code terminal coding agent perform the first ten minutes of an incident investigation while enforcing per-identity tool access at Kong AI Gateway 2.0. The gateway converts an existing REST ops API into MCP tools and filters tools/list per identity, so an oncall-investigator identity never sees the rollback-deployment tool while an oncall-operator does — a boundary the developer notes Muse's own sandbox and approval settings do not apply to MCP tool calls.", "body_md": "I wanted to find out whether a coding agent could do the first ten minutes of an incident investigation. Not fix anything. Just the part where you open four tabs, line up a deploy timeline against an error timeline, and work out which change to blame.\n\n**Muse Code**, Meta's terminal coding agent, connects to remote MCP servers over `streamable_http`. So the tools are easy. The part that stopped me is in Meta's own documentation: MCP tools operate outside the sandbox and approval mechanisms. They run as unrestricted child processes or network connections.\n\nThat matters more than it sounds. You can set `shell_execute` and `file_write` to `ask` in your Muse settings and feel covered. Those settings do not put a prompt in front of an MCP tool call. If one of your MCP tools can roll back a production deployment, the agent can roll back a production deployment, and nothing in the client will stop it.\n\nSo the question was never \"can the agent investigate\". It was \"where does the boundary go\". I put it in front of the tools, at **Kong AI Gateway**, and gave the same agent two identities to prove it holds.\n\nEverything below runs. The gateway config, the mock ops API, the verification script, and the real agent transcripts are all here: [github.com/tejakummarikuntla/kong-muse-oncall-toolkit](https://github.com/tejakummarikuntla/kong-muse-oncall-toolkit)\n\nI looked at writing a bespoke MCP server that checks a role on every handler. That works, and it means every permission decision lives in code I have to test, and the tool list is the same for everyone regardless of who is calling.\n\nI looked at just putting the tools behind an API key. That is all or nothing. Anyone with the key gets every tool.\n\nI went with Kong AI Gateway 2.0 because of one specific behavior: it converts an existing REST API into MCP tools, and it applies per-tool access control lists per identity, so `tools/list` itself is filtered. An investigator identity does not see a rollback tool. It cannot propose calling something it was never shown. That is a different property from \"the call gets rejected\", and I wanted both.\n\nKong AI Gateway sits between the agent and the operational APIs. It authenticates the caller, decides per tool whether that identity may use it, converts allowed calls into ordinary HTTP requests against the REST API, and writes an audit entry for every decision.\n\nTwo identities:\n\n| Tool | `oncall-investigator` | `oncall-operator` | \n|---|---|---|\n| `get-service-health` | visible | visible | \n| `list-deployments` | visible | visible | \n| `get-error-summary` | visible | visible | \n| `get-error-samples` | visible | visible | \n| `search-runbooks` | visible | visible | \n| `get-runbook` | visible | visible | \n| `create-incident` | visible | visible | \n| `rollback-deployment` | **not visible** | visible | \n\nNote that this is not a read versus write split. The investigator can write. It opens the incident. What it cannot do is change production. That distinction is the whole point, and a plain read-only key gets it wrong.\n\nThe ops API is an ordinary REST service. Deployments, error summaries, error samples, runbooks, incidents. Kong's `conversion-listener` maps each endpoint to an MCP tool, so I did not write an MCP server at all.\n\n```\nai_gateway_mcp_servers:\n  - ref: oncall-mcp\n    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}\n    name: oncall-mcp\n    display_name: \"On-call toolkit\"\n    type: conversion-listener\n    enabled: true\n    config:\n      url: http://host.docker.internal:9110/v1\n      route:\n        paths:\n          - /ops-mcp\n    tools:\n      - name: get-error-summary\n        description: >-\n          Error counts for a service in five-minute buckets, broken down by\n          error type. Use this to find the exact onset time and to see which\n          error type is actually growing rather than which is merely loudest.\n        method: GET\n        path: /ops-mcp/errors/summary\n        annotations:\n          read_only_hint: true\n          idempotent_hint: true\n        parameters:\n          - name: service\n            in: query\n            required: false\n            schema:\n              type: string\n            description: Service name. Defaults to \"checkout\".\n          - name: minutes\n            in: query\n            required: false\n            schema:\n              type: integer\n            description: Window length in minutes. Defaults to 60.\n```\n\nTwo things here are easy to get wrong.\n\nThe tool `path` must include the route prefix. Route `/ops-mcp` plus `url: .../v1` plus tool path `/ops-mcp/errors/summary` resolves to `http://host:9110/v1/errors/summary`. If you write `path: /errors/summary` you get a 404.\n\nThe `description` is not decoration. It is the only thing the agent reads when deciding which tool to reach for. I wrote each one as an instruction to a colleague, including when to use it, and the quality of the investigation moved noticeably.\n\n```\nai_gateway_auth_strategies:\n  - ref: oncall-key-auth\n    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}\n    name: oncall-key-auth\n    display_name: \"On-call Key Auth\"\n    type: key-auth\n    config:\n      key_names:\n        - apikey\n      key_in_header: true\n      hide_credentials: true\n\nai_gateway_consumers:\n  - ref: oncall-investigator\n    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}\n    name: oncall-investigator\n    display_name: \"On-call Investigator\"\n    type: api-key\n    credentials:\n      - ref: oncall-investigator-key\n        ai_gateway_consumer: !ref oncall-investigator#id\n        name: oncall-investigator-key\n        display_name: \"On-call Investigator Key\"\n        type: api-key\n        api_key: !secret {source: !env INVESTIGATOR_KEY}\n\nai_gateway_consumer_groups:\n  - ref: incident-response\n    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}\n    name: incident-response\n    display_name: \"Incident Response\"\n    consumers:\n      - !ref oncall-investigator#name\n```\n\nThe operator is the same shape, in a group called `sre-oncall`. Write-only fields like `api_key` need `!secret`, not a plain string.\n\nA baseline on the server, then an override on the one tool that changes production.\n\n```\n    access:\n      acl_attribute_type: consumer\n      auth_strategies:\n        - !ref oncall-key-auth#name\n      default_tool_acls:\n        allow:\n          - incident-response\n          - sre-oncall\n- name: rollback-deployment\n        description: >-\n          Roll a service back to the deployment that preceded the one named.\n          This changes production. Only run it when a runbook calls for it and\n          an incident is already open.\n        method: POST\n        path: /ops-mcp/deployments/{deployment_id}/rollback\n        annotations:\n          destructive_hint: true\n          idempotent_hint: false\n        access:\n          acls:\n            allow:\n              - sre-oncall\n        parameters:\n          - name: deployment_id\n            in: path\n            required: true\n            schema:\n              type: string\n            description: The deployment to roll back. It must currently be live.\n```\n\n`acl_attribute_type` is required whenever you set `access`. Leave it out and the apply is rejected.\n\nA per-tool ACL fully replaces `default_tool_acls` rather than merging with it. `rollback-deployment` lists only `sre-oncall`, so `incident-response` loses it. On `create-incident` I had to restate both groups for exactly this reason.\n\nApply it:\n\n```\nkongctl apply -f oncall-gateway.yaml --region in\n```\n\nUse `apply`, not `sync`. `apply` only creates and updates. `sync` reconciles every resource type named in the file, which means it deletes entities of those types that are not in your file.\n\n```\n    config:\n      logging:\n        audits: true\n        payloads: true\nai_gateway_policies:\n  - ref: oncall-audit-log\n    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}\n    name: oncall-audit-log\n    display_name: \"On-call audit log\"\n    type: http-log\n    enabled: true\n    global: false\n    config:\n      http_endpoint: http://host.docker.internal:9111\n      method: POST\n```\n\nI pointed `http-log` at a small local sink that prints the interesting fields. Every line below came from Kong.\n\nMuse Code reads `~/.config/muse/settings.json`:\n\n```\n{\n  \"schema_version\": 1,\n  \"provider\": \"meta\",\n  \"model\": \"muse-spark-1.3-contributor\",\n  \"mcp_servers\": {\n    \"oncall\": {\n      \"transport\": \"streamable_http\",\n      \"url\": \"http://localhost:8000/ops-mcp\",\n      \"headers\": {\n        \"apikey\": \"inv_your_investigator_key\"\n      },\n      \"enabled\": true,\n      \"mode\": \"required\"\n    }\n  }\n}\n```\n\nThen the prompt an on-call engineer would actually type:\n\n```\nmuse exec \"Checkout errors increased after the last deployment. Investigate why, using the on-call tools available to you. Show the evidence for your conclusion, then tell me what should be done about it.\"\n```\n\nNothing in that prompt names a tool, a service version, or a cause. Here is what came back, unedited:\n\n```\nConclusion: the last checkout deployment `dep-482` caused the spike.\n\nEvidence:\n\n- `checkout` is degraded since `2026-09-24T16:58:46Z`: error rate 8.4% vs SLO\n  0.5%, p99 4180ms vs baseline 310ms, live version `2026.9.24-3`.\n- Last deployment is `dep-482` at `16:55:46Z` — `2026.9.24-3`, \"checkout: reuse\n  pooled connections for payments client\" — ~3 min before degradation started.\n- Error breakdown: `PaymentProviderTimeout` was 0/5min before, then 125 in\n  `16:56:46-17:01:46Z`, rising to sustained 235/5min (1500 total). Other types\n  flat: `CardDeclined` ~55-60, `CartValidationError` 15, `SessionExpired` 10.\n- Samples on `2026.9.24-3`: \"payments client: pool exhausted, waited 3000ms for\n  a free connection\", `pool_max_size: 8, pool_active: 8, pool_idle: 0`, stack\n  `PaymentsClient.acquire (payments/pool.js:114)`.\n- Dependency ruled out: `payments-provider-api` healthy, error 0.1% vs SLO 1.0%,\n  p99 288ms vs baseline 275ms.\n- Runbook `rb-204` \"Payment provider timeouts in checkout\": timeouts here almost\n  always mean pool saturation, not provider down; if a deploy landed within\n  ~10 min of onset, treat it as cause.\n\nIncident opened: `INC-1001` (sev2).\n\nWhat should be done, per `rb-204`:\n\n1. Roll back `checkout` from `dep-482` to `dep-481` first, debug afterwards.\n   Rollback is safe/idempotent.\n2. Do not hot-patch pool sizing under load — a saturated pool does not recover\n   on its own while traffic continues.\n3. Do not page `#payments-oncall` / fail open to queued-authorization — that\n   path is only if provider latency is elevated, and it is not.\n```\n\nThe step I care about most is \"dependency ruled out\". The runbook says to check the provider before blaming it, and the agent checked the provider before blaming it. It also declined the escalation path on the grounds that its precondition was not met.\n\nAnd here is Kong's record of that same run:\n\n```\ntime      consumer              tool                     result\n17:31:47  oncall-investigator   initialize                     HTTP 200   6ms\n17:31:47  oncall-investigator   (tool list)              ALLOW HTTP 200   group incident-response\n17:31:52  oncall-investigator   get-error-summary        ALLOW HTTP 200   group incident-response\n17:31:52  oncall-investigator   list-deployments         ALLOW HTTP 200   group incident-response\n17:31:52  oncall-investigator   get-service-health       ALLOW HTTP 200   group incident-response\n17:31:57  oncall-investigator   get-service-health       ALLOW HTTP 200   group incident-response\n17:31:57  oncall-investigator   get-error-samples        ALLOW HTTP 200   group incident-response\n17:31:57  oncall-investigator   search-runbooks          ALLOW HTTP 200   group incident-response\n17:32:05  oncall-investigator   get-runbook              ALLOW HTTP 200   group incident-response\n17:32:23  oncall-investigator   create-incident          ALLOW HTTP 200   group incident-response\n```\n\nTen tool calls, every one allowed, and no `rollback-deployment` anywhere. The agent recommended a rollback in prose and never attempted the call, because Kong never showed it the tool.\n\nNow the same binary, same prompt style, operator key:\n\n```\nmuse exec \"Incident INC-1001 is open for the checkout service. Read the incident and runbook rb-204, confirm from the evidence whether a rollback is justified, and if it is, perform the rollback and report the result.\"\nphp\nRollback justified and completed.\n\nAction: rolled back `dep-482` -> `dep-481` (version 2026.9.24-2) under `INC-1001`.\nResult: `status: completed`, `rollback_id: rb-912229ed`, `restored: dep-481`,\ncompleted 2026-09-24T17:33:55Z. Checkout health now `recovering`.\n17:33:26  oncall-operator       (tool list)              ALLOW HTTP 200   group sre-oncall\n17:33:39  oncall-operator       get-service-health       ALLOW HTTP 200   group sre-oncall\n17:33:39  oncall-operator       get-runbook              ALLOW HTTP 200   group sre-oncall\n17:33:39  oncall-operator       list-deployments         ALLOW HTTP 200   group sre-oncall\n17:33:42  oncall-operator       get-error-samples        ALLOW HTTP 200   group sre-oncall\n17:33:55  oncall-operator       rollback-deployment      ALLOW HTTP 200   group sre-oncall\n17:33:57  oncall-operator       get-service-health       ALLOW HTTP 200   group sre-oncall\n```\n\nSame endpoint, same agent, different key. The eighth tool appears and works.\n\nTool filtering stops a cooperative agent from trying. It is not a security boundary on its own, because a client can call a tool name it was never offered. So I wrote a script that speaks the MCP handshake directly and deliberately calls `rollback-deployment` with the investigator key.\n\n``` php\n[anonymous] no credential\n  PASS  anonymous -> 401\n\n[investigator] key in group incident-response\n  PASS  tools/list shows 7 tools\n  PASS  rollback-deployment absent from tools/list\n  PASS  get-service-health -> degraded\n  PASS  create-incident -> allowed          | incident=INC-1001\n  PASS  rollback-deployment -> 403\n\n[operator] key in group sre-oncall\n  PASS  tools/list shows 8 tools\n  PASS  rollback-deployment -> completed    | restored=dep-481\n\n[audit] Kong's http-log output\n  PASS  exactly one deny recorded\n  PASS  deny names the caller               | oncall-investigator (identifier=username)\n  PASS  denied call never reached the ops API | upstream_status='' proxy_latency=-1\n  PASS  same tool allowed for sre-oncall\n\n16/16 passed\n```\n\nThe denial record is worth looking at directly:\n\n```\n\"ai\": {\n  \"mcp\": {\n    \"audit\": [\n      {\n        \"primitive_name\": \"rollback-deployment\",\n        \"primitive\": \"tool\",\n        \"consumer\": { \"name\": \"oncall-investigator\", \"identifier\": \"username\" },\n        \"action\": \"deny\"\n      }\n    ]\n  }\n},\n\"upstream_status\": \"\",\n\"latencies\": { \"proxy\": -1 }\n```\n\n`upstream_status` is empty and `latencies.proxy` is -1. Kong never opened a connection to the ops API. The refusal happened at the gateway, and the ops API has no record that anyone tried.\n\nThat is two layers doing two different jobs. Filtering keeps the agent from forming the intent. The ACL handles the client that forms it anyway.\n\n**Environment variables do not interpolate into MCP headers.** The Muse docs describe `${VAR}` interpolation. In Muse Code 1.3.0 it did not apply to MCP header values. My config sent the literal string `${INVESTIGATOR_KEY}`, Kong returned 401, and Muse reported this:\n\n```\nRequired MCP server `oncall` failed during startup: it requires an\nOAuth sign-in; run `muse mcp login oncall` and restart.\n```\n\nThere is no OAuth anywhere in this setup. Muse turns any 401 into an OAuth prompt, which sends you a long way in the wrong direction. Putting the literal key in the file connected immediately. If you hit that error against a key-auth server, check the header value before you touch OAuth.\n\n**There is no `--settings` flag.** Muse reads `~/.config/muse/settings.json` and nothing else. To run two identities without overwriting my real config, I used `XDG_CONFIG_HOME`, which it does respect:\n\n```\nXDG_CONFIG_HOME=/tmp/muse-investigator muse exec \"...\"\n```\n\nCopy `auth.json` and `trust.json` into that directory alongside your `settings.json` or the run will not authenticate.\n\n**Test the wiring with the echo provider.** `muse exec --provider echo \"ping\"` starts the session and connects every MCP server with `mode: required`, but never calls the model. Every configuration mistake above was found this way, at zero cost.\n\n**Request bodies collapse into one argument.** Path and query parameters get prefixed, so `deployment_id` in a path becomes `path_deployment_id` and `service` in a query becomes `query_service`. A request body does not get expanded into named arguments at all. It arrives as a single `body` object. Check `tools/list` before you assume an argument name.\n\n**The `request_body` shape is not in the docs.** The reference says tools accept `request_body` in OpenAPI JSON format and gives no example. The OpenAPI 3 `requestBody` object works:\n\n```\n        request_body:\n          required: true\n          content:\n            application/json:\n              schema:\n                type: object\n                required: [title]\n                properties:\n                  title:\n                    type: string\n```\n\n**`hide_credentials` does not apply to the log.** I had `hide_credentials: true` on the auth strategy, which strips the key from the request Kong forwards upstream. The `http-log` payload still contains the headers the *client* sent, raw key included. I found 26 copies of a working API key sitting in my log file. If you forward `http-log` to a hosted log service, that is where your keys end up. Scrub the sensitive headers at the sink, or before the sink:\n\n```\nSENSITIVE_HEADERS = (\"apikey\", \"authorization\", \"x-api-key\", \"cookie\")\n\ndef redact(entry):\n    for section in (\"request\", \"response\"):\n        headers = (entry.get(section) or {}).get(\"headers\")\n        if isinstance(headers, dict):\n            for h in list(headers):\n                if h.lower() in SENSITIVE_HEADERS:\n                    headers[h] = \"<redacted>\"\n```\n\n**An allow and a deny log different subjects.** An allow records the consumer group that granted access, with `identifier: consumer_group`. A deny records the calling consumer, with `identifier: username`, because no group matched. If you are building alerts on this, do not assume one field shape.\n\n**Each allowed tool produces two log entries.** One for the MCP call, one for the loopback HTTP request Kong makes to your REST API. They share no correlation id and the loopback is flushed first, so you cannot pair them by arrival order. I tried, and it confidently attached the wrong upstream call to the wrong tool. The `upstream_status` field on the MCP entry is the reliable signal.\n\n`create-incident` idempotent.`sre-oncall` can roll back anything at any time. Requiring a matching open incident would tie the permission to a live situation instead of a standing grant.`--base-url` flag that overrides the Meta provider endpoint, and Kong models accept an `upstream_url`, so the model traffic could run through the same gateway that holds the tools. I have not built that yet. It would put token spend and prompt content under the same control point as the actions.\nThe reusable shape is not the checkout scenario. It is this: your existing REST API already encodes operations at different risk levels, and an agent does not need all of them at once.\n\nTake any internal API you already run. Split its endpoints into diagnostics, safe writes, and production changes. Give the agent's identity the first two. Put the third behind a different identity. Then check the audit log for the tools it never got to call.\n\nThe thing that surprised me was how much the agent could conclude without any ability to act. It produced a complete, evidence-backed diagnosis with a specific recommendation, and the recommendation was correct. Withholding the destructive tool cost nothing in diagnostic quality.\n\nThe whole thing is on GitHub: [github.com/tejakummarikuntla/kong-muse-oncall-toolkit](https://github.com/tejakummarikuntla/kong-muse-oncall-toolkit)\n\n```\ngit clone https://github.com/tejakummarikuntla/kong-muse-oncall-toolkit\ncd kong-muse-oncall-toolkit\ncp .env.example .env     # add your gateway id and generate two keys\n./stack.sh               # ops API and audit sink\n./demo.sh up             # apply the gateway config\n./demo.sh verify         # 16 assertions, no model needed\n```\n\n`TRY.md` walks through poking the endpoint by hand with `./try.sh`, which does the MCP handshake for you, so you can watch a tool appear for one identity and vanish for the other before you spend a single agent prompt.\n\nIf you try this on your own ops API, I would like to know which endpoint you found hardest to classify. The diagnostics were obvious and the destructive ones were obvious. It was the middle tier, the writes that are safe until they are not, where I kept changing my mind.", "url": "https://wpnews.pro/news/give-muse-code-an-on-call-mcp-toolkit-with-kong-ai-gateway", "canonical_source": "https://dev.to/tejakummarikuntla/give-muse-code-an-on-call-mcp-toolkit-with-kong-ai-gateway-46mn", "published_at": "2026-09-30 17:37:08+00:00", "updated_at": "2026-09-30 17:47:13.309552+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "ai-tools", "ai-infrastructure", "developer-tools"], "entities": ["Muse Code", "Meta", "Kong AI Gateway", "Kong", "Model Context Protocol", "oncall-investigator", "oncall-operator", "rollback-deployment"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/give-muse-code-an-on-call-mcp-toolkit-with-kong-ai-gateway", "markdown": "https://wpnews.pro/news/give-muse-code-an-on-call-mcp-toolkit-with-kong-ai-gateway.md", "text": "https://wpnews.pro/news/give-muse-code-an-on-call-mcp-toolkit-with-kong-ai-gateway.txt", "jsonld": "https://wpnews.pro/news/give-muse-code-an-on-call-mcp-toolkit-with-kong-ai-gateway.jsonld"}}