Every tool you connect to an agent is two things at once: tokens the model reads on every run, and an action the model can take. So the same inventory answers your cost question and your permission question. I start every agent with that inventory, and with nothing connected that the task doesn't need.
I talked through this with Tom Smith on CoderLegion's Developer Stories (video, from 9:46). This post is the practical version.
An MCP client discovers tools by sending tools/list to each server. Each tool comes back with a name, a description and a JSON Schema for its input, and the client typically hands those definitions to the model. Anthropic published one example of what that costs: five servers (GitHub, Slack, Sentry, Grafana, Splunk) exposing 58 tools took about 55K tokens "before the conversation even starts" (Anthropic engineering). The same post names wrong tool selection as the most common failure when tools have similar names.
That is the cost side. The permission side is simpler: if a tool is in the list, the model can choose it.
Each agent gets its own environment with only the access we grant. Nothing is inherited from the developer's machine or from another agent. Zero trust applies inside the system, not only at its edge.
Instead of connecting every server the team uses, the task declares what it needs. An illustrative policy, not a format any tool enforces today:
task: triage-failing-build
allow:
- server: github
tools: [get_pull_request, list_check_runs, get_job_logs]
- server: sentry
tools: [get_issue]
deny_everything_else: true
run_budget:
max_tool_calls: 40
max_total_tokens: 200000
wall_clock_timeout: 10m
on_budget_exceeded: stop_and_report
Two things to notice. Write tools (create_comment, merge) aren't in the list, so this task can read and report but not act. And the budget is part of the same policy, because a loop is a cost problem before it is anything else.
Log each tool call with its arguments and result. The MCP spec already asks clients to do this: it says clients SHOULD "log tool usage for audit purposes" and "implement timeouts for tool calls," and that servers MUST "rate limit tool invocations" (MCP tools specification). The log is how you find where the agent went off the rails and what to roll back.
tools/list on every connected server and write down the count per server.
None of this needs a new framework. It is an allowlist, a budget and a log.
A fixed allowlist works when you know the task in advance. Many agent tasks don't announce what they'll need. Anthropic's answer is to let the model search for tools on demand instead of all of them; that addresses the token cost, but the permission question remains: a tool the model can find is a tool it can call. The options I'm weighing are a planner step that requests tools before execution, with a person or policy approving any escalation, versus narrow pre-built profiles per task type.
If you've built either, I'd like to hear how it held up.