# Introducing Arize AX MCP: When to use MCP, CLI, or skills

> Source: <https://arize.com/blog/extending-your-agent-with-arize-ax-mcp-cli-and-skills/>
> Published: 2026-10-02 15:05:45+00:00

Earlier this year our Chief Product Officer Aparna Dhinakaran [posted a hot take on X](https://x.com/aparnadhinak/status/2038798463098552727): “MCP = context tax. The future is CLI and files.” 

Last week we shipped a hosted MCP server for Arize AX, because that’s no longer true.

But both things were true when she said it. What changed is harnesses caught up, and once the context cost stops being the deciding factor, the real question comes into focus. It was never MCP or CLI. It is what abilities you are giving the agent, where the agent runs, and who is asking. MCP, CLI, and skills are three ways to extend an agent with the same platform, and this post is about how we think about each.

Want to try the new Arize AX MCP server?

Connect Arize AX to Claude Code, Cursor, Claude Desktop, and other MCP clients. [Read the MCP docs for setup instructions, authentication, and available tools](https://arize.com/docs/api-clients/mcp/overview).

### Build better agents with Arize

Trace, evaluate, and learn. Build agents that work with Arize AX and start tracing your runs today.

**Prefer open source?**
[Try Arize Phoenix for self-hosted, open source agent observability](https://arize.com/phoenix?utm_source=blog&utm_medium=referral&utm_campaign=ax-inline-cta&utm_content=extending-your-agent-with-arize-ax-mcp-cli-and-skills-inline-cta-phoenix).
      

## What we shipped

Let’s start with what an agent can do now that it couldn’t last month. From inside Claude Code, Cursor, or Claude Desktop, it can pull the ten slowest traces for a project and read the spans. It can diff two experiment runs and tell you which examples regressed. It can read the current version of an evaluator prompt, check which monitors are firing, and look up a dataset example, all without you switching windows, knowing the API, or installing the CLI.

The AX MCP server that makes this possible is hosted. There’s nothing to install. You simply point your client at one of three regional endpoints and pass your Arize API key as a bearer token:

```
{
  "mcpServers": {
    "arize-ax": {
      "type": "http",
      "url": "https://api.arize.com/mcp",
      "headers": {
        "Authorization": "Bearer ${ARIZE_API_KEY}"
      }
    }
  }
}
```

It exposes 45 tools plus `server_info`: `list_projects`, `list_traces`, `get_experiment`, etc. across projects, traces, datasets, experiments, evaluators, monitors, and prompts. Each is a real, named operation with its own filters, not a generic search-and-execute pair. Setup details for each client are in the [MCP docs](https://arize.com/docs/api-clients/mcp/overview).

We also ship the `ax` CLI: 22 command groups covering the full surface of the platform, from `ax experiments list` to `ax evaluators create` to `ax role-bindings`. On top of the CLI, a set of [Arize skills](https://github.com/Arize-ai/arize-skills) you install with `ax skills install` that teach an agent how to do a job with it.

Three surfaces. Same REST API underneath. The rest of this post is why we built it that way instead of picking one.

## The case against MCP was real

Let’s start with the criticism, because for most of the past year, it was correct.

MCP tool definitions used to load into context at the start of each agent session. Every tool, every parameter, every description, before the model had read your first message. [Community measurements](https://getunblocked.com/blog/mcp-token-budget-autopsy/) put a Claude Code session with five to ten servers installed at 50,000 to 67,000 tokens of overhead before the user typed anything. A GitHub + Slack + Sentry + Grafana + Splunk setup landed around 55,000 tokens in tool definitions alone. [One benchmark](https://www.scalekit.com/blog/mcp-vs-cli-use) put MCP at 4 to 32x the per-operation token cost of the equivalent CLI call. Every connected server was paying a tax whether or not you used it.

A CLI never had that problem. It costs roughly what the command string costs, a few dozen tokens, and agents discover the commands themselves. The model already knows git, docker, curl, and jq from pre-training, and when it doesn’t know a tool it can run `--help`. Output composes in the shell: pipe it, grep it, feed it to the next command. And a CLI runs anywhere there’s a shell and no chat client: CI pipelines, cron jobs, serverless functions.

So when Aparna posted that take, she was describing that context tax. But then, the harnesses fixed it.

Claude Code, Cursor, and Codex now do tool search. Only the names of connected servers sit in context at startup; the full tool schemas load when the model actually goes looking for one, and it searches before it executes. Anthropic’s engineering team went one step further in [Code execution with MCP](https://www.anthropic.com/engineering/code-execution-with-mcp): for a Drive-to-Salesforce task they cut token usage from roughly 150,000 to 2,000 by having the agent write code against the MCP servers instead of calling tools one at a time. Between deferred loading and code execution, the upfront cost that made “context tax” a fair label is mostly gone.

This changes the question. If MCP no longer costs you context you aren’t using, the reason to pick CLI over MCP, or the reverse, can’t be token math. It has to be about where the agent runs and who is asking.

## What MCP extends

Here is the part the CLI-maximalist take skips. A CLI assumes a shell. A shell assumes a terminal, a machine, and a person who is comfortable in both.

A lot of the people who want to ask questions of their AX data are not in that position. A PM in Claude Desktop asking “which of our experiments regressed on hallucination this week” has no terminal. A support engineer in Cursor wants to pull the last ten failed traces for a project without leaving the editor. Claude Desktop, Cursor, and Zed all speak MCP natively, and for those surfaces MCP is [the only option that doesn’t require the user to drop into a shell](https://getunblocked.com/blog/when-to-use-mcp-vs-cli/).

It is not only about people without a shell, though. [Thariq Shihipar](https://x.com/trq212) (@trq212) from Anthropic [put it plainly in September](https://x.com/trq212/status/2099958388230873165): he was not expecting it, but MCPs are now better than CLIs for most integrations. Models have gotten much better at tool calling, hosts defer tool loading, and MCP is stateless over HTTP. Composition moved too. Instead of piping output through grep, you put a query or filter parameter on the tool and the agent composes inside one call. Our “list” tools take filters for exactly this reason. So even an agent that has a shell, like Claude Code working in your repo, has a good reason to reach for the MCP first when the job is reading AX.

There is a governance argument too. An MCP server exposes a fixed set of named operations. A host can grant or deny each one, and the model never holds raw credentials in its context. Structured tool manifests also reduce errors: one comparison found agents using them made [38% fewer invocation errors](https://getunblocked.com/blog/when-to-use-mcp-vs-cli/) than agents composing free-form shell commands. When you are choosing among dozens of operations, a self-describing schema beats a prompt that lists every flag.

That is what the MCP server extends: any agent, in any client, gets AX as a set of named abilities it can discover and call. The CLI extends something different, which is the agent’s ability to script.

## The decision framework we use

Borrowing the [five-question shape](https://blog.mcpservers.org/posts/cli-vs-mcp) that has emerged this year, here’s a version tuned to our own experience:

| Question | If yes | Why | 
|---|---|---|
| Is the agent running somewhere without a shell? (Claude Desktop, a hosted assistant, Cursor chat) | **MCP** | There is no alternative. This is the whole reason MCP exists. | 
| Is the job a scripted pipeline, or is the agent chaining several commands into one workflow? | **CLI** | Scriptable, and the model already knows the shell idiom. It’s also the layer skills are built on. | 
| Do you need to scope permissions per operation, or keep credentials out of the model’s context? | **MCP** | Named tools with per-tool grants beat inherited shell permissions. | 
| Is this CI, a cron job, a Lambda, or a Docker entrypoint? | **CLI** | Non-interactive environments have nothing to hold a connection open. | 
| Is your token bill the problem? | **Check that tool search is on** | Modern harnesses load tool schemas on demand. If cost still dominates, move heavy workflows to code execution. The protocol is rarely the root cause. | 

Most production agents [use both](https://www.firecrawl.dev/blog/mcp-vs-cli): CLIs handle execution-heavy, scriptable work and provide a foundation for skills, while MCP supports integration and discovery for clients without shell access. Claude Code, Cursor, and Gemini CLI all do this. So do we.

## Skills are the part everyone leaves out

The MCP versus CLI framing misses a third layer, and it is the one we think matters most for getting an agent to be good at a job rather than merely able to call an API.

A skill is a `SKILL.md` file plus optional scripts and references, loaded by the agent when a task matches. It costs [roughly 100 tokens of context until it is invoked](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview), because only the name and description load at startup. The full instructions, the gotchas, the “run `ax spaces list -o json` first if you don’t know the space ID,” all of that arrives only when needed. Progressive disclosure, built in.

MCP gives an agent access to a live system. A skill teaches it how to do a task with that system. Our `arize-experiment` skill, for example, encodes the actual workflow: verify `ax --version` and the environment, resolve the space and project, list experiments with the right flags, then compare runs. That is procedural knowledge no tool schema carries.

This is also why the CLI is the right foundation for skills. Skills wrap commands. Commands are composable, produce text an agent can reason about, and chain into pipelines a skill can describe in a few lines.

The two surfaces are also starting to converge. The MCP project finalized a [Skills extension](https://github.com/modelcontextprotocol/ext-skills) in September that lets a server publish skills as resources: a client calls `skills/list`, gets back names and descriptions, and reads the full `SKILL.md` only when the model picks one. It is the same progressive disclosure skills already have on disk, with the added benefit that the instructions ship with the service they describe. Host support is still rolling out, so today our skills install through `ax skills install`. But the shape is clear: MCP handles discovery and access from any client, skills carry the procedural knowledge, and the CLI does the work where there is a shell. We expect the AX MCP server to serve our skills directly once the clients catch up.

## Best practices, as of right now

What we do (or recommend), drawn from our own setup and from what has held up in the industry over the past year:

- **Expose your actual tools.** There is a trend of hiding an MCP server behind two generic tools, search and execute. Max Stoiber at OpenAI[argues this is the wrong default](https://x.com/mxstbr/status/2093003978833277344) : clients like Codex and ChatGPT already do tool search and code mode across every server, and the models are trained on that path. A custom discovery layer per server makes the model navigate two. We ship 45 named tools and let the client find them.
- **Put filters on your tools.** If the agent needs to compose or narrow data, give the tool a query parameter rather than making it fetch everything and filter in context. One call with a filter beats three calls and a grep.
- **Annotate every tool.**`title` plus`readOnlyHint` or`destructiveHint` on every MCP tool. Clients use these to decide what needs per-call confirmation, and both the Claude and ChatGPT directories check for them.
- **Make sure tool search is on.** Modern harnesses load tool schemas on demand, but check yours does. If token cost still dominates, move heavy workflows to[code execution](https://mcp.directory/blog/mcp-context-bloat-fix-2026-tool-search-code-mode-progressive-disclosure) .
- **Wrap the CLI in skills.** Ship the procedural knowledge with the tool. A`SKILL.md` that says which command to run first is worth more than a longer tool description.
- **Same API under everything.** MCP, CLI, and SDKs should be thin transports over one REST surface. When behavior drifts between them, users notice before you do.
- **Bearer today, OAuth next.** Our MCP authenticates with an API key as a bearer token, passed through to the REST API. OAuth 2.1 with user consent is the next step, and it is what the connector directories require.

## Where this lands

We didn’t ship an MCP server because it is the future, and we didn’t skip it because CLIs are cheaper. We shipped one because it is the fastest way to give any agent, in any client, the ability to see what is happening in AX. We kept building the CLI and skills because scripting and procedural knowledge are abilities too, and they live better in a shell.

The context argument is settled. What’s left is deciding what you want your agent to be able to do, and picking the surface that gives it that where it runs.

Try it for yourself: [Get started with our MCP docs](https://arize.com/docs/api-clients/mcp/overview) 

You can also get started with the following CLI commands:

`pip install arize-ax-cli`

`ax skills install`
