# What AI agents actually do with your MCP server

> Source: <https://dev.to/krwndev/what-ai-agents-actually-do-with-your-mcp-server-14pp>
> Published: 2026-10-06 09:12:19+00:00

An MCP server can look healthy in every log and still be hard for agents to use. Requests come in, responses go out, nothing crashes. What you don't see is how the agent got there: what it asked for that you don't have, what it got wrong before it got it right, and where it got stuck.

These are the five things that only showed up once I measured them. The screenshots are from a demo flight-booking server with simulated traffic, so the numbers are examples; the patterns are the ones agents produce.

Agents guess tool names. Some guesses are near misses: `searchFlights` when the tool is `search_flights`, `seatMap` instead of `seat_map`. Others are a tool the agent assumed should exist, like `get_flight_price`.

Three things make this list useful rather than just a count of errors:

`searchFlights` is not a missing feature, it's a naming convention one client prefers. That's a description fix, not new code.`get_flight_price` most agents went on to 
`search_flights`, so they worked it out. After `export_itinerary` every session stopped: whatever the user wanted, they didn't get it.
When an agent sends arguments that don't match the input schema, the MCP SDK rejects them before your handler runs and sends the validation error back to the model as a normal result. Your handler never ran, so your code never saw a failure, and your error rate looks better than what agents experienced.

On this server, most of `book_flight`'s failures were refused arguments, not crashes. Which leads to the next point.

`book_flight` took `passengers`, described as "the passengers". Agents read that as a count and sent `2`. The schema wanted a list of names, so the SDK refused the call, the agent read the error and tried again with a list.

The fix was one sentence in the description. The dotted line on the chart is where it changed:

Errors drop right at the line, while the code didn't change at all. Rewording a description can change how agents use a tool more than a change to its logic, so it helps to see exactly when a description changed next to the numbers. The dashed lines are server releases, which is the other thing you'd want to rule out.

An agent tries `check_in`, gets "Check-in opens 24 hours before departure", and calls `check_in` again with exactly the same arguments. Then again.

Each call on its own is a normal, quick response. You only see the problem when you look at the session as a sequence, or count how many calls repeat the previous one. Usually it means the answer didn't tell the agent what to do next. Here, a better message would say when check-in opens and that retrying now won't help.

A tool that usually answers with 6 kB and sometimes with 650 kB looks fine on every latency chart, because it's fast. But that answer lands in the agent's context window, crowds out everything else, and costs tokens on every following turn.

The median and 95th percentile here are fine. The largest answer is a hundred times bigger. The usual cause is a query with no filters that returns everything. A default limit, or a summary with a way to ask for more, fixes it.

I built [mcpspan](https://github.com/mcpspan/mcpspan) for this: self-hosted analytics for MCP servers, MIT licensed. You run it with Docker and add one line to your server:

```
git clone https://github.com/mcpspan/mcpspan.git
cd mcpspan
docker compose up -d
js
import { instrument } from 'mcpspan';

instrument(server, {
  apiKey: process.env.MCPSPAN_API_KEY,
  endpoint: 'http://localhost:6271',
});
python
import os

import mcpspan

mcpspan.instrument(mcp, api_key=os.environ.get("MCPSPAN_API_KEY"), endpoint="http://localhost:6271")
```

There are SDKs for TypeScript, Python, Go, C#, Java, Rust, Ruby and PHP, with the same features in each. Parameter values never leave your server's process, and nothing is sent anywhere you didn't set up yourself.

It's an early version, so if there's something you'd want to see about how agents use your server, I'd like to hear it.
