# Counting MCP retries without counting pagination

> Source: <https://dev.to/getmcpulse/counting-mcp-retries-without-counting-pagination-88b>
> Published: 2026-10-09 06:01:30+00:00

A retry is the cheapest signal you'll ever get that a model didn't understand your tool.

It's also easy to count wrongly — and a wrong retry count is worse than no retry count, because it sends you off to rewrite descriptions that were fine.

Here's the definition that works, and why each part of it is there.

**Same tool, within 30 seconds, in the same session, with different arguments.**

Three conditions. Each one exists to throw out something that looks like a retry and isn't.

A model that calls `search_orders` and then `get_order` hasn't retried anything. It drilled down.

That's a tool *pair*, and it's a design note about your result shape — maybe `search_orders` should have returned enough detail that the second call wasn't needed — rather than a sign of confusion.

Only the same tool twice counts. When the model reaches for a *different* tool after a bad result, that shows up elsewhere: in empty answers, or in the pairs.

Long enough to cover a model reading a result, thinking, and calling again, with a slow tool on either side.

Short enough that the user's next, unrelated question doesn't land inside the window and get counted as a second attempt at the first one.

It's a judgement call. The important part is that it's stated rather than tuned per server, so a retry rate means the same thing everywhere.

This is the condition doing most of the work, and the one most likely to be got wrong.

**Same arguments inside thirty seconds is a repeat, not a retry.** Pagination. Polling. A client re-issuing a request it lost. A well-behaved paginating tool calls itself four times in ten seconds, and counting those as retries grades it as broken.

**Different arguments means the model looked at what came back, decided it was wrong, and changed what it asked for.** That's a rewording. A rewording is the model telling you it guessed.

Get this backwards and your best-behaved tools look like your worst.

If you're building this yourself and you don't want to store argument values — and for an MCP server you probably don't, since tool arguments are whatever the conversation contained — you need a hash.

Two things matter:

**Canonicalise first.** `{"a":1,"b":2}` and `{"b":2,"a":1}` are the same arguments and different bytes. Without sorted keys, every repeat reads as a rewording and your retry rate is noise. [RFC 8785](https://www.rfc-editor.org/rfc/rfc8785) defines exactly one byte sequence per JSON value — sorted keys, specified number format, specified escaping.

**Hash the wire bytes, not the parsed object.** By the time arguments reach your handler the SDK may have applied schema defaults and dropped unknown fields. Hash after that and a client who omits `limit` and a client who explicitly sends `limit: 50` produce the same hash — fine, arguably, until you change the default and every historical hash becomes incomparable with every new one. Silently.

What the client sent is what identifies the call.

Whether a call was retried depends on what came **after** it. That can't be known when the call arrives.

So retries are a batch job, not a property you can stamp on a request as it lands. Walk each session's calls in order, look forward thirty seconds.

One detail that's easy to miss: the pass has to read thirty seconds past the end of the period it's processing. A call at 23:59:50 can be retried at 00:00:05. Stop at midnight and the last call of every single day gets judged with its follow-up missing.

**A `bad_args`, then a successful call.** The model couldn't fill in your schema, read the validation error, and corrected itself.

This is the most fixable kind, and slightly embarrassing: the error message told the model what your schema should have told it up front. Usually a parameter with no description, or one whose description just restates its name.

**An empty answer, then a broader query.** The model got `[]`, assumed it had asked too narrowly, and loosened the filters.

Often the empty answer was *correct* — there genuinely were no matching records — and the description never said that nothing-matching was a possible outcome. One sentence fixes it.

**Two or three successful calls in a row with different arguments.** Every call returned 200. Every call returned data. The model still didn't get what it wanted.

This is the one no other metric catches. It looks perfectly green in every log you own. Usually it's a parameter whose meaning is ambiguous — a `customer_ref` that might be an ID, an email, or a name, and the model is working through the options.

Retries feed first-call success: a call followed by a retry isn't a first-call success, however it ended.

The useful view is **per tool and per client**. A server-wide retry rate tells you something is being guessed at somewhere, which is almost useless. The same rate isolated to one tool in one client tells you where to spend the afternoon.

Worth stating as a count rather than a percentage where you can. "Agents retried `search_orders` 2.4 times on average before getting a usable answer" lands harder than "62% first-call success", because a number of wasted attempts is easier to feel.

**A model that gives up rather than retrying leaves no trace.** It calls once, gets something unusable, and answers the user without your tool. That shows up as an empty answer at best and as nothing at worst.

**A retry across sessions isn't visible.** The user starts a new conversation and asks again — from inside the server, that's two unrelated first calls.

Both make the retry rate a floor rather than an estimate. Which is the right way round for a signal like this: when it's high, it's definitely telling you something.

I build [MCPulse](https://getmcpulse.com), which computes this from inside your server — two lines, never your arguments or your results.

*Originally published at [getmcpulse.com](https://getmcpulse.com/blog/retries).*
