I had a question that is annoying to answer in a spreadsheet and trivial to say out loud: if you bought an Indian stock index on any random day since 2005 and held for N years, how often did you lose money?
So I said it out loud to an AI assistant that had a Model Context Protocol connection to a server with daily index data back to 1996, and watched what it did. This post is about the plumbing, the one answer I was sure was wrong, and why that moment is the actual hard part of giving an agent real data.
If you want the investing takeaway, I wrote that up on Medium. This is the developer version.
I built the MCP server used here. Details at the end.
Take every trading day from April 2005 to October 2016 as a buy day for
Nifty 50, Nifty Midcap 150 and Nifty Smallcap 250. For each buy day,
check the index level 1, 2, 3, 4, 5, 6, 7, 8 and 10 years later. Tell
me what share of buy days ended with the index lower. Also give the
median 5-year and 10-year growth rate, and the single worst 10-year
buy day.
No code in the prompt. No mention of which tools exist. The assistant reads the tool schemas the server exposes and decides.
Four tool calls, in this order:
Result, in about forty seconds:
I then recomputed the whole thing in SQL against the same table. Every percentage matched to one decimal. Good.
I asked a follow-up: "What was the single worst 10-year buy day for each index?"
It answered 23 March 2010. For all three indices. The same date.
That looked like a data glitch to me. Three indices with different constituents, all bottoming out on the same Tuesday as the worst possible entry? I was sure a bad row was sitting in the table. I asked the agent to show me the closes around that date.
They were smooth. 5,205 on the 22nd, 5,225 on the 23rd, 5,260 on the 25th. Nothing wrong with the buy day at all.
Then I looked at the exit date. Ten years after 23 March 2010 is 23 March 2020. That was the bottom of the COVID crash. The Nifty 50 closed at 7,610 that day, down 13 percent from the day before and 38 percent from its February high.
So the "worst buy day" was never about the buy day. It was about where the exit landed. Any buy day in March 2010 would have scored almost identically, because they all exit into the same hole. The agent's answer was correct. My instinct that it was broken was the error.
Here is what bothers me about that. The agent gave me a true number with no explanation attached. A human analyst would have said "23 March 2010, but only because the exit falls on the COVID low". The agent said "23 March 2010" and stopped. I nearly threw away a correct answer because it arrived without its reason.
The fix is not a smarter model. It is a tool description that asks for the reason to travel with the number.
{
"name": "get_price_history",
"description": "Daily or weekly closes. When a caller derives a single extreme from this data (worst day, best day, peak, trough), report BOTH the entry and exit dates and check each against get_observation_status so a crash day or holiday boundary is named, not hidden."
}
One sentence in a schema. Agents follow instructions in tool descriptions more reliably than instructions in prompts, because the description is present on every call and the prompt is not.
I have not shipped that change yet. I am writing it here first because I want to see whether other people have hit the same shape of problem: a correct answer that looks wrong because the context that makes it correct was never requested.
1. Every extreme comes with both dates. A "worst 10-year window" is two dates, not one. If the tool only returns the entry, the caller will blame the entry.
2. Dates get a status, not just a value. Holiday, weekend, pre-listing, post-delisting, crash day. The server already exposes this as an observation-status tool. I did not ask for it, so the agent did not call it. The schema should tell it when to.
3. Identifiers are time-dependent. The index called "Nifty Midcap 150" today was not called that in 2005. Tata Motors became two companies in 2025. A tool that resolves names as-of a date is the difference between a right answer and a confident wrong one.
Four tool calls for the main question, two for the follow-up, two more to look at the closes around the suspicious date. Eight calls total against a free tier of fifty a day. The SQL I wrote to verify it took me longer than the agent took to produce it.
What is the equivalent of "check outliers first" in your domain, and have you put it in a schema yet? I would genuinely like to know what others are writing into tool descriptions.
The server is at tapetide.com/mcp: about 8,200 NSE and BSE stocks plus indices, MCP, free for 50 calls a day. Not investment advice.