Mappings don't carry meaning: Building a research agent that actually understands your data Nearform built a multi-agent research system that answers natural-language questions across an OpenSearch cluster indexing over 330,000 documents across roughly 30 indexes, after testing the OpenSearch MCP server and query-planning tool and finding that OpenSearch mappings carry no semantic meaning for an agent. The team presented the two-tier architecture at the Mastra TypeScript AI Conference in London in July 2026, following an eight-month build. The naive schema-to-DSL approach failed at scale because field names and types alone do not convey domain nuances such as tax-lot deduplication on ticker.keyword or month-end marketValue snapshots that differ between retirement and brokerage accounts. How we taught an agent to answer natural-language questions across hundreds of thousands of documents, why the obvious approach failed, and the two-tier architecture that fixed it. Many of the companies we work with at Nearform have the same starting point: they have years of data, ingested from many sources, sitting in a search cluster, sourced from multiple data stores or a GraphQL federation. Ideally, they want their users to just ask a question in plain English and get a correct answer back. Over the last eight months, one of our teams built exactly that: a multi-agent system where a research agent answers natural-language questions across a large OpenSearch cluster. We presented this architecture at the Mastra TypeScript AI Conference in London in July 2026. This article tells the story of that build. What looked easy, what broke, and the system that made it work in production. The promise: ask a question, get an answer On paper, the problem is simple. It takes three steps: 1. The user asks a question. They do not need to know where the answer lives or how the data is structured. 2. The agent builds an OpenSearch query. It reads the schema, converts the question into DSL, runs the query, and gets the data back. 3. The agent crafts an answer. The data is analysed, summarised, and given back to the user in plain language. That is the demo everyone has seen. A question goes in, an answer comes out, and it seems to work relatively well. Then, once you try to scale the amount of data, or the complexity of the data itself, this naive architecture just doesn't work anymore. To give a sense of scale, one real client system indexes over 330,000 OpenSearch documents, spread across roughly 30 separate indexes, with tens of thousands of PDF documents accessible via RAG search. Our research agent is one piece of a larger system with a main orchestrator agent that routes requests to specialised subagents, backed by persistent memory, an observability platform, and a full evaluation pipeline. The piece we focus on here is the data-retrieval subagent, which communicates with OpenSearch via a proxy. We did not jump straight to building our own architecture. First, we checked what already existed. If something off-the-shelf could answer questions across our cluster at scale, we would rather use it than build and maintain our own. So we ran a short research phase to see what the state of the art looked like. We tested the OpenSearch MCP server, which exposes the cluster to an agent so it can list indexes, read mappings, and run queries on its own. We tested the OpenSearch query-planning tool, which turns a plain-English question into DSL. We also tested a few community tools that wrap OpenSearch for agents in similar ways. Once we tried to make it work on real user queries, three problems showed up at once. The first is about meaning. The other two are about volume, and they share a single fix. In this article, we'll use a fictional financial example to explore each, but the same rules apply to pretty much every real-world domain. Problem 1 - mappings carry no meaning OpenSearch mappings list field names and types. That is all. They do not say what a field actually means, how its data was computed, when to use it, or the small nuances only a domain expert knows. We know what our fields mean because we built the system. The agent does not. Give it a field called marketValue of type double , and it will guess. It has no way to know the things we know, such as: - Each row in holdings is one tax lot, not one security. The same ticker appears multiple times. To count real positions, you have to deduplicate on ticker.keyword . - marketValue is a month-end snapshot, not a live number. Retirement accounts report it net of loans, brokerage accounts report it gross. The values are not comparable across account types. - Transaction amounts are signed. Dividends and interest book to a synthetic CASH holding, not to the security that paid them. - Transaction dates come from a monthly load. The day is always 01, so ordering within a month is lost. None of that lives in the mapping. So the agent guesses, picks the wrong fields, and writes queries that miss or return wrong data. The fix: a meta-index with a drift-guard We built a searchable, queryable meta-index with an entry for every field in the data. Next to each field's name and type, a subject-matter expert writes a clear, human description that conveys the meaning the field name does not. The agent reads it to learn what the field means, what values it can take, and how to query it correctly. The descriptions need upkeep, but they pay off. People maintain them over time, and in return the agent runs complex, domain-specific queries with confidence. The upkeep sounds scary, but you can automate much of it. The main risk is that someone adds or changes a field in the mapping and forgets to update its description in the meta-index. So we guard against that in CI. A test compares the live mapping against the meta-index on every change. If a field exists in the mapping but has no matching entry, or an entry points to a field that no longer exists, the build fails. The machine cannot write the domain knowledge for you, but it can force a human to fill the gap before the change ships. In this way, the meta-index never drifts silently out of sync. Here is what an entry looks like: { "index": "my investments", "field": "holdings.ticker", "type": "text", "nestedPaths": "holdings" , "hasKeyword": true, "description": "The security's ticker symbol, e.g. 'VTSAX', 'AAPL', 'BND'. The join/grouping key for a position. Use 'holdings.ticker.keyword' for exact term and aggregations, and the base field for full-text match. Query inside a 'nested' query on path 'holdings'." } The description is injected at runtime whenever the agent needs that field. It helps twice: once when building the query, and again when explaining the result, because the agent now knows what the data it retrieved actually is. This pattern is not OpenSearch-specific. If your agent queries a SQL database, a data warehouse, or any other data model, the same problem and the same fix apply. Problem 2 - too many indexes Real domains often span many indexes. In our demo domain, personal finance, there are 12: investments, transactions, accounts, budgets, recurring charges, net worth snapshots, bills, goals, merchants, categories, income, and insights. Each holds a different slice of the data. For every question, the agent first has to determine which indexes to even consider. And it is rarely just one. A single question often spans several indexes. "How much did I spend on subscriptions last month compared to my budget?" touches recurring charges, transactions, and budgets at once. So the agent cannot assume one question maps to one index. It has to discover every index the question needs and join the pieces together. Problem 3 - the mappings are too big The third problem makes the second one worse: some mappings are so large they break the model's context limit. You cannot just feed in every field of every index. Even when it fits, you would not want to. Context is not free; you pay by the token, and a context window full of irrelevant fields makes the model worse at picking the right ones. Frontier models now have enormous context windows, so why not just pour everything in? For starters, the model does its best work when you hand it a small, honed set of exactly the right fields, not a haystack it has to search first. Additionally, more context means higher costs, slower responses, and more chances to latch onto the wrong field. So we needed a way to surface only the fields that matter for each question - that constraint shaped the whole architecture. The architecture using two tiers The agent works in two tiers. Discovery comes first. It narrows the schema down in steps until it has a small, precise field reference, without ever loading the whole thing. Query and self-healing come second, building that reference into a query, running it, and repairing failures until it has an answer, or a "none found" when nothing fits. Each tier is its own agent with few tools and a strict prompt. That separation matters. A single agent that discovers, queries, heals, and answers all at once has too many jobs and drifts. Two focused agents, each with a single job, remain reliable. Tier 1 - progressive discovery The idea that makes this tier work is progressive disclosure : you do not load everything up front. You load a short, high-level summary first, and only pull in the details once you know you need them. If you have built agent skills, this will feel familiar. A skill shows the model a one-line description, and the full instructions load only when that skill is actually picked. We apply the same idea to every index. Each index carries two levels of description. A small one is a single sentence that says what the index holds and at what grain, cheap enough that the agent can scan all 12 at once and pick the right ones. A big one is the full detail: the field mappings, types, nested paths, and so on. The agent only ever reads the big description for the few indexes it has already chosen from the small ones. So the context window fills with what the question needs, and nothing else. The discovery agent narrows the search space in four steps, from the full catalogue down to the exact values it will filter on. It starts with listAvailableIndexes , which shows every index by what it holds, so the agent can pick the right ones by grain, not by name. For each index it picks, describeIndex reads the field mapping. Then describeFields resolves only the fields the query actually needs, in a single call reading the descriptions and details of the selected fields. Finally, sampleFieldValues pulls real values for those fields, so the agent sees the exact casing and the allowed options before it filters on anything. The search space shrinks at every step. From every index, to the chosen ones, to the needed fields, to the exact values. The last step, sampling values, earns its place. Suppose a field is an enum with 50 possible values. The agent needs to see real samples to learn the exact casing and the allowed options. Filtering on Crypto when the stored value is crypto returns zero hits, and the agent has no idea why. What discovery produces Take the question: "How much of my portfolio is in crypto right now?" Discovery runs its four tools and hands over a focused search space: - Target index: my investments , one doc per account, holdings nested - Fields that matter: holdings.assetClass keyword , holdings.marketValue float , holdings.ticker text - Sample values: holdings.assetClass : US Equity, Bond, ETF, Cash, Crypto, Commodity It carries only what this question needs. Budgets, bills, goals, merchants, and every unrelated field are left out. Complex questions may span more than one index, and the same flow handles that too. The discovery agent in practice Here is a simplified version of the discovery prompt: js const DISCOVERY PROMPT = You are the OpenSearch Discovery Agent. Your ONLY job is schema discovery: given a natural-language question, identify the single best index and the exact fields needed to answer it. You never write or run queries. WORKFLOW follow in order, stop as soon as you can answer : 1. listAvailableIndexes — pick the ONE index whose purpose/grain fits the question. 2. describeIndex — read that index's full mapping; note field types, .keyword subfields, and nested paths. 3. describeFields — confirm the specific fields you intend to use. 4. sampleFieldValues — for any field you'd filter on, verify the exact stored values casing/ enums . RULES: - Report the exact field paths from the mapping — never invent names. - Flag whether each chosen field is text/keyword/nested and whether it has a .keyword subfield. - Prefer the fewest fields that fully answer the question. - Do NOT compose DSL, aggregations, or run any query — discovery only. OUTPUT — a compact hand-off for the query agent: - index: