cd /news/ai-agents/mappings-don-t-carry-meaning-buildin… · home › topics › ai-agents › article
[ARTICLE · art-139840] src=nearform.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Mappings don't carry meaning: Building a research agent that actually understands your data

Nearform built a multi-agent research system that answers natural-language questions across an OpenSearch cluster indexing over 330,000 documents across roughly 30 indexes, after testing the OpenSearch MCP server and query-planning tool and finding that OpenSearch mappings carry no semantic meaning for an agent. The team presented the two-tier architecture at the Mastra TypeScript AI Conference in London in July 2026, following an eight-month build. The naive schema-to-DSL approach failed at scale because field names and types alone do not convey domain nuances such as tax-lot deduplication on ticker.keyword or month-end marketValue snapshots that differ between retirement and brokerage accounts.

by read23 min views19 publishedAug 27, 2026
Mappings don't carry meaning: Building a research agent that actually understands your data
Image: Nearform (auto-discovered)

How we taught an agent to answer natural-language questions across hundreds of thousands of documents, why the obvious approach failed, and the two-tier architecture that fixed it.

Many of the companies we work with at Nearform have the same starting point: they have years of data, ingested from many sources, sitting in a search cluster, sourced from multiple data stores or a GraphQL federation. Ideally, they want their users to just ask a question in plain English and get a correct answer back.

Over the last eight months, one of our teams built exactly that: a multi-agent system where a research agent answers natural-language questions across a large OpenSearch cluster. We presented this architecture at the Mastra TypeScript AI Conference in London in July 2026. This article tells the story of that build. What looked easy, what broke, and the system that made it work in production.

The promise: ask a question, get an answer #

On paper, the problem is simple. It takes three steps:

  1. The user asks a question. They do not need to know where the answer lives or how the data is structured.
  2. The agent builds an OpenSearch query. It reads the schema, converts the question into DSL, runs the query, and gets the data back.
  3. The agent crafts an answer. The data is analysed, summarised, and given back to the user in plain language.

That is the demo everyone has seen. A question goes in, an answer comes out, and it seems to work relatively well. Then, once you try to scale the amount of data, or the complexity of the data itself, this naive architecture just doesn't work anymore. To give a sense of scale, one real client system indexes over 330,000 OpenSearch documents, spread across roughly 30 separate indexes, with tens of thousands of PDF documents accessible via RAG search.

Our research agent is one piece of a larger system with a main orchestrator agent that routes requests to specialised subagents, backed by persistent memory, an observability platform, and a full evaluation pipeline. The piece we focus on here is the data-retrieval subagent, which communicates with OpenSearch via a proxy.

We did not jump straight to building our own architecture. First, we checked what already existed. If something off-the-shelf could answer questions across our cluster at scale, we would rather use it than build and maintain our own.

So we ran a short research phase to see what the state of the art looked like. We tested the OpenSearch MCP server, which exposes the cluster to an agent so it can list indexes, read mappings, and run queries on its own. We tested the OpenSearch query-planning tool, which turns a plain-English question into DSL. We also tested a few community tools that wrap OpenSearch for agents in similar ways.

Once we tried to make it work on real user queries, three problems showed up at once. The first is about meaning. The other two are about volume, and they share a single fix. In this article, we'll use a fictional financial example to explore each, but the same rules apply to pretty much every real-world domain.

Problem 1 - mappings carry no meaning #

OpenSearch mappings list field names and types. That is all. They do not say what a field actually means, how its data was computed, when to use it, or the small nuances only a domain expert knows.

We know what our fields mean because we built the system. The agent does not. Give it a field called marketValue of type double, and it will guess. It has no way to know the things we know, such as:

  • Each row in holdings is one tax lot, not one security. The same ticker appears multiple times. To count real positions, you have to deduplicate onticker.keyword .
  • marketValue is a month-end snapshot, not a live number. Retirement accounts report it net of loans, brokerage accounts report it gross. The values are not comparable across account types.
  • Transaction amounts are signed. Dividends and interest book to a synthetic CASH holding, not to the security that paid them.
  • Transaction dates come from a monthly load. The day is always 01, so ordering within a month is lost.

None of that lives in the mapping. So the agent guesses, picks the wrong fields, and writes queries that miss or return wrong data.

The fix: a meta-index with a drift-guard

We built a searchable, queryable meta-index with an entry for every field in the data. Next to each field's name and type, a subject-matter expert writes a clear, human description that conveys the meaning the field name does not. The agent reads it to learn what the field means, what values it can take, and how to query it correctly. The descriptions need upkeep, but they pay off. People maintain them over time, and in return the agent runs complex, domain-specific queries with confidence.

The upkeep sounds scary, but you can automate much of it. The main risk is that someone adds or changes a field in the mapping and forgets to update its description in the meta-index. So we guard against that in CI. A test compares the live mapping against the meta-index on every change. If a field exists in the mapping but has no matching entry, or an entry points to a field that no longer exists, the build fails. The machine cannot write the domain knowledge for you, but it can force a human to fill the gap before the change ships. In this way, the meta-index never drifts silently out of sync.

Here is what an entry looks like:

{
  "index": "my_investments",
  "field": "holdings.ticker",
  "type": "text",
  "nestedPaths": ["holdings"],
  "hasKeyword": true,
  "description": "The security's ticker symbol, e.g. 'VTSAX', 'AAPL', 'BND'. The join/grouping key for a position. Use 'holdings.ticker.keyword' for exact term and aggregations, and the base field for full-text match. Query inside a 'nested' query on path 'holdings'."
}

The description is injected at runtime whenever the agent needs that field. It helps twice: once when building the query, and again when explaining the result, because the agent now knows what the data it retrieved actually is.

This pattern is not OpenSearch-specific. If your agent queries a SQL database, a data warehouse, or any other data model, the same problem and the same fix apply.

Problem 2 - too many indexes #

Real domains often span many indexes. In our demo domain, personal finance, there are 12: investments, transactions, accounts, budgets, recurring charges, net worth snapshots, bills, goals, merchants, categories, income, and insights. Each holds a different slice of the data. For every question, the agent first has to determine which indexes to even consider.

And it is rarely just one. A single question often spans several indexes. "How much did I spend on subscriptions last month compared to my budget?" touches recurring charges, transactions, and budgets at once. So the agent cannot assume one question maps to one index. It has to discover every index the question needs and join the pieces together.

Problem 3 - the mappings are too big #

The third problem makes the second one worse: some mappings are so large they break the model's context limit. You cannot just feed in every field of every index. Even when it fits, you would not want to. Context is not free; you pay by the token, and a context window full of irrelevant fields makes the model worse at picking the right ones.

Frontier models now have enormous context windows, so why not just pour everything in? For starters, the model does its best work when you hand it a small, honed set of exactly the right fields, not a haystack it has to search first. Additionally, more context means higher costs, slower responses, and more chances to latch onto the wrong field.

So we needed a way to surface only the fields that matter for each question - that constraint shaped the whole architecture.

The architecture using two tiers #

The agent works in two tiers. Discovery comes first. It narrows the schema down in steps until it has a small, precise field reference, without ever the whole thing. Query and self-healing come second, building that reference into a query, running it, and repairing failures until it has an answer, or a "none found" when nothing fits.

Each tier is its own agent with few tools and a strict prompt. That separation matters. A single agent that discovers, queries, heals, and answers all at once has too many jobs and drifts. Two focused agents, each with a single job, remain reliable.

Tier 1 - progressive discovery #

The idea that makes this tier work is progressive disclosure: you do not load everything up front. You load a short, high-level summary first, and only pull in the details once you know you need them. If you have built agent skills, this will feel familiar. A skill shows the model a one-line description, and the full instructions load only when that skill is actually picked.

We apply the same idea to every index. Each index carries two levels of description. A small one is a single sentence that says what the index holds and at what grain, cheap enough that the agent can scan all 12 at once and pick the right ones. A big one is the full detail: the field mappings, types, nested paths, and so on. The agent only ever reads the big description for the few indexes it has already chosen from the small ones. So the context window fills with what the question needs, and nothing else.

The discovery agent narrows the search space in four steps, from the full catalogue down to the exact values it will filter on. It starts with listAvailableIndexes, which shows every index by what it holds, so the agent can pick the right ones by grain, not by name. For each index it picks, describeIndex reads the field mapping. Then describeFields resolves only the fields the query actually needs, in a single call reading the descriptions and details of the selected fields. Finally, sampleFieldValues pulls real values for those fields, so the agent sees the exact casing and the allowed options before it filters on anything.

The search space shrinks at every step. From every index, to the chosen ones, to the needed fields, to the exact values.

The last step, sampling values, earns its place. Suppose a field is an enum with 50 possible values. The agent needs to see real samples to learn the exact casing and the allowed options. Filtering on Crypto when the stored value is crypto returns zero hits, and the agent has no idea why.

What discovery produces

Take the question: "How much of my portfolio is in crypto right now?"

Discovery runs its four tools and hands over a focused search space:

  • Target index:my_investments , one doc per account, holdings nested
  • Fields that matter:holdings.assetClass (keyword),holdings.marketValue (float),holdings.ticker (text)
  • Sample values:holdings.assetClass : US_Equity, Bond, ETF, Cash, Crypto, Commodity

It carries only what this question needs. Budgets, bills, goals, merchants, and every unrelated field are left out. Complex questions may span more than one index, and the same flow handles that too.

The discovery agent in practice

Here is a simplified version of the discovery prompt:

const DISCOVERY_PROMPT = `You are the OpenSearch Discovery Agent. Your ONLY job is schema
discovery: given a natural-language question, identify the single best index and the exact
fields needed to answer it. You never write or run queries.

WORKFLOW (follow in order, stop as soon as you can answer):
1. listAvailableIndexes — pick the ONE index whose purpose/grain fits the question.
2. describeIndex — read that index's full mapping; note field types, .keyword subfields, and
nested paths.
3. describeFields — confirm the specific fields you intend to use.
4. sampleFieldValues — for any field you'd filter on, verify the exact stored values (casing/
enums).

RULES:
- Report the exact field paths from the mapping — never invent names.
- Flag whether each chosen field is text/keyword/nested and whether it has a .keyword subfield.
- Prefer the fewest fields that fully answer the question.
- Do NOT compose DSL, aggregations, or run any query — discovery only.

OUTPUT — a compact hand-off for the query agent:
- index: <chosen index>
- fields: <path — type — .keyword? — nested path (if any) — why it's needed>
- values: <exact sample values for each filtered field>
- notes: <object-vs-nested pitfalls, casing, or ambiguity>`;

And the agent itself, built with Mastra and running Claude on Amazon Bedrock:

export function createOpenSearchDiscoveryAgent(options: { awsRegion: string }) {
  const bedrock = createAmazonBedrock({
    region: options.awsRegion,
    credentialProvider: fromNodeProviderChain(),
  });

  return new Agent({
    id: "opensearch-discovery-agent",
    name: "OpenSearch Discovery Agent",
    description:
      "Read-only schema discovery for OpenSearch indexes: finds the best index " +
      "and the exact fields (types, .keyword subfields, nested paths, sample " +
      "values) needed to answer a question. Does not build or run queries.",
    instructions: DISCOVERY_PROMPT,
    model: bedrock(V2_SUBAGENT_MODEL_IDS.sonnet_4_6),
    tools: {
      listAvailableIndexes: listAvailableIndexesTool,
      describeIndex: describeIndexTool,
      describeFields: describeFieldsTool,
      sampleFieldValues: sampleFieldValuesTool,
    },
  });
}

Notice how little there is. Four read-only tools, one strict prompt, and a hard rule that this agent never writes or runs queries. The constraint is the design.

Tier 2 - query, execute, self-heal #

The execution phase was the trickiest part to get right. The execution agent receives the hand-off from discovery and works from broad to narrow:

  1. Run a broad query. Start with the core entity only, no filters.
  2. Response contains hits? If yes, apply the next filter from the target query and re-run. Each step narrows one filter closer to the fully filtered query. Keep going until every filter is applied and hits remain, then produce the business answer plus search notes.
  3. No hits, or an error? Diagnose, apply the fix, and re-run.

Starting broad may look wasteful, but it is the opposite. If you fire the full filtered query first and get zero hits, you learn nothing about which filter killed it. Adding filters one at a time means the moment results vanish, you know exactly which filter is wrong.

The self-healing playbook

As things break quite often in a non-deterministic system, the diagnose step has a playbook for each failure mode:

  • Wrong casing or value. Fix it once and re-run. If still empty, drop that filter and answer from the last result that had hits.
  • Malformed DSL or shard error. Correct the DSL and re-run. Never answer from the error.
  • Response too large. Leverage tools to manipulate the context in a code execution sandbox, with hard guardrails, so the context window does not get flooded.
  • Well-formed but empty. Report "none found". The absence is the answer.

Refusal is allowed only after the catalogue has been checked and nothing fits. The agent never gets to say "I don't have that data" based on the wording of the question alone. The tools we developed for self-healing explain to the agent what’s wrong and guide it in the self-healing process with specific instructions.

Teaching the model to write good DSL

The execution agent also carries a DSL_GUIDE prompt that teaches it OpenSearch query rules: when to use match versus term, why nested fields need a nested wrapper with the exact path from the mapping, why analysed text fields are lowercased at index time (the most common cause of zero hits), and why aggregation-only queries should set size: 0.

We did not write this from scratch. The OpenSearch ml-commons project includes an excellent query-planning prompt template in Java. We ported it into our Mastra setup and then refined it around the edge cases we hit in our own domain. If you are building anything similar, start there.

This is an extract from the full prompt:

Provide the OpenSearch _search API body as a JSON object (never a string).

QUERY TYPES
- match — full-text search on analyzed text fields.
- match_phrase — exact phrase; prefer for drug names and specific terms.
- term — exact value on keyword fields; use the .keyword subfield when the field is text-with-keyword.
- terms — multi-value exact match (OR).
- range — date/numeric comparisons (gt, gte, lt, lte).
- ... <SNIPPED>

OBJECT ARRAYS vs NESTED (check the mapping FIRST — wrong choice = 0 hits or a shard error / HTTP 500)
- Use a nested clause ONLY when the field's mapping type is literally "nested". Applying nested to a plain object array throws a shard error, not an empty-data signal.
- A plain object array is queried by its dotted leaf paths directly, no nested wrapper.
- ... <SNIPPED>

TEXT-FIELD CASING (most common 0-hit cause)
- Analyzed text fields are lowercased at index time; term/wildcard are NOT analyzed, so values must be lowercase to match.
- For exact original casing, target the .keyword subfield.

Permissions - keep them out of the agent #

Permissions are always hard to manage, and an agent with direct cluster access is a security incident waiting to happen. So our agent never talks to OpenSearch directly. Every query goes through a proxy:

Agent → OpenSearch proxy → OpenSearch cluster

The proxy handles role-based access control, rate limiting, a rules engine, session restrictions, business rules for filtering and enrichment, and runtime value overrides. We propagate the user's identity and permissions to it through Mastra's RequestContext. The agent stays simple and focuses on answering questions. The proxy enforces the rules in one place.

We built a harness specific to this system, drawing hard lines between what the model decides and what the code enforces. The agent gets to choose the parts that genuinely need judgement, like reading a question and deciding which fields answer it. But everything that can be done deterministically, we do deterministically, in code, outside the model. Permissions, rate limits, filtering rules, value overrides: none of that is left to the agent's discretion, because a hard rule enforced in code never drifts and never gets talked out of itself by a clever prompt.

A useful signal for this is your own system prompt. Every time you write an "always do X" or an "every time, Y" rule into the prompt, treat it as a smell. You are asking the model to reliably perform a mechanical task, which it will not always do. That is usually a chance to pull the rule out of the prompt and move it into the deterministic part of the system, where it holds every single time. The prompt gets shorter, the behaviour gets more reliable, and the model is left to do only what actually needs a model.

One technique we’ve used to improve the reliability of this system is explained in our article about re-anchoring agents via a transient reminder.

When a user lacks access to a field, the proxy does not fail the query - it strips the restricted values, keeps the data shape intact, and reports what it removed and why:

{
  "strippedFields": [
    "holdings.marketValue",
    "holdings.transactions.amount",
    "owners.fullName"
  ],
  "fieldAccess": [
    {
      "field": "holdings.marketValue",
      "reason": "Missing permission: premium feature not purchased",
      "code": "DENY_SCOPE_NOT_GRANTED"
    },
    {
      "field": "owners.fullName",
      "reason": "Missing permission: PII field requires ENTITLEMENT_PII_READ",
      "code": "DENY_ENTITLEMENT_MISSING"
    }
  ]
}

The hits come back with ticker and assetClass intact, but marketValue and owners.fullName set to null. The agent can then tell the user honestly: here is your answer, and here is what I could not include and why.

To be clear, stripping values but keeping the shape and telling the agent what was removed is a product decision, not the only option. In our case, part of the goal was upsell: when a user hits a field they do not have, we want the agent to say "this is a premium feature" so it becomes a natural prompt to upgrade. That only works if the agent knows the field exists and why it was blocked.

Your requirements may point the other way. If a field is sensitive, or its very existence leaks information, you may want the proxy to hide it entirely, so the agent never learns it exists. In that case, the meta-index and the discovery step are filtered by permission before the agent ever sees them, and the agent simply works with a smaller schema. Same proxy, different policy. Decide this based on your product and your security model, and the architecture will shift to match.

A complete real-world example #

Here is everything working together. A user on a Standard plan asks:

"What's my largest equity position across all my accounts, and how many distinct holdings do I have?"

The discovery agent:

  • listAvailableIndexes → chosemy_investments
  • describeIndex → six fields mapped
  • describeFields →holdings.ticker ,holdings.assetClass ,holdings.marketValue
  • sampleFieldValues → assetClass is one of {equity, fixed_income, cash}
  • Hand-off: index is my_investments ; group and dedup onholdings.ticker.keyword

The execution agent:

  • Built the DSL: size 0, nested terms aggregation on holdings.ticker.keyword
  • Validated the query, then executed it: 12 hits
  • The proxy stripped holdings.marketValue , so the agent ranked positions by lot count instead of dollar value, and noted it

The answer: your largest equity position is VOO, held in the Retirement Rollover IRA. Across all accounts, you hold seven distinct securities, deduplicated by ticker across 12 tax lots. A precise dollar ranking could not be computed, and the agent said so, with a note that valuations are limited by the user's subscription.

This one interaction shows the whole system: discovery finding the right fields, execution building and validating the query, the meta-index knowledge allowing the agent to deduplicate tax lots correctly, and the permissions proxy gracefully degrading while the agent stays honest about the limitation.

How we know it works: evaluating the query generation #

None of the architecture above earns trust on its own. An agent that writes OpenSearch queries can be wrong in ways that look right, so we measure it. Query generation is the hardest part of the whole system to evaluate, for one reason: the same question can produce many different but equally valid queries, and whether a query is correct depends on what is actually in the index. You cannot just diff the generated DSL against one "golden" query, because a different structure can return the same correct data.

So we do not grade the query. We grade the data it brings back. In our golden dataset, the expected result is a plain-language description of what correct retrieval looks like: which entities should come back, which fields, what the aggregation should total. It is neither a user-facing answer nor a fixed query. That lets any query shape score full marks as long as it retrieves the right data.

We check this in two places, and they work in tandem: in CI before anything ships, and on live traffic once it is running. CI has a golden answer to grade against, so it can score correctness. Production does not, so it grades a weaker but still useful signal: does the query even make sense?

In CI, we test against a frozen data snapshot. The OpenSearch data is version-controlled and fixed, so results are reproducible. This matters more than it sounds. An LLM judge is already non-deterministic, and if the underlying data also moved between runs, you would be stacking two sources of noise and could never tell which one caused a score to change. Freezing the data removes one of them, so a failing test points at the agent, not at the data. Two scorers run here:

  • Index Accuracy (deterministic). Did the agent query the right index? A portfolio question should hit the investments index, notmarket_rates . This is cheap, unambiguous, and catches routing mistakes before anything more expensive runs. It follows the same rule as the rest of the system: when the correct answer is exact and knowable, check it in code, not with a model.
  • Query Correctness (LLM-as-judge). This is the main CI scorer. Given the question and the data the agent retrieved, a judge model decides whether the agent retrieved the correct data, using that plain-language description of correct retrieval. We deliberately use a small, different model from the one that powers the agent, so the system does not grade its own homework.

In production, we score live traffic. Here there is no golden answer to compare against, so we cannot grade correctness. One scorer runs here:

  • Query Plausibility (LLM-as-judge). Instead of correctness, it asks whether the query even makes sense for the question: right fields, coherent filters, a structure that fits the question type, results that look related to what was asked. It does not pass or fail a request. It flags suspicious traces for a human to look at.

CI and production feed each other. When a production trace scores low on plausibility and the user is also unhappy, someone reviews it, determines what the correct data should have been, and adds it to the golden dataset. The live problem becomes a permanent CI test, so that class of failure cannot come back unnoticed. Monitoring detects new issues, and CI ensures old ones stay fixed.

If you want the full picture of how we build evals like these, including the scorers for the rest of the system, we wrote it up in From AI prototype to production: how to build evals for reliable agents.

Conclusion #

The demo version of "chat with your data" takes an afternoon to complete. The production version took a team of AI engineers eight months, and the difference is all in the parts you cannot see in a demo:

  • A meta-index that turns dead mappings into domain knowledge, maintained by the people who actually understand the data.
  • Progressive discovery, so the agent narrows the search space instead of drowning in it.
  • Two focused agents instead of one overloaded one.
  • A self-healing loop that treats failures as diagnosable, and "none found" as a valid answer.
  • A proxy that owns permissions, so the agent never has to.
  • An eval harness that grades the data, not the query, so the whole thing keeps earning trust as it changes.

If you are building a research agent over your own data, keep an eye out for the problems we hit, and steal whichever of these fixes apply. And if you could use the help of a partner experienced in taking large-scale agentic systems all the way to production, get in touch.

This article is based on the talk we gave at the Mastra TypeScript AI Conference. If you'd like to watch it in full, along with the other great talks from the event, the recording is below, queued up for our talk:

But wait - there's more. #

Nearform publishes real-world learnings on data & AI, engineering, and digital strategy - with more merged in weekly.

Insights

Perspectives on AI in engineering, product development, and strategy, for enterprise executives.

Community

Deep dives and tutorials by engineers, for engineers.

You may also like #

From AI prototype to production: how to build evals for reliable agents

Alfonso Graziano· 30 Mar 2026 · 10 min read

── more in #ai-agents 4 stories · sorted by recency
── more on @nearform 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mappings-don-t-carry…] indexed:0 read:23min 2026-08-27 · —