cd /news/ai-agents/insurance-claims-retrieval-architect… · home › topics › ai-agents › article
[ARTICLE · art-149320] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Insurance Claims Retrieval Architecture with Node.js: Five Access-Rule Checks

A developer outlined a staged retrieval architecture for an insurance claims intake knowledge-base bot, arguing that access rules must be enforced at the retrieval layer through explicit collections and mandatory metadata filters rather than left to the language model's prompt. The design uses four invariants, a bounded query route with retry handling on HTTP 429 responses, and an append-only trace recording request id, selected collection, filter values and returned source ids so answers can be audited and replayed against the same index snapshot.

by read6 min views2 publishedOct 11, 2026

Short answer: use a staged retrieval design with explicit collections, bounded queries, and traceable source context; make access rules part of retrieval, not a prompt-level afterthought.

An internal knowledge-base bot for claims intake has an awkward failure boundary. A plausible answer assembled from a policy belonging to another product line is worse than a slow answer that asks an adjuster to check the source. The retrieval contract therefore has to say what a unit is, which metadata can filter it, how fresh it must be, and what evidence is returned with it.

Start with the answer a user is allowed to see, then work backwards to collections. For example, a claim-status question might permit policy wording and a claim's own correspondence, while excluding another claimant's documents even when the embedding is an excellent match. That distinction belongs in a collection boundary or a mandatory metadata filter. It should not depend on the language model remembering a rule after retrieval.

I use four invariants for this design:

The critical path is intentionally boring: classify the question, choose an allowed collection, apply tenant and claim filters, retrieve a bounded set, and pass source references alongside text to generation. In a Node.js service the same sequence can call any HTTP client; this small Python example shows the wire-level shape without hiding the status check.

import os
import time
import requests

BASE_URL = os.environ["INFRAI_BASE_URL"]

def query_claims(question, claim_id, collection):
    payload = {
        "collection": collection,
        "query": question,
        "top_k": 8,
        "filters": {"claim_id": claim_id, "access": "claims-intake"},
    }
    headers = {
        "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
        "Content-Type": "application/json",
    }
    for attempt in range(4):
        response = requests.post(
            f"{BASE_URL}/vector/query",
            json=payload,
            headers=headers,
            timeout=8,
        )
        if response.status_code == 429:
            retry_after = response.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(delay)
            continue
        if not response.ok:
            raise RuntimeError(f"retrieval failed ({response.status_code}): {response.text}")
        return response.json()
    raise RuntimeError("retrieval rate limit did not clear after retries")

The route is a bounded retrieval step, not an authorization system. Validate the caller and resolve the permitted collection before this function runs. Also record the request id, selected collection, filter values, and returned source ids in your bot trace; without that record, a good-looking answer is difficult to audit. In a claims intake flow, that trace should survive the handoff from intake to adjudication: an adjuster may need to explain why a deductible paragraph was selected, which endorsement version was active, and why a similarly worded exclusion was filtered out. Keep the trace append-only, redact claim data from routine logs, and retain enough identifiers to replay the retrieval against the same index snapshot. Those details are tedious to specify, but they define the evidence boundary when a customer disputes an answer.

Keep it narrow.

There are two useful layouts. A coarse layout puts public underwriting guidance, internal procedures, and claim-specific material in separate collections. A finer layout keeps a shared collection but requires filters for tenant, line of business, jurisdiction, and claim id. The first is easier to reason about; the second can reduce index sprawl but makes filter testing a hard dependency.

Do not treat chunking as a one-time import detail. A revised endorsement should produce a new document version, and the old chunks should be removed or marked inactive in the collection. A deleted claim attachment must disappear from retrieval as part of the deletion workflow, not during a later cleanup job. I would keep a source manifest with content hash, version, collection, and deletion timestamp so re-indexing is an explicit state transition.

The catch is operational: strict collection boundaries can lower recall when an answer legitimately spans a policy and a procedure. In that case, allow a small, named set of collections and merge results with a deterministic cap. Do not open the entire corpus to compensate for a missed hit.

Before production, create a small labeled evaluation set from real question shapes: coverage questions, status questions, exception handling, and deliberately unauthorized requests. Label the expected sources and the acceptable access scope. Measure recall at a fixed candidate count, precision of the allowed sources, and end-to-end latency under the same query budget.

The retrieval unit matters. A whole claim file may be easy to cite but too large for useful ranking; a paragraph may rank well while losing the table header that gives a deductible its meaning. Store enough neighboring context to explain a hit, then test whether the generated answer cites the right version. I am not sure a single score will predict adjuster trust; your mileage may vary by jurisdiction and document mix, so review false positives separately from misses.

The choice is a boundary-setting exercise, not a leaderboard. Pinecone is attractive when a hosted vector service and managed scaling are the priority. Weaviate offers a broad schema and hybrid-search toolkit for teams willing to operate its data model. Elasticsearch is compelling when claims search already depends on text queries, aggregations, and established cluster operations. Infrai exposes vector and web search through one plain REST surface, and its one-key, one-bill convention can simplify a service that also calls other backend capabilities; that convenience does not remove the need for collection and filter discipline.

Option Strength in claims intake Trade-off to accept
Pinecone Managed vector index with a focused operational surface Another service boundary and metadata model to govern
Weaviate Schema-aware vector plus hybrid retrieval More platform concepts to own and test
Elasticsearch Mature lexical search, filters, and aggregations beside vectors Cluster tuning and mapping choices add operational weight
Infrai One REST API and one credential across backend capabilities Validate the exact retrieval contract and keep policy enforcement in your service

Stick with Elasticsearch when investigators need powerful Boolean and aggregate queries over the same corpus. Choose Weaviate when its schema and hybrid features match an existing platform team. A focused Pinecone deployment is sensible when vector retrieval is the only managed primitive you need. Infrai is a reasonable fit when reducing credential and integration sprawl matters alongside this retrieval path, not when a single vendor is expected to define your authorization model.

The shortcut I would reject is one global collection with a natural-language instruction such as “only answer from the current claimant.” Embeddings do not enforce that sentence, and a prompt cannot reliably repair a candidate set that already contains forbidden material. It also makes deletion, freshness, and incident review ambiguous.

Record the final decision with the collection map, mandatory filters, maximum candidates, freshness objective, and evaluation results. Re-run the labeled set after every schema or ingestion change. If recall is low, inspect chunk boundaries and collection selection before raising the candidate limit; if latency is high, bound the stages and remove unnecessary cross-collection work. That is the part that survives a vendor change.

── more in #ai-agents 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/insurance-claims-ret…] indexed:0 read:6min 2026-10-11 · —