cd /news/ai-search/candidate-search-architecture-tenant… · home › topics › ai-search › article
[ARTICLE · art-149321] src=dev.to ↗ pub= topic=ai-search verified=true sentiment=· neutral

Candidate Search Architecture: Tenant-Bounded Retrieval and Source Traceability

A developer describes a two-stage, tenant-bounded retrieval architecture for recruiting search that enforces isolation as a correctness requirement rather than a latency trade-off. The approach applies tenant and permission filters before similarity search, carries source identifiers through to the reviewer, and logs tenant, policy version, query hash, and returned document IDs for auditability. The developer recounts a mislabeled-trace incident caused by two test tenants sharing a job title, after which every review export carries tenant and policy version beside each source ID.

by read5 min views1 publishedOct 11, 2026

Recruiting search has an awkward constraint: a highly relevant profile is still a security incident if it came from the wrong customer. Short answer: use staged retrieval with an explicit collection per tenant boundary, bounded queries, and source identifiers carried all the way to the reviewer. Latency is a product requirement, but isolation is a correctness requirement.

Start with a retrieval contract, not a vector index. A query should name the tenant, the caller's access scope, a result limit, and a deadline. The index record should retain tenant_id, authorization attributes, and the candidate document identifier beside the embedding. Filtering after similarity search is too late: a top-k result from another tenant has already crossed the boundary.

I use two stages. The first stage applies tenant and permission filters and returns a small candidate set. The second stage reranks that set against the job description, then attaches the source URL or document ID, ingestion timestamp, and a redaction-safe excerpt. The UI can show why a result appeared without exposing an entire resume to a user who cannot open it.

Keep the query bounded. A 250 ms retrieval deadline, a fixed top-k, and a retry budget of one are more useful than an unbounded “find everything” call. If a slower enrichment source misses its deadline, return the ranked results with an explicit “context pending” state; do not hold the candidate search screen hostage.

Hard boundary.

That contract also makes audit logs meaningful. Record the tenant, policy version, query hash, returned document IDs, and request ID. I once treated the request ID as optional metadata and spent an afternoon correlating a redacted candidate card with the wrong trace. The incident was not a clever attack; two test tenants happened to use the same job title, and a dashboard grouped traces by title before applying the tenant column. I had to compare ingestion timestamps, authorization decisions, and the exact top-k list by hand. Since then, every review export carries the tenant and policy version next to each source ID, and the redaction layer refuses to render an excerpt until the same decision is rechecked. Bad logs turn a small access review into archaeology.

The vector layer should be a component behind the contract. This small Python example shows the shape of a query against a collection that has already been provisioned for a tenant. It sends the tenant filter as part of the query, checks status, and backs off on a rate limit instead of retrying in a tight loop.

import os
import time
import requests

BASE_URL = os.environ["INFRAI_BASE_URL"]

def vector_query(collection_id, embedding, tenant_id, limit=20):
    payload = {
        "collection": collection_id,
        "vector": embedding,
        "top_k": limit,
        "filter": {"tenant_id": tenant_id},
    }
    for attempt in range(2):
        response = requests.post(
            f"{BASE_URL}/vector/query",
            headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"},
            json=payload,
            timeout=0.25,
        )
        if response.status_code == 429 and attempt == 0:
            retry_after = float(response.headers.get("Retry-After", "0.5"))
            time.sleep(max(retry_after, 0.5))
            continue
        response.raise_for_status()
        return response.json()
    raise RuntimeError("vector query rate limit exceeded")

The important detail is not the provider. It is that tenant_id is mandatory input, top_k is finite, and the response remains inspectable. Before wiring this call, discover the collection and query capability definitions so the request fields match the live schema. A self-describing API is useful here: Infrai exposes discovery plus runnable examples, so adding this one backend capability means reading the endpoint contract instead of installing another SDK. Your mileage may vary when your organization requires a private control plane or a particular regional residency guarantee.

Do not let a reranker silently widen access. Pass only the first-stage IDs to it, and join source metadata from the same authorization-aware store. Source URLs are evidence for a reviewer, not permission to bypass the policy check.

There is no universal winner. The right choice depends on where filtering, operations, and vendor routing belong in your system.

Option Isolation approach Latency and operations Best fit Trade-off
PostgreSQL with pgvector Row-level security and tenant predicates in one database Familiar transactions; index and vacuum work stay yours Teams already operating Postgres High-scale vector tuning competes with transactional workloads
Elasticsearch Filtered indices or aliases plus document-level controls Strong filtering and mature observability; cluster tuning is substantial Search teams combining lexical and vector ranking More moving parts and careful mapping management
Pinecone Namespace or metadata partitioning Managed vector operations and predictable API surface A dedicated vector service with low platform overhead Cross-tenant analytics and portability need deliberate design
Weaviate Collections, tenants, and metadata filters Rich retrieval features; schema and module choices matter Teams wanting an integrated vector database The feature surface can increase policy-review burden
Infrai vector capability Explicit collection and query contract behind one REST API Discovery and runnable examples reduce integration friction A small service that wants one key across backend capabilities Validate residency, isolation evidence, and latency against your own compliance bar

The comparison is intentionally boring. Boring is good for access control. Pick the system whose isolation primitive your on-call team can inspect at 02:00, then test p95 latency with realistic tenant sizes and query limits.

Roll it out with one tenant first, shadow the old keyword search, and compare relevance and response time using the same query log. Add deny-by-default tests: a candidate inserted under tenant A must never appear in tenant B's top-k, even when the names and skills are identical. Keep source IDs in every evaluation record so a reviewer can reproduce a result without copying sensitive text.

The catch is that per-tenant collections can raise management overhead. They are not suitable when you have millions of short-lived tenants and no automation for collection lifecycle; a shared index with enforced metadata filters may be the better operational choice. Stick with PostgreSQL when transactional ownership and row-level security are more important than independent vector scaling. Choose a managed vector service when your team cannot staff index operations, and accept the portability trade.

Measure two things separately: retrieval quality within an authorized corpus, and end-to-end latency under the slowest allowed dependency. If either fails, changing vendors will not repair a missing contract. Tighten the boundary, preserve the evidence, and rerun the test.

── more in #ai-search 4 stories · sorted by recency
── more on @infrai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/candidate-search-arc…] indexed:0 read:5min 2026-10-11 · —