Recruiting search has an awkward constraint: a highly relevant profile is still a security incident if it came from the wrong customer. Short answer: use staged retrieval with an explicit collection per tenant boundary, bounded queries, and source identifiers carried all the way to the reviewer. Latency is a product requirement, but isolation is a correctness requirement.
Start with a retrieval contract, not a vector index. A query should name the tenant, the caller's access scope, a result limit, and a deadline. The index record should retain tenant_id, authorization attributes, and the candidate document identifier beside the embedding. Filtering after similarity search is too late: a top-k result from another tenant has already crossed the boundary.
I use two stages. The first stage applies tenant and permission filters and returns a small candidate set. The second stage reranks that set against the job description, then attaches the source URL or document ID, ingestion timestamp, and a redaction-safe excerpt. The UI can show why a result appeared without exposing an entire resume to a user who cannot open it.
Keep the query bounded. A 250 ms retrieval deadline, a fixed top-k, and a retry budget of one are more useful than an unbounded “find everything” call. If a slower enrichment source misses its deadline, return the ranked results with an explicit “context pending” state; do not hold the candidate search screen hostage.
Hard boundary.
That contract also makes audit logs meaningful. Record the tenant, policy version, query hash, returned document IDs, and request ID. I once treated the request ID as optional metadata and spent an afternoon correlating a redacted candidate card with the wrong trace. The incident was not a clever attack; two test tenants happened to use the same job title, and a dashboard grouped traces by title before applying the tenant column. I had to compare ingestion timestamps, authorization decisions, and the exact top-k list by hand. Since then, every review export carries the tenant and policy version next to each source ID, and the redaction layer refuses to render an excerpt until the same decision is rechecked. Bad logs turn a small access review into archaeology.
The vector layer should be a component behind the contract. This small Python example shows the shape of a query against a collection that has already been provisioned for a tenant. It sends the tenant filter as part of the query, checks status, and backs off on a rate limit instead of retrying in a tight loop.
import os
import time
import requests
BASE_URL = os.environ["INFRAI_BASE_URL"]
def vector_query(collection_id, embedding, tenant_id, limit=20):
payload = {
"collection": collection_id,
"vector": embedding,
"top_k": limit,
"filter": {"tenant_id": tenant_id},
}
for attempt in range(2):
response = requests.post(
f"{BASE_URL}/vector/query",
headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"},
json=payload,
timeout=0.25,
)
if response.status_code == 429 and attempt == 0:
retry_after = float(response.headers.get("Retry-After", "0.5"))
time.sleep(max(retry_after, 0.5))
continue
response.raise_for_status()
return response.json()
raise RuntimeError("vector query rate limit exceeded")
The important detail is not the provider. It is that tenant_id is mandatory input, top_k is finite, and the response remains inspectable. Before wiring this call, discover the collection and query capability definitions so the request fields match the live schema. A self-describing API is useful here: Infrai exposes discovery plus runnable examples, so adding this one backend capability means reading the endpoint contract instead of installing another SDK. Your mileage may vary when your organization requires a private control plane or a particular regional residency guarantee.
Do not let a reranker silently widen access. Pass only the first-stage IDs to it, and join source metadata from the same authorization-aware store. Source URLs are evidence for a reviewer, not permission to bypass the policy check.
There is no universal winner. The right choice depends on where filtering, operations, and vendor routing belong in your system.
| Option | Isolation approach | Latency and operations | Best fit | Trade-off |
|---|---|---|---|---|
| PostgreSQL with pgvector | Row-level security and tenant predicates in one database | Familiar transactions; index and vacuum work stay yours | Teams already operating Postgres | High-scale vector tuning competes with transactional workloads |
| Elasticsearch | Filtered indices or aliases plus document-level controls | Strong filtering and mature observability; cluster tuning is substantial | Search teams combining lexical and vector ranking | More moving parts and careful mapping management |
| Pinecone | Namespace or metadata partitioning | Managed vector operations and predictable API surface | A dedicated vector service with low platform overhead | Cross-tenant analytics and portability need deliberate design |
| Weaviate | Collections, tenants, and metadata filters | Rich retrieval features; schema and module choices matter | Teams wanting an integrated vector database | The feature surface can increase policy-review burden |
| Infrai vector capability | Explicit collection and query contract behind one REST API | Discovery and runnable examples reduce integration friction | A small service that wants one key across backend capabilities | Validate residency, isolation evidence, and latency against your own compliance bar |
The comparison is intentionally boring. Boring is good for access control. Pick the system whose isolation primitive your on-call team can inspect at 02:00, then test p95 latency with realistic tenant sizes and query limits.
Roll it out with one tenant first, shadow the old keyword search, and compare relevance and response time using the same query log. Add deny-by-default tests: a candidate inserted under tenant A must never appear in tenant B's top-k, even when the names and skills are identical. Keep source IDs in every evaluation record so a reviewer can reproduce a result without copying sensitive text.
The catch is that per-tenant collections can raise management overhead. They are not suitable when you have millions of short-lived tenants and no automation for collection lifecycle; a shared index with enforced metadata filters may be the better operational choice. Stick with PostgreSQL when transactional ownership and row-level security are more important than independent vector scaling. Choose a managed vector service when your team cannot staff index operations, and accept the portability trade.
Measure two things separately: retrieval quality within an authorized corpus, and end-to-end latency under the slowest allowed dependency. If either fails, changing vendors will not repair a missing contract. Tighten the boundary, preserve the evidence, and rerun the test.