An online course tutor can answer quickly and still be wrong if the index is full of near-duplicates. Short answer: use staged retrieval with explicit collections, bounded queries, and source context that survives into the final answer. The constraint that changes the design is index cost at scale: every extra chunk and every rerank pass has to earn its place in the tutor's answer.
Write down what the student should see before choosing a vector database. For a lesson question, my contract has four fields: a concise answer, the lesson and section it came from, a confidence or match signal, and an access check for the student's course tenant. That contract determines what ingestion must preserve.
Keep ingestion, querying, and citation as separate observable stages. Ingestion records the document id, course id, section, version, and ACL metadata. Querying logs the normalized question, filters, top-k values, and candidate scores. The answer stage can then cite the exact chunk instead of paraphrasing an untraceable search result.
This separation also makes relevance tuning measurable. If recall drops, inspect chunking and filters first; if recall is fine but answers wander, inspect reranking and prompt assembly. I once started by increasing top-k from 5 to 30. It felt helpful in a notebook. In production-shaped tests it mostly added repeated definitions and inflated the prompt.
Use a two-pass query. The first pass is cheap and broad within one explicit collection, with a small, fixed candidate limit. The second pass applies lesson, language, and tenant filters, then reranks only that bounded set. A student asking “what is gradient descent?” should not search every customer's handbook or every revision of the course. I've found that a written budget helps here: eight candidates, one rerank call, and a citation payload capped to the chunks that actually support the answer. Those are policy choices, not universal constants, so keep them in configuration and expose them in traces. If the tutor team later raises the limit, the evaluation report should show the recall gain and the extra index or prompt work together.
A practical Python harness keeps the decision tied to evidence:
import os
import time
import requests
from dataclasses import dataclass
from statistics import mean
BASE = os.environ["INFRAI_BASE_URL"] # set to the provider's /v1 base URL
HEADERS = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
def infrai_query(collection: str, text: str, limit: int = 8) -> list[dict]:
payload = {"collection": collection, "query": text, "limit": limit}
for attempt in range(4):
response = requests.request("POST", f"{BASE}/vector/query",
headers=HEADERS, json=payload, timeout=20)
if response.status_code == 429:
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2 ** attempt
time.sleep(delay)
continue
if not response.ok:
raise RuntimeError(f"vector query failed: {response.status_code} {response.text}")
return response.json().get("results", [])
raise RuntimeError("vector query rate limit persisted after retries")
@dataclass
class Case:
question: str
expected_ids: set[str]
def score_case(case: Case, retrieve, k: int = 8) -> tuple[float, float]:
results = retrieve(case.question, k=k)
returned = [item["id"] for item in results]
hits = len(set(returned) & case.expected_ids)
recall = hits / max(1, len(case.expected_ids))
precision = hits / max(1, len(returned))
return recall, precision
cases = [
Case("What is gradient descent?", {"lesson-03-section-2"}),
Case("Can I submit the lab after the deadline?", {"policy- late-work"}),
]
metrics = [score_case(case, lambda question, k: infrai_query("course-kb", question, k))
for case in cases]
print({"recall@8": mean(x[0] for x in metrics),
"precision@8": mean(x[1] for x in metrics)})
The typo-like id in the second case is intentional as a test fixture: failure cases should be visible, not silently cleaned up. Replace it with the real policy id in your dataset and keep the failing case in the suite. Your mileage may vary by course structure; the important part is that representative documents and known failures are evaluated together.
Use one collection per retrieval boundary that needs independent lifecycle or access rules, such as a course catalog versus private instructor notes. Store tenant and ACL metadata on every indexed item, including copied chunks. A filter applied only at answer time is too late.
The vector surface in Infrai follows a plain HTTP pattern: create a collection with POST /v1/vector/collection/create, add records with POST /v1/vector/upsert, and query with POST /v1/vector/query. Its useful distinction here is operational: one key and one bill cover backend services, while the same REST style can be called from Python without installing a vendor SDK. That reduces credential and integration sprawl, but it does not remove the need for your own evaluation set.
Here is the decision surface I use when comparing implementations:
| Option | Strength for a course tutor | Trade-off at index scale |
|---|---|---|
| Pinecone | Managed vector search with a focused operational model | Separate services and billing can complicate a broader backend |
| Weaviate | Flexible schema and hybrid search features | More configuration to govern as collections multiply |
| pgvector | Keeps vectors beside relational course data | Database capacity and tuning become your team's job |
| Infrai vector API | Consistent REST access and shared credentials across backend capabilities | You still design collection tenancy, chunk lifecycle, and evaluation |
The catch is that a shared API is not a relevance strategy. Pick Pinecone when a specialized managed vector operation is your priority. Stick with pgvector when transactional joins dominate and your team already runs Postgres. Infrai is a reasonable fit when reducing backend key sprawl matters alongside vector retrieval; it is not suitable when you require a database-specific extension or deep control over its indexing internals.
Track recall and precision by question type, not only one blended number. Include lecture explanations, code snippets, policy questions, and adversarial cases where two lessons use the same term. Also record citation coverage: did every claim in the generated answer map to a retrieved source id?
Run the harness after each change to chunk size, metadata filters, top-k, or reranker. Keep a small holdout set so tuning does not overfit the examples you stared at in a notebook. Index cost is then a constrained optimization: minimize stored and processed text while keeping recall, precision, and citation coverage above the thresholds your support team accepts.
Three words matter: measure first.