# Build a RAG pipeline from scratch — the actually simple version

> Source: <https://dev.to/aiunplugged/build-a-rag-pipeline-from-scratch-the-actually-simple-version-gen>
> Published: 2026-10-10 17:10:17+00:00

Originally published on [aiunplugged.in](https://aiunplugged.in/blog/build-rag-pipeline-from-scratch/) — cross-posting for the Dev.to community.

Every RAG tutorial online starts with LangChain, LlamaIndex, or a hosted vector database signup. None of that is necessary to understand what a RAG pipeline actually is. This walkthrough builds one in a single Python file with three libraries — no framework, no cloud accounts, no config files.

RAG (Retrieval-Augmented Generation) is a pattern where an LLM answers a question using text pulled from a private document store at query time, not from what the model memorized during training. It exists because pretrained LLMs don't know a team's docs, a product's changelog, or last week's meeting notes, and fine-tuning every time the source changes is impractical.

RAG makes sense when at least one of these is true:

If the entire dataset fits in a single 200K-token context window (roughly 500 pages) and doesn't change between requests, RAG is often overkill — sending the whole corpus with the question is simpler and sometimes cheaper. That trade-off is worth its own walkthrough, coming later in this series.

Every RAG system, regardless of framework or vendor, does some version of these four steps:

Everything else — reranking, hybrid search, query rewriting, evaluation — is optional refinement on top of this loop.

Three libraries, one pip install:

```
pip install chromadb sentence-transformers anthropic
```

`chromadb` — a small local vector database (no server, no signup, in-memory or disk-backed)`sentence-transformers` — the embedding model, downloaded once on first run`anthropic` — client for Claude, used for the answer step
The embedding model (`all-MiniLM-L6-v2`) has 22 million parameters (~90 MB download) and runs on CPU in milliseconds. The Claude client needs an API key from `console.anthropic.com` — Haiku 4.5 is $1 per million input tokens, so a full test session costs cents.

Set the key as an env var:

```
export ANTHROPIC_API_KEY=sk-ant-...
```

A "chunk" is a small passage the embedding model can turn into one vector. Chunks that are too big (thousands of tokens) blur meaning across topics. Too small (single sentences) lose surrounding context. A good starting point: 200-400 words per chunk with a 40-word overlap between adjacent chunks.

For this tutorial the source is a short list of imaginary product-doc passages held in a Python list:

```
docs = [
    "Aurora API rate limits: free tier is 60 requests per minute. Paid tiers start at 600 rpm.",
    "Aurora API authentication uses a bearer token in the Authorization header.",
    "Aurora supports Python, Go, and Node.js SDKs. Community SDKs exist for Ruby and Rust.",
    "Aurora billing is per-request. Failed requests (5xx) are not charged. 4xx requests are charged.",
    "Aurora response latency: p50 is 120ms, p99 is 800ms across all regions.",
    "Aurora data residency: EU customers can pin storage to Frankfurt. US customers use us-east-1 by default.",
]
```

Each entry is already a self-contained chunk, so no splitting is needed here. For real documents (Markdown files, PDFs, HTML pages), a character-based splitter works for a first pass:

``` php
def chunk_text(text: str, size: int = 1500, overlap: int = 200) -> list[str]:
    """Split text into overlapping character windows. Cheap and predictable."""
    chunks = []
    start = 0
    while start < len(text):
        chunks.append(text[start:start + size])
        start += size - overlap
    return chunks
```

Character splitting is imperfect — it can cut mid-sentence — but for a first pipeline it beats spending an hour picking the "right" splitter. Better strategies (heading-aware, semantic, sentence-boundary) get their own deep-dive later in this series.

An embedding is a fixed-length vector that represents the meaning of a text passage. Chunks with similar meanings sit close together in vector space; unrelated chunks are far apart. That distance is the entire mechanism behind retrieval.

``` python
from sentence_transformers import SentenceTransformer

embedder = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = embedder.encode(docs).tolist()
```

`all-MiniLM-L6-v2` outputs 384-dimensional vectors. The `.tolist()` call converts the NumPy array into plain Python lists so Chroma can serialize them.

First run downloads the model to `~/.cache/huggingface/` — a one-time 22 MB download.

Chroma runs in-process. `EphemeralClient()` keeps everything in RAM (fine for tutorials and tests). `PersistentClient(path="./chroma_data")` writes to disk and survives restarts.

``` python
import chromadb

client = chromadb.EphemeralClient()
collection = client.create_collection(name="aurora_docs")

collection.add(
    ids=[f"doc_{i}" for i in range(len(docs))],
    documents=docs,
    embeddings=embeddings,
)
```

Every stored item needs a unique `id`. Chroma keeps the raw document text alongside the embedding, so retrieval later returns both the vector match and the original text to feed to the LLM.

Retrieval is symmetric with storage: embed the user's question with the same model, ask Chroma for the closest N chunks, and hand them to the LLM as context.

``` php
import anthropic

def answer(question: str, k: int = 3) -> str:
    q_embedding = embedder.encode([question]).tolist()

    results = collection.query(
        query_embeddings=q_embedding,
        n_results=k,
    )
    context = "\n\n".join(results["documents"][0])

    llm = anthropic.Anthropic()
    reply = llm.messages.create(
        model="claude-haiku-4-5-20251001",
        max_tokens=400,
        messages=[{
            "role": "user",
            "content": (
                "Answer the question using ONLY the context below. "
                "If the answer is not in the context, say so.\n\n"
                f"Context:\n{context}\n\n"
                f"Question: {question}"
            ),
        }],
    )
    return reply.content[0].text
```

Two details worth pinning down:

Everything above, in one file (`rag.py`):

``` python
import chromadb
import anthropic
from sentence_transformers import SentenceTransformer

docs = [
    "Aurora API rate limits: free tier is 60 requests per minute. Paid tiers start at 600 rpm.",
    "Aurora API authentication uses a bearer token in the Authorization header.",
    "Aurora supports Python, Go, and Node.js SDKs. Community SDKs exist for Ruby and Rust.",
    "Aurora billing is per-request. Failed requests (5xx) are not charged. 4xx requests are charged.",
    "Aurora response latency: p50 is 120ms, p99 is 800ms across all regions.",
    "Aurora data residency: EU customers can pin storage to Frankfurt. US customers use us-east-1 by default.",
]

embedder = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = embedder.encode(docs).tolist()

client = chromadb.EphemeralClient()
collection = client.create_collection(name="aurora_docs")
collection.add(
    ids=[f"doc_{i}" for i in range(len(docs))],
    documents=docs,
    embeddings=embeddings,
)

def answer(question: str, k: int = 3) -> str:
    q_embedding = embedder.encode([question]).tolist()
    results = collection.query(query_embeddings=q_embedding, n_results=k)
    context = "\n\n".join(results["documents"][0])
    llm = anthropic.Anthropic()
    reply = llm.messages.create(
        model="claude-haiku-4-5-20251001",
        max_tokens=400,
        messages=[{
            "role": "user",
            "content": (
                "Answer the question using ONLY the context below. "
                "If the answer is not in the context, say so.\n\n"
                f"Context:\n{context}\n\nQuestion: {question}"
            ),
        }],
    )
    return reply.content[0].text

print(answer("Are failed requests charged?"))
```

Expected output:

```
No — failed requests that return a 5xx status are not charged. 4xx requests are charged.
```

That is a complete RAG loop.

The pipeline above works, but a production RAG system adds five things:

`EphemeralClient()` loses data on restart. Swap for `PersistentClient(path=...)`, or graduate to Qdrant, pgvector, or Pinecone when the collection grows past a few hundred thousand chunks. A comparison of vector databases is coming next in this series.`{"source": "changelog", "date": "2026-08-01"}` to each chunk lets retrieval filter by tag or date range, not similarity alone.
Everything past those is optimization. The four-step loop above is the whole idea.

RAG is four steps: chunk, embed, store, retrieve+answer. Everything else — LangChain, LlamaIndex, hosted vector DBs, hybrid search, reranking, query rewriting — is scaffolding built on top of this loop. Once the version above works on real documents, layering the extras one at a time reveals exactly what each one changes and why.

If you've worked through this and are now wrestling with the production side (chunking that doesn't lose context, picking vector DBs, modeling costs at scale), I'm writing a handbook on exactly that — ships next month. [More here](https://aiunplugged.in/products/rag-production-handbook/).
