{"slug": "build-a-rag-pipeline-from-scratch-the-actually-simple-version", "title": "Build a RAG pipeline from scratch — the actually simple version", "summary": "A developer published a walkthrough showing how to build a retrieval-augmented generation (RAG) pipeline from scratch in a single Python file using only three libraries — ChromaDB, sentence-transformers, and the Anthropic client — with no framework, cloud accounts, or config files. The tutorial walks through the four core RAG steps (chunking, embedding, retrieval, generation), recommends 200–400 word chunks with 40-word overlap, and uses the all-MiniLM-L6-v2 embedding model running on CPU. It argues that frameworks like LangChain and LlamaIndex are unnecessary for understanding the underlying pattern.", "body_md": "Originally published on [aiunplugged.in](https://aiunplugged.in/blog/build-rag-pipeline-from-scratch/) — cross-posting for the Dev.to community.\n\nEvery RAG tutorial online starts with LangChain, LlamaIndex, or a hosted vector database signup. None of that is necessary to understand what a RAG pipeline actually is. This walkthrough builds one in a single Python file with three libraries — no framework, no cloud accounts, no config files.\n\nRAG (Retrieval-Augmented Generation) is a pattern where an LLM answers a question using text pulled from a private document store at query time, not from what the model memorized during training. It exists because pretrained LLMs don't know a team's docs, a product's changelog, or last week's meeting notes, and fine-tuning every time the source changes is impractical.\n\nRAG makes sense when at least one of these is true:\n\nIf the entire dataset fits in a single 200K-token context window (roughly 500 pages) and doesn't change between requests, RAG is often overkill — sending the whole corpus with the question is simpler and sometimes cheaper. That trade-off is worth its own walkthrough, coming later in this series.\n\nEvery RAG system, regardless of framework or vendor, does some version of these four steps:\n\nEverything else — reranking, hybrid search, query rewriting, evaluation — is optional refinement on top of this loop.\n\nThree libraries, one pip install:\n\n```\npip install chromadb sentence-transformers anthropic\n```\n\n`chromadb` — a small local vector database (no server, no signup, in-memory or disk-backed)`sentence-transformers` — the embedding model, downloaded once on first run`anthropic` — client for Claude, used for the answer step\nThe embedding model (`all-MiniLM-L6-v2`) has 22 million parameters (~90 MB download) and runs on CPU in milliseconds. The Claude client needs an API key from `console.anthropic.com` — Haiku 4.5 is $1 per million input tokens, so a full test session costs cents.\n\nSet the key as an env var:\n\n```\nexport ANTHROPIC_API_KEY=sk-ant-...\n```\n\nA \"chunk\" is a small passage the embedding model can turn into one vector. Chunks that are too big (thousands of tokens) blur meaning across topics. Too small (single sentences) lose surrounding context. A good starting point: 200-400 words per chunk with a 40-word overlap between adjacent chunks.\n\nFor this tutorial the source is a short list of imaginary product-doc passages held in a Python list:\n\n```\ndocs = [\n    \"Aurora API rate limits: free tier is 60 requests per minute. Paid tiers start at 600 rpm.\",\n    \"Aurora API authentication uses a bearer token in the Authorization header.\",\n    \"Aurora supports Python, Go, and Node.js SDKs. Community SDKs exist for Ruby and Rust.\",\n    \"Aurora billing is per-request. Failed requests (5xx) are not charged. 4xx requests are charged.\",\n    \"Aurora response latency: p50 is 120ms, p99 is 800ms across all regions.\",\n    \"Aurora data residency: EU customers can pin storage to Frankfurt. US customers use us-east-1 by default.\",\n]\n```\n\nEach entry is already a self-contained chunk, so no splitting is needed here. For real documents (Markdown files, PDFs, HTML pages), a character-based splitter works for a first pass:\n\n``` php\ndef chunk_text(text: str, size: int = 1500, overlap: int = 200) -> list[str]:\n    \"\"\"Split text into overlapping character windows. Cheap and predictable.\"\"\"\n    chunks = []\n    start = 0\n    while start < len(text):\n        chunks.append(text[start:start + size])\n        start += size - overlap\n    return chunks\n```\n\nCharacter splitting is imperfect — it can cut mid-sentence — but for a first pipeline it beats spending an hour picking the \"right\" splitter. Better strategies (heading-aware, semantic, sentence-boundary) get their own deep-dive later in this series.\n\nAn embedding is a fixed-length vector that represents the meaning of a text passage. Chunks with similar meanings sit close together in vector space; unrelated chunks are far apart. That distance is the entire mechanism behind retrieval.\n\n``` python\nfrom sentence_transformers import SentenceTransformer\n\nembedder = SentenceTransformer(\"all-MiniLM-L6-v2\")\nembeddings = embedder.encode(docs).tolist()\n```\n\n`all-MiniLM-L6-v2` outputs 384-dimensional vectors. The `.tolist()` call converts the NumPy array into plain Python lists so Chroma can serialize them.\n\nFirst run downloads the model to `~/.cache/huggingface/` — a one-time 22 MB download.\n\nChroma runs in-process. `EphemeralClient()` keeps everything in RAM (fine for tutorials and tests). `PersistentClient(path=\"./chroma_data\")` writes to disk and survives restarts.\n\n``` python\nimport chromadb\n\nclient = chromadb.EphemeralClient()\ncollection = client.create_collection(name=\"aurora_docs\")\n\ncollection.add(\n    ids=[f\"doc_{i}\" for i in range(len(docs))],\n    documents=docs,\n    embeddings=embeddings,\n)\n```\n\nEvery stored item needs a unique `id`. Chroma keeps the raw document text alongside the embedding, so retrieval later returns both the vector match and the original text to feed to the LLM.\n\nRetrieval is symmetric with storage: embed the user's question with the same model, ask Chroma for the closest N chunks, and hand them to the LLM as context.\n\n``` php\nimport anthropic\n\ndef answer(question: str, k: int = 3) -> str:\n    q_embedding = embedder.encode([question]).tolist()\n\n    results = collection.query(\n        query_embeddings=q_embedding,\n        n_results=k,\n    )\n    context = \"\\n\\n\".join(results[\"documents\"][0])\n\n    llm = anthropic.Anthropic()\n    reply = llm.messages.create(\n        model=\"claude-haiku-4-5-20251001\",\n        max_tokens=400,\n        messages=[{\n            \"role\": \"user\",\n            \"content\": (\n                \"Answer the question using ONLY the context below. \"\n                \"If the answer is not in the context, say so.\\n\\n\"\n                f\"Context:\\n{context}\\n\\n\"\n                f\"Question: {question}\"\n            ),\n        }],\n    )\n    return reply.content[0].text\n```\n\nTwo details worth pinning down:\n\nEverything above, in one file (`rag.py`):\n\n``` python\nimport chromadb\nimport anthropic\nfrom sentence_transformers import SentenceTransformer\n\ndocs = [\n    \"Aurora API rate limits: free tier is 60 requests per minute. Paid tiers start at 600 rpm.\",\n    \"Aurora API authentication uses a bearer token in the Authorization header.\",\n    \"Aurora supports Python, Go, and Node.js SDKs. Community SDKs exist for Ruby and Rust.\",\n    \"Aurora billing is per-request. Failed requests (5xx) are not charged. 4xx requests are charged.\",\n    \"Aurora response latency: p50 is 120ms, p99 is 800ms across all regions.\",\n    \"Aurora data residency: EU customers can pin storage to Frankfurt. US customers use us-east-1 by default.\",\n]\n\nembedder = SentenceTransformer(\"all-MiniLM-L6-v2\")\nembeddings = embedder.encode(docs).tolist()\n\nclient = chromadb.EphemeralClient()\ncollection = client.create_collection(name=\"aurora_docs\")\ncollection.add(\n    ids=[f\"doc_{i}\" for i in range(len(docs))],\n    documents=docs,\n    embeddings=embeddings,\n)\n\ndef answer(question: str, k: int = 3) -> str:\n    q_embedding = embedder.encode([question]).tolist()\n    results = collection.query(query_embeddings=q_embedding, n_results=k)\n    context = \"\\n\\n\".join(results[\"documents\"][0])\n    llm = anthropic.Anthropic()\n    reply = llm.messages.create(\n        model=\"claude-haiku-4-5-20251001\",\n        max_tokens=400,\n        messages=[{\n            \"role\": \"user\",\n            \"content\": (\n                \"Answer the question using ONLY the context below. \"\n                \"If the answer is not in the context, say so.\\n\\n\"\n                f\"Context:\\n{context}\\n\\nQuestion: {question}\"\n            ),\n        }],\n    )\n    return reply.content[0].text\n\nprint(answer(\"Are failed requests charged?\"))\n```\n\nExpected output:\n\n```\nNo — failed requests that return a 5xx status are not charged. 4xx requests are charged.\n```\n\nThat is a complete RAG loop.\n\nThe pipeline above works, but a production RAG system adds five things:\n\n`EphemeralClient()` loses data on restart. Swap for `PersistentClient(path=...)`, or graduate to Qdrant, pgvector, or Pinecone when the collection grows past a few hundred thousand chunks. A comparison of vector databases is coming next in this series.`{\"source\": \"changelog\", \"date\": \"2026-08-01\"}` to each chunk lets retrieval filter by tag or date range, not similarity alone.\nEverything past those is optimization. The four-step loop above is the whole idea.\n\nRAG is four steps: chunk, embed, store, retrieve+answer. Everything else — LangChain, LlamaIndex, hosted vector DBs, hybrid search, reranking, query rewriting — is scaffolding built on top of this loop. Once the version above works on real documents, layering the extras one at a time reveals exactly what each one changes and why.\n\nIf you've worked through this and are now wrestling with the production side (chunking that doesn't lose context, picking vector DBs, modeling costs at scale), I'm writing a handbook on exactly that — ships next month. [More here](https://aiunplugged.in/products/rag-production-handbook/).", "url": "https://wpnews.pro/news/build-a-rag-pipeline-from-scratch-the-actually-simple-version", "canonical_source": "https://dev.to/aiunplugged/build-a-rag-pipeline-from-scratch-the-actually-simple-version-gen", "published_at": "2026-10-10 17:10:17+00:00", "updated_at": "2026-10-10 17:18:49.384045+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "generative-ai", "developer-tools"], "entities": ["ChromaDB", "sentence-transformers", "Anthropic", "Claude", "all-MiniLM-L6-v2", "LangChain", "LlamaIndex", "aiunplugged.in"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/build-a-rag-pipeline-from-scratch-the-actually-simple-version", "markdown": "https://wpnews.pro/news/build-a-rag-pipeline-from-scratch-the-actually-simple-version.md", "text": "https://wpnews.pro/news/build-a-rag-pipeline-from-scratch-the-actually-simple-version.txt", "jsonld": "https://wpnews.pro/news/build-a-rag-pipeline-from-scratch-the-actually-simple-version.jsonld"}}