Building a RAG Chatbot with FastAPI and ChromaDB (that runs locally, no API key) A developer open-sourced a RAG chatbot starter template built with FastAPI and ChromaDB that runs entirely locally without an API key. The project demonstrates retrieval-augmented generation for document Q&A, using the same embedding model for indexing and querying to avoid common bugs. The template is available on GitHub under an MIT license. Everyone wants a chatbot that can answer questions from their own documents — a company handbook, a contract, a research paper. The technique behind this is called RAG Retrieval-Augmented Generation . I build production RAG and LLM systems for a living, and I kept re-wiring the same foundation on every project. So I cleaned it up into a small, open-source starter. This post walks through how RAG actually works, the design decisions that matter, and how you can run the whole thing for free — no API key required. At the end there's a link to the full open-source repo you can clone and run in minutes. An LLM like GPT or Claude only knows what it was trained on. Ask it about a document sitting on your laptop and it will guess — or hallucinate. RAG fixes this with a simple idea: retrieve the relevant information first, then let the LLM answer using only that. The result is answers grounded in your actual documents, with citations you can trace back to a specific page. Path 1 — Indexing when a document comes in : Path 2 — Querying when a user asks something : After building a few of these, here's a stack that's easy to start with and holds up in production: One detail people miss: you can test the whole pipeline without any API key , either in a retrieval-only mode that returns the matched chunks, or fully local with Ollama so nothing leaves your machine. Here's roughly what the query side looks like — embed the question, search, hand the chunks to the LLM: python def answer question question: str, top k: int = 3 : 1. Embed the question with the same model used at indexing time query embedding = embed question 2. Retrieve the most similar chunks from the vector store chunks = vector store.search query embedding, top k=top k 3. Build a prompt grounded in those chunks, then ask the LLM context = "\n\n".join c "text" for c in chunks answer = llm.generate question=question, context=context return answer, chunks chunks double as citations The key constraint: use the same embedding model for the question and the documents. Vectors from different models aren't comparable, and this is a surprisingly common bug. From experience, whether a RAG system answers well comes down to a few things people underestimate: I packaged all of this into a starter template, open-sourced under MIT. Clone it, run one Docker command, and you have a working document Q&A chatbot with source citations. It's multilingual, and runs free with no API key. 👉 GitHub: https://github.com/panutpl/rag-chatbot-template-starter https://github.com/panutpl/rag-chatbot-template-starter git clone https://github.com/panutpl/rag-chatbot-template-starter cd rag-chatbot-template-starter cp .env.example .env docker compose up --build Then open http://localhost:8000/docs , upload a PDF, and start asking questions. If you're building RAG or just getting started, I'd love feedback — especially on chunking and retrieval strategies, which I'm still tuning myself. Drop a comment about what's working for you.