I Built a Python SDK to Debug RAG Pipelines A developer released rag-debugger, an open-source MIT-licensed Python SDK that intercepts and inspects retrieval-augmented generation (RAG) pipelines to surface missing knowledge-base coverage. The tool, installable via "pip install rag-debugger-amine," decomposes queries into atomic sub-intents, scores each independently, and routes borderline scores between 0.60 and 0.75 through a reranker LLM call; its GapDetector defaults to a 0.65 threshold. rag-debugger uses Google Gemini by default with the gemini-3.5-flash-lite LLM and gemini-embedding-001 embeddings, supports LangChain, LlamaIndex, and custom retrievers, and ships a local dashboard at http://localhost:7842 that auto-refreshes every 10 seconds. Intercept, inspect, and fix your RAG retrieval pipeline. Most RAG bugs aren't in your code — they're in your retrieval. Wrong chunks get selected, knowledge gaps go undetected, and you find out when users complain. rag-debugger gives you visibility into exactly what your vector DB returned, why it won, and what's missing from your knowledge base. pip install rag-debugger-amine For the local dashboard: pip install rag-debugger-amine dashboard python import rag debugger as rd rd.init project="my-rag-app" retriever = rd.wrap retriever your retriever Find what your knowledge base is missing before your users do: python from rag debugger import GeminiClient, GapDetector client = GeminiClient set GEMINI API KEY env var detector = GapDetector client default threshold is 0.65 chunks = your retriever.get relevant documents query report = detector.analyze query, {"content": c.page content} for c in chunks print report GAP DETECTED coverage=50% worst score=0.64 priority=0.50 Missing: refund policy, iOS-specific cancellation Fix: Add docs covering refund eligibility and iOS cancellation flow. ✓ 0.71 how to cancel ✗ 0.64 how to get a refund A query like "cancel my iOS subscription and get a refund" is really four questions. Standard RAG scores the whole query — if cancellation chunks score high, the query looks covered. rag-debugger decomposes it into atomic sub-intents and scores each one independently, so a missing refund policy is always caught even when the cancellation docs are excellent. Borderline scores 0.60–0.75 are passed through a reranker — a lightweight LLM call that asks "does this chunk actually answer this question?" — so semantically similar but irrelevant chunks don't pass as covered. Group multi-turn conversations under a single session to get a summary of retrieval quality across the whole interaction: with rd.session id="conv-123", user="user-42" as s: retriever.get relevant documents "first query" retriever.get relevant documents "follow-up query" summary = s.summary print summary Session conv-123 duration: 430ms events: 2 avg score: 0.741 worst score: 0.677 gaps: 0 / 2 Sessions are thread-safe — concurrent requests in a web app won't bleed into each other. Visualize retrieval events, chunk scores, and gap flags in a local web UI: rd.dashboard opens http://localhost:7842 The dashboard shows: - Per-session summary — avg score, worst score, gap count - Per-event chunk score bars with content preview - Gap flags with missing topics and fix suggestions - Auto-refreshes every 10 seconds Works with LangChain, LlamaIndex, and any custom pipeline: LangChain retriever = rd.wrap retriever vectorstore.as retriever , label="docs" LlamaIndex retriever = rd.wrap retriever index.as retriever Custom object retriever = rd.wrap retriever my retriever, method="fetch docs" rag-debugger uses Google Gemini by default free tier via Google AI Studio https://aistudio.google.com : python from rag debugger import GeminiClient client = GeminiClient api key="..." or set GEMINI API KEY env var Models used: - LLM: gemini-3.5-flash-lite - Embeddings: gemini-embedding-001 MIT