{"slug": "sledtrace-a-local-debugger-that-shows-why-your-rag-app-answered-wrong", "title": "SledTrace – A local debugger that shows why your RAG app answered wrong", "summary": "SledTrace, a local Python debugger for RAG pipelines, has been released and is installable via `pip install sledtrace`, with the dashboard served at http://127.0.0.1:4319. The tool records retrieval chunks, LLM prompts and responses, tool calls, timing and token usage to a local SQLite database at ~/.sledtrace/sledtrace.db, then runs seven deterministic local heuristic warnings — including no_retrieved_chunks, conflicting_chunks, numeric_mismatch and answer_not_grounded — to show why an answer diverged from retrieved context. SledTrace is Python-only with explicit trace, retrieval and llm calls, has no automatic LangChain or LlamaIndex integrations yet, and ships prebuilt binaries for Windows, macOS and Linux on x86-64 and ARM64.", "body_md": "**A local debugger for RAG pipelines.** When your app gives a wrong answer,\nSledTrace shows you why: what the retriever returned, what the model was given,\nwhat it said, and where those disagree.\n\nEverything runs on your machine. No account, no API key, nothing uploaded.\n\n```\npip install sledtrace\nsledtrace serve\n```\n\nYour browser opens the dashboard at `http://127.0.0.1:4319`. Now send it a\ntrace. Save this as `first_trace.py` and run it in another terminal:\n\n``` python\nfrom sledtrace import trace\n\nquestion = \"How many days do customers have to return items after delivery?\"\n\nwith trace(name=\"refund-question\", query=question) as t:\n    t.retrieval(\n        query=question,\n        chunks=[\n            {\"id\": \"policy-2024\", \"text\": \"Customers can return items within 30 days of delivery.\",\n             \"score\": 0.82, \"metadata\": {\"source\": \"refund_policy.md\"}},\n            {\"id\": \"policy-2021\", \"text\": \"Customers can return items within 14 days of delivery.\",\n             \"score\": 0.79, \"metadata\": {\"source\": \"legacy_refund_policy.md\"}},\n        ],\n    )\n    t.llm(model=\"demo-model\", prompt=question,\n          response=\"Customers have 45 days to return items after delivery.\")\n\nprint(t.flush())\n```\n\nRefresh the dashboard and open **refund-question**. SledTrace points out that\nthe two retrieved policies contradict each other, and that the answer's\n\"45 days\" isn't supported by either of them:\n\nReady to trace your own app? Follow the\n**[5-minute quickstart](https://github.com/Schromeo/SledTrace/blob/main/docs/QUICKSTART.md)**.\n\n| Warning | Meaning | \n|---|---|\n| `no_retrieved_chunks` | The retriever returned nothing usable | \n| `low_retrieval_score` | Even the best chunk scored low | \n| `duplicate_chunks` | The same text was retrieved more than once | \n| `weak_query_chunk_overlap` | Top chunks barely mention the question's key terms | \n| `conflicting_chunks` | Retrieved chunks disagree with each other | \n| `numeric_mismatch` | A number in the answer contradicts the retrieved context | \n| `answer_not_grounded` | A claim in the answer is weakly supported by the context | \n\nEvery warning shows the evidence behind it and what to check next. The rules\nare deterministic heuristics that run locally; no LLM judges your data. See\n[warning rules](https://github.com/Schromeo/SledTrace/blob/main/docs/demo/WARNING_RULES.md) for how each one works and where it\nfalls short.\n\nSledTrace also records tool calls, the final task result, timing, and LLM token usage. Values it doesn't know are shown as unknown, never as zero.\n\n```\nyour Python app ──(sledtrace SDK)──▶ local collector ──▶ SQLite\n                                        │\n                                        └──▶ dashboard in your browser\n```\n\nYou add a few calls to your request path (`trace`, `retrieval`, `llm`). The SDK\nsends each finished trace to the collector that `sledtrace serve` starts; the\ncollector runs the warning rules and stores everything in\n`~/.sledtrace/sledtrace.db`.\n\n- Python only, with explicit calls: there are no automatic LangChain or LlamaIndex integrations yet.\n- Warnings are heuristics built on English text patterns, not a correctness verdict.\n- Token usage is recorded only when you pass it (for example with\n`sledtrace.openai.record_response` ); cost estimates are indicative.\n- Local, single-user tool: no hosting, authentication or team features.\n- Prebuilt `sledtrace serve` for Windows, macOS and Linux (x86-64 and ARM64).\nOn other platforms,[run from source](https://github.com/Schromeo/SledTrace/blob/main/docs/DEVELOPMENT.md) .\n\n- [Quickstart](https://github.com/Schromeo/SledTrace/blob/main/docs/QUICKSTART.md) : instrument your own RAG app.\n- [Python SDK guide](https://github.com/Schromeo/SledTrace/blob/main/docs/integrations/PYTHON_SDK_GUIDE.md) : full API reference.\n- [Warning rules](https://github.com/Schromeo/SledTrace/blob/main/docs/demo/WARNING_RULES.md) : what each warning checks.\n- [Development setup](https://github.com/Schromeo/SledTrace/blob/main/docs/DEVELOPMENT.md) : run from source, Docker, demos, configuration.\n- [Contributing](https://github.com/Schromeo/SledTrace/blob/main/CONTRIBUTING.md) and[release notes](https://github.com/Schromeo/SledTrace/blob/main/docs/releases/V0_8_1.md) .\n- [Renaming from RAGLens](https://github.com/Schromeo/SledTrace/blob/main/docs/REBRANDING.md) :`raglens` imports still work.\n\nNamed after my husky. A RAG pipeline is like a sled team: retrievers, rerankers and LLMs all pulling together. When the sled goes off course, you read the tracks in the snow to find out which dog stumbled. SledTrace shows you the tracks.", "url": "https://wpnews.pro/news/sledtrace-a-local-debugger-that-shows-why-your-rag-app-answered-wrong", "canonical_source": "https://github.com/Schromeo/SledTrace", "published_at": "2026-10-02 22:31:55+00:00", "updated_at": "2026-10-02 23:06:28.170932+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "large-language-models", "ai-products"], "entities": ["SledTrace", "Python", "SQLite", "LangChain", "LlamaIndex", "sledtrace.openai.record_response"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/sledtrace-a-local-debugger-that-shows-why-your-rag-app-answered-wrong", "markdown": "https://wpnews.pro/news/sledtrace-a-local-debugger-that-shows-why-your-rag-app-answered-wrong.md", "text": "https://wpnews.pro/news/sledtrace-a-local-debugger-that-shows-why-your-rag-app-answered-wrong.txt", "jsonld": "https://wpnews.pro/news/sledtrace-a-local-debugger-that-shows-why-your-rag-app-answered-wrong.jsonld"}}