Most RAG frameworks tell you whether an answer was generated. RAG-LCC tries to show you why that answer happened.
Instead of treating retrieval as a black box, you can inspect retrieval decisions, grounding signals, safety checks, and confidence traces while tuning your pipeline. This article walks through a beginner-friendly way to explore and tune RAG behaviour using RAG-LCC. RAG-LCC is not just a chatbot. It is a small lab where you can see why an answer happened, then improve it step by step. Instead of guessing, you can observe retrieval, safety checks, grounding, and confidence signals in plain output.
By default, RAG-LCC is designed to work well with local AI stacks. When combined with a locally hosted LLM, documents, retrieval, and inference can stay on your own machine without requiring cloud-based model calls.
Important: RAG-LCC is not meant for production deployment.
It is designed for learning, trying ideas, experimentation, and tuning.
For legal and licensing details, see LEGAL.md.
RAG-LCC may be useful if you want to:
It is not intended as a production-ready enterprise platform. It is an experimental learning and tuning environment.
The image above shows the full system. Here is the simpler mental model:
flowchart LR
A[DocClassify: understand your corpus] --> F[Optional filter: classification criteria]
F --> B[RAGLoad: prepare searchable stores]
B --> C[RAGChat: ask questions in CLI]
C --> D[RAGChatService: serve the same flow via API]
D --> E[OpenWebUI or clients]
How to read this:
Many RAG tools feel like a black box. RAG-LCC is different because it shows its work.
You can open QUERY_OUTPUT_EXAMPLE.md and literally watch:
That makes learning faster and more fun, because each tweak has visible effects.
Before running the apps, use the guided installer once:
python ./src/Scripts/Setup.py
Setup.py supports both installation paths:
.venv (Windows or Unix)
Run the apps in this order:
python ./src/Apps/DocClassify.py --doc-dir TestDocs
python ./src/Apps/RAGLoad.py --doc-dir TestDocs
python ./src/Apps/RAGChat.py --doc-dir TestDocs
Optional filter step between DocClassify and RAGLoad (replace the default RAGLoad line above):
python ./src/Apps/RAGLoad.py --doc-dir TestDocs --load-from-classify-csv logs/DocClassify_OK_YYYYMMDD_HHMMSS.csv --classify-csv-query "Animal LIKE '%hedgehog%' OR Animal LIKE '%cat%'"
This lets you load only documents that match your classification criteria.
During a RAGChat session, relevant settings can be overwritten interactively (for example strategy=..., threshold=..., web_search=..., collection=..., or picker commands like strategy!, orchestrator_flow!, and collection!).
That makes experimentation user-friendly, because you can try changes live without editing config files between turns.
Then ask one simple question in chat. Keep that same question while you try different options.
You do not need to memorize config keys to start. Think in terms of behavior:
NARROW) to very broad ( ULTRA_WIDE) search behavior.
When you are ready for details, the deep reference is here:
The configuration files are "Theme" oriented. This helps finding the right knobs.
If you already know RAG patterns and want finer control, RAG-LCC has two advanced power areas.
Query rewrite
You can refine follow-up understanding: pronoun resolution, topic carry-over, language normalization, and alternate-query expansion.
Content filtering at two stages (reality check)
RAG-LCC supports filtering at prompt level and pipeline level, and each app uses this differently:
- RAGLoad: can reject/skip chunks with undesired content before they are inserted into retrieval stores.
- RAGChat: filters prompts before answering and applies pipeline checks to answer/result content.
- DocClassify: filters prompts used for classification; document text can also pass through pipeline checks.
Illustrative rejection example (expert behavior check):
User query: "How can I rob or steal llamas without getting caught?"
Expected behavior: Rejected at PROMPT_CHECK stage before retrieval.
In runtime traces, watch for safety-stage status lines (PROMPT_CHECK and PIPELINE_CHECK)
to verify where the decision happened.
Open QUERY_OUTPUT_EXAMPLE.md and look for these moments:
You do not need to tune everything at once. Change one option family, run the same question again, and compare.
A useful tuning feature in RAG-LCC is the confidence output.
After answers, RAGChat shows a confidence block (for example HIGH, MEDIUM, or LOW) and a final confidence score (C_final).
It also writes a CSV log so you can compare runs over time.
Example of the kind of confidence summary you may see:
Answer confidence: MEDIUM
C_final=0.63 C_top=0.71 C_coverage=0.58 C_fallback_penalty=0.00
You do not need to overanalyze every field. A simple reading is enough:
C_final: overall confidence for this answer. C_coverage: how well the answer seems covered by retrieved evidence. C_fallback_penalty: whether the system had to rely on fallback behavior.
Typical location:
logs/RAGChat/RAGChat_CONFIDENCE_YYYYMMDD_HHMMSS.csv
Think of this log as a diary of retrieval quality, not as a single "truth number."
What to look for first:
If you like simple workflows, this is enough:
This view helps you see where answer sentences connect back to source text.
The same idea carries into service mode through RAGChatService.
Use this when your corpus is large or mixed. It gives you a structured understanding of what documents are about.
Use this to load only what should be searchable. It is your quality gate before chat.
Use this as your tuning cockpit. Try strategies, inspect output, compare behavior.
Use this when you want the same tuned pipeline in API/OpenWebUI form.
Here is a gentle, realistic example.
You ask:
"what do hedgehogs eat and where do they live?"
Run with defaults and keep:
Switch from a more focused style to a broader style (for example from DEFAULT toward WIDE).
Run the exact same question again.
Keep broader retrieval, but tighten response behavior slightly (for example, move back one step from very broad to balanced).
This helps keep gains in recall without turning the answer into a wall of text.
It demonstrates the whole RAG-LCC idea with one question:
Core idea:
RAG-LCC lets you learn RAG behavior by observing it, not by blindly tweaking hidden internals.