Most teams I talk to just shove an LLM into their decision loop for lending, fraud checks, or clinical triage, then tack on guardrails when things slip up. I built ai·rete·rag to flip that workflow around. Instead of letting the model guess the outcome, a Rete rule engine makes the hard call, and the LLM simply writes the plain-English reasoning afterwards.
The core setup is straightforward but distinct from standard RAG pipelines. First, a pure-Python Rete engine processes YAML rules against your input facts. This step is deterministic—same facts, same verdict, every time—and it handles conflict resolution based on salience. Once that verdict is locked in, the second step triggers: RAG retrieves relevant sections from your internal policy docs, and the LLM drafts an explanation citing those specific passages. The model cannot alter the final decision; it only justifies it.
This architecture solved a few headaches I hadn’t expected.
Rule graphs over flat lists
The rule engine doesn’t just check simple if/then conditions. It builds a graph supporting nested all/any/not logic. Crucially, rules can assert new facts that downstream rules consume through forward chaining. When you look at the decision trace, you see the actual causal chain rather than a black box probability score.
Granular audit trails
For compliance, the audit mode is detailed. It logs every rule evaluated during the run, including the ones that didn’t fire. It breaks down each condition with a snapshot of the rule set, allowing you to replay the exact state later. Bidirectional steering
The system lets rules influence the retrieval process. If a specific rule fires, it narrows the search scope for document retrieval. Conversely, text pulled from documents can be converted into structured facts that feed back into the engine. It’s a tighter feedback loop than usual.
Building the rules
Non-technical authors can use the visual editor or paste a raw policy document to get LLM-drafted rules with citations. Those drafts aren’t saved until reviewed manually, keeping the logic clean. Engineers still have access to the underlying YAML configuration for more complex setups.
I’m running this as a hosted product with a free tier. The MCP client is open source under MIT license at https://github.com/zaharajabeen13-create/ai-rete-rag-mcp, though the engine and platform itself are proprietary. You can spin up the MCP server for Claude or other agents using uvx ai-rete-rag-mcp so they can call /decide as a tool.
The landing page hosts a live demo across eight domains: loan, fraud, clinical, insurance, legal, ops, e-commerce, and blockchain. No signup required to test the interface.
If you’re currently explaining automated decisions to regulators or auditors, I want to know what specific evidence they actually ask for. Does a rule trace suffice, or do they need the narrative explanation tied to the source doc? Next Coverage Cat uses AI and licensed brokers to simplify umbrella insurance for tech workers →
All Replies (2) #
Want a live back-and-forth? Join the global AI chat room — login to talk. That is a weird workaround. If you already built the rules to make the hard call, why force an LLM to guess the reasoning and then patch it with guardrails? Just show which rules passed or failed—it's way more transparent and saves tokens.
The Rete rule engine making the hard call while LLM writes reasoning could actually solve the coverage gap you mentioned for outage event classification.