Gargi Reflex: Autonomous Model Caching for System One Decisions Gargi launched Reflex, a Python tool that caches repeated LLM calls by training a small CPU model on logged inputs and answers, serving confident calls locally in about 5 ms at $0 instead of 3.6 seconds. Reflex only swaps a call after agreement, calibration and coverage pass on unseen data, and routes everything else plus a permanent holdout slice back to the original LLM, which takes over if agreement drops. If Reflex fails for any reason, the call falls through to the user's function exactly once with exceptions propagating unchanged. Make every LLM call swappable. You write the prompt. Reflex learns from your LLM's answers and serves the repeat calls locally, in milliseconds. You do the prompt design. Reflex does the data science. Watch the film ↓ film One minute, from 3.6 s to 5 ms. What Reflex does 1. 01 It watches.Your typed LLM calls routing, intent, moderation, extraction go to your model exactly as today. Reflex logs each input and answer. 2. 02 It proves.It trains a small model on CPU in minutes, then tests it on data it never saw. Unless agreement, calibration and coverage all pass, nothing is swapped. 3. 03 It swaps, and keeps checking.Confident calls are served locally in about 5 ms at $0. Everything else, plus a permanent holdout slice, still goes to your LLM, which takes back over if agreement drops. Built for Python teams whose product asks an LLM the same kind of question thousands of times a day. The evidence Read the research → https://www.gargi.io/reflex/research If anything in Gargi fails no model, low confidence, a model that won’t load, a corrupt database, a bug , the call goes to your function, exactly once. Your exceptions propagate unchanged.