{"slug": "oncall-memory-building-an-incident-response-agent-that-learns-from-production", "title": "OnCall Memory:Building an Incident Response Agent That Learns From Production", "summary": "A developer built OnCall Memory, an incident-response website that pairs a FastAPI backend with Groq's language models and the Hindsight persistent-memory layer to give AI diagnoses access to an organization's prior production incidents. The project uses a fictional fintech, Northwind Pay, to show side-by-side responses with and without retrieved incident history, and applies Hindsight's retain, recall and reflect operations to surface recurring root causes across incidents.", "body_md": "OnCall Memory: Building an Incident Response Agent That Learns From Production History\n\nProduction incidents are rarely completely new.\n\nA database connection pool can become exhausted again. A deployment configuration can break a service again. A payment provider can become rate-limited again. The symptoms may change, but the underlying patterns often repeat.\n\nThe problem is that a typical AI assistant starts each incident with very little knowledge of an organization's previous experiences.\n\nIt can analyze the logs in front of it and provide technically reasonable suggestions, but it does not automatically know what happened during the last incident, which fix worked, or which attempted solution failed.\n\nThat is the problem I wanted to address with OnCall Memory, an incident-response website designed around persistent AI memory.\n\nThe Idea:\n\nOnCall Memory is built for a fictional fintech company called Northwind Pay.\n\nWhen an engineer receives a production alert, the website provides two perspectives:\n\n->Without Memory a response generated from the current incident.\n\n->With Hindsight Memory a response generated after retrieving relevant historical incidents from the organization's memory.\n\nThis side-by-side comparison is the central idea of the website.\n\nThe goal isn't simply to make an AI response longer. It is to give the model access to something it normally doesn't have: the team's previous incident experience.\n\nBuilding the incident history\n\nThe project contains a collection of realistic Northwind Pay production incidents.\n\nThe incident dataset includes situations such as PostgreSQL connection pool exhaustion, Redis eviction, expired TLS certificates, bad deployment configuration, Kafka consumer lag, OOMKilled pods, and third-party payment API rate limiting.\n\nEach incident contains information such as the symptoms, logs, root cause, resolution steps, outcome, and whether attempted fixes worked.\n\nThis is important because a useful memory isn't just:\n\n\"There was a database incident.\"\n\nIt should capture the experience surrounding that incident.\n\nWhat happened?\n\nWhy did it happen?\n\nWhat was tried?\n\nWhat actually fixed it?\n\nDid the first solution fail?\n\nThat information becomes useful context for future incidents.\n\nHindsight as the memory layer\n\nHindsight is at the center of the architecture.\n\nThe website uses Hindsight to perform three important operations: retain, recall, and reflect.\n\nWhen useful incident information is available, it can be retained as memory.\n\nWhen a new incident arrives, the system recalls relevant historical incidents before generating the memory-enhanced diagnosis.\n\nFinally, the website can use reflection to look across the accumulated incident history and ask a broader question:\n\n**What patterns keep causing our incidents and what should we fix permanently?**\n\nThis moves the system beyond responding to individual alerts and toward learning from patterns across incidents.\n\nThe architecture\n\nThe website uses a lightweight architecture.\n\nThe frontend is a single-page HTML, CSS, and JavaScript dashboard.\n\nThe backend is implemented with FastAPI and handles the incident analysis workflow, memory operations, and language-model requests.\n\nGroq provides the language-model layer, while Hindsight provides persistent memory.\n\nThe project also separates configuration from application code. API credentials are kept locally in environment variables rather than being hardcoded into the source code, while `.env.example` provides the required configuration structure.\n\nMemory changes the workflow\n\nSuppose a new payment failure arrives.\n\nWithout historical context, an AI assistant might suggest checking the payment provider, network connectivity, credentials, rate limits, retries, and application logs.\n\nThose are reasonable suggestions, but the engineer still has to determine which one matches the organization's previous experience.\n\nWith Hindsight memory, the workflow becomes different.\n\nThe system first looks for similar historical incidents.\n\nIf a previous payment incident involved third-party API rate limiting, for example, that historical information can become part of the diagnosis.\n\nThe engineer can also see the recalled incidents rather than receiving a completely opaque recommendation.\n\nThat visibility was an important part of the design.\n\nLearning doesn't stop after the first answer\n\nOnCall Memory also includes a feedback loop.\n\nAfter applying a suggested fix, the engineer can indicate whether the fix worked or failed and provide the actual solution.\n\nThat information can then become another memory.\n\nThe intended loop is:\n\n**Incident → Recall → Diagnosis → Fix → Feedback → Retain → Better future diagnosis**\n\nThis creates a system where future incidents can benefit from experiences that were not available when the original system was built.\n\nThe current Hindsight memory bank contains **139 memories**.\n\nOne of the biggest lessons from building the website was that memory quality matters as much as memory quantity.\n\nAdding more memories does not automatically create a better assistant.\n\nA useful incident memory needs meaningful context: symptoms, root cause, attempted fixes, final resolution, and outcome.\n\nI also learned that memory should be visible to the user.\n\nIf an AI gives a very specific operational recommendation, an engineer should have some way to understand why that information was relevant.\n\nDisplaying recalled incidents makes the memory layer much easier to inspect.\n\nLimitations\n\nOnCall Memory currently uses a fictional Northwind Pay incident history rather than real production data.\n\nThat means the system demonstrates the workflow but should not be interpreted as a replacement for an organization's actual incident-management practices.\n\nThe quality of the recommendations also depends on the quality of the stored memories. Poor or incomplete incident records can lead to less useful historical context.\n\nFinal thoughts\n\nThe central idea behind OnCall Memory is simple:\n\n**An incident-response assistant should not forget what happened yesterday.**\n\nLanguage models are good at reasoning about the information they receive. Persistent memory adds another dimension: the ability to use accumulated organizational experience.\n\nBy combining Groq for reasoning with Hindsight for persistent memory, OnCall Memory turns incident history into a resource that can be recalled during future incidents and analyzed for recurring patterns.\n\nThe result is not an AI replacing an on-call engineer.\n\nIt is an AI assistant that can increasingly say:\n\n\"We've seen something like this before.\"", "url": "https://wpnews.pro/news/oncall-memory-building-an-incident-response-agent-that-learns-from-production", "canonical_source": "https://dev.to/saanvii248/oncall-memorybuilding-an-incident-response-agent-that-learns-from-production-401m", "published_at": "2026-09-29 18:39:28+00:00", "updated_at": "2026-09-29 18:46:36.445480+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "ai-infrastructure", "mlops"], "entities": ["OnCall Memory", "Hindsight", "Northwind Pay", "FastAPI", "Groq"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/oncall-memory-building-an-incident-response-agent-that-learns-from-production", "markdown": "https://wpnews.pro/news/oncall-memory-building-an-incident-response-agent-that-learns-from-production.md", "text": "https://wpnews.pro/news/oncall-memory-building-an-incident-response-agent-that-learns-from-production.txt", "jsonld": "https://wpnews.pro/news/oncall-memory-building-an-incident-response-agent-that-learns-from-production.jsonld"}}