OnCall Memory: Building an Incident Response Agent That Learns From Production History
Production incidents are rarely completely new.
A database connection pool can become exhausted again. A deployment configuration can break a service again. A payment provider can become rate-limited again. The symptoms may change, but the underlying patterns often repeat.
The problem is that a typical AI assistant starts each incident with very little knowledge of an organization's previous experiences.
It can analyze the logs in front of it and provide technically reasonable suggestions, but it does not automatically know what happened during the last incident, which fix worked, or which attempted solution failed.
That is the problem I wanted to address with OnCall Memory, an incident-response website designed around persistent AI memory.
The Idea:
OnCall Memory is built for a fictional fintech company called Northwind Pay.
When an engineer receives a production alert, the website provides two perspectives:
->Without Memory a response generated from the current incident.
->With Hindsight Memory a response generated after retrieving relevant historical incidents from the organization's memory.
This side-by-side comparison is the central idea of the website.
The goal isn't simply to make an AI response longer. It is to give the model access to something it normally doesn't have: the team's previous incident experience.
Building the incident history
The project contains a collection of realistic Northwind Pay production incidents.
The incident dataset includes situations such as PostgreSQL connection pool exhaustion, Redis eviction, expired TLS certificates, bad deployment configuration, Kafka consumer lag, OOMKilled pods, and third-party payment API rate limiting.
Each incident contains information such as the symptoms, logs, root cause, resolution steps, outcome, and whether attempted fixes worked.
This is important because a useful memory isn't just:
"There was a database incident."
It should capture the experience surrounding that incident.
What happened?
Why did it happen?
What was tried?
What actually fixed it?
Did the first solution fail?
That information becomes useful context for future incidents.
Hindsight as the memory layer
Hindsight is at the center of the architecture.
The website uses Hindsight to perform three important operations: retain, recall, and reflect.
When useful incident information is available, it can be retained as memory.
When a new incident arrives, the system recalls relevant historical incidents before generating the memory-enhanced diagnosis.
Finally, the website can use reflection to look across the accumulated incident history and ask a broader question:
What patterns keep causing our incidents and what should we fix permanently?
This moves the system beyond responding to individual alerts and toward learning from patterns across incidents.
The architecture
The website uses a lightweight architecture.
The frontend is a single-page HTML, CSS, and JavaScript dashboard.
The backend is implemented with FastAPI and handles the incident analysis workflow, memory operations, and language-model requests.
Groq provides the language-model layer, while Hindsight provides persistent memory.
The project also separates configuration from application code. API credentials are kept locally in environment variables rather than being hardcoded into the source code, while .env.example provides the required configuration structure.
Memory changes the workflow
Suppose a new payment failure arrives.
Without historical context, an AI assistant might suggest checking the payment provider, network connectivity, credentials, rate limits, retries, and application logs.
Those are reasonable suggestions, but the engineer still has to determine which one matches the organization's previous experience.
With Hindsight memory, the workflow becomes different.
The system first looks for similar historical incidents.
If a previous payment incident involved third-party API rate limiting, for example, that historical information can become part of the diagnosis. The engineer can also see the recalled incidents rather than receiving a completely opaque recommendation.
That visibility was an important part of the design.
Learning doesn't stop after the first answer
OnCall Memory also includes a feedback loop.
After applying a suggested fix, the engineer can indicate whether the fix worked or failed and provide the actual solution.
That information can then become another memory.
The intended loop is:
Incident → Recall → Diagnosis → Fix → Feedback → Retain → Better future diagnosis
This creates a system where future incidents can benefit from experiences that were not available when the original system was built.
The current Hindsight memory bank contains 139 memories.
One of the biggest lessons from building the website was that memory quality matters as much as memory quantity.
Adding more memories does not automatically create a better assistant.
A useful incident memory needs meaningful context: symptoms, root cause, attempted fixes, final resolution, and outcome.
I also learned that memory should be visible to the user.
If an AI gives a very specific operational recommendation, an engineer should have some way to understand why that information was relevant. Displaying recalled incidents makes the memory layer much easier to inspect.
Limitations
OnCall Memory currently uses a fictional Northwind Pay incident history rather than real production data.
That means the system demonstrates the workflow but should not be interpreted as a replacement for an organization's actual incident-management practices.
The quality of the recommendations also depends on the quality of the stored memories. Poor or incomplete incident records can lead to less useful historical context.
Final thoughts The central idea behind OnCall Memory is simple:
An incident-response assistant should not forget what happened yesterday.
Language models are good at reasoning about the information they receive. Persistent memory adds another dimension: the ability to use accumulated organizational experience.
By combining Groq for reasoning with Hindsight for persistent memory, OnCall Memory turns incident history into a resource that can be recalled during future incidents and analyzed for recurring patterns.
The result is not an AI replacing an on-call engineer.
It is an AI assistant that can increasingly say:
"We've seen something like this before."