How I Built an Incident Agent That Remembers Fixes With Hindsight A developer built On-Call Copilot, an incident-response agent that uses Hindsight as a persistent memory layer to recall project-specific fixes and anonymized shared lessons from past incidents. The system runs a tool-calling loop with four tools — inspect_alert, recall_project_memory, recall_shared_lessons, and ask_engineer_question — and writes each resolved incident back into a private per-project memory bank, optionally contributing anonymized lessons to a shared bank. How I Built an Incident Agent That Remembers Fixes With Hindsight The hardest part of incident response is rarely finding an answer once. It is making sure the next engineer does not have to rediscover the same answer six months later. I built On-Call Copilot around that problem: an incident-response agent that can inspect an alert, decide which information it needs, search project-specific incident memory, use anonymized lessons from elsewhere, ask an engineer for missing context, and retain the eventual resolution for the next incident. The problem was not diagnosis. It was forgetting. A conventional incident assistant can take an alert and produce a plausible explanation. That is useful, but it leaves out an important part of operational knowledge: what actually happened in this environment before. Consider an alert like this: ALERT prod service=search-api severity=P1 502 Bad Gateway rate: 210/min baseline: 1/min Logs: Timeout expired. The timeout period elapsed prior to obtaining a connection from the pool. A language model can recognize the connection-pool symptom. But recognition is not the same as operational knowledge. I wanted the agent to answer a more useful question: have we seen this in this project before, and what actually fixed it? That led me to use Hindsight GitHub as the memory layer rather than treating memory as a prompt-management problem. The important distinction is that the model does not own the memory. Hindsight does. The agent retrieves relevant facts when it needs them, and it writes the outcome back after an engineer resolves the incident. How the system hangs together The application is a Streamlit interface around an OnCallCopilot class. The main pieces are deliberately small: An autonomous AI incident-response copilot that investigates production alerts, uses persistent memory to learn from previous incidents, and recalls proven resolutions for future incidents. On-Call Copilot is an AI-powered incident-response system designed to assist engineers during production incidents. Traditional incident-response systems can detect alerts and provide static information, but they do not continuously learn from how engineers actually resolve incidents. On-Call Copilot addresses this by combining: The system can investigate an alert, inspect relevant information, retrieve previous lessons, ask an engineer for missing context, generate a diagnosis, and retain the resolved incident so that the knowledge can be reused later. During production incidents, engineers often need to: The dashboard handles the incident workflow. The agent owns the reasoning, Hindsight integration, tool execution, and memory lifecycle. Project configuration determines which private memory bank belongs to the active project. At runtime, the flow looks like this: Production alert │ ▼ Autonomous incident agent │ ├── inspect alert ├── recall project memory ├── recall shared lessons └── ask engineer question │ ▼ Diagnosis + immediate actions │ ▼ Engineer resolves incident │ ▼ Retain resolution in private Hindsight memory │ └── optionally create anonymized shared lesson The important design choice is that memory is scoped. Each project gets a private Hindsight bank containing detailed operational history. There is also a separate shared-lessons bank containing anonymized, generic lessons that a project has explicitly chosen to contribute. That means a shared lesson can inform an investigation without pretending that the incident happened in the current system. I stopped making the agent call every tool One of the more important changes in the project was moving from a fixed investigation sequence to an actual tool-calling loop. The agent exposes four tools: AGENT TOOLS = { "type": "function", "function": { "name": "inspect alert", "description": "Parse the incoming incident alert into structured facts before diagnosing it.", ... }, }, { "type": "function", "function": { "name": "recall project memory", "description": "Search THIS project's private Hindsight memory for similar incidents, fixes, owners and feedback.", ... }, }, { "type": "function", "function": { "name": "recall shared lessons", "description": "Search the anonymized shared Hindsight lessons.", ... }, }, { "type": "function", "function": { "name": "ask engineer question", "description": "Use when the alert is too vague to diagnose safely.", ... }, }, The LLM receives these tools and decides what to call next. for step in range max steps : response = self.groq.chat.completions.create model=GROQ MODEL, messages=messages, tools=self.AGENT TOOLS, tool choice="auto", temperature=0, msg = response.choices 0 .message calls = getattr msg, "tool calls", None or if not calls: The model has enough evidence to return the diagnosis. ... break I deliberately put a maximum step count around the loop. Autonomous tool calling without a boundary is a good way to turn a simple incident into an unpredictable chain of calls. The system prompt also tells the agent to use the minimum useful tools rather than blindly calling everything. That matters because an alert does not always need the same investigation. A detailed stack trace may only require alert inspection and project-memory recall. A vague report such as "customers say the site feels slow" needs a different response: ask the engineer for useful information instead of inventing a root cause. The tool layer keeps memory boundaries explicit The tool implementation makes the privacy boundary visible in code. if name == "recall project memory": query = arguments.get "query", "" :2000 return { "scope": "this project only", "facts": self. recall bank, query , } if name == "recall shared lessons": if not use shared: return { "scope": "shared lessons disabled", "facts": , } query = arguments.get "query", "" :2000 return { "scope": "anonymized shared lessons only", "facts": self. recall self.global bank, query , } I like this pattern because the distinction is not merely UI text. The agent gets different tools for different memory scopes, and the tool result explicitly tells it what kind of evidence it received. The diagnosis schema reinforces the same rule. Dates, owners, runbooks, and resolution times can only come from project memory. Shared lessons cannot suddenly become historical facts about the current project. That is an important guardrail for an incident system. A useful lesson from another service is not evidence that the same thing happened here. Hindsight is the part that makes the loop cumulative The Hindsight documentation describes memory as something an agent can retain and recall rather than simply stuffing more text into a context window. That maps well to how I structured this system. After an engineer resolves an incident, the resolution is written back into the current project's private memory: def resolve and retain self, alert text, resolution summary, diagnosis=None, bank id=None, : bank = bank id or self.bank id when = datetime.now timezone.utc .strftime "%Y-%m-%d %H:%M UTC" record = f"INCIDENT RESOLVED {when} \n" f"ALERT: {alert text}\n" f"RESOLUTION: {resolution summary}" if diagnosis and diagnosis.get "summary" : record += f"\nAGENT'S EARLIER ASSESSMENT: " f"{diagnosis 'summary' }" self. retain bank, record This is intentionally simple. I do not store only the model's diagnosis. I store the alert and the actual resolution supplied by the engineer. The difference matters because the model's initial explanation can be wrong while the engineer's resolution is the operational fact I want to remember. The system also records engineer feedback: verdict = "HELPFUL and worked" if helpful else "NOT helpful" record = f"ENGINEER FEEDBACK: For this alert: {alert text :300 }\n" f"The recommendation {steps} was rated {verdict}. {note}" .strip self. retain bank, record This gives future investigations another useful signal: not just what the agent suggested, but whether someone said that suggestion worked. That is where agent memory becomes more interesting than simply keeping a transcript. A connection-pool incident becomes reusable knowledge The recurring SQL connection-pool scenario is a good example. The alert reports 502 responses and a timeout while obtaining a database connection. In a cold project, the agent has to reason from the alert itself. After the incident has been resolved and retained, a later occurrence can retrieve the previous incident from project memory. The remembered resolution in the system is: The SQL connection pool was exhausted during the scheduled reindex job. The incident was resolved by staggering the reindex operation and increasing the connection pool capacity. After the change, the 502 error rate returned to normal. The important behavior is not that the model suddenly "knows" SQL connection pools. It already knew that. The useful change is that it now knows what this project did the last time the same failure appeared. That distinction also appears in the UI. The incident console separates: a memory hit in the current project, a lesson seen elsewhere, a genuinely new pattern, and an alert that needs more information. Those states are operationally different, so I did not want to collapse them into one generic confidence score. I also wanted the agent to admit when it does not know One of the easiest ways to make an incident assistant dangerous is to reward it for producing an answer every time. The diagnosis contract explicitly allows needs more info=true and requires clarifying questions when an alert is too vague. The system prompt says: If the alert is too vague to act on no service, symptom, error or metric : needs more info=true, give 3 specific clarifying questions, and still give 2-3 safe first checks in immediate actions. That gives the agent a useful middle state. It does not have to choose between "diagnose everything" and "do nothing." It can perform safe initial checks, explain what is missing, and ask the engineer for the information needed to continue. The same principle applies to memory. If nothing relevant is found, the agent is instructed to mark the pattern as new rather than inventing a historical match. Shared lessons are deliberately harder to publish The second Hindsight path is for knowledge that can be useful beyond one project. After a fix is retained privately, the system can generate an anonymized lesson: def sanitize lesson self, alert text, resolution, project name="" : raw = self. chat SANITIZE SYSTEM, f"ALERT:\n{alert text}\n\nRESOLUTION:\n{resolution}" terms = project name + service names alert text clean, n = redact raw or "" .strip , terms return {"lesson": clean, "redactions": n} But the generated lesson is not automatically published to the shared memory. The engineer reviews it first. The code also redacts project and service identifiers again after the human edit before retaining the lesson in the shared bank. That gives me a two-level memory model: Private project memory ├── alerts ├── resolutions ├── owners ├── runbooks └── engineer feedback Shared lessons └── anonymized, reviewed, generic knowledge I found this separation more useful than trying to make every memory globally searchable. Operational details are valuable inside the system that generated them. They are often inappropriate outside it. Asking the memory is a separate workflow Incident diagnosis is not the only reason to query operational history. The application also exposes a direct "Ask the memory" workflow. Questions such as: Which services fail most often and why? What recurring root causes should we fix permanently? Which past fixes were fastest and what made them fast? are sent to Hindsight's reflect capability for the current project's private bank. res = self.hindsight.reflect bank id=bank, query=question, return { "answer": getattr res, "text", None or str res , "method": "Hindsight reflect", } If reflection fails, the implementation falls back to recalling relevant memories and asking the language model to answer from that retrieved context. This makes the memory layer useful outside the immediate alert path. Incident history becomes something engineers can interrogate rather than a pile of old tickets nobody wants to read. What the system looks like during an incident The most useful interaction is a simple one. An engineer submits an alert. The agent starts by inspecting it. It may then recall private project history. If shared lessons are enabled, it can search those separately. If the alert is vague, it can ask a question instead. The UI exposes the resulting tool trace: Agent activity · autonomous tool calling Agent mode: autonomous tool calling The final diagnosis includes the severity, category, confidence, root-cause explanation, matched incidents, immediate actions, and longer-term fixes. After the engineer resolves the incident, they enter the actual resolution and retain it. From that point forward, the incident is no longer just something that happened once. It becomes evidence available to the next investigation. What I learned building it A transcript contains everything, including speculation. The resolution is much more valuable than another copy of the model's conversation. I therefore make the engineer-confirmed fix a first-class memory record. "Search memory" is too vague for an operational system. I needed to know whether a fact came from this project or from a shared lesson. Separate banks and separate tools made that boundary explicit. Letting the model choose tools is useful, but the model still needs a bounded loop, strict tool descriptions, structured output, and rules about what it is allowed to infer. Autonomy without constraints is just another failure mode. Past incidents are evidence, not certainty. The diagnosis still contains confidence and a root-cause explanation, and shared lessons are deliberately prevented from being represented as incidents in the current project. The most useful feedback is often binary and operational: "that worked" or "that did not work." Recording that feedback means the system can avoid repeating recommendations that engineers already rejected. Where I would take it next The core loop is now clear: Alert → investigate → recall → diagnose → engineer resolves → retain → recall next time The natural production integrations are around that loop: receiving alerts from monitoring systems, connecting incident records and ticketing systems, posting status updates to engineering communication channels, and automatically retaining confirmed resolutions after incidents close. The memory architecture does not need to change for those integrations. That is the part I like most about the design. The language model is replaceable. The dashboard is replaceable. The alert source is replaceable. The durable piece is the operational memory: what happened, what we tried, what actually fixed it, and what engineers learned afterward. That is a much more useful thing for an incident agent to remember than a conversation.