{"slug": "how-i-built-an-incident-agent-that-remembers-fixes-with-hindsight", "title": "How I Built an Incident Agent That Remembers Fixes With Hindsight", "summary": "A developer built On-Call Copilot, an incident-response agent that uses Hindsight as a persistent memory layer to recall project-specific fixes and anonymized shared lessons from past incidents. The system runs a tool-calling loop with four tools — inspect_alert, recall_project_memory, recall_shared_lessons, and ask_engineer_question — and writes each resolved incident back into a private per-project memory bank, optionally contributing anonymized lessons to a shared bank.", "body_md": "How I Built an Incident Agent That Remembers Fixes With Hindsight\n\nThe hardest part of incident response is rarely finding an answer once. It is making sure the next engineer does not have to rediscover the same answer six months later.\n\nI built On-Call Copilot around that problem: an incident-response agent that can inspect an alert, decide which information it needs, search project-specific incident memory, use anonymized lessons from elsewhere, ask an engineer for missing context, and retain the eventual resolution for the next incident.\n\nThe problem was not diagnosis. It was forgetting.\n\nA conventional incident assistant can take an alert and produce a plausible explanation. That is useful, but it leaves out an important part of operational knowledge: what actually happened in this environment before.\n\nConsider an alert like this:\n\nALERT [prod]\n\nservice=search-api\n\nseverity=P1\n\n502 Bad Gateway\n\nrate: 210/min\n\nbaseline: 1/min\n\nLogs:\n\nTimeout expired.\n\nThe timeout period elapsed prior to obtaining\n\na connection from the pool.\n\nA language model can recognize the connection-pool symptom. But recognition is not the same as operational knowledge.\n\nI wanted the agent to answer a more useful question: have we seen this in this project before, and what actually fixed it?\n\nThat led me to use Hindsight GitHub as the memory layer rather than treating memory as a prompt-management problem.\n\nThe important distinction is that the model does not own the memory. Hindsight does. The agent retrieves relevant facts when it needs them, and it writes the outcome back after an engineer resolves the incident.\n\nHow the system hangs together\n\nThe application is a Streamlit interface around an OnCallCopilot class. The main pieces are deliberately small:\n\nAn autonomous AI incident-response copilot that investigates production alerts, uses persistent memory to learn from previous incidents, and recalls proven resolutions for future incidents.\n\nOn-Call Copilot is an AI-powered incident-response system designed to assist engineers during production incidents.\n\nTraditional incident-response systems can detect alerts and provide static information, but they do not continuously learn from how engineers actually resolve incidents.\n\nOn-Call Copilot addresses this by combining:\n\nThe system can investigate an alert, inspect relevant information, retrieve previous lessons, ask an engineer for missing context, generate a diagnosis, and retain the resolved incident so that the knowledge can be reused later.\n\nDuring production incidents, engineers often need to:\n\nThe dashboard handles the incident workflow. The agent owns the reasoning, Hindsight integration, tool execution, and memory lifecycle. Project configuration determines which private memory bank belongs to the active project.\n\nAt runtime, the flow looks like this:\n\nProduction alert\n\n      │\n\n      ▼\n\nAutonomous incident agent\n\n      │\n\n      ├── inspect_alert\n\n      ├── recall_project_memory\n\n      ├── recall_shared_lessons\n\n      └── ask_engineer_question\n\n      │\n\n      ▼\n\nDiagnosis + immediate actions\n\n      │\n\n      ▼\n\nEngineer resolves incident\n\n      │\n\n      ▼\n\nRetain resolution in private Hindsight memory\n\n      │\n\n      └── optionally create anonymized shared lesson\n\nThe important design choice is that memory is scoped.\n\nEach project gets a private Hindsight bank containing detailed operational history. There is also a separate shared-lessons bank containing anonymized, generic lessons that a project has explicitly chosen to contribute.\n\nThat means a shared lesson can inform an investigation without pretending that the incident happened in the current system.\n\nI stopped making the agent call every tool\n\nOne of the more important changes in the project was moving from a fixed investigation sequence to an actual tool-calling loop.\n\nThe agent exposes four tools:\n\nAGENT_TOOLS = [\n\n    {\n\n        \"type\": \"function\",\n\n        \"function\": {\n\n            \"name\": \"inspect_alert\",\n\n            \"description\": \"Parse the incoming incident alert into structured facts before diagnosing it.\",\n\n            ...\n\n        },\n\n    },\n\n    {\n\n        \"type\": \"function\",\n\n        \"function\": {\n\n            \"name\": \"recall_project_memory\",\n\n            \"description\": \"Search THIS project's private Hindsight memory for similar incidents, fixes, owners and feedback.\",\n\n            ...\n\n        },\n\n    },\n\n    {\n\n        \"type\": \"function\",\n\n        \"function\": {\n\n            \"name\": \"recall_shared_lessons\",\n\n            \"description\": \"Search the anonymized shared Hindsight lessons.\",\n\n            ...\n\n        },\n\n    },\n\n    {\n\n        \"type\": \"function\",\n\n        \"function\": {\n\n            \"name\": \"ask_engineer_question\",\n\n            \"description\": \"Use when the alert is too vague to diagnose safely.\",\n\n            ...\n\n        },\n\n    },\n\n]\n\nThe LLM receives these tools and decides what to call next.\n\nfor step in range(max_steps):\n\n    response = self.groq.chat.completions.create(\n\n        model=GROQ_MODEL,\n\n        messages=messages,\n\n        tools=self.AGENT_TOOLS,\n\n        tool_choice=\"auto\",\n\n        temperature=0,\n\n    )\n\n```\nmsg = response.choices[0].message\ncalls = getattr(msg, \"tool_calls\", None) or []\n\nif not calls:\n    # The model has enough evidence to return the diagnosis.\n    ...\n    break\n```\n\nI deliberately put a maximum step count around the loop. Autonomous tool calling without a boundary is a good way to turn a simple incident into an unpredictable chain of calls.\n\nThe system prompt also tells the agent to use the minimum useful tools rather than blindly calling everything.\n\nThat matters because an alert does not always need the same investigation.\n\nA detailed stack trace may only require alert inspection and project-memory recall. A vague report such as \"customers say the site feels slow\" needs a different response: ask the engineer for useful information instead of inventing a root cause.\n\nThe tool layer keeps memory boundaries explicit\n\nThe tool implementation makes the privacy boundary visible in code.\n\nif name == \"recall_project_memory\":\n\n    query = arguments.get(\"query\", \"\")[:2000]\n\n    return {\n\n        \"scope\": \"this_project_only\",\n\n        \"facts\": self._recall(bank, query),\n\n    }\n\nif name == \"recall_shared_lessons\":\n\n    if not use_shared:\n\n        return {\n\n            \"scope\": \"shared_lessons_disabled\",\n\n            \"facts\": [],\n\n        }\n\n```\nquery = arguments.get(\"query\", \"\")[:2000]\nreturn {\n    \"scope\": \"anonymized_shared_lessons_only\",\n    \"facts\": self._recall(self.global_bank, query),\n}\n```\n\nI like this pattern because the distinction is not merely UI text.\n\nThe agent gets different tools for different memory scopes, and the tool result explicitly tells it what kind of evidence it received.\n\nThe diagnosis schema reinforces the same rule. Dates, owners, runbooks, and resolution times can only come from project memory. Shared lessons cannot suddenly become historical facts about the current project.\n\nThat is an important guardrail for an incident system. A useful lesson from another service is not evidence that the same thing happened here.\n\nHindsight is the part that makes the loop cumulative\n\nThe Hindsight documentation describes memory as something an agent can retain and recall rather than simply stuffing more text into a context window.\n\nThat maps well to how I structured this system.\n\nAfter an engineer resolves an incident, the resolution is written back into the current project's private memory:\n\ndef resolve_and_retain(\n\n    self,\n\n    alert_text,\n\n    resolution_summary,\n\n    diagnosis=None,\n\n    bank_id=None,\n\n):\n\n    bank = bank_id or self.bank_id\n\n    when = datetime.now(timezone.utc).strftime(\n\n        \"%Y-%m-%d %H:%M UTC\"\n\n    )\n\n```\nrecord = (\n    f\"INCIDENT RESOLVED [{when}]\\n\"\n    f\"ALERT: {alert_text}\\n\"\n    f\"RESOLUTION: {resolution_summary}\"\n)\n\nif diagnosis and diagnosis.get(\"summary\"):\n    record += (\n        f\"\\nAGENT'S EARLIER ASSESSMENT: \"\n        f\"{diagnosis['summary']}\"\n    )\n\nself._retain(bank, record)\n```\n\nThis is intentionally simple.\n\nI do not store only the model's diagnosis. I store the alert and the actual resolution supplied by the engineer. The difference matters because the model's initial explanation can be wrong while the engineer's resolution is the operational fact I want to remember.\n\nThe system also records engineer feedback:\n\nverdict = \"HELPFUL and worked\" if helpful else \"NOT helpful\"\n\nrecord = (\n\n    f\"ENGINEER FEEDBACK: For this alert: {alert_text[:300]}\\n\"\n\n    f\"The recommendation ({steps}) was rated {verdict}. {note}\"\n\n).strip()\n\nself._retain(bank, record)\n\nThis gives future investigations another useful signal: not just what the agent suggested, but whether someone said that suggestion worked.\n\nThat is where agent memory becomes more interesting than simply keeping a transcript.\n\nA connection-pool incident becomes reusable knowledge\n\nThe recurring SQL connection-pool scenario is a good example.\n\nThe alert reports 502 responses and a timeout while obtaining a database connection. In a cold project, the agent has to reason from the alert itself.\n\nAfter the incident has been resolved and retained, a later occurrence can retrieve the previous incident from project memory.\n\nThe remembered resolution in the system is:\n\nThe SQL connection pool was exhausted during the scheduled\n\nreindex job. The incident was resolved by staggering the\n\nreindex operation and increasing the connection pool capacity.\n\nAfter the change, the 502 error rate returned to normal.\n\nThe important behavior is not that the model suddenly \"knows\" SQL connection pools.\n\nIt already knew that.\n\nThe useful change is that it now knows what this project did the last time the same failure appeared.\n\nThat distinction also appears in the UI. The incident console separates:\n\na memory hit in the current project,\n\na lesson seen elsewhere,\n\na genuinely new pattern,\n\nand an alert that needs more information.\n\nThose states are operationally different, so I did not want to collapse them into one generic confidence score.\n\nI also wanted the agent to admit when it does not know\n\nOne of the easiest ways to make an incident assistant dangerous is to reward it for producing an answer every time.\n\nThe diagnosis contract explicitly allows needs_more_info=true and requires clarifying questions when an alert is too vague.\n\nThe system prompt says:\n\nIf the alert is too vague to act on\n\n(no service, symptom, error or metric):\n\nneeds_more_info=true,\n\ngive 3 specific clarifying_questions,\n\nand still give 2-3 safe first checks\n\nin immediate_actions.\n\nThat gives the agent a useful middle state.\n\nIt does not have to choose between \"diagnose everything\" and \"do nothing.\" It can perform safe initial checks, explain what is missing, and ask the engineer for the information needed to continue.\n\nThe same principle applies to memory. If nothing relevant is found, the agent is instructed to mark the pattern as new rather than inventing a historical match.\n\nShared lessons are deliberately harder to publish\n\nThe second Hindsight path is for knowledge that can be useful beyond one project.\n\nAfter a fix is retained privately, the system can generate an anonymized lesson:\n\ndef sanitize_lesson(self, alert_text, resolution, project_name=\"\"):\n\n    raw = self._chat(\n\n        SANITIZE_SYSTEM,\n\n        f\"ALERT:\\n{alert_text}\\n\\nRESOLUTION:\\n{resolution}\"\n\n    )\n\n```\nterms = [project_name] + _service_names(alert_text)\nclean, n = redact((raw or \"\").strip(), terms)\n\nreturn {\"lesson\": clean, \"redactions\": n}\n```\n\nBut the generated lesson is not automatically published to the shared memory.\n\nThe engineer reviews it first. The code also redacts project and service identifiers again after the human edit before retaining the lesson in the shared bank.\n\nThat gives me a two-level memory model:\n\nPrivate project memory\n\n    ├── alerts\n\n    ├── resolutions\n\n    ├── owners\n\n    ├── runbooks\n\n    └── engineer feedback\n\nShared lessons\n\n    └── anonymized, reviewed, generic knowledge\n\nI found this separation more useful than trying to make every memory globally searchable.\n\nOperational details are valuable inside the system that generated them. They are often inappropriate outside it.\n\nAsking the memory is a separate workflow\n\nIncident diagnosis is not the only reason to query operational history.\n\nThe application also exposes a direct \"Ask the memory\" workflow. Questions such as:\n\nWhich services fail most often and why?\n\nWhat recurring root causes should we fix permanently?\n\nWhich past fixes were fastest and what made them fast?\n\nare sent to Hindsight's reflect capability for the current project's private bank.\n\nres = self.hindsight.reflect(\n\n    bank_id=bank,\n\n    query=question,\n\n)\n\nreturn {\n\n    \"answer\": getattr(res, \"text\", None) or str(res),\n\n    \"method\": \"Hindsight reflect\",\n\n}\n\nIf reflection fails, the implementation falls back to recalling relevant memories and asking the language model to answer from that retrieved context.\n\nThis makes the memory layer useful outside the immediate alert path. Incident history becomes something engineers can interrogate rather than a pile of old tickets nobody wants to read.\n\nWhat the system looks like during an incident\n\nThe most useful interaction is a simple one.\n\nAn engineer submits an alert.\n\nThe agent starts by inspecting it. It may then recall private project history. If shared lessons are enabled, it can search those separately. If the alert is vague, it can ask a question instead.\n\nThe UI exposes the resulting tool trace:\n\nAgent activity · autonomous tool calling\n\nAgent mode: autonomous_tool_calling\n\nThe final diagnosis includes the severity, category, confidence, root-cause explanation, matched incidents, immediate actions, and longer-term fixes.\n\nAfter the engineer resolves the incident, they enter the actual resolution and retain it.\n\nFrom that point forward, the incident is no longer just something that happened once.\n\nIt becomes evidence available to the next investigation.\n\nWhat I learned building it\n\nA transcript contains everything, including speculation.\n\nThe resolution is much more valuable than another copy of the model's conversation. I therefore make the engineer-confirmed fix a first-class memory record.\n\n\"Search memory\" is too vague for an operational system.\n\nI needed to know whether a fact came from this project or from a shared lesson. Separate banks and separate tools made that boundary explicit.\n\nLetting the model choose tools is useful, but the model still needs a bounded loop, strict tool descriptions, structured output, and rules about what it is allowed to infer.\n\nAutonomy without constraints is just another failure mode.\n\nPast incidents are evidence, not certainty.\n\nThe diagnosis still contains confidence and a root-cause explanation, and shared lessons are deliberately prevented from being represented as incidents in the current project.\n\nThe most useful feedback is often binary and operational: \"that worked\" or \"that did not work.\"\n\nRecording that feedback means the system can avoid repeating recommendations that engineers already rejected.\n\nWhere I would take it next\n\nThe core loop is now clear:\n\nAlert\n\n  → investigate\n\n  → recall\n\n  → diagnose\n\n  → engineer resolves\n\n  → retain\n\n  → recall next time\n\nThe natural production integrations are around that loop: receiving alerts from monitoring systems, connecting incident records and ticketing systems, posting status updates to engineering communication channels, and automatically retaining confirmed resolutions after incidents close.\n\nThe memory architecture does not need to change for those integrations.\n\nThat is the part I like most about the design. The language model is replaceable. The dashboard is replaceable. The alert source is replaceable.\n\nThe durable piece is the operational memory: what happened, what we tried, what actually fixed it, and what engineers learned afterward.\n\nThat is a much more useful thing for an incident agent to remember than a conversation.", "url": "https://wpnews.pro/news/how-i-built-an-incident-agent-that-remembers-fixes-with-hindsight", "canonical_source": "https://dev.to/aashik/how-i-built-an-incident-agent-that-remembers-fixes-with-hindsight-4i2c", "published_at": "2026-09-28 19:25:40+00:00", "updated_at": "2026-09-28 19:50:20.818545+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "mlops", "artificial-intelligence"], "entities": ["On-Call Copilot", "Hindsight", "Streamlit", "OnCallCopilot"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-i-built-an-incident-agent-that-remembers-fixes-with-hindsight", "markdown": "https://wpnews.pro/news/how-i-built-an-incident-agent-that-remembers-fixes-with-hindsight.md", "text": "https://wpnews.pro/news/how-i-built-an-incident-agent-that-remembers-fixes-with-hindsight.txt", "jsonld": "https://wpnews.pro/news/how-i-built-an-incident-agent-that-remembers-fixes-with-hindsight.jsonld"}}