{"slug": "how-i-designed-an-ai-incident-response-agent-with-hindsight", "title": "How I Designed an AI Incident Response Agent with Hindsight", "summary": "A developer built an AI incident response agent that uses Hindsight as a persistent memory layer, letting the agent recall relevant past incidents alongside current evidence during an investigation and retain post-mortems for future use. The design keeps memory separate from the LLM's reasoning, with a small client wrapping the official hindsight-client SDK to expose retain and recall operations. The agent's flow runs from incident processing through AI investigation, recall of similar incidents, combined current and historical context, and finally retention of the post-mortem.", "body_md": "How I Designed an AI Incident Response Agent with Hindsight\n\nWhen an incident happens, the alert usually isn't the whole story.\n\nThere might be a similar incident from last week, a deployment that happened ten minutes earlier, or a previous post-mortem that already explains part of what we're seeing. If the agent only looks at the current alert, it misses that context.\n\nThat was the problem I wanted to solve with my AI Incident Response Agent.\n\nI wanted the agent to investigate the current incident, but also have a way to look back at previous incidents when that history was actually useful. For that, I used Hindsight as the persistent memory layer.\n\nThe architecture ended up being fairly simple. The interesting part was figuring out what to put into memory and how to bring the right memories back into an investigation.\n\nThe basic architecture\n\nAt a high level, the flow looks like this:\n\nSecurity Incident\n\n       ↓\n\nIncident Processing\n\n       ↓\n\nAI Investigation Agent\n\n       ↓\n\nRecall Relevant Incidents\n\n       ↓\n\nCurrent Evidence + Historical Context\n\n       ↓\n\nInvestigation / Decision\n\n       ↓\n\nRetain Post-mortem\n\nThere are really two directions here.\n\nThe first is from the past into the current investigation:\n\nPrevious incidents → Current investigation\n\nThe second is from the current investigation into the future:\n\nCurrent investigation → Future incidents\n\nThat second part is easy to overlook. Once an incident is finished, the investigation shouldn't necessarily disappear. Some of what we learned from it can become useful the next time something similar happens.\n\nThat's where Hindsight fits.\n\nWhy I didn't want a stateless agent\n\nA normal LLM workflow can do a good job if you give it all the information it needs in the prompt.\n\nBut incident response doesn't really work that way.\n\nImagine an alert like this:\n\nService: payments-api\n\nSymptoms:\n\nDeployment:\n\nv2.4.1 deployed 12 minutes ago\n\nThere is already quite a bit for the agent to investigate.\n\nBut now imagine that the same service had a similar incident a few weeks ago. During that investigation, the team discovered that a particular deployment change caused almost the same symptoms.\n\nThat information could be extremely useful.\n\nA stateless agent doesn't know about it unless we explicitly put it into the current context.\n\nThat's what I wanted to change.\n\nInstead of making the agent start from zero every time, I wanted it to be able to ask:\n\n\"Have we seen something like this before?\"\n\nThat sounds simple, but it changes the architecture quite a bit.\n\nKeeping memory separate from reasoning\n\nOne decision I made was to keep the memory layer separate from the LLM itself.\n\nThe LLM is responsible for reasoning about the incident.\n\nThe application is responsible for orchestrating the investigation.\n\nHindsight handles persistent memory.\n\nSo instead of putting memory logic everywhere in the application, I created a small client around it.\n\nThe rest of the incident-response code can then think in terms of:\n\nretain incident\n\nrecall incidents\n\nrather than worrying about the details of the memory backend.\n\nThat separation became useful because it keeps the responsibilities pretty clear.\n\nTurning an incident into memory\n\nThe first half of the problem is retention.\n\nWhen an incident has been investigated, I want to preserve the useful parts of that investigation.\n\nHere's the core of the Hindsight integration:\n\nclass HindsightMemoryClient:\n\n    \"\"\"Production wrapper around official hindsight-client SDK.\"\"\"\n\n``` php\ndef retain_incident(self, incident: Incident) -> RetainResponse:\n    \"\"\"Retain a post-mortem into Hindsight with rich context and metadata.\"\"\"\n    content = format_incident_as_memory(incident)\n    doc_id = f\"incident_{incident.incident_id}\"\n\n    return self.client.retain(\n        bank_id=self.bank_id,\n        content=content,\n        document_id=doc_id,\n        metadata={\n            \"incident_id\": incident.incident_id,\n            \"service\": incident.service,\n            \"severity\": str(incident.severity),\n            \"root_cause\": incident.root_cause or \"\",\n            \"runbook\": incident.runbook or \"\",\n        },\n        tags=[\n            f\"service:{incident.service}\",\n            f\"severity:{incident.severity}\",\n            f\"incident:{incident.incident_id}\",\n            \"type:incident\",\n        ],\n    )\n```\n\nThere are a few things I care about here.\n\nFirst, I'm not just throwing the raw incident object into memory. format_incident_as_memory() turns the incident into something that can actually make sense as a stored investigation.\n\nI'm also giving every incident a predictable document ID:\n\ndoc_id = f\"incident_{incident.incident_id}\"\n\nThen I attach metadata such as the service, severity, root cause, and runbook.\n\nThat matters because two incidents can have similar symptoms but completely different operational contexts.\n\nA critical incident in payments-api isn't necessarily comparable to a low-severity issue in an internal service just because both contain an authentication error.\n\nThe tags give me another layer of structure:\n\nservice:payments-api\n\nseverity:critical\n\nincident:1234\n\ntype:incident\n\nSo the memory isn't just a blob of text. It carries some of the context that explains what that memory actually represents.\n\nFinding the right memories\n\nStoring incidents is only half the problem.\n\nThe more interesting part is retrieval.\n\nWhen a new incident arrives, I don't want to dump the entire incident history into the LLM. That would be noisy and expensive, and most of it probably wouldn't be relevant anyway.\n\nInstead, I build a query from the current incident.\n\nHere's the recall side:\n\ndef recall_incidents(self, incident: Incident) -> list[RecalledIncidentMemory]:\n\n    \"\"\"Synthesize current symptoms, error logs, and deployment deltas into a semantic query.\"\"\"\n\n    query_str = (\n\n        f\"Service: {incident.service}. \"\n\n        f\"Symptoms: {'; '.join(incident.symptoms)}. \"\n\n        f\"Errors: {' '.join(incident.error_logs[:2])}. \"\n\n        f\"Deployment: {incident.deployment_version} ({incident.deployment_minutes_before}m ago).\"\n\n    )\n\n```\nresponse = self.client.recall(\n    bank_id=self.bank_id,\n    query=query_str,\n    max_tokens=4096\n)\n```\n\nI like this approach because the query represents the incident as a whole.\n\nIt's not just:\n\npayments-api\n\nIt's closer to:\n\npayments-api + current symptoms + recent errors + deployment information\n\nThat gives the recall operation something meaningful to work with.\n\nThe agent can then get historical incidents that actually have some relationship to the problem it is currently investigating.\n\nKeeping the recalled data useful\n\nOnce Hindsight returns results, I map them into an application-level object:\n\nreturn [\n\n    RecalledIncidentMemory(\n\n        id=item.id,\n\n        document_id=item.document_id,\n\n        content=item.text,\n\n        score=item.scores.get(\"total\") if item.scores else None,\n\n        tags=item.tags or [],\n\n        root_cause=item.metadata.get(\"root_cause\"),\n\n        resolution=item.metadata.get(\"resolution\"),\n\n    )\n\n    for item in response.results\n\n]\n\nOne detail I deliberately kept here is that I use the score returned by the memory system when it exists.\n\nI don't try to manufacture my own similarity score or make the result look more precise than it is.\n\nThe application gets the recalled content, metadata, tags, and available score and can then pass the relevant information into the investigation flow.\n\nThat keeps the memory layer fairly self-contained.\n\nWhat this looks like during an incident\n\nLet's take the payments-api example again.\n\nA deployment happens.\n\nTwelve minutes later, the service starts showing authentication failures and an elevated error rate.\n\nThe current incident contains:\n\nService: payments-api\n\nSymptoms:\n\nDeployment:\n\nv2.4.1\n\n12 minutes ago\n\nThe agent can investigate that information normally.\n\nBut before reaching a conclusion, it can also recall previous incidents with similar characteristics.\n\nSuppose it finds an earlier incident involving the same service where authentication failures started shortly after a deployment.\n\nThat doesn't mean the new incident has the same root cause.\n\nThat's important.\n\nThe previous incident is evidence worth considering, not an answer.\n\nThe agent still has to look at the current deployment, current logs, and current symptoms.\n\nThe useful context becomes:\n\nCurrent evidence\n\n        +\n\nRelevant historical incidents\n\n        ↓\n\n     Investigation\n\nThat's the behavior I wanted from the memory layer.\n\nWhat happens when memory doesn't help?\n\nThis was another thing I wanted to account for.\n\nNot every incident is going to have a useful historical match.\n\nSometimes the system will be dealing with something completely new.\n\nSometimes the previous incidents will look similar but turn out to have different causes.\n\nSometimes there simply won't be enough historical information.\n\nThat's fine.\n\nThe agent shouldn't become dependent on memory to function.\n\nThe fallback is straightforward:\n\nRelevant memory found\n\n    ↓\n\nCurrent evidence + historical context\n\nNo relevant memory\n\n    ↓\n\nCurrent evidence\n\nThe current incident always remains the primary source of evidence.\n\nHindsight adds context when that context exists.\n\nThe part I find most important\n\nThe interesting thing about this architecture isn't that the agent can search old incidents.\n\nIt's that completed investigations can become useful inputs for future investigations.\n\nThe loop looks like this:\n\nInvestigate\n\n    ↓\n\nLearn\n\n    ↓\n\nRetain\n\n    ↓\n\nRecall later\n\n    ↓\n\nInvestigate with more context\n\n    ↓\n\nLearn again\n\nThat gives the agent a kind of continuity.\n\nA new incident doesn't necessarily have to be the first time the system has encountered a particular pattern.\n\nWhat I learned building it\n\nIt is easy to say that an agent should \"remember everything.\"\n\nIn practice, that's not what I want.\n\nThe useful question is:\n\nWhat information could actually help with a future investigation?\n\nThat changes what gets retained and how the memory is structured.\n\nA generic memory search isn't very useful if it returns unrelated incidents.\n\nBuilding the recall query from the current service, symptoms, errors, and deployment context gives historical retrieval a much stronger connection to the investigation.\n\nThis is probably the most important rule in the system.\n\nA previous incident can suggest a direction for investigation.\n\nIt shouldn't determine the conclusion.\n\nThe current evidence still matters.\n\nThe HindsightMemoryClient gives the rest of the application a simple interface for retaining and recalling incidents.\n\nThat keeps the Hindsight-specific details in one place and makes the rest of the system easier to work with.\n\nA stateless agent answers the question in front of it.\n\nA memory-enabled agent can also use what it learned previously.\n\nFor incident response, that distinction is useful because investigations naturally produce knowledge that can matter again later.\n\nFinal Thoughts\n\nThe biggest architectural decision in my AI Incident Response Agent wasn't the LLM itself.\n\nIt was deciding what happens to the knowledge produced after an investigation is finished.\n\nIf that knowledge disappears, the next investigation starts from scratch.\n\nIf useful parts of it are retained, the next investigation has another source of context.\n\nThat's why Hindsight became an important part of the architecture.\n\nThe final loop is simple:\n\nIncident\n\n   ↓\n\nInvestigate\n\n   ↓\n\nRetain what was learned\n\n   ↓\n\nNew incident\n\n   ↓\n\nRecall relevant history\n\n   ↓\n\nInvestigate with current + historical context\n\nThe agent doesn't need to remember everything.\n\nIt needs to remember the right things, retrieve them when they are relevant, and still validate them against what is happening now.\n\nThat's the part of the system I'm most interested in: turning previous incident investigations from something that simply gets archived into something that can actively contribute to the next investigation.\n\nFor the memory layer, I used Hindsight on GitHub and the Hindsight documentation. The broader idea of agent memory from Vectorize is also useful for understanding why persistent memory is different from simply putting more text into an LLM prompt.\n\nThe agent doesn't have to start from zero every time.", "url": "https://wpnews.pro/news/how-i-designed-an-ai-incident-response-agent-with-hindsight", "canonical_source": "https://dev.to/guru06ashish/how-i-designed-an-ai-incident-response-agent-with-hindsight-1k45", "published_at": "2026-09-28 16:15:27+00:00", "updated_at": "2026-09-28 16:21:35.877923+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "large-language-models", "ai-tools"], "entities": ["Hindsight", "hindsight-client"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-i-designed-an-ai-incident-response-agent-with-hindsight", "markdown": "https://wpnews.pro/news/how-i-designed-an-ai-incident-response-agent-with-hindsight.md", "text": "https://wpnews.pro/news/how-i-designed-an-ai-incident-response-agent-with-hindsight.txt", "jsonld": "https://wpnews.pro/news/how-i-designed-an-ai-incident-response-agent-with-hindsight.jsonld"}}