{"slug": "i-gave-an-incident-agent-a-memory-with-hindsight", "title": "I Gave an Incident Agent a Memory With Hindsight", "summary": "A developer built MemoryOps, an AI incident response agent that pairs an LLM with the Hindsight memory layer so it can recall similar past incidents and retain new resolutions for future investigations. The system separates structured incident data in PostgreSQL from operational experience stored in Hindsight, creating a learning loop where resolved incidents inform later recommendations that a human still reviews before resolution.", "body_md": "Production incidents rarely happen in isolation.\n\nA payment API becomes slow. A Redis connection pool gets exhausted. A deployment introduces unexpected errors. Someone investigates the issue, finds the fix, resolves the incident, and moves on.\n\nBut when a similar incident happens again, a new investigation often starts from scratch.\n\nThat was the problem I wanted to solve with MemoryOps: an AI incident response agent that can investigate an incident, recall what happened in similar incidents, recommend an action based on that experience, and retain the new resolution for future incidents.\n\nThe key part is not just the AI model.\n\nIt is memory.\n\nA typical incident response workflow has access to a lot of information:\n\nThe difficult part is connecting the current incident with relevant historical experience.\n\nA language model can reason about the information given to it, but I wanted the agent to do something more useful:\n\n\"Have we seen something like this before, and what worked last time?\"\n\nThat is where Hindsight became the memory layer of the system.\n\nI designed MemoryOps around a simple separation of responsibilities.\n\n```\ntext\n                    ┌─────────────────────┐\n                    │   React Frontend    │\n                    │  Incident Dashboard │\n                    └──────────┬──────────┘\n                               │\n                               ▼\n                    ┌─────────────────────┐\n                    │     FastAPI API     │\n                    └──────────┬──────────┘\n                               │\n                               ▼\n                    ┌─────────────────────┐\n                    │  Incident Response │\n                    │       Agent        │\n                    └───────┬─────┬───────┘\n                            │     │\n                  ┌─────────┘     └──────────┐\n                  ▼                          ▼\n          ┌───────────────┐          ┌───────────────┐\n          │      LLM      │          │   Hindsight   │\n          │   Reasoning   │          │ Long-Term Mem │\n          └───────────────┘          └───────────────┘\n                            │\n                            ▼\n                    ┌─────────────────┐\n                    │   PostgreSQL    │\n                    │ Structured Data │\n                    └─────────────────┘\n\nPostgreSQL stores structured application information such as incident IDs, services, severity, status, logs, metrics, and deployment information.\nHindsight has a different responsibility.\nIt stores and recalls operational experience.\nThis distinction became important during development. I did not want to use the memory system as another traditional database. I wanted it to answer questions such as:\n\"What happened in previous incidents that looked similar to this one?\"\n\nThe Memory Loop\nThe core workflow is:\nIncident\n   ↓\nInvestigate\n   ↓\nRecall previous experience\n   ↓\nCombine current evidence + historical memory\n   ↓\nGenerate recommendation\n   ↓\nHuman reviews the recommendation\n   ↓\nResolve incident\n   ↓\nRetain the outcome\n   ↓\nFuture incidents can recall it\n\nThis creates a learning loop.\n![ ](https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/adns5uglgk209ubowp89.png)\n\nThe agent does not simply answer a question and forget the interaction.\nA resolved incident becomes useful information for a future investigation.\nUsing Hindsight for Incident Recall\nThe agent creates a query from the current incident.\nFor example:\nquery = f\"\"\"Incident: {incident.summary}Service: {incident.service}Logs:{incident.logs}Metrics:{incident.metrics}Deployment:{incident.deployment}\"\"\"memories = hindsight.recall_incidents(query)\n\nThe memory service sends this information to Hindsight and retrieves relevant previous experiences.\nThe agent then includes those memories when asking the LLM to analyze the incident.\nThe important part is that the model is not reasoning only from the current incident.\nIt is reasoning from:\nCurrent Incident\n       +\nCurrent System Evidence\n       +\nRelevant Historical Experience\n       ↓\n     Analysis\n\nThat changes the behavior of the system considerably.\nAn Example\nSuppose the current incident is a critical authentication-service failure.\nThe agent receives information such as:\nService: auth-service\nSeverity: CRITICAL\n\nSymptoms:\nHTTP 500 responses increased after deployment.\n\nMetrics:\nError rate increased significantly.\n\nLogs:\nConnection pool exhaustion detected.\n\nThe agent can then recall previous incidents.\nFor example, one historical incident involved a payment service where a Redis connection pool became exhausted. The resolution involved increasing the pool size and verifying the service after the change.\nThe historical incident does not automatically become the answer.\nInstead, it becomes evidence that the agent can consider.\nThe LLM combines that evidence with the current incident data and generates a structured response containing:\n- Root cause\n- Confidence\n- Recommended actions\n- Reasoning\n- Retrieved memories\nThis keeps the decision process understandable instead of simply returning a single opaque answer.\nTurning Resolutions Into Memory\nThe other half of the system is just as important.\nAfter an incident is resolved, the outcome is retained in Hindsight.\nmemory_text = f\"\"\"Incident {incident_id} was resolved.Service: {incident.service}Severity: {incident.severity}Summary:{incident.summary}Logs:{incident.logs}Metrics:{incident.metrics}Deployment:{incident.deployment}Resolution:{resolution}Outcome:Incident resolved successfully.\"\"\"hindsight.retain(    bank_id=bank_id,    content=memory_text,    context=\"Production incident response experience\",    document_id=incident_id,)\n\nNow the resolution becomes part of the agent's future experience.\nThis is the part I found most interesting.\nThe system is not just retrieving documentation.\nIt is building a collection of operational experiences from previous incidents.\nA Small Engineering Problem I Ran Into\nOne of the more useful lessons came from something that initially looked unrelated to memory.\nThe LLM was generating slightly different representations for confidence.\nAt one point, the model returned:\n{\n  \"confidence\": \"high\"\n}\n\nwhile the API expected a floating-point value.\nLater, another response returned:\n{\n  \"confidence\": 0.9\n}\n\nwhile the schema had temporarily been changed to expect a string.\nThe API validation failed because the generated output and the Pydantic schema disagreed.\nI fixed this by normalizing the model output before returning it from the agent.\nFor example:\nconfidence_map = {    \"very low\": 0.2,    \"low\": 0.4,    \"medium\": 0.6,    \"high\": 0.9,    \"very high\": 0.95,}\n\nThe final API schema uses:\nclass InvestigationOut(BaseModel):    incident_id: str    root_cause: str    confidence: float    recommendation: list[str]    reasoning: str    memories: list[dict]\n\nThis was a good reminder that LLM applications need strong boundaries between probabilistic model output and deterministic application code.\nWhy Human Approval Matters\nI deliberately did not design the agent to automatically execute arbitrary production commands.\nThe agent investigates the incident and recommends an action.\nA human can review the recommendation before anything operational is changed.\nThat gives the system this workflow:\nAI investigates\n      ↓\nAI recommends\n      ↓\nHuman reviews\n      ↓\nHuman approves\n      ↓\nResolution\n      ↓\nOutcome becomes memory\n\nFor an incident-response system, this separation is important because a recommendation and an actual production change are two different things.\nWhat I Learned\nBuilding MemoryOps changed how I think about memory in AI agents.\n1. Memory should have a purpose\nAdding a memory system does not automatically make an agent intelligent.\nThe memory needs to answer a useful question.\nFor this system, the question is:\n\"What previous incident experience is relevant to this incident?\"\n\n2. Structured Data and Agent Memory Are Different\nPostgreSQL is useful for structured application state.\nHindsight is useful for remembering and recalling experiences.\nKeeping those responsibilities separate made the architecture easier to reason about.\n3. Retrieval Is Only Useful When It Changes the Decision\nThe goal is not to retrieve the largest number of memories.\nThe goal is to retrieve memories that provide useful context for the current investigation.\n4. LLM Output Needs Validation\nEven when the model is instructed to return JSON, application code should still validate and normalize the result.\nThe confidence-field issue was a small example of why this matters.\n5. The Interesting Part Is the Learning Loop\nThe most useful property of the system is not simply that it can investigate an incident.\nIt is that:\nIncident → Investigation → Resolution → Memory\n                         ↑\n                         │\n                    Future Recall\n\nEvery resolved incident can potentially make future investigations more informed.\nWhat's Next\nThere are several areas I would explore next:\n- Connecting the agent to real observability platforms\n- Adding richer incident timelines\n- Improving memory retrieval and filtering\n- Tracking whether recommended actions actually solved incidents\n- Adding stronger evaluation for retrieved memories\n- Supporting multiple services and dependency relationships\n- Adding more detailed approval and audit workflows\nThe current implementation is intentionally focused on one core idea: giving an incident-response agent persistent operational memory.\nFinal Thoughts\nAI agents are often described in terms of reasoning, tools, and autonomous actions.\nI think memory deserves the same level of attention.\nAn agent that can investigate today's incident is useful.\nAn agent that can remember what happened yesterday, understand why a previous solution worked, and use that experience when investigating tomorrow's incident is a different kind of system.\nThat is what I wanted to explore with MemoryOps.\nHindsight provided the persistent memory layer that made this possible.\nThe project is available on GitHub:\nhttps://github.com/sriamsatwik2005/memoryops-ai-incident-response\nHindsight:\nhttps://github.com/vectorize-io/hindsight\nHindsight documentation:\nhttps://hindsight.vectorize.io/****\n```\n\n", "url": "https://wpnews.pro/news/i-gave-an-incident-agent-a-memory-with-hindsight", "canonical_source": "https://dev.to/sriram_satwik/i-gave-an-incident-agent-a-memory-with-hindsight-428g", "published_at": "2026-09-29 11:13:43+00:00", "updated_at": "2026-09-29 11:16:38.395761+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "artificial-intelligence", "mlops"], "entities": ["MemoryOps", "Hindsight", "PostgreSQL", "FastAPI", "React"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-gave-an-incident-agent-a-memory-with-hindsight", "markdown": "https://wpnews.pro/news/i-gave-an-incident-agent-a-memory-with-hindsight.md", "text": "https://wpnews.pro/news/i-gave-an-incident-agent-a-memory-with-hindsight.txt", "jsonld": "https://wpnews.pro/news/i-gave-an-incident-agent-a-memory-with-hindsight.jsonld"}}