cd /news/ai-agents/how-i-built-an-incident-agent-that-r… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-141252] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=↑ positive

How I Built an Incident Agent That Remembers Fixes With Hindsight

A developer built On-Call Copilot, an incident-response agent that uses Hindsight as a persistent memory layer to recall project-specific fixes and anonymized shared lessons from past incidents. The system runs a tool-calling loop with four tools β€” inspect_alert, recall_project_memory, recall_shared_lessons, and ask_engineer_question β€” and writes each resolved incident back into a private per-project memory bank, optionally contributing anonymized lessons to a shared bank.

by read11 min views2 publishedSep 28, 2026

How I Built an Incident Agent That Remembers Fixes With Hindsight

The hardest part of incident response is rarely finding an answer once. It is making sure the next engineer does not have to rediscover the same answer six months later.

I built On-Call Copilot around that problem: an incident-response agent that can inspect an alert, decide which information it needs, search project-specific incident memory, use anonymized lessons from elsewhere, ask an engineer for missing context, and retain the eventual resolution for the next incident.

The problem was not diagnosis. It was forgetting.

A conventional incident assistant can take an alert and produce a plausible explanation. That is useful, but it leaves out an important part of operational knowledge: what actually happened in this environment before.

Consider an alert like this:

ALERT [prod]

service=search-api

severity=P1

502 Bad Gateway

rate: 210/min

baseline: 1/min

Logs:

Timeout expired.

The timeout period elapsed prior to obtaining

a connection from the pool.

A language model can recognize the connection-pool symptom. But recognition is not the same as operational knowledge.

I wanted the agent to answer a more useful question: have we seen this in this project before, and what actually fixed it?

That led me to use Hindsight GitHub as the memory layer rather than treating memory as a prompt-management problem.

The important distinction is that the model does not own the memory. Hindsight does. The agent retrieves relevant facts when it needs them, and it writes the outcome back after an engineer resolves the incident.

How the system hangs together

The application is a Streamlit interface around an OnCallCopilot class. The main pieces are deliberately small:

An autonomous AI incident-response copilot that investigates production alerts, uses persistent memory to learn from previous incidents, and recalls proven resolutions for future incidents.

On-Call Copilot is an AI-powered incident-response system designed to assist engineers during production incidents.

Traditional incident-response systems can detect alerts and provide static information, but they do not continuously learn from how engineers actually resolve incidents.

On-Call Copilot addresses this by combining:

The system can investigate an alert, inspect relevant information, retrieve previous lessons, ask an engineer for missing context, generate a diagnosis, and retain the resolved incident so that the knowledge can be reused later.

During production incidents, engineers often need to:

The dashboard handles the incident workflow. The agent owns the reasoning, Hindsight integration, tool execution, and memory lifecycle. Project configuration determines which private memory bank belongs to the active project.

At runtime, the flow looks like this:

Production alert

  β”‚

  β–Ό

Autonomous incident agent

  β”‚

  β”œβ”€β”€ inspect_alert

  β”œβ”€β”€ recall_project_memory

  β”œβ”€β”€ recall_shared_lessons

  └── ask_engineer_question

  β”‚

  β–Ό

Diagnosis + immediate actions

  β”‚

  β–Ό

Engineer resolves incident

  β”‚

  β–Ό

Retain resolution in private Hindsight memory

  β”‚

  └── optionally create anonymized shared lesson

The important design choice is that memory is scoped.

Each project gets a private Hindsight bank containing detailed operational history. There is also a separate shared-lessons bank containing anonymized, generic lessons that a project has explicitly chosen to contribute.

That means a shared lesson can inform an investigation without pretending that the incident happened in the current system.

I stopped making the agent call every tool

One of the more important changes in the project was moving from a fixed investigation sequence to an actual tool-calling loop.

The agent exposes four tools:

AGENT_TOOLS = [

{

    "type": "function",

    "function": {

        "name": "inspect_alert",

        "description": "Parse the incoming incident alert into structured facts before diagnosing it.",

        ...

    },

},

{

    "type": "function",

    "function": {

        "name": "recall_project_memory",

        "description": "Search THIS project's private Hindsight memory for similar incidents, fixes, owners and feedback.",

        ...

    },

},

{

    "type": "function",

    "function": {

        "name": "recall_shared_lessons",

        "description": "Search the anonymized shared Hindsight lessons.",

        ...

    },

},

{

    "type": "function",

    "function": {

        "name": "ask_engineer_question",

        "description": "Use when the alert is too vague to diagnose safely.",

        ...

    },

},

]

The LLM receives these tools and decides what to call next.

for step in range(max_steps):

response = self.groq.chat.completions.create(

    model=GROQ_MODEL,

    messages=messages,

    tools=self.AGENT_TOOLS,

    tool_choice="auto",

    temperature=0,

)
msg = response.choices[0].message
calls = getattr(msg, "tool_calls", None) or []

if not calls:
    ...
    break

I deliberately put a maximum step count around the loop. Autonomous tool calling without a boundary is a good way to turn a simple incident into an unpredictable chain of calls.

The system prompt also tells the agent to use the minimum useful tools rather than blindly calling everything.

That matters because an alert does not always need the same investigation.

A detailed stack trace may only require alert inspection and project-memory recall. A vague report such as "customers say the site feels slow" needs a different response: ask the engineer for useful information instead of inventing a root cause.

The tool layer keeps memory boundaries explicit

The tool implementation makes the privacy boundary visible in code.

if name == "recall_project_memory":

query = arguments.get("query", "")[:2000]

return {

    "scope": "this_project_only",

    "facts": self._recall(bank, query),

}

if name == "recall_shared_lessons":

if not use_shared:

    return {

        "scope": "shared_lessons_disabled",

        "facts": [],

    }
query = arguments.get("query", "")[:2000]
return {
    "scope": "anonymized_shared_lessons_only",
    "facts": self._recall(self.global_bank, query),
}

I like this pattern because the distinction is not merely UI text.

The agent gets different tools for different memory scopes, and the tool result explicitly tells it what kind of evidence it received.

The diagnosis schema reinforces the same rule. Dates, owners, runbooks, and resolution times can only come from project memory. Shared lessons cannot suddenly become historical facts about the current project.

That is an important guardrail for an incident system. A useful lesson from another service is not evidence that the same thing happened here.

Hindsight is the part that makes the loop cumulative

The Hindsight documentation describes memory as something an agent can retain and recall rather than simply stuffing more text into a context window.

That maps well to how I structured this system.

After an engineer resolves an incident, the resolution is written back into the current project's private memory:

def resolve_and_retain(

self,

alert_text,

resolution_summary,

diagnosis=None,

bank_id=None,

):

bank = bank_id or self.bank_id

when = datetime.now(timezone.utc).strftime(

    "%Y-%m-%d %H:%M UTC"

)
record = (
    f"INCIDENT RESOLVED [{when}]\n"
    f"ALERT: {alert_text}\n"
    f"RESOLUTION: {resolution_summary}"
)

if diagnosis and diagnosis.get("summary"):
    record += (
        f"\nAGENT'S EARLIER ASSESSMENT: "
        f"{diagnosis['summary']}"
    )

self._retain(bank, record)

This is intentionally simple.

I do not store only the model's diagnosis. I store the alert and the actual resolution supplied by the engineer. The difference matters because the model's initial explanation can be wrong while the engineer's resolution is the operational fact I want to remember.

The system also records engineer feedback:

verdict = "HELPFUL and worked" if helpful else "NOT helpful"

record = (

f"ENGINEER FEEDBACK: For this alert: {alert_text[:300]}\n"

f"The recommendation ({steps}) was rated {verdict}. {note}"

).strip()

self._retain(bank, record)

This gives future investigations another useful signal: not just what the agent suggested, but whether someone said that suggestion worked.

That is where agent memory becomes more interesting than simply keeping a transcript.

A connection-pool incident becomes reusable knowledge

The recurring SQL connection-pool scenario is a good example.

The alert reports 502 responses and a timeout while obtaining a database connection. In a cold project, the agent has to reason from the alert itself.

After the incident has been resolved and retained, a later occurrence can retrieve the previous incident from project memory.

The remembered resolution in the system is:

The SQL connection pool was exhausted during the scheduled

reindex job. The incident was resolved by staggering the

reindex operation and increasing the connection pool capacity.

After the change, the 502 error rate returned to normal.

The important behavior is not that the model suddenly "knows" SQL connection pools.

It already knew that.

The useful change is that it now knows what this project did the last time the same failure appeared.

That distinction also appears in the UI. The incident console separates:

a memory hit in the current project,

a lesson seen elsewhere,

a genuinely new pattern,

and an alert that needs more information.

Those states are operationally different, so I did not want to collapse them into one generic confidence score.

I also wanted the agent to admit when it does not know

One of the easiest ways to make an incident assistant dangerous is to reward it for producing an answer every time.

The diagnosis contract explicitly allows needs_more_info=true and requires clarifying questions when an alert is too vague.

The system prompt says:

If the alert is too vague to act on

(no service, symptom, error or metric):

needs_more_info=true,

give 3 specific clarifying_questions,

and still give 2-3 safe first checks

in immediate_actions.

That gives the agent a useful middle state.

It does not have to choose between "diagnose everything" and "do nothing." It can perform safe initial checks, explain what is missing, and ask the engineer for the information needed to continue.

The same principle applies to memory. If nothing relevant is found, the agent is instructed to mark the pattern as new rather than inventing a historical match.

Shared lessons are deliberately harder to publish

The second Hindsight path is for knowledge that can be useful beyond one project.

After a fix is retained privately, the system can generate an anonymized lesson:

def sanitize_lesson(self, alert_text, resolution, project_name=""):

raw = self._chat(

    SANITIZE_SYSTEM,

    f"ALERT:\n{alert_text}\n\nRESOLUTION:\n{resolution}"

)
terms = [project_name] + _service_names(alert_text)
clean, n = redact((raw or "").strip(), terms)

return {"lesson": clean, "redactions": n}

But the generated lesson is not automatically published to the shared memory.

The engineer reviews it first. The code also redacts project and service identifiers again after the human edit before retaining the lesson in the shared bank.

That gives me a two-level memory model:

Private project memory

β”œβ”€β”€ alerts

β”œβ”€β”€ resolutions

β”œβ”€β”€ owners

β”œβ”€β”€ runbooks

└── engineer feedback

Shared lessons

└── anonymized, reviewed, generic knowledge

I found this separation more useful than trying to make every memory globally searchable.

Operational details are valuable inside the system that generated them. They are often inappropriate outside it.

Asking the memory is a separate workflow

Incident diagnosis is not the only reason to query operational history.

The application also exposes a direct "Ask the memory" workflow. Questions such as:

Which services fail most often and why?

What recurring root causes should we fix permanently?

Which past fixes were fastest and what made them fast?

are sent to Hindsight's reflect capability for the current project's private bank.

res = self.hindsight.reflect(

bank_id=bank,

query=question,

)

return {

"answer": getattr(res, "text", None) or str(res),

"method": "Hindsight reflect",

}

If reflection fails, the implementation falls back to recalling relevant memories and asking the language model to answer from that retrieved context.

This makes the memory layer useful outside the immediate alert path. Incident history becomes something engineers can interrogate rather than a pile of old tickets nobody wants to read.

What the system looks like during an incident

The most useful interaction is a simple one.

An engineer submits an alert.

The agent starts by inspecting it. It may then recall private project history. If shared lessons are enabled, it can search those separately. If the alert is vague, it can ask a question instead.

The UI exposes the resulting tool trace:

Agent activity Β· autonomous tool calling

Agent mode: autonomous_tool_calling

The final diagnosis includes the severity, category, confidence, root-cause explanation, matched incidents, immediate actions, and longer-term fixes.

After the engineer resolves the incident, they enter the actual resolution and retain it.

From that point forward, the incident is no longer just something that happened once.

It becomes evidence available to the next investigation.

What I learned building it

A transcript contains everything, including speculation.

The resolution is much more valuable than another copy of the model's conversation. I therefore make the engineer-confirmed fix a first-class memory record.

"Search memory" is too vague for an operational system.

I needed to know whether a fact came from this project or from a shared lesson. Separate banks and separate tools made that boundary explicit.

Letting the model choose tools is useful, but the model still needs a bounded loop, strict tool descriptions, structured output, and rules about what it is allowed to infer.

Autonomy without constraints is just another failure mode.

Past incidents are evidence, not certainty.

The diagnosis still contains confidence and a root-cause explanation, and shared lessons are deliberately prevented from being represented as incidents in the current project.

The most useful feedback is often binary and operational: "that worked" or "that did not work."

Recording that feedback means the system can avoid repeating recommendations that engineers already rejected.

Where I would take it next

The core loop is now clear:

Alert

β†’ investigate

β†’ recall

β†’ diagnose

β†’ engineer resolves

β†’ retain

β†’ recall next time

The natural production integrations are around that loop: receiving alerts from monitoring systems, connecting incident records and ticketing systems, posting status updates to engineering communication channels, and automatically retaining confirmed resolutions after incidents close.

The memory architecture does not need to change for those integrations.

That is the part I like most about the design. The language model is replaceable. The dashboard is replaceable. The alert source is replaceable.

The durable piece is the operational memory: what happened, what we tried, what actually fixed it, and what engineers learned afterward.

That is a much more useful thing for an incident agent to remember than a conversation.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @on-call copilot 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/how-i-built-an-incid…] indexed:0 read:11min 2026-09-28 Β· β€”