cd /news/ai-agents/building-incidentmind-an-ai-incident… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-141150] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=↑ positive

Building IncidentMind: An AI Incident Investigation Assistant That Learns From the Past

A developer built IncidentMind, an AI-powered incident investigation assistant that investigates software incidents, captures resolutions, and retains that experience in Hindsight memory to inform future investigations. Tested on a production-style authentication incident, the agent traced intermittent HTTP 401 errors to inconsistent JWT signing-key versions across auth-api instances, producing a structured investigation with evidence and likely root causes. The workflow records the confirmed resolution and runbook RB-AUTH-KEY-08 to close the incident-to-memory learning loop.

by read8 min views1 publishedSep 28, 2026

Incident response is rarely difficult because engineers lack the ability to investigate a problem. The harder problem is that teams often have to investigate similar failures repeatedly.

A database connection pool gets exhausted. A deployment introduces a configuration problem. A third-party service starts rate-limiting requests. A cache becomes stale. Engineers solve the problem, document the resolution, and move on. Months later, a similar incident happens and someone starts the investigation again from scratch.

I built IncidentMind to explore a different approach: an AI-powered incident investigation assistant that can investigate an incident, capture what was learned from the resolution, and retain that experience so it can become useful context for future investigations.

The important part isn't simply putting an LLM in front of an incident form. The goal is to create a learning loop around incident response.

Incident β†’ Investigation β†’ Resolution β†’ Memory β†’ Future Investigation

IncidentMind provides a workflow for submitting and investigating software incidents.

An incident contains information such as:

The investigation agent uses that context to produce a structured investigation, including evidence, likely root causes, and recommended actions.

Once the incident has been resolved, the engineer can record:

That experience can then be retained in Hindsight memory.

IncidentMind dashboard showing incident and retained-memory metrics.

The dashboard provides a simple view of the current incident state and the amount of incident experience that has been successfully retained.

For testing the workflow, I used a production-style authentication incident.

The affected service was an Authentication API. A new deployment introduced JWT validation middleware and changed the authentication service's signing-key configuration to use a centralized secrets provider.

After deployment, previously authenticated users were intermittently logged out and new authentication attempts returned HTTP 401 errors.

The logs contained evidence such as:

ERROR auth-api
JWT validation failed:
InvalidSignatureError: signature verification failed

WARN auth-api
POST /auth/refresh returned status=401

WARN auth-api
Instance=auth-api-7c8d9 signing_key_version=v2

WARN auth-api
Instance=auth-api-5f2a1 signing_key_version=v1

The interesting part of the incident is that the database itself remained healthy. The investigation therefore needed to connect the deployment change, authentication failures, and inconsistent configuration across instances.

An example incident submitted to IncidentMind with deployment context, symptoms, and error logs.

Rather than treating the logs as isolated errors, IncidentMind uses the surrounding incident context as part of the investigation.

The investigation agent analyzes the incident context and produces a structured investigation.

The resulting investigation identified the signing-key inconsistency as the likely cause: some authentication instances were using the new signing-key version while others were still using the previous version.

This is the part of the workflow where an AI assistant is useful. Engineers still need to verify the diagnosis, but the assistant can organize the evidence and connect information that might otherwise be spread across deployment history, logs, and symptoms.

After confirming the root cause, the resolution involved bringing all authentication instances onto the same signing-key version, restarting affected instances, and verifying login, token refresh, and existing-session validation.

The resolution was recorded together with a runbook:

RB-AUTH-KEY-08 β€” JWT Signing-Key Synchronization Failure

The lesson learned was also captured: authentication deployments should verify signing-key consistency across instances and include automated configuration checks and controlled key rotation.

This information is more valuable than simply marking the incident as resolved.

It represents operational experience.

The resolved incident is prepared for retention, including root cause, resolution, runbook, and lessons learned.

This is where Hindsight becomes an important part of IncidentMind.

I wanted the application to distinguish between two different kinds of persistence.

SQLite stores application state such as incidents and their status. Hindsight stores the experience that can be useful during future investigations.

When an incident is resolved, IncidentMind builds a structured memory record containing the incident context, root cause, resolution, runbook, and lesson learned.

The relevant implementation is intentionally small:

response = client.retain(
    bank_id=bank,
    content=memory_payload,
    metadata=metadata,
    tags=[service, severity, "incident_resolution"]
)

This shows IncidentMind stores the completed incident resolution in Hindsight.

The application also uses Hindsight's recall capability when looking for historical incident experience relevant to a new investigation.

recall_resp = client.recall(
    bank_id=bank,
    query=query,
    tags=[service] if service else None,
    budget="mid"
)

This shows IncidentMind searches Hindsight for relevant previous incidents when investigating a new problem.

The memory layer showing the Hindsight retain and recall integration.

The distinction matters.

The application database answers questions such as:

Which incidents exist?

The memory layer is intended to answer questions such as:

Have we seen an incident like this before, and what did we learn from it?

That makes retained incident experience part of the investigation workflow rather than just an archive.

The resulting workflow looks like this:

1. An incident occurs

An engineer provides the service, symptoms, logs, and recent changes.

2. IncidentMind investigates

The investigation agent analyzes the available evidence and produces a structured investigation.

3. The engineer confirms the resolution

The confirmed root cause and actual resolution are recorded.

4. The experience is retained

The incident resolution and lesson learned are stored in Hindsight.

5. Future investigations can recall relevant experience

When another incident occurs, historical incident memories can provide additional context.

That last step is what makes the system different from a simple incident tracker.

The goal is not merely to remember that an incident happened. The goal is to preserve useful experience from the incident.

One practical problem appeared while building the application.

Initially, successful Hindsight retention did not immediately translate into a useful dashboard metric. The application was tracking local fallback memories separately from cloud-retained memories, which meant the UI could report zero retained memories even after a successful Hindsight operation.

I changed the application so that successful retention is also reflected in the incident's application state.

The dashboard can therefore show that a resolved incident has actually contributed to the retained-memory count.

The dashboard after a successful retention, showing the updated retained-memory count.

This small implementation detail turned out to be important for observability. A memory operation should not only happen in the background; the application should make its state visible to the person using it.

IncidentMind currently has several relatively focused components.

Streamlit provides the application interface.

The investigation agent orchestrates the incident investigation workflow.

SQLite stores application-level incident state.

Hindsight provides persistent memory for incident experience.

The LLM layer provides the language-model capabilities used by the investigation workflow.

The architecture is intentionally straightforward:

                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚    Streamlit UI     β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ Investigation Agent β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
                            β”‚       β”‚
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       └─────────┐
                  β–Ό                           β–Ό
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”           β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚    SQLite    β”‚           β”‚   Hindsight  β”‚
          β”‚ App State    β”‚           β”‚   Memories   β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

The important architectural separation is between application state and operational memory.

SQLite provides predictable application persistence. Hindsight provides a memory layer specifically useful for retaining and recalling experience.

I also wanted to make sure the application behavior was not dependent entirely on a successful manual demo.

The project includes automated tests covering important application behavior, including incident handling and the retained-memory dashboard counter.

The final test suite passes with:

7 passed

This was particularly useful when changing the memory-retention flow because the dashboard behavior needed to remain correct without requiring a live Hindsight operation during every test.

The project also uses a local fallback memory path when Hindsight credentials are not configured, while the configured environment uses Hindsight Cloud.

One of the biggest lessons from building IncidentMind was that adding memory is not simply a matter of calling a retain API.

The surrounding application needs to answer several questions clearly:

What should be remembered?

For incident response, raw logs alone aren't necessarily the most useful memory. The confirmed cause, resolution, runbook, and lesson learned provide much more operational context.

When should something become memory?

IncidentMind retains experience after an incident has been resolved rather than treating every submitted incident as useful knowledge.

How do you know memory actually worked?

The application needs visible state and tests around the retention workflow. Otherwise, a successful API call can be difficult to distinguish from an application that merely assumes persistence happened.

How should historical memory influence a future investigation?

Recall should provide relevant context without replacing the current incident evidence. A previous incident can be useful evidence, but it should not automatically be treated as the explanation for a new failure.

That distinction is important for building a system engineers can actually trust.

There are several directions I would explore next.

First, I would make similarity-based incident recall more visible during the investigation itself, so an engineer can see which historical experiences influenced the investigation.

Second, I would add richer incident relationships. For example, multiple incidents could be associated with the same service, deployment, dependency, or failure pattern.

Third, I would add stronger evaluation around memory quality. It is not enough to know that something was retained; I want to measure whether recalled memories are actually relevant to the incident being investigated.

Finally, I would explore more structured operational knowledge around runbooks and prevention actions, allowing the system to move beyond incident diagnosis toward continuous improvement of incident-response practices.

IncidentMind started with a simple observation: engineering teams solve the same classes of problems more than once, but the experience from previous incidents is often difficult to bring into the next investigation.

The project explores a workflow where an AI investigation assistant doesn't stop when it identifies a root cause.

It can investigate the incident, record the confirmed resolution, preserve the lesson learned, and retain that experience for future use.

That creates a different model for incident response:

Don't just resolve the incident. Learn from it.

IncidentMind is an experiment in making that learning loop part of the incident investigation workflow, with Hindsight providing the persistent memory layer behind it.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @incidentmind 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/building-incidentmin…] indexed:0 read:8min 2026-09-28 Β· β€”