Picture the 3 a.m. version of this. An alert fires, you open your incident response tool, and the assistant says: "This looks like INC-214: connection pool exhaustion, fixed by raising max_connections." You go looking for INC-214 in your issue tracker. It doesn't exist.
That failure mode is why I built MemoryOps. To be clear, INC-214 is a hypothetical example of an LLM hallucination, not a captured model response: seed data in this repository spans INC-101 through INC-116, so any citation of INC-214 is invented. During an outage, a fabricated citation is worse than a generic answer because a citation reads like verified evidence. You either burn minutes verifying it, or you trust it and apply a fix that was never tested.
I reduced this risk by building persistent operational memory into the incident response lifecycle. Here is how the architecture and implementation work.
MemoryOps is an AI-powered incident response platform for DevOps and SRE teams. The frontend is built with React 19, Vite, and Tailwind CSS (providing Dashboard, Incident Creation, Investigation, and Memory Explorer views). The backend is FastAPI with SQLAlchemy over SQLite (data/incidentiq.db) as the system of record for live incident records.
Hindsight acts as the long-term persistent memory layer for resolved incident learnings, while Groq Cloud LLM (openai/gpt-oss-20b) serves as the AI reasoning engine. Nothing executes actions autonomously; the human engineer remains in full control.
React UI ββΊ FastAPI ββΊ SQLite (incidents, ai_recommendation)
ββββββββΊ Hindsight (RETAIN on resolve, RECALL on analyze, REFLECT)
ββββββββΊ Groq LLM (current incident + recalled memories)
MemoryOps system architecture β React frontend, FastAPI backend, SQLite database of record, Hindsight persistent memory layer, and Groq LLM reasoning engine.
The workflow begins when an engineer declares an incident via POST /api/v1/incidents. Navigating to the investigation page invokes POST /api/v1/incidents/{incident_id}/analyze. The backend recalls relevant past incidents from Hindsight, passes them alongside current symptoms to Groq LLM, and presents evidence-backed recommendations. When the incident is resolved via POST /api/v1/incidents/{incident_id}/resolve, its incident learnings are retained in Hindsight.
Instead of passing massive unstructured log streams to an LLM, MemoryOps uses Hindsight to store structured experience documents. SQLite answers "what is happening now," while Hindsight answers "what did we learn from past outages."
Incident Created βββΊ Investigation βββΊ Root Cause & Resolution βββΊ Hindsight RETAIN βββΊ Future RECALL
An incident is not retained when merely created, during unresolved investigation, or from AI guesses. When an engineer resolves an incident, HindsightService.aretain_incident() stores a complete experience document containing ID, service, error, symptoms, severity, root cause, resolution steps, and post-mortem.
response = await client.aretain(
bank_id=self.bank_id,
content=content_text,
metadata=metadata,
document_id=incident_id,
tags=[service, severity, outcome],
)
To support idempotent retention, document_id is set deterministically to incident.id (for example, INC-101). The Incident database model tracks a memory_retained boolean flag, which is flipped to True only after Hindsight confirms successful retention.
When an investigation is triggered, MemoryOps constructs a semantic search query from the current incident:
recall_query = f"Service: {incident.service} | Error: {incident.error} | Symptoms: {incident.symptoms}"
recalled = await hindsight_service.arecall_memories(query=recall_query, max_tokens=2048)
The router parses returned memories using parse_memory_item() and supplies them as grounded context to Groq.
MemoryOps also provides POST /api/v1/incidents/reflect and a Memory Explorer tab so engineers can query cross-incident patterns across historical outages.
In the investigation API response (IncidentInvestigationResponse), historical evidence and AI reasoning are kept strictly separate:
similar_historical_incidents: populated directly from Hindsight recalled memories.
ai_analysis: contains the structured JSON output returned by Groq LLM.
The UI renders these inputs as distinct pipeline stages so the engineer can evaluate historical evidence independently from LLM reasoning.
![MemoryOps investigation view]
β MemoryOps investigation view displaying the step-by-step pipeline, explicit memory status banner, recalled historical memories, and Groq AI recommendation.
MemoryOps explicitly exposes three distinct memory states in the API and UI:
memory_status = "ok": Hindsight successfully recalled relevant memories (β Historical Memory Used).
memory_status = "empty": Hindsight searched but found no matching memories (β No Relevant Historical Memory).
memory_status = "unavailable": Hindsight service was offline or unconfigured (β Historical Memory Unavailable).
Probable cause: Connection pool exhaustion. Matches INC-214, resolved by increasing max_connections to 100.
seed.py)
{
"probable_root_cause": "Database connection pool exhaustion",
"recommended_action": "Increase max_connections parameter from 20 to 100 and deploy session leak hotfix",
"confidence": "high",
"reasoning": "INC-101 historical incident showed identical Gateway Timeout symptoms and was resolved by expanding the pool.",
"supporting_historical_incidents": [
"INC-101: Payment API database connection timeout"
]
}
System prompt instructions direct Groq to cite only actual recalled memories. However, prompt instructions are guidance rather than mathematical guarantees. Application-side memory-status handling ensures that unavailable or empty Hindsight results are explicitly reported rather than silently represented as historical memory.
When external services fail, MemoryOps degrades gracefully without crashing.
If Hindsight returns an error or HTTP 402 insufficient credits, the backend sets memory_status = "unavailable", clears similar_historical_incidents, and continues investigation using current incident details alone. If Groq LLM is unconfigured, _fallback_analysis() applies rule-based heuristic analysis and sets analysis_status = "fallback" with confidence = "low".
![MemoryOps graceful degradation]
MemoryOps graceful degradation view displaying explicit memory status warning and low-confidence fallback heuristic reasoning when external services are unavailable.
The key lesson is that an incident-response agent does not need to change its model weights to learn from previous incidents. It can improve future investigations by retaining resolved operational experience, recalling relevant evidence when a new incident occurs, and clearly communicating when that memory layer is unavailable