Designing Hindsight Retain and Recall for SRE Incidents A developer built OpsMind, an AI SRE agent that uses the Hindsight persistent memory layer to recall relevant past incident experiences before diagnosis and retain resolved-incident outcomes afterward, positioning memory between current evidence collection and reasoning rather than as the first or final authority. The workflow gives current logs and metrics priority over historical incidents to prevent stale remediations from driving incorrect automation, and excludes the current incident from recall results. An incident-response system can retrieve the right information and still fail to learn anything. That was one of the design problems behind OpsMind. We wanted an AI SRE agent that could not only recall previous incident experiences while investigating a failure, but also retain the outcome of a resolved incident so that the experience could influence future investigations. The resulting workflow is built around two operations: Recall before diagnosis. Retain after resolution. Hindsight provides the persistent memory layer that connects those two stages. OpsMind receives an incident containing information such as its service, severity, symptoms, and current telemetry. The agent then gathers current evidence from logs and metrics. Only after establishing the current situation does it query Hindsight for historical context. The overall flow is: Incident → Evidence → Hindsight Recall → AI Diagnosis → Human Approval → Resolution → Hindsight Retain The important part is the position of memory in this workflow. Memory is not the first source the agent consults, and it is not the final authority. It sits between current evidence collection and reasoning as a source of additional operational context. Figure 1 — Hindsight connects historical incident experience to the current incident-response workflow. A useful memory system depends heavily on what you ask it to retrieve. For OpsMind, the recall query includes details from the current incident: For example, a payment API incident with high latency, HTTP 500 errors, and high database connection utilization should retrieve previous incidents involving similar operational behavior. The purpose is not to search for an identical incident ID. The purpose is to find relevant operational experience . The query therefore asks Hindsight to identify previous SRE incidents with similar symptoms, service behavior, performance problems, errors, resource saturation, remediation experiences, and outcomes. That gives the AI agent historical context without hardcoding a specific previous incident as the answer. The memory layer uses Hindsight's asynchronous API: python response = await hindsight client.arecall bank id=HINDSIGHT BANK ID, query=query, The result is then processed before being passed into the diagnosis pipeline. One important step is excluding the current incident from the historical results. The current incident should never appear as if it were already historical knowledge. OpsMind also extracts incident identifiers from returned memory and removes duplicates. This gives the reasoning layer a cleaner historical context. The result is effectively a list of previous incident experiences that may be relevant to the current investigation. One of the easiest mistakes in an AI incident-response system would be to let memory override current telemetry. Imagine that an old incident had a database connection problem and was resolved by increasing the connection pool. A new incident might also mention database latency, but its actual root cause could be completely different. If the agent blindly copies the old remediation, persistent memory becomes a source of incorrect automation. OpsMind therefore gives current logs and metrics priority. The system prompt explicitly instructs the AI agent to use current evidence as the primary basis for claims and historical incidents only as supporting context. This creates an important separation: Current telemetry answers: “What is happening now?” Hindsight answers: “What have we experienced before?” The AI uses both to reason about the incident. Recall solves only half of the problem. If OpsMind can retrieve historical experiences but never creates new ones, its knowledge becomes static. The second half of the workflow is therefore retention. After an incident is successfully resolved, OpsMind creates an incident learning record. The record contains: The retention call looks like this: python hindsight client.retain bank id=BANK ID, content=learning record, context="OpsMind SRE incident learning", metadata={ "incident id": str incident id , "service": str service , "type": "incident outcome", "successful": str successful .lower , }, The metadata is useful because it provides structured information alongside the learning record. INC-008 provides a concrete example. The Payment API experienced: OpsMind diagnosed database connection pool exhaustion. After the remediation was approved and simulated successfully, the outcome was retained in Hindsight. Figure 2 — The successful INC-008 resolution is retained as an organizational learning record. At that point, INC-008 changed from a current incident into historical organizational knowledge. That transition is the core behavior we wanted from persistent memory. The strongest test was to investigate another incident after INC-008 had been retained. When INC-007 was analyzed, Hindsight returned INC-008 as historical context. Figure 3 — A later investigation retrieves the newly retained INC-008 experience. This gave us a complete memory lifecycle: INC-008 → Resolve → Retain → INC-007 → Recall INC-008 The important observation is that the newly retained experience was not manually copied into the second investigation. It was available through the memory layer. The memory integration also exposed an asynchronous execution issue. The initial implementation produced: text Timeout context manager should be used inside a task The problem appeared when the synchronous Hindsight recall path interacted with the asynchronous FastAPI request environment. We changed the implementation to use an asynchronous Hindsight client and arecall inside a dedicated asynchronous function. The client is explicitly closed afterward: python await hindsight client.aclose This fixed the event-loop problem and also avoided leaving the underlying client session open. It was a useful reminder that an agent-memory integration has to fit the application's execution model, not just its logical architecture. Retrieval without learning produces a static knowledge base. Retention without retrieval produces information that cannot influence future reasoning. The value comes from connecting both. A vague query can return broadly related incidents. Including service behavior, symptoms, metrics, errors, and resource saturation gives the memory system more useful retrieval context. Previous incidents should inform diagnosis without becoming unquestioned instructions. The current incident remains the primary evidence source. OpsMind retains the outcome of successful remediation because that gives future investigations information about what previously worked. The central memory design in OpsMind is simple: Recall before reasoning. Retain after learning. Hindsight provides the persistent layer that connects those two operations. INC-008 demonstrated the complete lifecycle. Its symptoms were analyzed using current evidence and historical context. After successful resolution, the experience was retained. A later investigation could then recall INC-008 as historical context. That makes memory part of the incident-response lifecycle rather than a separate documentation system.