Building AI Support Agents with Hindsight Memory A developer built SupportMind AI, a demo B2B support agent that pairs a Next.js/React/TypeScript frontend and Python/FastAPI backend with Groq for LLM generation and Hindsight Cloud for persistent per-customer memory. The project uses Hindsight's REST endpoints for retain, recall, and reflect so the agent recalls a customer's environment, prior symptoms, and failed or successful fixes before answering, rather than treating each ticket as a blank slate. The implementation is intentionally explicit, calling /{bank_id}/memories, /{bank_id}/memories/recall, and /{bank_id}/reflect directly instead of using a specialized SDK. Building AI Support Agents with Hindsight Memory Support workflows are full of repeated questions: "What version are you on?", "What changed recently?", "Have you already tried the standard fix?" In a stateless AI agent, every ticket starts from zero. The system can answer the current message, but it cannot remember the customer's environment, failed attempts, or recurring issue patterns across tickets. That is the problem the repository in front of me is designed to solve. SupportMind AI is a demo application built with a Next.js/React/TypeScript frontend and a Python/FastAPI backend, with Groq for LLM generation and Hindsight Cloud for persistent memory. The seed data in backend/data/customers.py and backend/data/seed data.py is intentionally synthetic, but it reflects the real engineering challenge: a B2B support agent should not act like a blank slate every time a customer returns. The important distinction is between storing chat history and storing memory. A chat transcript is useful as context for one conversation. Memory is reusable across conversations, customer sessions, and support issues. In a technical support setting, that difference matters because the same customer often reappears with a related problem, a new environment change, or a previously failing workaround. The Problem: Support Without Memory A support agent that is only aware of the active thread is fundamentally limited. If a customer reports an API timeout again, the system does not know whether they are on the same stack, whether they recently upgraded PostgreSQL, whether a previous fix worked, or whether they already tried a cache clear and it failed. In a recurring support scenario, that quickly becomes frustrating and generic. The repository's seeded customer data makes this concrete. For example, acme-corp includes an Enterprise environment profile with Node.js 20.x, PostgreSQL 14, Redis 7.0, AWS ECS/Docker, and an issue history that explicitly records a connection-pool problem and its resolution. The repository also includes a later note that PostgreSQL was upgraded from 14 to 15, which is the kind of environment change that matters to a technical support conversation. This is not a social chat problem. It is an operational memory problem. The ideal support agent should remember: the customer's technical environment prior symptoms and their causes actions that failed actions that succeeded environment changes that may affect future incidents Without that memory, ordinary chatbot behavior is insufficient for recurring technical support. Why This Matters for SupportMind AI SupportMind AI is a small but real example of how memory changes the support workflow. The backend does not rely on a single conversation thread. Instead, it uses Hindsight as a memory layer for each customer. The core idea is straightforward: Recall relevant memories before answering. Generate the response with that context included. Retain new facts after the exchange. Reflect on accumulated patterns when needed. This structure is visible in backend/services/support agent.py, which is the operational heart of the project. The implementation is intentionally direct: it uses Hindsight Cloud as a REST API, not a specialized SDK. In backend/services/hindsight service.py, the repository defines explicit methods for retain, recall, and reflect using HTTP requests to /{bank id}/memories, /{bank id}/memories/recall, and /{bank id}/reflect. That is a useful engineering choice because it keeps the memory layer explicit and inspectable. The Architecture The app is divided into a browser UI, a Python API, and a Hindsight memory service. The frontend uses Next.js 14, React 18, and TypeScript, as shown in frontend/package.json. The backend uses FastAPI and Python, with environment variables configured in backend/config.py and .env.example. The LLM is configured through Groq with GROQ MODEL, defaulting to openai/gpt-oss-120b, and Hindsight is configured through HINDSIGHT BASE URL and HINDSIGHT API KEY. The repository is an architectural example, not a production benchmark. It is a clear demonstration of how an agent memory layer fits into a support system without pretending that the full support workflow has been solved. Making Hindsight the Memory Layer The memory layer is not an extra dashboard feature; it is the core of the support workflow. In backend/services/support agent.py, the service calls Hindsight before generating a response: recall data = await hindsight.recall bank id=customer id, query=message, types= "world", "experience", "observation" , budget="mid", tags= customer id , This is exactly the pattern used by the repo: a customer's bank is queried with the incoming ticket or issue text. The service then formats the result into a readable context string and passes it into the LLM. That is implemented in llm service.py: def build system message recalled context: Optional str - str: if not recalled context: return SYSTEM PROMPT return SYSTEM PROMPT + "\n\n--- CUSTOMER MEMORY CONTEXT ---\n" + recalled context + "\n--- END CONTEXT ---\n" + "\nUse the above context to personalize your response." This is significant because the memory does not sit outside the model as an invisible sidecar. It is injected into the model context, so the response can reference the customer's known environment and earlier support outcomes. The same repo also shows a second LLM call for fact extraction: async def extract facts to retain conversation turn: str, ai response: str - str: prompt = f"""Extract key facts from this support exchange that are worth remembering for future tickets. Focus on: environment details, problem description, solutions tried, outcomes, customer preferences. ... """ This is important because the system is not simply storing the raw chat transcript. It extracts facts such as environment details, problem description, solutions tried, outcomes, and customer preferences. That gives the memory layer more focused information to retrieve than storing every message verbatim. RECALL, LLM Response Generation, and RETAIN The workflow is explicit and sequential: RECALL — hindsight.recall ... LLM — chat completion ..., recalled context=... RETAIN — hindsight.retain ... This is visible in the main support turn function: ai reply = await chat completion messages=messages, recalled context=recalled context if recalled context else None, facts text = await extract facts to retain message, ai reply retain result = await hindsight.retain bank id=customer id, content=facts text, document id=f"session {session id}", context=f"Support session {session id}", metadata={"session id": session id, "customer id": customer id}, tags= customer id , The repository places RECALL before response generation and RETAIN after. That ordering matters. The agent should not answer from a blank context. It should answer from the customer's memory, then record the outcome for the next interaction. SupportMind AI support chat with Hindsight Memory Engine A Concrete Example from the Repo The strongest example in the repository is the seeded acme-corp data. The memory bank is seeded with several facts, including: customer profile: Node.js 20.x, PostgreSQL 14, Redis 7.0, AWS ECS with Docker prior API timeout issue connection pool exhaustion as the root cause a failed attempt: clearing the application cache did not resolve the problem a successful fix: increasing the connection pool size from 10 to 50 a later environment change: PostgreSQL upgraded from version 14 to version 15 These are not invented details. They are in backend/data/seed data.py. In fact, the code explicitly states: Ticket ACME-001: API requests timing out under load Root cause: database connection pool exhausted pool size was 10 Failed attempt: clearing application cache — did NOT resolve the issue Successful resolution: increased connection pool size from 10 to 50 This is exactly the kind of signal a support agent should remember. A later ticket asking about timeouts can pull relevant memories from the same bank, giving the support agent access to information from earlier incidents. The same pattern exists for other customers as well. For example, novatech has webhook delays and Kubernetes/GKE context; cloudpeak has OAuth token refresh issues on Azure AKS; datasphere has connection-pool and large-dataset query context. The project demonstrates that memory is not generic; it is customer-specific and issue-specific. Why Persistent Memory Matters More Than Chat History A chat transcript captures narrative. A memory bank captures facts and patterns. That distinction matters for support work. Chat history can answer: "What did this customer say in this thread?" Memory can answer: "What is the customer's environment, what has failed, what worked, and what changed?" This lets the agent recognize: the environment is not the same as last time the same user has had this issue before the last fix may no longer apply the customer prefers certain response styles the likely root cause is tied to previous incidents This makes the system useful because it keeps support context structured and retrievable across separate conversations. REFLECT: Memory as Insight The repository also includes a reflect method: async def reflect bank id: str, query: str, budget: str = "low", max tokens: int = 4096, tags: Optional list str = None, include facts: bool = True, - dict str, Any : payload = { "query": query, "budget": budget, "max tokens": max tokens, "include": {"facts": {} } if include facts else {}, } Screenshot 2026-09-29 163926 Support insights view showing REFLECT-generated patterns reflect is the portion of the system that synthesizes patterns from accumulated memory. The repo clearly names this as a support-insights workflow. The system collects memory across a bank and can reason over customer-specific patterns. For example, a stored fact like "PostgreSQL upgrade from 14 to 15" alongside earlier connection-pool incidents creates a pattern that is useful to the support team. It is not the same thing as raw chat logs. That is the value of memory: it supports both real-time support and longer-term operational understanding. ChatGPT Image Sep 29, 2026, 05 50 59 PM Genuine Engineering Lessons Customer memory should be isolated by bank. The repo uses bank id=customer id and tags like customer id . This avoids mixing Acme's tickets with other customers and is a clean multi-tenant pattern for demo purposes. Retrieval is not the same as storing transcript. The code intentionally extracts facts before retention, so the system isn't just dumping raw messages into a memory bank. That makes the memory layer more useful for targeted retrieval. Recall precedes generation for a reason. In support agent.py, the agent first retrieves relevant memory and only then calls the LLM. This separation is the core of the memory-aware support workflow. The direct HTTP interface is explicit and debuggable. The repo does not hide Hindsight behind a custom SDK; it calls the REST API directly in hindsight service.py, which makes request structure, payloads, and failures easy to inspect. REFLECT depends on the Hindsight service being properly configured. The feature is useful, but it is not free-standing and requires appropriate infrastructure setup. Limitations The repository is explicit about several constraints. The seeded data is synthetic; the application is a demo environment rather than a production support platform. The Hindsight API is used directly through HTTP, and extract facts to retain has a fallback to raw conversation text if fact extraction fails. That fallback is visible in the code: except Exception as e: logger.warning "Fact extraction failed, retaining raw turn: %s", e return f"Customer reported: {conversation turn :500 }" That is a useful engineering guardrail, but it also shows the system is designed for demonstration and iterative improvement rather than a fully production-hardened support memory stack. Where Hindsight Fits Hindsight is not a decoration in this project. It is the memory layer that turns a generic chatbot into a context-aware support agent. For technical support, that is the key difference between a conversational interface and a useful support system. The project's core idea is simple: support doesn't need a better general-purpose chatbot; it needs a memory-aware support system that remembers customer context, prior fixes, and failed attempts. That is why persistent memory is central to this workflow. The repository links to the official Hindsight resources directly, which is a good fit with the project's purpose: Hindsight GitHub Hindsight documentation What is Agent Memory? This is a technical support memory pattern built around realistic synthetic customer data, explicit recall, and durable retention.