{"slug": "teaching-a-ci-cd-failure-agent-to-remember-lessons-from-building-pipelinesage-on", "title": "Teaching a CI/CD Failure Agent to Remember: Lessons from Building PipelineSage on Hindsight", "summary": "A developer built PipelineSage, a Python and Streamlit agent that diagnoses CI/CD pipeline failures by recalling similar past incidents from the Hindsight agent memory system, re-ranking them, and grounding an LLM diagnosis (openai/gpt-oss-120b on Groq at temperature 0.1) only in those recalled memories. The design deliberately retains only structured, human-confirmed incident records rather than the model's own unverified recommendations, so that recalled memories act as evidence rather than compounding unverified precedent.", "body_md": "The third time a payment-service deploy died on a database migration timeout, the fix was sitting in a closed incident from weeks earlier, and nobody on call could find it. That's the whole motivation for PipelineSage: pipeline failures repeat, and the knowledge of how we fixed them last time is scattered across tickets, chat threads, and people's heads.\n\nI built an agent that diagnoses CI/CD failures. But the LLM call is the least interesting part. The real work went into the memory: what gets written, who is allowed to write it, and how it comes back out.\n\nWhat it does and how it hangs together\n\nPipelineSage is a small Python app with a Streamlit dashboard. You pick a failed deployment and the flow is:\n\nfailed deployment\n\n  → recall similar incidents from Hindsight\n\n  → re-rank the recalled memories\n\n  → LLM diagnosis grounded only in those memories\n\n  → recommended fix\n\n  → human confirms the outcome\n\n  → retain the confirmed outcome in Hindsight\n\nThe layout is deliberately boring:\n\napp.py                        # Streamlit dashboard\n\nagent/pipeline_agent.py       # recall → re-rank → prompt → diagnose\n\nmemory/hindsight_memory.py    # thin wrapper over Hindsight retain/recall\n\nservices/pipeline_service.py  # loads pipeline runs and incident history\n\nThe model is openai/gpt-oss-120b on Groq, run at temperature=0.1. Memory lives in Hindsight Cloud. The agent never touches the Hindsight client directly; it talks to a HindsightMemory class with two methods, retain_incident and recall. That boundary turned out to matter, and I'll come back to it.\n\nThe through-line: an agent's memory has to be evidence, not opinion\n\nThe design question that took the most thought was what the agent is allowed to remember.\n\nThe naive version is easy: after every diagnosis, store the failure log and the model's answer. I didn't, because of how recall feeds back into generation. If the agent stores its own unverified recommendations, the next similar failure recalls that recommendation as \"historical precedent.\" Confidence compounds; correctness doesn't.\n\nWhy Hindsight instead of my own vector table\n\nMy first instinct was a table of embeddings and a similarity query. I've built that before, and it works until you want anything past nearest-neighbor text matching. Then you own chunking, embedding refresh, dedup, and ranking, none of which is the thing I'm trying to build.\n\nHindsight, the open source agent memory system from Vectorize, gave me exactly two verbs that map onto the problem: retain to store an experience and recall to pull back what's relevant. If the concept is new to you, Vectorize has a solid explainer on what agent memory is and how it differs from stuffing chat history into a prompt, and the Hindsight documentation covers the API I use.\n\nThe wrapper is small. This is the recall side:\n\npython\n\ndef recall(self, query, limit=5):\n\n    result = self.client.recall(\n\n        bank_id=self.bank_id,\n\n        query=query,\n\n        max_tokens=4096,\n\n        budget=\"mid\",\n\n    )\n\n```\nmemories = []\nfor item in getattr(result, \"results\", []) or []:\n    text = getattr(item, \"text\", None)\n    if text:\n        memories.append(text)\nreturn memories[:limit]\n```\n\nWhat I store: structured incidents, not log dumps\n\nEvery retained incident is rendered into the same fixed shape: deployment, service, branch, environment, commit, status, failure, root cause, infrastructure change, resolution, outcome, and a pointer to the related historical incident.\n\npython\n\ncontent = f\"\"\"\n\nDevOps pipeline incident.\n\nDeployment: #{incident['deployment_id']}\n\nService: {incident['service']}\n\n...\n\nFailure:\n\n{incident['error']}\n\nRoot cause:\n\n{incident.get('root_cause', 'Not yet confirmed.')}\n\nResolution:\n\n{incident.get('resolution', 'Not yet resolved.')}\n\nOutcome:\n\n{incident.get('outcome', 'No outcome recorded.')}\n\nRelated historical incident:\n\n{incident.get('related_historical_incident', 'Not specified.')}\n\n\"\"\"\n\nTwo things fall out of this. First, the explicit Not yet confirmed. defaults mean a memory can honestly say it has no verified root cause, and the model can see that. Second, the Related historical incident field means a retained outcome links back to the precedent that produced it, so a chain like #1057 → #1017 stays legible when a human reads the memory in the dashboard.\n\nRecall is not the end of retrieval\n\nSemantic recall gets you candidates. It doesn't guarantee the top result is the one you should act on. A user-service missing-environment-variable incident can be semantically \"close\" to a payment-service failure just because both are production deploys that failed at startup.\n\nSo the agent issues a few differently-phrased queries built from the current incident's service and error text, de-duplicates the results, throws away the current deployment itself (an incident must never be its own precedent), and re-ranks with plain, inspectable signals:\n\npython\n\nif service and service.lower() in text:\n\n    score += 20          # same service\n\nif \"migration\" in text:\n\n    score += 15          # same failure pattern\n\nif \"timeout\" in text:\n\n    score += 10\n\nif \"succeeded\" in text:\n\n    score += 25          # prefer resolutions that worked\n\nif \"batch\" in text:\n\n    score += 15\n\nif \"dependency conflict\" in text:\n\n    score -= 30          # different failure family\n\nThis is crude and I know it. Substring scoring is a heuristic, and I'd rather have a heuristic I can read in thirty seconds and argue with than an opaque rerank I can't debug at 2 a.m. What matters is that successful, same-service, same-pattern incidents float to the top and unrelated families sink. The top five go into the prompt.\n\nConstraining the model to what memory actually says\n\nThe failure mode I was most worried about wasn't a missing answer. It was a plausible invented one: an LLM helpfully \"improving\" a documented fix. If memory says the fix was batches of 500 records, I do not want \"batches of 200–1000 records to keep transactions under the timeout.\" So the system prompt is mostly prohibitions:\n\nThe response is forced into three sections: Historical Evidence, Diagnosis, Recommended Fix. That separation is what lets an engineer check the claim against the evidence directly above it.\n\nWhat it looks like in use\n\nDeployment #1057 of payment-service fails: the migration times out while updating historical transaction rows. Earlier, deployment #1017 hit a similar timeout, and the memory in Hindsight records its root cause (one large transaction exceeding the deployment timeout), its resolution (split into batches of 500 records), and its outcome (deployment succeeded).\n\nWhen I run analysis on #1057, the dashboard shows the recalled memories first, then the three-part diagnosis. The recommendation is to process the historical transaction rows in batches of 500 records, citing #1017 as the evidence. After I confirm the fix worked, #1057 is retained with a link back to #1017. The next migration timeout now has two confirmed precedents instead of one.\n\nI want to be precise about the claim. This demonstrates the loop: an outcome retained on one day changes the diagnosis on another, with the evidence on screen. I haven't measured time-to-resolution, and I'd distrust any number I couldn't back up.\n\nLessons learned\n\nDecide who may write to memory before deciding what to store. Gating writes on human-confirmed outcomes was the most important decision in the project. An agent that remembers its own guesses builds a confident echo chamber.\n\nShow the recalled evidence next to the answer. It turns the agent from an oracle into something an on-call engineer can audit in seconds, and it makes bad recall obvious immediately.\n\nRetrieval needs a second pass you can read. Semantic recall gives you candidates; a small, explicit re-ranker gave me control over same-service and same-pattern preference, and excluding the incident under analysis from its own history.\n\nPrompt against embellishment, not just against silence. The dangerous output isn't \"I don't know.\" It's a fix with a number the memory never contained. Preserve documented values exactly and make \"insufficient evidence\" an acceptable answer.\n\nWhat's next\n\nProduction changes stay under human control. PipelineSage recommends; people decide. The next thing I want to capture is why a fix failed when it does, because that's the context hardest to reconstruct later and the most valuable to recall.", "url": "https://wpnews.pro/news/teaching-a-ci-cd-failure-agent-to-remember-lessons-from-building-pipelinesage-on", "canonical_source": "https://dev.to/medarapu_muralikrishna_b/teaching-a-cicd-failure-agent-to-remember-lessons-from-building-pipelinesage-on-hindsight-2jmc", "published_at": "2026-09-29 07:30:58+00:00", "updated_at": "2026-09-29 07:46:40.610463+00:00", "lang": "en", "topics": ["ai-agents", "mlops", "ai-tools", "developer-tools", "large-language-models"], "entities": ["PipelineSage", "Hindsight", "Vectorize", "Groq", "openai/gpt-oss-120b", "Streamlit", "Python"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/teaching-a-ci-cd-failure-agent-to-remember-lessons-from-building-pipelinesage-on", "markdown": "https://wpnews.pro/news/teaching-a-ci-cd-failure-agent-to-remember-lessons-from-building-pipelinesage-on.md", "text": "https://wpnews.pro/news/teaching-a-ci-cd-failure-agent-to-remember-lessons-from-building-pipelinesage-on.txt", "jsonld": "https://wpnews.pro/news/teaching-a-ci-cd-failure-agent-to-remember-lessons-from-building-pipelinesage-on.jsonld"}}