I Gave My Proposal Agent Hindsight of Every Lost Bid A developer built Aura Memory, an RFP proposal agent that records the outcome of every bid — won, lost or pending — and feeds those lessons back into future drafts using Hindsight, Vectorize's open-source agent memory system. The write-side design keys each memory to its source proposal via document_id so re-recorded outcomes upsert rather than duplicate, stores metadata as JSON-encoded strings, and writes memory content as prose for LLM extraction. The stated goal is attribution: reviewers can see which past deal drove a recommendation, such as a lost bid on lump-sum pricing, rather than generic 'best practices' claims. Aura Memory takes an RFP, works out what the client is asking for, writes an 11-section proposal, and then waits. When the deal closes, someone records the outcome: won, lost or pending, the factors that mattered, and a few lines from the debrief. The system turns that into a structured lesson and stores it. The next time a similar RFP arrives, those lessons come back and shape the new draft. The loop looks like this: RFP → analyze → recall past outcomes → reflect → generate → submit → record outcome → extract lessons → retain → next RFP recalls it The stack is deliberately boring. It's a TanStack Start app with server routes, one LLM behind an OpenAI-compatible client, and Redis for the application's own records. The memory layer is Hindsight, an open-source agent memory system from Vectorize. It does three things for me: retain stores an experience, recall retrieves relevant ones, and reflect reasons over them and returns an answer. I didn't want a vector store I'd have to babysit. A proposal outcome isn't a document you chunk and embed; it's an event with a result, a cause and a client attached. That's the gap between retrieval and agent memory as a distinct layer from RAG, and it's why I picked a memory system rather than building similarity search myself. The real problem: attribution Getting memories in was easy. The hard part, and the through-line of everything below, was attribution. When the agent says "lead with security architecture," I need to know which past deal taught it that. A reviewer who sees "because we lost Pacific National Bank on lump-sum pricing" trusts the draft. A reviewer who sees "best practices suggest…" doesn't. Hindsight doesn't store my records verbatim. On retain, it runs an LLM over the content and extracts facts, entities and relationships, which is what makes recall good. But it also means a recall result is a list of facts, not the proposal I stored. If I wanted citations, I had to design for them from the first write. Write-side: make every memory traceable Here's the retain call. Nearly every field is there for attribution: await hindsightFetch config, "/memories", { method: "POST", body: { async: false, items: { content: memoryToText memory , context: ${memory.industry} proposal outcome ${memory.outcome} , // One document per proposal: re-recording an outcome replaces the earlier memory. document id: memory.sourceProposalId ? proposal-${memory.sourceProposalId} : memory.id, timestamp: memory.createdAt, tags: "proposal-outcome", industry:${tagSlug memory.industry } , outcome:${memory.outcome} , metadata: { memory id: memory.id, title: memory.title, outcome: memory.outcome, lessons: JSON.stringify memory.lessons , // ...patterns, client preferences, recommendations }, }, , }, } ; Three decisions here made everything downstream easier: document id is derived from the proposal. People change their minds about outcomes: a "pending" becomes "won", or a debrief arrives a week late. Keying the document on the proposal makes re-recording an upsert, not a duplicate. The first version generated a fresh ID on every recording, so a deal recorded twice would have been recalled twice and cited twice. Metadata values are strings. Hindsight metadata is a string-to-string map, so arrays are JSON-encoded. It's a little ugly, but recall hands the metadata back on every extracted fact. That's what lets me rebuild the full record without a second lookup. content is prose, not JSON. memoryToText writes the outcome as sentences: "Outcome: LOST. What failed: Generic Pricing Language…" Hindsight's extraction works on natural language, so I give it sentences a person could read, not a serialized object. async: false matters too. In the live flow, someone records an outcome and often starts the next RFP right away. I'd rather they wait three or four seconds on retain than have the next recall miss the lesson they just typed. Read-side: turn facts back into deals A recall for a telehealth RFP can return a dozen facts from four different past proposals. The UI needs "St. Jude Health System, WON, 77% relevant," not twelve loose sentences. So recall groups facts by the memory id I planted: const data = await hindsightFetch<{ results?: HindsightFact } config, "/memories/recall", { method: "POST", body: { query: recallQueryText query , budget: "mid", max tokens: 4096, // Observations are Hindsight's cross-document summaries; they carry no document id or // metadata, so they can't be attributed to a specific past proposal. types: "world", "experience" , }, } ; // Hindsight returns extracted facts; group them back into the proposal experiences they came from. const groups = new Map ; for const fact of data.results ?? { const key = fact.metadata?. "memory id" || fact.document id || fact.id; groups.set key, ... groups.get key ?? , fact ; } The types filter came out of the first real integration test. Hindsight also produces observations: consolidated summaries it builds across memories, like "Successful proposal strategies for healthcare included…". They're genuinely useful, and reflect leans on them heavily. But they have no document ID or metadata, so they can't be tied to a specific deal. In recall they showed up as anonymous "Recalled experience" cards. For the "which deals are we learning from" view I only want facts I can trace to a source, so recall asks for world and experience facts. Observations still do their job inside reflect. Relevance comes from Hindsight's per-fact scores object. I use the semantic similarity, because the reranker's final score isn't calibrated for showing as a percentage. The groups are sorted by their best fact. Reasoning with receipts Recall tells me which deals are relevant. Reflect tells me what to do about them. I ask for structured output, so the response maps straight onto the UI: const data = await hindsightFetch config, "/reflect", { method: "POST", body: { query, budget: "mid", max tokens: 2048, response schema: REFLECT SCHEMA, include: { facts: {} } }, } ; const out = data.structured output ?? {}; const refs = referencedIds data.based on ; const cited = recalled.filter m = refs.has m.id || refs.has proposal-${m.sourceProposalId} || m.factIds?.some id = refs.has id , ; The schema asks for a strategy, patterns to repeat, pitfalls to avoid, and reasoning: bullets that each have to name the past proposal and its outcome. The bank's reflect mission says the same thing in plain English "always cite the specific past proposal each recommendation comes from" , and the two together hold up well. The subtle part is based on. With include.facts, reflect reports the facts it actually used, identified by fact ID rather than document ID. So each recalled experience keeps the fact IDs it was built from, and the "cited" list is the intersection. That's the difference between "here are five deals we found" and "here are the three deals this strategy actually rests on." Lessons are only as good as their extraction Retained memory is only useful if what goes in is specific. A debrief note like "pricing bad" helps nobody six months later. So before retain, an extraction step turns the proposal, the RFP, the outcome and the reviewer's notes into structured lessons, recommendations and client preferences. Two rules in that prompt earned their place: What it looks like in practice Here's a concrete run. A Northbay Medical Group RFP arrives: telehealth, HIPAA, Epic integration, six clinics, "per-clinic pricing breakdown." Recall's top result is a won healthcare deal St. Jude Health System , followed by experiences from other industries. The reasoning panel reads like a colleague who was in the debriefs: "St. Jude Health System win demonstrated that front-loaded security architecture and named workstream ownership are decisive…" "Pacific National Bank loss proved that lump-sum pricing and weak ROI justifications without performance baselines fail to satisfy executive evaluators." That second bullet is my favorite behavior in the system: a banking loss shaping a healthcare proposal. The draft's pricing section comes back as an itemized per-clinic table. That's a direct response to a lesson learned in a different industry, which no amount of same-industry similarity search would have surfaced. A second example: after a SaaS observability migration was lost because "the migration plan was generic — no per-wave rollback," the next SaaS migration RFP recalled that loss first. Its strategy opened with "granular, service-by-service migration plans." That deal was won. To check that memory was changing the output, not just decorating it, I wrote an evaluation that generates the same RFP twice: once with recall and once with memory switched off. It then looks for lessons that exist only in memory, things the RFP never asks for, like tying ROI to a measured baseline, a clinician training plan, or a named security lead. It's a crude keyword check and I treat it as a smoke test, not a benchmark. But the gap between the two drafts has been consistent and obvious: the baseline reads like a competent generic proposal, and the memory-informed one reads like it was written by someone who lost the last three deals and paid attention. Lessons I'd reuse Design for attribution on the write path. You can't cite what you can't trace. Stable document IDs, a record ID in metadata, and outcome tags cost nothing at write time, and they're painful to retrofit. Store outcomes, not artifacts. The proposal text is the least valuable thing to remember. What it caused won, lost, and why is what makes the next proposal better. Hindsight's retain and reflect APIs fit that shape naturally, because they're built around experiences rather than chunks. Don't let the memory layer fail silently. Early on, a failed recall quietly fell back to keyword matching over a local list, and I didn't notice for a while that my integration was hitting the wrong endpoints. Now, if Hindsight can't be reached, the user sees an error. An agent that appears to remember when it doesn't is worse than one that admits it's starting from scratch. Gate the input before you remember anything. Someone once pasted a personal health question into the RFP box, and the pipeline happily produced a "headache-tracking platform" proposal backed by five recalled deals. Analysis now decides whether the text is a real business request first, and rejects it if not. Memory makes a model confident, so the question of what deserves that confidence has to come earlier. Test the delta, not the output. "Does the proposal look good?" is unanswerable. "Does it contain things only memory could have told it?" is a test you can run on every change.