Omnichannel AI agents: sharing long-term memory between a voice and a chat agent with Amazon Bedrock AgentCore Memory, Strands and Amplify Gen 2 A developer built an omnichannel shopping assistant that shares long-term memory between a text chat agent and a real-time voice agent using Amazon Bedrock AgentCore Memory, Strands, and Amplify Gen 2. The system uses AgentCore Memory extraction strategies to distill raw conversation events into persistent preference and fact records, so a preference learned through one channel carries over to the other. The voice agent runs on Amazon Nova Sonic via Strands' BidiAgent, while the chat agent uses Amplify AI Kit with DynamoDB-backed history. In the first article https://dev.to/aws-builders/your-database-is-an-ai-tool-semantic-search-with-amazon-dynamodb-vector-search-46ff I built semantic product search on Amazon DynamoDB Vector Search and gave that capability to an AI agent as a tool. In the second one https://dev.to/aws-builders/deploying-a-real-time-voice-agent-with-agentcore-runtime-and-amplify-gen-2-45bl I deployed the voice agent to Amazon Bedrock AgentCore Runtime , inside the same Amplify Gen 2 backend. So now I have two agents that do the same job, help a user shop, through two different channels: a text chat Amplify AI Kit and a voice agent Strands BidiAgent using Amazon Nova Sonic . They work, but they are two strangers: tell the voice agent you are into ultralight camping gear, then open the chat and ask for a recommendation: it has no idea who you are. Each conversation starts from zero, and this article is about fixing that: giving both agents a shared memory so a preference learned in one channel shows up in the other. That is what turns "a few agents" into an omnichannel experience. I'll use Amazon Bedrock AgentCore Memory , and the key idea is deciding what the memory is keyed to. Let me walk through it. Companion posts: Your database is an AI tool: semantic search with Amazon DynamoDB Vector Search https://dev.to/aws-builders/your-database-is-an-ai-tool-semantic-search-with-amazon-dynamodb-vector-search-46ff Deploying a real-time voice agent with AgentCore Runtime and Amplify Gen 2 https://dev.to/aws-builders/deploying-a-real-time-voice-agent-with-agentcore-runtime-and-amplify-gen-2-45bl - Omnichannel agents: sharing memory across a voice and a text agent with Amazon Bedrock AgentCore Memory see blog/blog-3.md https://github.com/davide-desio-eleva/dynamodbvector/./blog/blog-3.md A sample application that shows how to use Amazon DynamoDB native vector search to build semantic search over application data, how to expose that capability to AI agents as a tool, how to deploy a real-time voice agent for it on Amazon Bedrock AgentCore Runtime , and how to give a voice agent and a text agent a shared memory so they behave as one omnichannel assistant — all inside a single AWS Amplify Gen 2 backend. It demonstrates the same… Before wiring anything, it helps to separate two things that both get called "memory". Short-term memory is the current conversation. The turns you and the agent just exchanged, so it can follow "make it cheaper" without asking cheaper than what. It lives and dies with the session. Long-term memory is what survives across sessions. Not the raw transcript, but distilled knowledge: "this customer likes ultralight gear", "their budget is around 150 euros", "they camp in winter". This is the part that makes an omnichannel experience possible, because it outlives any single conversation and any single channel. Amazon Bedrock AgentCore Memory gives me both. I write raw events short-term , and it runs extraction strategies in the background that distill those events into long-term records. I get to pick which strategies run: prefers ultralight gear , budget around 150 euros . bought a DayHike 25L Pack , camps in winter . There is also a Summarization strategy, but for a shopping assistant the preferences and facts are what matter, so I'll use those two. Here's a nice consequence of the stack I'm already on: short-term memory is basically handled for me on both channels, so the part I actually need to add is the long-term, cross-channel one. On the chat side, the Amplify AI Kit already persists the conversation to Amazon DynamoDB and replays it on every turn. Following "make it cheaper" within a conversation just works, the AI Kit stores and reloads the message history automatically, no AgentCore short-term events required. On the voice side, the BidiAgent keeps the live session context inside the open bidirectional stream with Nova Sonic. Within a single voice session the model already has everything it just heard, so per-session short-term memory isn't something the agent needs me to add either. So the gap that AgentCore Memory fills here is specifically the long-term, cross-session, cross-channel one: the distilled preferences and facts that must outlive any single conversation and travel between the two agents. That's the piece neither the AI Kit nor the BidiAgent gives me on its own, and it's what the rest of this article wires up. Here is the insight that makes or breaks the whole thing. AgentCore Memory organizes records under an actorId and a sessionId . The natural temptation is to let each agent use its own runtime session as the identity. If you do that, the voice agent remembers voice sessions and the chat agent remembers chat sessions, and they never meet. You would have two separate memories that happen to use the same service. For omnichannel, the memory has to be keyed to the user , not to the runtime session or the channel. My app already has a stable per-user identifier: the Amazon Cognito sub . The same user signs into the chat and the voice agent, so if both agents use the Cognito sub as the actorId , they read and write the same records. A preference the voice agent stored under sub=a2751... is exactly what the chat agent retrieves under sub=a2751... . So the design is one memory store, two agents, keyed by the Cognito sub : Because Amplify Gen 2 is CDK under the hood, the memory store is just another construct in backend.ts , next to the data, auth, and the voice runtime from the previous article. I use the L1 CfnMemory : for a service this new I want what I write to map one-to-one onto the CloudFormation resource, with no abstraction deciding things for me. js import { CfnMemory } from "aws-cdk-lib/aws-bedrockagentcore"; const agentMemory = new CfnMemory voiceStack, "ShoppingAgentMemory", { name: "shoppingAgentMemory", // Raw short-term events are kept for 30 days before expiring. eventExpiryDuration: 30, memoryExecutionRoleArn: memoryExecutionRole.roleArn, memoryStrategies: { userPreferenceMemoryStrategy: { name: "PreferenceLearner", namespaces: "/preferences/{actorId}/" , }, }, { semanticMemoryStrategy: { name: "FactExtractor", namespaces: "/facts/{actorId}/" , }, }, , } ; const memoryId = agentMemory.attrMemoryId; Two things worth calling out. The namespaces use a {actorId} template. AgentCore substitutes the real actorId at write and read time, so /preferences/{actorId}/ becomes /preferences/a2751.../ for that user. This is what physically separates one user's memories from another's, using the same key both agents share. The memoryExecutionRoleArn matters because long-term extraction runs Amazon Bedrock models on your behalf . The built-in strategies read your raw events and call a model to distill them, so the memory needs a role allowed to invoke Bedrock: js const memoryExecutionRole = new iam.Role voiceStack, "AgentMemoryRole", { assumedBy: new iam.ServicePrincipal "bedrock-agentcore.amazonaws.com", { conditions: { StringEquals: { "aws:SourceAccount": account } }, } , } ; memoryExecutionRole.addToPolicy new iam.PolicyStatement { actions: "bedrock:InvokeModel" , resources: "arn:aws:bedrock: ::foundation-model/ " , } ; Then both the voice runtime role and the chat handler role get read/write access to the memory CreateEvent , RetrieveMemoryRecords , ListMemoryRecords , and friends on agentMemory.attrMemoryArn , and both get MEMORY ID as an environment variable. Same store, same permissions, two consumers. The voice agent is a Strands BidiAgent . The first job is to make sure it keys memory to the Cognito sub , not to the runtime session. The frontend already authenticates the WebSocket to AgentCore with the user's Cognito token that was the whole point of the JWT authorizer in the previous article . The token is a JWT, and the sub is right there inside it. So I resolve the actorId from the connection: php def resolve actor id websocket: WebSocket - str: """The memory actorId is the Cognito sub , shared with the chat agent.""" headers = websocket.headers auth = headers.get "authorization" if auth: token = auth 7: if auth.lower .startswith "bearer " else auth sub = decode jwt sub token base64url-decode the JWT payload, read sub if sub: return sub custom = headers.get "x-amzn-bedrock-agentcore-runtime-custom-actorid" if custom: return custom return "anonymous" Now, a browser can't set arbitrary headers on a WebSocket handshake, and AgentCore only forwards headers to your container if they are on an allowlist . So I let the frontend pass the sub as a custom runtime header via a query parameter, and I allowlist it on the runtime: // backend.ts — on the CfnRuntime requestHeaderConfiguration: { requestHeaderAllowlist: "X-Amzn-Bedrock-AgentCore-Runtime-Custom-actorId" , }, js // frontend — the Cognito sub, passed as a custom runtime header const actorId = session.tokens?.idToken?.payload?.sub; url += &X-Amzn-Bedrock-AgentCore-Runtime-Custom-actorId=${encodeURIComponent actorId } ; Values sent as X-Amzn-Bedrock-AgentCore-Runtime-Custom- are delivered to the container as headers of the same name, and resolve actor id reads it. Now the voice agent and the chat agent agree on who the user is. For persistence, Strands and bedrock-agentcore offer a native integration: a session manager that transparently writes every turn to AgentCore Memory. I hand it the memory id, the session id, and, crucially, the shared actorId : python from bedrock agentcore.memory.integrations.strands.config import AgentCoreMemoryConfig from bedrock agentcore.memory.integrations.strands.session manager import AgentCoreMemorySessionManager, memory config = AgentCoreMemoryConfig memory id=MEMORY ID, session id=session id, unique per conversation actor id=actor id, the Cognito sub — shared across channels session manager = AgentCoreMemorySessionManager agentcore memory config=memory config, region name=MEMORY REGION, voice agent = BidiAgent model=sonic model, tools= search products, stop conversation , system prompt=build system prompt actor id , more on this in a second session manager=session manager, With the session manager attached, every turn of the conversation gets written to the memory store, and the background strategies distill preferences and facts from those turns. Writing is fully handled for me. Reading back is where it gets interesting, and where the two agents end up looking different. The native session manager's automatic retrieval applies to the standard Agent , not to the streaming BidiAgent that Nova Sonic uses. For a real-time voice agent, retrieval is not wired into the loop for you. So I retrieve the long-term records myself, at the start of the session, and inject them into the system prompt: php def retrieve memories actor id: str - list str : """Fetch this user's long-term preferences and facts, keyed by Cognito sub.""" namespaces = f"/preferences/{actor id}/", f"/facts/{actor id}/" context = for namespace in namespaces: records = memory client.retrieve memories memory id=MEMORY ID, namespace path=namespace, query="user preferences, interests and facts", top k=5, for record in records: text = record.get "content", {} .get "text", "" .strip if text: context.append text return context def build system prompt actor id: str - str: context = retrieve memories actor id if not context: return SYSTEM PROMPT remembered = "\n".join f"- {item}" for item in context return f"{SYSTEM PROMPT}\n\n" "Here is what you remember about this customer from previous " "conversations, across both voice and chat. Use it to personalize your " "suggestions, and confirm before assuming it still applies:\n" f"{remembered}" So on the voice side: the session manager writes, and I read. The write is native, the read is manual. The chat agent runs on the Amplify AI Kit, through a custom conversation handler. There is no magic session manager here either, so the pattern is symmetric with the voice agent's read path: I do the retrieve-and-inject myself, plus I persist the turn. The AI Kit passes the user's Cognito token on the conversation event headers, so I get the same sub the voice agent uses: function resolveActorId event: ConversationTurnEvent : string | undefined { const auth = event.request.headers "authorization" ; return decodeJwtSub auth ; // same base64url-decode → sub } Then the handler wraps the default AI Kit handler. Before the model runs, it retrieves the same namespaces and prepends what it finds to the system prompt. After, it writes the user's turn so the strategies can extract from it: js export const handler = async event: ConversationTurnEvent = { const actorId = resolveActorId event ; if memoryClient && MEMORY ID && actorId { const userText = await getLatestUserText event ; const preferences, facts = await Promise.all retrieveMemory actorId, "/preferences", userText , retrieveMemory actorId, "/facts", userText , ; const preamble = buildMemoryPreamble preferences, facts ; if preamble { event.modelConfiguration.systemPrompt = ${preamble}\n\n${event.modelConfiguration.systemPrompt} ; } if userText { await persistUserTurn actorId, event.conversationId, userText ; } } return handleConversationTurnEvent event ; }; Same store, same actorId , same namespaces. The only difference from the voice agent is that here I also write manually persistUserTurn calls CreateEvent , because there is no session manager doing it for me. This is the part I find genuinely interesting. The two agents talk to the same memory but integrate with it differently, and that is not a mistake, it's the reality of working across two runtimes: | | Voice agent Strands BidiAgent | Chat agent Amplify AI Kit | |---|---|---| | Write | Native session manager | Manual CreateEvent | | Read | Manual retrieve + inject into system prompt | Manual retrieve + inject into system prompt | | Identity | Cognito sub from JWT / custom header | Cognito sub from JWT | The takeaway: omnichannel memory is not about a single SDK that does everything for you. It's about agreeing on the key the user identity and the namespaces. Once both agents agree that memory is keyed to the Cognito sub and lives under /preferences/{actorId}/ and /facts/{actorId}/ , the plumbing on each side can differ. The memory is the contract; the integration is per-runtime. There was another perfectly valid way to do this, and it's worth naming. Instead of wiring each agent to AgentCore Memory through its own runtime integration, I could have built a single "memory" tool , a small function that reads and writes AgentCore Memory, and handed that same tool to every agent, exactly like searchProducts is shared today. Every agent would then remember and recall by calling the tool, the integration would be identical everywhere, and a third or fourth channel would just get the same tool. That approach is clean, uniform, and it's probably what I'd reach for if I had five channels instead of two. I chose the other path on purpose: I wanted to explore the native integration options each runtime offers, the Strands session manager on the voice side, and the Amplify AI Kit conversation handler on the chat side, and see how memory fits into each one's grain rather than bolting a uniform tool on top. That's also what surfaced the interesting asymmetry above native write, manual read for BidiAgent , which the shared-tool approach would have hidden. But the difference between the two isn't just uniformity, it's who decides when memory is used , and that's the part I find most important. With a memory tool , recall is agentic : the memory is one more tool in the agent's belt, and the LLM decides, turn by turn, whether to call it. That's flexible the agent can choose to look something up only when it seems relevant but it's also non-deterministic. The model might not call the tool when you'd want it to, so the user says "give me options" and the agent, having decided it didn't need memory this turn, answers as if it knows nothing about them. You're trusting the model's judgment about when to remember. With the native integration I used here, recall is deterministic . I retrieve the user's preferences and inject them into the system prompt at the start of every conversation, unconditionally. The model doesn't get a vote on whether to be aware of them; the context is simply always there. For a shopping assistant that should feel like it knows the returning customer, "always aware" is the behavior I want, not "aware if the model felt like calling a tool". So the trade-off is: a memory tool gives the LLM control and flexibility over recall; native injection gives you control and guarantees the context is present. Neither is universally right. Agentic recall shines when memory is large and lookups should be selective; deterministic injection shines when a small, high-value profile should shape every single response. So read this article as one of two good options. If you want maximum uniformity across many agents and you're comfortable letting the model decide when to recall, a shared memory tool is a great choice. If you want the context guaranteed on every turn and you want to understand how memory plugs into Strands and Amplify Gen 2 natively, this is that exploration. Either way, the design principle that matters, keying memory to the user, is the same. The test that matters is the bidirectional one. Chat, then voice. In the text chat I say I'm shopping for camping and I pick a DayHike 25L Pack. A minute later long-term extraction is asynchronous, it takes a moment , I open the voice agent and ask, in Italian, what it recommends for me. It brings up camping and the pack, without me repeating anything. It read what the chat agent wrote. Here is me asking via chat articles for un upcoming hiking in October in Iceland: the agent suggested me some useful ones. After that I've made a call to the voice agent, asking more information about those article. I've never mentioned Iceland again, thus confirming it got this information from the memory also I have logs . Voice, then chat. The reverse works the same way. A preference spoken to the voice agent surfaces in the next chat turn. One thing to keep in mind when you try this: long-term memory is extracted asynchronously . Right after a turn, the raw event exists but the distilled preference might not yet, so a retrieve one second later can come back empty. Give the extraction a moment. That is the nature of long-term memory: it's the slow, considered kind, not the immediate transcript. AgentCore Memory is serverless and consumption-based , there is no fixed monthly fee just for having a memory store. You pay on three axes: short-term events written, long-term records stored, and retrieval calls. For a demo like this it rounds to cents. The nice part is that the rest of the stack is the same kind of thing. The AgentCore Runtime in the serverless microVM mode we use bills CPU and memory only while a session is running, I/O wait is free, so with no one talking to it there is effectively no idle compute charge. The ECR image is just storage, a few cents per month. Left alone, this whole stack costs almost nothing; you pay when someone actually uses it. Here is where the design pays off. Once memory is keyed to the user and not to the channel, adding a third channel is mostly plumbing. The memory doesn't change at all. Imagine a WhatsApp channel using AWS End User Messaging Social . The shape would be: Amazon SNS to a Lambda. searchProducts tool the other two channels use. sub . So you need a mapping from phone number to your app's user identity, for example a small DynamoDB table populated during an opt-in or account-linking step. Once you resolve the phone number to the Cognito actorId . /preferences/{actorId}/ and /facts/{actorId}/ , inject them into the prompt, generate a reply, send it back through The user starts on WhatsApp on the train, continues by voice at home, finishes in the web chat, and the assistant remembers throughout. No channel owns the memory. The user owns the memory, and every channel is just a different door into the same context. That mapping step phone number to user identity is the only real new work. Everything else, the memory store, the namespaces, the retrieve-and-inject pattern, is already built. That is the point of keying memory to the user: new channels are additive, not a rewrite. Three things stood out building this. Identity is the design. The single most important decision wasn't which memory strategy to use or how to call the SDK. It was keying memory to the Cognito sub instead of the runtime session. Get that right and omnichannel falls out almost for free. Get it wrong and you have two agents with amnesia and no amount of SDK cleverness fixes it. One memory, many integrations. The voice agent and the chat agent integrate with AgentCore Memory differently, one uses a native session manager to write, the other writes manually, and both retrieve and inject by hand. That asymmetry is fine. The memory is the shared contract; how each runtime reads and writes it is a local detail. It's all one backend, still. The memory store, its IAM, its wiring into both the voice runtime and the chat handler, are all just CDK constructs sitting next to the data and auth. Adding cross-channel memory didn't mean a new system to operate. It meant a few more constructs in the same npx ampx sandbox deploy. Your Amazon DynamoDB database was an AI tool. The agent that talks to it became serverless. And now, whichever channel you reach for, it's the same assistant, and it remembers you. I'm D. De Sio https://www.linkedin.com/in/desiodavide and I work as a Head of Software Engineering in Eleva https://eleva.it/ . As of September 2026, I’m an AWS Certified Solution Architect Professional https://www.credly.com/badges/9929fdf2-7a3d-4013-9de6-57c80e4920b9/public url and AWS Certified DevOps Engineer Professional https://www.credly.com/badges/8c5a1487-191b-429e-8c2d-7cee43bf316b/public url , but also a User Group Leader in Pavia https://www.linkedin.com/company/aws-user-group-pavia/ , an AWS Community Builder and, last but not least, a serverless enthusiast. I just shared long-term memory with my AI agents. I'm sure they'll remember I'm the good guy when Skynet goes live. The full agenda for AWS Community Day Italy https://www.awscommunityday.it/ is out If you'd love to hear what the community has been working on, what they've learned, and what they want to share, come join us in Rome on October 2nd.