How to implement long-term AI agent memory in AlloyDB and Memorystore for Valkey A two-tier memory architecture combining Memorystore for Valkey for short-term session buffers and AlloyDB AI for long-term persistent memory can cut token spend by up to 70% for enterprise AI agents running multi-day workflows, according to Google Cloud. The approach replaces context stuffing, which can push response times past 30 seconds and triggers "lost in the middle" degradation, and lossy LLM rolling summaries that drop earlier user constraints. Memorystore for Valkey handles sub-millisecond session-state lookups while AlloyDB AI stores facts, preferences and episodic data across sessions with transactional integrity and hybrid relational-vector retrieval. Enterprise AI agents need persistent memory to execute complex, multi-day workflows and long-horizon tasks. In this blog, we examine how a 2-tier memory architecture using Memorystore for Valkey https://docs.cloud.google.com/memorystore/docs/valkey for short-term buffer memory and AlloyDB AI https://cloud.google.com/alloydb/ai for long-term persistent memory can help reduce token spend by up to 70%, while maintaining critical data and enterprise guardrails. Imagine building a personalized travel agent designed to help users book vacations. The user interacts with the agent many times over the course of several days, asking questions that range from brainstorming itineraries to actual purchase intent. To provide a truly seamless experience, this agent must remember flight preferences e.g. “I only want non-stop flights” , hotel budgets, and dietary restrictions e.g. “I need Gluten Free dining options” established in previous sessions. More importantly, it has to hold onto these core facts even when the conversation gets deep into the weeds of sightseeing recommendations and itinerary planning. The agent must ensure that all the follow-up questions and exciting details about places to visit doesn’t cause it to forget or overwrite the user's fundamental requirements and decisions. AI agents are being used for multi-turn conversations and long-running workflows, but large language models LLMs remain stateless across sessions. When a user returns to an agent days later, the model starts with an empty context window. Unless your application is built to reconstruct past context using long-term agent memory, your users have to explain their goals and context all over again, resulting in a frustrating and fragmented experience. With million-token context windows now the norm, a common shortcut to this problem is "context stuffing" - dumping everything you can fit into the prompt at every turn, from raw chat histories to tool execution logs. In fact, this was a common pattern in the early days of AI model usage. But this shortcut quickly creates issues at scale: token costs multiply with every message, response times can drag out past 30+ seconds for otherwise simple prompts, and the model starts suffering from " lost in the middle https://arxiv.org/abs/2307.03172 " degradation, overlooking critical instructions buried in mountains of prompt text. Another common workaround is to use rolling summaries; asking an LLM to periodically compress older messages into a summary paragraph. While this trims prompt size, LLM-based summarization is inherently lossy. After a few rounds of compression, subtle but important details get filtered out as background noise. A few turns later, your agent quietly breaks the exact constraints you set earlier. Clearly, you need a more scalable approach for keeping a memory from previous conversations or multi-turn tasks, but without degrading the experience or creating new bottlenecks. To build reliable, cost-effective enterprise agents that respect your guardrails and constraints, we recommend a 2-tier memory architecture : Short-term session buffer : Caches active conversation turns in memory using a token-bounded sliding window, allowing the context window size to remain stable across multiple rounds, even across devices, while keeping the latest messages fresh. This tier requires sub-millisecond, high-throughput lookups on every turn, making Memorystore for Valkey https://docs.cloud.google.com/memorystore/docs/valkey well suited for maintaining active session state. Long-term persistent memory : Stores important facts, user preferences, and episodic facts across sessions. This tier requires transactional integrity, data governance, and hybrid retrieval across relational data and vectors - capabilities provided natively by AlloyDB AI https://cloud.google.com/alloydb/ai . While it’s possible to store both long-term memory and active session buffers directly in a relational database, a two-tier architecture with an in-memory cache provides better performance and scalability. Active session buffers are highly ephemeral and require sub-millisecond, high-throughput updates on every single conversation turn. Handling these rapid-fire writes in an in-memory key-value cache prevents write amplification and table bloat in your relational database, which would otherwise require frequent row deletions and intensive vacuuming. This is analogous to adding a caching tier in front of your database to offload high-frequency lookups for hot rows. The division of labor keeps your primary database lean and responsive, allowing it to focus its resources on what it is designed for: transactional consistency, complex hybrid vector search, and long-term analytical query execution. By pairing short-term caching in Memorystore for Valkey with native AlloyDB AI capabilities, you can run entity extraction, memory compaction, cross-session memory, and hybrid retrieval directly inside the database tier, keeping active prompts lean, fast, and cost-efficient. To organize long-term agent state effectively, we further divide agent memory into four complementary types across the short-term and long-term storage tiers we described above: | Memory type | What it stores | Storage layer | Lifespan | |---|---|---|---| | Buffer Short-term | Recent raw conversation turns | Memorystore for Valkey | Active session | | Summary memory | Compressed history of older turns, commonly referred to as “compaction” | Memorystore for Valkey | Multi-turn window | | Episodic memory | Past actions, events, and tool outputs | AlloyDB for PostgreSQL Hybrid retrieval with structured SQL + full-text search + vector | Permanent | | Entity & rule memory | User preferences, constraints, and vetoes | AlloyDB for PostgreSQL Hybrid retrieval with structured SQL + full-text search + vector | Permanent | By isolating short-term conversation context from structured long-term rules, your agent retrieves relevant context on demand without filling token windows with raw interaction logs. The diagram below outlines the read and write paths connecting the application orchestration layer, the Memorystore for Valkey short-term buffer, and the AlloyDB AI long-term repository: The system operates across two coordinated execution paths: The read path : When a user asks a question, the application fetches the active sliding window from Memorystore for Valkey , runs in-database query normalization using ai.generate , and queries AlloyDB using ai.hybrid search to retrieve scoped entity rules and relevant episodic facts. The write path : After generating the response, the turn is immediately cached in Memorystore for Valkey . An asynchronous background queue worker extracts structured entities from the exchange and writes them directly into AlloyDB , where transactional auto-embeddings immediately compute and store vector representations in the database. In our benchmark testing across multi-turn development dialogues 45+ turns with heavy tool executions , separating short-term caching from long-term persistence delivered measurable cost and performance improvements compared to naive context stuffing: | Metric / dimension | Naive context stuffing | Tiered memory AlloyDB + Valkey | Net business impact | |---|---|---|---| | Active prompt size turn 45 | 747,033 tokens | 83,262 tokens | 88.9% smaller prompt | | Turn 45 response latency | 33.5 seconds | 6.7 seconds | 80.0% faster response | | Per-turn response wait time | 33.5 seconds | 4.2s – 6.7s | 36% to 80% reduction | | Cumulative session tokens | 17.9M tokens | 4.09M tokens | 72.0% token & cost savings | | Rule & constraint recall | Degrades over turns | Does not degrade | ACID-preserved recall | Note: The metrics above reflect internal benchmark results from a simulated developer workload that mimics a real-world enterprise AI pair-programming assistant interacting with a developer over multiple sessions, projects, and context switches. Actual savings and latencies vary based on prompt structure, query frequency, and data volume. These results demonstrate how tiered memory changes agent unit economics: instead of an escalating cost curve on every additional turn, prompt sizes remain bounded, reducing ongoing LLM API expenses while keeping response times fast. AlloyDB AI reduces the operational overhead of implementing persistent agent memory by embedding core AI functions directly into the database engine: Transactional in-database auto-embeddings ai.initialize embeddings : AlloyDB automatically generates vector embeddings for text columns using a native integration with Agent Platform https://cloud.google.com/products/gemini-enterprise-agent-platform?e=48754805 formerly Vertex AI , generating up to 3,000 embeddings per second. Using incremental refresh mode = 'transactional' , AlloyDB keeps embeddings up to date as source data changes within the same transaction, removing the need for custom embedding pipelines, external schedulers, or complex retry logic. In-database generative AI functions ai.generate : AlloyDB allows you to execute foundation models, such as Gemini, directly from SQL queries. You can use this for in-database query decomposition - breaking compound user questions into single-aspect sub-queries and resolving relative time phrases like "last session" into explicit identifiers - without making separate roundtrips from your application. Native hybrid search with built-in Reciprocal Rank Fusion ai.hybrid search : AlloyDB provides a built-in SQL function that executes Reciprocal Rank Fusion RRF directly inside the engine. It combines vector cosine similarity <= over HNSW or ScaNN indexes with PostgreSQL full-text search using BM25, RUM, or GIN in a single database call, blending semantic matching with exact keyword retrieval while supporting metadata filter pushdown filter condition for improved performance and deterministic scope isolation. Direct Agent Platform integration with IAM credentials : AlloyDB connects directly to Agent Platform foundation models over Google Cloud's private network using Google Cloud IAM https://cloud.google.com/products/iam service account roles and database authentication, avoiding the need to store, rotate, or pass API keys in application code. Unified operational, vector, and governance engine : AlloyDB consolidates relational business data, vector embeddings, full-text indexes, and enterprise permissions in a single ACID-compliant PostgreSQL database, avoiding data drift and integration complexity across separate operational and vector databases. Below are the core database patterns used to configure the 2-tier memory architecture. For complete, runnable Python and SQL scripts, refer to the companion AlloyDB Agent Memory Codelab https://codelabs.developers.google.com/alloydb-agentic-tiered-memory . In AlloyDB , install the necessary extensions and define the agent entities table with structured metadata, a generated tsvector column for full-text search, and a vector embedding column. Then, generate embeddings using ai.initialize embeddings . Using transactional mode ensures the embeddings are kept up to date as source data changes. Finally, create the HNSW vector index and the RUM full-text search index to ensure your hybrid searches are fast and efficient. On the read path, retrieve relevant long-term entities using ai.hybrid search . This native SQL function executes Reciprocal Rank Fusion RRF directly in AlloyDB, seamlessly reranking and combining vector similarity search and full-text keyword search results in a single database query: To manage long-term storage growth without writing custom pruning scripts, you can run automated extraction and compaction queries directly in AlloyDB to identify the important parts which can benefit from being stored. This pattern uses a SQL common table expression CTE with ai.generate to consolidate older episodic entries into a high-density summary: For example, a long interaction regarding the complexities of changing flights with kids can result in a summary of “I prefer non-stop flights”. With the core 2-tier memory architecture configured, you can now extend the default ADK Memory provider to use this 2-tier memory architecture as shown in the accompanying Codelab https://codelabs.developers.google.com/alloydb-agentic-tiered-memory 14 e.g. ADKTieredMemoryProvider . To attach the long-term memory to an ADK agent, you simply provide it as a tool e.g. longterm memory tool like this: Running agent memory in enterprise production environments requires strict security boundaries and access controls. First of all, we need to maintain scope and multi-tenant isolation. By indexing user id , project id , and scope columns in agent entities and enforcing PostgreSQL Row-Level Security RLS , you can isolate memory stores across departments, teams, and individual users within the same database cluster. Parameterized Secure Views https://docs.cloud.google.com/alloydb/docs/parameterized-secure-views-overview PSV offer another layer of deterministic application-level security, helping you protect against malicious prompts and overly-broad SQL queries. In addition, we need to provide automated memory lifecycle management as shown in the implementation pattern above. Combining scheduled SQL compaction queries with time-based partition pruning helps maintain predictable database footprint and query latencies over time. See the accompanying Codelab https://codelabs.developers.google.com/alloydb-agentic-tiered-memory for more details on this approach. Decoupling active context windows from persistent storage is a practical approach to building production-ready AI agents. By pairing Memorystore for Valkey for sub-millisecond session caching with AlloyDB AI for transactional long-term storage, you can achieve substantial token cost savings and faster response times while maintaining strict business rules throughout long-horizon tasks and many-turn agentic experiences. To get started: Step through the complete hands-on tutorial in the companion AlloyDB Agent Memory Codelab https://codelabs.developers.google.com/alloydb-agentic-tiered-memory to deploy the working 2-tier memory architecture. Learn more about database-side machine learning features in the AlloyDB AI documentation https://cloud.google.com/alloydb/docs/ai/overview . Explore guides on generating auto vector embeddings https://docs.cloud.google.com/alloydb/docs/ai/generate-manage-auto-embeddings-for-tables and running hybrid vector search https://docs.cloud.google.com/alloydb/docs/ai/run-hybrid-vector-similarity-search .