As customers modernize to lakehouse architectures, they are standardizing on open formats such as Apache Iceberg to create a shared data estate across compatible engines. This enables you to build AI-native lakehouses that turn data to semantic knowledge, enable proactive action, and operate at an agentic scale.
One of the key components of a lakehouse is the catalog, and in the Apache Iceberg environment, that usually means the Iceberg REST Catalog. An Apache Iceberg catalog is responsible for maintaining table pointers, handling atomic commits, and serving as the single source of truth for table locations. But as large enterprise organizations modernize to lakehouses, they have begun to realize that they need a highly scalable and available managed catalog as a part of their lakehouse architecture. This becomes even more important as querying scales with agents. To support a large amount of repeated small queries from agents, you will need to build on top of a managed catalog that provides atomicity, consistency, availability and concurrency at massive scale.
In this blog, we explore the challenges a managed catalog faces in modern cloud environments at agent scale, and show you how Google Cloud’s serverless Lakehouse runtime catalog can help address them. Powered by Spanner, Google Cloud’s always-on database with virtually unlimited scale, and built to meet the open Apache Iceberg REST catalog specification, the Lakehouse runtime catalog is the highly scalable and available foundation you need for the agentic era.
When speaking with data engineers and infrastructure leads running production analytics at scale in a Lakehouse, the following core pain points consistently emerge and need to be solved by a managed catalog:
Atomic commits and concurrency control: Iceberg guarantees ACID transactions via optimistic concurrency control (OCC). A managed catalog must implement a bulletproof atomic compare-and-swap (CAS) operation to swap the current metadata pointer.
High availability and operational maintenance: Because queries fail immediately if the catalog is down, a managed catalog becomes a critical Tier-1 service.
Scaling the database backing the catalog: Catalog architects typically face a difficult trade-off when choosing a backing database for table metadata and state. Traditional scale-up relational databases provide SQL and ACID transactions, but hit vertical CPU, memory, storage and connection limits under heavy concurrent read/write loads unless manually sharded which incurs a huge operational overhead; while scale-out database systems are either eventually consistent, hard to manage, not enterprise-ready, or all of the above.
Table maintenance coordination: A catalog alone does not optimize data; you must build and operate ancillary pipelines for compaction, snapshot expiration, manifest rewriting, and orphan file cleanup.
Governance and security: A catalog acts as the security gatekeeper. The catalog must implement and maintain:
Authentication protocols (e.g., OAuth2 token exchange, IAM federation)
Access control down to namespace and table levels
Vended storage credentials (e.g., generating short-lived tokens so query engines don't need broad, direct storage credentials)
To solve these challenges, we built the Lakehouse runtime catalog (GA) with support for Iceberg Rest Catalog. The Lakehouse runtime catalog is a fully serverless, highly available, and unified metadata registry designed from the ground up to support modern open table formats like Apache Iceberg. By natively implementing the Apache Iceberg REST Catalog specification**,** the Lakehouse runtime catalog decouples metadata discovery from compute engines, helping ensure multiple Iceberg-compatible engines can access a shared data estate and enabling you to take your workloads to production sooner. We’ve helped many customers streamline the migration of their managed catalogs. For example, Etsy migrated its catalog to Lakehouse runtime catalog, joining data in place to accelerate pipeline queries by 60%.
This approach offers a number of architectural benefits:
Multi-engine interoperability: Once registered, tables are immediately discoverable and queryable across Google Cloud Managed Service for Apache Spark, BigQuery, and open-source engines via standard REST interfaces.
Read/write interoperability for Iceberg tables: Leverage Iceberg-compatible engines such as BigQuery, Managed Spark to write to Iceberg tables registered in the Lakehouse runtime catalog. Customers can also use Managed Spark to write to Iceberg tables in external catalogs.
Fully managed Iceberg storage with enterprise-grade features: Use Google's differentiated infrastructure to run analytics with performance on Iceberg tables. This gives you the benefits of open-source flexibility plus performance, scale, governance, and multimodal processing.
Zero data copy: Table definitions point directly to your existing data in the underlying object store. You do not move, rewrite, or duplicate your underlying data.
Bi-directional catalog federation across clouds: Access data from Databricks Unity, Snowflake Horizon and AWS Glue with support for vended credentials and OIDC token exchange. This lets you bring Google AI directly to your AWS and Azure data.
Secure access using credential vending: The catalog supports multiple authorization mechanisms, letting you choose between credential vending and end-user credentials. This means that you can access tables with modern mechanisms such as credential vending without needing direct access to the files in the underlying object store (Cloud Storage, AWS S3, Azure Blob Storage).
AI-powered context and governance: The Lakehouse runtime catalog integrates directly with Knowledge Catalog and Cloud IAM, allowing you to define trusted context for your agents, and apply table-level security consistently across all compute engines. Get out-of-the-box search, lineage, and insights for Iceberg tables in the catalog.
Atomic commits and concurrency control, high availability and scalability: Backed by Google’s planet-scale infrastructure and Spanner, you get the high availability, concurrency and scale you need for your metadata to scale with your data. Support for Cloud Storage dual-region and multi-region buckets enables failover use cases. It also provides reduced TCO due to serverless and no-ops environments, and scalability for any workload size.
The Lakehouse runtime catalog is a highly available, concurrent and scalable catalog with strong consistency guarantees because it is built on top of Spanner. Unlike traditional scale-up relational databases that hit vertical single-node ceilings, Spanner combines full relational SQL semantics and multi-table ACID transactions with the horizontal scale-out elasticity for both reads and writes of a NoSQL system.
Spanner makes the Lakehouse runtime catalog highly available through Spanner’s regional configurations with up to 99.99% availability. Spanner also delivers out-of-the-box scalability for Lakehouse runtime catalog: As a horizontally scalable database, Spanner does not require manual sharding and scales compute and storage independently and transparently. Spanner dynamically monitors data volume and query load, splitting and redistributing data ranges across nodes. Compute nodes scale dynamically based on CPU utilization and storage thresholds. Spanner automatically detects split-level overload and moves heavy splits away from overloaded nodes. Spanner also lets Lakehouse runtime catalog users eliminate the traditional trade-offs between relational consistency and distributed scalability, delivering Lakehouse runtime catalog’s industry-leading consistency guarantees. Lakehouse transactions are serializable — the order of transactions within the database is the same as the order in which clients observe the transactions to have been committed. This foundation allows Lakehouse users to operate at agent-scale.
Then, to further power agentic use cases, the Lakehouse runtime catalog integrates directly with Knowledge Catalog to easily discover lakehouse Iceberg tables and provide trusted context to agents. Knowledge Catalog leverages an efficient combination of full-text search and native vector search provided by Spanner; this approach enables better recall for search retrieval, pairing lexical keyword searches with semantic embeddings in a single query. Because both index types are built on the identical base dataset, they update with strict, transactional ACID consistency alongside base table DML operations. This removes operational overhead such as managing sync pipelines, and external-vector and full-text search systems. In short, the Lakehouse runtime catalog provides faster time-to-market for your agentic use cases.
With Google Cloud’s borderless Lakehouse based on Apache Iceberg, you can combine your analytical data with your operational workloads. Use cases span combining data assets from your Lakehouse with OLTP data (from Spanner) for analytics, to low-latency serving applications where your Lakehouse assets are accessible in an operational database such as Spanner, to conversational analytics in first-party and third-party agents.
Below, in an example, you can see the Lakehouse runtime catalog with Spanner in action. Here, we combine analytical data for taxi trips in Manhattan (backed by Apache Iceberg) with operational data for taxi zones in Spanner to find the most congested traffic routes. The example also shows that you can also use a Conversational Analytics Agent to access the same data and get second-order insights.
Modernizing to Google Cloud’s Lakehouse minimizes data silos across analytics engines and agents, unifies multi-engine governance, provides trusted context to your agents and slashes operational TCO. Leveraging the Spanner-based Lakehouse runtime catalog helps prepare your modern cloud environments to operate at agent-scale. To learn more and get started with a free trial, visit the