Today’s data lakehouse is no longer mere data repository, but increasingly a system of action, actively executing tasks via always-on, autonomous AI agents. Rather than waiting for static reports, these agents run continuous, real-time reasoning loops, monitoring supply chains, flagging anomalies, and executing business workflows. To scale this model, AI agents need to access your entire data estate and the right context to understand what data to use and when. However, traditional data architectures, with their high costs, fragmented security, and weak governance, don’t make it easy.
Today at Next Tokyo, we’re introducing enhancements to our borderless Lakehouse. Built on open Apache Iceberg, it connects your on-premises, cross-cloud operational systems, and SaaS application clouds, so you can activate and query your data wherever it lives, without moving it.
The borderless Lakehouse enables Gemini Enterprise and conversational agents to analyze and act on data regardless of its physical location. Built on the Iceberg REST catalog, it lets you discover and query remote data instantly, eliminating the high costs and delays of building data pipelines. This connectivity is made possible with catalog federation (now in preview) for AWS Glue, Databricks Unity, and Snowflake Horizon, providing secure, bi-directional access via BigQuery, Managed Service for Apache Spark, and any Iceberg-compatible engine.
This open approach also extends to the application layer, with zero copy data integrations to major SaaS applications like SAP, Salesforce, and Workday. BigQuery can now directly query live application data in these platforms without complex ETL pipelines, while they can run BigQuery's AI engines on their own data in-place to easily unify finance, HR, and customer data. These capabilities unlock three significant benefits:
Zero-copy, cross-cloud analytics: Instantly discover and query enterprise data across platforms without duplicating files, allowing various data teams to analyze the exact same copy of Apache Iceberg data.
Bidirectional interoperability: Read and write across environments, querying external tables from BigQuery or Managed Spark, then sharing derived datasets enriched by Google AI back to partner systems for downstream action.
Unified governance and access control: Get out-of-the-box governance with trusted context and Gemini insights for your agents. Secure access control at the table level with support for credential vending, regardless of which platform initiates the query.
The borderless Lakehouse expands to Google Cloud’s database portfolio which lets you integrate transactional systems with data lakehouses: Spanner Omni lets you run the highly scalable database in any environment outside of Google Cloud, and Lakehouse Federation for AlloyDB allows transactional systems to directly query warehouses. By eliminating costly data movement, your teams can securely and efficiently analyze live operational and historical data together in real time. Now your lakehouse is truly borderless.
Historically, running advanced analytics or training ML models across clouds meant a cross-cloud tax: high egress fees, network latency, and fragile ETL pipelines. The borderless Lakehouse solves this with Cross-Cloud Interconnects. These private, dedicated links deliver consistent bandwidth and lower latency than the public internet, at a fraction of the cost of traditional cross-cloud connections.
The borderless Lakehouse supports zero variable egress costs when accessing your data from AWS1. With Partner Cross-Cloud Interconnect pricing, you get predictable monthly costs with an SLA-backed connection. This managed, private connectivity simplifies provisioning from 1G to 100G using a flat-rate, subscription-based model.
In addition, intelligent cross-cloud caching in the borderless Lakehouse securely stores remote data fragments temporarily inside Google Cloud to eliminate repeated, costly transfers for subsequent ad-hoc or BI queries. These capabilities, combined with BigQuery's vectorized processing, BigQuery AI functions for multimodal analysis, Spark's Lightning Engine runtime, and scalable metadata storage, scale to petabytes of data without sacrificing performance and enable:
In-place AI and machine learning: Apply powerful AI models and Gemini directly to AWS data and Azure without migration. Ingesting metadata and context right where it lives helps guarantee high-accuracy grounding and a faster time-to-market.
Avoid the cost and delay of data copying: Query live data stored in other clouds directly and bring your data closer to your agents.
Unified analytics experience: Deliver consistent, hardware-optimized performance across BigQuery, Spark, or open-source engines by centralizing multi-cloud compute back to Google Cloud’s infrastructure.
To prevent hallucinations, AI agents need more than raw technical metadata; they need deep business context. The borderless Lakehouse relies on Knowledge Catalog, our always-on agentic context engine, to establish a unified view of your enterprise context across clouds, without moving physical files.
To support this, the borderless Lakehouse’s runtime catalog automatically synchronizes with AWS Glue, Databricks Unity Catalog, and Snowflake Horizon, all in preview. Knowledge Catalog then ingests these feeds to aggregate, extract, and index their metadata, translating raw schemas into clear business terminology and searchable column-level lineage. As schemas change, Knowledge Catalog instantly updates their business meaning, delivering:
Lower costs and overhead: Minimize expensive, time-consuming data migration pipelines while gaining a consolidated view of your entire multi-cloud estate.
Trustworthy AI decisions: Provide agents with a clear semantic layer and lineage tracking so they can quickly discover, trust, and accurately interpret data.
Automated governance: Embed security directly into the metadata layer, helping ensure AI agents strictly respect compliance guardrails and access permissions.
By pairing the open-source Google Cloud Data Agent Kit with the Conversational Analytics API, you can build and publish custom data agents that operate on the borderless Lakehouse directly into Gemini Enterprise. This allows business users to talk to their data using natural language. The Data Agent Kit is a full-stack collection of agent building capabilities — meeting developers inside their favorite IDEs (like VS Code) — to package pre-codified analytical skills and Model Context Protocol (MCP) tools. With the borderless Lakehouse, your agents operate across your entire data estate and help with:
Integrated data and agent ecosystem: Built-in MCP tools establish secure, direct connections to BigQuery, Managed Spark, and Cloud Storage, eliminating the need to write complex pipeline code or copy-paste massive table schemas into LLM prompts.
Self-service analytics: Business users bypass static dashboards, querying and visualizing multi-cloud datasets instantly in plain language within the Gemini Enterprise interface.
Grounded, high-accuracy agent results: Agents run on top of the Knowledge Catalog, ensuring that natural-language-to-SQL translations are anchored in curated business schemas, highly secure, and strictly governed.
The borderless Lakehouse redefines enterprise AI economics by delivering compounded savings across data transfer, compute, and token consumption. By utilizing Cross-Cloud Interconnects and zero-copy sharing, you bypass fragile data pipelines and unpredictable egress fees to query remote datasets at a flat, predictable rate. Knowledge Catalog filters and delivers the precise, minimal business context required for each prompt, preventing token bloat and eliminating unnecessary reasoning loops. BigQuery AI prevents runaway agent billing through built-in token controls that let you estimate token usage pre-query, enforce strict limits, and leverage an optimized mode that automatically uses smaller, distilled models. In fact, customers are seeing 230x reduction in token consumption using BigQuery’s cost-optimized, built-in AI functions.
The future belongs to the system of action, and the borderless Lakehouse allows your AI agents to query, reason, and act on your data wherever it lives — safely, instantly, and cost-effectively. To start building, check out the official Google Cloud Lakehouse About Guide and explore our Building a borderless Lakehouse codelab.
1. Customers are required to pay an hourly fee for interconnection service.