{"slug": "alloydb-a-unified-database-engine-for-hybrid-search", "title": "AlloyDB: A unified database engine for hybrid search", "summary": "Google's AlloyDB AI now offers a unified hybrid search capability that combines vector search and full-text search in a single SQL function, ai.hybrid_search, using Reciprocal Rank Fusion (RRF) to eliminate manual score normalization. The update adds a RUM extension for low-latency relevance ranking and phrase matching, a natively supported BM25 index for keyword scoring, and an external search Foreign Data Wrapper (FDW) that queries Elasticsearch, OpenSearch, and Solr from within AlloyDB. The change matters because prior hybrid search required multi-step orchestration, cross-service joins, and brittle score-normalization logic in the application layer.", "body_md": "For modern AI and RAG applications, achieving high search relevance requires that you combine at least two techniques: vector search for semantic context, and full-text search (FTS) for keyword precision. While AlloyDB for PostgreSQL supports both of these capabilities, managing them has traditionally required a more hands-on operational approach to ensure peak performance.\n\nThe challenge is not executing the searches, but the subsequent fusion of result sets. Merging results from the vector query (distance scores) and the FTS query (relevance scores) requires complex SQL queries or custom code in the application layer. This often means maintaining a separate system for fusion, score normalization, and re-ranking.\n\nThis article details how AlloyDB AI's hybrid search eliminates this complexity. We will explore how recent updates allow you to:\n\n**Simplify hybrid search:** Consolidate complex SQL queries, or multi-step application workflows into a single, high-performance SQL function powered by Reciprocal Rank Fusion ([RRF](https://docs.cloud.google.com/alloydb/docs/ai/run-hybrid-vector-similarity-search#rank-fusion)).\n\n**Optimize FTS performance:** Use the new RUM extension to achieve low-latency relevance ranking and efficient phrase matching by storing word positions directly in the index.\n\n**Leverage industry-standard ranking:** Utilize the [new natively supported BM25 index](https://cloud.google.com/blog/products/databases/native-bm25-search-in-alloydb-and-cloud-sql) for superior keyword-based scoring directly out of the box.\n\n**Expand search versatility:** Execute queries against specialized external clusters, including Elasticsearch, OpenSearch, and Solr, using the new external search Foreign Data Wrapper (FDW) without leaving the AlloyDB environment.\n\nBefore AlloyDB AI's native solution, achieving robust hybrid search was very demanding, especially for developers trying to keep this logic within the database using standard SQL. This approach required multi-step orchestration that was not only difficult to manage and maintain for two sources, but became virtually impossible to scale as additional sources were added. These steps included:\n\n**Executing vector search:** Run a query using a vector column or vector index (for AlloyDB this can be a ScaNN index) to find top k results, generating vector scores.\n\n**Executing the FTS query:** Run FTS query e.g. by using a generalized inverted index, or GIN, to find top k results.\n\n**Normalizing scores (the brittle step):** Write complex SQL logic or custom application code to map both result sets onto a common scale. This logic is prone to breaking when data distributions change.\n\n**Performing cross-service joins and re-rankings:** Use application memory to perform a complex FULL OUTER JOIN on document IDs, apply a weighted summation of the normalized scores, and finally sort the combined results in cases where the FTS search was done in a separate system than the one used for vector search.\n\nThis decentralized approach led to fragile score logic, increased latency, high operational load, and a dependency on application expertise for maintaining search quality.\n\nThe key to simplifying this is to adopt RRF, which, because it is inherently rank-based, cleverly bypasses the brittle step of score normalization entirely. RRF consists of two parts:\n\nInstead of a multi-step fetch and join process, you simply call one SQL function, providing your search components as a declarative JSON array:\n\nSQL\n\n(Note: The unified `ai.hybrid_search` API isn't limited to just two components. It also natively supports external search sources via the new FDW).\n\nThe hybrid_search() function executes the entire workflow in a single query plan, minimizing overhead and helping ensure transaction consistency. It uses:\n\n**Dynamic CTE generation:** The function constructs dynamic SQL, creating Common Table Expressions (CTEs) for each component. Each CTE is responsible for calculating the positional rank (ROW_NUMBER()) of its results.\n\n**Kernel-level fusion:** All ranked component results are immediately combined using a FULL OUTER JOIN based on the document ID.\n\n**Final RRF score:** The combined ranks are used to calculate the final, unified score using the RRF formula:\n\nBy taking this approach, brittle score calculations and application-side joins become unnecessary. While Reciprocal Rank Fusion (RRF) is our current ranking algorithm, we intend to introduce additional merging and ranking options down the road.\n\nMoving from complex external workflows to native SQL functions provides immediate, measurable benefits:\n\n| **Metric** | **Multi-step application workflow (simulated by manual SQL)** | **AlloyDB AI-native hybrid_search() UDF** | \n| **Code complexity** | Complex score normalization functions, service calls, and application joins. | Zero external logic; a single, declarative SQL call. | \n| **Maintenance** | Constant tuning of normalization formulas. | No tuning required; RRF is rank-based and distribution-agnostic. | \n\nAlloyDB AI's hybrid_search() function is designed to deliver comprehensive search relevance by combining high-performance vector search with FTS. While AlloyDB's native FTS capabilities are powerful, reliance on the standard PostgreSQL GIN index for the text component can present a bottleneck for advanced operations.\n\nThe challenge with GIN indexes is that they do not store the positional information of words. This limitation forces costly table scans to re-analyze content for:\n\n**Relevance ranking:** Calculating search result scores based on word proximity and frequency\n\n**Phrase searching:** Finding exact word order, which requires positional information.\n\nThe **RUM extension** is a powerful index access method based on GIN that directly resolves these performance issues.\n\n**RUM's core advantage:** Unlike the GIN index (which maps word -> [docID]), the RUM index stores the positional information of each word directly within the index (e.g., word -> [(docID1, [pos])], ...).\n\n**Performance benefits:** This allows RUM to perform complex operations like ranking and phrase matching predominantly within the index itself, avoiding expensive heap scans. RUM provides significantly faster relevance ranking and efficient phrase and proximity searches.\n\n**Great for hybrid search:** RUM is a critical complement to vector search (like ScaNN) within a hybrid search framework, as GIN’s potential latency makes it less suitable for a real-time hybrid approach.\n\nThe RUM extension is an excellent choice for ranking-heavy or high-concurrency search applications. By storing word positions directly in the index, RUM cuts out the need to re-scan table pages during ranking, giving you fast and sorted results. However, it comes with some trade-offs: slower index builds and a larger disk footprint. If your workload prioritizes quick queries and precise ranking over write speeds and storage density, RUM is an investment that's well worth it.\n\nAdopting RUM directly bolsters the performance of FTS in AlloyDB, showing considerable gains for raw FTS queries, as well as gains on hybrid search that includes FTS queries. Below are performance statistics showing the speedup between using RUM compared to GIN on various [BEIR](https://github.com/beir-cellar/beir) datasets.\n\nRUM integrates with the specialized <=> distance operator for search and ranking, which is supported for use in hybrid search SQL calls, delivering the best of both worlds. The following [codelab](https://codelabs.developers.google.com/next26/alloydb-ai-hybrid-search) shows how to configure and use the RUM index with the Hybrid Search UDF. \n\nWe recently introduced the [BM25 index](https://docs.cloud.google.com/alloydb/docs/ai/create-bm25-index) to CloudSQL and  AlloyDB in preview, bringing the industry-standard relevance ranking via the pg_textsearch extension. BM25 uses [TF-IDF](https://en.wikipedia.org/wiki/Tf%E2%80%93idf), which accounts for term frequency saturation and document length normalization, delivering significantly higher precision and quality for keyword-based search queries. By leveraging native BM25 indexes within AlloyDB AI, you can achieve superior ranking accuracy out of the box, so that exact-term matches and critical keywords score well without having to implement external search engines or complex custom scoring logic.\n\nTo further extend search capabilities, we also introduced the external search Foreign Data Wrapper ([FDW](https://docs.cloud.google.com/alloydb/docs/reference/extensions#fdw-extensions)) in AlloyDB AI. This lets you execute full-text searches against specialized external clusters, starting with Elasticsearch, OpenSearch, and Solr, and provides several key architectural advantages:\n\n**Optimized retrieval:** Leverage ranking algorithms and a richer feature set from dedicated search backends without leaving the AlloyDB environment.\n\n**Unified SQL interface:** Interact with external data, perform joins, and meld results using standard PostgreSQL SQL without losing the expressiveness of advanced FTS queries.\n\n**Strong portability:** Maintain existing search infrastructure while benefiting from the simplified hybrid architecture offered by AlloyDB AI.\n\nThe following codelabs are end-to-end code guides walking through hybrid search with [Elasticsearch](https://codelabs.developers.google.com/next26/alloydb-elastic?hl=en#0) and [Solr](https://codelabs.developers.google.com/alloydb-solr-fdw?hl=en#0) integrations.\n\nThe true challenge of modern search is not technical in nature, but in achieving architectural simplicity and sustained operational stability. AlloyDB AI addresses this with a unified platform built on three core innovations:\n\n**Simplifying hybrid search:** AlloyDB AI’s hybrid search function, powered by RRF, transforms a fragile, multi-step application workflow into a single, high-performance SQL call. This native implementation eliminates the need for complex score normalization and application-side joins, drastically lowering the operational and engineering cost while delivering consistently accurate and fast hybrid relevant results.\n\n**Optimizing FTS performance:** To ensure the FTS component of Hybrid Search meets diverse application demands, AlloyDB AI offers distinct full-text search options. The RUM extension optimizes for low-latency performance by storing positional data directly in the index to enable faster relevance ranking and efficient phrase searches **—** important when you need fast query speeds. Alternatively, the BM25 index provides industry-standard relevance ranking, and is the preferred choice when prioritizing keyword scoring precision. You can select between these options depending on whether your focus is search latency or ranking precision.\n\n**Enhancing versatility through external search:** The addition of the external search through FDW extends AlloyDB’s reach to specialized search backends like Elasticsearch. This lets you leverage the superior scale and advanced retrieval features of dedicated search clusters while maintaining a familiar PostgreSQL interface. By integrating these external results directly into the hybrid search framework, AlloyDB AI ensures that you can combine even the most massive text repositories with vector-based semantic insights.\n\nBy consolidating the complexity of scoring, joining, and re-ranking within the database kernel, and offering a dual-path approach to FTS — leveraging RUM for optimized, low-latency, internal FTS or external search for specialized, scalable backends — AlloyDB AI delivers a robust and self-contained search foundation. This coexistence is essential to the overall hybrid search story, providing the flexibility to choose the optimal path based on workload, data volume, and existing infrastructure. Now, you can focus on building intelligent application features, confident that your search architecture is both highly performant and easy to maintain.\n\nWatch how AlloyDB is the [ultimate hybrid search engine](https://www.youtube.com/watch?v=vYB0P9rPXt8).\n\nLearn about [BM25 support in AlloyDB](https://www.youtube.com/watch?v=-JxQb-kjFHk). \n\nReady to bring more speed and cost-efficiency to your AI workloads?\n\n**New to AlloyDB?** Discover AlloyDB with a [30-day free trial](https://docs.cloud.google.com/alloydb/docs/free-trial-cluster).\n\n**Getting started with hybrid search**: Set up a [text-search index](https://docs.cloud.google.com/alloydb/docs/ai/full-text-search-overview) and select a [vector search index](https://docs.cloud.google.com/alloydb/docs/ai/choose-index-strategy). Once these are created, you can find some [examples](https://docs.cloud.google.com/alloydb/docs/ai/run-hybrid-vector-similarity-search) to perform various hybrid search queries.\n\n**External search**: Create a foreign data wrapper and foreign table in AlloyDB to query external data from [Elasticsearch](https://docs.cloud.google.com/alloydb/docs/elastic-search), [Solr](https://docs.cloud.google.com/alloydb/docs/solr-search), or [OpenSearch](https://docs.cloud.google.com/alloydb/docs/opensearch).", "url": "https://wpnews.pro/news/alloydb-a-unified-database-engine-for-hybrid-search", "canonical_source": "https://cloud.google.com/blog/products/databases/simplify-ai-search-with-alloydb-hybrid-search-and-rrf/", "published_at": "2026-10-06 16:00:00+00:00", "updated_at": "2026-10-06 16:17:55.444700+00:00", "lang": "en", "topics": ["ai-search", "ai-infrastructure", "large-language-models", "generative-ai", "developer-tools"], "entities": ["Google", "AlloyDB", "AlloyDB AI", "PostgreSQL", "Elasticsearch", "OpenSearch", "Solr", "Reciprocal Rank Fusion"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/alloydb-a-unified-database-engine-for-hybrid-search", "markdown": "https://wpnews.pro/news/alloydb-a-unified-database-engine-for-hybrid-search.md", "text": "https://wpnews.pro/news/alloydb-a-unified-database-engine-for-hybrid-search.txt", "jsonld": "https://wpnews.pro/news/alloydb-a-unified-database-engine-for-hybrid-search.jsonld"}}