{"slug": "nearest-by-join-scaling-vector-search-in-databricks-runtime", "title": "NEAREST BY Join: Scaling Vector Search in Databricks Runtime", "summary": "Databricks built vector search directly into Databricks Runtime as a first-class SQL join called NEAREST BY, replacing its earlier VECTOR_SEARCH implementation that federated requests to an external real-time endpoint and was capped by endpoint sizing rather than cluster size. The new stack lowers the top-k ranking join onto SIMD-accelerated distance functions, a bounded top-k aggregate, and a fused Photon operator with a custom GEMM kernel, plus an optional IVF index stored as an ordinary liquid-clustered Delta table so APPROX queries score a fraction of the base vectors. Databricks cites batch workloads such as a payments company matching 100M+ daily transactions against 140M merchant embeddings and a quantitative fund running million-query batches against a 50M-vector corpus as the target use cases.", "body_md": "How we built vector search into Databricks as a first-class SQL join, with deep kernel optimizations in Photon and a vector index in an open storage format.\n\nby  [Zero Qu (Zero)](https://www.databricks.com/blog/author/zero-qu-zero), [Alexis Schlomer](https://www.databricks.com/blog/author/alexis-schlomer), [Akash Nayar](https://www.databricks.com/blog/author/akash-nayar), [Yingyi Bu](https://www.databricks.com/blog/author/yingyi-bu)  and  [Sergei Tsarev](https://www.databricks.com/blog/author/sergei-tsarev) \n\nVector search originated as a serving problem. The classical use case is a chatbot or a search bar: one query embedding arrives, and the system is optimized to return the top-k nearest documents within tens of milliseconds.\n\nHowever, a good share of the vector search workloads on our platform are inherently batch-oriented — precomputing exact or approximate nearest neighbors offline rather than looking them up at request time. A payments company matches 100M+ daily transactions against 140M merchant embeddings for entity resolution; a data firm enriches tens of millions of historical records nightly; a quantitative fund runs million-query batches against a 50M-vector corpus for taxonomy tagging.\n\nEntity resolution, deduplication, semantic tagging, classification, record enrichment, batch recommendations — these are fundamentally batch workloads: millions of queries against millions to billions of vectors on a schedule, measured by whether the job completes within its SLA at a reasonable cost rather than by the latency of a single lookup. These workloads deserve a very different architecture for better performance, reliability, and cost efficiency — so we went back to first principles.\n\nThe Databricks Runtime fits these requirements perfectly — a distributed, fault-tolerant, elastic execution engine built on top of Spark and Photon, a vectorized native C++ query engine. This is exactly why we decided to build vector search directly as an engine native feature rather than relying on separate infrastructure.\n\nOur first version of the [VECTOR_SEARCH](https://docs.databricks.com/aws/en/sql/language-manual/functions/vector_search) SQL function was designed to federate requests to an external real-time Vector Search endpoint. It was implemented as a Generate node streaming one query row at a time: every row incurred a network request, a response to deserialize, and possibly retries. It worked, but exposed a performance ceiling — throughput was capped by real-time endpoint sizing rather than runtime cluster size, with the runtime engine reduced to a dispatcher. It also missed the true shape of the query. A batch vector search is not a million small searches. It is one large query: for each row on the left, find the k nearest rows on the right — a top-k ranking join. Executing enormous joins is exactly what the runtime engine excels at.\n\nImplementing vector search natively in the runtime engine pays off from two angles.\n\nThis resulted in a deliberately small yet deep stack: a new join syntax, [NEAREST BY](https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-syntax-qry-select-nearest-by), that makes the top-k ranking join a first-class relational operation; a rewrite that lowers it onto three primitives — SIMD-accelerated distance functions and a bounded top-k aggregate; a fused Photon operator that replaces the plan's entire middle with a custom GEMM kernel; and an optional IVF index built as an ordinary liquid-clustered Delta table, which lets APPROX queries score a fraction of the base vectors with the same kernels.\n\nExisting engines converged on two interface shapes. Postgres with pgvector and Snowflake compose distance operators with ORDER BY … LIMIT — batch then needs a LATERAL subquery per driving row, and the optimizer lacks a pattern to recognize for differentiating KNN and ANN queries. That recognition is also fragile: drift from the expected query shape and the fast path silently disappears. BigQuery exposes a table-valued function — batch is first-class, but column references are strings the parser can't validate.\n\nStructurally, batch vector search is a binary relational operation: two table inputs, an output combining both, a per-left-row top-k connecting them. The syntax encodes that structure as a native top-k ranking join:\n\nThe join is asymmetric, similar to LATERAL: the left side drives, the right side is searched. Ranking direction is explicit: BY SIMILARITY descending, BY DISTANCE ascending. LEFT OUTER keeps query rows with no candidates, and the BY expression is pluggable: any orderable scalar over both sides works, so other scoring expressions can reuse the same clause later.\n\nAPPROX and EXACT encode a semantic contract. EXACT guarantees the true top-k by exhaustive evaluation; APPROX lets the optimizer substitute an approximate strategy, such as an ANN index, where one applies. So creating or dropping an index can never silently change query results: only queries that specify APPROX consent to approximation.\n\nNEAREST BY parses into a logical join node, which the optimizer lowers onto standard relational operators: the rewrite tags each query row with a generated id, scores every (query, base) pair, keeps the k best per id with a grouped top-k, and inlines the kept rows back out:\n\nSemantically, this encapsulates the entire feature: a cross join, a scalar scoring expression, and a grouped top-k aggregate. Because every operator is an ordinary relational one, the plan distributes, spills, and retries like any other — correctness and fault tolerance come for free. What the rewrite really isolates is the two primitives all execution time flows through: the distance function that scores a pair, and the aggregate that keeps each group's k best.\n\nWe implemented every operator in this plan natively in Photon, plus one additional fused operator, purposely built for vector search, that collapses the plan's middle section entirely with a more performant, batch-friendly kernel.\n\nThe core building blocks are a family of vector SQL functions over ARRAY<FLOAT> columns. Three of them are responsible for similarity and distance computation:\n\n| **SQL Function** | **Computes** | **Nearer Means** | \n|---|---|---|\n| [vector_inner_product(a, b)](https://docs.databricks.com/aws/en/sql/language-manual/functions/vector_inner_product) |  | Higher (BY SIMILARITY) | \n| [vector_cosine_similarity(a, b)](https://docs.databricks.com/aws/en/sql/language-manual/functions/vector_cosine_similarity) |  | Higher (BY SIMILARITY) | \n| [vector_l2_distance(a, b)](https://docs.databricks.com/aws/en/sql/language-manual/functions/vector_l2_distance) |  | Lower (BY DISTANCE) | \n\nAlongside the similarity and distance functions, we shipped two norm helpers — [vector_norm](https://docs.databricks.com/aws/en/sql/language-manual/functions/vector_norm) and [vector_normalize](https://docs.databricks.com/aws/en/sql/language-manual/functions/vector_normalize), and two aggregates, [vector_sum](https://docs.databricks.com/aws/en/sql/language-manual/functions/vector_sum) and [vector_avg](https://docs.databricks.com/aws/en/sql/language-manual/functions/vector_avg). Together they cover both query and index construction: the distance functions score queries and assign rows to their nearest centroid, while the aggregates and normalizers recompute those centroids during k-means.\n\nPhoton runs the whole family as native SIMD kernels. Every metric is multiply-adds at the core, and a single fused multiply-add (FMA) instruction computes 𝑎 · 𝑏 + 𝑐 across an entire vector register per issue. The kernels are implemented based on four deliberate design choices.\n\nIn plain SQL, grouped top-k is a window function: `ROW_NUMBER() OVER (PARTITION BY query ORDER BY score).` This sorts every partition in full, then throws away everything but the top k ranks. We instead extended the existing [max_by](https://docs.databricks.com/aws/en/sql/language-manual/functions/max_by) / [min_by](https://docs.databricks.com/aws/en/sql/language-manual/functions/min_by) aggregates with a third K parameter overload. The implementation is built around four key properties.\n\nWith all the native kernels above, the query plan is fully Photonized, yet still far from optimal at batch scale. The reason is theoretical, not implementational, and the [roofline model](https://people.eecs.berkeley.edu/~kubitron/cs252/handouts/papers/RooflineVyNoYellow.pdf) is a simple and effective way to visualize it.\n\n𝑃*<sub>𝑎𝑡𝑡𝑎𝑖𝑛𝑎𝑏𝑙𝑒</sub>* = 𝑚𝑖𝑛(𝑃*<sub>𝑝𝑒𝑎𝑘</sub>*, 𝐴𝐼 × 𝐵𝑊)\n\n𝑃*<sub>𝑝𝑒𝑎𝑘</sub>* is the hardware's peak compute throughput (FLOPs/s), *BW* is memory bandwidth (bytes/s), and 𝐴𝐼 is the kernel's arithmetic intensity (FLOPs performed per byte moved). Plot 𝑃*<sub>𝑎𝑡𝑡𝑎𝑖𝑛𝑎𝑏𝑙𝑒</sub>* against 𝐴𝐼 and we get the roofline: a diagonal memory roof meeting a horizontal compute roof at the ridge point — the minimum 𝐴𝐼 at which a kernel can be compute-bound. Left to the ridge, only moving fewer bytes per FLOP helps; right of it, the kernel is compute-bound and the arithmetic units themselves are the limit. Concretely, on a reference m6i.2xlarge machine (illustrative, the constants shift with the hardware):\n\n|  | **Per core** | \n|---|---|\n| **Compute throughput roof** |  | \n| **DRAM bandwidth roof** |  | \n| **Ridge point** |  | \n\n**L1/L2 cache roofs sit 30-60x above the DRAM fair share, but the base table is far larger than any cache, so every base byte crosses the DRAM boundary at least once.*\n\n**The compute throughput roof scales linearly with cores. A cluster of 64 m6i.2xlarge executors (4 physical cores each) peaks at 64 x 4 x 204.8 GFLOP/s ≈ 52.5 TFLOP/s of fp32 FMA.*\n\nScoring two aligned columns yields one score per row — every vector is used once, and no reuse exists. The NEAREST BY shape yields 𝑛<sub>𝑞</sub> × 𝑛<sub>𝑏</sub> scores from only 𝑛<sub>𝑞</sub> + 𝑛<sub>𝑏</sub> distinct vectors, since each base vector is needed by every query.\n\nNow compare what vector search moves under the naive plan vs. a reuse-aware one. The FLOPs are identical under both: 2 × 𝑑 × 𝑛<sub>𝑞</sub> × 𝑛<sub>𝑏</sub>, where 𝑛<sub>𝑞</sub> and 𝑛<sub>𝑏</sub> are the query and base-side cardinality and is the embedding dimension.\n\n| **Plan** | **Bytes moved** | **Arithmetic Intensity (AI)** | **Scaling** | \n|---|---|---|---|\n| **Pairwise cross join** — every operand loaded fresh, used only once | 2 × 4 × 𝑑 × 𝑛 <sub>𝑞</sub> × 𝑛<sub>𝑏</sub> | 1/4 | *0* (1) | \n| **Fused GEMM** — buffer , stream once | 4 × 𝑑 × 𝑛 <sub>𝑏</sub> | 𝑛 <sub>𝑞</sub> /2 | *0* (𝑛<sub>𝑞</sub> ) | \n\nPairwise scoring is pinned to the DRAM slope at 0.8% of the peak, whereas the fused GEMM kernel’s arithmetic intensity grows with the batch size: 𝐴𝐼 = 𝑛<sub>𝑞</sub>/2 — letting it cross the ridge at 𝑛<sub>𝑞</sub> = 64 and ride the compute roof at larger batches. The vector distance and similarity kernels are near-optimal for their input shape, but the naive plan shape has no data reuse to exploit, and batching cannot help it: the pair explosion multiplies FLOPs and bytes equally. The fix is a fused operator plan, let the kernel exploit data reuse, and the arithmetic intensity follows.\n\nThe fused operator is a single Photon execution node that replaces the cross join, the distance / similarity projection, and optionally the partial top-k. It buffers the smaller input, streams the other in batches, scores query-by-base tiles with a custom GEMM kernel, and maintains per-query top-k state across tiles when k is small enough. With the top-k fused, its output is at 𝑛<sub>𝑞</sub> × 𝑘 most candidate rows rather than 𝑛<sub>𝑞</sub> × 𝑛<sub>𝑏</sub> pairs, consumed by the downstream [max_by](https://docs.databricks.com/aws/en/sql/language-manual/functions/max_by) / [min_by](https://docs.databricks.com/aws/en/sql/language-manual/functions/min_by) merge kernel. Peak memory per task is the buffered side + one in-flight batch + *0*(𝑛<sub>𝑞</sub> × 𝑘) selection state. Distribution follows the standard join strategies — broadcast the smaller side when it fits, otherwise partition both sides and run block-cartesian.\n\nAs shown in the roofline model, fusion changes the bytes we move in memory, not the FLOPs we perform. Each base vector is scored against all buffered queries once loaded, so arithmetic intensity grows linearly with the batch and crosses the reference ridge point at 𝑛<sub>𝑞</sub> = 64 — well below the batch sizes of typical production workloads. The same holds when the base side is smaller, arithmetic intensity scales with whichever side stays resident.\n\nCrossing the ridge is necessary, but not sufficient. The 𝑛<sub>𝑞</sub>/2 byte count is the direct result of the blocked GEMM kernel below. Clearing the DRAM roof only moves the bottleneck down, to cache bandwidth and then FMA latency.\n\nConcretely, the kernel is a classic blocked GEMM over the score matrix 𝐷 = 𝑄 · 𝐵<sup>𝑇</sup>. The outer loop advances over the base a panel of vectors at a time and packs each panel dimension-major for reuse across query tiles. Because the embedding dimension is itself blocked into panels, the middle loop sweeps every buffered query tile against the packed panel before the next one is fetched. The innermost loop accumulates a small output tile in registers across one dimension-panel; for embeddings wider than a single panel, the running partials are parked in the score buffer and read back to continue the next panel. Each finished tile is written into a bounded, recycled score buffer that streaming top-k consume in place, so the full 𝑛<sub>𝑞</sub> × 𝑛<sub>𝑏</sub> score matrix is never written to DRAM.\n\nThere are two levels of blocking to improve data reuse. Packed base panels are reused across query tiles to encourage CPU cache hits, while register blocking lets each loaded value contribute to multiple FMAs. Each tile size balances two competing considerations:\n\nMost production workloads run with a relatively small k, so we optimized the streaming top-k for that case. Per-query selection state stays inside the fused operator, each score tile merges into the selections while still in cache, and the full 𝑛<sub>𝑞</sub> × 𝑛<sub>𝑏</sub> matrix never reaches DRAM. Once a query holds k entries, its worst score becomes the admission threshold, letting the merge reject multiple scores per vectorized compare. Only surviving rows are gathered to output. When k is large enough that this state becomes memory pressure, the operator runs the blocked GEMM alone and hands scored tiles to the existing [max_by](https://docs.databricks.com/aws/en/sql/language-manual/functions/max_by) / [min_by](https://docs.databricks.com/aws/en/sql/language-manual/functions/min_by) partial. Under this fallback, emitting one 𝑓𝑝32 score per pair costs 4 𝑏𝑦𝑡𝑒𝑠 against 2𝑑 𝐹𝐿𝑂𝑃𝑠, so 𝐴𝐼 = 𝑑/2— still far past the ridge for any realistic dimension.\n\nEverything so far accelerates exhaustive KNN (k-nearest-neighbor) search; none of it changes that exhaustive search is *0*(𝑛<sub>𝑞</sub> × 𝑛<sub>𝑏</sub>) — a million queries against a billion rows is 10<sup>15</sup> dot products, and no register tile amortizes away an exponent. That is what APPROX and the vector index are for.\n\nThe index is a classic IVF (inverted file) design. Indexing trains k-means over a sample of the corpus to produce a set of centroids. Every base row is assigned to its nearest centroid, and at query time each query scores only the vectors in its nearest clusters. The pruning compounds with 𝑛<sub>𝑏</sub>: at billions scale, a query probes ≤ 0.1% of the corpus, orders of magnitude less distance work than brute force. We chose IVF over graph indexes because independent cluster scans parallelize across executors and the layout is naturally in columnar storage, whereas graph traversal is a chain of serial lookups.\n\nPhysically, the index is an ordinary Delta table. Assignment rows carry the vector next to its centroid id, and the table is liquid-clustered by centroid id. Each cluster's candidates sit contiguous on blob storage, and files whose clusters have no query probed are pruned before they are read. Refresh is transactional and incremental.\n\nAn APPROX query rewrites onto the same primitives: probe the centroids with the same top-k NEAREST BY join, equi-join on centroid id to restrict each query to its probed clusters, score and top-k, then merge. Files added since the last refresh are brute-forced in a compensation branch and unioned into the same merge, so a stale index prunes less, but never degrades search quality.\n\nQuery execution is shaped by query and base cardinalities.\n\n| **Scenario** | **Join strategy** | **Key characteristics** | \n|---|---|---|\n| **Small index table, small query table** | BNLJ, broadcast index | Everything in memory | \n| **Small index table, large query table** | BNLJ, broadcast index | Queries stay partitioned and stream, broadcast index to all | \n| **Large index table, small query table** | BNLJ, broadcast queries | Index partitions stream, broadcast probed queries | \n| **Large index table, large query table** | Equi-join by centroid ID, with in-place shuffle for the index | Shuffle probed queries by centroid ID; leverage the index’s liquid clustering to avoid a full shuffle of the index | \n\nThe large–large case is where liquid clustering pays off hardest: only the query side moves; each probed query is shuffled to its clusters' partitions, while the index side directly scans files for the matching centroid ids. Each partition probes exactly its own clusters' candidates.\n\nWithin a partition, each cluster's candidates arrive as a dense contiguous block, so the same GEMM kernel applies. The exact path runs it once globally; the approximate path runs it once per cluster. Local top-k per query, regroup, merge the partials: the aggregate's partial/merge contract doing exactly what it was built for.\n\nWe evaluated NEAREST BY on canonical workloads observed from customers, with base tables ranging from 100,000 to 5 billion vectors and query batches of up to 10 million vectors, targeting at least 96% recall@K. Using indexed ANN:\n\nBatch vector search did not need a new system — it needed to become a first-class citizen of the one that already holds the data. NEAREST BY expresses the workload as what it structurally is: a top-k ranking join. The roofline model explains why the naive plan is memory-bound and what any faster kernel must do about it. The fused operator and its blocked GEMM turn that analysis into arithmetic intensity, and the vector index scales it further while remaining an ordinary liquid-clustered Delta table. The new machineries are deliberately designed to be simple, narrow, but deep and effective: one join clause, seven vector functions, one aggregate overload, one fused operator, and one storage layout decision. Everything else, from shuffles and spills to governance and autoscaling, came with the runtime engine, built and hardened over more than a decade.\n\nIf building hard computer systems for large scale AI/ML workloads at Databricks sounds interesting to you, [come build with us](https://www.databricks.com/company/careers)!\n\nSubscribe to our blog and get the latest posts delivered to your inbox.", "url": "https://wpnews.pro/news/nearest-by-join-scaling-vector-search-in-databricks-runtime", "canonical_source": "https://www.databricks.com/blog/nearest-join-scaling-vector-search-databricks-runtime", "published_at": "2026-10-05 20:00:00+00:00", "updated_at": "2026-10-05 20:19:11.332782+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "developer-tools", "mlops"], "entities": ["Databricks", "Databricks Runtime", "Photon", "NEAREST BY", "VECTOR_SEARCH", "Spark", "Delta", "pgvector"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/nearest-by-join-scaling-vector-search-in-databricks-runtime", "markdown": "https://wpnews.pro/news/nearest-by-join-scaling-vector-search-in-databricks-runtime.md", "text": "https://wpnews.pro/news/nearest-by-join-scaling-vector-search-in-databricks-runtime.txt", "jsonld": "https://wpnews.pro/news/nearest-by-join-scaling-vector-search-in-databricks-runtime.jsonld"}}