Databricks Adds NEAREST BY Join for Batch Vector Search
The native SQL operation ranks matches for every query row in exact or approximate mode, bringing offline vector workloads into the same engine that holds the source tables.
Edited by Tyronne Panaino
Databricks introduced `NEAREST BY` on October 5, 2026 as a native SQL join for batch vector search. The operation finds the top matching rows for every row on the query side by a distance or similarity expression, giving data teams a relational way to run large offline matching jobs without turning each input row into a separate request to a serving endpoint.
The Databricks SQL reference says the clause applies to Databricks SQL and Databricks Runtime 18 LTS or later. The feature offers explicit `EXACT` and `APPROX` modes, making the accuracy contract part of the query instead of an optimizer detail that can change silently.
A vector query becomes a top-k join
`NEAREST BY` treats batch matching as a join between a query table and a target table. For each row on the left, it returns up to the requested number of target rows, ranked either by the smallest distance or the largest similarity. Databricks allows a result count from one through 100,000 and accepts any orderable scalar expression that can score a pair of rows.
That structure is useful for offline work where the unit of success is a whole job rather than the latency of one search request. Databricks positions the operation for large scheduled workloads such as matching and enrichment, where many query vectors must be compared with a large reference table. The company says embeddings can remain in Delta tables instead of being synchronized to a separate vector-serving system.
The two modes deserve different expectations. `EXACT` exhaustively returns the true top-k under the ranking expression. `APPROX` permits the optimizer to use a faster approximate strategy, which can trade perfect ranking agreement for less work. Because the choice is written in the query, creating an approximate index does not by itself change a query that asked for exact results.
Databricks moves the work into its runtime
In its engineering explanation, Databricks says the implementation combines vector functions, a bounded top-k aggregate and a fused Photon operator designed for batch scoring. For approximate queries, the company describes an inverted-file index stored as a liquid-clustered Delta table so the engine can avoid reading partitions outside the selected clusters.
This is a product architecture claim from Databricks, not an independent comparison with a dedicated vector database. It establishes how the company designed the feature and where it runs, but it does not prove that every workload will be faster, cheaper or easier to operate after migration. Teams still need to test their own vector dimensions, table shapes, recall requirements and job-level cost.
Exactness, scale and current limits
Databricks reports that its own evaluation included one million queries against one billion reference vectors completing in minutes under its indexed approximate path. That result is useful as a statement of intended scale, but the source does not provide an independently reproduced benchmark for this run. Hardware, cluster sizing, index settings, data distribution and the chosen recall target can all affect an observed result.
The SQL reference also records an important boundary: `NEAREST BY` is not supported on streaming DataFrames or Datasets. It is therefore a batch operation, not a replacement for every low-latency serving or continuous-stream use case. The next useful evidence will be customer-run comparisons that publish configuration, recall, total job cost and end-to-end time against existing pipelines.
Status
Confirmed. Databricks has documented `NEAREST BY` for Databricks SQL and Runtime 18 LTS or later. Internal confidence is medium because availability, architecture and performance evidence come from Databricks and were not independently reproduced in this run.
Sources
Update note: Last reviewed 2026-10-06. We will revise this post if Databricks changes availability, adds streaming support or publishes independently reproducible workload evidence.
Sources
- Databricks — NEAREST BY engineering overview — official
- Databricks SQL reference — NEAREST BY clause — official
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.