Lakebase Search: the retrieval primitive for agents on Neon Neon announced that Lakebase Search, its vector and BM25 full-text search primitive for AI agents running on Lakebase Postgres, is now generally available. Lakebase Search ships as two Postgres extensions, lakebase_vector for semantic search and lakebase_text for BM25, and Neon says it led the tested systems on price-performance in VectorDBBench tests on LAION-100M. The durable index lives in object storage so compute can scale to zero without rebuilding the index, and Neon states pgvector users can switch by adding a lakebase_ann index without migrating data or changing queries. Lakebase Search is now generally available Lakebase Search just reached GA. Read the Databricks announcement https://www.databricks.com/blog/lakebase-search-state-art-full-text-and-vector-search-postgres . Lakebase Search https://neon.com/docs/ai/lakebase-search gives agents fast vector and BM25 search inside Lakebase Postgres https://neon.com/docs/introduction/neon-and-lakebase , without moving data to a separate search system. Combined with Neon Functions https://neon.com/docs/compute/functions/overview , Object Storage https://neon.com/docs/storage/overview , and AI Gateway https://neon.com/docs/ai-gateway/overview , it supports the full path from ingesting documents to serving hybrid search, with compute that scales independently of the index. High-quality search was once limited to companies such as Google and Amazon. Building it required search specialists and dedicated infrastructure. It was hard to get right, but small improvements in recall could materially affect the quality of the results. Today, search is a core primitive for apps and AI agents. It connects operational data to models and underpins memory, personalization, and context. The quality of an agentic experience often depends on how well the agent can query that data. Building retrieval for agents requires three components: - The search engine: Runs fast, efficient search at scale. - The ingestion layer: Turns raw data and documents into searchable indexes. - The serving layer: Exposes the right query tools to agents through a consistent interface. Neon provides the primitives https://neon.com/blog/neon-backend-is-ga to build this pipeline around your data: Lakebase Postgres https://neon.com/docs/introduction/neon-and-lakebase for storage and search, Functions https://neon.com/docs/compute/functions/overview for TypeScript ingestion and serving, Object Storage https://neon.com/docs/storage/overview for raw files, and AI Gateway https://neon.com/docs/ai-gateway/overview for generating embeddings. Fast, scalable search in Postgres Lakebase Search https://neon.com/docs/ai/lakebase-search brings vector, keyword, and hybrid search to Neon through two Postgres extensions: - lakebase vector for semantic search - lakebase text for BM25 full-text search You can use them together to build fast, scalable agents while keeping search alongside your operational data in Lakebase Postgres. Vector search with lakebase vector Standard HNSW indexes perform best when their working set remains in memory. As the corpus grows, this ties search capacity to the memory available on the serving compute. lakebase vector is designed for Lakebase Postgres's separated compute and storage architecture. Its durable index lives in object storage, while RAM and local storage cache frequently accessed data. To work efficiently with this storage hierarchy, lakebase vector combines hierarchical IVF with RaBitQ quantization. IVF narrows the search to a small set of relevant clusters, while RaBitQ compresses vectors so more index data fits in the faster cache layers. Fully compatible with scale to zero. The index stays in object storage when compute suspends. When traffic returns, compute reconnects to the same index instead of rebuilding it. This keeps idle compute costs down without making you rebuild the index before search can resume. Compatible with pgvector. If you already use pgvector, switching to lakebase vector does not require migrating data or changing queries. You add a lakebase ann index while keeping the same vector types, distance operators, and query syntax. Built for better price-performance. Because the durable index lives in object storage and compute scales independently, you do not have to size always-on compute around the full index. In our VectorDBBench tests on LAION-100M, Lakebase Search led the tested systems on price-performance: For more details on results and methodology, see the Databricks announcement https://www.databricks.com/blog/lakebase-search-state-art-full-text-and-vector-search-postgres . Full-text search with lakebase text Agents also need another type of search: full-text search. Vector search finds similar meaning, while full-text search finds exact words and phrases. lakebase text adds the lakebase bm25 index type for BM25 keyword search. It preserves standard Postgres tsvector types and query operators while adding corpus-wide BM25 ranking and top-K pushdown. Applications can combine vector and keyword results in one SQL query, giving agents both semantic matches and exact-term results. Build the ingestion pipeline with Neon Functions Before data can be searched, it must be parsed, chunked, embedded, and indexed. A TypeScript Neon Function https://neon.com/docs/compute/functions/overview can run these steps next to the data in Lakebase Postgres. AI Gateway provides an OpenAI-compatible embeddings endpoint, so the Function can generate embeddings without configuring a separate model provider: Index documents when they are uploaded With Function Triggers https://neon.com/docs/compute/functions/triggers/object-storage , raw files uploaded to Object Storage can become searchable automatically. A storage object created trigger invokes the ingestion Function with the bucket and object key: Expose retrieval as one tool The final layer determines how the agent reaches your data. Another Neon Function https://neon.com/docs/compute/functions/overview can wrap the hybrid query and expose it as one tool call. That call runs vector and keyword search in Postgres and returns ranked passages with their sources: Let the agent search in a loop Agents change the retrieval pattern. Instead of relying on one hand-tuned query, the agent can search, inspect the results, and decide whether to proceed or try again with a sharper query, tighter filters, or a different angle. The loop has three steps: - Search: Run a vector, keyword, or hybrid query. - Evaluate: Let the agent inspect the returned passages. - Retry: Rewrite the query or adjust its filters when the results are not good enough. This pattern depends on search that can handle repeated queries without requiring permanently provisioned compute. With Lakebase Search, compute follows active demand and can suspend between periods of activity. The indexes remain durable in storage and are available when compute restarts. Build the retrieval stack around your data Neon brings together the primitives needed for retrieval: Lakebase Postgres https://neon.com/docs/introduction/neon-and-lakebase , Object Storage https://neon.com/docs/storage/overview , Neon Functions https://neon.com/docs/compute/functions/overview , and AI Gateway https://neon.com/docs/ai-gateway/overview . Applications can store operational data, build vector and BM25 indexes, run ingestion code, and expose retrieval tools from the same backend. The result is a retrieval stack built with familiar tools: Postgres for data and search, and TypeScript for ingestion and serving. Lakebase Search also preserves the price-performance benefits described above: the durable index lives in object storage while compute scales with query demand and can suspend when idle.