# The Company Behind Cursor's Vector Search Just Killed the Vector-First Database

> Source: <https://dev.to/jamilxt/the-company-behind-cursors-vector-search-just-killed-the-vector-first-database-3e4j>
> Published: 2026-10-03 12:04:25+00:00

Last month I watched a teammate spend two days migrating 4 million embeddings from Postgres into a dedicated vector database because a benchmark blog post scared him. The migration added a second database, a sync job, and a new failure mode. His queries got 30 ms faster. Nobody using the product noticed.

I thought about him when I read turbopuffer's September 30 post, [RIP, vector database](https://turbopuffer.com/blog/rip-vector-database). This is not a hot take from a startup nobody uses. turbopuffer is the search layer behind Cursor and Notion. Cursor manages billions of vectors across millions of codebases on it. Notion runs more than 10 billion vectors over millions of namespaces. If anyone has earned the right to say what a vector database should be, it is the company serving those workloads. And their conclusion is blunt: they are removing the vector index as the primary data structure in their next major version.

If the most successful vector database of this cycle is demoting vectors, the question worth asking is the one your architecture should have answered in the first place: do you actually need a dedicated vector database, or was Postgres fine all along?

I have not used turbopuffer myself. Everything about their architecture below comes from their own blog posts and the Hacker News thread, linked inline. What I bring to the table is the Postgres side of this decision, which I have lived.

Read the title and you could easily assume vector search is dying. It is not. The announcement is narrower and more interesting.

So the real headline: even a company whose entire original identity was vector-first storage hit a wall where treating vectors as special made everything else worse. Storage amplification for multi-vector documents, write amplification on every rebalance, and query plans locked to cluster-sized blocks of 100 to 200 documents, when engines like DuckDB work in 2,048-row batches and ClickHouse up to around 65k. That is the engineering argument. Now the decision it should trigger in your stack.

Now the part where I stop quoting benchmarks and show you the queries. Because in my experience the benchmarks matter less than one question nobody's blog post asks: what does the rest of your query look like?

Here is the situation that keeps repeating across teams. You need "find me similar documents, but only in this workspace, only in this category, that the user has not already dismissed, joined back to the documents table for display."

In a dedicated vector database, that query shape is awkward by design. Vector engines are built around one primary query: nearest neighbors, optionally with attribute filters. The join back to your relational data happens in application code, often as a second round trip. In Postgres it is one query:

``` js
SELECT d.id, d.title, d.snippet,
       e.embedding <=> $1 AS distance
FROM embeddings e
JOIN documents d ON d.id = e.document_id
WHERE d.workspace_id = $2
  AND d.category = ANY($3)
  AND NOT EXISTS (
      SELECT 1 FROM dismissed_items x
      WHERE x.user_id = $4
        AND x.document_id = d.id
  )
ORDER BY e.embedding <=> $1
LIMIT 20;
```

Filters, joins, permissions, and ranking in one statement, inside one transaction, backed by one backup and restore story. This is the workload where Postgres does not just compete with dedicated vector databases, it is the better tool by design.

The common counter-argument is "but the ANN index gets slower when filters are selective." True, and there is a standard fix:

```
CREATE INDEX ON embeddings
USING hnsw (embedding vector_cosine_ops)
WHERE workspace_id = 42;
```

A partial HNSW index per high-traffic workspace keeps the graph small and the recall high. pgvector's partial index support is the underrated feature in this entire debate. For the full pattern, pgvector's README documents iterative index scans and exact fallback when post-filtering would otherwise starve your LIMIT, so filtered queries stay correct, not just fast.

One more thing the benchmarks never model: the embeddings table is not the only thing touching those rows. When the document changes, you want to re-embed it and delete stale vectors in the same transaction as the content update. In two-database architectures, that consistency is a distributed systems problem. In Postgres, it is a WHERE clause.

Honesty requires the other side of the ledger. Three situations where Postgres is the wrong choice, and I say that as someone who argues for Postgres by default.

Notice the pattern in that list. These are scale problems. If you are arguing about them before launch, you are arguing about the wrong thing. Here is the checklist I now apply before any migration discussion.

`hnsw.ef_search`, use partial indexes for hot tenants. Benchmark with your own data, not blog numbers.
The threshold that matters most is not vector count. It is the question from earlier: does your retrieval query look like SQL, or does it look like nearest-neighbors-plus-a-prayer? Benchmarks measure the second shape. Most applications need the first.

The vector database gold rush of 2023 sold "your data is different now, you need new infrastructure." Three years later, the leading vector-first database is re-architecting around exactly the design Postgres people always argued for: stable document IDs as the primary key, vectors as one index among several. Turbopuffer arrived at the relational model from the vector side. Postgres arrived at vector search from the relational side. They are meeting in the middle, and that convergence tells you where the industry actually landed.

Meanwhile the durable trend is consolidation. Supabase, Neon, and the managed Postgres providers all ship pgvector by default now, and the default answer to "where do my embeddings live" keeps collapsing into "next to the rest of my data." The teams still running separate vector infrastructure are increasingly the ones with the three scale signatures above, not the ones who read an exciting benchmark in 2024.

That teammate from the opening? His migration turned out fine, technically. But when I asked what he would do differently, he said he wishes he had run the filtered query shape first, because that was his actual workload. Measure your query, not the marketing.

I write about databases, AI infrastructure, and the engineering decisions behind the hype every week. Subscribe, it is free.

Over to you: are you running embeddings in Postgres, or did you end up on a dedicated vector database? If you migrated in either direction, what triggered it? I read every comment.
