{"slug": "rip-vector-database", "title": "RIP, Vector Database", "summary": "Turbopuffer is replacing its primary ANN vector index with a new storage architecture informally called \"turbopuffer v3,\" demoting ANN to \"just another\" secondary index, the company said in a technical post. The serverless vector database, whose early customers include Cursor and Notion, said the change will let it serve many more query plans at much greater scale, after pushing its primary ANN index as far as it can. Turbopuffer's v1 stored documents as only an ID and a vector using hierarchical clustering indexes (SPANN, later SPFresh) keyed by an ANN address of ClusterId and LocalId, with attribute filtering and full-text search marking the informal v1-to-v2 transition.", "body_md": "We are changing turbopuffer's storage architecture. The new engine, informally called \"turbopuffer v3\", changes how documents and indexes are laid out, written, compacted, and queried in turbopuffer. It will allow us to serve many more query plans, at much greater scale.\n\nturbopuffer launched as a\n[serverless vector database](https://x.com/Sirupsen/status/1709602855160869275),\nhighly specialized to the task of serving extremely cheap and reasonably fast\nvector searches. Object storage as the source of truth gave the economics, and\ntiered NVMe SSD/memory caches gave the performance. The value of these\nparticular tradeoffs was validated by our earliest customers, including Cursor\nand Notion.\n\nOver time, turbopuffer has become a generalized database for search (and\n[other things](https://linear.app/now/rebuilding-delta-sync-read-path)). The\nquery engine has evolved along the way to support all of these query plans, but\nthe storage architecture has remained largely unchanged: the\nANN vector index was and still is\nthe primary index around which all other indexes and query plans revolve.\n\nWe've pushed the primary ANN index as far as we can, but it's time to move on. We're in the process of moving to a new primary index, and making ANN \"just another\" secondary index. We thought it might be fun to open up the doors and let you follow along.\n\nFor this first update, we'll set the stage with why we're doing this in the first place. Walk with me on a short journey from tpuf v1 to today.\n\nIn the first version of turbopuffer, documents consisted of nothing but an ID\nand a vector. The prevailing wisdom at the time was graph-based vector indexes,\nbut a hierarchical clustering index plays better with object storage. We started\nwith [SPANN](https://arxiv.org/abs/2111.08566), and eventually migrated to\n[SPFresh](https://arxiv.org/abs/2410.14452) to support incremental indexing.\nVectors are clustered into groups, whose centroids are clustered in turn,\nrepeated to form a tree with a single root.\n\n```\n                      ┌───────────────────┐\n                      │   root centroid   │\n                      └───────────────────┘\n                     ╱          │          ╲\n                    ╱           │           ╲\n┌───────────────────┐ ┌───────────────────┐ ┌───────────────────┐\n│   leaf centroid   │ │   leaf centroid   │ │   leaf centroid   │\n└───────────────────┘ └───────────────────┘ └───────────────────┘\n      ╱       ╲             ╱       ╲             ╱       ╲\n     ╱         ╲           ╱         ╲           ╱         ╲\n┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐\n│ vector │ │ vector │ │ vector │ │ vector │ │ vector │ │ vector │\n└────────┘ └────────┘ └────────┘ └────────┘ └────────┘ └────────┘\n┌───────────────┐\n      │ root centroid │\n      └───────────────┘\n        ╱     │     ╲\n       ╱      │      ╲\n┌────────┐┌────────┐┌────────┐\n│  leaf  ││  leaf  ││  leaf  │\n│centroid││centroid││centroid│\n└────────┘└────────┘└────────┘\n   ╱  ╲      ╱  ╲      ╱  ╲\n┌───┐┌───┐┌───┐┌───┐┌───┐┌───┐\n│vec││vec││vec││vec││vec││vec│\n└───┘└───┘└───┘└───┘└───┘└───┘\n```\n\nWe implemented this on top of a storage layer presenting as a key-value map,\nwith sorted and unique keys. Each cluster is given a `ClusterId`, and vectors\nwithin each cluster are given a dense `LocalId`.\n\n```\n// leaf vectors\nK::Vector(C0L0) = vec![0.45, 0.32, ...]\nK::Id(C0L0) = 7\nK::Vector(C0L1) = vec![-0.28, 0.96, ...]\nK::Id(C0L1) = 13\n\n// cluster centroid for C0 is itself clustered at the next level of the tree\nK::Vector(C1L4) = vec![0.64, -0.48, ...]\nK::Id(C1L4) = C0\n```\n\nAs you can see above, everything is keyed by `ClusterId` and `LocalId` (e.g.\n`C0L1`), which together we call the **ANN address**. This is what we mean when\nwe say the ANN index is the primary index.\n\nTwo new query plans marked the informal transition from turbopuffer v1 → v2: attribute filtering and full-text search.\n\nNaturally, customers wanted to be able to add attribute values and filter vector\nsearches on them. To\n[make filtering fast and high-recall](https://turbopuffer.com/blog/native-filtering), we modeled these\nas an inverted index that maps an attribute value to the ANN address of the\ndocuments that contain it.\n\n``` php\nK::AttrIndex(\"family\", \"Alcidae\") -> vec![C0L3, C1L2, C1L3, ...]\nK::AttrIndex(\"genus\", \"Fratercula\") -> vec![C0L3, C1L2, C1L9, ...]\n```\n\nFor projections (`include_attributes`), we also stored the document attributes\nalongside the ID and the vector.\n\n```\nK::Vector(C0L0) = vec![0.45, 0.32, ...]\nK::Id(C0L0) = 7\nK::Attr(C0L0, \"family\") = \"Alcidae\"\nK::Attr(C0L0, \"genus\") = \"Fratercula\"\n```\n\nBM25 full-text search was another obvious and much-demanded query plan. Similar\nto attribute search, full-text search works by first finding the documents that\nhave the query term present (commonly called \"postings\"). For an FTS index, we\nalso include the `(term count, document length)` metadata necessary for BM25\nscoring:\n\n``` php\nK::FTS(\"description\", \"Atlantic\") -> vec![(C0L0, 2, 37), (C9L4, 1, 42), ...]\nK::Attr(C0L0, \"description\") -> \"A sharply dressed black-and-white seabird with a \\\nhuge, multicolored bill, the Atlantic Puffin is often \\\ncalled the clown of the sea. It breeds in burrows on \\\nislands in the North Atlantic, and winters at sea.\"\n```\n\nOver time, we've shipped several other index structures and query engines:\n[aggregations](https://turbopuffer.com/docs/query#aggregations),\n[regex search](https://turbopuffer.com/docs/query#param-Regex),\n[fuzzy matching](https://turbopuffer.com/docs/fts#fuzzy-matching),\n[sparse vector search](https://turbopuffer.com/docs/query#sparse-vector-search), and\n[attribute ordering](https://turbopuffer.com/docs/query#ordering-by-attributes) — all built around the\nsame vector-primary storage layout.\n\nThe ANN primary index has largely remained intact until today for one simple\nreason: it works really, really well for ANN search on object storage. On top of\nthis architecture, we've pushed vector search to\n[single indexes of 100B+ vectors serving 200 ms p99 reads at 1k+ QPS](https://turbopuffer.com/blog/ann-v3).\nAny significant change here risks introducing regressions in ANN performance.\n\nHowever, this layout holds us back from being state-of-the-art for the non-vector query shapes we support, in three main ways: storage amplification, write amplification, and limited vectorization.\n\nAs described above, turbopuffer currently puts the full contents of each document under its ANN address. When there is only one vector, the non-vector data is stored alongside the vector only once.\n\nHowever, for multi-vector representations of a document, such as document\nnesting or late interaction, this means we have to duplicate the contents for\neach vector. This is the reason for some of our more unfortunate\n[limits](https://turbopuffer.com/docs/limits#limit-max-vector-columns-per-namespace).\n\nAny time a document is inserted, updated, or deleted, SPFresh may rebalance the vectors to ensure they remain well clustered (otherwise recall may suffer). Because everything in a document is stored keyed by the ANN address of the document's vector, this rebalancing cascades to moving the full document contents, as well as any inverted (attribute and FTS) indexes that reference it. Updating just one vector can move hundreds of attributes and their indexes.\n\nThis write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.\n\nModern query engines are [vectorized](https://turbopuffer.com/blog/fts-v2-maxscore): they run tight\nloops over blocks of values, which amortizes fixed per-block costs, compresses\nbetter, keeps the CPU pipeline full, and unlocks SIMD. DuckDB, for example,\nworks in batches of 2,048 rows, ClickHouse up to ~65k, Lucene's posting blocks\nare 256 docs, and our ANN index works best with clusters of around 100–200\ndocuments. Every query plan has an optimal block size, but today they are all\nconstrained by the ANN primary index. A plan that wants blocks of thousands of\ndocuments to keep the CPU saturated is still stuck at 100–200.\n\nWe've already documented how much this matters in turbopuffer. Our first version\nof full-text search partitioned posting lists along ANN cluster boundaries, and\nthe median block held just ~1.5 postings. FTS v2\n[reworked postings](https://turbopuffer.com/blog/fts-v2-postings) into fixed blocks of ~256, and the\nindex got 10x smaller and queries got up to 20x faster. Posting lists could do\nthat because they're stored separately and point at documents, so their layout\ndoesn't have to follow the clusters. Aggregations and other scans read the\ndocuments themselves, and those are stored one block per cluster. As long as the\nANN address is the primary key, their block size is constrained to the cluster\nsize, even if they'd prefer something larger.\n\nThe solution to these problems is simple: don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.\n\nWe recently hit a major milestone: As of earlier this month, 100% of CI passes on turbopuffer v3. However, it represented a significant performance regression from production turbopuffer. That's not a surprise! So far, we've focused primarily on the foundational design and correctness, and we have just started tuning it. Watching the numbers go down is great fun, so we wanted to get you in at day zero of perf grinding.\n\nWe'll add results to [the benchmarks](https://turbopuffer.com/v3) as we write up what changed. Along\nthe way, we'll dive into detail about the new architecture and the optimizations\nthat will ultimately make turbopuffer faster for more query plans in prod.\n\nturbopuffer is a fast search engine that hosts 1T+ documents, handles 10M+ writes/s, and serves 25k+ queries/s. We are ready for far more. We hope you'll trust us with your queries.", "url": "https://wpnews.pro/news/rip-vector-database", "canonical_source": "https://turbopuffer.com/blog/rip-vector-database", "published_at": "2026-09-30 14:46:37+00:00", "updated_at": "2026-09-30 14:50:16.857319+00:00", "lang": "en", "topics": ["ai-infrastructure", "artificial-intelligence", "ai-search"], "entities": ["turbopuffer", "Cursor", "Notion", "SPANN", "SPFresh"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/rip-vector-database", "markdown": "https://wpnews.pro/news/rip-vector-database.md", "text": "https://wpnews.pro/news/rip-vector-database.txt", "jsonld": "https://wpnews.pro/news/rip-vector-database.jsonld"}}