Vector Search Is Now A Data Type, But The Money Math Has Not Gone Away Approximate nearest-neighbor vector search has collapsed into existing databases — Postgres via pgvector 0.5.0 (August 2023), Elasticsearch, OpenSearch, ClickHouse 25.8, MongoDB Atlas (GA December 2023) and Redis — rather than remaining a standalone category, but memory cost still governs the bill, with OpenSearch documenting an in-memory HNSW footprint of 1.1 × (4 × dimension + 8 × M) bytes per vector. One million 1,536-dimension float32 vectors, the output size of OpenAI's text-embedding-3-large, consume about 6.1 GB before graph overhead, and ten million require roughly 60 GB for the vectors alone. The consolidation removes an operational surface but not the physics: bytes per vector, compression aggressiveness and what falls out of RAM set cost regardless of vendor. STORE Vector Search Is Now A Data Type, But The Money Math Has Not Gone Away For a few years it looked like every AI stack would carry a dedicated vector database next to its primary store for the few who needed it . That moment has passed. Approximate nearest-neighbor search is now everywhere, and is more often than not a column type in the databases that teams already run – Postgres, Elasticsearch, OpenSearch, ClickHouse, MongoDB, Redis – and the “do I need a new database for this” question mostly answers itself: No, you don’t. What did not collapse is the cost math. Vector search as a data type means storing dense embeddings in an existing engine and querying them with approximate nearest-neighbor ANN indexes alongside the normal filters, joins, and lexical scoring that engine already does, instead of standing up a separate system. That removes an operational surface. It does not remove the physics. At scale, your bill is set by how many bytes each vector occupies in memory, how aggressively you compress them, and where you let them fall out of RAM – and almost none of that depends on which vendor’s logo is on the box. The Category Collapsed Into The Engines You Already Run The dedicated vector database was a real category with real engineering behind it. Pinecone launched its managed service in 2021 and effectively named the segment; Milvus, Weaviate, and Qdrant arrived around the same window. The pitch was sound for its time: general-purpose databases had no good ANN index, so semantic search needed a purpose-built home. The incumbents closed that gap fast. Elasticsearch shipped approximate k-Nearest Neighbor k-NN over the Hierarchical Navigable Small World HNSW graph algorithm, building on Lucene’s HNSW implementation. OpenSearch has carried a k-NN plugin since its 1.0 release. And pgvector added HNSW in 0.5.0 August 2023 , turning a stock Postgres instance into a credible vector store. MongoDB Atlas Vector Search reached general availability in December 2023. Redis exposed vector similarity in RediSearch 2.4 back in 2022. ClickHouse made its HNSW vector similarity index generally available in 25.8. The same algorithm – HNSW – shows up everywhere, because it is the one that delivers 95+ percent recall without a rebuild step. The reason this consolidation stuck is not feature parity on paper. It is that production retrieval is rarely pure vector lookup. Useful systems combine lexical scoring, structured filters tenant, language, ACL, freshness , and aggregations with dense retrieval, then often rerank the merged set. All of that wants to live next to the vectors. A separate vector database forces dual writes, a second consistency model, and a second thing to page someone about at 3 AM. When the engine you already operate can hold the embedding as one more field, the standalone system has to justify itself on more than recall – and usually can’t. What Actually Drives Cost: Memory Is The Whole Story HNSW is a graph that has to be traversed with pointer chasing, and that traversal is only fast when the graph and its vectors sit in RAM. The moment nodes spill to disk, latency is dominated by page faults rather than distance computation. So the dominant cost of vector search at scale is memory, and memory scales with the product of vector count and dimensionality. The per-vector memory footprint for an in-memory HNSW index is well defined. OpenSearch documents it as 1.1 × 4 × dimension + 8 × M bytes per vector: the 4 × dimension term is the float32 vector itself, 8 × M is the graph’s per-node connection overhead, and the 1.1 is roughly 10 percent slack. The consequence is that dimensionality, not document count alone, sets your hardware tier. The raw numbers are blunt. One million 1,536-dimension float32 vectors – the output size of OpenAI’s text-embedding-3-large – take about 6.1 GB before graph overhead. Ten million of them need roughly 60 GB just for the vectors https://bigdataboutique.com/blog/sparse-vs-dense-vectors-how-lexical-and-semantic-search-actually-work , all of it RAM-resident for predictable latency. That is the figure that turns a vector feature into a capacity-planning problem. Graph structure compounds it. HNSW stores neighbor lists at every layer, so the index can run 2-5× the size of the raw vectors https://bigdataboutique.com/blog/hnsw-vs-ivfflat-how-to-choose-the-right-vector-index , where IVFFlat’s flat per-cell storage stays near 1.1×. This is why the instinctive response to slow vector search - more shards, higher recall settings, bigger ef search – so often makes things worse. Each lever adds memory pressure, fan-out, or per-query CPU, and at scale the slowest shard sets your p99. The recall-versus-throughput curve is steep and non-linear: On standard ANN benchmarks, pushing HNSW from 95 percent to 100 percent recall can cost roughly 7× in throughput https://github.com/apache/lucene/issues/10976 . Spending the last two points of recall is usually the most expensive thing in the whole pipeline, and rarely the thing users notice. The Levers That Bend The Curve: Quantization And Tiering If memory is the cost, we can definitely use compression to reduce it. Quantization shrinks each vector’s footprint by trading a controllable amount of precision for a large reduction in bytes, usually paired with a rescoring step that re-ranks the top candidates using full-precision vectors held on cheaper storage. This is where real cost reduction happens, and it is mostly orthogonal to vendor choice. The techniques stack up by aggressiveness: The numbers behind these are concrete. Lucene’s int8 scalar quantization cuts memory by about 75 percent https://www.elastic.co/search-labs/blog/scalar-quantization-in-lucene and keeps raw vectors on disk for rescoring. Binary quantization, which collapses each dimension to a single bit, reaches 32× and works well on high-dimensional embeddings with oversampling - Qdrant reports https://qdrant.tech/articles/binary-quantization/ 0.98 recall@100 on 1536-dim OpenAI embeddings with 4× oversampling. Elasticsearch’s Better Binary Quantization takes a 138M × 1024-dim dataset from roughly 535 GB to about 19 GB https://www.elastic.co/search-labs/blog/better-binary-quantization-lucene-elasticsearch . The other lever is admitting that not all vectors deserve RAM. Disk-based ANN keeps the bulk of the index on SSD with only a compressed copy in memory: Microsoft’s DiskANN https://www.microsoft.com/en-us/research/publication/diskann-fast-accurate-billion-point-nearest-neighbor-search-on-a-single-node/ indexes a billion points on a single 64 GB machine plus SSD at 95+ percent recall, serving thousands of queries per second. Object-storage tiers push this further - Amazon’s S3 Vectors targets up to 90% lower cost https://aws.amazon.com/s3/features/vectors/ on large vector datasets by trading latency for price. The trade is real and measurable: OpenSearch’s disk-based mode showed p90 latency of 96 ms versus 24 ms https://opensearch.org/blog/reduce-cost-with-disk-based-vector-search/ for the in-memory path. For agentic retrieval inside a multi-second reasoning loop that is invisible; for interactive autocomplete it is disqualifying. The design move at billion scale vector search https://bigdataboutique.com/blog/scaling-vector-search-performance-from-millions-to-billions-8d50a1 is tiering - hot vectors in RAM-backed HNSW, cold vectors quantized on disk or in object storage, whereas some solutions like AWS S3 Vectors and Turbopuffer still allow to enjoy both worlds – large scale vector search at decent performance and low cost. Where The Self-Managed Versus Managed Line Really Falls Because the engines have converged on the same algorithm and the same compression toolkit, the self-managed versus managed decision is no longer about capability. Both can do HNSW, both can quantize, both can tier. The line falls on who owns three things: the memory-sizing math, the rescoring and latency cliff, and the routing between hot and cold tiers. Managed services earn their keep by absorbing the provisioning problem. They size the RAM, isolate vector traffic onto dedicated nodes so it does not contend with operational queries – MongoDB shipped dedicated Search Nodes https://www.mongodb.com/blog/post/dedicated-search-nodes-vector-search-now-in-general-availability alongside its Vector Search GA for exactly this reason – and hide the cold-tier fetch behind a single API. You give up access to some of the knobs and you pay a margin, but you stop owning the failure modes. Self-managed gives you every knob and the bill that comes with using them wrong. You set M, ef construction, and ef search; you choose the quantization scheme and the oversampling ratio; you decide the rebuild cadence. You also inherit the sharp edges: cold queries against an on-disk HNSW index that page-fault on every traversal step, recall that drifts https://bigdataboutique.com/blog/hnsw-vs-ivfflat-how-to-choose-the-right-vector-index as data distribution changes, and the fan-out math that makes over-sharding a tail-latency trap. The decision driver is rarely the technology. It is two questions. Does the vector data already live in an operational store you run – in which case adding a column avoids a whole second system – and what is your latency SLO against your cost ceiling? A team with a strict sub-20 millisecond interactive budget and a small ops crew is usually better served by a managed hot tier. A team that already runs Postgres or OpenSearch at scale, has the headcount to own capacity planning, and wants to control the quantization-and-tiering curve directly will get more out of self-managing it. Pick the engine for the workload, not the other way around. Here are the key takeaways: · The standalone vector database category has largely dissolved. ANN is now a native data type in Postgres pgvector , the Lucene engines, ClickHouse, MongoDB, and Redis, because production retrieval needs vectors co-located with lexical scoring and filters. · Memory is the dominant cost. An in-memory HNSW index needs about 1.1 × 4 × dimension + 8 × M bytes per vector, and 10M 1536-dim vectors run roughly 60 GB before graph overhead. · The recall-versus-throughput curve is non-linear. Chasing the last few points of recall can cost multiples in throughput; “turn the knobs up” usually backfires.· Quantization and tiering are the levers that move cost, and they are mostly vendor-independent: int8 4× , Matryoshka truncation 3-12× , binary quantization 32× , and disk/object-storage tiers up to ~90% cheaper, at a latency cost . · Self-managed versus managed is about ownership, not capability. Managed absorbs the sizing math and isolation; self-managed gives you every knob and every failure mode. Decide based on whether the data already lives in your store and on your latency budget versus cost ceiling. Itamar Syn-Hershko is a long-time reader and intelligent commenter at The Next Platform. Yes, I am aware that the new CMS we have has destroyed our comments. I am working with the team to try to fix this. Syn-Hershko was a software developer at Yayasoft back in the Great Recession, and was a co-manager and developer on the CLucene search engine project as well as a software engineer at Hibernating Rhinos, where he was the lead developer for RavenDB. He was then a senior staff software engineer for Forter, building the Elastic-based identity portion of Forter’s real-time fraud prevention system. In 2016, he was the founder and chief technology officer at BigData Boutique, which helps customers with data analytics workloads and which is still on ongoing concern, as well as being the founder in 2024 of NeverBlink AI, which helps customers modernize their databases so they can work with agentic AI applications. If you have a burr under your saddle about something and you have the technical chops, I am always happy to let you have the soapbox for a story. Reach out.