{"slug": "vector-search-is-now-a-data-type-but-the-money-math-has-not-gone-away", "title": "Vector Search Is Now A Data Type, But The Money Math Has Not Gone Away", "summary": "Approximate nearest-neighbor vector search has collapsed into existing databases — Postgres via pgvector 0.5.0 (August 2023), Elasticsearch, OpenSearch, ClickHouse 25.8, MongoDB Atlas (GA December 2023) and Redis — rather than remaining a standalone category, but memory cost still governs the bill, with OpenSearch documenting an in-memory HNSW footprint of 1.1 × (4 × dimension + 8 × M) bytes per vector. One million 1,536-dimension float32 vectors, the output size of OpenAI's text-embedding-3-large, consume about 6.1 GB before graph overhead, and ten million require roughly 60 GB for the vectors alone. The consolidation removes an operational surface but not the physics: bytes per vector, compression aggressiveness and what falls out of RAM set cost regardless of vendor.", "body_md": "STORE\n\n# Vector Search Is Now A Data Type, But The Money Math Has Not Gone Away\n\nFor a few years it looked like every AI stack would carry a dedicated vector database next to its primary store (for the few who needed it). That moment has passed. Approximate nearest-neighbor search is now everywhere, and is more often than not a column type in the databases that teams already run – Postgres, Elasticsearch, OpenSearch, ClickHouse, MongoDB, Redis – and the “do I need a new database for this” question mostly answers itself: No, you don’t.\n\nWhat did not collapse is the cost math. Vector search as a data type means storing dense embeddings in an existing engine and querying them with approximate nearest-neighbor (ANN) indexes alongside the normal filters, joins, and lexical scoring that engine already does, instead of standing up a separate system. That removes an operational surface. It does not remove the physics. At scale, your bill is set by how many bytes each vector occupies in memory, how aggressively you compress them, and where you let them fall out of RAM – and almost none of that depends on which vendor’s logo is on the box.\n\n### The Category Collapsed Into The Engines You Already Run\n\nThe dedicated vector database was a real category with real engineering behind it. Pinecone launched its managed service in 2021 and effectively named the segment; Milvus, Weaviate, and Qdrant arrived around the same window. The pitch was sound for its time: general-purpose databases had no good ANN index, so semantic search needed a purpose-built home.\n\nThe incumbents closed that gap fast. Elasticsearch shipped approximate k-Nearest Neighbor (k-NN) over the Hierarchical Navigable Small World (HNSW) graph algorithm, building on Lucene’s HNSW implementation. OpenSearch has carried a k-NN plugin since its 1.0 release. And pgvector added HNSW in 0.5.0 (August 2023), turning a stock Postgres instance into a credible vector store. MongoDB Atlas Vector Search reached general availability in December 2023. Redis exposed vector similarity in RediSearch 2.4 back in 2022. ClickHouse made its HNSW vector similarity index generally available in 25.8. The same algorithm – HNSW – shows up everywhere, because it is the one that delivers 95+ percent recall without a rebuild step.\n\nThe reason this consolidation stuck is not feature parity on paper. It is that production retrieval is rarely pure vector lookup. Useful systems combine lexical scoring, structured filters (tenant, language, ACL, freshness), and aggregations with dense retrieval, then often rerank the merged set. All of that wants to live next to the vectors. A separate vector database forces dual writes, a second consistency model, and a second thing to page someone about at 3 AM. When the engine you already operate can hold the embedding as one more field, the standalone system has to justify itself on more than recall – and usually can’t.\n\n### What Actually Drives Cost: Memory Is The Whole Story\n\nHNSW is a graph that has to be traversed with pointer chasing, and that traversal is only fast when the graph and its vectors sit in RAM. The moment nodes spill to disk, latency is dominated by page faults rather than distance computation. So the dominant cost of vector search at scale is memory, and memory scales with the product of vector count and dimensionality.\n\nThe per-vector memory footprint for an in-memory HNSW index is well defined. OpenSearch documents it as 1.1 × (4 × dimension + 8 × M) bytes per vector: the 4 × dimension term is the float32 vector itself, 8 × M is the graph’s per-node connection overhead, and the 1.1 is roughly 10 percent slack. The consequence is that dimensionality, not document count alone, sets your hardware tier.\n\nThe raw numbers are blunt. One million 1,536-dimension float32 vectors – the output size of OpenAI’s text-embedding-3-large – take about 6.1 GB before graph overhead. Ten million of them need [roughly 60 GB just for the vectors](https://bigdataboutique.com/blog/sparse-vs-dense-vectors-how-lexical-and-semantic-search-actually-work), all of it RAM-resident for predictable latency. That is the figure that turns a vector feature into a capacity-planning problem.\n\nGraph structure compounds it. HNSW stores neighbor lists at every layer, so the index can run [2-5× the size of the raw vectors](https://bigdataboutique.com/blog/hnsw-vs-ivfflat-how-to-choose-the-right-vector-index), where IVFFlat’s flat per-cell storage stays near 1.1×. This is why the instinctive response to slow vector search - more shards, higher recall settings, bigger ef_search – so often makes things worse. Each lever adds memory pressure, fan-out, or per-query CPU, and at scale the slowest shard sets your p99. The recall-versus-throughput curve is steep and non-linear: On standard ANN benchmarks, pushing HNSW from 95 percent to 100 percent recall can cost roughly [7× in throughput](https://github.com/apache/lucene/issues/10976). Spending the last two points of recall is usually the most expensive thing in the whole pipeline, and rarely the thing users notice.\n\n### The Levers That Bend The Curve: Quantization And Tiering\n\nIf memory is the cost, we can definitely use compression to reduce it. Quantization shrinks each vector’s footprint by trading a controllable amount of precision for a large reduction in bytes, usually paired with a rescoring step that re-ranks the top candidates using full-precision vectors held on cheaper storage. This is where real cost reduction happens, and it is mostly orthogonal to vendor choice.\n\nThe techniques stack up by aggressiveness:\n\nThe numbers behind these are concrete. Lucene’s int8 scalar quantization cuts memory by [about 75 percent](https://www.elastic.co/search-labs/blog/scalar-quantization-in-lucene) and keeps raw vectors on disk for rescoring. Binary quantization, which collapses each dimension to a single bit, reaches 32× and works well on high-dimensional embeddings with oversampling - [Qdrant reports](https://qdrant.tech/articles/binary-quantization/) 0.98 recall@100 on 1536-dim OpenAI embeddings with 4× oversampling. Elasticsearch’s Better Binary Quantization takes a 138M × 1024-dim dataset from [roughly 535 GB to about 19 GB](https://www.elastic.co/search-labs/blog/better-binary-quantization-lucene-elasticsearch).\n\nThe other lever is admitting that not all vectors deserve RAM. Disk-based ANN keeps the bulk of the index on SSD with only a compressed copy in memory: Microsoft’s [DiskANN](https://www.microsoft.com/en-us/research/publication/diskann-fast-accurate-billion-point-nearest-neighbor-search-on-a-single-node/) indexes a billion points on a single 64 GB machine plus SSD at 95+ percent recall, serving thousands of queries per second. Object-storage tiers push this further - Amazon’s S3 Vectors targets [up to 90% lower cost](https://aws.amazon.com/s3/features/vectors/) on large vector datasets by trading latency for price. The trade is real and measurable: OpenSearch’s disk-based mode showed [p90 latency of 96 ms versus 24 ms](https://opensearch.org/blog/reduce-cost-with-disk-based-vector-search/) for the in-memory path. For agentic retrieval inside a multi-second reasoning loop that is invisible; for interactive autocomplete it is disqualifying. The design move at [billion scale vector search](https://bigdataboutique.com/blog/scaling-vector-search-performance-from-millions-to-billions-8d50a1) is tiering - hot vectors in RAM-backed HNSW, cold vectors quantized on disk or in object storage, whereas some solutions like AWS S3 Vectors and Turbopuffer still allow to enjoy both worlds – large scale vector search at decent performance and low cost.\n\n### Where The Self-Managed Versus Managed Line Really Falls\n\nBecause the engines have converged on the same algorithm and the same compression toolkit, the self-managed versus managed decision is no longer about capability. Both can do HNSW, both can quantize, both can tier. The line falls on who owns three things: the memory-sizing math, the rescoring and latency cliff, and the routing between hot and cold tiers.\n\nManaged services earn their keep by absorbing the provisioning problem. They size the RAM, isolate vector traffic onto dedicated nodes so it does not contend with operational queries – MongoDB shipped [dedicated Search Nodes](https://www.mongodb.com/blog/post/dedicated-search-nodes-vector-search-now-in-general-availability) alongside its Vector Search GA for exactly this reason – and hide the cold-tier fetch behind a single API. You give up access to some of the knobs and you pay a margin, but you stop owning the failure modes.\n\nSelf-managed gives you every knob and the bill that comes with using them wrong. You set M, ef_construction, and ef_search; you choose the quantization scheme and the oversampling ratio; you decide the rebuild cadence. You also inherit the sharp edges: cold queries against an on-disk HNSW index that page-fault on every traversal step, [recall that drifts](https://bigdataboutique.com/blog/hnsw-vs-ivfflat-how-to-choose-the-right-vector-index) as data distribution changes, and the fan-out math that makes over-sharding a tail-latency trap.\n\nThe decision driver is rarely the technology. It is two questions. Does the vector data already live in an operational store you run – in which case adding a column avoids a whole second system – and what is your latency SLO against your cost ceiling? A team with a strict sub-20 millisecond interactive budget and a small ops crew is usually better served by a managed hot tier. A team that already runs Postgres or OpenSearch at scale, has the headcount to own capacity planning, and wants to control the quantization-and-tiering curve directly will get more out of self-managing it. Pick the engine for the workload, not the other way around.\n\nHere are the key takeaways:\n\n· The standalone vector database category has largely dissolved. ANN is now a native data type in Postgres (pgvector), the Lucene engines, ClickHouse, MongoDB, and Redis, because production retrieval needs vectors co-located with lexical scoring and filters.\n\n· Memory is the dominant cost. An in-memory HNSW index needs about 1.1 × (4 × dimension + 8 × M) bytes per vector, and 10M 1536-dim vectors run roughly 60 GB before graph overhead.\n\n· The recall-versus-throughput curve is non-linear. Chasing the last few points of recall can cost multiples in throughput; “turn the knobs up” usually backfires.· Quantization and tiering are the levers that move cost, and they are mostly vendor-independent: int8 (4×), Matryoshka truncation (3-12×), binary quantization (32×), and disk/object-storage tiers (up to ~90% cheaper, at a latency cost).\n\n· Self-managed versus managed is about ownership, not capability. Managed absorbs the sizing math and isolation; self-managed gives you every knob and every failure mode. Decide based on whether the data already lives in your store and on your latency budget versus cost ceiling.\n\n*Itamar Syn-Hershko* *is a long-time reader and intelligent commenter at The Next Platform. (Yes, I am aware that the new CMS we have has destroyed our comments. I am working with the team to try to fix this.) Syn-Hershko was a software developer at Yayasoft back in the Great Recession, and was a co-manager and developer on the CLucene search engine project as well as a software engineer at Hibernating Rhinos, where he was the lead developer for RavenDB. He was then a senior staff software engineer for Forter, building the Elastic-based identity portion of Forter’s real-time fraud prevention system. In 2016, he was the founder and chief technology officer at BigData Boutique, which helps customers with data analytics workloads and which is still on ongoing concern, as well as being the founder in 2024 of NeverBlink AI, which helps customers modernize their databases so they can work with agentic AI applications.*\n\nIf you have a burr under your saddle about something and you have the technical chops, I am always happy to let you have the soapbox for a story. Reach out.", "url": "https://wpnews.pro/news/vector-search-is-now-a-data-type-but-the-money-math-has-not-gone-away", "canonical_source": "https://www.nextplatform.com/store/2026/09/22/vector-search-is-now-a-data-type-but-the-money-math-has-not-gone-away/5298393", "published_at": "2026-09-22 17:06:09+00:00", "updated_at": "2026-10-06 17:47:19.135418+00:00", "lang": "en", "topics": ["ai-infrastructure", "artificial-intelligence", "machine-learning", "large-language-models"], "entities": ["Pinecone", "Milvus", "Weaviate", "Qdrant", "Postgres", "pgvector", "Elasticsearch", "OpenSearch"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/vector-search-is-now-a-data-type-but-the-money-math-has-not-gone-away", "markdown": "https://wpnews.pro/news/vector-search-is-now-a-data-type-but-the-money-math-has-not-gone-away.md", "text": "https://wpnews.pro/news/vector-search-is-now-a-data-type-but-the-money-math-has-not-gone-away.txt", "jsonld": "https://wpnews.pro/news/vector-search-is-now-a-data-type-but-the-money-math-has-not-gone-away.jsonld"}}