#
🏎️ Qdrant — The Hot Hatchback That Carries More Than It Looks
Sharp acceleration, precise handling — and a trunk noticeably bigger than it looks from the outside.
Under the hood: a Rust core shipped as a single binary or container (the Python client even has an embedded/local mode for experiments). Apache 2.0. The signature feature — filterable HNSW: the graph is built with payload indexes in mind, so filters participate in graph traversal instead of being glued on top.
Where it shines:
- Remember pgvector's three filter tools? Here filtering is part of the index itself: "similar, but only year=2024" is this car's home turf
Hybrid out of the box: sparse vectors since v1.7, the Query API with prefetch and fusion (RRF/DBSF) since v1.10 — dense + sparse in a single request #
Quantization is our "lightweight alloys": int8 shrinks vectors 4×, binary 32×. Originals can live on disk while only compressed copies stay in RAM. Hundreds of millions of vectors on a single node is a real setup, not marketing #
Multivectors (ColBERT-style): * several vectors per point, late interaction — something very few databases in this showroom offer Operations: Prometheus metrics, snapshots, API keys + JWT RBAC
Where it stalls:
Bring your own embeddings — it's a database, not an embedder (contrast with last issue's scooter). Clients for Python/JS/Go/Rust, LangChain/LlamaIndex integrations included
- Distributed mode (sharding, replication, Raft) is real ops work — simpler than the freight train arriving later in this series, heavier than Postgres
- Backups are manual snapshots. No WAL, no point-in-time recovery like the one you're used to in Postgres
- Tuning is part of the contract: m, ef_construct, hnsw_ef, quantization with rescoring. Defaults are decent, but a hot hatch likes to be tuned
What breaks if you skip the manual:
The distance function (COSINE/DOT/EUCLID/MANHATTAN) is locked at collection creation. Wrong pick = full reindex — there is no "alter on the fly" #
Payload indexes: forget to create them and filtered search degrades to scans. The right order: collection → payload indexes → bulk upload. Adding an index afterwards rebuilds the HNSW graph in the background to wire up the filterable edges #
Benchmark trap #1: while a segment stays below the optimizer'sindexing_threshold , no HNSW index is built for it — search falls back to brute force, fast and with perfect recall. A naive benchmark on a small collection measures exact search, not HNSW. Honest numbers require production-like volume — or explicitly loweringindexing_threshold #
Quantization without rescoring eats recall — and binary quantization is model-sensitive: brilliant on some embeddings, poor on others. Ground truth for comparison: search with exact=true. Showroom thesis: vendor benchmarks are ads — verify on your own data #
Named vectors: if the collection was created with named vectors, queries must pass the vector name via the using parameter — otherwise a400
Cost of ownership: free under Apache 2.0; Qdrant Cloud (managed & hybrid) exists. RAM appetite is flexible — quantization + on-disk originals instead of "6 GB per million."
Mechanics & parts: Qdrant GmbH (Berlin) behind it; the docs are among the best in the industry. "Qdrant engineers" are rare on the market, but this is a database a backend developer can drive — not a twenty-pod zoo.
✅ Take it if: millions to hundreds of millions of vectors, filter-heavy RAG, self-hosting, and you want speed without building railroads
❌ Pass if: your logic lives in SQL and you need joins (the station wagon does those better), or you're counting billions with GPUs — the train arrives later in this series
Test drive:
A hatchback doesn't pretend to be a station wagon and doesn't build railroads — it just carries fast. And when your data outgrows the trunk, the API doesn't change: you move to a cluster via snapshots, and the application never notices.