cd /news/ai-infrastructure/vector-database-showroom-part-3-qdra… · home › topics › ai-infrastructure › article
[ARTICLE · art-147844] src=dev.to ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Vector database showroom. Part 3 Qdrant — The Hot Hatchback That Carries More Than It Looks

A developer's hands-on review of the Qdrant vector database details its filterable HNSW index, which builds payload filters into graph traversal rather than applying them after search, along with out-of-the-box hybrid dense-plus-sparse queries via the Query API and int8/binary quantization that shrinks vectors 4x and 32x respectively. The writeup warns that the distance function is locked at collection creation, that missing payload indexes degrade filtered search to scans, and that segments below the optimizer's indexing_threshold skip HNSW entirely, so small-collection benchmarks measure brute-force exact search rather than HNSW. It concludes Qdrant suits filter-heavy RAG at millions to hundreds of millions of vectors on self-hosted nodes, while noting manual snapshots, no point-in-time recovery, and real operational work for distributed mode.

by read3 min views1 publishedOct 8, 2026

#

🏎️ Qdrant — The Hot Hatchback That Carries More Than It Looks

Sharp acceleration, precise handling — and a trunk noticeably bigger than it looks from the outside.

Under the hood: a Rust core shipped as a single binary or container (the Python client even has an embedded/local mode for experiments). Apache 2.0. The signature feature — filterable HNSW: the graph is built with payload indexes in mind, so filters participate in graph traversal instead of being glued on top.

Where it shines:

  • Remember pgvector's three filter tools? Here filtering is part of the index itself: "similar, but only year=2024" is this car's home turf

Hybrid out of the box: sparse vectors since v1.7, the Query API with prefetch and fusion (RRF/DBSF) since v1.10 — dense + sparse in a single request #

Quantization is our "lightweight alloys": int8 shrinks vectors 4×, binary 32×. Originals can live on disk while only compressed copies stay in RAM. Hundreds of millions of vectors on a single node is a real setup, not marketing #

Multivectors (ColBERT-style): * several vectors per point, late interaction — something very few databases in this showroom offer Operations: Prometheus metrics, snapshots, API keys + JWT RBAC

Where it stalls:

Bring your own embeddings — it's a database, not an embedder (contrast with last issue's scooter). Clients for Python/JS/Go/Rust, LangChain/LlamaIndex integrations included

  • Distributed mode (sharding, replication, Raft) is real ops work — simpler than the freight train arriving later in this series, heavier than Postgres
  • Backups are manual snapshots. No WAL, no point-in-time recovery like the one you're used to in Postgres
  • Tuning is part of the contract: m, ef_construct, hnsw_ef, quantization with rescoring. Defaults are decent, but a hot hatch likes to be tuned

What breaks if you skip the manual:

The distance function (COSINE/DOT/EUCLID/MANHATTAN) is locked at collection creation. Wrong pick = full reindex — there is no "alter on the fly" #

Payload indexes: forget to create them and filtered search degrades to scans. The right order: collection → payload indexes → bulk upload. Adding an index afterwards rebuilds the HNSW graph in the background to wire up the filterable edges #

Benchmark trap #1: while a segment stays below the optimizer'sindexing_threshold , no HNSW index is built for it — search falls back to brute force, fast and with perfect recall. A naive benchmark on a small collection measures exact search, not HNSW. Honest numbers require production-like volume — or explicitly loweringindexing_threshold #

Quantization without rescoring eats recall — and binary quantization is model-sensitive: brilliant on some embeddings, poor on others. Ground truth for comparison: search with exact=true. Showroom thesis: vendor benchmarks are ads — verify on your own data #

Named vectors: if the collection was created with named vectors, queries must pass the vector name via the using parameter — otherwise a400

Cost of ownership: free under Apache 2.0; Qdrant Cloud (managed & hybrid) exists. RAM appetite is flexible — quantization + on-disk originals instead of "6 GB per million."

Mechanics & parts: Qdrant GmbH (Berlin) behind it; the docs are among the best in the industry. "Qdrant engineers" are rare on the market, but this is a database a backend developer can drive — not a twenty-pod zoo.

✅ Take it if: millions to hundreds of millions of vectors, filter-heavy RAG, self-hosting, and you want speed without building railroads

❌ Pass if: your logic lives in SQL and you need joins (the station wagon does those better), or you're counting billions with GPUs — the train arrives later in this series

Test drive:

A hatchback doesn't pretend to be a station wagon and doesn't build railroads — it just carries fast. And when your data outgrows the trunk, the API doesn't change: you move to a cluster via snapshots, and the application never notices.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @qdrant 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/vector-database-show…] indexed:0 read:3min 2026-10-08 · —