# Vector database showroom. Part 2: Chroma — The E-Scooter That Starts in 90 Second

> Source: <https://dev.to/silver_dev/vector-database-showroom-part-2-chroma-the-e-scooter-that-starts-in-90-second-23gn>
> Published: 2026-10-03 10:30:05+00:00

## 
  
  
  🛴 Chroma — The E-Scooter That Starts in 90 Seconds

`pip install chromadb` — and you're already riding. No server, no ports, no API keys.

Under the hood: an open-source embedding database (Apache 2.0, Python-first, Rust core since v0.4). Embedded mode is the default: use `PersistentClient` to write your data to SQLite + Parquet right on disk (while the basic `Client()` is ephemeral and loses data on restart — a classic first-day surprise). Need to share it with the team? One command turns the scooter into an HTTP server or a container.

## 
  
  
  Where it shines:

- 
**Three lines to your first search, embeddings included** : a local ONNX model runs by default — no OpenAI key, no ongoing network calls (the model downloads once, ~80 MB, on first run), no bills
- 
**The default of every RAG tutorial** : first-class LangChain and LlamaIndex integrations
- Metadata filters (`where` ) and document-content filters (`where_document` )
- Cosine instead of the default L2 — a single parameter
- 
**Your data is actually yours** :`collection.get()` dumps everything. Lock-in is measured in hours, not months

## 
  
  
  Where it stalls:

- 
**Single node by design** — there is no distributed mode. Single-digit millions of vectors: comfortable. Tens of millions: an expedition with duct tape
- 
**The HNSW index lives in RAM** : 1M × 1536-dim vectors ≈ 6 GB of raw data — plus roughly the same again for the graph itself (same math as pgvector)
- 
**Prototype-grade durability** : SQLite + Parquet segments on disk. A hard kill of the process can cost you the latest writes
- No BM25 hybrid, no quantization, no RBAC — invisible in a prototype, painfully visible in production

## 
  
  
  What breaks if you skip the manual:

- 
**The scooter that drove itself to production** . The classic arc: the prototype grows, "it's basically done, let's ship it" — and at 5M vectors you're looking for a cluster that doesn't exist. Migrating to a "real" database is a day's work; replanning the architecture mid-incident is the day you don't have
- 
**No alarm system** : token auth at best, TLS is DIY through a reverse proxy. "It's only reachable from the office network" — famous last words
- The storage format has changed between major versions before: read the changelog before upgrading

**Cost of ownership**: free under Apache 2.0. Chroma Cloud exists, but the scooter's honest habitat is "zero infrastructure at all."

**Mechanics & parts**: job postings for a "Chroma engineer" don't exist — because they're not needed. Any Python developer can fix a scooter with a screwdriver. The community is one of the largest among vector databases.

✅ **Take it if**: you're prototyping RAG, learning embeddings, running evals in CI, or building a personal tool that will never see the internet

❌ **Pass if**: this database is about to hold production with paying users — rent a car, not a scooter

Test drive — 90 seconds:

That's the whole vehicle. If it feels too easy — that's the point. Just remember: a scooter's job is to one day step aside for a real car.
