Hybrid-retrieval memory layer for AI agents DeepMem, a drop-in AI memory layer compatible with the Mem0 API, claims 2× faster response times and 10× lower cost, with self-hosted options at $0 and cloud plans starting at $1.9/month compared to Mem0's $19/month. The open-source server, available on GitHub, offers hybrid retrieval (vector + BM25 + entity boost + time-decay), semantic caching, and GDPR controls, and can be migrated from Mem0 by changing one import line. Drop-in AI memory layer with 2× faster response and 10× lower cost. Fully compatible with Mem0 API. Migrate in 5 minutes - one import line. Self-hostable. No auth, no payment, no lock-in. Or use the managed cloud at deepmem.dev https://deepmem.dev . Cloud https://deepmem.dev · Self-host quick-start · Benchmarks benchmarks · Reproduce them /deepmemteam/deepmem/blob/main/benchmarks/README.md Migrate from Mem0 in one line - same MemoryClient , same method signatures: python Before - Mem0 from mem0 import MemoryClient client = MemoryClient api key="m0-..." After - DeepMem only the import changes from deepmem import MemoryClient client = MemoryClient api key="dm live-..." get a key at deepmem.dev pip install deepmem-client Turn conversations into searchable long-term memory: a FastAPI HTTP API in front of a Qdrant vector store, with LLM fact extraction, hybrid retrieval vector + BM25 + entity boost + time-decay , semantic caching, async batched distillation, GDPR controls, and a built-in MCP server. It runs in open mode - no API key, no user registration - so you can deploy it for your own agents in minutes. Multi-tenant isolation is driven by user id in the request body. Prefer not to self-host?DeepMem Cloudis the managed version of this exact engine at- same API, no infra. Sign up, grab a key deepmem.dev dm live ... , point your base URL at https://deepmem.dev , done. The cloud and the open-source server speak the same Mem0-compatible API, so client code is identical. Already using Mem0? Switch to DeepMem cloud in one line. The deepmem-client package mirrors mem0.MemoryClient - same class name, same method signatures, same filters={"user id": ...} style - so everything after the import stays untouched. pip install deepmem-client python before Mem0 from mem0 import MemoryClient client = MemoryClient api key="m0-..." client.add messages, user id="alex" client.search "What can Alex cook?", filters={"user id": "alex"} after DeepMem cloud - change one import line from deepmem import MemoryClient client = MemoryClient api key="dm live ..." key at https://deepmem.dev client.add messages, user id="alex" identical calls client.search "What can Alex cook?", filters={"user id": "alex"} Behavioral notes - it returns add infer=True the default is asynchronous on DeepMem cloud pending=True with results= and extracted facts land a few seconds later. Mem0 cloud's add is async too - it returns PENDING . Pass infer=False for synchronous raw-text storage that's immediately searchable. No graph relations - DeepMem uses hybrid vector retrieval vector + BM25- time-decay , so relations is always . Mem0's graph features aren't replicated. - time-decay , so - Mem0's is account-wide; DeepMem's is per- reset differs user id with a confirm guard. DeepMem Cloud is 10x cheaper than Mem0 at every paid tier - the same shape of plans, a tenth of the price. | Tier | DeepMem | Mem0 cloud | |---|---|---| | Hobby | Free | Free | | Starter | $1.9/mo | $19/mo | | Growth | $7.9/mo | $79/mo | | Professional | $24.9/mo | $249/mo | Self-host instead and it's $0 - you pay only your own LLM/embedding provider the same LLM cost Mem0 charges on top of its plan price , with no memory-service markup. Batched distillation also cuts LLM calls ~80%, so even your provider bill is smaller than per-message extractors. Plans and limits: deepmem.dev https://deepmem.dev · mem0.ai https://mem0.ai/pricing . No cherry-picked headline. The scripts and workload ship in /benchmarks /deepmemteam/deepmem/blob/main/benchmarks/README.md - run them yourself. Here's what we measured and the exact config that produced it: | Metric | DeepMem self-hosted ¹ | DeepMem cloud | Mem0 cloud | |---|---|---|---| Search p50 | 73 ms | 643 ms | 653 ms | Search p95 | 86 ms | 811 ms | 710 ms | Search hits 40 queries | - | 195 | 84 | Add p50 raw store | 899 ms ² | 792 ms | 695 ms ³ | ¹ BGE-M3 on a GTX 1070 GPU 2016-era , local file Qdrant, infer=False , 100 ops, concurrency 1. ² Dominated by local-file Qdrant I/O - a Qdrant server cuts this sharply. ³ Mem0 has no raw-store mode; add always runs LLM extraction, so this row isn't apples-to-apples. Self-hosted is where intrinsic latency lives - no internet RTT, your embedder, your Qdrant. 73 ms p50 search on an old consumer GPU. DeepMem cloud beats Mem0 cloud on search p50 643 ms vs 653 ms and returns ~2.3x more candidates per search 195 vs 84 hits across 40 queries . Cloud latency is RTT-dominated - both cloud columns were measured through a proxy from mainland China; run-to-run jitter is ~±10%. Runfrom a low-RTT location for your own numbers. /benchmarks Agent frameworks keep re-discovering that they need persistent, retrievable memory. The hosted options bill per call and send your data to someone else's cloud. DeepMem is the self-hostable alternative: the same Mem0-shaped API you can drop in, but it runs on your box, with your embedder, your LLM key, and your Qdrant - and the code is right here to verify it. How does DeepMem compare to other Mem0 alternatives? Most are hosted-only or layer memory on top of someone else's vector DB. DeepMem combines three things at once: it's self-hostable your data stays on your box - $0 beyond your own LLM key , MCP-native Claude Desktop / Cursor read and write memories directly as tools , and fully open-source - and the managed cloud runs the exact same engine, so cloud and self-host are one API, not two products. | Without DeepMem | With DeepMem | |---|---| | Re-explain who you are and what you're working on every session | The agent recalls identity, projects, and preferences automatically | | Lose debugging and research context between sessions | Past root causes, dead ends, and findings are recalled, so work isn't repeated | | Manually restate preferences every session | Preferences persist across sessions, agents, and projects | | Hosted memory services that bill per call and hold your data | Self-host on your infra, or use the cloud - your call, same API | Hybrid retrieval, not a knowledge graph. Search fuses vector similarity, BM25 keyword match, entity boost, and time-decay scoring. There is no temporal graph layer; if that's what you need, look at Zep. Stores preferences, not code dumps. Large fenced code blocks are stripped before LLM extraction, so the store fills with durable user/project facts instead of pasted implementations. BYOK, multi-provider. Bring your own LLM OpenAI / Anthropic / any OpenAI-compatible endpoint and embedding BGE-M3 / Google / OpenAI-compatible . MCP-native. Ships an MCP server so Claude Desktop / Cursor can read and write memories directly. Three ways to run. All speak the same Mem0-compatible API. export DEEPMEM API KEY=dm live ... from https://deepmem.dev curl https://deepmem.dev/v1/memories \ -H "Authorization: Bearer $DEEPMEM API KEY" -H "Content-Type: application/json" \ -d '{"messages": {"role":"user","content":"I am Pat, I live in Lisbon."} ,"user id":"pat","infer":false}' curl https://deepmem.dev/v1/memories/search \ -H "Authorization: Bearer $DEEPMEM API KEY" -H "Content-Type: application/json" \ -d '{"query":"Where does Pat live?","user id":"pat"}' cp .env.example .env add an LLM key docker compose up --build DeepMem HTTP :8000 + MCP :8001 + Qdrant sidecar curl http://localhost:8000/health Or pull the published image: docker pull langdeepmem/deepmem:latest docker run -p 8000:8000 -p 8001:8001 -e DEEPSEEK API KEY=sk-... langdeepmem/deepmem:latest The image exposes :8000 HTTP and . The :8001 MCP Dockerfile and docker-compose.yml cover the GPU variant CUDA torch + BGE DEVICE=cuda and BGE-M3 model-download options HF mirror, proxy, or local mount . git clone https://github.com/deepmemteam/deepmem.git && cd deepmem pip install -r requirements.txt cp .env.example .env add an LLM key + embedder config python server/start.py HTTP :8000 + MCP :8001 Write and search in three lines: python import httpx httpx.post "http://localhost:8000/v1/memories", json={"messages": {"role":"user","content":"I'm Pat, I live in Lisbon."} , "user id":"pat"} print httpx.post "http://localhost:8000/v1/memories/search", json={"query":"Where does Pat live?","user id":"pat"} .json "results" user id is optional defaults to "default" ; send different user id s to isolate end-users. infer: false stores raw text immediately test-friendly ; the default infer: true queues for LLM fact extraction. Multi-provider by config, not code. Both layers switch on env vars: | Layer | Options | Selector | |---|---|---| LLM fact extraction | OpenAI · Anthropic native SDK · any OpenAI-compatible DeepSeek / vLLM / Ollama / Groq / LM Studio | LLM PROVIDER + LLM API KEY / ANTHROPIC API KEY / DEEPSEEK API KEY | Embeddings | BGE-M3 local, GPU/CPU · Google Gemini · any OpenAI-compatible | EMBEDDING PROVIDER + BGE M3 PATH / GOOGLE API KEY / OPENAI API KEY | BGE DEVICE=auto|cpu|cuda picks GPU when available, else CPU force cpu on small-VRAM cards to avoid multi-process contention . BYOK overrides the LLM per-request. Hybrid retrieval. Vector similarity + BM25 keyword + entity boost + time-decay, fused into one score. Over-fetch, re-rank, return. Async batched distillation. Writes queue behind a silence window and are extracted in batches - ~80% fewer LLM calls than per-message extraction. Semantic cache. Repeat adds/searches hit a similarity-gated cache and return cached facts without re-embedding or re-querying Qdrant. Stores preferences, not code. extraction filter strips large fenced code blocks before LLM extraction, so the store fills with durable facts, not pasted implementations. MCP server. deepmem write / deepmem search / deepmem delete tools for Claude Desktop, Cursor, and any MCP client. GDPR. Soft-delete with retention window, hard-delete reset , SHA-256 id masking in logs, export/import for portability. | Method | Path | Description | |---|---|---| | POST | /v1/memories | Write messages; LLM-extract facts infer=false stores raw | | POST | /v1/memories/search | Semantic search vector + BM25 + entity + time-decay | | GET | /v1/memories | List all for a user id paginated | | GET | /v1/memories/{id} | Get one by ID | | PUT | /v1/memories/{id} | Update one memory's text | | DELETE | /v1/memories/{id} | Soft-delete one | | DELETE | /v1/memories | Soft-delete all for a user id GDPR | | GET | /v1/memories/{id}/history | ADD/UPDATE/DELETE audit log | | POST | /v1/reset | Hard-delete all + history needs confirm user id | | GET | /v1/export · POST /v1/import | Portable JSON export / import | | GET | /health · /ready | Liveness / readiness probes | agent id / run id optionally scope writes/reads mirrors Mem0's three-level isolation: user - agent - run . Interactive docs at /docs . { "mcpServers": { "deepmem": { "command": "python", "args": "server/mcp server.py" , "env": { "DEEPMEMORY BASE URL": "http://localhost:8000" } } } } Open mode needs no API key - DEEPMEMORY API KEY stays empty. php client HTTP / MCP - FastAPI :8000 - rate-limit middleware per-IP token bucket on add/search - SemanticCache.check return cached facts on similarity hit - AsyncBatchDistiller.enqueue POST /v1/memories, infer=true ↳ silence-window or max batch triggers on batch ready ↳ VectorStore.process batch LLM extraction - Qdrant upsert - VectorStore.search vector + BM25 + entity + time-decay - LLM: OpenAI / Anthropic / OpenAI-compatible fact extraction - BGE-M3 / Gemini / OpenAI-compatible embeddings - Qdrant vectors + SQLite audit history - MCP server :8001 deepmem write / deepmem search / deepmem delete Single Qdrant collection memories , hard-filtered by user id payload. TenantValidator NFC-normalizes and enforces a A-Za-z0-9. :- {1,256} charset on user id - never trust the raw request value. Config loads once from env vars .env , auto-loaded config.json defaults. Key variables: | Variable | Required | Description | |---|---|---| DEEPSEEK API KEY / LLM API KEY / ANTHROPIC API KEY | one LLM key | LLM for fact extraction | LLM PROVIDER | no | auto / openai / anthropic / openai compatible default auto | EMBEDDING PROVIDER | no | bge-m3 / google / openai default bge-m3 | BGE M3 PATH | no | local BGE-M3 dir or HF model id default BAAI/bge-m3 | BGE DEVICE | no | auto / cpu / cuda default auto | QDRANT URL / QDRANT API KEY | no | remote Qdrant; omit for local file store | CORS ORIGINS | yes | comma-separated allowed origins no in prod | RATE LIMIT ADD / RATE LIMIT SEARCH | no | per-minute limits default 30 / 60 | Backends auto-switch on env vars - no code changes: QDRANT URL set - remote Qdrant; unset - local file Qdrant under ./data/qdrant . CORS ORIGINS= is refused at boot unless DEEPMEMORY DEBUG=1 . php systemctl restart deepmem.service scripts/start.sh - uvicorn :8000 journalctl -u deepmem -f For HTTPS, put Caddy or Nginx in front; scripts/start.sh runs under systemd or any process manager. Two reproducible scripts in /benchmarks /deepmemteam/deepmem/blob/main/benchmarks/README.md : Cloud vs cloud - DeepMem cloud vs Mem0 cloud. Register keys at deepmem.dev https://deepmem.dev + mem0.ai https://mem0.ai , then python benchmarks/benchmark cloud.py . Self-hosted - your DeepMem, your hardware no key, no external service : python benchmarks/run benchmark.py . Both ship a self-contained workload and report P50/P95/P99 + throughput. The numbers at the top of this README were produced with these scripts - rerun them and read your own percentiles. Do I need the cloud? No. The open-source server is fully functional on its own. The cloud deepmem.dev https://deepmem.dev is the zero-ops option - same API. Does it work offline? Retrieval and raw-store infer=false work with no network. LLM fact extraction infer=true needs an LLM key - or run a local OpenAI-compatible model Ollama / vLLM / LM Studio and point LLM BASE URL at it. Where is my data? In your Qdrant local file or server + a SQLite audit log. Nothing leaves your machine except the LLM/embedding calls you configure. Multi-tenant? Yes - single Qdrant collection hard-filtered by user id . agent id / run id add agent and session scope. Is it production-ready? Used in production under systemd with a remote Qdrant. Local file Qdrant is fine for dev/single-worker; use a Qdrant server for multi-worker or high-throughput. pip install -r requirements.txt DEEPMEMORY DEBUG=1 python -m uvicorn server.main:app --reload --host 0.0.0.0 --port 8000 pytest tests/ -x -v MIT /deepmemteam/deepmem/blob/main/LICENSE .