{"slug": "hybrid-retrieval-memory-layer-for-ai-agents", "title": "Hybrid-retrieval memory layer for AI agents", "summary": "DeepMem, a drop-in AI memory layer compatible with the Mem0 API, claims 2× faster response times and 10× lower cost, with self-hosted options at $0 and cloud plans starting at $1.9/month compared to Mem0's $19/month. The open-source server, available on GitHub, offers hybrid retrieval (vector + BM25 + entity boost + time-decay), semantic caching, and GDPR controls, and can be migrated from Mem0 by changing one import line.", "body_md": "**Drop-in AI memory layer with 2× faster response and 10× lower cost.\nFully compatible with Mem0 API. Migrate in 5 minutes - one import line.**\n\nSelf-hostable. No auth, no payment, no lock-in. Or use the managed cloud at [deepmem.dev](https://deepmem.dev).\n\n[Cloud](https://deepmem.dev)\n·\n[Self-host](#quick-start)\n·\n[Benchmarks](#benchmarks)\n·\n[Reproduce them](/deepmemteam/deepmem/blob/main/benchmarks/README.md)\n\n**Migrate from Mem0 in one line** - same `MemoryClient`\n\n, same method signatures:\n\n``` python\n# Before - Mem0\nfrom mem0 import MemoryClient\nclient = MemoryClient(api_key=\"m0-...\")\n\n# After - DeepMem (only the import changes)\nfrom deepmem import MemoryClient\nclient = MemoryClient(api_key=\"dm_live-...\")   # get a key at deepmem.dev\npip install deepmem-client\n```\n\nTurn conversations into searchable long-term memory: a FastAPI HTTP API in\nfront of a Qdrant vector store, with LLM fact extraction, hybrid retrieval\n(vector + BM25 + entity boost + time-decay), semantic caching, async batched\ndistillation, GDPR controls, and a built-in MCP server. It runs in **open\nmode** - no API key, no user registration - so you can deploy it for your own\nagents in minutes. Multi-tenant isolation is driven by `user_id`\n\nin the\nrequest body.\n\nPrefer not to self-host?DeepMem Cloudis the managed version of this exact engine at- same API, no infra. Sign up, grab a key ([deepmem.dev]`dm_live_...`\n\n), point your base URL at`https://deepmem.dev`\n\n, done. The cloud and the open-source server speak the same Mem0-compatible API, so client code is identical.\n\nAlready using Mem0? Switch to DeepMem cloud in one line. The `deepmem-client`\n\npackage mirrors `mem0.MemoryClient`\n\n- same class name, same method signatures,\nsame `filters={\"user_id\": ...}`\n\nstyle - so everything after the import stays\nuntouched.\n\n```\npip install deepmem-client\npython\n# before (Mem0)\nfrom mem0 import MemoryClient\nclient = MemoryClient(api_key=\"m0-...\")\nclient.add(messages, user_id=\"alex\")\nclient.search(\"What can Alex cook?\", filters={\"user_id\": \"alex\"})\n\n# after (DeepMem cloud) - change one import line\nfrom deepmem import MemoryClient\nclient = MemoryClient(api_key=\"dm_live_...\")        # key at https://deepmem.dev\nclient.add(messages, user_id=\"alex\")                # identical calls\nclient.search(\"What can Alex cook?\", filters={\"user_id\": \"alex\"})\n```\n\n## Behavioral notes\n\n- it returns`add(infer=True)`\n\n(the default) is asynchronous on DeepMem cloud`pending=True`\n\nwith`results=[]`\n\nand extracted facts land a few seconds later. (Mem0 cloud's`add`\n\nis async too - it returns`PENDING`\n\n.) Pass`infer=False`\n\nfor synchronous raw-text storage that's immediately searchable.**No graph relations**- DeepMem uses hybrid vector retrieval (vector + BM25- time-decay), so\n`relations`\n\nis always`[]`\n\n. Mem0's graph features aren't replicated.\n\n- time-decay), so\n- Mem0's is account-wide; DeepMem's is per-`reset`\n\ndiffers`user_id`\n\nwith a confirm guard.\n\nDeepMem Cloud is **10x cheaper than Mem0** at every paid tier - the same shape\nof plans, a tenth of the price.\n\n| Tier | DeepMem | Mem0 cloud |\n|---|---|---|\n| Hobby | Free | Free |\n| Starter | $1.9/mo |\n$19/mo |\n| Growth | $7.9/mo |\n$79/mo |\n| Professional | $24.9/mo |\n$249/mo |\n\nSelf-host instead and it's **$0** - you pay only your own LLM/embedding\nprovider (the same LLM cost Mem0 charges on top of its plan price), with no\nmemory-service markup. Batched distillation also cuts LLM calls ~80%, so even\nyour provider bill is smaller than per-message extractors.\n\nPlans and limits: [deepmem.dev](https://deepmem.dev) · [mem0.ai](https://mem0.ai/pricing).\n\nNo cherry-picked headline. The scripts and workload ship in\n[ /benchmarks](/deepmemteam/deepmem/blob/main/benchmarks/README.md) - run them yourself. Here's what we\nmeasured and the exact config that produced it:\n\n| Metric | DeepMem self-hosted ¹ | DeepMem cloud | Mem0 cloud |\n|---|---|---|---|\nSearch p50 |\n73 ms |\n643 ms |\n653 ms |\nSearch p95 |\n86 ms | 811 ms | 710 ms |\nSearch hits (40 queries) |\n- | 195 |\n84 |\nAdd p50 (raw store) |\n899 ms ² | 792 ms | 695 ms ³ |\n\n¹ BGE-M3 on a GTX 1070 GPU (2016-era), local file Qdrant,\n\n`infer=False`\n\n, 100 ops, concurrency 1. ² Dominated by local-file Qdrant I/O - a Qdrant server cuts this sharply. ³ Mem0 has no raw-store mode;`add`\n\nalways runs LLM extraction, so this row isn't apples-to-apples.\n\n**Self-hosted is where intrinsic latency lives**- no internet RTT, your embedder, your Qdrant. 73 ms p50 search on an old consumer GPU.** DeepMem cloud beats Mem0 cloud on search p50**(643 ms vs 653 ms) and returns**~2.3x more candidates per search**(195 vs 84 hits across 40 queries).** Cloud latency is RTT-dominated**- both cloud columns were measured through a proxy from mainland China; run-to-run jitter is ~±10%. Runfrom a low-RTT location for your own numbers.`/benchmarks`\n\nAgent frameworks keep re-discovering that they need persistent, retrievable memory. The hosted options bill per call and send your data to someone else's cloud. DeepMem is the self-hostable alternative: the same Mem0-shaped API you can drop in, but it runs on your box, with your embedder, your LLM key, and your Qdrant - and the code is right here to verify it.\n\n**How does DeepMem compare to other Mem0 alternatives?** Most are hosted-only or layer memory on top of someone else's vector DB. DeepMem combines three things at once: it's **self-hostable** (your data stays on your box - $0 beyond your own LLM key), **MCP-native** (Claude Desktop / Cursor read and write memories directly as tools), and **fully open-source** - and the managed cloud runs the exact same engine, so cloud and self-host are one API, not two products.\n\n| Without DeepMem | With DeepMem |\n|---|---|\n| Re-explain who you are and what you're working on every session | The agent recalls identity, projects, and preferences automatically |\n| Lose debugging and research context between sessions | Past root causes, dead ends, and findings are recalled, so work isn't repeated |\n| Manually restate preferences every session | Preferences persist across sessions, agents, and projects |\n| Hosted memory services that bill per call and hold your data | Self-host on your infra, or use the cloud - your call, same API |\n\n**Hybrid retrieval, not a knowledge graph.** Search fuses vector similarity, BM25 keyword match, entity boost, and time-decay scoring. There is no temporal graph layer; if that's what you need, look at Zep.**Stores preferences, not code dumps.** Large fenced code blocks are stripped before LLM extraction, so the store fills with durable user/project facts instead of pasted implementations.**BYOK, multi-provider.** Bring your own LLM (OpenAI / Anthropic / any OpenAI-compatible endpoint) and embedding (BGE-M3 / Google / OpenAI-compatible).**MCP-native.** Ships an MCP server so Claude Desktop / Cursor can read and write memories directly.\n\nThree ways to run. All speak the same Mem0-compatible API.\n\n```\nexport DEEPMEM_API_KEY=dm_live_...      # from https://deepmem.dev\ncurl https://deepmem.dev/v1/memories \\\n  -H \"Authorization: Bearer $DEEPMEM_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"messages\":[{\"role\":\"user\",\"content\":\"I am Pat, I live in Lisbon.\"}],\"user_id\":\"pat\",\"infer\":false}'\ncurl https://deepmem.dev/v1/memories/search \\\n  -H \"Authorization: Bearer $DEEPMEM_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"query\":\"Where does Pat live?\",\"user_id\":\"pat\"}'\ncp .env.example .env          # add an LLM key\ndocker compose up --build     # DeepMem (HTTP :8000 + MCP :8001) + Qdrant sidecar\ncurl http://localhost:8000/health\n```\n\nOr pull the published image:\n\n```\ndocker pull langdeepmem/deepmem:latest\ndocker run -p 8000:8000 -p 8001:8001 -e DEEPSEEK_API_KEY=sk-... langdeepmem/deepmem:latest\n```\n\nThe image exposes ** :8000 (HTTP)** and\n\n**. The**\n\n`:8001`\n\n(MCP)`Dockerfile`\n\nand `docker-compose.yml`\n\ncover the GPU variant (CUDA torch + `BGE_DEVICE=cuda`\n\n)\nand BGE-M3 model-download options (HF mirror, proxy, or local mount).\n\n```\ngit clone https://github.com/deepmemteam/deepmem.git && cd deepmem\npip install -r requirements.txt\ncp .env.example .env          # add an LLM key + embedder config\npython server/start.py        # HTTP :8000 + MCP :8001\n```\n\nWrite and search in three lines:\n\n``` python\nimport httpx\nhttpx.post(\"http://localhost:8000/v1/memories\",\n    json={\"messages\":[{\"role\":\"user\",\"content\":\"I'm Pat, I live in Lisbon.\"}],\n          \"user_id\":\"pat\"})\nprint(httpx.post(\"http://localhost:8000/v1/memories/search\",\n    json={\"query\":\"Where does Pat live?\",\"user_id\":\"pat\"}).json()[\"results\"])\n```\n\n`user_id`\n\nis optional (defaults to`\"default\"`\n\n); send different`user_id`\n\ns to isolate end-users.`infer: false`\n\nstores raw text immediately (test-friendly); the default`infer: true`\n\nqueues for LLM fact extraction.\n\n**Multi-provider by config, not code.** Both layers switch on env vars:\n\n| Layer | Options | Selector |\n|---|---|---|\nLLM (fact extraction) |\nOpenAI · Anthropic (native SDK) · any OpenAI-compatible (DeepSeek / vLLM / Ollama / Groq / LM Studio) | `LLM_PROVIDER` + `LLM_API_KEY` / `ANTHROPIC_API_KEY` / `DEEPSEEK_API_KEY` |\nEmbeddings |\nBGE-M3 (local, GPU/CPU) · Google Gemini · any OpenAI-compatible | `EMBEDDING_PROVIDER` + `BGE_M3_PATH` / `GOOGLE_API_KEY` / `OPENAI_API_KEY` |\n\n`BGE_DEVICE=auto|cpu|cuda`\n\npicks GPU when available, else CPU (force `cpu`\n\non small-VRAM cards to avoid multi-process contention). BYOK overrides the\nLLM per-request.\n\n**Hybrid retrieval.** Vector similarity + BM25 keyword + entity boost +\ntime-decay, fused into one score. Over-fetch, re-rank, return.\n\n**Async batched distillation.** Writes queue behind a silence window and are\nextracted in batches - ~80% fewer LLM calls than per-message extraction.\n\n**Semantic cache.** Repeat adds/searches hit a similarity-gated cache and\nreturn cached facts without re-embedding or re-querying Qdrant.\n\n**Stores preferences, not code.** `extraction_filter`\n\nstrips large fenced\ncode blocks before LLM extraction, so the store fills with durable facts,\nnot pasted implementations.\n\n**MCP server.** `deepmem_write`\n\n/ `deepmem_search`\n\n/ `deepmem_delete`\n\ntools\nfor Claude Desktop, Cursor, and any MCP client.\n\n**GDPR.** Soft-delete with retention window, hard-delete `reset`\n\n, SHA-256\nid masking in logs, export/import for portability.\n\n| Method | Path | Description |\n|---|---|---|\n| POST | `/v1/memories` |\nWrite messages; LLM-extract facts (`infer=false` stores raw) |\n| POST | `/v1/memories/search` |\nSemantic search (vector + BM25 + entity + time-decay) |\n| GET | `/v1/memories` |\nList all for a `user_id` (paginated) |\n| GET | `/v1/memories/{id}` |\nGet one by ID |\n| PUT | `/v1/memories/{id}` |\nUpdate one memory's text |\n| DELETE | `/v1/memories/{id}` |\nSoft-delete one |\n| DELETE | `/v1/memories` |\nSoft-delete all for a `user_id` (GDPR) |\n| GET | `/v1/memories/{id}/history` |\nADD/UPDATE/DELETE audit log |\n| POST | `/v1/reset` |\nHard-delete all + history (needs `confirm_user_id` ) |\n| GET | `/v1/export` · POST `/v1/import` |\nPortable JSON export / import |\n| GET | `/health` · `/ready` |\nLiveness / readiness probes |\n\n`agent_id`\n\n/ `run_id`\n\noptionally scope writes/reads (mirrors Mem0's three-level\nisolation: user -> agent -> run). Interactive docs at `/docs`\n\n.\n\n```\n{\n  \"mcpServers\": {\n    \"deepmem\": {\n      \"command\": \"python\",\n      \"args\": [\"server/mcp_server.py\"],\n      \"env\": { \"DEEPMEMORY_BASE_URL\": \"http://localhost:8000\" }\n    }\n  }\n}\n```\n\nOpen mode needs no API key - `DEEPMEMORY_API_KEY`\n\nstays empty.\n\n``` php\nclient (HTTP / MCP)\n  -> FastAPI (:8000)\n       -> rate-limit middleware (per-IP token bucket on add/search)\n       -> SemanticCache.check            (return cached facts on similarity hit)\n       -> AsyncBatchDistiller.enqueue    (POST /v1/memories, infer=true)\n            ↳ silence-window or max_batch triggers on_batch_ready\n                ↳ VectorStore.process_batch  (LLM extraction -> Qdrant upsert)\n       -> VectorStore.search             (vector + BM25 + entity + time-decay)\n  -> LLM: OpenAI / Anthropic / OpenAI-compatible (fact extraction)\n  -> BGE-M3 / Gemini / OpenAI-compatible (embeddings)\n  -> Qdrant (vectors)  +  SQLite (audit history)\n  -> MCP server (:8001)  deepmem_write / deepmem_search / deepmem_delete\n```\n\nSingle Qdrant collection (`memories`\n\n), hard-filtered by `user_id`\n\npayload.\n`TenantValidator`\n\nNFC-normalizes and enforces a `[A-Za-z0-9._:-]{1,256}`\n\ncharset on `user_id`\n\n- never trust the raw request value.\n\nConfig loads once from env vars (`.env`\n\n, auto-loaded) > `config.json`\n\n> defaults.\nKey variables:\n\n| Variable | Required | Description |\n|---|---|---|\n`DEEPSEEK_API_KEY` / `LLM_API_KEY` / `ANTHROPIC_API_KEY` |\none LLM key | LLM for fact extraction |\n`LLM_PROVIDER` |\nno | `auto` / `openai` / `anthropic` / `openai_compatible` (default `auto` ) |\n`EMBEDDING_PROVIDER` |\nno | `bge-m3` / `google` / `openai` (default `bge-m3` ) |\n`BGE_M3_PATH` |\nno | local BGE-M3 dir or HF model id (default `BAAI/bge-m3` ) |\n`BGE_DEVICE` |\nno | `auto` / `cpu` / `cuda` (default `auto` ) |\n`QDRANT_URL` / `QDRANT_API_KEY` |\nno | remote Qdrant; omit for local file store |\n`CORS_ORIGINS` |\nyes | comma-separated allowed origins (no `*` in prod) |\n`RATE_LIMIT_ADD` / `RATE_LIMIT_SEARCH` |\nno | per-minute limits (default 30 / 60) |\n\nBackends auto-switch on env vars - no code changes:\n\n`QDRANT_URL`\n\nset -> remote Qdrant; unset -> local file Qdrant under`./data/qdrant`\n\n.`CORS_ORIGINS=*`\n\nis refused at boot unless`DEEPMEMORY_DEBUG=1`\n\n.\n\n``` php\nsystemctl restart deepmem.service     # scripts/start.sh -> uvicorn :8000\njournalctl -u deepmem -f\n```\n\nFor HTTPS, put Caddy or Nginx in front; `scripts/start.sh`\n\nruns under systemd\nor any process manager.\n\nTwo reproducible scripts in [ /benchmarks](/deepmemteam/deepmem/blob/main/benchmarks/README.md):\n\n**Cloud vs cloud**- DeepMem cloud vs Mem0 cloud. Register keys at[deepmem.dev](https://deepmem.dev)+[mem0.ai](https://mem0.ai), then`python benchmarks/benchmark_cloud.py`\n\n.**Self-hosted**- your DeepMem, your hardware (no key, no external service):`python benchmarks/run_benchmark.py`\n\n.\n\nBoth ship a self-contained workload and report P50/P95/P99 + throughput. The numbers at the top of this README were produced with these scripts - rerun them and read your own percentiles.\n\n**Do I need the cloud?** No. The open-source server is fully functional on its\nown. The cloud ([deepmem.dev](https://deepmem.dev)) is the zero-ops option -\nsame API.\n\n**Does it work offline?** Retrieval and raw-store (`infer=false`\n\n) work with no\nnetwork. LLM fact extraction (`infer=true`\n\n) needs an LLM key - or run a local\nOpenAI-compatible model (Ollama / vLLM / LM Studio) and point `LLM_BASE_URL`\n\nat it.\n\n**Where is my data?** In your Qdrant (local file or server) + a SQLite audit\nlog. Nothing leaves your machine except the LLM/embedding calls you configure.\n\n**Multi-tenant?** Yes - single Qdrant collection hard-filtered by `user_id`\n\n.\n`agent_id`\n\n/ `run_id`\n\nadd agent and session scope.\n\n**Is it production-ready?** Used in production under systemd with a remote\nQdrant. Local file Qdrant is fine for dev/single-worker; use a Qdrant server\nfor multi-worker or high-throughput.\n\n```\npip install -r requirements.txt\nDEEPMEMORY_DEBUG=1 python -m uvicorn server.main:app --reload --host 0.0.0.0 --port 8000\npytest tests/ -x -v\n```\n\n[MIT](/deepmemteam/deepmem/blob/main/LICENSE).", "url": "https://wpnews.pro/news/hybrid-retrieval-memory-layer-for-ai-agents", "canonical_source": "https://github.com/deepmemteam/deepmem", "published_at": "2026-08-12 11:14:59+00:00", "updated_at": "2026-08-12 11:41:30.847695+00:00", "lang": "en", "topics": ["ai-tools", "ai-infrastructure", "ai-agents", "developer-tools"], "entities": ["DeepMem", "Mem0", "Qdrant", "FastAPI", "MCP", "deepmem.dev"], "alternates": {"html": "https://wpnews.pro/news/hybrid-retrieval-memory-layer-for-ai-agents", "markdown": "https://wpnews.pro/news/hybrid-retrieval-memory-layer-for-ai-agents.md", "text": "https://wpnews.pro/news/hybrid-retrieval-memory-layer-for-ai-agents.txt", "jsonld": "https://wpnews.pro/news/hybrid-retrieval-memory-layer-for-ai-agents.jsonld"}}