We achieved 94.7% in Longmemeval with no hacks lol, everything open sourced Swafra, an open-source semantic memory system for AI assistants, achieved 94.7% recall_all@10 on the LongMemEval benchmark for long-term memory. The tool works as an MCP server with Claude Desktop, Claude Code, and VS Code Copilot, using local embeddings and hybrid retrieval without requiring a database or cloud service. Semantic memory for AI — ingest anything, retrieve what matters. 94.7% recall all@10 on LongMemEval — the standard benchmark for long-term memory in AI assistants. Works as an MCP server with Claude Desktop, Claude Code, VS Code Copilot, and any MCP-compatible AI. pip install swafra Or with Node.js: npm install -g swafra Add to ~/Library/Application Support/Claude/claude desktop config.json : { "mcpServers": { "swafra": { "command": "swafra" } } } Restart Claude Desktop — the tools appear automatically. claude mcp add swafra swafra Add to .vscode/mcp.json in your project: { "servers": { "swafra": { "command": "swafra" } } } Once connected, Claude can use swafra to remember and retrieve anything: "Remember this meeting transcript: ..." "What did we decide about the API design?" "What are my editor preferences?" "Forget everything from the project X sessions" | Tool | What it does | |---|---| add knowledge | Store text — chunked, embedded, and graph-linked | search knowledge | Find relevant chunks by natural language query | get context | Search + graph walk combined recommended | graph walk | Explore connected chunks from a starting point | list sources | See everything stored | delete source | Remove a source and all its data | 1. Chunking Text is split into semantically coherent chunks using Leiden community detection — a graph algorithm that groups sentences by topic. Falls back to conversation-aware chunking if Leiden deps aren't available Python 3.13+ . 2. Local embeddings Uses fastembed https://github.com/qdrant/fastembed ONNX, CPU-only, no API key with BAAI/bge-small-en-v1.5 . Falls back to deterministic hash vectors if fastembed isn't installed. 3. Hybrid retrieval 4-signal fused scoring: BM25 + vector cosine + entity/date overlap + character n-gram. Returns the best chunk per source so you get diverse, non-redundant context. 4. Knowledge graph Chunks are connected with sequential next/prev , similarity, and entity co-occurrence edges. Graph walk expands retrieval beyond what search alone finds. 5. Storage JSON files in ~/.scimap/ . No database, no server, no cloud. 94.7% recall all@10 on LongMemEval-S — 500 questions across 6 categories, 53 sessions each. | Category | recall all@10 | |---|---| | knowledge-update | 100.0% | | single-session-user | 100.0% | | single-session-preference | 100.0% | | single-session-assistant | 100.0% | | temporal-reasoning | 99.2% | | multi-session | 93.3% | Full benchmark details and reproduction steps → /kunal12203/swafra/blob/master/BENCHMARK.md | Variable | Default | Description | |---|---|---| SCIMAP DATA DIR | ~/.scimap | Where knowledge is stored | SCIMAP EMBED MODEL | BAAI/bge-small-en-v1.5 | Embedding model |