# We achieved 94.7% in Longmemeval with no hacks lol, everything open sourced

> Source: <https://github.com/kunal12203/swafra>
> Published: 2026-07-26 11:28:31+00:00

Semantic memory for AI — ingest anything, retrieve what matters.

**94.7% recall_all@10 on LongMemEval** — the standard benchmark for long-term memory in AI assistants.

Works as an MCP server with Claude Desktop, Claude Code, VS Code Copilot, and any MCP-compatible AI.

```
pip install swafra
```

Or with Node.js:

```
npm install -g swafra
```

Add to `~/Library/Application Support/Claude/claude_desktop_config.json`

:

```
{
  "mcpServers": {
    "swafra": {
      "command": "swafra"
    }
  }
}
```

Restart Claude Desktop — the tools appear automatically.

```
claude mcp add swafra swafra
```

Add to `.vscode/mcp.json`

in your project:

```
{
  "servers": {
    "swafra": {
      "command": "swafra"
    }
  }
}
```

Once connected, Claude can use swafra to remember and retrieve anything:

```
"Remember this meeting transcript: ..."
"What did we decide about the API design?"
"What are my editor preferences?"
"Forget everything from the project X sessions"
```

| Tool | What it does |
|---|---|
`add_knowledge` |
Store text — chunked, embedded, and graph-linked |
`search_knowledge` |
Find relevant chunks by natural language query |
`get_context` |
Search + graph walk combined (recommended) |
`graph_walk` |
Explore connected chunks from a starting point |
`list_sources` |
See everything stored |
`delete_source` |
Remove a source and all its data |

**1. Chunking**
Text is split into semantically coherent chunks using Leiden community detection — a graph algorithm that groups sentences by topic. Falls back to conversation-aware chunking if Leiden deps aren't available (Python 3.13+).

**2. Local embeddings**
Uses [fastembed](https://github.com/qdrant/fastembed) (ONNX, CPU-only, no API key) with `BAAI/bge-small-en-v1.5`

. Falls back to deterministic hash vectors if fastembed isn't installed.

**3. Hybrid retrieval**
4-signal fused scoring: BM25 + vector cosine + entity/date overlap + character n-gram. Returns the best chunk per source so you get diverse, non-redundant context.

**4. Knowledge graph**
Chunks are connected with sequential (next/prev), similarity, and entity co-occurrence edges. Graph walk expands retrieval beyond what search alone finds.

**5. Storage**
JSON files in `~/.scimap/`

. No database, no server, no cloud.

94.7% recall_all@10 on LongMemEval-S — 500 questions across 6 categories, 53 sessions each.

| Category | recall_all@10 |
|---|---|
| knowledge-update | 100.0% |
| single-session-user | 100.0% |
| single-session-preference | 100.0% |
| single-session-assistant | 100.0% |
| temporal-reasoning | 99.2% |
| multi-session | 93.3% |

[Full benchmark details and reproduction steps →](/kunal12203/swafra/blob/master/BENCHMARK.md)

| Variable | Default | Description |
|---|---|---|
`SCIMAP_DATA_DIR` |
`~/.scimap` |
Where knowledge is stored |
`SCIMAP_EMBED_MODEL` |
`BAAI/bge-small-en-v1.5` |
Embedding model |
