{"slug": "dropbox-witchcraft", "title": "Dropbox Witchcraft", "summary": "Dropbox released Witchcraft 0.2.0, a from-scratch safe-Rust reimplementation of Stanford's XTR-Warp semantic search engine that runs stand-alone on a single-file SQLite database with no API keys or vector database, hitting 14ms p95 end-to-end search latency on NFCorpus at 34% NDCG@10 on an Apple MacBook Pro M4 Max — more than twice as fast as the original XTR-WARP on server-class hardware. The release adds native Windows GPU support without OpenVINO, a more compact on-disk index format, better index-update scalability and search accuracy, and updates to the Pickbrain agent memory tool including launchd service registration on macOS for background index updates. The default build uses Dropbox's ModernBERT retrieval model with 96-dimensional token embeddings and token gating, fine-tuned from IBM's Granite Embedding English R2 under Apache 2.0, and requires uv, make, and the Rust toolchain.", "body_md": "(formerly known as Rust-Warp)\n\nThis is a from-scratch reimplementation of Stanford's XTR-Warp semantic search\nengine ( [https://github.com/jlscheerer/xtr-warp](https://github.com/jlscheerer/xtr-warp) ) in safe rust, using a\nsingle-file SQLite database as backing storage, making it suitable for\nclient-side deployment. It runs completely stand-alone on your device, needs no\nAPI keys, no vector database, no chunking strategy, no fancy re-rankers, and it\nis lightning fast (14ms p.95 end-to-end search latency on NFCorpus, at 34%\nNDCG@10, on an Apple Macbook Pro M4 Max, more than twice as fast as the\noriginal XTR-WARP on server-class hardware.)\n\nVersion 0.2.0 adds native support for Windows GPUs without relying on OpenVINO etc, a much more compact on-disk index format, better scalability with index updates, better search accuracy, and many updates to the Pickbrain agent memory tool, among them the ability to register as a launchd service on MacOS, so that the index is kept updated in the background, reducing the risk of having to wait on index updates during queries.\n\nNeeds uv, make, and the rust toolchain.\n\nThe default build uses our ModernBERT retrieval model with 96-dimensional\ntoken embeddings and token gating. Make downloads a compressed archive of the\nsafetensors weights, config, and tokenizer from the\n[versioned model release](https://github.com/dropbox/witchcraft/releases/tag/modernbert-96d-gated-v1)\ninto `assets/` and derives the quantized GGUF model locally:\n\n```\nmake warp-cli\n```\n\nDownloads use public URLs with `curl`; no GitHub account or GitHub CLI is required.\n\nThe same weights support the unquantized backend\n(`make warp-cli ENCODER=modernbert`). The weights derive from IBM's\n[Granite Embedding English R2](https://huggingface.co/ibm-granite/granite-embedding-english-r2),\nlicensed under Apache 2.0. Our fine-tuning and other modifications are\nCopyright (c) 2026 Dropbox Inc. and covered by the repository's Apache 2.0 license.\nThe archive includes the repository `LICENSE`, the upstream `LICENSE.granite`,\na `NOTICE` identifying our modifications, and `SHA256SUMS` for verifying downloads.\nExisting local assets are kept. To use your own checkpoint, export it with\n`env/bin/python scripts/export_modernbert.py <checkpoint> assets`, then quantize\nit with `cargo run -p quantize --release -- assets/modernbert.safetensors assets/modernbert.gguf`.\n\nFor Google's original XTR model, run `make warp-cli ENCODER=t5-quantized`;\nthe included Python scripts download its weights from Hugging Face and\nquantize them to GGUF.\n\nFor testing, we used the BEIR download script from XTR-Warp to download nfcorpus and check that we could replicate their results. For your convenience, nfcorpus.tsv is included here, so you can run:\n\n``` bash\n$ make nfcorpus\n```\n\nWith all the nfcorpus documents imported, embeddings will be created, and the index updated with them.\n\nAll state gets persisted in mydb.sqlite, and you can abort the indexer and it will pick up where it left off. To start over, you can just delete mydb.sqlite.\n\nYou can rerun the nfcorpus scoring result with:\n\n``` bash\n$ make nfcorpus-score\n```\n\nWhen you have the index, you can query it with:\n\n```\n$ ./warp-cli query \"does milk intake cause acne in teenagers?\"\n```\n\nAnd hopefully get a bunch of relevant answers. You can also try other variations of this, instead of \"query\" you can also use \"hybrid\", which combines semantic search with the BM25 search functionality that comes standard with sqlite.\n\nIncluded as an example is **pickbrain** (screenshot above), a CLI that indexes\nyour Pi, Claude Code, and OpenAI Codex session transcripts, memory files, and\nauthored documents into a Witchcraft database for fast semantic search. Ever\nwondered \"what was that conversation where I fixed the auth middleware?\" —\npickbrain finds it, and lets you resume the session directly.\n\n```\nmake pickbrain\n./pickbrain auth middleware fix    # search across all sessions (auto-ingests new sessions)\n./pickbrain --session <UUID> auth  # search within one session\n./pickbrain --dump <UUID>          # print full conversation\n```\n\nSet `PRE_INGEST_COMMAND` to a shell command that runs before pickbrain checks for new sessions (for example, an `rsync` from another host). Set `EXTRA_CODEX_DIRS`, `EXTRA_CLAUDE_DIRS`, or `EXTRA_PI_DIRS` to additional data directories separated by `:` on Unix or `;` on Windows. Each entry should have the same layout as `~/.codex`, `~/.claude`, or `~/.pi/agent`, respectively. Pickbrain scans these alongside the default directories.\n\n`scripts/sync_codex_sessions.sh` takes an SSH host and an absolute local directory. It copies only session `*.jsonl` files and writes the host to `pickbrain.remote`. Pickbrain labels those results as downloaded; press `r` in the browser to see the SSH command for that host.\n\n```\nPRE_INGEST_COMMAND=\"$PWD/scripts/sync_codex_sessions.sh your-ssh-alias $HOME/.pickbrain/remote-codex\" \\\nEXTRA_CODEX_DIRS=\"$HOME/.pickbrain/remote-codex\" ./pickbrain auth middleware fix\n```\n\nFor an existing extra Codex directory, put its SSH host name in `<extra-codex-dir>/pickbrain.remote` to tag those sessions.\n\nThe source lives in `examples/pickbrain/` and demonstrates how to use\nWitchcraft as a library: document ingestion, embedding, indexing, and hybrid\nsearch. To install pickbrain as a skill/extension for Pi and as a skill for both Claude Code and Codex:\n\n```\nmake pickbrain-install\n```\n\nFor automatic ingestion, run Pickbrain in watcher mode:\n\n```\nmake pickbrain\n./pickbrain --watch\n```\n\n`pickbrain --watch` watches local Claude, Codex, and Pi directories, configured extra\ndirectories, and Slack's IndexedDB blobs on macOS. It uses native notifications,\nwaits for 30 seconds without changes, and caps the delay at five minutes. Set\n`--delay SECONDS` and `--max-delay SECONDS` to adjust these intervals. It checks\nfor outstanding work at startup and rescans when notifications are lost. New or\nrecreated source directories are detected automatically.\n\nOn macOS, FSEvents can defer notifications for session files kept open by Codex. The watcher also checks session file sizes and change times every quiet-delay interval (30 seconds by default). This reads metadata only; ingestion starts when changes are detected and the debounce delay expires.\n\nThe watcher spawns its own executable with `--ingest-only --quiet`, with one child at a time.\nIt logs each completed run as `pickbrain: no changes` or\n`pickbrain: ingested and indexed N documents`. Counts refer to processed documents\n(conversation turns and files), including documents reread from changed sessions.\nThe summary appears after indexing succeeds; errors remain visible.\nThe watcher retains changes arriving during ingestion for another pass. Failed children\nare retried after the quiet delay. Model memory and GPU resources belong to the\nchild and are released when it exits; watcher mode enters before database,\nmodel, or GPU initialization and waits between notifications and metadata checks. Pickbrain's database, logs, and watermarks\ndo not trigger ingestion. The normal ingestion lock also protects against\ninteractive Pickbrain runs. Incomplete headless ingestion is retried even if\nsource watermarks have already advanced.\n\n`make pickbrain-install` installs the single binary in `~/bin`, then registers or\nupdates the background watcher. To register an existing binary directly, run:\n\n```\npickbrain --register\n```\n\nRegistration starts the watcher immediately and at login, using a macOS user\nLaunchAgent (`com.dropbox.pickbrain`), Linux user systemd service\n(`pickbrain.service`), or Windows Task Scheduler logon task (` Pickbrain-<user SID>`).\nIt replaces the same job on subsequent registrations. `--delay` and `--max-delay`\nalso work with `--register`. Registration saves the current working directory,\n`PATH`, source-directory settings, `PICKBRAIN_DIR`, `WARP_ASSETS`, and\n`PRE_INGEST_COMMAND` in `~/.pickbrain/watch.json`, so the background job can use the\nsame settings outside your shell. Re-register to change them. Logs, including\ningestion summaries and errors, go to `~/.pickbrain/watch.log`:\n\n```\ntail -f ~/.pickbrain/watch.log\n```\n\nOn macOS, stop the job with\n`launchctl bootout gui/$(id -u)/com.dropbox.pickbrain`; remove\n`~/Library/LaunchAgents/com.dropbox.pickbrain.plist` to prevent launch at the next\nlogin. On Linux use `systemctl --user disable --now pickbrain.service`; on Windows,\nstop and delete the named task in Task Scheduler. Linux registration requires a\nrunning user systemd manager. On Windows, `USERPROFILE` supplies the home directory\nwhen `HOME` is unset. The watcher uses its own executable path, so symlinks and\n`PATH` resolution need no configuration. Remote hosts still need a separate\nperiodic synchronization job, since local filesystem notifications cannot detect\nchanges on those hosts.\n\nThis puts the binary on your `PATH` and installs the skill/extension definitions so\nyou can use pickbrain directly from Pi, Claude Code, or Codex to answer questions requiring\nglobal knowledge of all your projects:\n\nWhen building, exactly one encoder backend must be enabled:\n\n- `t5-quantized` -- GGUF quantized weights via candle\n- `modernbert-quantized` -- GGUF ModernBERT weights (default)\n- `modernbert` -- full-precision ModernBERT weights\n\nOther flags:\n\n- `metal` -- Candle Metal acceleration\n- `neso-metal` -- Neso-generated Metal kernels\n- `neso-d3d12` -- Neso-generated D3D12 kernels\n- `fbgemm` -- fbgemm-rs packed GEMM (bf16 weights, faster on x86)\n- `hybrid-dequant` -- F32 attention + Q4K FFN with fused gated-gelu (x86, requires`fbgemm` )\n- `napi` -- Node.js native module via napi-rs\n- `python` -- Python extension module via PyO3/maturin\n- `embed-assets` -- bake weights into binary\n- `progress` -- progress bars for CLI\n\nNeso GPU kernels are cached as compressed archives in `kernels/out/` and embedded\nin the binary. Cargo rebuilds them when the Neso environment is available (at\n`../neso`, or `NESO_DIR`); otherwise it uses the checked-in caches, including their\nRust loaders. This fallback needs no Python, Neso, or shader compiler. Metal uses\nthe scalar kernels on both Mac architectures; Windows uses the DXIL cache.\nTo refresh the caches after changing kernels, run `make neso-kernels-metal-nosimd`\nand `make neso-kernels-hlsl` with Neso installed. Cache generation also needs `zstd`.\n\nPlatform-specific recommended features (these are what `make` uses automatically):\n\n- **Apple Silicon** :`modernbert-quantized,neso-metal`\n- **Intel Mac (x86_64)** :`modernbert-quantized,neso-metal`\n- **Intel Windows (x86_64)** :`modernbert-quantized,neso-d3d12`\n- **Linux x86_64 (CPU)** :`modernbert-quantized,fbgemm,hybrid-dequant`\n- **Linux x86_64 (CUDA)** :`modernbert-quantized,cuda`\n- **Linux ARM (Graviton, Pi, Ampere)** :`modernbert-quantized` (`fbgemm` /`hybrid-dequant` are x86-only)\n\nBuild a wheel for your current platform (auto-selects the right backend features):\n\n```\nmake python-wheel\n```\n\nOr install directly into the repo-managed virtualenv for development:\n\n```\nmake python-dev\n```\n\nThe Makefile picks the correct feature set automatically (`metal` on Apple Silicon,\n`fbgemm,hybrid-dequant` on Intel, `cuda` on Linux with a GPU). The wheel works on\nPython 3.8+, including 3.14+.\n\nTo build/install the extension and run the Python test suite:\n\n```\nmake python-test\n```\n\nOverride `PYTHON_TEST_ARGS` to pass custom pytest arguments, for example\n`make python-test PYTHON_TEST_ARGS=\"-q tests/test_witchcraft.py\"`.\n\nThen in Python:\n\n``` python\nimport witchcraft\n\nwc = witchcraft.Witchcraft('/path/to/db.sqlite', '/path/to/assets')\n\n# Add documents (fire-and-forget, processed by background thread)\nwc.add('550e8400-e29b-41d4-a716-446655440000',\n       '2024-01-15T10:00:00Z',\n       '{\"source\": \"dropbox\"}',\n       'The document text goes here')\n\n# Build pending embeddings and index data; blocks until complete\nwc.index()\n\n# Hybrid semantic + BM25 search\nresults = wc.search('does milk intake cause acne?', threshold=0.3, top_k=5)\nfor r in results:\n    print(r['score'], r['body'])\n\n# Score individual sentences against a query\nscores = wc.score('acne and diet', ['milk causes acne', 'exercise helps skin'])\n\n# Shut down the background indexer cleanly\nwc.shutdown()\n```\n\n`search` returns a list of dicts with keys: `score`, `metadata`, `body`, `idx`, `date`.\n\n```\nmake module\n```\n\nBuilds a universal macOS binary at `target/release/warp-macos-universal.node`\n(lipo'd from aarch64 + x86_64 builds with platform-appropriate features).\n\nTo use in another project:\n\n```\ncd /path/to/your-project\nnpm install /path/to/witchcraft\n```\n\nThen in JavaScript:\n\n``` js\nconst { Witchcraft } = require('warp');\nconst wc = new Witchcraft('/path/to/db.sqlite', '/path/to/assets');\ncargo install cargo-nextest\ncargo install cargo-llvm-cov\ncargo install llvm-tools-preview\nmake test\n```\n\nNOTICE that nextest is necessary, simply using \"cargo test\" will lead to random test failures, because individual tests run in the same process, leading to \"history effects\".\n\nBefore submitting a PR, please make sure that\n\n```\nmake test\n```\n\nand\n\n```\nmake nfcorpus\n```\n\nRun, and that\n\n```\nmake nfcorpus-score\n```\n\nRuns and scores in the 0.31-0.33 range\n\nNOTICE that nextest is necessary, simply using \"cargo test\" will lead to random test failures, because individual tests run in the same process, leading to \"history effects\".\n\nTRECCOVID is prepared from the BEIR `trec-covid` archive on demand:\n\n```\nmake treccovid-files\nmake treccovid\nmake treccovid-score\n```\n\n`make treccovid` builds `mydb.sqlite` from `datasets/treccovid.tsv`, and\n`make treccovid-score` runs `treccovid-score.sh` against the test queries and\nprints NDCG@10.\n\nScoring runs entirely in Rust, without Python or `pytrec_eval`:\n\n```\nmake trec-score\ntarget/release/trec-score output.txt testset/nfcorpus/collection_map.json testset/nfcorpus/qrels.test.json\n```\n\nThe tool reads `querycsv` results and prints mean NDCG@10 with trec_eval's\nlinear relevance gains. An optional fourth argument changes the cutoff.\nUse a JSON `null` collection map when results already contain original document\nIDs. Like the previous scorer, it averages submitted queries that have qrels;\nempty results count as zero and missing or unjudged queries are excluded.\nThe standalone workspace tool builds without encoder features, weights, or GPU\ndependencies. NFCorpus, TRECCOVID, SciFact, and Spotlight scoring use it.\n\nUnless otherwise noted:\n\n```\nCopyright (c) 2026 Dropbox Inc.\n\nLicensed under the Apache License, Version 2.0 (the \"License\");\nyou may not use this file except in compliance with the License.\nYou may obtain a copy of the License at\n\n    http://www.apache.org/licenses/LICENSE-2.0\n\nUnless required by applicable law or agreed to in writing, software\ndistributed under the License is distributed on an \"AS IS\" BASIS,\nWITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\nSee the License for the specific language governing permissions and\nlimitations under the License.\n```\n\n", "url": "https://wpnews.pro/news/dropbox-witchcraft", "canonical_source": "https://github.com/dropbox/witchcraft", "published_at": "2026-10-10 18:53:50+00:00", "updated_at": "2026-10-10 19:17:04.934674+00:00", "lang": "en", "topics": ["ai-search", "ai-tools", "developer-tools", "natural-language-processing", "ai-agents"], "entities": ["Dropbox", "Witchcraft", "XTR-Warp", "Stanford", "Pickbrain", "ModernBERT", "IBM", "Granite Embedding English R2"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/dropbox-witchcraft", "markdown": "https://wpnews.pro/news/dropbox-witchcraft.md", "text": "https://wpnews.pro/news/dropbox-witchcraft.txt", "jsonld": "https://wpnews.pro/news/dropbox-witchcraft.jsonld"}}