Dropbox Witchcraft Dropbox released Witchcraft 0.2.0, a from-scratch safe-Rust reimplementation of Stanford's XTR-Warp semantic search engine that runs stand-alone on a single-file SQLite database with no API keys or vector database, hitting 14ms p95 end-to-end search latency on NFCorpus at 34% NDCG@10 on an Apple MacBook Pro M4 Max — more than twice as fast as the original XTR-WARP on server-class hardware. The release adds native Windows GPU support without OpenVINO, a more compact on-disk index format, better index-update scalability and search accuracy, and updates to the Pickbrain agent memory tool including launchd service registration on macOS for background index updates. The default build uses Dropbox's ModernBERT retrieval model with 96-dimensional token embeddings and token gating, fine-tuned from IBM's Granite Embedding English R2 under Apache 2.0, and requires uv, make, and the Rust toolchain. formerly known as Rust-Warp This is a from-scratch reimplementation of Stanford's XTR-Warp semantic search engine https://github.com/jlscheerer/xtr-warp https://github.com/jlscheerer/xtr-warp in safe rust, using a single-file SQLite database as backing storage, making it suitable for client-side deployment. It runs completely stand-alone on your device, needs no API keys, no vector database, no chunking strategy, no fancy re-rankers, and it is lightning fast 14ms p.95 end-to-end search latency on NFCorpus, at 34% NDCG@10, on an Apple Macbook Pro M4 Max, more than twice as fast as the original XTR-WARP on server-class hardware. Version 0.2.0 adds native support for Windows GPUs without relying on OpenVINO etc, a much more compact on-disk index format, better scalability with index updates, better search accuracy, and many updates to the Pickbrain agent memory tool, among them the ability to register as a launchd service on MacOS, so that the index is kept updated in the background, reducing the risk of having to wait on index updates during queries. Needs uv, make, and the rust toolchain. The default build uses our ModernBERT retrieval model with 96-dimensional token embeddings and token gating. Make downloads a compressed archive of the safetensors weights, config, and tokenizer from the versioned model release https://github.com/dropbox/witchcraft/releases/tag/modernbert-96d-gated-v1 into assets/ and derives the quantized GGUF model locally: make warp-cli Downloads use public URLs with curl ; no GitHub account or GitHub CLI is required. The same weights support the unquantized backend make warp-cli ENCODER=modernbert . The weights derive from IBM's Granite Embedding English R2 https://huggingface.co/ibm-granite/granite-embedding-english-r2 , licensed under Apache 2.0. Our fine-tuning and other modifications are Copyright c 2026 Dropbox Inc. and covered by the repository's Apache 2.0 license. The archive includes the repository LICENSE , the upstream LICENSE.granite , a NOTICE identifying our modifications, and SHA256SUMS for verifying downloads. Existing local assets are kept. To use your own checkpoint, export it with env/bin/python scripts/export modernbert.py