{"slug": "show-hn-ai-search-for-every-photo-and-every-frame-of-video-on-macos", "title": "Show HN: AI search for every photo and every frame of video on macOS", "summary": "A developer released SCM 0.2.4, a local-first AI search app for macOS that indexes every photo and every frame of video in watched folders using on-device vision models, OCR via Tesseract, and dialogue search via Whisper. The app ships as SCM-0.2.4.dmg and SCM-0.2.4.zip, installs via a Homebrew cask (Apple Silicon, macOS 12+), and downloads roughly 435MB of CLIP weights on first use, after which all inference runs offline with no accounts, cloud, or uploads.", "body_md": "Deep AI search for every photo and every frame of video in any folder on macOS. Local-first — no accounts, no cloud, no uploads. Inference runs on your Mac.\n\n**What makes it different**\n\n- **Search like you think** — describe a memory in plain language; a local\nvision model does the rest.\n- **Video, down to the moment** — scenes are segmented and embedded, so you\nland on the shot, not just the file.\n- **Text and dialogue too** — OCR over visible text; exact spoken-line search\nvia Whisper, each as its own mode.\n- **Your tabs, your prompts** — save any query as a tab; Screenshots and\nEmail tabs are toggleable.\n- **Self-maintaining library** — watched folders auto-import, content hashes\ndedupe renames, and model switches re-embed in the background without\nblocking search.\n- **Truly private** — your media never leaves the machine. Weights download\nonce; everything after that is offline.\n\n| Mode | Finds | \n|---|---|\n| **Files** | Whole photos/videos by meaning — vision rank with filename and phrase boosts | \n| **Scenes** | Moments inside video — search a shot, jump to its timecode | \n| **OCR** | Text visible in images and frames, matched literally (Tesseract; eng + 35 language toggles) | \n| **Dialogue** | Exact spoken words in videos (Whisper), tiered exactness | \n| **LLMs***(opt-in)* | Local chat over the dialogue, OCR, and filenames your Mac already extracted — cited answers | \n\n- macOS (packaged with electron-builder; menu-bar/tray features are macOS-only)\n- [Bun](https://bun.sh) — the project uses`bun` as package manager and\nrunner\n- Node modules installed: `bun install`\n- First use of a model downloads its weights (~435MB for the default CLIP); after that, fully offline.\n\nThe easiest install is via Homebrew (Apple Silicon, macOS 12+). The tap's cask clears the macOS quarantine flag automatically on every install and upgrade, so the app launches with no manual Gatekeeper steps:\n\n```\nbrew tap allenv0/scm\nbrew trust allenv0/scm\nbrew install --cask allenv0/scm/scm\n```\n\nUpgrades keep the same behavior:\n\n```\nbrew upgrade --cask allenv0/scm/scm\n```\n\nPrefer least privilege? Trust just the cask instead of the whole tap:\n\n```\nbrew tap allenv0/scm\nbrew trust --cask allenv0/scm/scm\nbrew install --cask scm\n```\n\nThe tap lives at\n[allenv0/homebrew-scm](https://github.com/allenv0/homebrew-scm).\n\n```\nbun run dev     # build the renderer bundle, then launch the Electron app\nbun start       # launch the Electron app without rebuilding\nbun run build   # just rebuild the renderer bundle into dist/\nbun run dist            # signed if an identity is in the keychain\nbun run dist:unsigned   # skip code-sign discovery\n```\n\nThis runs two steps in sequence:\n\n1. **`vite build`** — compiles the React renderer into` dist/` (picks up all\nchanges under`src/` ).\n2. **`electron-builder --mac`** — packages the app. It bundles the fresh` dist/` bundle together with`main.js` ,`preload.js` ,`main-lib/` , and`indexer/` (the file list is configured under`build.files` in`package.json` ), then produces the installers.\n\n**Output:** the installers land in `dist-app/` (see `build.directories.output`\nin `package.json`) — look for `SCM-0.2.4.dmg` and `SCM-0.2.4.zip`.\n\nTyping starts an instant filename-keyword pre-pass, then the vision model takes over: results are scored by cosine similarity against image embeddings, with gated phrase and filename boosts, an honesty floor calibrated per model, and a near-duplicate diversity filter. Every tile carries a \"why it matched\" badge (Visual match / Filename match / …) and a hover tooltip with the per-component score breakdown. CJK queries search as overlapping bigrams (\"台北車站\" also matches 台北, 車站).\n\nEvery scene segment across all videos is scored, so a hit lands on the exact shot: tiles show the scene poster with a timecode badge, and opening the video jumps straight to that moment. A noise gate returns \"no scene match\" instead of flooding the grid with gibberish, and each video contributes at most 3 scenes.\n\nMatches the fraction of query tokens literally visible in each image's OCR text — the filename is ignored and no vision model is involved, so it works even while the AI engine is warming up or offline. Matched words are boxed in amber on tiles and in the lightbox.\n\nExact literal retrieval over Whisper transcripts — no embeddings, no\nthresholds, works with the AI engine down. Results come in three tiers:\n**Exact line** (contiguous phrase in one utterance), **Exact words** (all\nwords in one utterance or an ≤8s window), and **Words spoken** (all words in\nthe same video). Matching words are highlighted in a speech snippet; opening\na result seeks straight to the line.\n\nOpt-in — nothing downloads or runs until enabled in Settings → LLMs Chat. A\nllama.cpp sidecar bound to loopback answers your question from evidence the\napp already extracted — dialogue lines, OCR text, and filename keyword hits\n— with numbered citations you can click, streamed token-by-token with a live\ntok/s readout. Leading `/screenshots`, `/videos`, `/email` narrow the\ncorpus; Stop keeps the partial answer; empty evidence short-circuits before\nthe model ever runs.\n\n| Chat model | Size | Notes | \n|---|---|---|\n| **Qwen3 1.7B** (default) | ~1.1GB | Fast everyday chat; fits 8GB Macs | \n| **Llama 3.2 3B** | ~2GB | Stronger long answers; needs headroom | \n\n- Built-in browse tabs: **All** ,**Videos** , plus**Screenshots** and**Email** — the latter two toggleable in Settings → Smart Tabs. Selecting\nVideos auto-enables Scenes mode.\n- **Save any query as a tab** : the pin pill under the search bar saves the\ncurrent prompt with its mode (Files/Scenes/OCR/Dialogue) — up to 20 tabs,\nrenameable, each restored exactly as saved.\n- Five semantic views behind remappable shortcuts (⌘1–⌘5 by default), plus ⌘I import / AI Insights, ⌘, for Settings.\n- **Search stays in the selected tab** : pick Screenshots, Email, Videos, or\nany saved tab and results are filtered to it — scope first, then search.\nIn LLMs chat the same idea is explicit: leading`/screenshots` ,`/videos` ,`/email` narrow the corpus before the model ever runs.\n\nSurfaces photos whose visible OCR text contains an email address — an overlapping view (a photo keeps its category too). Detection is OCR-tolerant: it reassembles addresses Tesseract fractures across word boxes, and handles comma-for-dot noise (\"gmail,com\"), split TLDs (\"gmail. com\"), bracketed obfuscation (\"allen [at] gmail [dot] com\"), and dictated addresses (\"allen at gmail dot com\"). Tiles show a contact strip; expand it to copy or compose.\n\nScreenshot classification is rename-proof. Four signals, in priority order: a manual override (right-click any tile) → filename vocabulary (30+ localized OS screenshot names in 20+ languages) → a PNG/JPEG metadata probe (reads \"screenshot\" from PNG text chunks / EXIF UserComment, so a renamed Bildschirmfoto still classifies) → source-folder hint. Everything else lands in Projects.\n\nFour switchable models via ONNX Runtime; the active one is chosen per library:\n\n| Model | Role | Speed (CPU) | Download | \n|---|---|---|---|\n| **CLIP ViT-L/14@336** (default) | Best real-world video scene-search | ~480–570ms/img | ~435MB | \n| **SigLIP-2-B/16** | Fastest bulk import | ~50–100ms/img | ~412MB | \n| **SigLIP-2-L/16@256** | High-detail (1024-dim) — small objects, signs, on-screen text | ~200ms/img | ~850MB | \n| **SigLIP-B/16@384** | Maximum detail | ~480ms/img | ~214MB | \n\nSwitching models re-embeds the whole library: the flip lands instantly with the tail filled in the background, and search falls back to filename keywords until it completes. Per-model text-mean centering de-biases text embeddings so similarity scores stay honest across models.\n\nffmpeg scans each video for shot boundaries and builds a segment plan sampled to the density you pick in Settings → Video Search — each preset shows its measured time and disk cost before you commit:\n\n| Preset | Seconds per point | Segment budget | \n|---|---|---|\n| Eco | 60 | 4–32 | \n| **Balanced** (default) | 30 | 8–128 | \n| Detailed | 15 | 12–256 | \n| Ultra | 5 | 16–1024 | \n| Ultra Pro | 2.5 | 24–2048 (confirm required) | \n\nEach segment embeds its midpoint frame and keeps a poster; shot plans are cached per file (path + size + mtime + config fingerprint), so re-imports skip detection entirely.\n\n**Dialogue transcription:** Whisper `tiny.en` (~150MB, default) or `base.en`\n(~300MB) — switching re-transcribes every video. Whole videos embed three\nframes (20/50/80%) averaged; GIFs embed an average of middle frames.\n\nTesseract runs in its own worker, separate from the vision model. English is always on; 35 more languages are toggleable in Settings → Photo Search (default: Simplified + Traditional Chinese, Japanese, Korean). Each language pack downloads once (~2.4–5MB; ~17MB for the default set), then everything is offline. Word boxes are stored with the text so matches highlight in place; CJK text is joined without spaces and email fragments fractured across word boxes are reassembled.\n\n- **Import** via ⌘I, drag-and-drop, or watched folders — importing a folder\nstarts watching it (live`fs.watch` plus a re-sync at every launch).\nProblem files retry up to 3 times, then sit out watch-syncs until they\nchange.\n- **Rename-proof dedupe** : every file is content-hashed (SHA-256) before\ncopy, and lying extensions are normalized by MIME sniffing.\n- **Named embedding versions** (Settings → Library): point-in-time snapshots\nof the entire searchable state — the index, every model's embedding bins,\nscene and transcript sidecars — with restore (auto-backup first) and a\nFresh Start danger zone. Cap: 10.\n- Everything lives under `~/Library/Application Support/scm` (`MEMORIES_DATA_DIR` overrides it): the index JSON, Float32 embedding bins\nper model, scene and transcript sidecars, thumbnails, and posters.\n\n- The renderer is a sandboxed `app://` bundle —`contextIsolation` , OS\nsandbox, and a CSP pinned to`'self'` (+ Google Fonts CDN for display\ntype, with a monospace fallback when offline).\n- Only main-process workers ever download, once per thing: vision weights (Hugging Face), OCR language packs (Tesseract CDN), Whisper weights, and — only if you opt in — the llama.cpp sidecar and GGUF chat models (GitHub + Hugging Face), sha256-verified at download time.\n- Media is copied into the app-managed library and streamed from disk. No telemetry, no accounts, no uploads.\n\nA macOS-style settings sheet with ten panels (Library, Appearance, Grid, Smart Tabs, Photo Search, Video Search, LLMs Chat, Global Shortcut, Keyboard, Menu Bar). Around the core: light/dark/system theme, a CRT screen effect for the lightbox, 6 alternate app icons, menu-bar-only mode, a recordable global shortcut, a first-run onboarding tour, background-work trays (scenes / transcripts / OCR) with a global pause, and a status bar with version and indexed-video counts.\n\n- **First build is slow:** electron-builder downloads the Electron binary and\nffmpeg once on its first run; subsequent builds are much faster.\n- **Code signing:** without an Apple Developer identity configured in the\nkeychain, the DMG builds unsigned. The app still runs locally, but macOS may\nrequire right-click → Open the first time it's launched.\n- **Rebuilds pick up changes automatically:** since`main.js` ,`preload.js` ,`main-lib/` , and`indexer/` are packaged from source (not cached), a`bun run dist` after editing any of them produces a fresh package.\n\n```\nbun run test:all         # the full verification battery (unit + smoke suites)\nbun run smoke:indexer    # headless smoke test of the CLIP indexer\nbun test test/           # unit tests (pure modules, no Electron needed)\nbun run lint             # eslint\nbun run format:check     # prettier\n```\n\n- **Unit tests** cover the pure cores: ranking, dialogue exact-match, CJK\ntokens, the Whisper model ladder, MIME sniffing, transcript fusion, and\nmore (see the`test:*` scripts in`package.json` ).\n- **E2E smoke tests** run inside the Electron app via`ELECTRON_SMOKE_*` environment variables — a dozen-plus scenarios from boot/protocol checks\nto search matrices, model migration, and Ask mode (drivers in`scripts/e2e/` , dispatched in`main.js` ; e.g.`bun run smoke:ask` ,`smoke:deep` ,`smoke:grid` ).\n- **Benchmarks:**`bun run bench:inference` ,`bench:enrich` , and`bench:detect` write JSON reports into`MDs/bench-*` .\n\n```\nmain.js             Electron main process: library, IPC, indexer worker pool, app:// protocol\npreload.js          contextBridge — exposes window.memories to the renderer\nmain-lib/           main-process modules split out of main.js (settings, library store,\n                    rank search, Ask retrieval, LLM sidecar config, embedding versions, …)\nindexer/            vision/OCR/ASR workers (utility processes) + video utils + model registry\nsrc/                React renderer (grid, search modes, lightbox, tabs, settings, onboarding)\nscripts/            bench scripts, E2E drivers, the smoke battery (smoke-all.sh)\ntest/               unit + integration tests\nMDs/                design docs, bench reports, plans\ndist/               vite build output (renderer bundle)\ndist-app/           electron-builder output (DMG / ZIP)\n```\n\n", "url": "https://wpnews.pro/news/show-hn-ai-search-for-every-photo-and-every-frame-of-video-on-macos", "canonical_source": "https://github.com/allenv0/SCM", "published_at": "2026-10-04 09:24:52+00:00", "updated_at": "2026-10-04 09:42:00.391597+00:00", "lang": "en", "topics": ["ai-search", "ai-tools", "computer-vision", "natural-language-processing", "ai-products"], "entities": ["SCM", "macOS", "Homebrew", "Electron", "Bun", "Tesseract", "Whisper", "CLIP"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-ai-search-for-every-photo-and-every-frame-of-video-on-macos", "markdown": "https://wpnews.pro/news/show-hn-ai-search-for-every-photo-and-every-frame-of-video-on-macos.md", "text": "https://wpnews.pro/news/show-hn-ai-search-for-every-photo-and-every-frame-of-video-on-macos.txt", "jsonld": "https://wpnews.pro/news/show-hn-ai-search-for-every-photo-and-every-frame-of-video-on-macos.jsonld"}}