SCM is a macOS Electron application built with Bun that provides local, AI‑powered search across photos and video frames without any cloud uploads or account requirements. On first launch it downloads a CLIP ViT‑L/14@336 vision model (~435 MB) via Hugging Face; subsequent operation is fully offline. The app indexes media by computing image embeddings, OCR text with Tesseract (English plus 35 toggleable language packs, each ~2.4‑5 MB), and Whisper‑transcribed dialogue (tiny.en ~150 MB or base.en ~300 MB). Video files are segmented by ffmpeg shot detection; users choose a preset (Eco, Balanced, Detailed, Ultra, Ultra Pro) that defines seconds per point and segment budget, and each segment’s midpoint frame is embedded for scene‑level search. Search results are scored by cosine similarity against the embeddings, with gated phrase and filename boosts, an honesty floor, and a near‑duplicate diversity filter; matched items display a “why it matched” badge and per‑component score breakdown. OCR matches literal text (highlighted in amber), while dialogue search offers three tiers—exact line, exact words within an ≤8 s window, and any words spoken in the video—using exact string matching over Whisper transcripts. An opt‑in LLM chat mode runs a llama.cpp sidecar (Qwen3 1.7B ≈ 1.1 GB or Llama 3.2 3B ≈ 2 GB) over the extracted dialogue, OCR, and filename keywords, providing numbered, streamed answers. The library watches folders, deduplicates files via SHA‑256 content hashes, and re‑embeds in the background when switching among four ONNX‑Runtime vision models (CLIP, SigLIP‑2‑B/16, SigLIP‑2‑L/16@256, SigLIP‑B/16@384). All data resides under ~/Library/Application Support/scm, and the app can be installed or upgraded via Homebrew cask with automatic quarantine‑flag clearing.
Read original
hackernews