A Mac Can Now Search Every Frame of Your Video Without the Cloud
SCM indexes photos, video scenes, on-screen text and speech locally on macOS. It shows how far private, on-device AI search has come.
A developer has released SCM (Screen Memories), an open-source macOS app that lets you describe a memory in plain language and lands you on the exact shot in a video, not just the file. According to the project's page, inference runs on your Mac, with no accounts, no uploads and no telemetry. Model weights download once; after that, the app works offline.
The Signal: Search Is Moving to the Edge
The interesting part is not the app but the stack beneath it. Embedding models, speech recognition and a small chat model now run acceptably on a consumer laptop, using free, off-the-shelf components: CLIP and SigLIP vision models via ONNX Runtime, Whisper for speech, Tesseract for text, ffmpeg for video, and llama.cpp for chat. A single developer can assemble what once required a cloud vendor. The source also lists a default chat model of about 1.1GB that it says fits 8GB Macs. If that holds up in practice, the privacy argument for local search no longer costs users much in capability.
How It Works
The core idea is the embedding: a vision model converts an image into a long list of numbers, and converts your text query into the same kind of list. Images whose numbers sit closest to your query's numbers (measured by cosine similarity) rank highest. That is why you can search "red umbrella on a rainy street" without anyone having tagged the photo.
Video is harder because a file is thousands of frames. SCM uses ffmpeg to detect shot boundaries, splits each video into segments, embeds one midpoint frame per segment, and keeps a poster image. Users choose a density preset, from one sample per 60 seconds (Eco) to one per 2.5 seconds (Ultra Pro, which requires confirmation). The project says each preset shows its measured time and disk cost before you commit, which is the real constraint: finer search means more compute and storage.
Text is handled separately. OCR (optical character recognition) reads words visible in images and frames, matched literally. Dialogue search uses Whisper transcripts and also matches literally, in three tiers from an exact line down to all words appearing somewhere in the same video. Neither depends on the vision model, so they work even while it is warming up.
There is also an opt-in chat mode that answers questions only from evidence the app already extracted, with numbered citations, and short-circuits before the model runs if there is no evidence. That design choice limits, though does not eliminate, the risk of invented answers.
The Tradeoffs Built Into It
Several details show where local search still hurts. Four vision models trade speed for detail: the source lists roughly 50 to 100 milliseconds per image for the fastest, and about 480 to 570 milliseconds for the default CLIP. Switching models re-embeds the whole library in the background, with search falling back to filename keywords meanwhile. Switching the Whisper model re-transcribes every video. Your media is also copied into an app-managed library, so you are duplicating your files to make them searchable. The project's privacy claims, including its sandboxed renderer and no-telemetry stance, are the developer's own and not independently audited here.
Questions You Should Be Asking
- If the index of your photos, screenshots and spoken dialogue sits in one folder on disk, how is that folder protected, and who can read it if the Mac is lost or shared?
- The app OCRs everything, including email addresses, and has a tab built to surface them. What happens when a library contains documents, IDs or other people's information you never meant to make searchable?
- The project states search scores are "honest" via per-model calibration. Where is the evidence that a no-match result is reliable, and how would you test it on your own footage?
- Copying all media into a managed library doubles storage. Is that cost acceptable for your archive, and what is the exit path if the project is abandoned?
- Weights come from third-party hosts at install time. Who verifies them, and what is your policy for open-source tools that download models on first run?
What To Watch Next
The signal is whether operating systems and mainstream photo apps ship this natively, with scene-level video search and speech search running on-device by default. If they do, standalone tools like this become reference designs; if they stay cloud-first, projects like SCM mark the privacy-conscious alternative. Watch for independent benchmarks of accuracy and speed on large libraries, and for whether the Homebrew-first install model, which clears macOS quarantine flags automatically, draws scrutiny.
- 1Install SCM on your Mac to search video content by natural language description without uploading data to cloud servers.
- 2Use local AI models like CLIP, Whisper, and llama.cpp to process sensitive video files privately on your own device.
- 3Leverage open-source tools (ONNX Runtime, ffmpeg, Tesseract) to build your own edge-based video search without vendor lock-in.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
