Open-source project
ssrajadh/sentrysearch avatar
ssrajadh/sentrysearch

SentrySearch: type what you saw, get the clip back

Semantic search over videos using Gemini Embedding 2 or Qwen3-VL.

4,526 stars426 forksPythonApache-2.0

At a glance

What is it?
SentrySearch is an Apache-2.0 Python tool for semantic search over video footage, splitting videos into overlapping chunks, embedding them with Google's Gemini Embedding API, Alibaba's qwen-cloud or a local Qwen3-VL model, and storing vectors in ChromaDB, so a text or image query returns a trimmed clip of the best match. It runs entirely local with an MLX backend twice as fast on Apple Silicon, and adds companion tools for stitching and redaction.
Who is it for?
Use SentrySearch when hours of footage, dashcams, security cameras, drone files, need finding by description rather than scrubbing, and the trade of per-chunk embedding costs is acceptable for the retrieval speed afterward. Choose the local or MLX backend when footage cannot leave the machine or API costs are unwanted, accepting the hardware requirements.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Chunks, embeddings, and a trimmed clip

The pipeline is one paragraph long and complete. SentrySearch splits videos into overlapping chunks, embeds each chunk as video using Google's Gemini Embedding API, Alibaba DashScope under the qwen-cloud backend, or a local Qwen3-VL model, and stores the vectors in a local ChromaDB database. When you search, the text query or an image is embedded into the same vector space and matched against the stored video embeddings, and the top match is automatically trimmed from the original file and saved as a clip. The design's consequence is worth stating, indexing is the expensive, one-time phase proportional to footage length, and searching is cheap and local, the shape that suits growing archives queried many times after a single indexing pass. A demo video in the docs directory shows the loop end to end.

uv tool install, and the Python 3.12 pin

Getting started installs uv first, through the astral.sh install script on macOS and Linux:

bash
curl -LsSf https://astral.sh/uv/install.sh | sh

or the PowerShell one-liner on Windows, then clones the repository and installs:

bash
git clone https://github.com/ssrajadh/sentrysearch.git
cd sentrysearch
uv tool install .

The Python constraint is stated with the reason attached, requires Python 3.11 or 3.12, because PyTorch wheels do not yet support 3.13 and later, and the remedy is shown, uv python install 3.12 followed by uv tool install with the python flag pinning the managed interpreter. Setup continues with sentrysearch init, which prompts for the Gemini API key, writes it to .env and validates it with a test embedding, then sentrysearch index on the footage directory, then sentrysearch search with a query. A manual path exists for those skipping init, copying .env.example and adding a key from aistudio.google.com, and ffmpeg comes either system-wide or through the bundled imageio-ffmpeg automatically.

The index command's dials

Indexing exposes its cost levers as flags. --chunk-duration sets seconds per chunk, defaulting in the shown example output to files of four chunks each, and --overlap sets the overlap between chunks, the two together deciding granularity against embedding volume. Preprocessing downsizes before embedding, with --target-resolution 480 setting target height and --target-fps 5 the frame rate, and --no-preprocess sends raw chunks when fidelity matters more than cost. --no-skip-still embeds every chunk including ones with no visual change, the flag for dashcam footage where parked minutes would otherwise be deduplicated. --rpm caps requests per minute to the cloud API for rate-limited free tier keys, and the backend flags select local or mlx models instead of Gemini. The example output shows the accounting, indexed 12 new chunks from 3 files, totals accumulating across runs.

Search results with confidence gates

Search output shows ranked matches with scores and time ranges, the example listing three results from three camera files with scores from 0.87 down to 0.61, followed by the saved clip's filename encoding source and range. A confidence threshold guards against silent bad matches, default 0.41, or 0.35 on the mlx backend, and below it the tool prompts, no confident match found, show results anyway, rather than trimming. Options reshape the output, --results for count, --output-dir, --no-trim to show low-confidence results with a note instead of prompting, --threshold adjusting the cutoff, --save-top saving the top N clips, and --dedupe setting the cosine similarity ceiling at which a result is dropped as too similar to an already-kept pick, on by default so overlapping chunks of one moment do not fill the list. The --rerank flag asks a VLM to re-rank candidates before trimming, and backend and model auto-detect from the index.

Four backends, including two local ones

Beyond the default Gemini backend, three alternatives exist. The qwen-cloud backend runs through Alibaba DashScope with a DASHSCOPE_API_KEY in .env, and a LiteLLM integration serves AI gateways, with SDK mode needing the litellm extra while proxy mode needs no extra. The local backend runs Qwen3-VL on the machine with no API key, and the MLX backend runs the same local model through Apple Silicon's MLX framework, described as faster and lighter than local. The dependency metadata documents the engineering behind the local path, transformers pinned at 5.4 or newer with a comment explaining that 5.3.0 drops the per-frame video grid expansion and raises StopIteration on multi-frame chunks, fixed in 5.4.0, and a local-quantized extra adding bitsandbytes. Version pins carry reasons, the mark of a project that has been burned and recorded it.

MLX on Apple Silicon, quantized on Metal

The MLX extra's comments are the most technical writing in the repository, and they explain the headline claim, runs the local 2B model twice as fast in under half the memory at the same accuracy. The mechanism is platform-specific quantization, MLX quantizes on Metal, which the PyTorch path cannot do since its only quantization route, bitsandbytes, is CUDA-only, so the 2B model runs at 4-bit in roughly 1.8 GB with no measured loss of retrieval accuracy. The dependency upper bound is deliberate, the video shim in mlx_embedder.py depends on mlx-vlm internals, and the comment instructs maintainers to widen the bound only after re-running the mlx embedder test against the newer release. The platform marker restricts installation to darwin on arm64, Apple Silicon only by construction.

SentryMerge, SentryBlur, and Tesla overlays

The tool's scope extends past search into the post-processing that footage workflows need. Stitch with SentryMerge combines clips, and Redact with SentryBlur handles redaction, both named in the usage table of contents, while a Tesla metadata overlay feature addresses dashcam files specifically, the SentrySearch user's most common source. A tesla optional dependency adding geopy confirms the location-aware aspect of the overlay. The companion tools make the output of search directly usable, a found moment can be assembled into a longer sequence or have faces and plates blurred before sharing, keeping the pipeline inside one tool rather than handing off to an editor for each step. The project's own benchmark release tag, benchmark-clip v1, marked as an asset rather than software, shows the author validating retrieval quality against a fixed clip set.

Supply chain caution, in both directions

Two security postures are visible in the repository. Outbound, the important notice states that github.com/ssrajadh/sentrysearch is the only official home, other sites republishing or mirroring the project are not affiliated with or endorsed by the maintainer, always download from this repository, a warning against the SEO clones that parasite on popular tools. Inbound, the litellm extra's floor is 1.101.0, documented as the tested release and chosen to keep installs clear of 1.82.7 and 1.82.8, the two releases compromised on PyPI, a supply chain incident encoded as a version constraint. The README also lists known warnings as harmless, tracks failed chunks with retrying, documents cache and state files, and offers verbose mode, the operational hygiene of a tool meant to run unattended over large media collections.

Editorial conclusion

Use SentrySearch when hours of footage, dashcams, security cameras, drone files, need finding by description rather than scrubbing, and the trade of per-chunk embedding costs is acceptable for the retrieval speed afterward. Choose the local or MLX backend when footage cannot leave the machine or API costs are unwanted, accepting the hardware requirements. Before deploying, pin Python 3.11 or 3.12 since PyTorch wheels do not yet support 3.13, set an API spending limit when using the Gemini backend, prefer the MLX backend on Apple Silicon for the doubled speed at half the memory, and mind the important notice that the GitHub repository is the only official source, with mirrors unaffiliated.

Frequently asked questions

What is SentrySearch?

SentrySearch is an Apache-2.0 Python CLI for semantic search over video footage. It splits videos into overlapping chunks, embeds them with Google Gemini, Alibaba qwen-cloud or a local Qwen3-VL model into a local ChromaDB, and a text or image query returns ranked matches with the best one automatically trimmed into a clip.

Does SentrySearch work without an API key?

Yes, the local backend runs a Qwen3-VL model on your machine with no API key, and on Apple Silicon the MLX backend runs the same 2B model twice as fast in under half the memory through 4-bit quantization. The qwen-cloud backend needs a DashScope key, and the default Gemini backend needs a Gemini API key.

Which Python versions does SentrySearch support?

Python 3.11 or 3.12, because PyTorch wheels do not yet support 3.13 and later. If your default Python is newer, install a managed 3.12 with uv and pin the tool install with uv tool install --python 3.12.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. ssrajadh/sentrysearch on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ssrajadh-sentrysearch.svg)](https://hysenlabs.com/projects/ssrajadh-sentrysearch)