Model or dataset
NVIDIA-AI-Blueprints/video-search-and-summarization avatar
NVIDIA-AI-Blueprints/video-search-and-summarization

NVIDIA VSS Blueprint: A Video Search and Summarization Agent on NIM Microservices

NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

1,880 stars396 forksPythonNOASSERTION

At a glance

What is it?
The NVIDIA AI Blueprint for Video Search and Summarization is a GPU-accelerated reference architecture for building video agents that search archives, verify alerts, and summarize long recordings. It is a deployable stack, not a library, and that shapes everything about how you adopt it.
Who is it for?
Adopt the VSS Blueprint if you already run NVIDIA GPUs and want a working multi-service video agent rather than a research notebook: the repository gives you real-time video intelligence, downstream analytics, and an MCP-based agent layer in one tree, with tagged releases v3.2.1, v3.2.0 and v3.1.0 to diff against.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the VSS Blueprint actually solves

Most teams that want to ask questions of video end up gluing together a detector, a tracker, a captioning model, a vector store and a chat interface. Each piece works. The seams do not. The NVIDIA AI Blueprint for Video Search and Summarization (VSS) exists to remove those seams by shipping the whole pipeline as a reference architecture you deploy rather than assemble.

The README describes the target plainly: the blueprint addresses "the challenge of deploying visual agents capable of interacting with large volumes of video data, both stored and streamed." That sentence is the whole scope. It is not a model, and it is not a SaaS product. It is a set of microservices plus agent workflows, published on GitHub under NVIDIA-AI-Blueprints and also reachable as a hosted demo at build.nvidia.com.

The intended users are named in the README as two groups: video analysts and IT engineers who want one-click deployment and plug-and-play models, and GenAI developers or ML engineers who need to modify pipelines for their own datasets and fine-tune LLMs. Those two audiences pull in opposite directions, and the repository structure reflects the tension. There is a deploy directory for people who want to stand it up, and there are libs, services, skills and tools directories for people who want to take it apart.

How the VSS architecture moves a frame to an answer

The blueprint is organized into three processing layers, and the README is explicit about the data flow between them.

The first layer is real-time video intelligence. It "extracts rich visual features, semantic embeddings, and contextual understanding from video data in real-time, publishing results to a message broker for downstream analytics and agentic workflows." The message broker is the hinge of the design. Nothing in this layer answers a question directly; it emits structured output and moves on. The README states it provides three core microservices for processing video streams, though the README does not name them.

The second layer, downstream analytics, consumes that broker stream and enriches it, "transforming raw detections into actionable insights and verified alerts." This is where object detection and tracking become trajectories and incidents. The alert verification workflow is the clearest illustration: perception generates alerts, and a VLM then re-examines them to cut false positives.

The third layer is the agent. It uses the Model Context Protocol (MCP) to reach video analytics data, incident records and vision processing through what the README calls "a unified tool interface." The tools behind that interface include VLM-based video understanding, semantic search over video embeddings, long video summarization, and snapshot or clip retrieval. The model components listed are Cosmos3 Nano Reasoner as the VLM and NVIDIA Nemotron 3.5 Lightning 30B A3B as the LLM, both served as NIM microservices.

The design consequence is worth stating directly: because the agent talks to MCP tools rather than to a database, you can swap the retrieval strategy without rewriting the agent, and you can call the same tools from a different agent. That is the real product here.

Installing the VSS Blueprint and running a first workflow

The README points to a Quickstart Guide anchor and to docs.nvidia.com/vss/latest/quickstart.html for the Q&A and report generation workflow, which it describes as video retrieval, VLM-based Q&A, and report generation on short video clips. The README does not include the literal shell commands, so the honest position is that you should take the commands from the quickstart page rather than from this article.

What the README does establish is the shape of the setup. Before you run anything, it directs you to Prerequisites and Hardware Requirements sections, which sit above the Quickstart Guide in the table of contents. Given that the stack is GPU-accelerated and built on NIM microservices, those two sections are where the deployment either proceeds or stops.

The repository layout tells you where the pieces live. Configuration and deployment material sits under deploy/, service code under services/, shared code under libs/, agent tooling under tools/ and skills/, and the prose documentation under docs/ with a fern/ directory suggesting the docs site is generated. A release_metadata.yaml at the top level is the file to read if you want to know what a given tag contains.

If you want to see the system before deploying it, the README offers the hosted experience at build.nvidia.com/nvidia/video-search-and-summarization. That is the cheapest way to judge whether the Q&A and summarization output is good enough for your footage.

For version pinning, the releases are tagged v3.2.1 (2026-07-23), v3.2.0 (2026-06-16) and v3.1.0 (2026-03-18). Checking out a tag rather than the develop branch is the difference between a reproducible deployment and a moving one, since develop is the default branch.

The VSS workflows, and where each one breaks down

The README lists five agent workflows, and they are not equally mature. Reading the table carefully is the fastest way to avoid disappointment.

The Q&A and Report Generation workflow is the quickstart and operates on short video clips. Alert Verification runs real-time perception to generate alerts and then uses a VLM to verify them, which is a sensible way to spend GPU time only on candidates. Real-Time Alerts pushes streams continuously through a VLM for anomaly detection, which is the most expensive pattern in the set because every frame is a candidate. Long Video Summarization handles extended recordings "through chunking and aggregation of dense captions," an approach that inherits the classic summarization failure mode: errors in individual chunk captions compound when you aggregate, and the final summary can be confidently wrong about a moment that was captioned badly.

Video Search is marked alpha in the README's own table. It performs natural language search across archives using video embeddings. An alpha label on the retrieval path matters, because retrieval quality is what determines whether the agent is useful or merely impressive in a demo.

The broader limitation is the one the README admits in its Target Audience section: the blueprint is "designed for ease of setup with extensive configuration options, requiring technical expertise." Those two claims sit awkwardly together. Extensive configuration options are not ease of setup; they are a large surface area. If your team has no one comfortable with container orchestration and GPU allocation, the one-click framing will not survive contact with your environment.

How VSS differs from assembling your own video RAG stack

The obvious alternative is to build the same thing yourself from open components: a detector such as YOLO or a DETR variant, a tracker, a captioning or VLM model, a vector database, and an orchestration layer you write. Plenty of teams do exactly this.

The difference is not model quality. It is the interfaces. In a self-assembled stack, you own the schema that carries detections from the detector to the analytics layer, and you own the contract between analytics and the agent. Every component upgrade risks breaking that contract. The VSS Blueprint fixes those boundaries in advance: a message broker between real-time intelligence and downstream analytics, and MCP between the agent and the data. The cost of that decision is that you inherit NVIDIA's choices, including the NIM microservices and the specific models listed (Cosmos3 Nano Reasoner and Nemotron 3.5 Lightning 30B A3B).

A second alternative is to use only the real-time layer and skip the agent entirely, publishing to your own broker consumer. The README explicitly supports this, saying the architectures can be used "in existing applications, as standalone microservices, or as part of a larger vision agent." That partial adoption path is the most underrated option in the repository, and it is the one to consider if your existing analytics platform already handles incidents.

What you give up either way is portability. A self-built stack can run on whatever accelerator you have. This one is tied to the NVIDIA stack by construction.

Maintenance, releases and the licence question

The repository is not archived, and its last push was on 2026-09-10, which is recent. The release cadence visible in the tags is roughly every one to two months across 2026: v3.1.0 in March, v3.2.0 in June, v3.2.1 in July. That is a fast-moving reference architecture, and it has a practical consequence: pinning to a tag is not optional if you want reproducible builds, because develop is the default branch and is where work lands first.

Upgrade cost is the part the README does not document. There is no rollback procedure described, no migration guide between v3.1.0 and v3.2.0 visible in the repository, and no compatibility matrix for the NIM microservices. If you deploy this, budget for reading release_metadata.yaml at each tag and for testing the agent workflows you depend on, particularly the alpha video search path.

The licence is the sharpest edge. The GitHub metadata reports the licence as NOASSERTION, meaning no standard SPDX identifier was detected. The repository carries a LICENSE file plus LICENSE-3rd-party.txt and a separate LICENSE.DATA, which strongly suggests the code, the bundled third-party components, and the data are governed by different terms. That is common for blueprints that ship sample footage or model artefacts. It also means you cannot assume an Apache-2.0-style grant covers everything in the tree. Read all three files before you ship anything derived from this repository, and route the question to your own counsel. Nothing here is legal advice.

Editorial conclusion

Adopt the VSS Blueprint if you already run NVIDIA GPUs and want a working multi-service video agent rather than a research notebook: the repository gives you real-time video intelligence, downstream analytics, and an MCP-based agent layer in one tree, with tagged releases v3.2.1, v3.2.0 and v3.1.0 to diff against. Do not adopt it if you need a single pip install, a CPU-only path, or a permissive licence you have already cleared, because the LICENSE file is NOASSERTION and ships alongside LICENSE-3rd-party.txt and LICENSE.DATA. Before committing, read the Hardware Requirements and Prerequisites sections of the README against your actual GPU inventory, and check the docs.nvidia.com/vss pages for the workflow you intend to run, since the README notes that video search is still alpha.

Frequently asked questions

How do I summarize a video with the NVIDIA VSS Blueprint?

The README lists a Long Video Summarization workflow that analyzes extended recordings through chunking and aggregation of dense captions. It also lists a Q&A and Report Generation quickstart for shorter clips, which the documentation covers at docs.nvidia.com/vss/latest/quickstart.html.

What is a summary video in the context of NVIDIA VSS?

The blueprint does not use the phrase summary video. Its Long Video Summarization workflow takes extended recordings, chunks them, aggregates dense captions, and produces a summary of the footage rather than a separate edited video file.

Can AI watch a video and summarize it in the NVIDIA VSS Blueprint?

Yes, that is one of the blueprint's stated capabilities: the README lists summarizing hours of video alongside natural language search, visual Q&A, and alert verification. The summarization path works by chunking long recordings and aggregating dense captions rather than by feeding the whole video to a model at once.

What are some good apps for summarizing videos compared with the NVIDIA VSS Blueprint?

The blueprint is not a consumer app; it is a reference architecture you deploy, built on NIM microservices, a VLM and an LLM, with an MCP-based agent layer. If you want a hosted look at its summarization and Q&A output before deploying, the README points to build.nvidia.com/nvidia/video-search-and-summarization.

What is the NVIDIA video search and summarization (VSS) blueprint?

It is a GPU-accelerated reference architecture from NVIDIA for building video analytics agents, combining vision language models, LLMs, RAG and NVIDIA NIM microservices. The repository implements the blueprint across three layers: real-time video intelligence, downstream analytics, and agentic and offline processing.

Official sources

  1. Issues
  2. NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-ai-blueprints-video-search-and-summarization.svg)](https://hysenlabs.com/projects/nvidia-ai-blueprints-video-search-and-summarization)