Vision-language models have made it possible to build visual AI agents that understand video at production scale. The harder problem is turning that capability into a maintainable system that combines ingestion, stream processing, event detection, retrieval, summarization, and reporting.
The NVIDIA Metropolis Blueprint for Video Search and Summarization (VSS) and its agent skills help developers build visual AI agents faster. VSS connects vision-language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, retrieval-augmented generation (RAG), and Model Context Protocol (MCP) tools to turn live and recorded video into natural-language search, visual Q&A, verified alerts and automated reporting.
VSS Blueprint 3.3 also reduces costs from both sides: faster application composition with the new Build Vision Agent skill (vss-build-vision-ai), and cheaper VLM processing at runtime with Adaptive Efficient Video Sampling (EVS).
This post covers both sides of that cost reduction. On the development side, a single prompt builds and deploys a bottling-line overflow agent in under 30 minutes, with a few dollars of coding-agent usage. On the runtime side, Adaptive EVS delivers 80% fewer VLM input tokens for a 60-minute summary and 46% more concurrent streams on the same GPU.
To learn more, join us live on Oct. 1 at 9 a.m. PT, where we will build a visual AI agent from a single prompt.
Why visual AI agents are expensive #
A production visual AI agent usually spans more than one workflow. A smart city application may need vehicle detection, collision alerting, searchable incident clips, hourly summaries, and an operator report. A warehouse application may need people and forklift tracking, near-miss alerts, SOP checks, and follow-up Q&A. Each workflow is useful by itself, but real deployments become valuable when those workflows work together.
That creates 3 recurring cost drivers:
- Development cost: Teams must choose and connect microservices, shared services such as Kafka, Redis, Elasticsearch, and Video IO and Storage (VIOS), model endpoints, environment variables, and APIs without duplicating infrastructure.
- Operating cost: Visual AI workloads can lead to heavy token usage, every additional stream, frame window, prompt, and visual token can increase GPU usage, queueing delay, and end-to-end summarization latency.
- Change cost: Teams must move from proof of concept to production, add capabilities, and keep configuration, documentation, and operations aligned.
VSS 3.3 updates for building and running video analytics AI agents #
VSS Agent Skills let coding agents such as Claude Code, Codex, or any agentskills.io-compatible agent deploy and operate VSS from natural-language requests. VSS 3.3 adds two updates that reduce cost on both sides of a deployment:
- Build Vision Agent skill (vss-build-vision-ai): Combines VSS workflows such as alerting, search, and summarization into one application for use cases such as SOP compliance or traffic management, and extends a running deployment without rebuilding the whole stack.
- Adaptive EVS: This new feature reduces redundant VLM processing by pruning the visual tokens for parts of a frame that did not change compared to the previous one, and by batching VLM work around the moments when something happens. EVS already ships in vLLM and the Cosmos NIM microservices at a fixed pruning rate. The adaptive version in VSS 3.3 is integrated into the real-time VLM microservice and decides which tokens to keep per patch and per frame.
The Build Vision Agent skill reduces development and change costs by composing and extending deployments. Adaptive EVS lowers operating cost by reducing VLM tokens and GPU time spent on unchanged video.
Reduce development cost with the Build Vision Agent skill #
Earlier VSS skills handled individual operations such as deployment, camera setup, summarization, search, alerts, and analytics. VSS 3.3 organizes them as deployment skills, operation skills, tools, and benchmarks, with vss-build-vision-ai composing the rest.
Developers describe the application they want, and the Build Vision Agent skill translates that intent into a deployment plan spanning profiles, microservices, configuration, and runtime operations.
Rather than generate a deployment from scratch, the skill starts with the closest of four validated developer profiles, each a complete, tested stack for one workflow (see Table 1, below). The skill calls that starting profile the Foundation and changes only what the request requires.
| Profile | Capability |
| base | VLM dense captioning and Q&A on clips |
| alerts | Real-time VLM alerting, or RT-CV detection with behavior analytics and VLM alert verification |
| lvs | Long video summarization |
| search | Object and video embeddings with agentic search |
Table 1: The four developer profiles a build can start from. The skill selects one as the Foundation and computes the smallest delta on top of it
The skill then computes the smallest delta: adds or removes only exact service keys, keeps only the services that a requested capability actually reaches, and converges shared roles onto one instance.
Two capabilities that both need a detector get one detector. Two that both need Kafka and Elasticsearch share one message bus and one Elasticsearch deployment, each writing its own indices. When the rules cannot settle a choice, the skill asks one structured question instead of guessing.
What the skill automates
- Maps the application goal to the required VSS workflows and microservices.
- Combines workflows such as alerting, search, and summarization in one deployment plan.
- Reuses shared infrastructure, including VIOS, Kafka, Redis, Elasticsearch, HAProxy ingress, and MCP services.
- Generates the deployment as a self-contained build: _builds/<name>/override.env (the Foundation, the effective Compose profiles, and only the settings you changed), compose.yml, and resolved.yml, one flattened Compose file produced with docker compose config that deploys on its own. The repository’s deploy/docker/ tree is never modified.
- Shows an architecture diagram for review before anything is written or deployed, then runs validation, deployment, and readiness checks so the developer can verify the stack before building application logic on top.
- Asks whether to deploy an agent harness. The default is NemoClaw , a host-side sandbox with the VSS skills installed; answering no produces a headless stack driven by the VSS CLI.
- Extends a running deployment through a smaller delta that reuses existing services.
This shortens discovery, makes composition repeatable, and avoids duplicate ingestion, storage, messaging, and analytics infrastructure across workflows.
Build a bottling line visual AI agent with VSS #
For an orange juice bottling line, the team wants an agent that watches filler and capper cameras, alerts on overflows or spills, searches past incidents, and generates shift reports.
Manual development would connect event detection, alert verification, storage, search ingestion, summarization, and reporting. With the Build Vision Agent skill, development begins with the desired outcome.
Sample prompt
Build a VSS vision agent for an orange juice bottling line. Use two RTSP cameras on the filler and capper. Detect bottle overflows and juice spills, verify each alert with the VLM, make alert clips searchable, and generate a shift report for the line supervisor.
What the agent assembles
- VIOS-backed ingestion for RTSP cameras and recorded clips.
- Real-time detection, tracking, captioning, or VLM-based alerts, depending on the Foundation profile.
- Behavior analytics or rules for overflow, spill, and line-stoppage events.
- VLM verification that confirms alerts and explains its reasoning.
- Natural-language search across validated clips and indexed video.
- Summarization and reporting for operator handoff and incident review.
- Shared messaging, storage, APIs, and observability.
On a two-GPU RTX PRO 6000 Blackwell host, the skill combines search and alerts by reusing existing services and adding only an alert bridge and real-time VLM. FP8 Cosmos 3 Nano shares the detector’s GPU, avoiding duplication. A recorded alerts build reached a live, previewable deployment in under 30 minutes.
The result is a reusable pattern: ingest video once, share evidence across workflows, and give operators natural-language search, alerts, summaries, and reports.
How adaptive EVS reduces operating cost #
Once the bottling line agent from the example above is running, its cameras keep producing video around the clock, and most of each frame never changes: the filler, the guards, the floor. Only the bottles moving through it do, and an overflow is rare.
Every frame window still becomes visual context for the VLM, so the model spends most of its compute re-reading regions that look exactly like the frame before. That VLM processing is one of the main runtime cost drivers in a deployed visual AI agent, and it is the cost Adaptive EVS goes after.
What Adaptive EVS changes in the pipeline
- Dynamic pruning. Each patch is compared with the prior frame using cosine similarity; unchanged patches are dropped before reaching the language model.
- Event-aware batching. Token retention indicates activity: clips above about 70% are batched as events, those below about 30% are dropped or flushed, and the rest run normally.
Performance impact #
On an NVIDIA RTX PRO 6000 Blackwell running Cosmos 3 Super FP8, Adaptive EVS:
- Cut alert contextualization latency 17%, from 1,021 ms to 844 ms, while similarly reducing token usage.
- Increased concurrent real-time VLM streams 46%, from 13 to 19.
- Summarized a 60-minute video in about half the time with 80% fewer VLM input tokens.
Results vary with scene motion, chunk length, and similarity threshold; benchmark representative footage before choosing production defaults.
Adaptive EVS is most useful when VLMs read many frames and produce short responses, as in dense captioning, long-video summarization, and alert verification. It offers less benefit for long outputs from few frames, runs inside the RT-VLM container rather than against remote endpoints, and is optional. Enable it in override.env, then benchmark accuracy, throughput, and latency on representative footage:
VIA_EVS_SESSION=true
VLM_VIDEO_PRUNING_RATE=0.5 # 0.0 to 1.0; higher prunes more
VLLM_EVS_SIMILARITY_THRESHOLD=0.2
Cost impact: What changes for teams #
- Lower integration effort through natural-language composition of multi-workflow applications.
- Less duplicate infrastructure across alerting, search, summarization, reporting, and Q&A.
- Higher GPU efficiency by pruning unchanged visual regions.
- Lower summarization and incident-review latency by skipping uneventful video.
- Easier extension through incremental deltas that reuse a running deployment.
Together, these changes reduce upfront development work and the GPU work required to process video with VLMs across environments and applications.
Getting started With VSS 3.3 #
- Clone the VSS Blueprint repository and check out the branch containing the 3.3 skills.
- Install the VSS skills in your coding agent’s standard skills directory.
- Describe the desired agent, including video sources, workflows, and deployment constraints, or say “build a vision agent” for guidance.
- Review the architecture diagram and _builds/<name>/override.env, especially GPU placement, model endpoints, ports, storage, and security boundaries.
- For RT-VLM workloads, enable Adaptive EVS, tune the pruning rate, and benchmark accuracy, throughput, and latency on representative video.
- Deploy on a trusted, isolated network with authentication, TLS, rate limiting, and external controls.
Example setup commands
git clone https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git
cd video-search-and-summarization
Ask your coding agent to install the skills:
Read skills/README.md and every SKILL.md under skills/. Install each skill for this host
using the standard skills directory, symlinking rather than copying so a git pull keeps
them current.
After the skills are installed, you can start with a prompt such as:
Build a VSS vision agent that combines alert verification, natural-language video search,
and hourly summarization for my warehouse cameras. Reuse existing Kafka and Elasticsearch
services where possible, and produce a deployment plan before running Docker Compose.
Going further #
- Clone the VSS Blueprint repository
- Install the VSS Agent Skills
- Try the Build Vision Agent skill with the sample prompt above on your own cameras