Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3 NVIDIA released VSS Blueprint 3.3, adding a Build Vision Agent skill (vss-build-vision-ai) and Adaptive Efficient Video Sampling (EVS) that cut VLM input tokens by 80% for a 60-minute summary and raise concurrent streams by 46% on the same GPU. NVIDIA said the Build Vision Agent skill lets coding agents such as Claude Code and Codex build and deploy a bottling-line overflow agent from a single prompt in under 30 minutes for a few dollars of coding-agent usage. The blueprint connects vision-language models including NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, retrieval-augmented generation and Model Context Protocol tools for video search, visual Q&A, alerts and reporting. Vision-language models have made it possible to build visual AI agents that understand video at production scale. The harder problem is turning that capability into a maintainable system that combines ingestion, stream processing, event detection, retrieval, summarization, and reporting. The NVIDIA Metropolis Blueprint for Video Search and Summarization VSS https://build.nvidia.com/nvidia/video-search-and-summarization and its agent skills help developers build visual AI agents https://www.nvidia.com/en-us/use-cases/video-analytics-ai-agents/ faster. VSS connects vision-language models VLMs such as NVIDIA Cosmos https://www.nvidia.com/en-us/ai/cosmos/ , LLMs such as NVIDIA Nemotron https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/ , retrieval-augmented generation RAG , and Model Context Protocol MCP tools to turn live and recorded video into natural-language search, visual Q&A, verified alerts and automated reporting. VSS Blueprint 3.3 also reduces costs from both sides: faster application composition with the new Build Vision Agent skill vss-build-vision-ai , and cheaper VLM processing at runtime with Adaptive Efficient Video Sampling EVS . This post covers both sides of that cost reduction. On the development side, a single prompt builds and deploys a bottling-line overflow agent in under 30 minutes, with a few dollars of coding-agent usage. On the runtime side, Adaptive EVS delivers 80% fewer VLM input tokens for a 60-minute summary and 46% more concurrent streams on the same GPU. To learn more, join us live https://www.youtube.com/watch?v=PQJKs1dyK7I on Oct. 1 at 9 a.m. PT, where we will build a visual AI agent from a single prompt. Why visual AI agents are expensive A production visual AI agent https://www.nvidia.com/en-us/use-cases/video-analytics-ai-agents/ usually spans more than one workflow. A smart city application may need vehicle detection, collision alerting, searchable incident clips, hourly summaries, and an operator report. A warehouse application may need people and forklift tracking, near-miss alerts, SOP checks, and follow-up Q&A. Each workflow is useful by itself, but real deployments become valuable when those workflows work together. That creates 3 recurring cost drivers: - Development cost: Teams must choose and connect microservices, shared services such as Kafka, Redis, Elasticsearch, and Video IO and Storage VIOS , model endpoints, environment variables, and APIs without duplicating infrastructure. - Operating cost: Visual AI workloads can lead to heavy token usage, every additional stream, frame window, prompt, and visual token can increase GPU usage, queueing delay, and end-to-end summarization latency. - Change cost: Teams must move from proof of concept to production, add capabilities, and keep configuration, documentation, and operations aligned. VSS 3.3 updates for building and running video analytics AI agents VSS Agent Skills let coding agents such as Claude Code, Codex, or any agentskills.io-compatible agent deploy and operate VSS from natural-language requests. VSS 3.3 adds two updates that reduce cost on both sides of a deployment: - Build Vision Agent skill vss-build-vision-ai : Combines VSS workflows such as alerting, search, and summarization into one application for use cases such as SOP compliance or traffic management, and extends a running deployment without rebuilding the whole stack. - Adaptive EVS: This new feature reduces redundant VLM processing by pruning the visual tokens for parts of a frame that did not change compared to the previous one, and by batching VLM work around the moments when something happens. EVS already ships in vLLM and the Cosmos NIM microservices at a fixed pruning rate. The adaptive version in VSS 3.3 is integrated into the real-time VLM microservice and decides which tokens to keep per patch and per frame. The Build Vision Agent skill reduces development and change costs by composing and extending deployments. Adaptive EVS lowers operating cost by reducing VLM tokens and GPU time spent on unchanged video. Reduce development cost with the Build Vision Agent skill Earlier VSS skills handled individual operations such as deployment, camera setup, summarization, search, alerts, and analytics. VSS 3.3 organizes them as deployment skills, operation skills, tools, and benchmarks, with vss-build-vision-ai composing the rest. Developers describe the application they want, and the Build Vision Agent skill translates that intent into a deployment plan spanning profiles, microservices, configuration, and runtime operations. Rather than generate a deployment from scratch, the skill starts with the closest of four validated developer profiles, each a complete, tested stack for one workflow see Table 1, below . The skill calls that starting profile the Foundation and changes only what the request requires. | Profile | Capability | | base | VLM dense captioning and Q&A on clips | | alerts | Real-time VLM alerting, or RT-CV detection with behavior analytics and VLM alert verification | | lvs | Long video summarization | | search | Object and video embeddings with agentic search | Table 1: The four developer profiles a build can start from. The skill selects one as the Foundation and computes the smallest delta on top of it The skill then computes the smallest delta: adds or removes only exact service keys, keeps only the services that a requested capability actually reaches, and converges shared roles onto one instance. Two capabilities that both need a detector get one detector. Two that both need Kafka and Elasticsearch share one message bus and one Elasticsearch deployment, each writing its own indices. When the rules cannot settle a choice, the skill asks one structured question instead of guessing. What the skill automates - Maps the application goal to the required VSS workflows and microservices. - Combines workflows such as alerting, search, and summarization in one deployment plan. - Reuses shared infrastructure, including VIOS, Kafka, Redis, Elasticsearch, HAProxy ingress, and MCP services. - Generates the deployment as a self-contained build: builds/