{"slug": "physical-ai-smart-spaces-a-large-scale-benchmark-for-multi-camera-3d-perception", "title": "Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces", "summary": "NVIDIA released Physical AI Smart Spaces, a benchmark it calls the first to simultaneously provide large-scale, multi-class, multi-camera 3D perception data for indoor smart spaces, containing over 280 hours of synchronized 1080p footage from nearly 1,800 cameras in warehouses, hospitals, and retail venues. The benchmark includes automatic annotations for multi-camera identities, 2D and 3D bounding boxes, camera calibration, and depth, and spans Isaac Sim synthetic generation, Cosmos Transfer appearance augmentation, and real-world Sim2Real evaluation across two warehouse deployments. Its central contribution is a 3D instantiation of Higher Order Tracking Accuracy (HOTA), extending 2D box-based tracking evaluation to 3D locations and 3D boxes, with the release available at https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces.", "body_md": "arXiv:2610.02580v1 Announce Type: new \nAbstract: Physical AI Smart Spaces is, to the best of our knowledge, the first benchmark to simultaneously provide large-scale, multi-class, and multi-camera 3D perception data for indoor smart spaces. It contains over 280 hours of synchronized 1080p footage captured by nearly 1,800 cameras in warehouses, hospitals, retail venues, and similar settings, together with automatic annotations for multi-camera identities, 2D bounding boxes, 3D bounding boxes, camera calibration, and depth where available. The benchmark spans Isaac Sim synthetic generation, Cosmos Transfer appearance augmentation, and real-world Sim2Real evaluation. For the real-world target, we include two warehouse deployments with time-synchronized streams, automatic VGGT-based calibration, and a 3D labeling interface that projects world-frame 3D boxes into each view for cross-camera verification. We describe the dataset scope, annotation and calibration schema, generation workflow, benchmark protocols, and official evaluation system, which standardizes submission format, and leaderboard reporting. A central contribution is a 3D instantiation of Higher Order Tracking Accuracy (HOTA), extending the usual 2D box-based tracking evaluation to 3D locations and 3D boxes. We further report empirical baselines from the AI City Challenge leaderboards, showing how methods evolve from person-only 3D location tracking to multi-class 3D box tracking under realistic smart-space constraints. The release is available at https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces.", "url": "https://wpnews.pro/news/physical-ai-smart-spaces-a-large-scale-benchmark-for-multi-camera-3d-perception", "canonical_source": "https://arxiv.org/abs/2610.02580", "published_at": "2026-10-05 04:00:00+00:00", "updated_at": "2026-10-05 04:13:48.589091+00:00", "lang": "en", "topics": ["computer-vision", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["NVIDIA", "Physical AI Smart Spaces", "Isaac Sim", "Cosmos Transfer", "HOTA", "AI City Challenge", "Hugging Face", "VGGT"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/physical-ai-smart-spaces-a-large-scale-benchmark-for-multi-camera-3d-perception", "markdown": "https://wpnews.pro/news/physical-ai-smart-spaces-a-large-scale-benchmark-for-multi-camera-3d-perception.md", "text": "https://wpnews.pro/news/physical-ai-smart-spaces-a-large-scale-benchmark-for-multi-camera-3d-perception.txt", "jsonld": "https://wpnews.pro/news/physical-ai-smart-spaces-a-large-scale-benchmark-for-multi-camera-3d-perception.jsonld"}}