{"slug": "rosa-a-robotics-foundation-model-serving-system-for-robot-factories", "title": "ROSA: A Robotics Foundation Model Serving System for Robot Factories", "summary": "ROSA, a robotics foundation model serving system for robot factories, improves factory productivity by up to 12.06x over conventional dedicated serving systems, according to a paper proposing the system. ROSA uses shared GPU-pool serving, robotics-aware programming abstractions, and factory-objective-driven scheduling to support multi-model pipelines and per-task performance requirements. The system is implemented on Ray Serve with vLLM, PyTorch, and JAX backends and evaluated on real robots and synthetic workloads.", "body_md": "Robotics foundation models (RFMs) are making general-purpose robots increasingly practical for factory deployments. While RFM serving systems are central to this vision, existing systems are largely shaped by a single-robot, single-model assumption: inference is treated as an edge-computing problem handled by an on-robot or dedicated nearby GPU, and the serving objective is to minimize the latency of a single action model. In this paper, we propose ROSA, an RFM serving system for robot factories designed around three key principles. First, ROSA adopts shared GPU-pool serving, allowing a fleet of robots to access powerful server-class GPUs over the network in order to improve inference performance, battery duration, and GPU utilization. Second, ROSA provides a robotics-aware programming abstraction and system design that supports multi-model pipelines, per-task performance requirements, and failure handling. Third, ROSA uses factory-objective-driven scheduling to maximize SLO-qualified factory productivity rather than minimizing individual request latency. We implement ROSA on top of Ray Serve for distributed orchestration, with vLLM, PyTorch, and JAX as model-serving backends, and evaluate it on both real robots and synthetic large-scale workloads. The results show that ROSA improves factory productivity by up to 12.06x over conventional dedicated serving systems.", "url": "https://wpnews.pro/news/rosa-a-robotics-foundation-model-serving-system-for-robot-factories", "canonical_source": "https://research.nvidia.com/publication/2026-07_rosa-robotics-foundation-model-serving-system-robot-factories", "published_at": "2026-08-05 20:50:22+00:00", "updated_at": "2026-08-09 12:16:23.541735+00:00", "lang": "en", "topics": ["robotics", "ai-infrastructure"], "entities": ["ROSA", "Ray Serve", "vLLM", "PyTorch", "JAX"], "alternates": {"html": "https://wpnews.pro/news/rosa-a-robotics-foundation-model-serving-system-for-robot-factories", "markdown": "https://wpnews.pro/news/rosa-a-robotics-foundation-model-serving-system-for-robot-factories.md", "text": "https://wpnews.pro/news/rosa-a-robotics-foundation-model-serving-system-for-robot-factories.txt", "jsonld": "https://wpnews.pro/news/rosa-a-robotics-foundation-model-serving-system-for-robot-factories.jsonld"}}