- World Labs says its Real-to-Sim-to-Real engine converts one physical task into thousands of variations covering object configuration, clutter, lighting, physics, robot state and camera viewpoint. [1] - Policies trained entirely in simulation ran autonomously for one hour on RB-Y1, YAM, Flexiv and xArm systems; ALOHA was used for separate training and evaluation demonstrations. [1][2] - The company evaluated policy ranking with 2,000 simulated and 100 real cube-handover trials per checkpoint, but did not disclose success rates for the one-hour runs or provide an independent replication.
[1] World Labs says it has built a robot-training system that turns one physical task into thousands of controllable simulations, then transfers policies trained without real-world data onto physical robots. In demonstrations published July 28, the policies ran autonomously for one hour on tasks involving cables, power cords, test tubes and thin objects in clutter. [1]
The startup, founded by Fei-Fei Li, calls the system Real-to-Sim-to-Real, or R2S2R. It comes from SceniX, a robotics and simulation company that joined World Labs on July 21. [3]
From one task to thousands of variations #
R2S2R begins with a physical robot, its sensors, the surrounding environment, the objects involved and a task demonstration. World Labs says it reconstructs those elements as an interactive simulation that preserves task-relevant observations and physical interactions, then varies appearance, object configuration, clutter, friction, robot state and camera viewpoint. [1]
The company showed examples involving rigid, articulated and deformable objects. They included cable sliding and plugging with a YAM robot, elastic-cable insertion with ALOHA, power-cord routing around a refrigerator with RB-Y1, and a bimanual cube handover with ALOHA. [1]
This differs from World Labs’ earlier public robotics work with Marble, its generative 3D-world product. Marble’s published case study describes generating scenes and exporting Gaussian splats and collision meshes into tools including NVIDIA Isaac Sim, MuJoCo and RoboSuite. R2S2R instead focuses on reconstructing a particular physical task closely enough to train and evaluate robot policies. [4]
The hardware demonstrations #
World Labs says policies trained entirely in simulation transferred directly to physical systems and ran without human intervention for one hour. The hour-long demonstrations covered power-cord manipulation with RB-Y1, sustained cable manipulation with YAM, tight-tolerance test-tube transfer with Flexiv, and marker and pencil singulation with xArm. ALOHA was used for the bimanual box-packing and cube-handover demonstrations, but the release does not identify an ALOHA one-hour run. [1][2]
The release gives task descriptions and run durations, but it does not provide success rates, cycle counts, interruption counts or failure distributions for those one-hour demonstrations. The evidence establishes extended autonomous runs under the company’s test conditions; it does not show how often the robots would complete the same tasks across repeated trials or in less controlled environments. [1]
That distinction matters because simulation-trained policies can still degrade on real hardware when dynamics, sensor noise or environmental conditions differ. NVIDIA’s SPARR project, for example, describes the sim-to-real gap as a continuing problem and adds a real-world residual policy to correct discrepancies after simulation pretraining. [5]
A narrower evaluation claim #
World Labs separately tested whether its simulator could predict which policies would perform better on real hardware. The test used an ALOHA cube-handover task and compared policy architectures, training configurations and checkpoints in simulated and physical rollouts. Each checkpoint received 2,000 simulated trials—1,000 in-distribution and 1,000 out-of-distribution—and 100 real trials split evenly between the two settings. [1]
The company says simulation preserved the relative ranking of policies, tracked improvements and plateaus during training, and identified similar spatial regions of success and failure. It also showed paired near-boundary successes and matching failures between simulation and reality. [1]
Those results support a practical use for the system: screening weak checkpoints in simulation and reserving physical tests for the most promising candidates. They do not show that simulation reproduces real-world success rates exactly. World Labs explicitly frames the simulator’s value as supporting decisions such as ranking policies and locating failure regions. [1]
The July 28 post does not link to a peer-reviewed paper, public benchmark code or an independent replication. Whether the approach transfers to longer-horizon tasks, unfamiliar objects, outdoor settings or everyday clutter remains unresolved.
Companies mentioned #
Further sources #
[1] World Labs’ July 28, 2026 technical release describing the Real-to-Sim-to-Real … ↗ [2] The Decoder’s August 15, 2026 report clarifying that ALOHA was one test platfor… ↗
[[3] World Labs’ July 21, 2026 announcement that SceniX joined the company, and Worl… ↗](https://www.worldlabs.ai/blog/scenix)
[[4] World Labs’ November 12, 2025 case study describing Marble’s generated environm… ↗](https://www.worldlabs.ai/case-studies/1-robotics)
[[5] NVIDIA Research’s SPARR project page describing continuing sim-to-real degradat… ↗](https://research.nvidia.com/labs/srl/projects/sparr/)
The stories that matter, in one email. Free — unsubscribe anytime.