The Machines that Make the Machines NVIDIA's Seattle Robotics Lab, working with the NVIDIA Isaac engineering team and contract manufacturer Foxconn, has demonstrated that robots can assemble GB300 tester trays, tackling busbar assembly and multi-connector insertion after nine years of lab research. The lab said the tasks required flexible rather than fixed automation because part geometry, appearance and dynamics are not precisely known, and that robotics research success rates of 50%-60%, or 80%-90% to call a task solved, fall short of the performance needed for production. The work challenges the lab's assumptions about where end-to-end learning is most powerful versus a modular stack. How we taught robots to assemble GB300 tester trays and what it taught us about robot learning, mechanical intelligence, and good old-fashioned engineering The NVIDIA Grace Blackwell GB300 superchip https://developer.nvidia.com/blog/inside-nvidia-blackwell-ultra-the-chip-powering-the-ai-factory-era/ nvidia grace blackwell ultra superchip is the engine of AI training and inference for modern foundation models. However, the process of assembling GB300 trays https://docs.nvidia.com/dgx/dgxgb200-user-guide/hardware.html requires skilled physical labor in factories across the world. The NVIDIA Seattle Robotics Lab SRL https://research.nvidia.com/labs/srl/ , in collaboration with the NVIDIA Isaac https://developer.nvidia.com/isaac engineering team, has asked the central question: Can we build intelligent robots to assemble such complex systems? The answer is yes , but the path has been challenging. Enabling robots to assemble GB300 trays is one of the hardest problems SRL has tackled in nine years as a highly-prolific research lab https://research.nvidia.com/labs/srl/publication/ . The problem has also challenged our assumptions about where end-to-end learning is most powerful, and where a modular stack might be more effective. Problems at hand: Moravec returns We focused on assembling GB300 tester trays , which are used to verify the functionality and performance of GB300 compute modules prior to shipping and deployment. Based on recommendations from the NVIDIA Operations Team and Foxconn, a GB300 contract manufacturer, we zeroed in on two critical tasks: 1 busbar assembly and 2 multi-connector insertion. In busbar assembly , a long, heavy busbar i.e., a metallic bar that transmits high current must be grasped, transported to the tray, inserted, and screwed in at 16 locations. Fixtures and clamps must also be inserted and removed. In multi-connector insertion , two large and two small cable-mounted electrical connectors must be lifted from the tray and inserted into tight-clearance sockets. The real-world experimental setup is shown below. Skilled factory workers perform these two tasks reliably and fluidly, without significant mental load. However, in another example of Moravec’s paradox https://en.wikipedia.org/wiki/Moravec%27s paradox , the dexterity required for such tasks can be exceptionally difficult to automate. The first key challenge is uncertainty . Part geometry, appearance, and dynamics aren’t precisely known, and initial part poses can vary substantially. In addition, the electrical connectors are attached to cables that deform along their entire length, have significant manufacturing variation, and even change mechanically through plastic deformation and fatigue. In many manufacturing operations, uncertainty is reduced via fixed automation , where the environment is carefully controlled. Fixtures and adapters are designed to constrain parts to submillimeter accuracy, and robots can operate as repetitive motion machines. In GB300 tester-tray assembly, uncertainty must be addressed via flexible automation , as rapid design cycles and low volumes make fixed automation infeasible. Robots must be inherently adaptive and robust to uncertainty. Given the growing demand for AI infrastructure, this paradigm is a strategic priority for NVIDIA. The second key challenge is performance . In robotics research, 50%-60% success rates are often sufficient, and 80%-90% success rates can be enough to call a task solved . Cycle time, motion smoothness, and gentle handling are often secondary. In addition, statistical rigor https://developer.nvidia.com/blog/how-to-evaluate-general-purpose-robot-policies-for-real-world-deployment/ is often insufficient; 10 to 20 trials are often executed per task, whereas reasonable confidence intervals may require far more evaluations. Our performance requirements were determined in direct collaboration with the NVIDIA Operations Team and Foxconn to ensure that solved really means solved. As of the publication date of this article, we are rapidly approaching their thresholds. Technical strategy: The right tool at the right time In robotics research, a premium is placed on technical novelty; this tradition promotes invention but creates little incentive to optimize baselines that may achieve higher performance. As a result, it can be easy to walk away from a research paper without understanding why prior approaches didn’t work. Our strategy for GB300 tester-tray assembly was the right tool at the right time . For each subproblem, we considered and usually implemented classical baselines first. When they fell short, we explored imitation learning, simulation-based reinforcement learning RL with sim-to-real transfer, real-world RL, and vision-language-action VLA models. This progression allowed us to identify the strengths and weaknesses of each approach and justify innovation and complexity. We now discuss busbar assembly and multi-connector insertion in detail; we aim for these tasks to serve as case studies for other research and engineering groups solving their own real-world manipulation problems. Assembling a busbar: Revisiting the classics Busbar assembly consists of the following steps: 1. Grasp and insert a limit fixture i.e., a particular fixture used to assist with assembly into the tray 2. Grasp a busbar with an attached clamp and insert it into the tray 3. Drive 16 screws to fasten the busbar to the tray 4. Release and remove the clamp attached to the busbar 5. Remove the limit fixture Our initial conditions were consistent with the manufacturing setting. The limit fixture and the busbar with attached clamp were initially placed at unconstrained locations on a flat surface next to the robot, and the screws initially rested within the busbar. The manufacturers requested us to achieve a 99.5% success rate and to complete the task in no more than twice the time of skilled workers in this case, no more than 124 seconds . No unintended collisions were allowed with the tray, as even minor damage could be grounds for discarding the entire system. We began our work by implementing a classical pipeline “ good old-fashioned engineering https://www.science.org/doi/10.1126/scirobotics.aea7390 ” consisting of separate perception, planning, and control modules, expecting a pivot to end-to-end learning afterward. As it turns out, the classical pipeline was highly effective, and no pivot was needed. The modular structure of our solution also made the system straightforward to debug and tune, as well as distribute across robot arms. For perception, we initially used NVIDIA FoundationPose https://nvlabs.github.io/FoundationPose/ , a generalist model that estimates an object’s 6D pose from its mesh, an RGB image, and a depth image. However, when an object wasn’t sufficiently visible, estimates could be off by 90 or 180 degrees. We thus extended the pipeline to ingest multiple views, produce one estimate per view, and select the estimate with the higher confidence score, which mitigated errors. Note: We later transitioned to a specialist model called DOPER, described later. For planning, we used a fast waypoint planner in free space to reduce our cycle times, as well as manipulation primitives https://www.nature.com/articles/s42256-025-01045-3 based on Lissajous curves https://en.wikipedia.org/wiki/Lissajous curve and other simple paths to align parts that may be initially misaligned. For control, we implemented a high-performance impedance controller to ensure gentle and stable contact while applying sufficient force during insertion. Whereas modern learning pipelines often implement the simplest-possible controllers, our solution leveraged automatic damping design and inertial compensation. We view high-performance control as a powerful, but underutilized tool in robot learning. To reduce cycle time, we divided work among three robot arms. One Flexiv Rizon 4S positioned a camera for pose estimation; a second Flexiv Rizon 4S inserted the limit fixture and busbar and later removed the busbar clamp and fixture; and a Universal Robots UR10e used an OnRobot Screwdriver for 16 screwdriving operations. In agreement with the manufacturer, the busbar clamp was lightly modified so that a single arm could unlatch it. Our busbar assembly solution achieved success rates above 95%, which we are actively pushing higher. Remaining failures come primarily from grasping errors and in-hand slippage during insertion. The limit fixture and busbar subtasks met their cycle-time requirements; however, screwdriving didn’t, causing a full cycle time of 160 seconds compared to the 124-second target. Sequential, device-specific screwdriver operations are the main bottleneck, leaving a straightforward path for optimization. Interlude 1: The unreasonable effectiveness of infrastructure On challenging problems like busbar assembly , well-designed real-world robotics infrastructure has enabled us to rapidly develop, deploy, and inspect our systems. The first key piece of infrastructure is robotics services . As roboticists know, experimenting with multiple libraries often devolves into dependency hell https://en.wikipedia.org/wiki/Dependency hell . We convert libraries into independently-deployable web services that run on edge, on-prem, or in the cloud. The services are packaged as Docker containers with FastAPI/HTTP endpoints; researchers can use a lightweight Python client to compose services from different languages and frameworks. We use a cuRobo https://nvlabs.github.io/curobo service for motion planning, a FoundationPose https://nvlabs.github.io/FoundationPose/ service for pose estimation, and a FoundationStereo https://nvlabs.github.io/FoundationStereo/ service for depth estimation, among others. The second key piece of infrastructure is the Task and Agent Lifecycle Orchestration System TALOS . TALOS uses ros2 control https://control.ros.org/rolling/index.html as a backbone and provides motion planners, manipulation primitives, and high-performance controllers impedance, admittance, and force . The library can seamlessly sequence these components and learned policies, as well as execute controllers in real-time loops faster than 500 Hz. TALOS supports the Franka FR3, Flexiv Rizon 4S, Universal Robots UR10e, and SO-101. We use the library for all our real-world robotics applications, including kinesthetic teaching https://pakt-website.github.io/pakt-website/ , teleoperation, bimanual coordination, and motion planning. TALOS is also exposed through standard web interfaces. Coding agents can thus access robotics services and physical robots, allowing the agents to compose perception, planning, and control capabilities during both development and deployment. Our thinking is aligned with Anthropic’s Model Hardware Standard MHS https://www.anthropic.com/news/model-hardware-standard-research-preview , a model-agnostic specification for making physical devices discoverable and safely operable by AI agents, and our open-source DROID+ project https://github.com/NVlabs/DroidPlus exemplifies this approach. Empowered by our robotics infrastructure and our progress on busbar assembly , we proceeded to our second critical task: multi-connector insertion . Multi-connector insertion: ‘Inverse bitter lesson’ Multi-connector insertion can be decomposed into the following steps. For each electrical connector 2 large and 2 small : - Grasp the attached cable to lift the electrical connector from the board - Grasp the connector at the end of the cable - Insert the connector into its corresponding socket on the board Our initial conditions were again consistent with the manufacturing setting. All four cables were connected to the GB300 board at one end, the cables weren’t heavily tangled, and the cable-mounted connectors could be placed arbitrarily on the board. The manufacturers again requested us to achieve a success rate of at least 99.5% and to complete the multi-connector insertion task in no more than twice the time it took skilled human workers in this case, no more than 72 seconds for all four cables . As before, no unintended collisions were allowed with the tray. Although describing the task is straightforward, performing multi-connector insertion is exceptionally challenging due to three factors: 1 the cables are irregular, plastically deformable, and susceptible to wear-and-tear, 2 the connectors are small and nearly textureless, and 3 the connectors and their sockets have tight clearances with little margin for error. We began by again implementing a classical pipeline. To grasp each cable, we used the Segment Anything Model SAM3 https://ai.meta.com/research/sam3/ to segment cables within a predefined region of an overhead RGB image. We fit a centerline to each mask, commanded the robot to grasp its midpoint at a designated height, and commanded the robot upward to expose the connector. This simple approach produced almost no observed failures. Grasping the exposed connectors was far more challenging. Pose estimation with a generalist model was ineffective; the connectors’ small size, textureless appearance, and variable attachment regions to the cables caused large errors. Consequently, motion and grasp planning were unreliable. We thus pivoted to learning perception, motion, and grasping end-to-end. This path was also a dead end. Behavior cloning with action-chunking transformers ACT or diffusion policies achieved low success rates; these approaches benefit from high-quality data at scale, which we struggled to acquire. Cable shape and stiffness varied significantly across samples, and only a small number of these specialized cables were available. Each cable could be grasped and manipulated only a limited number of times before permanently deforming, which no longer represents deployment conditions. We view this circumstance as an “inverse bitter lesson”: in certain real-world situations, scale can be intractable due to physical, financial, and logistical phenomena. Classical approaches that utilize domain knowledge and require less interaction data become worth investigating again. We returned to pose estimation, but with a different set of requirements: specialize to the parts of interest, avoid relying on accurate or distinctive visual textures, and improve rapidly with real-world data. With colleagues in NVIDIA Research, we developed a perception framework called Deep Object Pose Estimation Revisited DOPER . DOPER was inspired by the NVIDIA Deep Object Pose Estimation DOPE https://arxiv.org/pdf/1809.10790 , one of the first papers to investigate using large-scale synthetic data to enable real-world perception. In simulation, a DOPER model is first trained by ingesting a CAD model of a specific part, rendering RGB images from various views, and automatically generating keypoint labels via farthest-point sampling on the part’s convex hull, filtered by the visibility and saliency of the points. A fast, lightweight convolutional network is then trained to predict bounding boxes, segmentation masks, and the keypoint labels from RGB images. In the real world, a 3D neural reconstruction https://docs.nvidia.com/nurec/ of the part is captured. The pretrained DOPER model is executed on different views of the neural reconstruction with randomized backgrounds, generating pseudolabels; low-confidence predictions are corrected by projecting high-confidence predictions into the view of interest. The DOPER model is then fine-tuned on these labels. Finally, at inference, the DOPER model is executed in the scene of interest, and a pose estimate is extracted from the predicted keypoints. This process can be executed at the camera rate of 30 Hz or faster. Figure 7. Left An animation rendered from a neural reconstruction of the DC-SCI connector, which is used to fine-tune a DOPER model. Right A DOPER model at inference time DOPER reinvigorated our belief in the potential of pose estimation. The framework allowed us to train accurate pose estimators specialized to our connectors and deploy these models in real-time. Using the resulting estimates, one robot arm could translate and rotate each connector into a pose from which a second arm could reliably grasp it. Interlude 2: Mechanical design as uncertainty reduction As described earlier, due to rapid design cycles and low volumes, GB300 tester-tray assembly requires flexible automation, where robots are adaptive and robust to uncertainty in parts and poses. At the same time, flexible systems can still benefit from innovations within fixed automation, such as using mechanical design to reduce uncertainty. Unfortunately, mechatronics researchers who focus on fixed automation and computer scientists who focus on flexible automation have long operated in different spheres. We instead advocate for a full-stack approach, with a holistic consideration for hardware, software, and AI research. During connector grasping, contact during gripper closure could perturb the connectors into unintended or unstable poses. Although tactile sensing https://ieeexplore.ieee.org/abstract/document/10912733 , regrasping, and/or in-hand manipulation can be effective solutions, the most straightforward https://en.wikipedia.org/wiki/KISS principle and cycle-time-optimal solution in our case was simply better mechanical design. We designed two sets of multi-purpose gripper fingers that enabled our robots to grasp the busbar, limit fixture, cables, and large and small connectors. The finger geometry constrained part motion during contact, making grasp poses far more repeatable. These fingers can be reproduced on almost any hobby-level 3D printer. The final stage: A contact sport After grasping each connector, the final step of multi-connector insertion was inserting the connector into its corresponding socket on the GB300 board. Compared to inserting the busbar, inserting the connectors was more challenging due to delicate conductive pins, tight clearances, and minimal guiding surfaces during contact. Under such conditions, classical approaches are known to struggle. Variability in socket geometry, accumulated errors in robot control, or improper insertion-force profiles can cause the connector to wedge within the socket, miss the socket, or even collide with surrounding components. Compliant controllers are helpful, but lack robustness to large spatial errors and require careful tuning for each task. We thus turned towards learning-based solutions. Building on our line of sim-to-real research 1 https://arxiv.org/abs/2305.17110 , 2 https://arxiv.org/abs/2407.08028 , 3 https://arxiv.org/abs/2408.04587 , 4 https://arxiv.org/abs/2408.06506 , 5 https://arxiv.org/abs/2503.04538 , we used model-free RL to train policies in the NVIDIA Isaac Lab https://developer.nvidia.com/isaac/lab that could insert the electrical connectors from a wide range of initial poses, and we deployed these policies in the real world. Videos 9-12. Parallelized RL-based training of connector-insertion policies in Isaac Lab . Left Partially-trained policies for the DC-SCI and MCIO connectors. Right Fully-trained policies Sim-to-real transfer achieved high success rates for the large connector in the real world. However, the approach was less effective for the small connector. The policies consumed pose estimates, and even minute estimation errors could cause the connector to slip against the socket edge. Even for the large connector, reliability still didn’t meet industrial standards. We thus reached our final technical conclusion: there is no substitute for real-world data. More precisely, most real-world deployments aim to achieve optimal performance, such as maximum reliability, maximum speed, or minimum intervention rate. As mismatch will always exist between simulation and the real world, policies trained in simulation may easily be suboptimal. Real-world deployments can thus benefit from learning directly from real-world data, ideally from online interaction whenever possible. In our case, we improved our insertion policies by extending our prior work SPARR https://arxiv.org/pdf/2602.23253 . In SPARR, a state-based base policy is first trained via RL in simulation, and an image-based residual policy is then trained via RL in the real world. Rollouts from the base policy bootstrap real-world learning to encourage efficient learning and safe exploration. Because the connectors were visually occluded during insertion, we trained the residual policy with force-torque inputs rather than images. Our final multi-connector insertion solution integrated our approaches to cable grasping, connector grasping, and connector insertion to assemble all four connectors from start to finish. The system achieved success rates of 90-95%; remaining failures come primarily from multi-arm coordination issues during the connector grasping process. Cycle times are an average of 40 seconds per cable; the primary bottleneck is slow, discrete adjustment of the poses of exposed connectors prior to grasping them. Although performance falls short of deployment requirements, ongoing efforts are continuing to close the gap. Why it matters: Towards a virtuous cycle Automating the production and testing of GB300 systems and their successors https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/ is invaluable to NVIDIA and the AI ecosystem, as these systems are the computational engine of foundation models. Farther out, complementing skilled labor with intelligent automation is pivotal to domestic manufacturing: 1.9 million manufacturing jobs may be unfilled https://www.deloitte.com/global/en/alliances/workday/perspectives/deloitte-manufacturing-trends-from-a-workday-lens.html by 2033. In the long run, robotic assembly of AI hardware can fulfill an even grander vision: a future where, under human supervision, AI-powered robots assemble AI hardware, which then train the next generation of AI-powered robots in a safe and virtuous cycle. Blackwell’s fables: Morals of the story Tackling GB300 tester-tray assembly was one of the most informative research exercises we have undertaken, reinforcing the value of task-driven research http://joschu.net/blog/opinionated-guide-ml-research.html . In research, inventing technical approaches in abstraction is a common and valuable practice, as it can pay dividends years later as applications become clear. Our experience with GB300 assembly makes the case for a complementary approach: identify a near-impossible problem of substantial real-world value, expose the capabilities and limits of existing methods, and innovate where there is no clear path forward. Our work has also taught us seven key technical lessons: 1. Pose estimation can be highly effective. At the beginning of the project, we neither believed this conclusion nor even wanted it to be true. However, our DOPER framework produced fast and accurate pose-estimation models for industrial parts. The key was revisiting pose estimation with modified requirements: training specialist models, avoiding reliance on accurate or distinctive visual textures, and enabling rapid fine-tuning on real-world data. 2. Mechanical design can efficiently reduce uncertainty. This conclusion is obvious to the automation community, but may be less apparent in robot learning. Mechanical design of 3D-printed gripper fingers allowed us to reliably grasp clamps, cables, and fixtures without introducing complexity through tactile sensors, regrasp mechanisms, or in-hand manipulation. Robotics is a full-stack problem, and the best results can be achieved with both artificial and mechanical intelligence. 3. High-performance control shouldn’t be overlooked. Researchers have worked for decades to develop high-performance robot controllers, but most robot learning papers implement the simplest-possible control laws. High-performance controllers with automatic damping design and inertial compensation can ensure smooth free-space motion, facilitate gentle and stable contact interactions, and improve the success rates and cycle times of behaviors learned via RL. 4. Scale-bottlenecked regimes should be examined further. The Bitter Lesson http://www.incompleteideas.net/IncIdeas/BitterLesson.html is supported by remarkable empirical evidence; nevertheless, in what we call the “inverse bitter lesson,” scale can be intractable when interaction data is scarce, parts are difficult to simulate, replacements are expensive, or repeated use degrades the objects. Classical approaches can still be powerful, and learning-based research may benefit from further emphasizing structure and sample efficiency. 5. Cycle time is a worthy counterpart to success rate. A typical robot-learning system may combine 1 kHz proprioception, a 100 Hz policy, 30 Hz camera streams, and 5 Hz gripper control, with differing real-time guarantees. Certain subtasks might be difficult or impossible to parallelize. Increased speeds must preserve control accuracy, ensure human and robot safety, and avoid part damage. Researchers should view cycle time and success rate with equal priority. 6. Research infrastructure accelerates development and deployment. In our work, robotics services and TALOS were critical for avoiding dependency conflicts, making different planners and controllers interchangeable, and transferring code across robots. Whereas startup companies and corporations may invest deeply in infrastructure, research labs typically don’t; community efforts to develop research-tailored infrastructure may accelerate progress in the field. 7. Both synthetic data and real-world data play a critical role. DOPER used large-scale, part-specific synthetic data to initiate the training process, and our connector-insertion approach leveraged simulation-based RL to pretrain policies. Nevertheless, real-world data still proved essential. The strict demands of industrial robotics can’t afford to leave performance on the table, and synthetic and real-world data must be combined to deliver real-world value. Closing the gap between lab and factory In robotics, the true measure of research success is translation to the real world. Our next goal is to deploy our technology into the factories where NVIDIA contract manufacturers produce AI hardware; we aim to meet or exceed their performance requirements, and in doing so, accelerate the production of the world’s most advanced compute. Nevertheless, real-world deployment comes with notable challenges. Foremost is the following: even if a robotics solution performs well in laboratory trials, how does one guarantee its performance before deployment? The factory will have distinct ambient and workcell conditions, and handling domain mismatch is a near-universal weakness of learning-based robots. Real production lines are also subject to product design changes, another degree of freedom that can’t be observed in the lab. In addition, even if the system is robust or adaptive enough, its performance must be verified. To support a success probability of at least 99.5% with a one-sided 95% confidence bound, zero failures must be observed over at least 598 trials https://developer.nvidia.com/blog/how-to-evaluate-general-purpose-robot-policies-for-real-world-deployment/ . In our case, due to wear-and-tear, such evaluations may require hundreds of duplicate GB300 components. Accurately simulating geometrically and mechanically irregular components is also an open challenge, despite recent progress in physics simulation https://developer.nvidia.com/blog/newton-adds-contact-rich-manipulation-and-locomotion-capabilities-for-industrial-robotics/ and interactive world models https://www.yixuanwang.me/interactive world sim/ . We propose drawing inspiration from the autonomous vehicles AV industry. In AV, the boundaries between offline training and online deployment are blurred: initially, humans carefully supervise and correct vehicle behavior, but as time passes, the car becomes more and more autonomous. Similarly, a robot learning system should be deployed with strong, but imperfect performance. Humans or human-supervised agents should seamlessly record interventions, and the system should leverage them to continuously improve. As time passes, enough data would be collected in real-world deployment conditions to provide strong statistical guarantees about performance. The second key challenge of real-world deployment is, after deploying a robotics solution to solve a particular task, how can the next task take less effort to solve? This goal can be summarized by two words: generalization and adaptivity . A strong form of generalization is when a robot trained on a set of tasks can solve a new one well. If a robot knows how to insert dowel pins into holes, assemble a set of gears, and tighten screws, it should have the high-level common sense and low-level dexterity to plug in a USB-C cable. A strong form of adaptivity is when a robot doesn’t know how to solve a new task, but can learn it efficiently. A robot that can assemble gears with 0.5 mm clearances should rapidly learn to assemble gears with 0.1 mm clearances. In line with research consensus, we believe that the long-term path towards generalization and adaptivity is through models that learn from internet-scale data, including vision-language-action models VLAs https://mbreuss.github.io/blog post iclr 26 vla.html , world-action models WAMs https://developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action-models/ , and robotics-native foundation models. These models have shown remarkable signs of physical common-sense, fast adaptation, and in-context learning. However, there is an open secret in the robotics community: we are still a long way from capturing the near-infinite diversity of the physical world and learning from it efficiently. As a thought experiment, an AV company may collect 100 TB to 1 PB of useful data per day, but autonomous driving is just one task that a generalist robot should know how to perform. Furthermore, modular systems in reasonably-constrained environments often generalize better and faster than end-to-end models e.g., 1 https://arxiv.org/abs/2603.09971 , 2 https://arxiv.org/abs/2511.04758 . We see three complementary paths to generalization and adaptivity in robotics foundation models. The first path is the data flywheel : build continuously-improving robotic systems in deployment conditions, collect data during deployment, and use the data to mid-train or post-train foundation models. The key limitation is coverage: each deployment represents a narrow interaction distribution, and foundation models may struggle to interpolate. The second path is world models . Defining them as an “action-conditioned multimodal prediction model,” these models can be trained on almost any physical data source. Researchers may eventually use them as a simulator for all physical interaction, enabling diverse post-training of robotics foundation models. However, future world models must evolve over time; as discussed earlier, object dynamics can be highly non-stationary. The third path is agentically-driven robotics . As suggested in recent breakthroughs e.g., 1 https://arxiv.org/abs/2607.05369 , 2 https://arxiv.org/abs/2607.00272 , 3 https://arxiv.org/abs/2606.19980 , agents not only exhibit increasing amounts of human common sense, including spatial reasoning, but can also propose effective actions in simulators, world models, and the real world. Agent-guided actions might generate post-training data, and agent-guided interventions might eventually substitute for human corrections on novel tasks. We conclude with a call to action: come join us on our journey. Robotic assembly of AI hardware not only represents one of the most impactful problems of the modern era, but is also a compelling North Star for fundamental and applied robotics research. We are excited to see the community make rapid advances in this direction. Dive deeper We are preparing a whitepaper that describes our methods and latest experimental results in significant detail. We are also working toward releasing DOPER and TALOS for community use, along with reference workflows for connector insertion built in collaboration with the NVIDIA Isaac engineering team. We will update this post with links as each release goes live. Check out our NVIDIA Robotics Research and Development Digest R