Figure AI's Helix 2.5 Robot Made Beds in 30 Homes It Had Never Seen Figure AI unveiled Helix 2.5 on September 17, 2026, a humanoid neural network that achieved a 56% success rate across 237 completed tasks out of 420 attempts in 30 Bay Area homes its robots had never entered, using a single fixed checkpoint shared across all homes. An otherwise identical model trained from scratch on the same tasks managed just 9%, and Figure says Helix 2.5 needed 50% less task-specific data than its predecessor Helix 02. CEO Brett Adcock called zero-shot generalization "the most important project we've ever taken on at Figure," while task-level results ranged from 67% for bed making (94 of 140) to 40% for tidying toys (56 of 140). Figure AI says a robot that had never set foot in a house can walk in and start cleaning, using a single frozen brain shared across all 30 homes it was tested in. On September 17, 2026, Figure AI unveiled Helix 2.5, a humanoid neural network the company sent into 30 homes across the Bay Area that its robots had never entered and where no training data had ever been collected. The result: a 56% success rate across 237 completed tasks out of 420 attempts, covering bed making, towel folding, and tidying toys in living rooms. An otherwise identical model trained from scratch on the same tasks managed just 9%. That gap is the whole story. Figure didn't rebuild the robot's hardware or hand-tune it for each house. It used one fixed checkpoint, the same weights, in every single home, and let the robot work out unfamiliar furniture, unfamiliar toys, and unfamiliar towels on its own. None of the evaluation objects, not the toys, the bedding, or the towels, ever appeared in the data the model trained on. Break the numbers down by task and the picture gets more honest. Bed making hit 67%, with 94 of 140 attempts succeeding, the strongest of the three. Towel folding came in at 62%, 87 of 140. Tidying toys back into a basket lagged well behind at just 40%, 56 of 140. Full completion was required for credit. A robot that made most of a bed but left a pillow crooked got nothing, and Figure counted every safety intervention as a failure rather than a gray area. "The most important project we've ever taken on at Figure." That's how CEO Brett Adcock described Helix 2.5. Zero-shot generalization is his phrase for it: a humanoid walking into a house it's never seen and getting straight to work. He calls it the holy grail of the industry. "We rented 30 homes in the Bay Area and are doing tasks without any new training," Adcock wrote on X ahead of the official release. Figure's robot livestream raises the bar for humanoid proof https://startupfortune.com/figures-robot-livestream-raises-the-bar-for-humanoid-proof/ Figure AI says its Figure 03 humanoids sorted packages autonomously with no teleoperation during a livestream that extended well beyond the original eight-hour target. The test is a meaningful endurance signal, but buyers still need independent proof of uptime, maintenance, supervision and cost. - humanoid robot livestream autonomous warehouse test https://startupfortune.com/figures-robot-livestream-raises-the-bar-for-humanoid-proof/ - proving AI robots can work independently in factories https://startupfortune.com/figures-robot-livestream-raises-the-bar-for-humanoid-proof/ Figure's bet is simple: pretraining on huge amounts of general human behavior beats narrow task-specific practice. That's what lets a robot cope with a house it's never seen. The company built that pretraining data through Index, a separate product it took out of stealth on August 25, 2026. Figure calls it the most diverse robot training dataset ever assembled. Index pulls in roughly 35 minutes of human experience footage every second, uploaded by tens of thousands of weekly contributors through an app. The payoff: Figure says Helix 2.5 needed 50% less task-specific data than its predecessor, Helix 02, to hit stronger results across three times as many homes. That's a real, checkable claim in a field that mostly runs on choreographed demo videos. No cherry-picked reel. Figure logged safety interventions as failures and imposed timeouts on every trial, rather than letting the robot try indefinitely until it happened to succeed. The rest of the field isn't standing still Figure isn't the only company chasing this number. Sunday Robotics recently reported a 99.1% success rate, 778 of 785 attempts, for its ACT-2 system folding laundry in unseen environments, according to a report from Humanoids Daily. The two results aren't directly comparable. Different tasks, different grading rules, different definitions of "unseen." But the contrast shows how crowded and how noisy the humanoid benchmark race has become, and why Figure leaned so hard on publishing raw attempt counts instead of a single flattering percentage. The timing isn't an accident either. Figure closed more than $1 billion in a Series C round and struck a partnership with Brookfield earlier this month, giving it, by Adcock's own account, the strongest balance sheet in humanoid robotics. Money like that buys runway. Adcock has said Figure moved its home robot timeline up by two years, and Helix 2.5 is the evidence the company wants on hand when investors and the public ask whether that acceleration is real or just marketing. Whether 56% is impressive depends entirely on what you measure it against. Against a robot trained from nothing, it's a six-fold jump. Against the bar for a robot you'd actually trust alone in your living room, it's still a machine that fails at picking up toys more often than it succeeds. Also read: Researchers Used Anthropic's Claude to Hack Into OpenAI's Own Systems https://startupfortune.com/researchers-used-anthropics-claude-to-hack-into-openais-own-systems/ • The FAA Bet $875 Million on an AI Startup to Fix Air Traffic Control https://startupfortune.com/the-faa-bet-875-million-on-an-ai-startup-to-fix-air-traffic-control/ • Bank of Japan Raises Interest Rate to 1.25%, a 31-Year High https://startupfortune.com/bank-of-japan-raises-interest-rate-to-125-a-31-year-high/ Figure AI will test humanoid autonomy in an eight hour livestream https://startupfortune.com/figure-ai-will-test-humanoid-autonomy-in-an-eight-hour-livestream/ Figure AI CEO Brett Adcock promised an eight hour autonomous humanoid livestream after a public challenge from robotics veteran Scott Walter. The test could help show whether humanoids are moving from polished demos toward real labor economics. - Figure AI humanoid robot eight hour livestream test https://startupfortune.com/figure-ai-will-test-humanoid-autonomy-in-an-eight-hour-livestream/ - can humanoid robots work autonomously without human intervention https://startupfortune.com/figure-ai-will-test-humanoid-autonomy-in-an-eight-hour-livestream/ This article is posted in AI News https://startupfortune.com/category/ai/ , check it out for more related stories. Join the discussion Open in the community → https://startupfortune.com/community/ Almost there. Sign in and your reply posts straight away.