cd /news/robotics/general-purpose-physical-ai-may-need… · home topics robotics article
[ARTICLE · art-139175] src=humanoidanalytics.com ↗ pub= topic=robotics verified=true sentiment=· neutral

General-Purpose Physical AI May Need More Data and Compute Than LLMs

Figure reported that pretraining on its Index dataset raised full-task success in 30 previously unseen homes from 9% to 56%, a 47-percentage-point gain, in a September 17, 2026 evaluation that held architecture and downstream training data fixed. The result strengthens the case that useful physical data, not compute, is the better-evidenced bottleneck in robot learning, though disclosed training runs remain far below LLM scale: NVIDIA reported roughly 50,000 H100 GPU hours for GR00T-N1-2B pretraining versus Meta's 30.84 million H100 GPU hours for Llama 3.1 405B, about 0.16%. Figure said in August 2026 that Index had paid $15 million to contributors after outside vendors failed to meet its requirements, and Generalist announced $400 million in new funding on June 4, 2026.

by read6 min views2 publishedSep 24, 2026
General-Purpose Physical AI May Need More Data and Compute Than LLMs
Image: Humanoidanalytics (auto-discovered)

Useful physical data remains the better-evidenced bottleneck in robot learning. Figure’s Helix 2.5 results strengthen that assessment: the company reports that pretraining on its Index dataset raised full-task success in unfamiliar homes from 9% to 56%. [1]

For investors, the infrastructure thesis is becoming more concrete. Developers are paying to acquire experience and committing to larger computing systems. The evidence supports a market for collecting, processing and learning from physical data. It does not yet establish that general-purpose robotics requires more training compute than frontier language models. The distinction matters. Demand for infrastructure can become commercially significant before compute becomes the main technical constraint.

Data has a measurable effect on generalization #

Figure’s September 17, 2026 evaluation covered three learned household tasks across 30 previously unseen homes. Its comparison held architecture, downstream training data and other training settings fixed while changing Index pretraining. The 47-percentage-point improvement therefore provides unusually direct evidence for the value of physical pretraining in that test. [1]

The commercial implication is lower adaptation effort: learning that transfers between environments could reduce the work required to prepare each new customer location.

Earlier academic work points in the same direction. DROID’s 350-hour dataset improved out-of-distribution success by 17 percentage points over the next-best method in its comparison, which included Open-X co-training and no co-training. [2]

The scale has since changed. In November 2025, Generalist reported that GEN-0 used more than 270,000 hours of manipulation data, with collection growing by 10,000 hours weekly at that time. Its reported model-size results were specific: 1B-parameter models struggled to absorb further information, 6B models benefited from pretraining, and 7B-plus models transferred learning with limited post-training. [3]

That finding suggests the constraint can shift. Once a developer builds a sufficiently productive data operation, model capacity and training resources may become more important. The investment case concerns useful experience and the ability to learn from it, rather than raw recording volume alone.

Disclosed training runs remain far below LLM scale #

NVIDIA reported roughly 50,000 H100 GPU hours for GR00T-N1-2B pretraining. Meta disclosed 30.84 million H100 GPU hours for Llama 3.1 405B. Dividing those figures gives approximately 0.16%. This compares the disclosed runs, with different accounting boundaries: GR00T also inherits a previously trained vision-language backbone. It is not a comparison of complete development costs. [4, 5]

Epoch AI’s August 2025 analysis provides broader historical support. It estimated that the largest manipulation models typically used about 1% of the training compute of frontier models in other domains. Its accounting included upstream model training, while excluding simulation and data generation. Its survey predates GEN-0 and Helix 2.5, so it cannot settle the September 2026 comparison. [6]

Body control adds workloads outside that accounting. Berkeley researchers trained a humanoid locomotion controller across thousands of GPU-simulated environments before transferring it to a real robot, illustrating why simulation demand needs separate measurement. [7]

Follow the buyers and their suppliers #

Figure provides a direct data-spending example. In August 2026, it said Index had paid $15 million to contributors. It also said outside vendors had failed to meet its requirements, prompting an internal collection system. That supports demand for paid collection, while highlighting the risk that robot developers capture the service internally. Investors should track repeat contributor payments and the cost of accepted, useful data. [8]

Generalist offers a financing signal. On June 4, 2026, it announced $400 million in new funding, led by Radical Ventures, to expand models, its physical data operation and computing infrastructure. The announcement does not allocate the proceeds among those uses. Its GEN-0 disclosure discusses data-foundry partners without naming them. Named supplier contracts and recurring purchases would turn this financing signal into a more investable services-market map. [3, 9]

Compute demand now has a named buyer and provider. On September 3, Figure and Nscale announced a signed agreement with an initial $3.5 billion compute commitment and potential deployment of up to 100,000 NVIDIA Vera Rubin GPUs. Initial deployment is targeted for the second half of 2027 in Barstow, Texas. [10]

Nscale also announced a strategic investment in Figure. The relationship therefore combines supplier revenue potential with equity exposure to its customer. The next milestones are commissioned capacity, Figure’s paid utilization and disclosed revenue conversion. A future commitment is not evidence that the capacity is operating today. [10]

These disclosures support distinct opportunities: contributors and collection services supplying physical experience, and infrastructure providers supplying processing, simulation and training. Their economics must be assessed separately.

What the evidence does not show #

The company tests do not establish a general-purpose humanoid or sustained customer economics. Figure’s 56% result covers a defined evaluation, and Generalist’s model-size findings do not establish a universal 7B threshold. Figure itself describes both data and compute as constraints, which argues against treating data scarcity as the sole bottleneck everywhere. [1, 3, 10]

The long-term hypothesis remains that general-purpose physical AI could require more data, and possibly more training compute, than frontier LLMs. NVIDIA’s Cosmos announcement cited 20 million video hours and around 9,000 trillion tokens. Its paper distinguishes that raw collection from curated training clips. Video tokens and text tokens are not directly comparable measures of learning requirements. [11, 12]

Efficiency is a credible alternative to continuously escalating training budgets. NVIDIA tested constrained GR00T fine-tuning on a single A6000 GPU. Shared models and targeted adaptation could let downstream developers expand deployments without replicating a foundation lab’s spending. [4]

The current investment case is clearest in paid physical-data acquisition, with compute an increasingly substantial bet on scaling. One result would change that assessment: a controlled, independently validated study showing that additional training compute on a fixed physical dataset produces a large, repeatable reduction in customer-site interventions. That would provide stronger evidence that compute has become the next constraint.

Sources:

  1. Figure, “Helix 2.5: Zero-Shot 30-Home Generalization” Source type: Tier 3, detailed first-party disclosure. Company-controlled experiments; September 17, 2026.https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization
  2. DROID Dataset Team, “DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset” Source type: Tier 2, strong independent evidence. Academic research; author-reported experiments.https://droid-dataset.github.io/
  3. Generalist, “GEN-0 / Embodied Foundation Models That Scale with Physical Interaction” Source type: Tier 3, detailed first-party disclosure. Company-controlled results and data-operation figures; November 4, 2025.https://generalistai.com/blog/gen-0
  4. NVIDIA, “GR00T N1: An Open Foundation Model for Generalist Humanoid Robots” Source type: Tier 3, detailed first-party disclosure. Company-authored paper, version 1; March 18, 2025.https://arxiv.org/html/2503.14734v1
  5. Meta, “Llama 3.1 Model Card” Source type: Tier 3, detailed first-party disclosure. Company-controlled model documentation and training-resource figures.https://github.com/meta-llama/llama-models/blob/main/models/llama3_1/MODEL_CARD.md
  6. Epoch AI, “Compute is not a bottleneck for robotic manipulation” Source type: Tier 2, strong independent evidence. Independent analysis published August 8, 2025; survey described as current through July 2025.https://epoch.ai/data-insights/compute-for-robotic-manipulation
  7. University of California, Berkeley research team, “Learning Humanoid Locomotion with Transformers” Source type: Tier 2, strong independent evidence. Academic research; author-reported simulation and real-robot experiments.https://humanoid-transformer.github.io/
  8. Figure, “Introducing Index: Building The World’s Largest and Most Diverse Physical Dataset” Source type: Tier 3, detailed first-party disclosure. Company-controlled account of contributor payments and collection strategy; August 25, 2026.https://www.figure.ai/news/introducing-index
  9. Generalist, “Accelerating the Next Phase of Physical AI” Source type: Tier 3, detailed first-party disclosure. Company-controlled funding announcement and intended uses; June 4, 2026.https://generalistai.com/blog/accelerating-the-next-phase-of-physical-ai
  10. Nscale, “Nscale and Figure Sign Strategic Partnership to Power the Next Generation of Physical AI” Source type: Tier 3, detailed first-party disclosure. Company-controlled agreement announcement; September 3, 2026. Nscale is both the prospective provider and an announced investor in Figure.https://www.nscale.tech/press-releases/nscale-and-figure
  11. NVIDIA, “NVIDIA Makes Cosmos World Foundation Models Openly Available to Physical AI Developer Community” Source type: Tier 3, detailed first-party disclosure. Company-controlled article; January 6, 2025.https://blogs.nvidia.com/blog/cosmos-world-foundation-models/
  12. NVIDIA, “Cosmos World Foundation Model Platform for Physical AI”

Source type: Tier 3, detailed first-party disclosure. Company-authored paper, version 1; January 7, 2025.https://arxiv.org/html/2501.03575v1 Related Analysis: Figure’s Helix 2.5 Brings Humanoid Robotics Closer to Its ChatGPT Moment

── more in #robotics 4 stories · sorted by recency
── more on @figure 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/general-purpose-phys…] indexed:0 read:6min 2026-09-24 ·