cd /news/robotics/facet-0-a-robotic-foundation-model-f… · home topics robotics article
[ARTICLE · art-125086] src=arxiv.org ↗ pub= topic=robotics verified=true sentiment=↑ positive

Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation

Researchers introduced Facet-0, a robotic foundation model for contact-rich precise manipulation, which achieved 82% mean success on five sub-millimeter computer-assembly tasks compared with 15% for the strongest baseline, with 0.5 mm placement accuracy and 50 ms command latency. The model, trained on the 1,000-hour ManuFacet-1K corpus spanning three embodiments, unifies multimodal representation learning and reinforcement learning around a joint action-wrench proposal to predict and value contact consequences.

read2 min views2 publishedSep 9, 2026
Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation
Image: source
  [Submitted on 1 Sep 2026]


[View PDF](/pdf/2609.01596)

[HTML (experimental)](https://arxiv.org/html/2609.01596v1)

Abstract:Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. We present Facet-0, a robotic foundation model that predicts and values the contact consequences of its actions. Facet-0 unifies multimodal representation learning and reinforcement learning (RL) post-training around a joint action-wrench proposal: a causal wrench history is aligned with vision-language semantics and kinematic state, and flow matching generates each action chunk together with the future wrist-wrench profile it is expected to induce. Deployment rollouts train a distributional Action-Wrench Critic to distinguish motions with similar task progress but different contact outcomes, while phase-aware rewards and contact-selective credit concentrate policy improvement on decisive interactions. To accommodate part-specific dynamics, a lightweight bounded actor reuses the frozen representation for on-robot adaptation; RL remains defined over executable Cartesian actions, while an auxiliary wrench head preserves predictive, non-commanded action-contact coupling. Trained on ManuFacet-1K, a 1,000-hour force-synchronized corpus spanning three embodiments and multiple manufacturing cells, the bounded task-adapted system reaches 82% mean success on five sub-millimeter computer-assembly tasks, compared with 15% for the strongest baseline, with 0.5 mm placement accuracy and 50 ms command latency.

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?) alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?) Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?) arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #robotics 4 stories · sorted by recency
── more on @facet-0 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/facet-0-a-robotic-fo…] indexed:0 read:2min 2026-09-09 ·