arXiv:2609.12155v1 Announce Type: new Abstract: Understanding person-level bi-manual interactions requires not only detecting hands, but also identifying which two hands belong to the same person and what each hand interacts with. Existing hand--object interaction methods are mostly hand-centric: they treat each hand as an independent instance, which can lead to ambiguous ownership in multi-person scenes. We propose a person-centric formulation in which a single query predicts a structured output for one person, including the human box, body pose, hand boxes and states, and interaction targets. We introduce part-aware deformable attention to allocate attention across human, hand, and pose-specific reference regions, enabling one query to capture the full person structure. We further unify detection and interaction reasoning with a hand-to-query relationship matrix, where each hand selects its interaction target from the detected query set plus a learnable off token, directly recovering the target's box and class without separate object regression. We build a COCO-based dataset with person-centric bi-manual interaction annotations and define structured metrics for evaluating hand states and complete hand--object tuples. Experiments with a transformer-based detector show that our formulation improves person-level bi-manual interaction parsing and provides an effective unified framework for joint detection, pose estimation, and hand reasoning.
Single-Query Person-Centric Bimanual Hand-Object Interaction Detection
Researchers posting to arXiv (paper 2609.12155v1) proposed a person-centric formulation for bimanual hand-object interaction detection in which a single query predicts a structured output for one person, including the human box, body pose, hand boxes and states, and interaction targets. The method introduces part-aware deformable attention across human, hand, and pose-specific reference regions and a hand-to-query relationship matrix that lets each hand select its interaction target from the detected query set plus a learnable off token. Experiments with a transformer-based detector on a new COCO-based dataset with person-centric bi-manual interaction annotations showed the formulation improves person-level bi-manual interaction parsing and provides a unified framework for joint detection, pose estimation, and hand reasoning.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.