{"slug": "agenthands-generating-interactive-hand-gestures-for-spatially-grounded-agent-in", "title": "AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR", "summary": "Google XR researchers Xun Qian and Ruofei Du introduced AgentHands, an LLM-powered extended reality (XR) prototype that generates synchronized, expressive hand gestures for spatially grounded agent conversations, published at CHI 2026. The system uses a multi-dimensional taxonomy and a gesture library to map LLM reasoning into real-time physical motions, aiming to enhance user engagement in physical tasks by bridging the mental mapping gap in immersive platforms like Android XR.", "body_md": "August 25, 2026\n\nXun Qian, Research Scientist, and Ruofei Du, Interactive Perception & Graphics Lead, Google XR\n\nAgentHands is an LLM-powered XR prototype that augments conversational agents with synchronized, expressive hand gestures to provide spatially grounded guidance, bridging the mental mapping gap and enhancing user engagement in physical tasks.\n\nAs AI assistants evolve from simple text interfaces to multimodal companions, we are seeing a shift toward more proactive, situated assistance. Recent innovations like [Project Astra](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-universal-ai-assistant/#live-capabilities) and [Gemini 3.1 Flash Live](https://blog.google/innovation-and-ai/technology/developers-tools/build-with-gemini-3-1-flash-live/) already allow users to discuss their physical surroundings in real time, often utilizing visual bounding box overlays to identify objects in a camera feed. While these overlays are highly effective for 2D screens, the transition to immersive platforms like [Android XR](https://www.android.com/xr/) presents a unique challenge: how do we move beyond flat UI to create a truly embodied, spatially aware dialogue?\n\nTo bridge this gap, we introduce [AgentHands](https://research.google/pubs/agenthands-generating-interactive-hands-gestures-for-spatially-grounded-agent-conversations-in-xr/), published at [CHI 2026](https://chi2026.acm.org/), a research prototype that brings the power of co-speech gestures to the 3D world. In human communication, our hands do more than just point; they describe shapes, mimic actions, and emphasize points, all synchronized with our voice. By leveraging the spatial understanding capabilities of [Extended Reality](https://www.android.com/xr/) (XR), AgentHands replicates this natural synergy. Following up our prior research in [Human I/O](https://research.google/blog/human-io-detecting-situational-impairments-with-large-language-models/) and [Sensible Agent](https://research.google/blog/sensible-agent-a-framework-for-unobtrusive-interaction-with-proactive-ar-agents/), AgentHands further equips AI agents with expressive, synchronized hand gestures that transform abstract verbal instructions into intuitive, physical demonstrations, making conversations about your surroundings more natural and engaging.\n\nTo start, we conducted a formative study with XR and [human–computer interaction](https://en.wikipedia.org/wiki/Human%E2%80%93computer_interaction) (HCI) experts at Google to determine what makes a virtual hand “legible” in a 3D environment. We distilled these insights into a multi-dimensional taxonomy that defines how an agent should use its hands to ground a conversation within a user's physical space.\n\nThe core innovation of AgentHands is its ability to map the high-level reasoning of LLMs into precise, real-time physical motions that match the agent's “voice” and the user's XR environment. We introduce the following key steps to compose the AgentHands workflow.\n\nThe system begins with a lightweight object registration module. Using eye gaze and scene reconstruction, users can quickly “tag” items — like an orchid or a laptop — creating a spatial registry with 3D bounding boxes that the agent can reference.\n\nWe constructed a library of hand gesture behaviors across three semantic categories: a) *deictic* for referencing, b) *iconic* for depicting actions or forms, and c) *expression* for conveying social cues and emotion.\n\nWhen a user asks a question, the backend LLM generates a response that includes inline GestureEvents. Each event is attached to specific trigger words and encodes the primitives for a hand behavior following the taxonomy dimensions.\n\nA local parser on the XR headset coordinates the [text-to-speech](https://cloud.google.com/text-to-speech) (TTS) playback with the animation engine. By using word-level timestamps, the agent’s hands perform co-speech gestures in perfect sync with the spoken words, providing clear, expressive spatial references.\n\nBy integrating these modules, AgentHands creates a seamless bridge between linguistic intent and physical action. The system transforms a standard LLM output into a rich, multimodal performance where the agent's generated responses are manifested through both speech and spatially accurate movement, allowing for complex instructions to be demonstrated exactly where they occur in the user's environment.\n\nWe demonstrated how these embodied gestures, paired with the spatial awareness of XR, enhance our understanding of our physical surroundings.\n\n*Interactive tutoring:* In an orchid-care scenario, the agent doesn’t just say “check the roots”; it moves its hands to the base of the plant and outlines the air roots while explaining their function.\n\n*Technical walkthroughs:* For 3D printer operations, the agent can demonstrate the exact ''turn and click'' sequence needed to navigate control knobs and select files, making complex physical interface steps intuitive.\n\n*Lifestyle companionship:* The agent can serve as a wellness coach that interacts with your physical choices. For instance, the agent can perform an interactive “warning” gesture by holding the user’s hand and a visual effect to caution the user against unhealthy behavior.\n\nTo evaluate the impact of these gestures, we conducted a within-subjects study (N = 12) comparing AgentHands to a speech-only baseline. Both conditions used the same researcher-scripted verbal content, ensuring the only difference was the presence of the embodied hands and their synchronized gestures. Participants completed two procedural tasks that balanced everyday care with technical operation.\n\nThe results confirmed that the combination of XR and co-speech gestures is highly effective for spatially grounded interactions. We analyzed the data across several key metrics of communication effectiveness.\n\nAgentHands represents a step toward a future where AI systems aren’t just analyzing our world, but dynamically operating within it. By leveraging co-speech gestures and the spatial power of XR to ground conversation in physical movement, we can reduce the cognitive load of complex tasks and make spatial computing more accessible and human-centric.\n\nAs we continue to develop for the Android XR ecosystem, we are exploring ways to make these gestures even more personalized, adapting to a user’s dominant hand or learning their specific spatial routines, to create an even more seamless human-AI collaboration.\n\n*This research was primarily conducted by Ziyi Liu during his Student Researcher tenure at Google, as part of a joint collaboration across multiple teams. We extend our sincere gratitude to key contributors David Li, Zhongyi Zhou, and David Kim for their support, and to Adarsh Kowdle, Guru Somadder, and Shahram Izadi for their strategic guidance and thoughtful reviews.*", "url": "https://wpnews.pro/news/agenthands-generating-interactive-hand-gestures-for-spatially-grounded-agent-in", "canonical_source": "https://research.google/blog/agenthands-generating-interactive-hand-gestures-for-spatially-grounded-agent-conversations-in-xr/", "published_at": "2026-08-25 19:10:59+00:00", "updated_at": "2026-08-25 19:15:09.769593+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "generative-ai", "ai-research", "ai-products"], "entities": ["Google XR", "Xun Qian", "Ruofei Du", "AgentHands", "CHI 2026", "Android XR", "Project Astra", "Gemini 3.1 Flash Live"], "alternates": {"html": "https://wpnews.pro/news/agenthands-generating-interactive-hand-gestures-for-spatially-grounded-agent-in", "markdown": "https://wpnews.pro/news/agenthands-generating-interactive-hand-gestures-for-spatially-grounded-agent-in.md", "text": "https://wpnews.pro/news/agenthands-generating-interactive-hand-gestures-for-spatially-grounded-agent-in.txt", "jsonld": "https://wpnews.pro/news/agenthands-generating-interactive-hand-gestures-for-spatially-grounded-agent-in.jsonld"}}