cd /news/artificial-intelligence/agenthands-generating-interactive-ha… · home topics artificial-intelligence article
[ARTICLE · art-110741] src=research.google ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

Google XR researchers Xun Qian and Ruofei Du introduced AgentHands, an LLM-powered extended reality (XR) prototype that generates synchronized, expressive hand gestures for spatially grounded agent conversations, published at CHI 2026. The system uses a multi-dimensional taxonomy and a gesture library to map LLM reasoning into real-time physical motions, aiming to enhance user engagement in physical tasks by bridging the mental mapping gap in immersive platforms like Android XR.

read5 min views5 publishedAug 25, 2026
AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR
Image: Google Research

August 25, 2026

Xun Qian, Research Scientist, and Ruofei Du, Interactive Perception & Graphics Lead, Google XR

AgentHands is an LLM-powered XR prototype that augments conversational agents with synchronized, expressive hand gestures to provide spatially grounded guidance, bridging the mental mapping gap and enhancing user engagement in physical tasks.

As AI assistants evolve from simple text interfaces to multimodal companions, we are seeing a shift toward more proactive, situated assistance. Recent innovations like Project Astra and Gemini 3.1 Flash Live already allow users to discuss their physical surroundings in real time, often utilizing visual bounding box overlays to identify objects in a camera feed. While these overlays are highly effective for 2D screens, the transition to immersive platforms like Android XR presents a unique challenge: how do we move beyond flat UI to create a truly embodied, spatially aware dialogue?

To bridge this gap, we introduce AgentHands, published at CHI 2026, a research prototype that brings the power of co-speech gestures to the 3D world. In human communication, our hands do more than just point; they describe shapes, mimic actions, and emphasize points, all synchronized with our voice. By leveraging the spatial understanding capabilities of Extended Reality (XR), AgentHands replicates this natural synergy. Following up our prior research in Human I/O and Sensible Agent, AgentHands further equips AI agents with expressive, synchronized hand gestures that transform abstract verbal instructions into intuitive, physical demonstrations, making conversations about your surroundings more natural and engaging.

To start, we conducted a formative study with XR and human–computer interaction (HCI) experts at Google to determine what makes a virtual hand “legible” in a 3D environment. We distilled these insights into a multi-dimensional taxonomy that defines how an agent should use its hands to ground a conversation within a user's physical space.

The core innovation of AgentHands is its ability to map the high-level reasoning of LLMs into precise, real-time physical motions that match the agent's “voice” and the user's XR environment. We introduce the following key steps to compose the AgentHands workflow.

The system begins with a lightweight object registration module. Using eye gaze and scene reconstruction, users can quickly “tag” items — like an orchid or a laptop — creating a spatial registry with 3D bounding boxes that the agent can reference.

We constructed a library of hand gesture behaviors across three semantic categories: a) deictic for referencing, b) iconic for depicting actions or forms, and c) expression for conveying social cues and emotion.

When a user asks a question, the backend LLM generates a response that includes inline GestureEvents. Each event is attached to specific trigger words and encodes the primitives for a hand behavior following the taxonomy dimensions.

A local parser on the XR headset coordinates the text-to-speech (TTS) playback with the animation engine. By using word-level timestamps, the agent’s hands perform co-speech gestures in perfect sync with the spoken words, providing clear, expressive spatial references.

By integrating these modules, AgentHands creates a seamless bridge between linguistic intent and physical action. The system transforms a standard LLM output into a rich, multimodal performance where the agent's generated responses are manifested through both speech and spatially accurate movement, allowing for complex instructions to be demonstrated exactly where they occur in the user's environment.

We demonstrated how these embodied gestures, paired with the spatial awareness of XR, enhance our understanding of our physical surroundings.

Interactive tutoring: In an orchid-care scenario, the agent doesn’t just say “check the roots”; it moves its hands to the base of the plant and outlines the air roots while explaining their function.

Technical walkthroughs: For 3D printer operations, the agent can demonstrate the exact ''turn and click'' sequence needed to navigate control knobs and select files, making complex physical interface steps intuitive.

Lifestyle companionship: The agent can serve as a wellness coach that interacts with your physical choices. For instance, the agent can perform an interactive “warning” gesture by holding the user’s hand and a visual effect to caution the user against unhealthy behavior.

To evaluate the impact of these gestures, we conducted a within-subjects study (N = 12) comparing AgentHands to a speech-only baseline. Both conditions used the same researcher-scripted verbal content, ensuring the only difference was the presence of the embodied hands and their synchronized gestures. Participants completed two procedural tasks that balanced everyday care with technical operation.

The results confirmed that the combination of XR and co-speech gestures is highly effective for spatially grounded interactions. We analyzed the data across several key metrics of communication effectiveness.

AgentHands represents a step toward a future where AI systems aren’t just analyzing our world, but dynamically operating within it. By leveraging co-speech gestures and the spatial power of XR to ground conversation in physical movement, we can reduce the cognitive load of complex tasks and make spatial computing more accessible and human-centric.

As we continue to develop for the Android XR ecosystem, we are exploring ways to make these gestures even more personalized, adapting to a user’s dominant hand or learning their specific spatial routines, to create an even more seamless human-AI collaboration.

This research was primarily conducted by Ziyi Liu during his Student Researcher tenure at Google, as part of a joint collaboration across multiple teams. We extend our sincere gratitude to key contributors David Li, Zhongyi Zhou, and David Kim for their support, and to Adarsh Kowdle, Guru Somadder, and Shahram Izadi for their strategic guidance and thoughtful reviews.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google xr 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agenthands-generatin…] indexed:0 read:5min 2026-08-25 ·