cd /news/artificial-intelligence/puppeteer-object-grounded-posture-aw… · home topics artificial-intelligence article
[ARTICLE · art-118516] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

Researchers introduced Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model that operates in a causal latent space, enabling the synthesis of physically consistent gestures conditioned on speech, motion history, initial posture, and object geometry. The model decomposes gestures into structured primitives encoded as temporally ordered latent tokens, supporting tasks like gesture in-betweening and completion, and outperforms prior methods in diversity and temporal synchronization. The team also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures with corresponding 3D objects, to facilitate object-grounded gesture generation.

read1 min views1 publishedSep 2, 2026

arXiv:2609.00369v1 Announce Type: new Abstract: Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @puppeteer 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/puppeteer-object-gro…] indexed:0 read:1min 2026-09-02 ·