{"slug": "egoplay-event-triggered-video-editing-for-egocentric-streams", "title": "EgoPlay: Event-Triggered Video Editing for Egocentric Streams", "summary": "Researchers introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams that jointly learns event recognition, temporal restraint, and pixel-level editing in a single end-to-end model. On the Ego4D benchmark, EgoPlay outperforms the state-of-the-art baseline EgoEdit with relative gains of 17.7%, 16.9%, and 16.4% in editing quality, visual quality, and background consistency, while using less than half the GPU memory.", "body_md": "arXiv:2607.24560v1 Announce Type: cross\nAbstract: We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form \"when X happens, do Y,\" EgoPlay infers whether and when event X occurs, preserves pre-event frames, and applies edit Y only to the post-event continuation. Rather than cascading a separate event detector with an editor, EgoPlay learns event recognition, temporal restraint, and pixel-level editing jointly in a single end-to-end model, while also handling negative and multi-event prompts. To support this, we construct a large-scale dataset of 106K event-triggered clip-prompt pairs spanning positive triggers, fabricated-trigger negatives, and multi-event prompts. We then train a bidirectional video diffusion editor with event-triggered supervision and derive a causal variant for chunk-by-chunk streamable inference. We further introduce an event-aware evaluation protocol that separately measures post-trigger editing quality, pre-trigger preservation, and false-trigger robustness. On the Ego4D benchmark, EgoPlay substantially outperforms EgoEdit, the state-of-the-art instruction-based egocentric video editing baseline, with relative gains of 17.7%, 16.9%, and 16.4% in editing quality, visual quality, and background consistency. It also surpasses a VLM-guided detector-editor baseline by 15.7%, 14.5%, and 13.5% on the same metrics, while using less than half the GPU memory.", "url": "https://wpnews.pro/news/egoplay-event-triggered-video-editing-for-egocentric-streams", "canonical_source": "https://www.machinebrief.com/news/egoplay-event-triggered-video-editing-for-egocentric-streams-t9ge", "published_at": "2026-07-28 04:00:00+00:00", "updated_at": "2026-07-28 04:57:40.743720+00:00", "lang": "en", "topics": ["artificial-intelligence", "computer-vision", "generative-ai", "ai-research"], "entities": ["EgoPlay", "Ego4D", "EgoEdit", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/egoplay-event-triggered-video-editing-for-egocentric-streams", "markdown": "https://wpnews.pro/news/egoplay-event-triggered-video-editing-for-egocentric-streams.md", "text": "https://wpnews.pro/news/egoplay-event-triggered-video-editing-for-egocentric-streams.txt", "jsonld": "https://wpnews.pro/news/egoplay-event-triggered-video-editing-for-egocentric-streams.jsonld"}}