{"slug": "crayotter-learning-long-horizon-video-editing-agents-via-group-relative", "title": "Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation", "summary": "Researchers introduced Group-Relative Preference Backpropagation (GRPB), a method that converts same-task rankings into zero-sum advantages to train long-horizon video editing agents, and released the resulting 9B Crayotter model, which surpasses several proprietary systems on AgenticVBench. The team manually constructed a project-disjoint, horizon-stratified suite of realistic editing tasks for training and evaluation, and made code and materials publicly available at https://github.com/idwts/Crayotter.", "body_md": "arXiv:2608.02694v1 Announce Type: new\nAbstract: Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. Yet editing quality is subjective, admits multiple valid solutions, and is not meaningfully calibrated across heterogeneous requests, making a global scalar objective both ambiguous and temporally uninformative. Our key observation is that fixing the request, materials, and production constraints converts this subjective objective into an ordinal comparison among directly comparable alternatives. We introduce Group-Relative Preference Backpropagation (GRPB), which transforms same-task rankings into zero-sum advantages and redistributes them as bounded credit over semantic editing segments. A lagged allocator and guarded transmission prevent current judgments or unreliable estimates from directly shaping the same rollout group. We manually construct a project-disjoint, horizon-stratified suite of realistic editing tasks for training and controlled evaluation. Across matched baselines, credit interventions, external benchmarking, and blinded human evaluation, GRPB improves both editing behavior and rendered products. The resulting 9B Crayotter model surpasses several proprietary systems on AgenticVBench, supporting task-local preference reduction as a practical approach to learning from subjective, delayed outcomes. Code and all supporting materials are publicly available at https://github.com/idwts/Crayotter.", "url": "https://wpnews.pro/news/crayotter-learning-long-horizon-video-editing-agents-via-group-relative", "canonical_source": "https://arxiv.org/abs/2608.02694", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 04:03:39.547562+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-agents", "ai-research"], "entities": ["Crayotter", "Group-Relative Preference Backpropagation", "AgenticVBench", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/crayotter-learning-long-horizon-video-editing-agents-via-group-relative", "markdown": "https://wpnews.pro/news/crayotter-learning-long-horizon-video-editing-agents-via-group-relative.md", "text": "https://wpnews.pro/news/crayotter-learning-long-horizon-video-editing-agents-via-group-relative.txt", "jsonld": "https://wpnews.pro/news/crayotter-learning-long-horizon-video-editing-agents-via-group-relative.jsonld"}}