{"slug": "bytedance-adapts-grpo-for-enhanced-post-training-model-capabilities", "title": "ByteDance adapts GRPO for enhanced post-training model capabilities", "summary": "ByteDance's Seed team, working with University of Hong Kong researchers, adapted DeepSeek's Group Relative Policy Optimization (GRPO) technique into a visual-generation framework called DanceGRPO, reporting benchmark gains of up to 181% on HPS-v2.1 and CLIP Score across text-to-image, text-to-video, and image-to-video tasks. A follow-on project, BranchGRPO, announced around September 2025, reported alignment score improvements of up to 16% while cutting training time by 55%. ByteDance also developed DAPO with Tsinghua University's AIR lab, which scored 50 on the AIME 2024 math reasoning benchmark using 50% fewer training steps than a comparable DeepSeek-R1 setup.", "body_md": "# ByteDance adapts GRPO for enhanced post-training model capabilities\n\nThe TikTok parent's research team reworked a DeepSeek technique to align visual generation models with human preferences, posting benchmark gains of up to 181%.\n\nByteDance’s AI research division has taken a reinforcement learning technique originally designed for large language models and retrofitted it for visual generation, producing a framework called DanceGRPO that delivers significant quality improvements across text-to-image, text-to-video, and image-to-video tasks.\n\nThe work, developed by ByteDance’s Seed team in collaboration with researchers at the University of Hong Kong, represents one of the more ambitious attempts to solve a persistent headache in generative AI: getting diffusion models and rectified flow models to actually produce what humans want.\n\n## From language to visuals\n\nGroup Relative Policy Optimization, or GRPO, first appeared in April 2024 as part of DeepSeek’s DeepSeekMath research. Its core innovation was elegant. Instead of training a separate critic model to evaluate outputs (the standard approach in reinforcement learning from human feedback), GRPO scores outputs relative to a group of samples. That architectural shortcut made the whole training process cheaper and more efficient.\n\nByteDance’s contribution was figuring out how to apply that same logic to visual generation, which is a fundamentally different problem. Language models produce tokens in sequence. Diffusion models generate images and videos through iterative denoising, a process governed by stochastic differential equations. DanceGRPO reworks the sampling process to make GRPO’s group-based relative scoring compatible with these visual pipelines.\n\nThe result is a framework that works across multiple foundational models, including Stable Diffusion, FLUX, HunyuanVideo, and SkyReels-I2V. It supports at least five different reward types, giving researchers flexibility in how they define “good” output.\n\n## The numbers behind the improvement\n\nBenchmark improvements tell the story most clearly. On metrics like HPS-v2.1 and CLIP Score, which measure how well generated images align with text prompts and human preferences, DanceGRPO recorded enhancements of up to 181%.\n\nA follow-on project called BranchGRPO, announced around September 2025, pushed the approach further. It reported alignment score improvements of up to 16% while cutting training time by 55%.\n\n## Beyond visual generation\n\nByteDance also developed DAPO, short for Decoupled Clip and Dynamic Sampling Policy Optimization, in collaboration with Tsinghua University’s AIR lab. DAPO targets reasoning capabilities in language models.\n\nDAPO achieved a score of 50 on the AIME 2024 benchmarks, which test mathematical reasoning ability, while using 50% fewer training steps than a comparable DeepSeek-R1 setup.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/bytedance-adapts-grpo-for-enhanced-post-training-model-capabilities", "canonical_source": "https://cryptobriefing.com/bytedance-dancegrpo-post-training-visual-generation/", "published_at": "2026-09-10 14:58:21+00:00", "updated_at": "2026-09-10 15:18:17.646572+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "ai-research", "computer-vision"], "entities": ["ByteDance", "Seed", "University of Hong Kong", "DeepSeek", "GRPO", "DanceGRPO", "BranchGRPO", "DAPO"], "alternates": {"html": "https://wpnews.pro/news/bytedance-adapts-grpo-for-enhanced-post-training-model-capabilities", "markdown": "https://wpnews.pro/news/bytedance-adapts-grpo-for-enhanced-post-training-model-capabilities.md", "text": "https://wpnews.pro/news/bytedance-adapts-grpo-for-enhanced-post-training-model-capabilities.txt", "jsonld": "https://wpnews.pro/news/bytedance-adapts-grpo-for-enhanced-post-training-model-capabilities.jsonld"}}