cd /news/artificial-intelligence/bytedance-adapts-grpo-for-enhanced-p… · home topics artificial-intelligence article
[ARTICLE · art-125914] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

ByteDance adapts GRPO for enhanced post-training model capabilities

ByteDance's Seed team, working with University of Hong Kong researchers, adapted DeepSeek's Group Relative Policy Optimization (GRPO) technique into a visual-generation framework called DanceGRPO, reporting benchmark gains of up to 181% on HPS-v2.1 and CLIP Score across text-to-image, text-to-video, and image-to-video tasks. A follow-on project, BranchGRPO, announced around September 2025, reported alignment score improvements of up to 16% while cutting training time by 55%. ByteDance also developed DAPO with Tsinghua University's AIR lab, which scored 50 on the AIME 2024 math reasoning benchmark using 50% fewer training steps than a comparable DeepSeek-R1 setup.

read2 min views2 publishedSep 10, 2026
ByteDance adapts GRPO for enhanced post-training model capabilities
Image: Cryptobriefing (auto-discovered)

The TikTok parent's research team reworked a DeepSeek technique to align visual generation models with human preferences, posting benchmark gains of up to 181%.

ByteDance’s AI research division has taken a reinforcement learning technique originally designed for large language models and retrofitted it for visual generation, producing a framework called DanceGRPO that delivers significant quality improvements across text-to-image, text-to-video, and image-to-video tasks.

The work, developed by ByteDance’s Seed team in collaboration with researchers at the University of Hong Kong, represents one of the more ambitious attempts to solve a persistent headache in generative AI: getting diffusion models and rectified flow models to actually produce what humans want.

From language to visuals #

Group Relative Policy Optimization, or GRPO, first appeared in April 2024 as part of DeepSeek’s DeepSeekMath research. Its core innovation was elegant. Instead of training a separate critic model to evaluate outputs (the standard approach in reinforcement learning from human feedback), GRPO scores outputs relative to a group of samples. That architectural shortcut made the whole training process cheaper and more efficient.

ByteDance’s contribution was figuring out how to apply that same logic to visual generation, which is a fundamentally different problem. Language models produce tokens in sequence. Diffusion models generate images and videos through iterative denoising, a process governed by stochastic differential equations. DanceGRPO reworks the sampling process to make GRPO’s group-based relative scoring compatible with these visual pipelines.

The result is a framework that works across multiple foundational models, including Stable Diffusion, FLUX, HunyuanVideo, and SkyReels-I2V. It supports at least five different reward types, giving researchers flexibility in how they define “good” output.

The numbers behind the improvement #

Benchmark improvements tell the story most clearly. On metrics like HPS-v2.1 and CLIP Score, which measure how well generated images align with text prompts and human preferences, DanceGRPO recorded enhancements of up to 181%.

A follow-on project called BranchGRPO, announced around September 2025, pushed the approach further. It reported alignment score improvements of up to 16% while cutting training time by 55%.

Beyond visual generation #

ByteDance also developed DAPO, short for Decoupled Clip and Dynamic Sampling Policy Optimization, in collaboration with Tsinghua University’s AIR lab. DAPO targets reasoning capabilities in language models.

DAPO achieved a score of 50 on the AIME 2024 benchmarks, which test mathematical reasoning ability, while using 50% fewer training steps than a comparable DeepSeek-R1 setup.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @bytedance 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/bytedance-adapts-grp…] indexed:0 read:2min 2026-09-10 ·