ByteDance’s AI research division has taken a reinforcement learning technique originally designed for large language models and retrofitted it for visual generation, producing a framework called DanceGRPO that delivers significant quality improvements across text-to-image, text-to-video, and image-to-video tasks.
The work, developed by ByteDance’s Seed team in collaboration with researchers at the University of Hong Kong, represents one of the more ambitious attempts to solve a persistent headache in generative AI: getting diffusion models and rectified flow models to actually produce what humans want.
From language to visuals
Group Relative Policy Optimization, or GRPO, first appeared in April 2024 as part of DeepSeek’s DeepSeekMath research. Its core innovation was elegant. Instead of training a separate critic model to evaluate outputs (the standard approach in reinforcement learning from human feedback), GRPO scores outputs relative to a group of samples. That architectural shortcut made the whole training process cheaper and more efficient.
ByteDance’s contribution was figuring out how to apply that same logic to visual generation, which is a fundamentally different problem. Language models produce tokens in sequence. Diffusion models generate images and videos through iterative denoising, a process governed by stochastic differential equations. DanceGRPO reworks the sampling process to make GRPO’s group-based relative scoring compatible with these visual pipelines.
The result is a framework that works across multiple foundational models, including Stable Diffusion, FLUX, HunyuanVideo, and SkyReels-I2V. It supports at least five different reward types, giving researchers flexibility in how they define “good” output.
The numbers behind the improvement
Benchmark improvements tell the story most clearly. On metrics like HPS-v2.1 and CLIP Score, which measure how well generated images align with text prompts and human preferences, DanceGRPO recorded enhancements of up to 181%.
A follow-on project called BranchGRPO, announced around September 2025, pushed the approach further. It reported alignment score improvements of up to 16% while cutting training time by 55%.
Beyond visual generation
ByteDance also developed DAPO, short for Decoupled Clip and Dynamic Sampling Policy Optimization, in collaboration with Tsinghua University’s AIR lab. DAPO targets reasoning capabilities in language models.
DAPO achieved a score of 50 on the AIME 2024 benchmarks, which test mathematical reasoning ability, while using 50% fewer training steps than a comparable DeepSeek-R1 setup.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

41 minutes ago
9









English (US) ·