ByteDance Seed and Tsinghua open-source DAPO for scalable LLM RL
ByteDance Seed and Tsinghua AIR open-sourced DAPO (Decoupled Clip and Dynamic Sampling Policy Optimization): algorithm, verl training stack, DAPO-Math-17k, and Qwen2.5-32B weights. Trained from scratch, DAPO-Qwen-32B hit 50% on AIME 2024 with about half the steps of DeepSeek-R1-Zero-Qwen-32B.