字节跳动种子基金与清华大学AIR实验室联合发布了开源强化学习系统DAPO1。该系统基于解耦裁剪和动态采样策略优化算法1,包含完整的算法、代码基础设施和数据集的开源实现1。
在AIME 2024数学竞赛上,DAPO-Qwen-32B模型取得50分的成绩1,超越了此前DeepSeek-R1-Zero-Qwen-32B的表现且训练用时减少了50%1。系统的开源版本包括DAPO-Math-17k训练数据集、AIME 2024验证集以及完整的训练脚本和基础设施1,该项目基于verl框架构建,并在火山引擎机器学习平台进行了实验验证1。
系统的发展经历了两个阶段的迭代:2025年3月发布的早期版本在AIME 2024上达到44%的成绩1,而2025年5月更新的完整版本则达到50%以上,并提供了评估指南1。
ByteDance's Seed Fund and Tsinghua University's AIR Laboratory have jointly released DAPO, an open-source reinforcement learning system designed for mathematical reasoning tasks 1. The system employs decoupled actor pruning and dynamic sampling policy optimization algorithms to enhance performance on complex problem-solving benchmarks 1.
DAPO-Qwen-32B achieved a score of 50 on the AIME 2024 mathematics competition, surpassing the previous result from DeepSeek-R1-Zero-Qwen-32B while reducing computational time by 50 percent 1. The system's development followed two release milestones: an early version released in March 2025 that reached 44 percent on AIME 2024, followed by the complete version in May 2025 that achieved over 50 percent 1. The open-source release includes the DAPO-Math-17k training dataset, the AIME 2024 validation set, complete training scripts, and infrastructure built on the verl framework, with experimental validation conducted on Volcengine's machine learning platform 1.
评论
还没有评论,欢迎留下第一条。