Open source · Sep 24, 2026 · as partner
Peking University DCAI team open-sources DataFlex-RL, a data strategy framework for RL post-training
Peking University DCAI Team, with University of Chinese Academy of Sciences, Shanghai Institute of Algorithm Innovation and Zhongguancun Academy, has open-sourced DataFlex-RL, a data-strategy framework for RL post-training built as a verl plugin. It turns data selection, sample reweighting and domain ratio adjustment into configurable components, so developers can test and compare strategies inside a single RLVR/GRPO pipeline.
The paper reports 591 experiments across Qwen2.5-7B-base, Llama-3.1-8B-base and math, logic and science tasks, and ranked second on Hugging Face's daily paper chart. Licence was not stated.
Original sources (Chinese)
日榜第二!北大开源,强化学习数据策略接入更简单、对比更公平