← Back to the directory

University / Lab · Works with Peking University DCAI Team, University of Chinese Academy of Sciences

Zhongguancun Academy

中关村学院


Coverage1

Open source · Sep 24, 2026 · as partner

Peking University DCAI team open-sources DataFlex-RL, a data strategy framework for RL post-training

Peking University DCAI Team, with University of Chinese Academy of Sciences, Shanghai Institute of Algorithm Innovation and Zhongguancun Academy, has open-sourced DataFlex-RL, a data-strategy framework for RL post-training built as a verl plugin. It turns data selection, sample reweighting and domain ratio adjustment into configurable components, so developers can test and compare strategies inside a single RLVR/GRPO pipeline.

The paper reports 591 experiments across Qwen2.5-7B-base, Llama-3.1-8B-base and math, logic and science tasks, and ranked second on Hugging Face's daily paper chart. Licence was not stated.

Original sources (Chinese)

日榜第二!北大开源,强化学习数据策略接入更简单、对比更公平aiera