DAPO introduces decoupled clipping and dynamic sampling for LLM RL, achieving 50 on AIME 2024 with Qwen2.5-32B while fully open-sourcing code, data, and the verl-based training system.
Training language models to follow instructions with human feedback
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2025 2representative citing papers
citing papers explorer
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
DAPO introduces decoupled clipping and dynamic sampling for LLM RL, achieving 50 on AIME 2024 with Qwen2.5-32B while fully open-sourcing code, data, and the verl-based training system.
- MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent