DHRCL, a staged RL curriculum with dense hierarchical rewards and token-level credit redistribution, improves code LLM Pass@1 by about 1 point over the strongest baseline in controlled Qwen3 experiments.
InTheThir- teenth International Conference on Learning Representa- tions, ICLR 2025, Singapore, April 24-28,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning
DHRCL, a staged RL curriculum with dense hierarchical rewards and token-level credit redistribution, improves code LLM Pass@1 by about 1 point over the strongest baseline in controlled Qwen3 experiments.