EduAlign trains a three-dimensional reward model (HPC-RM) and uses GRPO to fine-tune Qwen2.5-72B, reporting improved helpfulness, personalization, and creativity on its own and public benchmarks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning
EduAlign trains a three-dimensional reward model (HPC-RM) and uses GRPO to fine-tune Qwen2.5-72B, reporting improved helpfulness, personalization, and creativity on its own and public benchmarks.