A two-stage SFT plus DPO pipeline on synthetic pedagogical data improves factual accuracy and tutoring quality in LLM math tutors over base and prior models.
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
citing papers explorer
-
Towards Pedagogically Aligned LLM Tutors for Math Mistake Remediation
A two-stage SFT plus DPO pipeline on synthetic pedagogical data improves factual accuracy and tutoring quality in LLM math tutors over base and prior models.
- SocialCoach: Personalized Social Skill Learning with Agentic Tutoring and Practice