PLAN-TUNING post-trains small LLMs to generate step-by-step plans before answering, improving math benchmark accuracy over standard SFT and GRPO baselines.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
PLAN-TUNING post-trains small LLMs to generate step-by-step plans before answering, improving math benchmark accuracy over standard SFT and GRPO baselines.