Sequential SFT followed by RL, guided by the Plasticity-Ceiling Framework, achieves higher performance ceilings in LLM mathematical reasoning than synchronized methods by optimizing data scale and transition timing.
Implicit reward as the bridge: A unified view of sft and dpo connections
3 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
LoRR augments preference optimization methods like DPO with high-replay training, periodic resets to initial data/policy, and a hybrid objective to improve sample efficiency and reduce primacy bias on math and reasoning tasks.
CADFT improves supervised fine-tuning of large language models by dynamically down-weighting training samples whose low model-likelihood indicates high gradient variance, yielding better stability and generalization.
citing papers explorer
-
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
Sequential SFT followed by RL, guided by the Plasticity-Ceiling Framework, achieves higher performance ceilings in LLM mathematical reasoning than synchronized methods by optimizing data scale and transition timing.
-
Sample-efficient LLM Optimization with Reset Replay
LoRR augments preference optimization methods like DPO with high-replay training, periodic resets to initial data/policy, and a hybrid objective to improve sample efficiency and reduce primacy bias on math and reasoning tasks.
-
Compatibility-Aware Dynamic Fine-Tuning for Large Language Models
CADFT improves supervised fine-tuning of large language models by dynamically down-weighting training samples whose low model-likelihood indicates high gradient variance, yielding better stability and generalization.