A three-phase post-training pipeline (LP-FFT-RFT) aligns a recommender foundation model with business metrics and outperforms single-phase fine-tuning and direct reward-model ranking in offline and online experiments.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training
A three-phase post-training pipeline (LP-FFT-RFT) aligns a recommender foundation model with business metrics and outperforms single-phase fine-tuning and direct reward-model ranking in offline and online experiments.