A single federated reward model can personalize faster than group-specific models when preference groups are balanced, and group-debiased sampling restores this under imbalance.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning
A single federated reward model can personalize faster than group-specific models when preference groups are balanced, and group-debiased sampling restores this under imbalance.