PAFO applies Pareto fairness optimization and group-specialized distillation to produce a single personalized reward model that improves accuracy for both majority and minority preference groups without requiring group labels at inference.
Multi-objective alignment of large language models through hypervolume maximization
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
MOCHA combines Chebyshev scalarization with exponential annealing to optimize LLM agent skills across performance and platform constraints, improving mean correctness by 7.5% over baselines on six tasks while finding more Pareto-optimal variants.
Focal Reward balances rubric-based RL by saturation-aware reweighting derived from inverse reward projection, outperforming static aggregation on 18 model-benchmark pairs.
citing papers explorer
-
PAFO: Pareto Fairness Optimization for Personalized Reward Modeling
PAFO applies Pareto fairness optimization and group-specialized distillation to produce a single personalized reward model that improves accuracy for both majority and minority preference groups without requiring group labels at inference.
-
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
MOCHA combines Chebyshev scalarization with exponential annealing to optimize LLM agent skills across performance and platform constraints, improving mean correctness by 7.5% over baselines on six tasks while finding more Pareto-optimal variants.
-
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
Focal Reward balances rubric-based RL by saturation-aware reweighting derived from inverse reward projection, outperforming static aggregation on 18 model-benchmark pairs.