Pith. sign in

REVIEW 4 cited by

Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.05469 v1 pith:5SYIOFBQ submitted 2024-12-06 cs.LG

classification cs.LG
keywords humanmoahfmulti-objectivepreferencesalignmentdiversehypervolumelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-objective alignment from human feedback (MOAHF) in large language models (LLMs) is a challenging problem as human preferences are complex, multifaceted, and often conflicting. Recent works on MOAHF considered a-priori multi-objective optimization (MOO), where human preferences are known at training or inference time. In contrast, when human preferences are unknown or difficult to quantify, a natural approach is to cover the Pareto front by multiple diverse solutions. We propose an algorithm HaM for learning diverse LLM policies that maximizes their hypervolume. This is the first application of a-posteriori MOO to MOAHF. HaM is computationally and space efficient, and empirically superior across objectives such as harmlessness, helpfulness, humor, faithfulness, and hallucination, on various datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs

    cs.AI 2025-11 conditional novelty 6.0 of 10

    Benign PEFT fine-tuning changes LLM safety and fairness: adapter-based methods (LoRA, IA3) preserve alignment better than prompt-based methods, and the base model strongly moderates outcomes.

  2. Multi-objective Large Language Model Alignment with Hierarchical Experts

    cs.CL 2025-05 conditional novelty 6.0 of 10

    HoE claims to align a single LLM to any preference vector over multiple objectives using training-free LoRA experts, lightweight trained routers, and nearest-neighbor preference routing.

  3. AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models

    cs.LG 2025-06 reject novelty 5.0 of 10

    AMoPO uses the model's own token probabilities to define Gaussian-sampled weights, combining per-dimension SimPO-style losses for reference-free multi-objective alignment.

  4. Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models

    cs.CL 2025-07 reject novelty 4.0 of 10

    GAPO combines multiple-gradient descent with gradient rescaling to balance helpfulness and harmlessness in RLHF, and P-GAPO adds user preference weights.

Pith tools