Pith. sign in

REVIEW 2 cited by

Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00530 v2 pith:DKGK4ZLJ submitted 2024-03-31 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords preferencejointpreferencesinstruction-responsehumanllmsoptimizationaligning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

A common technique for aligning large language models (LLMs) relies on acquiring human preferences by comparing multiple generations conditioned on a fixed context. This method, however, relies solely on pairwise comparisons, where the generations are evaluated within an identical context. While effective to such conditional preferences often fail to encompass the nuanced and multidimensional nature of human preferences. In this work, we revisit the traditional paradigm of preference acquisition and propose a new axis based on eliciting preferences jointly over the instruction-response pairs. Unlike prior preference optimizations, which are designed for conditional ranking protocols (e.g., DPO), we propose Joint Preference Optimization (JPO), a new preference optimization objective that upweights the joint probability of the chosen instruction-response pair over the rejected instruction-response pair. Interestingly, LLMs trained with joint instruction-response preference data using JPO outperform LLM trained with DPO by $5.2\%$ and $3.3\%$ win-rate for summarization and open-ended dialogue datasets, respectively. Our findings reveal that joint preferences over instruction and response pairs can significantly enhance the alignment of LLMs by tapping into a broader spectrum of human preference elicitation. The data and code is available at https://github.com/Hritikbansal/dove.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LPOI: Listwise Preference Optimization for Vision Language Models

    cs.CV 2025-05 conditional novelty 7.0 of 10

    LPOI reduces VLM hallucination by training the model to prefer the original image over progressively masked versions of the same image, using a listwise ranking loss built from pairwise preference data.

  2. ModelCitizens: Representing Community Voices in Online Safety

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A community-annotated toxicity dataset with conversational context shows that models trained on ingroup labels outperform state-of-the-art moderation APIs.

Pith tools