Pith. sign in

REVIEW 1 cited by

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.04567 v1 pith:H2WOAKZ7 submitted 2025-02-06 cs.AI

classification cs.AI
keywords completionsdispreferredmc-pomodelpreferenceproposealgorithmcontrastive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing studies on preference optimization (PO) have centered on constructing pairwise preference data following simple heuristics, such as maximizing the margin between preferred and dispreferred completions based on human (or AI) ranked scores. However, none of these heuristics has a full theoretical justification. In this work, we develop a novel PO framework that provides theoretical guidance to effectively sample dispreferred completions. To achieve this, we formulate PO as minimizing the negative log-likelihood (NLL) of a probability model and propose to estimate its normalization constant via a sampling strategy. As we will demonstrate, these estimative samples can act as dispreferred completions in PO. We then select contrastive divergence (CD) as the sampling strategy, and propose a novel MC-PO algorithm that applies the Monte Carlo (MC) kernel from CD to sample hard negatives w.r.t. the parameterized reward model. Finally, we propose the OnMC-PO algorithm, an extension of MC-PO to the online setting. On popular alignment benchmarks, MC-PO outperforms existing SOTA baselines, and OnMC-PO leads to further improvement.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. York's Cavity Formalism and Quantum Modified Thermodynamics of (2+1)D Black Holes

    gr-qc 2025-06 reject novelty 4.0 of 10

    The paper claims Barrow entropy corrections reshape BTZ black hole thermodynamics in a cavity, but the free energy analysis rests on a wrong extrinsic curvature term and an incorrect zero-crossing interpretation.

Pith tools