Pith. sign in

REVIEW 2 cited by

Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.14723 v2 pith:RZ2VDEEX submitted 2024-04-23 cs.CL

classification cs.CL
keywords alignmentmodelinsightsinstructionmathematicalmethodsperformanceproblem-solving
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study evaluates Direct Preference Optimization (DPO) and its variants for aligning Large Language Models (LLMs) with human preferences, testing three configurations: (1) with Supervised Fine Tuning (SFT), (2) without SFT, and (3) without SFT but using an instruction tuned model. We further investigate how training set size influences model performance. Our evaluation spans 13 benchmarks covering dialogue, reasoning, mathematical problem-solving, question answering, truthfulness, MT-Bench, Big Bench, and the Open LLM Leaderboard. We find that: (1) alignment methods often achieve near optimal performance even with smaller subsets of training data; (2) although they offer limited improvements on complex reasoning tasks, they enhance mathematical problem-solving; and (3) using an instruction tuned model improves truthfulness. These insights highlight the conditions under which alignment methods excel, as well as their limitations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BPO: Revisiting Preference Modeling in Direct Preference Optimization

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Replacing DPO's relative reward margin with min(r_w, -α r_l) improves math benchmark accuracy by 5-10 points in this paper, though the exact gain figure is misreported and α is tuned on test data.

  2. Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M

    cs.CL 2025-09 reject novelty 3.0 of 10

    For OPT-350M on the Anthropic HH-RLHF set, SFT plus DPO gives the highest combined helpfulness/harmlessness score, but not the highest safety score.

Pith tools