Pith. sign in

REVIEW 2 cited by

Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.16681 v2 pith:NI4ABZHO submitted 2024-05-26 cs.CL

classification cs.CL
keywords optimizationpreferencepointsreasoningachievesalignmentdesignedimprovements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement Learning with Human Feedback (RLHF) enhances the alignment of Large Language Models (LLMs). However, its limitations have led to the development of Direct Preference Optimization (DPO), an RL-free approach designed to overcome these shortcomings. While studies have shown that DPO improves instruction-following capabilities, it negatively impacts the reasoning ability of LLMs. Additionally, DPO is highly sensitive to judgment noise in preference datasets and the size of the training set. Although several modifications to DPO have been proposed, they still fail to fully resolve these issues. To address these limitations, we propose Triple Preference Optimization (TPO), a new preference learning method designed to enhance both reasoning and instruction-following abilities through one-step optimization. We compare TPO against DPO and its recent variants using state-of-the-art training setups, including both base and instruction-tuned models such as Mistral and Llama 3. Our evaluation covers a comprehensive range of chat-based and reasoning benchmarks. The results demonstrate that TPO achieves significant improvements over existing methods without substantially increasing response length across different dataset sizes. Specifically, TPO outperforms DPO and SimPO by up to 7.0% and 7.3% points on Arena-Hard, 12.2% and 13.3% points on MixEval-Hard, 10.4% and 10.1% points on MMLU-Pro, and 19.0% and 19.2% points on GSM8K, respectively. Furthermore, TPO achieves these improvements while requiring less data than DPO.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ThinkTuning: Instilling Cognitive Reflections without Distillation

    cs.AI 2025-08 conditional novelty 5.0 of 10

    ThinkTuning uses teacher feedback grafted onto GRPO rollouts to teach small language models reflective reasoning, improving MATH-500, AIME, and GPQA-Diamond accuracy over vanilla GRPO.

  2. PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving

    cs.CL 2025-07 conditional novelty 4.0 of 10

    PLAN-TUNING post-trains small LLMs to generate step-by-step plans before answering, improving math benchmark accuracy over standard SFT and GRPO baselines.

Pith tools