Pith. sign in

REVIEW 3 cited by

Binary Classifier Optimization for Large Language Model Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.04656 v2 pith:SW5UKBBS submitted 2024-04-06 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords binaryclassifieralignmentfeedbacklossmodeloptimizationdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In real-world services such as ChatGPT, aligning models based on user feedback is crucial for improving model performance. However, due to the simplicity and convenience of providing feedback, users typically offer only basic binary signals, such as 'thumbs-up' or 'thumbs-down'. Most existing alignment research, on the other hand, relies on preference-based approaches that require both positive and negative responses as a pair. We propose Binary Classifier Optimization (BCO), a technique that effectively aligns LLMs using only binary feedback. BCO trains a binary classifier, where the logit serves as an implicit reward, effectively minimizing the Direct Preference Optimization (DPO) loss. We demonstrate that the binary cross-entropy loss employed in classifier training acts as an upper bound for the DPO loss. Additionally, a novel reward shift technique further minimizes the gap between the losses. We validate our methodology in two settings: first, on a paired preference dataset, where our method performs on par with DPO; and second, on a Likert-5 scale annotation dataset which stems from real users' queries. Our model consistently demonstrates effective and robust alignment across four base LLMs and three different datasets, showcasing the strength of our approach to learning from binary signals.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment

    cs.CY 2025-07 conditional novelty 6.0 of 10

    CALMA is a grounded-theory, participatory method for deriving community-specific language model alignment axes from open-ended user interactions and group discussion, piloted with two small groups.

  2. MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment

    cs.LG 2025-05 conditional novelty 6.0 of 10

    The paper introduces TRADE, an online-only MCP attack, and RAG-Pref, a retrieval-based preference alignment method that together with DPO improves strict refusal of falsely benign MCP exploits from 6.7% to 24.1% on average.

  3. Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization

    cs.CL 2025-05 conditional novelty 6.0 of 10

    OTPO uses unbalanced optimal transport to reweight token-level log-likelihood ratios in DPO, reporting up to 10.9% higher length-controlled win rate on AlpacaEval2 than DPO.

Pith tools