REVIEW 3 cited by
Binary Classifier Optimization for Large Language Model Alignment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In real-world services such as ChatGPT, aligning models based on user feedback is crucial for improving model performance. However, due to the simplicity and convenience of providing feedback, users typically offer only basic binary signals, such as 'thumbs-up' or 'thumbs-down'. Most existing alignment research, on the other hand, relies on preference-based approaches that require both positive and negative responses as a pair. We propose Binary Classifier Optimization (BCO), a technique that effectively aligns LLMs using only binary feedback. BCO trains a binary classifier, where the logit serves as an implicit reward, effectively minimizing the Direct Preference Optimization (DPO) loss. We demonstrate that the binary cross-entropy loss employed in classifier training acts as an upper bound for the DPO loss. Additionally, a novel reward shift technique further minimizes the gap between the losses. We validate our methodology in two settings: first, on a paired preference dataset, where our method performs on par with DPO; and second, on a Likert-5 scale annotation dataset which stems from real users' queries. Our model consistently demonstrates effective and robust alignment across four base LLMs and three different datasets, showcasing the strength of our approach to learning from binary signals.
Forward citations
Cited by 3 Pith papers
-
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
CALMA is a grounded-theory, participatory method for deriving community-specific language model alignment axes from open-ended user interactions and group discussion, piloted with two small groups.
-
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
The paper introduces TRADE, an online-only MCP attack, and RAG-Pref, a retrieval-based preference alignment method that together with DPO improves strict refusal of falsely benign MCP exploits from 6.7% to 24.1% on average.
-
Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization
OTPO uses unbalanced optimal transport to reweight token-level log-likelihood ratios in DPO, reporting up to 10.9% higher length-controlled win rate on AlpacaEval2 than DPO.
Discussion (0). Continue with ORCID to comment.