Pith. sign in

REVIEW 1 cited by

Active Learning for Direct Preference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.01076 v1 pith:MKKCRB52 submitted 2025-03-03 cs.LG stat.ML

classification cs.LGstat.ML
keywords feedbackhumanlearningactivealgorithmscollectdirectinformative
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Direct preference optimization (DPO) is a form of reinforcement learning from human feedback (RLHF) where the policy is learned directly from preferential feedback. Although many models of human preferences exist, the critical task of selecting the most informative feedback for training them is under-explored. We propose an active learning framework for DPO, which can be applied to collect human feedback online or to choose the most informative subset of already collected feedback offline. We propose efficient algorithms for both settings. The key idea is to linearize the DPO objective at the last layer of the neural network representation of the optimized policy and then compute the D-optimal design to collect preferential feedback. We prove that the errors in our DPO logit estimates diminish with more feedback. We show the effectiveness of our algorithms empirically in the setting that matches our theory and also on large language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Reviews to Dialogues: Active Synthesis for Zero-Shot LLM-based Conversational Recommender System

    cs.IR 2025-04 conditional novelty 5.0 of 10

    Active sample selection over review, metadata, and collaborative seed data plus LLM-generated synthetic dialogues improves fine-tuned conversational recommendation on ReDial and INSPIRED, though not uniformly across a...

Pith tools