Pith. sign in

REVIEW 1 cited by

BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02479 v2 pith:IUD63DKH submitted 2024-02-04 cs.LG cs.AIcs.CLcs.HC

classification cs.LGcs.AIcs.CLcs.HC
keywords methodsdistributionreward-conditionedamortizedbayesianbraindistributionalfeedback
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PIPA: Preference Alignment as Prior-Informed Statistical Estimation

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A unified maximum-likelihood framework with prior constraints that recovers DPO and KTO as special cases and yields new PIPA-M/PIPA-N losses with 3-10% gains on GSM8K and MATH.

Pith tools