Pith. sign in

REVIEW 2 cited by

Learning from negative feedback, or positive feedback or both

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.04166 v3 pith:YTFS5M7K submitted 2024-10-05 cs.LG stat.ML

classification cs.LGstat.ML
keywords feedbacknegativepositivelearningapproachavailableexamplesmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing preference optimization methods often assume scenarios where paired preference feedback (preferred/positive vs. dis-preferred/negative examples) is available. This requirement limits their applicability in scenarios where only unpaired feedback--for example, either positive or negative--is available. To address this, we introduce a novel approach that decouples learning from positive and negative feedback. This decoupling enables control over the influence of each feedback type and, importantly, allows learning even when only one feedback type is present. A key contribution is demonstrating stable learning from negative feedback alone, a capability not well-addressed by current methods. Our approach builds upon the probabilistic framework introduced in (Dayan and Hinton, 1997), which uses expectation-maximization (EM) to directly optimize the probability of positive outcomes (as opposed to classic expected reward maximization). We address a key limitation in current EM-based methods: they solely maximize the likelihood of positive examples, while neglecting negative ones. We show how to extend EM algorithms to explicitly incorporate negative examples, leading to a theoretically grounded algorithm that offers an intuitive and versatile way to learn from both positive and negative feedback. We evaluate our approach for training language models based on human feedback as well as training policies for sequential decision-making problems, where learned value functions are available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PIPA: Preference Alignment as Prior-Informed Statistical Estimation

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A unified maximum-likelihood framework with prior constraints that recovers DPO and KTO as special cases and yields new PIPA-M/PIPA-N losses with 3-10% gains on GSM8K and MATH.

  2. Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

    cs.LG 2025-07 conditional novelty 5.0 of 10

    SFT on curated data is a lower bound on a sparse-reward RL objective, and an importance-weighted variant, iw-SFT, tightens the bound and beats plain SFT on AIME 2024 and GPQA.

Pith tools