Pith. sign in

This retains the positive gradient properties 14 of logistic log-likelihood, allowing standard optimizers to converge efficiently

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO

cs.LG · 2025-06-10 · conditional · novelty 5.0

GFRIEND generates chain-of-thought preference judgments, scores them by perplexity, and uses weighted multi-level preference optimization so a reward model trained on 3,000 samples rivals models trained on much larger datasets.

citing papers explorer

Showing 1 of 1 citing paper.

  • GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO cs.LG · 2025-06-10 · conditional · none · ref 12

    GFRIEND generates chain-of-thought preference judgments, scores them by perplexity, and uses weighted multi-level preference optimization so a reward model trained on 3,000 samples rivals models trained on much larger datasets.