Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CL 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

Weighted-Reward Preference Optimization for Implicit Model Fusion

cs.CL · 2024-12-04 · conditional · novelty 6.0

WRPO tunes an 8B chat model by combining its own preferred responses (on-policy) with high-reward responses from ten heterogeneous source LLMs (off-policy) using an annealed weight, beating prior fusion and preference-optimization baselines on AlpacaEval-2, Arena-Hard, and MT-Bench.

citing papers explorer

Showing 1 of 1 citing paper.

  • Weighted-Reward Preference Optimization for Implicit Model Fusion cs.CL · 2024-12-04 · conditional · none · ref 4

    WRPO tunes an 8B chat model by combining its own preferred responses (on-policy) with high-reward responses from ten heterogeneous source LLMs (off-policy) using an annealed weight, beating prior fusion and preference-optimization baselines on AlpacaEval-2, Arena-Hard, and MT-Bench.