Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Training Language Model to Critique for Better Refinement

cs.CL · 2025-06-27 · conditional · novelty 7.0

RCO trains critic LLMs using refinement utility, the extent to which a critique improves a revised answer, as the reward, outperforming direct critique preference methods across five tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Training Language Model to Critique for Better Refinement cs.CL · 2025-06-27 · conditional · none · ref 5

    RCO trains critic LLMs using refinement utility, the extent to which a critique improves a revised answer, as the reward, outperforming direct critique preference methods across five tasks.