Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Unlocking Recursive Thinking of LLMs: Alignment via Refinement

cs.CL · 2025-06-06 · conditional · novelty 6.0

An offline alignment pipeline using reward-filtered self-refinement data and long chain-of-thought SFT raises an 8B model's AlpacaEval 2 win rate from 25.0% to 51.0% with roughly 14k training examples.

citing papers explorer

Showing 1 of 1 citing paper.

  • Unlocking Recursive Thinking of LLMs: Alignment via Refinement cs.CL · 2025-06-06 · conditional · none · ref 5

    An offline alignment pipeline using reward-filtered self-refinement data and long chain-of-thought SFT raises an 8B model's AlpacaEval 2 win rate from 25.0% to 51.0% with roughly 14k training examples.