Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Generalist Reward Models: Found Inside Large Language Models

cs.CL · 2025-06-29 · conditional · novelty 4.0

The probability an LLM assigns to a response is treated as a reward, formally justified via an equivalence between next-token prediction and offline inverse RL, and then used to fine-tune the model itself.

citing papers explorer

Showing 1 of 1 citing paper.

  • Generalist Reward Models: Found Inside Large Language Models cs.CL · 2025-06-29 · conditional · none · ref 1

    The probability an LLM assigns to a response is treated as a reward, formally justified via an equivalence between next-token prediction and offline inverse RL, and then used to fine-tune the model itself.