Pith. sign in

A general theoretical paradigm to understand learning from human preferences

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

ShiQ: Bringing back Bellman to LLMs

cs.LG · 2025-05-16 · conditional · novelty 6.0

ShiQ is a Bellman-derived loss that makes LLM logits behave as Q-values, enabling off-policy token-level and multi-turn reinforcement-learning fine-tuning with the softmax over logits as the policy.

citing papers explorer

Showing 1 of 1 citing paper.

  • ShiQ: Bringing back Bellman to LLMs cs.LG · 2025-05-16 · conditional · none · ref 3

    ShiQ is a Bellman-derived loss that makes LLM logits behave as Q-values, enabling off-policy token-level and multi-turn reinforcement-learning fine-tuning with the softmax over logits as the policy.