ShiQ is a Bellman-derived loss that makes LLM logits behave as Q-values, enabling off-policy token-level and multi-turn reinforcement-learning fine-tuning with the softmax over logits as the policy.
A general theoretical paradigm to understand learning from human preferences
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ShiQ: Bringing back Bellman to LLMs
ShiQ is a Bellman-derived loss that makes LLM logits behave as Q-values, enabling off-policy token-level and multi-turn reinforcement-learning fine-tuning with the softmax over logits as the policy.