Leave-one-turn deletion attribution, in which replacing a search turn with [DELETE] and measuring the drop in gold-answer likelihood yields a sign-gated, self-generated process reward, improves multi-turn search RL by 0.053 average EM over IGPO.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning
Leave-one-turn deletion attribution, in which replacing a search turn with [DELETE] and measuring the drop in gold-answer likelihood yields a sign-gated, self-generated process reward, improves multi-turn search RL by 0.053 average EM over IGPO.