Step-level self-distilled advantage weights, applied only to incorrect trajectories, make GRPO training of deep web-search agents roughly twice as sample-efficient on Qwen3-8B.
SFT memorizes, RL generalizes
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
Step-level self-distilled advantage weights, applied only to incorrect trajectories, make GRPO training of deep web-search agents roughly twice as sample-efficient on Qwen3-8B.