In contrast, GRPO-Full averages informative and weakly informative trajectories uniformly at every step, diluting the retrieval-specific supervision signal
This reflects a reinforcing selection dynamic: as training progresses, the model naturally produces more deep-search trajectories, CuSearch continues to concentrate gradient updates on those trajectories, progressively strengthening the
· 2022
1 Pith paper cite this work. Polarity classification is still indexing.