← back to paper
arxiv: 2608.00536 · 2 revisions
DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards