← back to paper
arxiv: 2509.16456 · 2 revisions
GPO: Learning from Critical Steps to Improve LLM Reasoning