Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T20:45:34.358581Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2606.18954.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T20:45:34.358581Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8dff9ab2-4629-462e-8d72-07cc2c53e6e5 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8eb38e-11c7-4454-9b7e-5af16271b794 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Kimi K2: Open Agentic Intelligence
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10e64322-081a-4b21-bf6d-5ef9da81a23f · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Qwen3 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6186bd2b-0d40-4153-97eb-ad116e9d2736 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Troll: Trust regions improve reinforcement learning for large language models.The Fourteenth International Conference on Learning Representations, 2026
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5f35f13-6697-4efe-939c-53eda433a491 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Geometric-mean policy optimization.The Fourteenth International Conference on Learning Representations, 2026
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 186abeb2-b2bf-46bb-825e-8d1bd640faaa · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Dapo: An open-source llm reinforcement learning system at scale.The Thirty-Ninth Annual Conference on Neural Information Processing Systems, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4efbfb10-56ea-4be6-a2d3-6b62be733658 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Vineppo: Refining credit assignment in rl training of llms.Forty-Second International Conference on Machine Learning, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b02761-259f-438c-8a17-e2fafa6fa6d1 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Rewarding progress: Scaling automated process verifiers for llm reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c085b3-aa18-4734-b883-6e45e2fae371 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Treerpo: Tree relative policy optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation acb03cb7-9e4e-4464-b7f3-e144367ad852 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Pros: Towards compute-efficient rlvr via rollout prefix reuse
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f47b9ba8-f051-415c-9919-7d43b6af8005 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Let’s verify math questions step by step
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfa5b26d-4c54-4a3b-81e1-98cbd7147f95 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3f39c5-b15b-4fb6-845e-67da2186bd5a · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models The lessons of developing process reward models in mathematical reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da18bede-5249-4ce9-bc84-fccd1ff1cb3e · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Alphamath almost zero: process supervision without process.Advances in Neural Information Processing Systems, 37:27689– 27724, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 836e8882-9224-4902-a3fb-4fa7e9148f62 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Mutual rea- soning makes smaller llms stronger problem-solver
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35bb6f08-2990-436d-b05a-bd0046532499 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Process reinforcement through implicit rewards
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42f5a30-dc60-41d4-9404-c94ffdd0109b · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Group-in-group policy optimization for llm agent training
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecce0114-b2bd-4bf6-aed9-1ccbf7985667 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Treepo: Enhancing policy efficacy and inference efficiency with tree modeling.OpenReview preprint, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e2732ea-080b-42ba-8ceb-759f4f9aa2ca · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Tree search for llm agent reinforcement learning.The Fourteenth International Conference on Learning Representations, 2026
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ec75b2-e39e-4919-8cd1-3a3d255ab08e · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Scheduling your llm reinforcement learning with reasoning trees.The Fourteenth International Conference on Learning Representations, 2026
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6965f263-65d0-4ed9-af76-1ef22fb2016f · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Treerl: Llm reinforce- ment learning with on-policy tree search
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d013ac53-d5bf-458f-8b51-778ec161b81a · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Segment policy optimization: Effective segment-level credit assignment in rl for large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 162318d0-7d84-4a9d-89e0-b6d07c6a8a0b · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc01ea30-bac9-4079-835b-c5c2020e3223 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f79709e-c423-4c08-af2e-6aa2b519d306 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Trust, but verify: A self-verification approach to reinforcement learning with verifiable rewards
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a747e75-e1c5-44b2-985c-093e772b4535 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Self-aligned reward: Towards effective and efficient reasoners
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c355564d-ee87-4d03-a596-caff5e02d204 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Lookahead Tree- Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc422f2d-f099-4766-87dd-68e38d4eb3ee · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Monte carlo planning with large language model for text- based game agents
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9987b6af-ba42-4086-bacf-a3295a1599a1 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Tree-opo: Off-policy monte carlo tree- guided advantage optimization for multistep reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d471033c-c55d-4639-abba-a464500ab285 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Qwen2.5 Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e4cadbb-64c8-4ab5-befe-874b5049d2a1 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Sfr-embedding-2: Advanced text embedding with multi-stage training, 2024.URL https://huggingface
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4280ba47-1a47-4c75-aa8e-b92753d9cb03 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Mirb: Mathematical information retrieval benchmark
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5c81dfd-292c-4304-9cc5-b66c06108c54 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Measuring mathematical problem solving with the math dataset
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b8ef730-76a6-4a1e-b505-0d3b40019dc1 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ba5668f-0211-472a-8b22-70e6fb6b3e6a · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Rouge: A package for automatic evaluation of summaries
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c77f631-e089-48e8-b859-2dde477567a4 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation febc5756-6855-4390-8cba-ba02069345e2 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cebc2f51-77a2-4280-8743-a833f62dc8c8 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8428aa38-5a3a-461e-89d5-a23352a8bea3 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Gaia: a benchmark for general ai assistants
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5fa9e74-5bc6-4582-891c-a33e8f274bcb · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Webwalker: Benchmarking llms in web traversal
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7db7f2f3-3662-41d6-a3d9-2304da408e02 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec24d86c-2fba-4017-9059-2d1d5c438c71 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 746c2ad8-3d3c-4f91-84fb-eb4681e81b9c · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models ReAct: Synergizing Reasoning and Acting in Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation baa2e284-b63d-4b09-a156-b1d6afd451d4 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models HybridFlow: A Flexible and Efficient RLHF Framework
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 908ee632-e64d-4de9-a39c-f3d80b968738 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09f2649a-192c-4916-a237-73606cc87ee8 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1e84055-8f76-4c8f-b2cc-b92a9b9f7422 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d045f32-4ff9-4b3f-86cc-1072348a3461 · outbound
GraphPO: Graph-based Policy Optimization for Reasoning Models WebDancer: Towards Autonomous Information Seeking Agency
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.