Pith. sign in

Paper Citation Record · LEDGER

GraphPO: Graph-based Policy Optimization for Reasoning Models

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2606.18954.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.18954 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T20:45:34.358581Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact13
  • verified fuzzy0
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8dff9ab2-4629-462e-8d72-07cc2c53e6e5 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638, 2025.

GraphPO: Graph-based Policy Optimization for Reasoning Models Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:ce6a496fa9a7560a1e11c601f169601578a5de5ce3c30306967541495cb3a087

Observation 5c8eb38e-11c7-4454-9b7e-5af16271b794 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

GraphPO: Graph-based Policy Optimization for Reasoning Models Kimi K2: Open Agentic Intelligence

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.131847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:a87c63a68948c65d2ed093e7d35fd16fe32d9c29a2ebe453f3b8d24d356e28cb

Observation 10e64322-081a-4b21-bf6d-5ef9da81a23f · outbound

This paper cites Qwen3 Technical Report.

GraphPO: Graph-based Policy Optimization for Reasoning Models Qwen3 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.136079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:ceb206859dbd66cbde40b96388e8d818c4b80c83aa54c5466752affd1ac37dc5

Observation 6186bd2b-0d40-4153-97eb-ad116e9d2736 · outbound

This paper cites Troll: Trust regions improve reinforcement learning for large language models.The Fourteenth International Conference on Learning Representations, 2026.

GraphPO: Graph-based Policy Optimization for Reasoning Models Troll: Trust regions improve reinforcement learning for large language models.The Fourteenth International Conference on Learning Representations, 2026

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:68752e7287ee448f3be6a13eb19d57d246163637528abf5b652fba5e237f9405

Observation d5f35f13-6697-4efe-939c-53eda433a491 · outbound

This paper cites Geometric-mean policy optimization.The Fourteenth International Conference on Learning Representations, 2026.

GraphPO: Graph-based Policy Optimization for Reasoning Models Geometric-mean policy optimization.The Fourteenth International Conference on Learning Representations, 2026

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:744db3cdef38b85f1c40dbef470dae17e949becb4e059efbac3e87d21e9f4d93

Observation 186abeb2-b2bf-46bb-825e-8d1bd640faaa · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale.The Thirty-Ninth Annual Conference on Neural Information Processing Systems, 2025.

GraphPO: Graph-based Policy Optimization for Reasoning Models Dapo: An open-source llm reinforcement learning system at scale.The Thirty-Ninth Annual Conference on Neural Information Processing Systems, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:dcb71f760699732797b0912ffac9760ee743cd36857e61690ce3fd0905c5aaab

Observation 4efbfb10-56ea-4be6-a2d3-6b62be733658 · outbound

This paper cites Vineppo: Refining credit assignment in rl training of llms.Forty-Second International Conference on Machine Learning, 2025.

GraphPO: Graph-based Policy Optimization for Reasoning Models Vineppo: Refining credit assignment in rl training of llms.Forty-Second International Conference on Machine Learning, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:1382dfb7420562c1d6d2f8ddd5dfe36c53d6751b93d05d17de35f1af7b19ec45

Observation 76b02761-259f-438c-8a17-e2fafa6fa6d1 · outbound

This paper cites Rewarding progress: Scaling automated process verifiers for llm reasoning.

GraphPO: Graph-based Policy Optimization for Reasoning Models Rewarding progress: Scaling automated process verifiers for llm reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:4ece329235515c1bedce51a147d453afaf29914f9c9e7f8da6b2c0b58ba0020e

Observation f5c085b3-aa18-4734-b883-6e45e2fae371 · outbound

This paper cites Treerpo: Tree relative policy optimization.

GraphPO: Graph-based Policy Optimization for Reasoning Models Treerpo: Tree relative policy optimization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:21.125901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:320516e00d8e31a21723068225aa8df0c7270da261d9d82fee2ad3206048d9c1

Observation acb03cb7-9e4e-4464-b7f3-e144367ad852 · outbound

This paper cites Pros: Towards compute-efficient rlvr via rollout prefix reuse.

GraphPO: Graph-based Policy Optimization for Reasoning Models Pros: Towards compute-efficient rlvr via rollout prefix reuse

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:06e55d883f60ff185ed8d79f8212d79ea85aae67f3fab68bba7708ae341cb49c

Observation f47b9ba8-f051-415c-9919-7d43b6af8005 · outbound

This paper cites Let’s verify math questions step by step.

GraphPO: Graph-based Policy Optimization for Reasoning Models Let’s verify math questions step by step

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:5f06ba331ed7ad82d10e0d9cc90050294c3d61672a4976d3ae095e7caa7d8b9e

Observation bfa5b26d-4c54-4a3b-81e1-98cbd7147f95 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

GraphPO: Graph-based Policy Optimization for Reasoning Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:de12b0494663c6ec2b2725f2ecd2b41ad2a8029ec9985c878224d48830374ce6

Observation ee3f39c5-b15b-4fb6-845e-67da2186bd5a · outbound

This paper cites The lessons of developing process reward models in mathematical reasoning.

GraphPO: Graph-based Policy Optimization for Reasoning Models The lessons of developing process reward models in mathematical reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:953ec274fdada314b1caf08860a1800da676f041ea40932c65a63056dbfe132b

Observation da18bede-5249-4ce9-bc84-fccd1ff1cb3e · outbound

This paper cites Alphamath almost zero: process supervision without process.Advances in Neural Information Processing Systems, 37:27689– 27724, 2024.

GraphPO: Graph-based Policy Optimization for Reasoning Models Alphamath almost zero: process supervision without process.Advances in Neural Information Processing Systems, 37:27689– 27724, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:e2c0fac8196a8f8a579a04dfd95a253e571479fbf115099a248ab442f07ac96c

Observation 836e8882-9224-4902-a3fb-4fa7e9148f62 · outbound

This paper cites Mutual rea- soning makes smaller llms stronger problem-solver.

GraphPO: Graph-based Policy Optimization for Reasoning Models Mutual rea- soning makes smaller llms stronger problem-solver

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:d72bd522810cdf0598c6cef69a4796b206eec2f8ea013cd549cd54ab3991e3f2

Observation 35bb6f08-2990-436d-b05a-bd0046532499 · outbound

This paper cites Process reinforcement through implicit rewards.

GraphPO: Graph-based Policy Optimization for Reasoning Models Process reinforcement through implicit rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:666acfdc902a98ec4bd0c6325c1cdd28676c6c6a9f4378dd6a41f3f0efed8378

Observation c42f5a30-dc60-41d4-9404-c94ffdd0109b · outbound

This paper cites Group-in-group policy optimization for llm agent training.

GraphPO: Graph-based Policy Optimization for Reasoning Models Group-in-group policy optimization for llm agent training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:d3c8ff7ad015a7fae0011e763e230a0d1e1056e9ee5413027db6931cfc5f2a9c

Observation ecce0114-b2bd-4bf6-aed9-1ccbf7985667 · outbound

This paper cites Treepo: Enhancing policy efficacy and inference efficiency with tree modeling.OpenReview preprint, 2025.

GraphPO: Graph-based Policy Optimization for Reasoning Models Treepo: Enhancing policy efficacy and inference efficiency with tree modeling.OpenReview preprint, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:48423c3753b6c13588a08c148d561ca58adf2e25a50d23fd713763df1de065dd

Observation 9e2732ea-080b-42ba-8ceb-759f4f9aa2ca · outbound

This paper cites Tree search for llm agent reinforcement learning.The Fourteenth International Conference on Learning Representations, 2026.

GraphPO: Graph-based Policy Optimization for Reasoning Models Tree search for llm agent reinforcement learning.The Fourteenth International Conference on Learning Representations, 2026

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:f105d7942a355bece1cbe2535cbe69ab527acce726ecafce0a6346a4ecd55e05

Observation 56ec75b2-e39e-4919-8cd1-3a3d255ab08e · outbound

This paper cites Scheduling your llm reinforcement learning with reasoning trees.The Fourteenth International Conference on Learning Representations, 2026.

GraphPO: Graph-based Policy Optimization for Reasoning Models Scheduling your llm reinforcement learning with reasoning trees.The Fourteenth International Conference on Learning Representations, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:62b3657ca7eb754edf48b7de4ec14626ca53a1c3b240d847114211334d07f54f

Observation 6965f263-65d0-4ed9-af76-1ef22fb2016f · outbound

This paper cites Treerl: Llm reinforce- ment learning with on-policy tree search.

GraphPO: Graph-based Policy Optimization for Reasoning Models Treerl: Llm reinforce- ment learning with on-policy tree search

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:26f74dd0c979a32e2b17fee99d95a715fae5bd751cbc5f4a669591b4c0e6b8ac

Observation d013ac53-d5bf-458f-8b51-778ec161b81a · outbound

This paper cites Segment policy optimization: Effective segment-level credit assignment in rl for large language models.

GraphPO: Graph-based Policy Optimization for Reasoning Models Segment policy optimization: Effective segment-level credit assignment in rl for large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:25e705f1411f48cef92017c05ad64f5f9eb15b38c16a4f82ff754001d7dcbb52

Observation 162318d0-7d84-4a9d-89e0-b6d07c6a8a0b · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

GraphPO: Graph-based Policy Optimization for Reasoning Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:7e1308f00a6d8400dd1e41cd7e759aeb54fb0bd131a726265da02197f2cc59d1

Observation bc01ea30-bac9-4079-835b-c5c2020e3223 · outbound

This paper cites Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms.

GraphPO: Graph-based Policy Optimization for Reasoning Models Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:e6a2abc954eeb6082ed125e1e3ecf0b3735c11e6b41d0c2cc67558654216317c

Observation 8f79709e-c423-4c08-af2e-6aa2b519d306 · outbound

This paper cites Trust, but verify: A self-verification approach to reinforcement learning with verifiable rewards.

GraphPO: Graph-based Policy Optimization for Reasoning Models Trust, but verify: A self-verification approach to reinforcement learning with verifiable rewards

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:84d33dee542984eb15708bf6b12e7bbc8ef24658504893dbb904b0fc09b27215

Observation 2a747e75-e1c5-44b2-985c-093e772b4535 · outbound

This paper cites Self-aligned reward: Towards effective and efficient reasoners.

GraphPO: Graph-based Policy Optimization for Reasoning Models Self-aligned reward: Towards effective and efficient reasoners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:7f188fb34dddfb23bd1ffb4a4a19fc2c51f3cecaff70507074dab3075cbf194a

Observation c355564d-ee87-4d03-a596-caff5e02d204 · outbound

This paper cites Lookahead Tree- Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards.

GraphPO: Graph-based Policy Optimization for Reasoning Models Lookahead Tree- Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:17a7c325ac5136a8d6549230471ac8c0409907e0e79263a24976fa14703fe91f

Observation cc422f2d-f099-4766-87dd-68e38d4eb3ee · outbound

This paper cites Monte carlo planning with large language model for text- based game agents.

GraphPO: Graph-based Policy Optimization for Reasoning Models Monte carlo planning with large language model for text- based game agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:fdfcca30278c003de7451d664d926333dd35c898e157c5189db8fe781322ac1b

Observation 9987b6af-ba42-4086-bacf-a3295a1599a1 · outbound

This paper cites Tree-opo: Off-policy monte carlo tree- guided advantage optimization for multistep reasoning.

GraphPO: Graph-based Policy Optimization for Reasoning Models Tree-opo: Off-policy monte carlo tree- guided advantage optimization for multistep reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:ab0e2a52f03879068c8b378f09e4c8e6a14e4beeb9a60bab5fbcc19c1976e2d8

Observation d471033c-c55d-4639-abba-a464500ab285 · outbound

This paper cites Qwen2.5 Technical Report.

GraphPO: Graph-based Policy Optimization for Reasoning Models Qwen2.5 Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.117938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:f9b9a8d3cbe52cbfd2771e06f0dda54ac072149e465ee1461363112040836647

Observation 5e4cadbb-64c8-4ab5-befe-874b5049d2a1 · outbound

This paper cites Sfr-embedding-2: Advanced text embedding with multi-stage training, 2024.URL https://huggingface.

GraphPO: Graph-based Policy Optimization for Reasoning Models Sfr-embedding-2: Advanced text embedding with multi-stage training, 2024.URL https://huggingface

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:5ff101ef411dd4b2203d3f0be325aad3a3b131c72fa6096e54f528f099ee4c6c

Observation 4280ba47-1a47-4c75-aa8e-b92753d9cb03 · outbound

This paper cites Mirb: Mathematical information retrieval benchmark.

GraphPO: Graph-based Policy Optimization for Reasoning Models Mirb: Mathematical information retrieval benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:50d2846862d37d0ffb62d3ef091db4a7d37ac7b8e26dea8732a57851f886c949

Observation d5c81dfd-292c-4304-9cc5-b66c06108c54 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

GraphPO: Graph-based Policy Optimization for Reasoning Models Measuring mathematical problem solving with the math dataset

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:0452320338f71bd4db23388a52f546b062338e7c464f2d8bd87c40365b49e575

Observation 6b8ef730-76a6-4a1e-b505-0d3b40019dc1 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

GraphPO: Graph-based Policy Optimization for Reasoning Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.109577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:6b0c16e9026315e7a556b62b35611e025fdda991b4916ae54fc4ee9380d0cda9

Observation 0ba5668f-0211-472a-8b22-70e6fb6b3e6a · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

GraphPO: Graph-based Policy Optimization for Reasoning Models Rouge: A package for automatic evaluation of summaries

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:a0c0a8e94d371266f697206698ba88719e34116d54bd02351affbb4a2619cff1

Observation 6c77f631-e089-48e8-b859-2dde477567a4 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

GraphPO: Graph-based Policy Optimization for Reasoning Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.096630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:c77a3f2e4761b4d705d6d7c8124118ea810c7482f5abdc173dd5dcb51d103dfa

Observation febc5756-6855-4390-8cba-ba02069345e2 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

GraphPO: Graph-based Policy Optimization for Reasoning Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.091999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:0e90a3a7887a9171f00df62b175b3bf342dee04a06cf17748b8ca6b34f645268

Observation cebc2f51-77a2-4280-8743-a833f62dc8c8 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

GraphPO: Graph-based Policy Optimization for Reasoning Models Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:18d4a969acd3ad695979d00891af0365a4bd17b267827fd2736625ad539ab1d4

Observation 8428aa38-5a3a-461e-89d5-a23352a8bea3 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

GraphPO: Graph-based Policy Optimization for Reasoning Models Gaia: a benchmark for general ai assistants

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:110f9bdfb93aa72c952d8634e9d7922f08d990f98c35e14ef8cd665d1d06c4a5

Observation b5fa9e74-5bc6-4582-891c-a33e8f274bcb · outbound

This paper cites Webwalker: Benchmarking llms in web traversal.

GraphPO: Graph-based Policy Optimization for Reasoning Models Webwalker: Benchmarking llms in web traversal

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:03360bde14a2afb624a8e76237b12a54fac55e3d214cf731a02d9f120ecd5cf4

Observation 7db7f2f3-3662-41d6-a3d9-2304da408e02 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

GraphPO: Graph-based Policy Optimization for Reasoning Models BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.100303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:8abc9a21018778231c86717138fdf038f3eaf39863f24e10d11ba273a42b2651

Observation ec24d86c-2fba-4017-9059-2d1d5c438c71 · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

GraphPO: Graph-based Policy Optimization for Reasoning Models xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:21.105637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:e1f071fdb5455c1113aa6917f1917d7511c2876e798d1da7ea4d13b6e31b4c8d

Observation 746c2ad8-3d3c-4f91-84fb-eb4681e81b9c · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

GraphPO: Graph-based Policy Optimization for Reasoning Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.113786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:1adb04cd078d984576aedb2d72f326e1773c08be42c0e739c3dabd6d8ec1acb5

Observation baa2e284-b63d-4b09-a156-b1d6afd451d4 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

GraphPO: Graph-based Policy Optimization for Reasoning Models HybridFlow: A Flexible and Efficient RLHF Framework

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.121536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:1f5bbb7e8f036f3635cc6e037d3259e3a49db93a76278656b3a48653b0e5747f

Observation 908ee632-e64d-4de9-a39c-f3d80b968738 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

GraphPO: Graph-based Policy Optimization for Reasoning Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T20:45:34.358581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:adeb0db9355facd1240a1418ad62fe3ac4f2f69c8f5fed883a95be04b65f3255

Observation 09f2649a-192c-4916-a237-73606cc87ee8 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

GraphPO: Graph-based Policy Optimization for Reasoning Models Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:21.078049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:e9e6e8d02af7af1618b82854e743fc31f8d6ab9f2f8c2d55690a4897e56e9e40

Observation e1e84055-8f76-4c8f-b2cc-b92a9b9f7422 · outbound

This paper cites Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl.

GraphPO: Graph-based Policy Optimization for Reasoning Models Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:21.082677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:b8b2f4767b18f917c1c4ecc1f9a99fc06dc35dbd0008b414b781b4ff97acd333

Observation 4d045f32-4ff9-4b3f-86cc-1072348a3461 · outbound

This paper cites WebDancer: Towards Autonomous Information Seeking Agency.

GraphPO: Graph-based Policy Optimization for Reasoning Models WebDancer: Towards Autonomous Information Seeking Agency

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:59:21.087774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:45:34.358581Z digest=sha256:1ae2a6b7466654ef635e1a81260de505a9d52a7c6df824035ed84789283e31ee

Pith citing papers

No inbound Pith citation observations are available.