Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T10:00:58.600743Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2605.12058.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T10:00:58.600743Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 43ad087b-a796-4fbd-bd7c-0bc027932762 · outbound
Holder Policy Optimisation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7696505d-910a-4feb-b9f7-e793877836a0 · outbound
Holder Policy Optimisation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da38957e-b774-450f-b086-fb3495993e88 · outbound
Holder Policy Optimisation Advances in neural information processing systems , volume=
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d6aa43b-2e71-4ae0-8ccc-c467fa2c8197 · outbound
Holder Policy Optimisation 2025 , eprint=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8eecf019-2b66-413e-a18e-11570fe6e97c · outbound
Holder Policy Optimisation Proximal Policy Optimization Algorithms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 85bf3601-d69f-4fa6-985a-7b05e12c1893 · outbound
Holder Policy Optimisation Understanding R1-Zero-Like Training: A Critical Perspective
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ebf9a6a4-d103-4eaa-bb8a-0caef8b179f6 · outbound
Holder Policy Optimisation Geometric-mean policy optimization
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 67d3a995-a91f-497c-ba8b-df8e075f6283 · outbound
Holder Policy Optimisation Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a4fc8100-a057-4936-a2b3-8b5f06026b57 · outbound
Holder Policy Optimisation Measuring Mathematical Problem Solving With the MATH Dataset
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a853cf97-ac1c-4dfd-b423-57e28360ab54 · outbound
Holder Policy Optimisation Advances in neural information processing systems , volume=
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dde29968-9a11-4325-a85c-b14a9e7a9a9b · outbound
Holder Policy Optimisation O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 78ec5504-d929-4418-b1a2-737a6c468d1d · outbound
Holder Policy Optimisation 2024 , publisher =
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 932c3309-ef12-4c71-892f-c26a2bb94164 · outbound
Holder Policy Optimisation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6a9db7fa-3258-46cc-8aa8-eaaf37d197de · outbound
Holder Policy Optimisation Machine learning , volume=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7f4aa4cc-8de8-4688-ab08-da870ce8a983 · outbound
Holder Policy Optimisation Advances in Neural Information Processing Systems (NIPS) , volume=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b5cb384f-eea6-4a83-b717-307099f0a248 · outbound
Holder Policy Optimisation Group Sequence Policy Optimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 378f4177-f5b9-47c0-9efc-92acda06d027 · outbound
Holder Policy Optimisation Group-in-Group Policy Optimization for LLM Agent Training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa40f318-7828-4858-bb10-15bf92e381c8 · outbound
Holder Policy Optimisation Advances in neural information processing systems , volume=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e61f42e1-8c5c-4827-8c7b-f6674608dfbe · outbound
Holder Policy Optimisation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5c3243e8-fcb1-4ec2-b813-bd874299d5a8 · outbound
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 43d32857-e4bc-489a-aefc-046b29bb3201 · outbound
Holder Policy Optimisation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1405764-6658-43c9-b323-c1f35785062c · outbound
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d96dd91-0bcd-49bf-b628-94cdaf424d23 · outbound
Holder Policy Optimisation arXiv preprint arXiv:2504.02546 , year=
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4fbe9a3d-06bb-4289-9c29-eb8d6bc4630b · outbound
Holder Policy Optimisation DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9f4b2d20-ede4-4577-9da3-8fbe75fe7436 · outbound
Holder Policy Optimisation AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9e69fa3-e40f-4cfb-8334-32dbe2cca0f6 · outbound
Holder Policy Optimisation BNPO: Beta Normalization Policy Optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2871403a-483e-406b-bdba-5763fb16b9f3 · outbound
Holder Policy Optimisation Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f00ade20-5fac-4ddc-94e6-b81b96dd10f9 · outbound
Holder Policy Optimisation Process Reinforcement through Implicit Rewards
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1394ac26-6e52-457a-a261-5452c7f99f99 · outbound
Holder Policy Optimisation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a2279bdc-0f69-4eab-a8e5-b328bef51afe · outbound
Holder Policy Optimisation Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0d5b02c-ba47-40fd-b114-e14189b52187 · outbound
Holder Policy Optimisation SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cdafa392-70f2-41ba-b055-d46aef9d1d2e · outbound
Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 86272908-79cd-42bc-8097-fc47a2786e69 · outbound
Holder Policy Optimisation SIAM review , volume=
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6a706d4a-8c50-420d-8d9e-eb83ec5f6d74 · outbound
Holder Policy Optimisation International conference on machine learning , pages=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 551248a5-5879-44a9-a38d-5a11abab795c · outbound
Holder Policy Optimisation 1976 , publisher=
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 139656f4-b693-4d5d-aba9-0fe108be2a5f · outbound
Holder Policy Optimisation arXiv preprint arXiv:2601.22521 , year=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9c65e001-2487-45a7-9e10-2bf270810a5d · outbound
Holder Policy Optimisation ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation adff6dcc-face-4c8f-8ede-3d4135922d51 · outbound
Holder Policy Optimisation arXiv preprint arXiv:2508.03772 , year=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a36adce7-c8bd-4d28-a25e-a4c7dc97f3c4 · outbound
Holder Policy Optimisation Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 05e24fbc-19e0-48c3-ba6a-0b47d5be5639 · outbound
Holder Policy Optimisation arXiv preprint arXiv:2506.08440 , year=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ea8c1185-787b-401f-98a2-5f755ff490cd · outbound
Holder Policy Optimisation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 11c3f6df-09db-45b7-a07d-6dbe8609d1d5 · outbound
Holder Policy Optimisation arXiv preprint arXiv:2510.03669 , year=
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 14660e6d-04b5-42bf-996f-eea9bc9e65f9 · outbound
Holder Policy Optimisation arXiv preprint arXiv:2510.09369 , year=
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 07d4c7bf-ee9e-416c-aa48-3605d010d34e · outbound
Holder Policy Optimisation Advances in Neural Information Processing Systems , volume=
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 68a22f9a-02cc-4f3f-9f4d-60c29e2d3a47 · outbound
Holder Policy Optimisation The Twelfth International Conference on Learning Representations , year=
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd9f59b6-256b-449c-9320-f14e2e7f3116 · outbound
Holder Policy Optimisation Does Reinforcement Learning Really Incentivize Reasoning Capacity in
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a24fb90-8044-4eab-b67e-2ba6e1484065 · outbound
Holder Policy Optimisation Rewarding the Unlikely: Lifting
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 909e2619-478b-49af-8258-f852037d735c · outbound
Holder Policy Optimisation Advances in Neural Information Processing Systems , volume=
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6253ef90-5c84-40ab-b456-be1062ce4bc0 · outbound
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cc13b490-a75e-4caf-a2a7-f983d863add8 · outbound
Holder Policy Optimisation Advances in Neural Information Processing Systems , volume=
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ef676581-d826-4d64-99cc-092eb718d330 · outbound
Holder Policy Optimisation arXiv preprint arXiv:2510.06870 , year=
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6d678c24-3e73-4b0a-95e6-a526d8b673f7 · outbound
Holder Policy Optimisation On-Policy RL with Optimal Reward Baseline
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94c188f2-641f-4ea3-8252-6faeb9c094f9 · outbound
Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3ee601fd-7c70-42c0-9605-704c0a0ecb64 · outbound
Holder Policy Optimisation 2018 , publisher=
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f3df1a4c-fa53-4f7f-8200-56b1edba6982 · outbound
Holder Policy Optimisation International conference on machine learning , pages=
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 726a0883-aeaf-4dd3-a01c-6345737dee68 · outbound
Holder Policy Optimisation Spurious Rewards: Rethinking Training Signals in RLVR
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 650509d9-5b85-4be6-a670-b921a6214281 · outbound
Holder Policy Optimisation 2018 , publisher=
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 29c6a257-8b60-4b5d-aa7d-3466eecbd7f3 · outbound
Holder Policy Optimisation Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e645e442-1a9a-48fa-8968-5adc9d756b3c · outbound
Holder Policy Optimisation Transformer Circuits Thread , year=
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.