Pith. sign in

Paper Citation Record · LEDGER

Holder Policy Optimisation

As of 5 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2605.12058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12058 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T10:00:58.600743Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact13
  • verified fuzzy22
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch24

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43ad087b-a796-4fbd-bd7c-0bc027932762 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Holder Policy Optimisation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.172738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:ae9b7f35248973eb3dffcf48949327cfb13c7cfe28a5302a09c475e135bfea76

Observation 7696505d-910a-4feb-b9f7-e793877836a0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Holder Policy Optimisation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.006094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:d05cf89ee9b0611b1d73bccf06eda69217b3c5987c16d01308c8af0f0a43e569

Observation da38957e-b774-450f-b086-fb3495993e88 · outbound

This paper cites Advances in neural information processing systems , volume=.

Holder Policy Optimisation Advances in neural information processing systems , volume=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.094020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:6af77c743a53c31ef69b8c1b5024310601f8b3186c71881cb22f6d591263651b

Observation 7d6aa43b-2e71-4ae0-8ccc-c467fa2c8197 · outbound

This paper cites 2025 , eprint=.

Holder Policy Optimisation 2025 , eprint=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.098372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:ccb4b7683e5bd0e483fa8deed74346a2bc2ec14848d89e22cde8c1095ef73366

Observation 8eecf019-2b66-413e-a18e-11570fe6e97c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Holder Policy Optimisation Proximal Policy Optimization Algorithms

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.017387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:9e5b59b5b2c6e4e6aa694f7c803eae6f7bd7ada42d5cb1866fedfef050ed7086

Observation 85bf3601-d69f-4fa6-985a-7b05e12c1893 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Holder Policy Optimisation Understanding R1-Zero-Like Training: A Critical Perspective

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.027656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:c30c630712afc11ec1d4905be593575f7266b5edd729cf4382bcb4059608a9f7

Observation ebf9a6a4-d103-4eaa-bb8a-0caef8b179f6 · outbound

This paper cites Geometric-mean policy optimization.

Holder Policy Optimisation Geometric-mean policy optimization

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T10:01:23.033372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:86fd00f92f4053ff96e5198033013f736d87465f340fab9c79f1edad0a5428d3

Observation 67d3a995-a91f-497c-ba8b-df8e075f6283 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Holder Policy Optimisation Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.011593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:72c787b9d99c0627cfeab7aba3861d1d11b84e015dd617163e8d62bbdaa6c12f

Observation a4fc8100-a057-4936-a2b3-8b5f06026b57 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Holder Policy Optimisation Measuring Mathematical Problem Solving With the MATH Dataset

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.038187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:111bebb2193fa9d8e7e055b8d0ba338772c624ec8675f1a514375ac08d64c664

Observation a853cf97-ac1c-4dfd-b423-57e28360ab54 · outbound

This paper cites Advances in neural information processing systems , volume=.

Holder Policy Optimisation Advances in neural information processing systems , volume=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.086187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:dc7071be2a0d97e6d645bff0952ccea9205c72dbcc5a6ff2ed65bb47ff6c406a

Observation dde29968-9a11-4325-a85c-b14a9e7a9a9b · outbound

This paper cites O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Holder Policy Optimisation O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 11

Resolution
verified exact
doi, observed 2026-05-22T10:01:22.026817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:e574c3f41cabfe2730c98dd84a56013ebd60dafeb2856d77b3d176c6810bc94e

Observation 78ec5504-d929-4418-b1a2-737a6c468d1d · outbound

This paper cites 2024 , publisher =.

Holder Policy Optimisation 2024 , publisher =

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.089952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:804a4aa557d6a9a6274eb245fd42e86c197372a06fbca6153c0cf882afe52fb3

Observation 932c3309-ef12-4c71-892f-c26a2bb94164 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Holder Policy Optimisation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.136075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:56bc6295e1a2882a67e39f387e0ed7d11d80d356196c29b4cc642108b53ce75a

Observation 6a9db7fa-3258-46cc-8aa8-eaaf37d197de · outbound

This paper cites Machine learning , volume=.

Holder Policy Optimisation Machine learning , volume=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.081524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:1ab4b7c5f4a3fa6dafb7680521b1a292df9c7e082f91aa26b76793ff87ef76df

Observation 7f4aa4cc-8de8-4688-ab08-da870ce8a983 · outbound

This paper cites Advances in Neural Information Processing Systems (NIPS) , volume=.

Holder Policy Optimisation Advances in Neural Information Processing Systems (NIPS) , volume=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.077169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:33126cc684ab4ce3568f1f2d1b60bb0f618e933287bd37f5cc6ac21dc9582d23

Observation b5cb384f-eea6-4a83-b717-307099f0a248 · outbound

This paper cites Group Sequence Policy Optimization.

Holder Policy Optimisation Group Sequence Policy Optimization

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.121758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:994efc9953efeebc70a3e19f5f5a596029aa054eda343afcb8293dda44d25269

Observation 378f4177-f5b9-47c0-9efc-92acda06d027 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Holder Policy Optimisation Group-in-Group Policy Optimization for LLM Agent Training

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:22.975399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:34c3ce63ae1f9c189719cf17fa07cf382925b316c3278977b168bd292ed48aa3

Observation fa40f318-7828-4858-bb10-15bf92e381c8 · outbound

This paper cites Advances in neural information processing systems , volume=.

Holder Policy Optimisation Advances in neural information processing systems , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.072211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:01aad320bcc31f47dd9d551df3ed12d7b0920bec33c662d60b35e14178b5f404

Observation e61f42e1-8c5c-4827-8c7b-f6674608dfbe · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

Holder Policy Optimisation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.109055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:f0a965b8b0a7d2f4eda025346439cef626e501c4ca32580a1dd983fea0239278

Observation 5c3243e8-fcb1-4ec2-b813-bd874299d5a8 · outbound

This paper cites OpenAI o1 System Card.

Holder Policy Optimisation OpenAI o1 System Card

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.141829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:e3d4abc5c2082d4b896ec1c3d1bed266faf1fee067c16f3b17e44ca9d60d6675

Observation 43d32857-e4bc-489a-aefc-046b29bb3201 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Holder Policy Optimisation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:22.987983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:aad6ca0b18e5f29102c1a0036829e976dde94b2ba97aa3599d78a199ae294e95

Observation b1405764-6658-43c9-b323-c1f35785062c · outbound

This paper cites Qwen3 Technical Report.

Holder Policy Optimisation Qwen3 Technical Report

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.000838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:00922bd633fd85e4c9bb9f2012f0b097d4a6bcccfaf7d567e97b6519221a71b7

Observation 0d96dd91-0bcd-49bf-b628-94cdaf424d23 · outbound

This paper cites arXiv preprint arXiv:2504.02546 , year=.

Holder Policy Optimisation arXiv preprint arXiv:2504.02546 , year=

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.154216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:c32f7c3030c557072b5c88c0dc35ef33f98584007e766da829c6316caf9d2e08

Observation 4fbe9a3d-06bb-4289-9c29-eb8d6bc4630b · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Holder Policy Optimisation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.022545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:aaafa82d67413040ce65133f31f0bd83ddc4b4cea49c1d0f60dc990d3ce398e2

Observation 9f4b2d20-ede4-4577-9da3-8fbe75fe7436 · outbound

This paper cites AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin.

Holder Policy Optimisation AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.048828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:c85535ea1ae2e79b177e9ed4bdecd734287cca92e4a56b6811385f9ca71adf26

Observation b9e69fa3-e40f-4cfb-8334-32dbe2cca0f6 · outbound

This paper cites BNPO: Beta Normalization Policy Optimization.

Holder Policy Optimisation BNPO: Beta Normalization Policy Optimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.160661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:9117452d8edf918d1da4af7dd61d2df167b5a0b3c7b00e47029b70e4b12fc385

Observation 2871403a-483e-406b-bdba-5763fb16b9f3 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Holder Policy Optimisation Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.043308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:10bf5c5d1902381e3931f28449cc7709e7f7c604a6b87c093850a418a9bb3d95

Observation f00ade20-5fac-4ddc-94e6-b81b96dd10f9 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Holder Policy Optimisation Process Reinforcement through Implicit Rewards

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:22.981624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:b91310ae50f44c1bfc331fb7b82df1081aaf92ed7264bb15abda903291f905bc

Observation 1394ac26-6e52-457a-a261-5452c7f99f99 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Holder Policy Optimisation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T10:01:22.994765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:2b85d5b6860c3dd596ff265acb61127a803fcdf57a9c40848af52d299bb59ea5

Observation a2279bdc-0f69-4eab-a8e5-b328bef51afe · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Holder Policy Optimisation Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.068373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:9f7c7de49164a3582c04cc540682fa2f2a53159e1f823b17141a0570bafc2bb1

Observation d0d5b02c-ba47-40fd-b114-e14189b52187 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Holder Policy Optimisation SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.061118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:8bef6c649ba13617140b74c6c7cb6d88fd54c4254e2f7838d86af65c964df66a

Observation cdafa392-70f2-41ba-b055-d46aef9d1d2e · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.055182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:48339fab78e06e708c05735a32f117bb90d3eda56241463129d2c0b73240225d

Observation 86272908-79cd-42bc-8097-fc47a2786e69 · outbound

This paper cites SIAM review , volume=.

Holder Policy Optimisation SIAM review , volume=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.064561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:d6b1471246c0ead7b7833691de74278a758a9754deb7faf4effa68f0f7087ea9

Observation 6a706d4a-8c50-420d-8d9e-eb83ec5f6d74 · outbound

This paper cites International conference on machine learning , pages=.

Holder Policy Optimisation International conference on machine learning , pages=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.057269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:b638f378ce33b158c012269f9e7d32887268ab623d94cf614ca1e3f9d65fa715

Observation 551248a5-5879-44a9-a38d-5a11abab795c · outbound

This paper cites 1976 , publisher=.

Holder Policy Optimisation 1976 , publisher=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.060875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:932673a5b70d9d284dab0914762815aab966802a5411a256ad5ef2a6969bd468

Observation 139656f4-b693-4d5d-aba9-0fe108be2a5f · outbound

This paper cites arXiv preprint arXiv:2601.22521 , year=.

Holder Policy Optimisation arXiv preprint arXiv:2601.22521 , year=

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:22.964122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:c2d20499c44ce23248ef62468ed5c3edd3054027d3ea02fcadcf327423efc444

Observation 9c65e001-2487-45a7-9e10-2bf270810a5d · outbound

This paper cites ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models.

Holder Policy Optimisation ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:22.969538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:59c3df69ae1a64cfda9953ea8c79131a9de7e89b32f627b9bcd1963f87dfd75f

Observation adff6dcc-face-4c8f-8ede-3d4135922d51 · outbound

This paper cites arXiv preprint arXiv:2508.03772 , year=.

Holder Policy Optimisation arXiv preprint arXiv:2508.03772 , year=

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.067064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:5f96724c87d5e4bd1e76e5f96c58ee8e3e1293debe4d22a55c99644c387bc9d5

Observation a36adce7-c8bd-4d28-a25e-a4c7dc97f3c4 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Holder Policy Optimisation Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.166511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:68fad7406f89ab6327ad20df5096374767ed4c4f208258ba94d7fb4b0da69dac

Observation 05e24fbc-19e0-48c3-ba6a-0b47d5be5639 · outbound

This paper cites arXiv preprint arXiv:2506.08440 , year=.

Holder Policy Optimisation arXiv preprint arXiv:2506.08440 , year=

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.148052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:5cbc2941a7ff6a0075aab7ea37783061de0794f6a111c25ec95b9e2d6ee106f7

Observation ea8c1185-787b-401f-98a2-5f755ff490cd · outbound

This paper cites Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs.

Holder Policy Optimisation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:22.958441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:8ac8fb11a6b1271be4804c2249353d84bac578a49f2b6aeb72ef7862ae0bae7a

Observation 11c3f6df-09db-45b7-a07d-6dbe8609d1d5 · outbound

This paper cites arXiv preprint arXiv:2510.03669 , year=.

Holder Policy Optimisation arXiv preprint arXiv:2510.03669 , year=

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.091395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:8b5c4f7a68a3507d57ee9515eb3749d110d7ebcf1f02becd9ba798e1bb6d115b

Observation 14660e6d-04b5-42bf-996f-eea9bc9e65f9 · outbound

This paper cites arXiv preprint arXiv:2510.09369 , year=.

Holder Policy Optimisation arXiv preprint arXiv:2510.09369 , year=

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.097593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:215aa8265c770335e0538c2c905197a5b3b7cb317b8d67ee0bf19d0dba93f63a

Observation 07d4c7bf-ee9e-416c-aa48-3605d010d34e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Holder Policy Optimisation Advances in Neural Information Processing Systems , volume=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.044510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:7fabdc85f08dd10f5c2ad9e38c48b8258874f4e593be9ba9218277de257b94b0

Observation 68a22f9a-02cc-4f3f-9f4d-60c29e2d3a47 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Holder Policy Optimisation The Twelfth International Conference on Learning Representations , year=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.049376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:6d06ee5cc2618d958a2b3db65242706aec898594a04305c42689c859c6b89822

Observation fd9f59b6-256b-449c-9320-f14e2e7f3116 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in.

Holder Policy Optimisation Does Reinforcement Learning Really Incentivize Reasoning Capacity in

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.053655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:d13c8fbcdeca8ce279b426fa5a54a264c4d1fdb3da8a6a076a946a020d7d3775

Observation 7a24fb90-8044-4eab-b67e-2ba6e1484065 · outbound

This paper cites Rewarding the Unlikely: Lifting.

Holder Policy Optimisation Rewarding the Unlikely: Lifting

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.031924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:6062d4cb456d0e1ecd722b4f47f4b7ff1d739363427767a5d4eec285e9843853

Observation 909e2619-478b-49af-8258-f852037d735c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Holder Policy Optimisation Advances in Neural Information Processing Systems , volume=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.102710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:834795b0663685e1a09c4c53f681b9bdf1e42781b86b7bb38baad25cf1fc2e42

Observation 6253ef90-5c84-40ab-b456-be1062ce4bc0 · outbound

This paper cites Let's Verify Step by Step.

Holder Policy Optimisation Let's Verify Step by Step

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.102975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:5cc42a374376a2f0d84c22e09b954b15de9fe41c896cb2085fd35ab9afcd40eb

Observation cc13b490-a75e-4caf-a2a7-f983d863add8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Holder Policy Optimisation Advances in Neural Information Processing Systems , volume=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.036302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:705d66a1f994855f36d72bbb6a3f4988c9118edfe0964ce18ee95b0f39767caa

Observation ef676581-d826-4d64-99cc-092eb718d330 · outbound

This paper cites arXiv preprint arXiv:2510.06870 , year=.

Holder Policy Optimisation arXiv preprint arXiv:2510.06870 , year=

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:22.946551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:71191391597f465b4e203be362924af6c967ed207013e6f0a7fefdba277a64b7

Observation 6d678c24-3e73-4b0a-95e6-a526d8b673f7 · outbound

This paper cites On-Policy RL with Optimal Reward Baseline.

Holder Policy Optimisation On-Policy RL with Optimal Reward Baseline

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.127974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:3b95a32fdbdf5545bf338a2d65f1e69f9dfeb559999f9515601bf035aa310761

Observation 94c188f2-641f-4ea3-8252-6faeb9c094f9 · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.115200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:a31224b6ef16adf3151c216fca729816e98895f420c7decaf01312b60b88eed4

Observation 3ee601fd-7c70-42c0-9605-704c0a0ecb64 · outbound

This paper cites 2018 , publisher=.

Holder Policy Optimisation 2018 , publisher=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.027694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:1e61ddeca259f314cfe575f2bc52a2cebdfc33d3df8c4624f31f578415f94570

Observation f3df1a4c-fa53-4f7f-8200-56b1edba6982 · outbound

This paper cites International conference on machine learning , pages=.

Holder Policy Optimisation International conference on machine learning , pages=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.040599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:4bffaa34149f0bc48b9348897f6f5f8470bbdcd3383c82aeccb6e23f46691754

Observation 726a0883-aeaf-4dd3-a01c-6345737dee68 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Holder Policy Optimisation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:22.952772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:55992a072e603da102bbd316ace7a88fc12a0110ee6f70c9cfa411157708529a

Observation 650509d9-5b85-4be6-a670-b921a6214281 · outbound

This paper cites 2018 , publisher=.

Holder Policy Optimisation 2018 , publisher=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.015496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:ce55969148118486efd82b81f454504a329cccd11155a509c565f9537808ec49

Observation 29c6a257-8b60-4b5d-aa7d-3466eecbd7f3 · outbound

This paper cites Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=.

Holder Policy Optimisation Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.019501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:e582d12f920b83b381ce94197d4c2ff488d12c99e7da45a1fd42de562008cd59

Observation e645e442-1a9a-48fa-8968-5adc9d756b3c · outbound

This paper cites Transformer Circuits Thread , year=.

Holder Policy Optimisation Transformer Circuits Thread , year=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:01:24.023424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:1c8b13731cc4f4142bcdd7b14341c2c967605fa07f1b18cd7fcfac4e939787c1

Pith citing papers

No inbound Pith citation observations are available.