Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T19:56:39.465133Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2605.00365.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T19:56:39.465133Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T21:54:44.345264Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-04T21:54:44.621660Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8c63f10a-6c93-4e41-b446-d0fd7b3d0cb2 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Nature , year=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed441f2d-4a7c-4f56-8876-1442f660c6dc · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2024 , eprint=
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c80f782c-49df-404c-ac90-0055ef923aae · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity arXiv preprint arXiv:2601.15609 , year=
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f08b654c-3f9a-47fa-a0a7-ad00e088a2ed · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Advances in neural information processing systems , volume=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c695e834-270d-4a59-82a6-38dbfe8e6e63 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2026 , eprint=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 537ecb71-4d46-4791-93e6-b26b58eb06c8 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Agentic Reinforced Policy Optimization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a14abe58-6121-4983-9cbf-3aae2077b979 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9ddb812-b7d7-4af8-838d-3c78e5830b14 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity arXiv preprint arXiv:2509.25133 , year=
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9c90e26-a085-44e3-9e8a-ea9752a63180 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2021 , eprint=
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24de10b9-817f-4028-82f3-d3a849216c59 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Diversity-incentivized exploration for versatile reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54e1a03b-6357-4945-85a8-13093bc99c7f · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84330745-ad0b-4b9d-ac2b-3be2bc8f9003 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e32dc371-2dc2-44a8-9eff-409e458b0ba4 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a7e5f71-1920-4c91-8646-0bd053605a74 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffe8d3ad-fd93-423f-9af3-dca52795b780 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Reasoning with Sampling: Your Base Model is Smarter Than You Think
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8622f2a9-dbf4-40ce-a373-990bef82e04f · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Enhancing efficiency and exploration in reinforcement learning for llms.arXiv preprint arXiv:2505.18573
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b171ce6b-071d-4c54-8e73-8c5ae6b9eccf · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04663399-ea7e-4ef9-b83e-8582e627ac35 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Notion Blog , year=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 984540f4-233e-4855-ab98-9c3de7490746 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity NeurIPS , year=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 282031cd-6c11-4d82-b87c-0e9cc8b09057 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2bb6b0f-2b72-42da-aa49-42e91c9e71fb · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb631d45-676c-41d6-adb8-2f2b66495092 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Bowman , booktitle=
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc507de8-a56e-4be9-9afa-962b35aa80f9 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Advances in Neural Information Processing Systems , volume=
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a1c6b22-a6c6-4cc4-bae0-0713e68beac9 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Qwen2.5: A Party of Foundation Models , url =
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd0037da-7843-47a4-9362-e2b42f68043d · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f9a438e-d767-480b-9e90-66e45d68f524 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 784e7443-325a-4bcf-9740-4902538e4ec0 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Differential smooth- ing mitigates sharpening and improves llm reasoning.arXiv preprint arXiv:2511.19942
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 820c6cfe-775d-41e3-8e00-a5fdce9977f2 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Transactions on Machine Learning Research , volume=
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de914d03-9ffa-4678-a00e-b646b81acce5 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Hugging Face repository , volume=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65d3827a-0af7-4dd5-8a12-8ecfc1a8951b · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0171a140-eb49-405c-851a-3a052f9101c8 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7081bfc4-1cd2-4b53-9ab6-7d92e432a443 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ebee529-2a35-42a0-acdf-3447cf0af6e0 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9ba9776-59af-4dd6-a264-345e83e61ca9 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Proximal Policy Optimization Algorithms
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c77f848-df2a-4b1f-84ae-2047ba73f338 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc8eb1c4-2005-4bcb-8c0a-d00995cf9619 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Outcome-based Exploration for LLM Reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b16acb1-5208-4a5e-b25c-d926a4b56189 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity ACL 2026 , year=
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3633c4fc-9d87-493b-af8a-864014a4dd39 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Advances in neural information processing systems , volume=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad0f8c17-b2b0-4238-9e09-7a9f03c753de · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1549946c-0990-42cc-9b04-83df2f79463d · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Does Reinforcement Learning Really Incentivize Reasoning Capacity in
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba8c1a5a-9146-4eaa-97a1-427fb0dae1fc · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity 2025 , eprint=
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 405b115a-8a38-42b3-add9-9e8eb14af53e · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Reasoning with Exploration: An Entropy Perspective
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 025d54a3-7039-44ad-a8b4-e51c809af7c7 · outbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 490492f5-7b91-4f81-b8db-d1e281bbd6c6 · inbound
Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.