Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T01:41:55.435903Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2605.11491.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T01:41:55.435903Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e2f35f59-3f5b-42a7-8ac4-4f3d2411b021 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f6d63a7-2c17-411a-9fc9-ac3db64bf02e · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8bfce8e4-612f-4640-a818-e1abc0c59a8f · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87a02fee-f3f3-4fe7-8db5-a601d0647c6e · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9909891d-4097-44cf-8064-9ff913c47680 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Reasoning with Exploration: An Entropy Perspective
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e40d96f9-c037-4971-9feb-e59ff445b2cb · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 409e0c23-fae1-4209-9552-8fb1b02c8fd8 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Bapo: Stabilizing off-policy reinforcement learning for llms via balanced policy optimization with adaptive clipping.arXiv preprint arXiv:2510.18927
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a26ff55e-ddb4-49a8-867c-f615d48eddf6 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Machine learning , volume=
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bc3c8fd7-f75e-4c7e-86eb-c4b302036285 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 03fd08bf-1307-4fe7-b355-87ef112d9050 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0bac2b1-d013-4666-839b-b59a9a17aed6 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization OpenAI o1 System Card
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d94312e7-493d-454c-9bd2-786486a3718b · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c291eec-5958-4556-ada1-4f94cc98cd5c · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Qwen3 Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0c2751b-e4fa-4cd0-ab9b-d6b9880ce0c3 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5d6be59-57a9-4e0e-a3ae-b0ce35276256 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Qwen2.5 Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d8666732-3695-4570-a77c-4331444b0776 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Skywork Open Reasoner 1 Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 593fc83e-68ac-48f9-8324-5ed8eeed4b69 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Measuring Mathematical Problem Solving With the MATH Dataset
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 12bdea9b-e524-41c4-adb4-279bc59039e3 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Hugging Face repository , volume=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 393d51c7-e30f-4bdd-ba3a-94ad9f71949c · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Proceedings of the Twentieth European Conference on Computer Systems , pages=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7cd5a674-0e18-416b-9796-9344f8ec0fcb · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Advances in neural information processing systems , volume=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 512754f8-c52a-47f8-a672-582603f64108 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization GPT-4 Technical Report
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b923da60-a2e4-41d2-836a-d81b5c193d53 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53e77c18-dc63-48ca-824c-5709335d6baf · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DeepSeek-V3 Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f8a7630c-e457-43e5-b772-208d9c711fbc · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 339f55b3-9a09-42a7-871f-07a1d9229b69 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Zhihu Zhuanlan , year=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 64135884-6e2f-49d2-9914-5d90975be742 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Proximal Policy Optimization Algorithms
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7703f9f8-8a80-44f9-af87-8c5dc26d385d · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Machine Learning , volume=
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d643fba9-3ac5-4661-bda5-ce0b9609a086 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Advances in neural information processing systems , volume=
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ab2c24a-6a69-42d1-81b1-7b2d204f26e5 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 138d3726-b1f2-4d58-9e97-e6ba8612f863 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization International conference on machine learning , pages=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 49ff38a0-1757-4a3c-a194-b628e658636f · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization International conference on machine learning , pages=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f63ea4b3-fee7-4416-9b64-cae9c7f20040 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DCPO: Dynamic Clipping Policy Optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7a042dc-bc83-46d0-8a15-13d0b8cf2450 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c2f054b-53b2-4ebc-84ea-7a283a39378f · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization 2024 , organization =
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f8bde296-48e9-448e-982e-23f602f12800 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b411e255-e80b-45f4-a5fc-aa0596bb71d1 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e115b3b2-c420-4da2-824f-a708c6261f90 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd7ad71b-9101-475c-a584-a2de4c4843d3 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2980fee8-513b-4ed6-9c10-0537c68478ed · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Decoupled Weight Decay Regularization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfaef6e5-7d2a-476b-9dd6-e4182d3eb683 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization 2017 , url=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb1413e4-7d89-463c-8777-db47a92cb050 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization On-Policy RL with Optimal Reward Baseline
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d38b3bcc-7b45-4ace-803e-c8c417aa7d82 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Prosperity before collapse: How far can off-policy rl reach with stale data on llms?
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f247a322-b1c6-4be2-9a99-b97a6a08b6c7 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization When Maximum Entropy Misleads Policy Optimization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79c5031c-f0ec-4fe3-9a85-60c213abefb0 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a14d97cd-74c9-44fa-afc8-c8ba359d7fc7 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Proceedings of the twelfth international conference on machine learning , pages=
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0137a0e8-448b-4674-8edf-de7e6228ecc3 · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Advances in neural information processing systems , volume=
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a008f9aa-afcd-449e-9171-86d2674feece · outbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization A Bayesian Perspective on Generalization and Stochastic Gradient Descent
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.