Pith. sign in

Paper Citation Record · LEDGER

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization

As of 5 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2605.11491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.11491 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:41:55.435903Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact5
  • verified fuzzy14
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch28

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2f35f59-3f5b-42a7-8ac4-4f3d2411b021 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.811425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:5d89cd36931433b57f25f4b7dd1328fcef2355d17ec9399a9ebf58e2427c8a96

Observation 6f6d63a7-2c17-411a-9fc9-ac3db64bf02e · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.800543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:dc3851cfd1c4e6297fa9c9ebda1859c189848b204895f3cce3e02472a8c8c363

Observation 8bfce8e4-612f-4640-a818-e1abc0c59a8f · outbound

This paper cites Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.795189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:3453fc60cd389f71b5f1b039b5b747f9ee71a2535504f74b42258f7f289c1840

Observation 87a02fee-f3f3-4fe7-8db5-a601d0647c6e · outbound

This paper cites SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:05:34.479506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:03d121e4b73629797b0ecc18d7ee7b38b2c7a95e7e4263a63407f00b8bdb21c2

Observation 9909891d-4097-44cf-8064-9ff913c47680 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Reasoning with Exploration: An Entropy Perspective

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:20:56.524904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:7f3d03b371c5c8a75e7f425e115647ee2bce2554e266ff3813dfd2728b955417

Observation e40d96f9-c037-4971-9feb-e59ff445b2cb · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.774087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:a8295eedf7021c6c6b06ba256830a8fa9acd98af91617179b6d11bd3d1e947ab

Observation 409e0c23-fae1-4209-9552-8fb1b02c8fd8 · outbound

This paper cites Bapo: Stabilizing off-policy reinforcement learning for llms via balanced policy optimization with adaptive clipping.arXiv preprint arXiv:2510.18927.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Bapo: Stabilizing off-policy reinforcement learning for llms via balanced policy optimization with adaptive clipping.arXiv preprint arXiv:2510.18927

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:02.790418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:59a7a2d692a82ad505bf2357c9cb66ffa2305a1fba5af28a739bc78e319f9f68

Observation a26ff55e-ddb4-49a8-867c-f615d48eddf6 · outbound

This paper cites Machine learning , volume=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Machine learning , volume=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.455607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:59c40cedfc09c899f3e41c9d22ec077d79f88ff5b99539db8dded76d6f85e19a

Observation bc3c8fd7-f75e-4c7e-86eb-c4b302036285 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.816167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:27ae6195941847cfb3ff7c7d364233517dd93294510513214db749b55d9039dd

Observation 03fd08bf-1307-4fe7-b355-87ef112d9050 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.820384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:b463c3f9a37a4c926c7357e5dbca019f1c9fa2aa2edb14a48e872e6031bca15e

Observation b0bac2b1-d013-4666-839b-b59a9a17aed6 · outbound

This paper cites OpenAI o1 System Card.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization OpenAI o1 System Card

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.825386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:503d04cdb0460276cb1e2fa545c855cdb7210348ffda2a27b3d0f7098c49a630

Observation d94312e7-493d-454c-9bd2-786486a3718b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.836357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:e884e04ed151cf0200e88bd44ff0e46168e1993b0da9811f3fbb21c5cc942fae

Observation 7c291eec-5958-4556-ada1-4f94cc98cd5c · outbound

This paper cites Qwen3 Technical Report.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Qwen3 Technical Report

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.841611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:caf3cf911e17fe01ed0bd0ffb8b5283b7ff0596796e3dc38721995666a6f586f

Observation a0c2751b-e4fa-4cd0-ab9b-d6b9880ce0c3 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:42:02.779591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:4c529c2ffee303c09ad3a3f687327079f3a4bd4cdff734c392a1b28a1e215470

Observation a5d6be59-57a9-4e0e-a3ae-b0ce35276256 · outbound

This paper cites Qwen2.5 Technical Report.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Qwen2.5 Technical Report

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.806304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:71cf1335fe5d475b1b05e088dec3c7eaaa93eb60c4077358aa49a6e2803fab50

Observation d8666732-3695-4570-a77c-4331444b0776 · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Skywork Open Reasoner 1 Technical Report

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:26:47.432522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:7d6bde985cd859559cc962468a29011db67d7059378855e30e702130bfa1bc59

Observation 593fc83e-68ac-48f9-8324-5ed8eeed4b69 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.878781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:014effd684abe2d103cc38a5f2c8f87bbad0fa3cbd1c0808f70ddcecc105be7f

Observation 12bdea9b-e524-41c4-adb4-279bc59039e3 · outbound

This paper cites Hugging Face repository , volume=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Hugging Face repository , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.444244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:6c243b72230a68d19ed7cc75651ff8988da7071b1824eedd5f0a7a8b0a5727dc

Observation 393d51c7-e30f-4bdd-ba3a-94ad9f71949c · outbound

This paper cites Proceedings of the Twentieth European Conference on Computer Systems , pages=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Proceedings of the Twentieth European Conference on Computer Systems , pages=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.447795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:a424c9df7f9329afeaed0fc5a8d297030a2506511bc834be354ee76248fd154a

Observation 7cd5a674-0e18-416b-9796-9344f8ec0fcb · outbound

This paper cites Advances in neural information processing systems , volume=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Advances in neural information processing systems , volume=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.451592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:840f2954bdfb3e6d34372a859a4745ac5ed7464be2e41cc118d947f0de6caa7c

Observation 512754f8-c52a-47f8-a672-582603f64108 · outbound

This paper cites GPT-4 Technical Report.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization GPT-4 Technical Report

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.873990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:b38e8cf7920b210d4441d2e959566e8362ca69df0bfb688846a5455a48481a1b

Observation b923da60-a2e4-41d2-836a-d81b5c193d53 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.883491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:5b6155353771a3ce166082d1ab4b5990868ec9520b4790a2f1316cc2e0e01152

Observation 53e77c18-dc63-48ca-824c-5709335d6baf · outbound

This paper cites DeepSeek-V3 Technical Report.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DeepSeek-V3 Technical Report

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.846948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:19f3dc1e12e5284ffa6c75d18fa329cf5b920d3c5e0bcbe52a5d1c64e015a380

Observation f8a7630c-e457-43e5-b772-208d9c711fbc · outbound

This paper cites On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:02.919880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:91fb78d088a8892176c7a2ea576f792fdfea6f6420bae250e0a8ebb4f143c3d9

Observation 339f55b3-9a09-42a7-871f-07a1d9229b69 · outbound

This paper cites Zhihu Zhuanlan , year=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Zhihu Zhuanlan , year=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.439655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:1659d41070c0aba6b122da10dd9a0792de94c4b09bd4d1f9f94a12cedcec88c5

Observation 64135884-6e2f-49d2-9914-5d90975be742 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Proximal Policy Optimization Algorithms

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.902885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:83605f0c617973621b794f531899a469f214233f67409bbe3ecbf23559414b63

Observation 7703f9f8-8a80-44f9-af87-8c5dc26d385d · outbound

This paper cites Machine Learning , volume=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Machine Learning , volume=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.412266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:aee5abac5fab93d77f01bc577899407c6919bab246b99eebac464f1f3f252e84

Observation d643fba9-3ac5-4661-bda5-ce0b9609a086 · outbound

This paper cites Advances in neural information processing systems , volume=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Advances in neural information processing systems , volume=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.404701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:9b84d1698d8418c675d88547b1b56a4521929a47d04ca25fdbcf6f6b65c7944c

Observation 6ab2c24a-6a69-42d1-81b1-7b2d204f26e5 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.415706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:64ecb8874c98a2af6b8ec87f7fc380081a15602a856935a63e7639c3ba1de72b

Observation 138d3726-b1f2-4d58-9e97-e6ba8612f863 · outbound

This paper cites International conference on machine learning , pages=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization International conference on machine learning , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.419590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:bc83186eb0803ffdc73733ed593d64b2c237fbe79bad96be44c45d929b01bb29

Observation 49ff38a0-1757-4a3c-a194-b628e658636f · outbound

This paper cites International conference on machine learning , pages=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization International conference on machine learning , pages=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.431823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:16e49aec2332692926ffcbb8dcba75df1b9150e354d5ccd595cfcd7c1a8ec099

Observation f63ea4b3-fee7-4416-9b64-cae9c7f20040 · outbound

This paper cites DCPO: Dynamic Clipping Policy Optimization.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DCPO: Dynamic Clipping Policy Optimization

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:02.915611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:8c2a4166ba0afa4c3822fe20e46b2e81ba299ea44efe34b5983034991b2207e1

Observation e7a042dc-bc83-46d0-8a15-13d0b8cf2450 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.938987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:13c376afd9f033063717ca43dae73e94c57e827d8fe3ed1191be7677a9c8e526

Observation 3c2f054b-53b2-4ebc-84ea-7a283a39378f · outbound

This paper cites 2024 , organization =.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization 2024 , organization =

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.435358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:be6a318bbce2d169116d0f4763f36ac9ee41d065bb74ac0a9a834fb57686b2d5

Observation f8bde296-48e9-448e-982e-23f602f12800 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:02.863401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:fed4510d6e3bf21cdc135bfd48253d13b7c1aba5a6742b2d3c22d29874bbaa78

Observation b411e255-e80b-45f4-a5fc-aa0596bb71d1 · outbound

This paper cites Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:02.852730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:508b70f14445cb70d3d00d225318bbbe41db27d4957f0f5fc856ba53355bec3f

Observation e115b3b2-c420-4da2-824f-a708c6261f90 · outbound

This paper cites Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.139688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:6e7d905f8298eb506286201c7728c1cc231c36ecb0fc19f98659ff1d349cbee2

Observation cd7ad71b-9101-475c-a584-a2de4c4843d3 · outbound

This paper cites Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:02.924935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:698e855938fabfbb2bb60012748fceb6913bceaa70284803ce79d6fe2498ecb3

Observation 2980fee8-513b-4ed6-9c10-0537c68478ed · outbound

This paper cites Decoupled Weight Decay Regularization.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Decoupled Weight Decay Regularization

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:42:02.888634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:108685d88deacd66538896a15dbdc4d595aefa623972401b2ddbb33c47883b7e

Observation cfaef6e5-7d2a-476b-9dd6-e4182d3eb683 · outbound

This paper cites 2017 , url=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization 2017 , url=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.408478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:c7d1149fc8ba50da954f8938dda6c56865a41fbfe9e9e592a0b92d8882ecd262

Observation bb1413e4-7d89-463c-8777-db47a92cb050 · outbound

This paper cites On-Policy RL with Optimal Reward Baseline.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization On-Policy RL with Optimal Reward Baseline

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:02.898305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:24a08be62e612f8ae49938a2e01144f6a15d6e02b1f459eec8385afbf711b2a3

Observation d38b3bcc-7b45-4ace-803e-c8c417aa7d82 · outbound

This paper cites Prosperity before collapse: How far can off-policy rl reach with stale data on llms?.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Prosperity before collapse: How far can off-policy rl reach with stale data on llms?

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:02.910063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:4b39d47a265e79f6b18047ef2350c2fd348d5e95d203e716f607425c2048f5b9

Observation f247a322-b1c6-4be2-9a99-b97a6a08b6c7 · outbound

This paper cites When Maximum Entropy Misleads Policy Optimization.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization When Maximum Entropy Misleads Policy Optimization

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:02.858360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:463722d26e78e50947e02705ebecfd5a5e0a77624aacf8cd81bdbbd4e8930f4b

Observation 79c5031c-f0ec-4fe3-9a85-60c213abefb0 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:5e048e719d15e1d1ee3c37ccd190f3c0f68b6f1f9f2daa80def6f162cd53ac97

Observation a14d97cd-74c9-44fa-afc8-c8ba359d7fc7 · outbound

This paper cites Proceedings of the twelfth international conference on machine learning , pages=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Proceedings of the twelfth international conference on machine learning , pages=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.423892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:83b6e2d080249114409fe29250d76a54679767db152ea2abc72b7c8ff766023e

Observation 0137a0e8-448b-4674-8edf-de7e6228ecc3 · outbound

This paper cites Advances in neural information processing systems , volume=.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization Advances in neural information processing systems , volume=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:53.427777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:9b8ed7cb1785f6f5ed5b14fed094bf73f97244e666070fafc29add47708cb27d

Observation a008f9aa-afcd-449e-9171-86d2674feece · outbound

This paper cites A Bayesian Perspective on Generalization and Stochastic Gradient Descent.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization A Bayesian Perspective on Generalization and Stochastic Gradient Descent

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:02.934283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:aa5fb7f319f53c4e374d33cf32b6748432259dd1fd7ce4e4c8f18595eb7af2c3

Pith citing papers

No inbound Pith citation observations are available.