Pith. sign in

Paper Citation Record · LEDGER

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning

As of 21 July 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2512.05591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.05591 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T00:53:51.251900Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T05:53:19.306726Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact15
  • verified fuzzy7
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47852254-67b9-432a-9c82-8b5371fca63c · outbound

This paper cites Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.282148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:4d9fb331aaaa3f42c41dbcce35914347b31a479e5234667acb945444570fcf53

Observation 051254d6-6815-46c0-8514-79096d9ba67c · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.311837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:fea113ada049f2fe048eda0d1857af4453f9e335aa6c50ae82354e85397a4d89

Observation c919d9ec-072b-4ebf-8890-1ff2d662a3e8 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Reasoning with Exploration: An Entropy Perspective

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.297219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:429fd489eb8282457e21266ff450fd41872d34f23e4704c3efec064c95a46a20

Observation 3b62c073-3194-45b3-ac47-4d13381b48bb · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.317503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:bcbe8bd600d61bf4122d4e7f2f695eeaadfb849c6427f1e2df4d991c38e61b49

Observation abb0e17e-99ae-4f2e-a9c5-1c8be70e4e92 · outbound

This paper cites doi: 10.1038/s41586-025-09422-z.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning doi: 10.1038/s41586-025-09422-z

Reference 5

Resolution
verified exact
doi, observed 2026-05-17T00:58:46.269326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:623cc421704e4810454cc87a5e86421ebde0ef1d15844047c61f6e1fb7f2f2a5

Observation cb0b87c3-9ee5-414b-9332-75dfe5c9a4d4 · outbound

This paper cites O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 6

Resolution
verified exact
doi, observed 2026-05-17T00:58:46.314464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:dd460e1461fbd7fb4d1976db9a18b0aedfda480b47069418d36e76d4c2459bca

Observation c886c028-3b9c-4a41-9651-98be435abcdd · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Skywork Open Reasoner 1 Technical Report

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:26:47.432522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:bd845a7a4fc954af4253b7aa07d18925c3bb4dd34df7efe7f7aff02657723d3b

Observation afc627c3-5eb9-40a8-a8b5-3bccf0bd1985 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T00:58:46.264356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:951977d1c852d80a42d7dfa64da3d85d832d8af950bd75d0cec7ec387d4feeba

Observation a7e1ecc4-a593-490a-b305-abab8f8dc1b9 · outbound

This paper cites an unresolved cited work.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Unresolved cited work

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.668878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:449ba71c6f785366fc5a8e5154ec5ae28bd53ce41f57f85a55e5a1f441acab38

Observation a9364254-2f08-4f78-8056-1a594ccbc92d · outbound

This paper cites an unresolved cited work.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Unresolved cited work

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.663624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:5aa056e8e400a2f50c73d75b79b483845ca6508d3a5ee6ab58f7cb468e0535e2

Observation fd087267-0ab3-4580-ab85-1469ec6bfc2b · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:d23c53990e9a7c0fcca234034d5978ab049a0404e4feee5fcff937811a984531

Observation 15fae4bd-f790-4413-9a3b-86e863e6460c · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.658483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:1d42717ef118ca8fb82f3446b4495ceb4ddc556d50e04d2958bd79a9d0fc50ac

Observation a57e4718-d462-4a58-9cda-d0ff0864524b · outbound

This paper cites an unresolved cited work.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Unresolved cited work

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.666124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:a1634ee3a360bdeb8f7955f2f14136e3a1050654a928f2cfc8f5946d5ad8e4d5

Observation 4ee32e20-5fb6-4714-94c4-524904d7a74f · outbound

This paper cites Jordan, and Philipp Moritz.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Jordan, and Philipp Moritz

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.660989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:b7e38bf63ba601b15c6437bb1bc69fdc3765a5187028e64a1fcb6f351d046626

Observation c72b767c-337d-4011-bb60-3b840ac5fbf3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.653290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:4205db659b6d4101c0d04cd0c1da5aae9fa4946904dac1d9eeb277de89c87b49

Observation b199663b-922a-475f-b8a3-dbfb90c1aa89 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T00:58:46.308207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:57bd958e955706eb575315ddb97ac5750a8259ec9ac91080237b069e29978683

Observation 9536a65b-18f1-47b4-ae0c-dff7513a7d6e · outbound

This paper cites Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.289894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:3a64b9383e6e491fdaf7b4dca9eaa9dc162b0e47075cef81c0263bff2323c0b5

Observation 91f99259-325e-4266-b969-c308027e011b · outbound

This paper cites CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.300735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:84fc28c7f8cbf11af94b954b616983f1740566f66b553448ae71b5efde84f2e1

Observation 2d15c3b6-b9a6-4533-b4ce-7948d24c60b2 · outbound

This paper cites Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.293783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:0da15a03bed575f78118e4d617ea95898a389d2345a0b91c1a1c2f9a2dcda372

Observation 48abf2b6-0eab-43be-a8cc-8b44acbb1dff · outbound

This paper cites Qwen3 Technical Report.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Qwen3 Technical Report

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T00:58:46.286019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:0653f2e558257b34e7ede78db4cdae0e8634f7cf441669cf6faa1200bd11c529

Observation f27d17e7-0154-4e1e-8b83-2a0947408e08 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.657275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:7f9aa7239c60e76a893a1fb95014c515b7dd2aced76b77d9639530a53f989241

Observation e7572d50-345c-45d3-8c97-6f31eabbdcfe · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.277967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:bb267fccf3e847f8edc327303f596de844f1fe130c9e077e258257a7a2dd336f

Observation 69fa3b8a-b794-4c14-988d-08f8db7cede4 · outbound

This paper cites Group Sequence Policy Optimization.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Group Sequence Policy Optimization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.260517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:ce2260061bbfc12aa24fe18d84707f65b4f2583905332d6d4653f93c7dde0470

Observation dc28c5fb-2647-4ff9-a55f-0232c3e93694 · outbound

This paper cites online" 'onlinestring :=.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning online" 'onlinestring :=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.672015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:dabadd882f4ab1f7b16327dd8ac1df334e025326ee5be66f78e1b77a2c0c6269

Observation f27b4428-a805-490c-b6de-c5335be8975b · outbound

This paper cites write newline.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning write newline

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.675613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:975833d590f397403d91725c92b747f123c23a62458e97c6fb293042ca6f7b37

Pith citing papers

Observation a0ef91d8-d78b-4cb0-b86a-450a434dcc53 · inbound

GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment cites this paper.

GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:22.205659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-20T05:53:19.306726Z digest=sha256:afedffa93b5669c0e010f0ea3fd9f1a05c10bd4921c4299160dddb5f48230620