Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T00:53:51.251900Z
Paper Citation Record · LEDGER
As of 21 July 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2512.05591.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T00:53:51.251900Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T05:53:19.306726Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 47852254-67b9-432a-9c82-8b5371fca63c · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 051254d6-6815-46c0-8514-79096d9ba67c · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c919d9ec-072b-4ebf-8890-1ff2d662a3e8 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Reasoning with Exploration: An Entropy Perspective
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 3b62c073-3194-45b3-ac47-4d13381b48bb · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation abb0e17e-99ae-4f2e-a9c5-1c8be70e4e92 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning doi: 10.1038/s41586-025-09422-z
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation cb0b87c3-9ee5-414b-9332-75dfe5c9a4d4 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c886c028-3b9c-4a41-9651-98be435abcdd · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Skywork Open Reasoner 1 Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation afc627c3-5eb9-40a8-a8b5-3bccf0bd1985 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a7e1ecc4-a593-490a-b305-abab8f8dc1b9 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a9364254-2f08-4f78-8056-1a594ccbc92d · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fd087267-0ab3-4580-ab85-1469ec6bfc2b · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 15fae4bd-f790-4413-9a3b-86e863e6460c · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a57e4718-d462-4a58-9cda-d0ff0864524b · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4ee32e20-5fb6-4714-94c4-524904d7a74f · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Jordan, and Philipp Moritz
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c72b767c-337d-4011-bb60-3b840ac5fbf3 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b199663b-922a-475f-b8a3-dbfb90c1aa89 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 9536a65b-18f1-47b4-ae0c-dff7513a7d6e · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 91f99259-325e-4266-b969-c308027e011b · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2d15c3b6-b9a6-4533-b4ce-7948d24c60b2 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 48abf2b6-0eab-43be-a8cc-8b44acbb1dff · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Qwen3 Technical Report
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f27d17e7-0154-4e1e-8b83-2a0947408e08 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e7572d50-345c-45d3-8c97-6f31eabbdcfe · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 69fa3b8a-b794-4c14-988d-08f8db7cede4 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Group Sequence Policy Optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation dc28c5fb-2647-4ff9-a55f-0232c3e93694 · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning online" 'onlinestring :=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f27b4428-a805-490c-b6de-c5335be8975b · outbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning write newline
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a0ef91d8-d78b-4cb0-b86a-450a434dcc53 · inbound
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.