Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:23.492554Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 3 inbound Pith citation observations for arXiv:2506.14731.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:23.492554Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:07:39.150321Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
15 of 15 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 791f41b8-8b5a-4841-9f32-c105a5483f26 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a0e87ad-35f9-427c-8aae-3ec1f1dfbbb9 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a24296e-cae0-4a09-88b7-873cf6066d1d · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be78314-a998-474a-996f-16531294b857 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dc87b8bc-e519-4efd-a889-a6939d8f7889 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dcd56729-79a2-468d-8c68-3dd13c2d5a48 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce9eb9b-0084-4264-ae49-e22ba1e4d821 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Are Your LLMs Capable of Stable Reasoning?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ad01978-3460-4b95-bf1d-3383f778d694 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Decoupled Weight Decay Regularization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4dc78d-d00c-4968-94e7-79b7d3d07e1f · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Qwen2.5 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd343ea1-b0af-4958-ac07-206f0ff553d3 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfda8c5e-4bc2-4b25-bbef-9a520947345d · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7369c7f0-44e4-4cd5-96dd-5f2271220783 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs TACO: Topics in Algorithmic COde generation dataset
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a844fb5d-128d-452a-bc81-47e67470c87b · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3cd57e7-3742-41e8-bc7a-9d202c092f4b · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Skywork Open Reasoner 1 Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83135e4b-12e8-4cb8-a729-be7995bc8e86 · outbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Nemotron-crossthink: Scaling self-learning beyond math reasoning.arXiv preprint arXiv: 2504.13941,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9fd62bf-9c62-42d2-950a-509aa958bef2 · inbound
Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d15c3b6-b9a6-4533-b4ce-7948d24c60b2 · inbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c822d43c-a000-4fea-8898-e4eba2f046e8 · inbound
Stabilizing Policy Optimization via Logits Convexity Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.