Pith. sign in

Paper Citation Record · LEDGER

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

As of 20 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 3 inbound Pith citation observations for arXiv:2506.14731.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14731 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:23.492554Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:07:39.150321Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 791f41b8-8b5a-4841-9f32-c105a5483f26 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.426243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.426243Z digest=sha256:a49ceccd1456e2e33a09f8da7452c3b27647c3cfe8fb535a42c32a00485cbb26

Observation 2a0e87ad-35f9-427c-8aae-3ec1f1dfbbb9 · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.519933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.519933Z digest=sha256:c4d23941515b410170eca48de52e73d69a8774c6e7c534ca816a9028a6ed94c0

Observation 5a24296e-cae0-4a09-88b7-873cf6066d1d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.609418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.609418Z digest=sha256:182a503c7a2fbc23e604c577bf4fc1c33d323b216fb7420bbf9e333f2ff162c7

Observation 0be78314-a998-474a-996f-16531294b857 · outbound

This paper cites an unresolved cited work.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:14:24.243668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:14:22.791494Z digest=sha256:21a3e9feab985888290d851a03228415879e175edf27efc72e6159fa202c78fa

Observation dc87b8bc-e519-4efd-a889-a6939d8f7889 · outbound

This paper cites Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:24.075438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:14:22.858003Z digest=sha256:702bdfcc506aaad6ee2b527fc61d5bf25628f7b63daddb3ae2d64cfdca6080b2

Observation dcd56729-79a2-468d-8c68-3dd13c2d5a48 · outbound

This paper cites Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.999127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.999127Z digest=sha256:9ced363052b30316669430514f960d8ecfbf64296ca36d56a59e205b3d978bcb

Observation 0ce9eb9b-0084-4264-ae49-e22ba1e4d821 · outbound

This paper cites Are Your LLMs Capable of Stable Reasoning?.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Are Your LLMs Capable of Stable Reasoning?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:23.045879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:23.045879Z digest=sha256:edcb363afe19ef08a29d7246d0e32c2af75fccbeb0ffcdffbdcd6d2b6e0012b5

Observation 3ad01978-3460-4b95-bf1d-3383f778d694 · outbound

This paper cites Decoupled Weight Decay Regularization.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Decoupled Weight Decay Regularization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:23.120263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:23.120263Z digest=sha256:9e12c1c166f507c061e5684e196e9090cc91cd987695606fb2bb7fa0447d2e5a

Observation 4d4dc78d-d00c-4968-94e7-79b7d3d07e1f · outbound

This paper cites Qwen2.5 Technical Report.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Qwen2.5 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:23.190529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:23.190529Z digest=sha256:2043eb4de14a9d835863680a98615416a1330008be0930414d64229e9bab54a7

Observation fd343ea1-b0af-4958-ac07-206f0ff553d3 · outbound

This paper cites an unresolved cited work.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:23.365603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:23.365603Z digest=sha256:c57b20d50d4ba6063285e76007c3957a28bafc2f204cb9d3bbcd577cf647c6e8

Observation cfda8c5e-4bc2-4b25-bbef-9a520947345d · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:23.492554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:23.492554Z digest=sha256:d94dc11b7f7ae5a9629df2acff66b373de6a062c7203b93aa0be0c6fd61d433e

Observation 7369c7f0-44e4-4cd5-96dd-5f2271220783 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs TACO: Topics in Algorithmic COde generation dataset

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.930873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.930873Z digest=sha256:8a77c828e794dbc4ba1ec9b1516099b7b04bd90a7ab55185d51b23a210543c3c

Observation a844fb5d-128d-452a-bc81-47e67470c87b · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:23.277165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:23.277165Z digest=sha256:ea059f979b366998f1ba039ee319f5cb6691f1aa9ab9e428fccf2a06b46d7afd

Observation d3cd57e7-3742-41e8-bc7a-9d202c092f4b · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Skywork Open Reasoner 1 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.710839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.710839Z digest=sha256:bf4cc7ee3ebda504aef9d48728e75fcd76f0bf73513f285466cc7288425d1e77

Observation 83135e4b-12e8-4cb8-a729-be7995bc8e86 · outbound

This paper cites Nemotron-crossthink: Scaling self-learning beyond math reasoning.arXiv preprint arXiv: 2504.13941,.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs Nemotron-crossthink: Scaling self-learning beyond math reasoning.arXiv preprint arXiv: 2504.13941,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.327769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.327769Z digest=sha256:7b086c597b5b91a70aa697f5d5ba5a508a8f9990783831fc4c02d8b601d42900

Pith citing papers

Observation a9fd62bf-9c62-42d2-950a-509aa958bef2 · inbound

Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning cites this paper.

Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:39.150321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:39.150321Z digest=sha256:53df4a765f2cdbafcfbf75792ad59fddc66650dfb5fc25cd33b85236d445bc55

Observation 2d15c3b6-b9a6-4533-b4ce-7948d24c60b2 · inbound

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning cites this paper.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.293783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:1f0e0bd72c1f56a18762da31b7b31a9c05c938fe6d58292c202a1bec6757d3ab

Observation c822d43c-a000-4fea-8898-e4eba2f046e8 · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.502936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.502936Z digest=sha256:c16b2772eb897f08ffdde645da4db9d429bbe041581943b18be238cd69009d44