Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:07:05.170825Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2506.13358.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:07:05.170825Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 43422694-5fcf-4c9a-8120-5b6f48571be6 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e02623a-a6b6-4c28-aa55-e5469dce6ef9 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c59f1e-7650-492d-b43e-f500f4b2b15b · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Understanding R1-Zero-Like Training: A Critical Perspective
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20db525f-4548-4f5c-8824-2f87469da143 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fbcdbc4-e47f-431c-92ff-fdabb68813f8 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77a8b28-5f4d-4a83-996b-b80cfc92e079 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff884a3e-0cc3-486c-8992-0839e1b10db6 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130bbb4e-c597-480b-8b5d-2e281be011ba · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1645e46-7316-44d4-a1a7-504f8ffb1115 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f6ed1be-52b1-4235-9376-f13db9efdac8 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50851a1d-8eed-4cbe-a43b-c5ded3787f5e · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33be6667-9526-4d82-9891-906ba17f7781 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation OpenAI o1 System Card
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e1e4f6-7bdf-49d3-ac38-d51304b4600b · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Gemini: A Family of Highly Capable Multimodal Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce31fc07-e1ba-4d81-9984-d17d18c164d2 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation The claude 3 model family: Opus, sonnet, haiku
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c857d0ab-e713-46b6-af00-64917fb18109 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13f9409a-e21a-40e5-83c7-f305397d52fa · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3729b848-3a45-4c30-913e-ed62d1d62d20 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation RLHF Workflow: From Reward Modeling to Online RLHF
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7577c151-8838-4f33-a636-e8f41ab32d04 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06283bf8-86ae-4bad-b5b4-55fab025bcaa · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Self- refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ed0976-138b-4920-bb1e-91c25288168c · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8687d532-d921-48a9-a935-675ad2f9d2e0 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Constitutional ai: Harmlessness from ai feedback, 2022
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b287acd4-044a-4b8a-b0af-c882bbcc7c44 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b9f6000-6e9b-48f9-87c5-38a7c22a23d0 · outbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.