Pith. sign in

Paper Citation Record · LEDGER

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation

As of 17 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2506.13358.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13358 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:07:05.170825Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43422694-5fcf-4c9a-8120-5b6f48571be6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.088037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.088037Z digest=sha256:fefb458ace07b79457eb19887dd2702b421232c06e694c46da209903658863b2

Observation 8e02623a-a6b6-4c28-aa55-e5469dce6ef9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.093314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.093314Z digest=sha256:bc8d935877ac54f368d88b000c1412ec57f82c59d171c1a9886a30c2e9c40246

Observation 47c59f1e-7650-492d-b43e-f500f4b2b15b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Understanding R1-Zero-Like Training: A Critical Perspective

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.097482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.097482Z digest=sha256:e1f8394900a4dea950503d6d4deb929e7482371ab03d43125e52855a1407ff16

Observation 20db525f-4548-4f5c-8824-2f87469da143 · outbound

This paper cites GPT-4 Technical Report.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.101687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.101687Z digest=sha256:c46c70c458bcdc255cdf1687f56d40f3030551afed4896e45a37040896370370

Observation 3fbcdbc4-e47f-431c-92ff-fdabb68813f8 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.105719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.105719Z digest=sha256:9dd8abc582f86f05e7aa0cf69d6df8782eaaf070d24cab649ca65687833b9771

Observation f77a8b28-5f4d-4a83-996b-b80cfc92e079 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.110212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.110212Z digest=sha256:0494bc5db266daba10866be3fa5570006a1e223c2bff276f90e4cd8846704992

Observation ff884a3e-0cc3-486c-8992-0839e1b10db6 · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.114855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.114855Z digest=sha256:00b6c55ae44ee48cec340dee91e01b2f635cbcfaed22b57201b50708ed04aec5

Observation 130bbb4e-c597-480b-8b5d-2e281be011ba · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.118325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.118325Z digest=sha256:a7ab2d936ed5fa5b08cf161b392b6671289b150f60785d446eabd3abc6efd6ff

Observation c1645e46-7316-44d4-a1a7-504f8ffb1115 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.122252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.122252Z digest=sha256:5336166cafa87814e57d204f66e0c66e6dc46a49259be81181d60c56b9cb3b8e

Observation 7f6ed1be-52b1-4235-9376-f13db9efdac8 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.125995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.125995Z digest=sha256:e81984c5f6faa321ec3bb6359641fb12851b99341b678e120bb999fed49ebbfa

Observation 50851a1d-8eed-4cbe-a43b-c5ded3787f5e · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.129778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.129778Z digest=sha256:8adc4e6772afaaa62ae15b2513c377f53291ae8f4d3dd96044aa61b387f8aab8

Observation 33be6667-9526-4d82-9891-906ba17f7781 · outbound

This paper cites OpenAI o1 System Card.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.133492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.133492Z digest=sha256:643335da535d25025794e2c54d84f6e617a8d38f0b9e593e3ac801b8b1c72152

Observation e4e1e4f6-7bdf-49d3-ac38-d51304b4600b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.137196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.137196Z digest=sha256:f106614d1fdfab4f0a5248a2c204627348f660c558cded18ce8ad70dca3c5424

Observation ce31fc07-e1ba-4d81-9984-d17d18c164d2 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation The claude 3 model family: Opus, sonnet, haiku

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.141038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.141038Z digest=sha256:46ccfdf6b935617e4bcd193cd4964c81d92511bc0a98730108d97a728615e19a

Observation c857d0ab-e713-46b6-af00-64917fb18109 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.144490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.144490Z digest=sha256:9a6121fbbd30be6ed596058b25b64533ae336ed18f0ee6e24fe4703bb17d2551

Observation 13f9409a-e21a-40e5-83c7-f305397d52fa · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:07:05.449047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:07:05.147982Z digest=sha256:104e318584ef80c97756b646361acb918fcbe6742050383831bef3620ae01d08

Observation 3729b848-3a45-4c30-913e-ed62d1d62d20 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation RLHF Workflow: From Reward Modeling to Online RLHF

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.151276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.151276Z digest=sha256:407008963bfa418c7e9b1941d86d38374f210fc75580be5cea2c4a0b1c11c11a

Observation 7577c151-8838-4f33-a636-e8f41ab32d04 · outbound

This paper cites Rlaif: Scaling reinforcement learning from human feedback with ai feedback.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Rlaif: Scaling reinforcement learning from human feedback with ai feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.154786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.154786Z digest=sha256:e122e267598695d6d9f9e11619583217d4c97c1e61b057711318475d7e7c8dcf

Observation 06283bf8-86ae-4bad-b5b4-55fab025bcaa · outbound

This paper cites Self- refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Self- refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.157904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.157904Z digest=sha256:eb844289a916b9a7e311e40da7846e83e19265092bb86c38136963896f004e45

Observation 20ed0976-138b-4920-bb1e-91c25288168c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.161052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.161052Z digest=sha256:1f0f96f9f62d0d2340f84e708f93aa5d1b50ed2a2dfb2c0ba158a1829869213c

Observation 8687d532-d921-48a9-a935-675ad2f9d2e0 · outbound

This paper cites Constitutional ai: Harmlessness from ai feedback, 2022.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Constitutional ai: Harmlessness from ai feedback, 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:07:05.416091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:07:05.164067Z digest=sha256:db8fd4602d132a3ed5d7e823d4bebd5d0d7093401a826d76093fb16326f08c3c

Observation b287acd4-044a-4b8a-b0af-c882bbcc7c44 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.167331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.167331Z digest=sha256:1c1be47de843d748d575114a3428cc9a97b4428951512d3baf126209e395f29c

Observation 0b9f6000-6e9b-48f9-87c5-38a7c22a23d0 · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.170825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.170825Z digest=sha256:07e499c2db4f84b30713f7e1bdff3692c355cc06e37aa417e49377f8ba8539b7

Pith citing papers

No inbound Pith citation observations are available.