Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T04:04:13.942267Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.07646.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T04:04:13.942267Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9e5a1bfb-af03-4cbb-aa3e-1ee2e1e513c2 · outbound
RL Post-Training Builds Compositional Reasoning Strategies Reasoning with Exploration: An Entropy Perspective
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9b34299-e08b-4ff3-99f0-4fbe54755ddf · outbound
RL Post-Training Builds Compositional Reasoning Strategies SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 886e0802-4b12-49da-8643-0c054b72ab5a · outbound
RL Post-Training Builds Compositional Reasoning Strategies Training Verifiers to Solve Math Word Problems
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d1161cd2-2b67-4b46-a2a0-96cc8d128ea2 · outbound
RL Post-Training Builds Compositional Reasoning Strategies DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4c210a0-bc67-4d85-8272-b4298fb9f1cd · outbound
RL Post-Training Builds Compositional Reasoning Strategies Let's Verify Step by Step
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f5c25b2d-c968-4333-b469-e58d2bcee014 · outbound
RL Post-Training Builds Compositional Reasoning Strategies ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a4720dae-dd96-4b8a-a84b-9300af08d97d · outbound
RL Post-Training Builds Compositional Reasoning Strategies Rethinking sample polarity in reinforcement learning with verifiable rewards
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 91d72329-834d-4aa9-85a3-a02ad3547861 · outbound
RL Post-Training Builds Compositional Reasoning Strategies Solving math word problems with process- and outcome-based feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1af94325-0b81-4800-a5d4-2bc2ea721aa8 · outbound
RL Post-Training Builds Compositional Reasoning Strategies Emergent hierarchical reasoning in llms through reinforcement learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2a0565b5-129a-4dc0-886a-f67595930fbc · outbound
RL Post-Training Builds Compositional Reasoning Strategies The invisible leash: Why rlvr may or may not escape its origin
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6299d2f4-a2e8-4d50-99c9-9b112124ab68 · outbound
RL Post-Training Builds Compositional Reasoning Strategies A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da7a4ad2-faa1-42e5-942e-73e99f1353fb · outbound
RL Post-Training Builds Compositional Reasoning Strategies arXiv preprint arXiv:2509.25123 , year=
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19bbfce2-7ce2-4954-9f7a-bed918ce15fb · outbound
RL Post-Training Builds Compositional Reasoning Strategies Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94236739-64e4-4465-84ec-8c5cb5f8ea07 · outbound
RL Post-Training Builds Compositional Reasoning Strategies STaR: Bootstrapping Reasoning With Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e419ae80-2f9d-411a-bc8c-2a8f4b32b3d2 · outbound
RL Post-Training Builds Compositional Reasoning Strategies Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.