Pith. sign in

Paper Citation Record · LEDGER

RL Post-Training Builds Compositional Reasoning Strategies

As of 5 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.07646.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07646 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T04:04:13.942267Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact12
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e5a1bfb-af03-4cbb-aa3e-1ee2e1e513c2 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

RL Post-Training Builds Compositional Reasoning Strategies Reasoning with Exploration: An Entropy Perspective

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T04:05:55.478521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:7f9592b560aeef3abdc3eaad33356018e124b28a098069003cbe8a6aaa3adedd

Observation b9b34299-e08b-4ff3-99f0-4fbe54755ddf · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

RL Post-Training Builds Compositional Reasoning Strategies SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T04:05:55.481095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:ca1f0ea4f26c986f3c12b70640d79b75d6362fd6fddd32ee34284173a537c5f4

Observation 886e0802-4b12-49da-8643-0c054b72ab5a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RL Post-Training Builds Compositional Reasoning Strategies Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.491553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:bc9f764a094b9b0d941f3cbe0e8f441db2e165eecd31b037ea5ccb5e9b22c0da

Observation d1161cd2-2b67-4b46-a2a0-96cc8d128ea2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RL Post-Training Builds Compositional Reasoning Strategies DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.493284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:ea0ca3c60b11eba876c62df7090af4aeac102524d919e2915128e21b227d4d34

Observation e4c210a0-bc67-4d85-8272-b4298fb9f1cd · outbound

This paper cites Let's Verify Step by Step.

RL Post-Training Builds Compositional Reasoning Strategies Let's Verify Step by Step

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.461276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:66996c2e1100e9f6c1d7fa671f26f248b62cf18caf83d8cd7ac46afa027ce042

Observation f5c25b2d-c968-4333-b469-e58d2bcee014 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

RL Post-Training Builds Compositional Reasoning Strategies ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.490553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:25c64867504d442f460dc786cac917adc4f5faf8deb1da5c9a7253ea11607ef1

Observation a4720dae-dd96-4b8a-a84b-9300af08d97d · outbound

This paper cites Rethinking sample polarity in reinforcement learning with verifiable rewards.

RL Post-Training Builds Compositional Reasoning Strategies Rethinking sample polarity in reinforcement learning with verifiable rewards

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T04:05:55.475237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:1efd23c5e967298cc7cf426b2443450819bd582c709c0d49ff01f544d63c977f

Observation 91d72329-834d-4aa9-85a3-a02ad3547861 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

RL Post-Training Builds Compositional Reasoning Strategies Solving math word problems with process- and outcome-based feedback

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.471938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:61f327a9951472223841a608078679511dce841073982f8c81fe2902d43b630a

Observation 1af94325-0b81-4800-a5d4-2bc2ea721aa8 · outbound

This paper cites Emergent hierarchical reasoning in llms through reinforcement learning.

RL Post-Training Builds Compositional Reasoning Strategies Emergent hierarchical reasoning in llms through reinforcement learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-09T04:05:55.484824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:956ba0cea47d30d1b1a71690b5fa5c1626ee928c97d34a4cffdc279bd6334afe

Observation 2a0565b5-129a-4dc0-886a-f67595930fbc · outbound

This paper cites The invisible leash: Why rlvr may or may not escape its origin.

RL Post-Training Builds Compositional Reasoning Strategies The invisible leash: Why rlvr may or may not escape its origin

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-09T04:05:55.474654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:c5d0e077e921af148b04d0b036722b4adfaf140bd50db594185a3a3a889132ab

Observation 6299d2f4-a2e8-4d50-99c9-9b112124ab68 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

RL Post-Training Builds Compositional Reasoning Strategies A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.457832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:c00c3b34efaa84aded23224045b9f671c3b83f03ecc1813f5cf95517f24badf4

Observation da7a4ad2-faa1-42e5-942e-73e99f1353fb · outbound

This paper cites arXiv preprint arXiv:2509.25123 , year=.

RL Post-Training Builds Compositional Reasoning Strategies arXiv preprint arXiv:2509.25123 , year=

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-09T04:05:55.478757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:51d3a0ef3540e2e8ec2aeea65686bbbc537341ef4a5ccb281af02aa8b731c9cf

Observation 19bbfce2-7ce2-4954-9f7a-bed918ce15fb · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

RL Post-Training Builds Compositional Reasoning Strategies Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.489088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:c0882a2aa2cfaacb6c4ddda84c416b5e30a05fec04aa5ed7c4170991c02ba6ff

Observation 94236739-64e4-4465-84ec-8c5cb5f8ea07 · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

RL Post-Training Builds Compositional Reasoning Strategies STaR: Bootstrapping Reasoning With Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.483803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:3e2826d415f0f790f226ed373b40b77e015167d8c742995c3cdab349bb5fa796

Observation e419ae80-2f9d-411a-bc8c-2a8f4b32b3d2 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

RL Post-Training Builds Compositional Reasoning Strategies Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.494223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:bc724dacc1f395ee737082fc7e91e412b673b518d0ca33b0178d826561719353

Pith citing papers

No inbound Pith citation observations are available.