Pith. sign in

Paper Citation Record · LEDGER

Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2504.07912.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07912 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:37.961065Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 928eb557-583d-4526-8c3c-d1d4ea2ada9b · inbound

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling cites this paper.

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:40:46.496771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T23:39:39.018498Z digest=sha256:4ccc37d2a99884cd0391baa8c3a7479cc58f8cdec64e8c60e4c67ab5dcb85548

Observation 1c8fce14-0754-438c-bcef-c2de1fb7f8e5 · inbound

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning cites this paper.

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:56:55.052234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T23:56:47.274169Z digest=sha256:580e38f65039b04ef784db117085e4807c84e7439466c805b8cd85cb04f8ce97

Observation ce4dbadb-8bdf-48de-bfa1-e2464ee72378 · inbound

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration cites this paper.

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:36:53.734614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T22:33:01.074518Z digest=sha256:c9663f974cb85c0e932fc84b4ce060140de47ebcf41023407b702428f4302f0f

Observation cbe4cbfe-156e-46e5-874d-4a5c2542018f · inbound

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought cites this paper.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.961065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.961065Z digest=sha256:5a27ddcef72eb5fbb1f08a87abd93293bb6d4c86956b0f6ab89bd5ff65858b98

Observation 956c47bb-bede-428d-b782-c4f032804a4c · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 235

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:48.002445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:48.002445Z digest=sha256:a3067524f630bcccbbf8ce77323fba6e428ab81eb48b2f344b4df479b35e8e58

Observation f5cb03da-797e-420f-9839-d10cf6717f17 · inbound

Polychromic Objectives for Reinforcement Learning cites this paper.

Polychromic Objectives for Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.958923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:163667310f4a151506fd71bdbcde7dd0e1fccec67d6a7fddd353b9815fd5a288

Observation 9d30284e-7c29-4951-bac7-482e391095df · inbound

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts cites this paper.

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:29.812004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:39:29.812004Z digest=sha256:a59e9e2711d91e19f5537c2a9d6499d9f4173642ef8f9443ef6809cd2438de04

Observation 822b0dd1-0a95-4955-87e4-6c232c4c4d42 · inbound

Value Drifts: Tracing Value Alignment During LLM Post-Training cites this paper.

Value Drifts: Tracing Value Alignment During LLM Post-Training Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.061716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.061716Z digest=sha256:77bbcafd0e2f684c22db4e4b9f6fb9b586966a168005496d7d26a9bdd4afa0d3

Observation 3ca369b3-50c4-4a06-bb76-78579f3f88a7 · inbound

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning cites this paper.

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 1378

Resolution
malformed identifier
no resolver link, observed 2026-08-03T05:52:38.582979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:52:38.582979Z digest=sha256:5feee16c604f87aa5f862bf62ed3ad933905ba05b49ff3ba13d1a46fccdf6f3e

Observation 4c196a99-ed67-402f-8213-df1727a0d125 · inbound

Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning cites this paper.

Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:27:36.462167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:27:20.282859Z digest=sha256:5172cb4d37ece47884cdbe5ffc8b55405b7507e2dc887390d477c704515ebc68

Observation 48e5d4e4-67be-461b-a6bc-b400c58cffb1 · inbound

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics cites this paper.

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T23:11:58.245425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:11:58.245425Z digest=sha256:d4e0c74752f847a09c3205cfb6fcab3fb3e890dc2da37cf0acb908d3e06e024c

Observation 06b01db8-7aa0-4d64-b9d3-7cfa93ece865 · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:2cfa6296d49454a9f4967f7575d3911ef29772a2704dd6408f1a5cbb0086cb55

Observation b9ddf425-8719-445e-847c-40d4f1f00e86 · inbound

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning cites this paper.

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:13.937770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T04:29:21.897215Z digest=sha256:f2dded548b7b245ef25cbe9064b2bdc9682300493a0c3d1d62c8e334ec22e151

Observation 8b2ea63d-c95a-43b5-80ab-c311e81f9c11 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:06:31.196542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:ba016ffa01fe6f689918f444eb1b3af0f8ac555f46afd3afbdeffe21d01e1f1a

Observation a63f88fc-7293-424e-be01-d1f745737a87 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:07:42.335525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:9795e1f31a6a0adecdfa0fd49f536c5cdf14c182fbae73c425316b6f6e8bd4f3

Observation a8ae377b-f8f8-4fa1-9e10-a451ea56ac09 · inbound

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR cites this paper.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.445555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:d19757467f690d4a13e72cc3dde370720c61176e96cd63911c97aa5ac1fb3a9a

Observation 2409b916-7add-4a41-9b0d-35e670a02646 · inbound

Reasoning Can Be Restored by Correcting a Few Decision Tokens cites this paper.

Reasoning Can Be Restored by Correcting a Few Decision Tokens Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:57:46.848006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T20:56:48.771058Z digest=sha256:980ac92ab534ed38f8ee5b8d79741a1c8a44b20c4571ac0cb8e6f1ea2cdb7cb8

Observation 56707dc9-0357-4dfe-9935-42d6a653c18d · inbound

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning cites this paper.

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-05T10:20:57.200486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-05T10:18:18.717871Z digest=sha256:2994b1f0b602cdd49ef96de156e8ad05f7e35f5b9289ed826552bf5daae8ac44

Observation f67dd987-2349-4996-afd1-ab5850dba0dc · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T01:20:20.569507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:70e090a21bb4a27f0df5c42729b7463e29e32d907957b752d2009047e3fa1a18

Observation e419ae80-2f9d-411a-bc8c-2a8f4b32b3d2 · inbound

RL Post-Training Builds Compositional Reasoning Strategies cites this paper.

RL Post-Training Builds Compositional Reasoning Strategies Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.494223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:9863e5ec983d233acb1770cc230b35ddfb5254168a60f135c4bb5610116b87ef

Observation fa86c831-0274-4832-9eba-19bc8f4e8c0e · inbound

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion cites this paper.

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T07:14:53.084480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:14:53.084480Z digest=sha256:f4d1723711c8fd35cd4f651b761032aaba52246fb35a869cb5d98a44da53d9b6