Pith. sign in

Paper Citation Record · LEDGER

Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2504.07912.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07912 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:13.065149Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b1265abc-8c31-453d-8e78-17fdf1016739 · inbound

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning cites this paper.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:42.861182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:42.861182Z digest=sha256:bedfc1a4a117a943f0b668d8132103a6aafaaa0e6fc40d0a7402a3dc11578032

Observation 8b14d4b8-a2bd-44ab-8590-7b87811c4b6e · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.977166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.977166Z digest=sha256:b61413eb96a90194e7a6b8e5d9628c852291c4e218010cb229445c7fa46353b2

Observation a25873b9-4bbb-48a7-b9d8-600f23a38b99 · inbound

Accelerating RL for LLM Reasoning with Optimal Advantage Regression cites this paper.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.838631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.838631Z digest=sha256:17b0158279f9c7b82b3a6904744a6efd1e90d7e9be91258aa613c216496e7e07

Observation bb07e421-7e4c-483a-8836-43bc5d0f129b · inbound

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning cites this paper.

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:24.730994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:24.730994Z digest=sha256:fab9bf404222b571d9627956122b7800e1315ee89112e9e0d4b3d9722852d5c4

Observation 88fa932b-ec92-4839-a37e-5fcf057a7eeb · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.584200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.584200Z digest=sha256:03835f5115c0cb01403b0e11df21e9498c3b17943b2de4e075b3a64ece73c29e

Observation 7b7724ee-3cc2-4c91-8daa-3cf638a5b3d5 · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.749398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.749398Z digest=sha256:b2b59fbe9ba3547826235d34acc268aa28b958cca1a6ec6a880abbf6518b5fbe

Observation 77f34f5e-32a2-486f-a497-2323514f9ea2 · inbound

Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning cites this paper.

Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:08.980643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:39:08.980643Z digest=sha256:67099c797a6ea764cc73e05bf0142a44fa5b92b266a1bcdf4f023a011afa06f1

Observation 3f71c7fa-5d4b-4ef8-a912-a461dd82f61c · inbound

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models cites this paper.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.795096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.795096Z digest=sha256:3e4a62a4b9f548709a8f8a032cc756eaaa88aded72ec4c20d9b66fe51c14516a

Observation c34a1c97-537c-4a4d-a78e-267fce97c439 · inbound

RAST: Reasoning Activation in LLMs via Small-model Transfer cites this paper.

RAST: Reasoning Activation in LLMs via Small-model Transfer Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:59.473464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:59.473464Z digest=sha256:32138744789f94fdefe43c1033de439485575a70e100a187a79c77ea72ca175d

Observation cbb9ee45-ebdf-46fe-9dd0-c32d2aec739b · inbound

Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess cites this paper.

Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:15.416519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:13:15.416519Z digest=sha256:7e916e128f7bca7ca456543a99a8ca163d2d7e69307677faa844cd733178eb9c

Observation 928eb557-583d-4526-8c3c-d1d4ea2ada9b · inbound

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling cites this paper.

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:40:46.496771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T23:39:39.018498Z digest=sha256:6b9cfece73dff2fada8e979035293d5d827900b999ca86e53cb8ee9c7a1a707e

Observation 85c4f4bd-1b2b-4aca-830a-17452bf3b613 · inbound

CTR-Guided Generative Query Suggestion in Conversational Search cites this paper.

CTR-Guided Generative Query Suggestion in Conversational Search Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:55.387382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:55.387382Z digest=sha256:d67b3a267cc426e452f8db345fe4c25ddb47196127683ff0c06c852250aef2ed

Observation 6bc7292e-75cf-4773-bb90-8c07c170c430 · inbound

Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them cites this paper.

Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.701241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:01.701241Z digest=sha256:928cee81dd1d7b74abf4827cdcfef60316ae357b82c2c4c118c69d9d3e2b6239

Observation 1c8fce14-0754-438c-bcef-c2de1fb7f8e5 · inbound

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning cites this paper.

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:56:55.052234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:56:47.274169Z digest=sha256:e24354f8b2decce73f136c41032aa462a43e685947fa57453d25109f564cf171

Observation ce4dbadb-8bdf-48de-bfa1-e2464ee72378 · inbound

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration cites this paper.

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:36:53.734614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T22:33:01.074518Z digest=sha256:c40d8246888fe19273c87b1597b05598e9be11d39b8d392472fb2bd230931eb7

Observation cbe4cbfe-156e-46e5-874d-4a5c2542018f · inbound

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought cites this paper.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.961065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.961065Z digest=sha256:ef821b2b44fbd73f975e674c1985d3a959c4ac2424fc6dde4df2fce6e37e9011

Observation 956c47bb-bede-428d-b782-c4f032804a4c · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 235

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:48.002445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:48.002445Z digest=sha256:dcde8bd6451d031f1dbc6bd04c09c3096b8ffebb35722fef1ce895ece4ed5d51

Observation f5cb03da-797e-420f-9839-d10cf6717f17 · inbound

Polychromic Objectives for Reinforcement Learning cites this paper.

Polychromic Objectives for Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.958923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:ab162f053935cd6fd293880d2b8fc822e79e688f01a08f59ff001faa4cf880c8

Observation 9d30284e-7c29-4951-bac7-482e391095df · inbound

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts cites this paper.

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:29.812004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:39:29.812004Z digest=sha256:1a3e822d2e0c9d3494246936e43fbb21c094d6b07196db5382b9a1d65fc9c33a

Observation 822b0dd1-0a95-4955-87e4-6c232c4c4d42 · inbound

Value Drifts: Tracing Value Alignment During LLM Post-Training cites this paper.

Value Drifts: Tracing Value Alignment During LLM Post-Training Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.061716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.061716Z digest=sha256:70316879f7534e5bf49893cc56e932a8ab2855b5e444ba6b265165232dc3168f

Observation 3ca369b3-50c4-4a06-bb76-78579f3f88a7 · inbound

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning cites this paper.

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 1378

Resolution
malformed identifier
no resolver link, observed 2026-08-03T05:52:38.582979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:52:38.582979Z digest=sha256:54e684d2eaeef2f53e50ab41e8f18c1c7051ab60f5ba0cbc1add6836f322563d

Observation 4c196a99-ed67-402f-8213-df1727a0d125 · inbound

Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning cites this paper.

Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:27:36.462167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T08:27:20.282859Z digest=sha256:ce19a6d6d0e21240de4b7c7a468461216e3ddcc9aae3aff7f987d365da828fb2

Observation 48e5d4e4-67be-461b-a6bc-b400c58cffb1 · inbound

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics cites this paper.

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T23:11:58.245425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:11:58.245425Z digest=sha256:39cd337a1952c5498552d1232cca8907fc0b89cda685c01ca794750e10297017

Observation 06b01db8-7aa0-4d64-b9d3-7cfa93ece865 · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:9b707b0973a3070a9134726139f8e86f1dcfff170962daa14b65948d772d40e9

Observation b9ddf425-8719-445e-847c-40d4f1f00e86 · inbound

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning cites this paper.

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:13.937770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T04:29:21.897215Z digest=sha256:5c4788c95e5f56977307a252bb83743e64951668859a96c2b25e5d6c4aa623bc

Observation 8b2ea63d-c95a-43b5-80ab-c311e81f9c11 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:06:31.196542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:6ee8987e195b1f741e45d73688a9607ea85737d3cce9f2de355437bf43b0e36d

Observation a63f88fc-7293-424e-be01-d1f745737a87 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:07:42.335525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:eeeeb3b6b75e31f3289312f4956e3c74cd0623cfdd5b4b22a867d446866b72c4

Observation a8ae377b-f8f8-4fa1-9e10-a451ea56ac09 · inbound

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR cites this paper.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.445555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:b7aa4bcba36651dbebd0298c3531f72061475621eb6fc5eec31fdb40f01231b3

Observation 2409b916-7add-4a41-9b0d-35e670a02646 · inbound

Reasoning Can Be Restored by Correcting a Few Decision Tokens cites this paper.

Reasoning Can Be Restored by Correcting a Few Decision Tokens Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:57:46.848006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T20:56:48.771058Z digest=sha256:99eff22c9534fa814993176e26d390df729d8517c6bad4f454ceef095926ace3

Observation 56707dc9-0357-4dfe-9935-42d6a653c18d · inbound

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning cites this paper.

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-05T10:20:57.200486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-05T10:18:18.717871Z digest=sha256:59352d6358061a95a6367ddfb1ccb1939f5c34948f93eb0cfa2a80ac2dca86df

Observation f67dd987-2349-4996-afd1-ab5850dba0dc · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T01:20:20.569507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:2049809d35264860e9570d398e038062c88604a7cd8dd028b349d56474d2772c

Observation e419ae80-2f9d-411a-bc8c-2a8f4b32b3d2 · inbound

RL Post-Training Builds Compositional Reasoning Strategies cites this paper.

RL Post-Training Builds Compositional Reasoning Strategies Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.494223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:c20c2c51725f462d51c7c37c1a2005858e9bff7e8e1fa14726ccc02aabde1fce

Observation fa86c831-0274-4832-9eba-19bc8f4e8c0e · inbound

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion cites this paper.

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T07:14:53.084480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:14:53.084480Z digest=sha256:6f9a1a27d06462ba0c0d6e428842aa289129862b33370aa9721b3e19c2eb5d7f

Observation 15009b28-60d0-476b-927b-4e68acfec58a · inbound

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models cites this paper.

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 148

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:40.951940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:40.951940Z digest=sha256:0e7656718a864f694cabd05ab92691ae445a1dcfaf8f11a3483b7d740c5ab04a

Observation 91df57ff-200d-4d8b-b1c3-ad18867ce443 · inbound

Parameter Exploration for RLVR via Variational Learning cites this paper.

Parameter Exploration for RLVR via Variational Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.129972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.129972Z digest=sha256:7b1305924aa54b0bd14adb270b4a8926d866d5b534f0d095bd758e10fd9a9d89

Observation 01c6da4f-5b98-4b78-ba78-d2c75dd6599c · inbound

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure cites this paper.

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:13.065149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:13.065149Z digest=sha256:0248c3f1f02f32daabec7167dca49f05904003d6514b58dced1b7e4c0d9b2f81