Pith. sign in

Paper Citation Record · LEDGER

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

As of 5 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 9 inbound Pith citation observations for arXiv:2605.10781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10781 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:20:19.940462Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:13:29.068490Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T09:07:48.371751Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact23
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40413594-1191-4c95-90c6-05c7e2125c11 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Evaluating Large Language Models Trained on Code

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:22.432836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:99bff0b37419bea620952f7f5995ac07cec2e8a7cc16e76da5a1181c5ecbb3a6

Observation a6579577-05d4-4a44-a098-bf48c50f0bf6 · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.408581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:07729e4434380ea29450f2dcfebc9bb56d8f0f43a5f67a60f11ec85f670f5276

Observation 787862a4-87a9-4fb8-be27-4a27dd5daae7 · outbound

This paper cites Reasoning with exploration: An entropy perspective.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Reasoning with exploration: An entropy perspective

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:36:39.744841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:ba0fe753c61c234a94a2fe113cab1dffcbc32b3b93cf54f6e697e1cb320d46d6

Observation 3e8d6b83-a83c-44ff-9c65-09d5159c3fb7 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:31:34.026635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:08d43c3938b1c760d78a2d55859f1cf5a50224ddadd53eedd9de8a57402ecac9

Observation c2b54c09-ad9a-4653-8e53-a8fa3a7e11df · outbound

This paper cites Improving rl exploration for llm reasoning through retrospective replay.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Improving rl exploration for llm reasoning through retrospective replay

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:36:39.748064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:cf1cb8070b2f65c39912b5b31abb923875ac4bfdc9aecd5c8dcee7b84fe7ef2a

Observation 53bc1824-b14b-441b-a67d-e0d5ba1b610b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:22.426152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:dc57733fd7c04e01fe5c46b3ca7065d33912a4819103d4c5bd5230056e02a9b3

Observation d1709089-fb85-4a3b-975d-f7c060c4a330 · outbound

This paper cites Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:22.421664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:623ad2d4dc68d0f81dc0a43859ecaad58719dfe546363089e4b1a66e1d4011df

Observation 5074bb95-ae66-47c5-acf7-e5d60ab7decd · outbound

This paper cites Diversity-incentivized exploration for versatile reasoning.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Diversity-incentivized exploration for versatile reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.400659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:3d76ce61464961e188b228fad3949f340646bb9145f76dcfb34dd4584a532308

Observation 97f269f9-68c5-4c0c-81c2-051aabfa3216 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Reinforcement Learning via Self-Distillation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:29:18.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:8cb4a4c292fb0325fa05eb2c7faef1a4c33a6d9d397e6359a097c9f8f2e3b8b1

Observation 4824a536-d164-45ab-b032-e2b0e7132d6a · outbound

This paper cites Revisiting Entropy in Reinforcement Learning for Large Reasoning Models.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Revisiting Entropy in Reinforcement Learning for Large Reasoning Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:22.581915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:5366e7f0eecd82d8b77db409df29b0ec6ca5341732cfe4aa9691f2da19df712c

Observation 38c14afe-d670-4d9a-b719-93cf3d30400e · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:22.561357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:482c18947b1cf562b7ce953757d7de2e661d017ab7625f9dcdb43442a4f61098

Observation 9048021f-686b-4274-97fa-80bb9d3544c5 · outbound

This paper cites Unifying group-relative and self-distillation policy optimization via sample routing.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Unifying group-relative and self-distillation policy optimization via sample routing

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.530633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:d6e5eab9a6318a3f56f2cc6df661102a5b12e80b9aebd5db86e28bd3d73754ea

Observation 6761fb64-1cde-421d-af1b-53fe8db4d361 · outbound

This paper cites Exploratory memory-augmented LLM agent via hybrid on- and off-policy optimization.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Exploratory memory-augmented LLM agent via hybrid on- and off-policy optimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:36:39.750999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:955fc44e29af9b869bf24300a559ce9922d0caf72d2bba93ee3bb2fa6f3516a3

Observation ca70ce8b-6d95-4af4-b2f7-4f8b63a971f3 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Understanding r1-zero-like training: A critical perspective

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:36:39.755353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:f65e43c1389daf8087a1a07412ad1e6728e6b6146c0fc7f551985246d467a88e

Observation 50442307-3d06-4151-8e2a-febfdc15df2a · outbound

This paper cites Sparse but critical: A token-level analysis of distributional shifts in rlvr fine-tuning of llms.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Sparse but critical: A token-level analysis of distributional shifts in rlvr fine-tuning of llms

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.556249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:5347426186b6ec7d9dff3877933eee120ac56b615feb8e7909113fa5c14a14bc

Observation a4a13f01-28d1-4714-bdd5-7a4dc296a789 · outbound

This paper cites Fightin’words: Lexical feature selection and evaluation for identifying the content of political conflict.Political Analysis, 16 (4):372–403.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Fightin’words: Lexical feature selection and evaluation for identifying the content of political conflict.Political Analysis, 16 (4):372–403

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:36:39.758265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:badfd44060393938e2b2b7809bdf80ca47a4179fed72a858d34c639896819f86

Observation 6c42a37a-6815-4f75-b397-29c11480c2f4 · outbound

This paper cites arXiv:2510.02230 , year=.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR arXiv:2510.02230 , year=

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:21:22.545861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:d93bb418a8c7b1f38a8118a117d25172d6e715b1a8f92ee6f11ab0ac240bcbd9

Observation 81a7bed7-b543-475d-b62b-a9fae62b5b12 · outbound

This paper cites Clip-low increases entropy and clip-high decreases entropy in reinforcement learning of large language models.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Clip-low increases entropy and clip-high decreases entropy in reinforcement learning of large language models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.476543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:e8e7aa2f1506eee441b2292aba8ffd9aeccda8e213435219f7a8789c5e2d1f5e

Observation a886ae07-ac25-4491-9d2d-fbaa23a784c7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:22.491218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:433999c369cd41df988cb78d73bb5fdb1686abfba3fb3c78aa6705eba4ae3f25

Observation b9c49be3-95d3-43e3-a9b0-f9fc007a2373 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Self-Distillation Enables Continual Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:52.723793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:c5d01d1a62f8ee706b067008889704db9ef9ccb11c4863c47a5f84d9d336c1b2

Observation 3a0da2d4-b844-4ad7-966f-e3e171b50b9c · outbound

This paper cites Outcome-based Exploration for LLM Reasoning.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Outcome-based Exploration for LLM Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.471904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:8be01fae186bf74ae84e98546909a60515330a8ce5a2e93f2a79c0a04222b57b

Observation 7633ff53-f2bd-4ca5-bed0-7c9458b0f803 · outbound

This paper cites arXiv preprint arXiv:2602.02482 , year=.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR arXiv preprint arXiv:2602.02482 , year=

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.438495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:6235369e99533e87b643fdc3d3cf46598fe8807fc62ddde42591e6298c428c34

Observation 534b2987-ec12-4f09-8e94-1421a2318905 · outbound

This paper cites DSDR: Dual-scale diversity regularization for exploration in LLM reasoning.arXiv preprint arXiv:2602.19895.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR DSDR: Dual-scale diversity regularization for exploration in LLM reasoning.arXiv preprint arXiv:2602.19895

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.464592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:031b65eeaebccccaa6ec7a414d8b0d5a5ada80d597154436fd60972cf997423c

Observation 83783fba-31ff-496f-9dee-c05295e34c33 · outbound

This paper cites SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:05:34.479506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:fe007f7610cc66e20fd529e4b95759a58c0914e40ad1120e1de0c2850247980c

Observation 64dcb783-2730-420a-a05b-c05026d81050 · outbound

This paper cites Self-Distilled RLVR.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Self-Distilled RLVR

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:22.536857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:0ffe534e4fe0f25c49ae436b513ad2a4fcb19a27c87917a4fdbdea4095e9edd6

Observation de6ef254-7b89-450a-ae72-e740d0c0774f · outbound

This paper cites The debate on rlvr reasoning capability boundary: Shrinkage, expansion, or both? a two-stage dynamic view.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR The debate on rlvr reasoning capability boundary: Shrinkage, expansion, or both? a two-stage dynamic view

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.516189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:a58b4cf3a1d547d283bb670788d29d424e595d49435e36959c68188f9da02084

Observation 3ba93bfc-cbed-4bcc-95f8-7b9e32c1eeb2 · outbound

This paper cites On-Policy Context Distillation for Language Models.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR On-Policy Context Distillation for Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:48:41.492604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:a461a5223f4d4d78f31d0a0ed7bf6c26ab6e9ea68f4bbac3336cce70bace7fe0

Observation e40e8aba-ea15-4627-9908-61ed7fe4b6bc · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:22.448956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:6778f67a736aa57af7f5075a1e304614c6ea1b93057cf5e3c6a65197c817b063

Observation 62754914-43b4-40c2-91bc-45e4a2ab91b6 · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:36:39.761025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:44e3d8013c0670a46f413c6f67f68d1a96356a5c304ffeadfab7b330b30c7f1b

Observation 12be5d22-6e51-4c80-9d1e-e88a68e85fff · outbound

This paper cites On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.572306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:0f022309a9d3ab6f5e455909cd390e3121daf85337ab52b4ef92a8925b68f5b4

Observation a8ae377b-f8f8-4fa1-9e10-a451ea56ac09 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.445555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:bcf4a6dc59bbe85eddfcad0447737c5837c66382811108d8f1eb3c36026332a0

Observation 24103ca4-07c8-4bc4-b7e9-6782130182a4 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 32

Resolution
malformed identifier
local_arxiv, observed 2026-05-12T04:21:22.566603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:2469b60c07e3da44bdda9bb01a081ae62d8399e854b21b9252659f6551a3b588

Pith citing papers

Observation da622c26-1105-4028-9666-30f98f602728 · inbound

A Brief Overview: On-Policy Self-Distillation In Large Language Models cites this paper.

A Brief Overview: On-Policy Self-Distillation In Large Language Models Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:08:10.001047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T09:05:30.262601Z digest=sha256:bafb944eb2fec48b5b3f3d9370930b8a89dc49c6eddb91e91d9b5ee22a74dd91

Observation 22b40ec6-eaf0-49cf-8d18-de10ffbe0479 · inbound

A Brief Overview: On-Policy Self-Distillation In Large Language Models cites this paper.

A Brief Overview: On-Policy Self-Distillation In Large Language Models Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:54:47.024463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T09:51:52.886663Z digest=sha256:a81e043a9e1d8fb4323fce21f28ffac115e158234c10fcd7d5345d6b3fdffcad

Observation 83c53ce3-d5ce-427a-80ef-693339f060f2 · inbound

CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO cites this paper.

CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:26:01.075423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:25:04.067339Z digest=sha256:75355bf974fed805740343aed5a2b0a7f79b815f178f8204bede8f73d699aa1e

Observation 14c82400-5449-413d-aabe-a7f9e0864b7e · inbound

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation cites this paper.

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:07:48.373003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T10:28:18.490452Z digest=sha256:b05b4a6004f4596cbae8063301c2a6cb97664d26d4383f9ed0db2b4de5d54b80

Observation d2401d52-7387-44be-98fa-11b6596a74ac · inbound

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation cites this paper.

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:21.994570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T07:15:44.387347Z digest=sha256:f35a31b75ac9610064923a1d9c65cf507fd69f2bbfff980cde4a2e237ba30917

Observation 3648714b-4cc3-4d6f-9b40-dd8d5d9759e5 · inbound

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation cites this paper.

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T09:08:42.969885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:08:42.969885Z digest=sha256:d482562b4030c340b4ef0f256df91092e742d2d68fb9a4f38c6bc09dffe93bc6

Observation 1bfa63be-856f-48aa-9bee-a8b619557631 · inbound

CriPO: Enhancing Rubric-based RL via Self-Distillation cites this paper.

CriPO: Enhancing Rubric-based RL via Self-Distillation Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T16:14:24.250799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:14:24.250799Z digest=sha256:cb294c28920f851b4f2c49f92d1d9b4c8caef9f35e824bf6a6418be8f27ad9f8

Observation 09e9f91c-b592-4930-a6e2-15ef02a5b423 · inbound

CriPO: Enhancing Rubric-based RL via Self-Distillation cites this paper.

CriPO: Enhancing Rubric-based RL via Self-Distillation Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:13:29.068490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:13:29.068490Z digest=sha256:1dbbc30bc8c6b571cc37e30817ae9bb9eec0b8f552490d2de16676d049c1d840

Observation 4c2a502f-0aea-4f6b-b807-7a21c373f0f5 · inbound

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks cites this paper.

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:59.655967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:59.655967Z digest=sha256:ee26680f61ad71edba8e3c0007abb1be39faa9165434f09592eacb623c185028