Pith. sign in

Paper Citation Record · LEDGER

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2608.06243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06243 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:38:13.296565Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved48
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4219dd4-1dc8-48ae-a32d-8e59cf97dc34 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.106699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.106699Z digest=sha256:69ef96fdc786ae019db90fe12428deafda44c30353bcfa7ad488a8942ee97145

Observation 726d5435-1ecd-4f66-8769-a2242d3f5264 · outbound

This paper cites 2025 , doi =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , doi =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.112655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.112655Z digest=sha256:32c804d42cbaa7c77b3bac16a7b02b5834cb7f8212c7ebf19b46019ec97d9bdd

Observation eda955a7-ecd4-4746-9535-02c278dd2022 · outbound

This paper cites International Conference on Learning Representations , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models International Conference on Learning Representations , year =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.116032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.116032Z digest=sha256:91b0decd5b918ee7c75589e60db6e517c9f938315e82a01ba37b7a7a9ea8d6b5

Observation 35218719-cf3c-48af-84a0-7bf1d25e9e91 · outbound

This paper cites 2025 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , eprint =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.119605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.119605Z digest=sha256:6a8ff12dd03f758dfa25e85f614cca45a109f8cb2103dfe6dc3793a83b35936e

Observation 77be2d5b-91d3-4f50-af29-38a046dded11 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.123104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.123104Z digest=sha256:df2633a640b96606649d9e7560407476f9b6346285ec548b012c74e1f570decf

Observation 4aa4d116-78de-46fd-8b3b-50f1884c3425 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.126568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.126568Z digest=sha256:7922e6b78ede41047bac10e30cb4c3324d212306a3fbb8b276028e1840f08dc4

Observation a3d96e27-24ea-4989-a382-7cebde7f76b4 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.130376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.130376Z digest=sha256:ca8663c53287b50c26b7942510773256c2e4b71a7527c923e87493d5caf10e7b

Observation 3496a0ea-259d-4abf-8081-293c11409170 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Entropy-Aware On-Policy Distillation of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.133996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.133996Z digest=sha256:ac98d4a9a7d4a50fdaf54cb9045ccc7d51ce05710cf6f5dd2e21a8abcb68d28a

Observation 658e2172-3811-4392-90c3-fc95df007c96 · outbound

This paper cites When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.138616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.138616Z digest=sha256:ceb24673fc080f23355bb1a26449b7917c5a75c196b0eaba76db56768dbfb8f6

Observation b1d36558-4772-43ab-a2c1-95bc733479a7 · outbound

This paper cites On the Position Bias of On-Policy Distillation.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models On the Position Bias of On-Policy Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.143580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.143580Z digest=sha256:944ba3a4c257141242cf9d39423d73fd154b2933426bfc831b4f8bcd322d095c

Observation 65004d52-a603-482a-8250-49eb9f656b22 · outbound

This paper cites Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.148001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.148001Z digest=sha256:24d295088e80be439f9acbf580a7b556056fc9f40433a5befec862e9504dd33b

Observation b8f8facd-2fb7-4aa1-9442-792adc2bb604 · outbound

This paper cites 2025 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , eprint =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.152075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.152075Z digest=sha256:46d4453a14bb0046a18880d4087b02c34b47942f7f713e4d082860dfc98e0e4f

Observation 1486c8b0-c60f-423b-b72d-96011c922606 · outbound

This paper cites 2025 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , eprint =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.155967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.155967Z digest=sha256:500afaaa46d39fb8512c12f0f85e06d5594cfc183101aa1dac41e245ab8a51cf

Observation 707e60fc-85eb-40d6-8cf2-3be294fd1c8c · outbound

This paper cites 2026 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2026 , eprint =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.159977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.159977Z digest=sha256:ecd3626a7db17778c7825a1b75e110dd88d6a66e3b3726fb697c7de9e9c8a35f

Observation 18d30972-ef0d-46dc-986c-1014289d4d7f · outbound

This paper cites Purified.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Purified

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.164056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.164056Z digest=sha256:21dc2020d5a6ac91903d0f3d9c17dd9a9b37c9abadc88c3ae3be6e0e74c5d1c6

Observation 799d723b-8736-4285-a3a5-546de5330d24 · outbound

This paper cites and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.167823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.167823Z digest=sha256:40c9cccf0198cf90a2f55e25b81d8158ccedacbbe70dc233a06df4bb0f96fc42

Observation e5570b07-990d-4a48-85d0-6965be7c2fb2 · outbound

This paper cites 2026 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2026 , eprint =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.171465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.171465Z digest=sha256:195174804f9d8046f25c54b664b7298ddb13886c5465781fcf02a59917042701

Observation a463fb26-6db5-4f4b-8b61-2a110149adc1 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 18

Resolution
parse uncertain
no resolver link, observed 2026-08-15T14:38:13.175594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.175594Z digest=sha256:4eafbca824063b64f8e16f45719e066433955e685d54c0dbe2cf95f164074ecd

Observation d8e611ed-ac46-4b1d-a7b4-d2322e8cd760 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 19

Resolution
parse uncertain
no resolver link, observed 2026-08-15T14:38:13.179253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.179253Z digest=sha256:30d2a6497c9319cb8e7d9fa67282195ba976aa44acf998eee3b32834d7ededde

Observation 1673e8b1-bf0e-469c-9678-1629ef476dbe · outbound

This paper cites and Liu, Alisa and Dziri, Nouha and Lyu, Shane and Gu, Yuling and Malik, Saumya and Graf, Victoria and Hwang, Jena D.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models and Liu, Alisa and Dziri, Nouha and Lyu, Shane and Gu, Yuling and Malik, Saumya and Graf, Victoria and Hwang, Jena D

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.182936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.182936Z digest=sha256:a93b2a979674c3312fccbdff150dc669ada58627b2ac80fbcda7a784351dfbef

Observation 70769e61-e37a-41f5-be4c-86b17454c191 · outbound

This paper cites 2025 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , eprint =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.186410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.186410Z digest=sha256:ad4b6b8050da63281ec18aad4a1c32af158b0e86c52f10490e8ef0d804d21a35

Observation c0d0ca39-527d-4c4f-b579-af30d59a96cf · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.189982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.189982Z digest=sha256:53b17f30372b79ae0ef29a5182287494a9d83da29d018ab810cdb23cd1f9b160

Observation e021e681-8e98-4066-a4bd-602d2adb5631 · outbound

This paper cites Proximal Policy Optimization Algorithms.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.193577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.193577Z digest=sha256:95cd077a48083687ce2fe5e3e4b3e9d9734bcbdd15eb3cde0d49cbdf1ef3d302

Observation 35ef7025-daf5-410b-9df5-b12d96148eb5 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.197421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.197421Z digest=sha256:bf19b1cd4338ec514b836e75878f25281120a2aa3fc94dea1f9aace9db8d15cd

Observation 8c4fa34c-b90c-4ec8-8dda-a5b5ec07bbf2 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.201026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.201026Z digest=sha256:f8535ab8f32810f43ba05f3b740aeeb7d3db0e04de124c9b99cd86bb0f124708

Observation 38eb1fa2-53bb-40ea-b757-be4a789de237 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Training Verifiers to Solve Math Word Problems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.205106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.205106Z digest=sha256:b060262875f40c617c7c73d88261e20ded4e455191c9560580ef7f9c98be31b5

Observation 2e230541-cadb-4756-9b70-d2c741a229b6 · outbound

This paper cites Measuring Mathematical Problem Solving with the.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Measuring Mathematical Problem Solving with the

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.208929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.208929Z digest=sha256:30db191c3175be2c4af24c5c1aab0446f65e34c9cc5f14adc375bcfd3fd6a93c

Observation 1af62569-06fb-4acd-8790-49d4c7f075c3 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Evaluating Large Language Models Trained on Code

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.213714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.213714Z digest=sha256:451fb8d07b64919fe9a4aca4642668d7f8b10cb3fe90f13661052a459af50af5

Observation 9e9f466b-31ca-452a-bddb-66650b598c04 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.217919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.217919Z digest=sha256:d2f4ef2e67cb21b2c47123f213016270b3ddd368d8df82846c4dfdea3452f4ec

Observation 449bc022-0803-4ee7-a5f8-0853ac5c1422 · outbound

This paper cites , booktitle =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models , booktitle =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.222227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.222227Z digest=sha256:06befa989d767f9ee04f086cb2cede9a1791fbfc5979fe55fa0c6c835fde3183

Observation 9402f2e0-8cfe-4004-bd22-d54e179147fc · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Solving math word problems with process- and outcome-based feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.226014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.226014Z digest=sha256:1c210bd0fb9c2980c2723761f81007ed70ad979b84544918e85cc534a0240777

Observation cf01f4ab-af70-47fd-999c-d4e52b0ae2fd · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.229462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.229462Z digest=sha256:bcb1955f728ca010b065fc3dfc100e97f4bfb18595fa12372682fdb855e5ad5c

Observation 0b3253d8-0166-4546-93b5-3bddc95a8dac · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Rewarding Progress: Scaling Automated Process Verifiers for

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.232686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.232686Z digest=sha256:f785271217847a81a617136c9f81ac0dbfc5a4f549f9b380425be744a9d7c9e8

Observation 053bdf03-8ec8-4c7a-9bcc-f397aa2347a0 · outbound

This paper cites 2024 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2024 , eprint =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.237285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.237285Z digest=sha256:f273706a9c0fc92e48e7ea9e93e89d17751690b040fcea8ad6960c8774751411

Observation dfba234b-9852-4766-93de-bdd7310f3e7e · outbound

This paper cites Machine Learning , volume =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Machine Learning , volume =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.241163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.241163Z digest=sha256:5c08020ba0ef8ed5d3a7fbf3aaca320b14ecbdffe8d33a7b8554e859bf7219c7

Observation 0c6fdf8c-64b4-40e2-a061-c0651c902dd1 · outbound

This paper cites International Conference on Learning Representations , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models International Conference on Learning Representations , year =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.244783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.244783Z digest=sha256:24b6627f539ecbcb8c122ae03cce4836fd6b17393776610f9bb499ae10680aec

Observation 15552ae1-12b0-43bf-bdb1-155e39000a17 · outbound

This paper cites Machine Learning , volume =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Machine Learning , volume =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.248531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.248531Z digest=sha256:871e4bd910a22091417ea2e4c9005f18067d038c8e7a3f0f67b7d1c76e855ed4

Observation c2daa83c-01a0-4451-bf9e-cb6bf2d32177 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.252229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.252229Z digest=sha256:7c11ffef542052d41b190d8880f6d9c9ae1a5d11ac0ed60dcc1016e97ef671a5

Observation f720ca1a-1281-41e2-90e7-3871528697f9 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Distilling the Knowledge in a Neural Network

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.255795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.255795Z digest=sha256:78fa2bc52d0d64372f5d0deee9e9bb73259f3cf02184d04eca031a1844994206

Observation 6227b58b-9615-42d4-819c-f3e6439ce9cd · outbound

This paper cites Proceedings of the Conference on Empirical Methods in Natural Language Processing , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Proceedings of the Conference on Empirical Methods in Natural Language Processing , year =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.259231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.259231Z digest=sha256:f173d909b5d8d18026ca4ab4670c54d006b29a7206fa5a81d9f9981afa4ed551

Observation 65642262-2d28-4179-b206-d993d7c2a956 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.262599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.262599Z digest=sha256:8b464fab9c026f5d661a0938136c8daed3d54c4436ffb832122ab80060a87f5a

Observation 70374d51-c51e-4658-ab42-99885ab85a10 · outbound

This paper cites Proceedings of the International Conference on Artificial Intelligence and Statistics , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Proceedings of the International Conference on Artificial Intelligence and Statistics , year =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.266232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.266232Z digest=sha256:af362dff19f3bdafc2b1c5d1434ebecf181b41aca53e7e1cf8a0397be0fd170f

Observation 9874e403-50a8-4014-aac4-aa1940815ee4 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.269547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.269547Z digest=sha256:90c9aac1feae256fe080a7671af9120fcab85ff52c078db2253fbb268634f857

Observation 071c1af5-d5ed-495a-869f-c18f20be8eb3 · outbound

This paper cites Neural Networks , volume =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Neural Networks , volume =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.273289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.273289Z digest=sha256:30203b598bff2ccfd8dba03b94931108b0e887395a4b62815e2137758de2b190

Observation 4671c1ef-63a1-4416-9fd9-f645f8167aa9 · outbound

This paper cites International Conference on Learning Representations , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models International Conference on Learning Representations , year =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.277474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.277474Z digest=sha256:9e506e60f4331218ba3fbe8c4d7c213943a9213ab69be27b5a1e3d9a86f92284

Observation 8587597b-2499-4693-acd1-2be192138a50 · outbound

This paper cites Robotics: Science and Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Robotics: Science and Systems , year =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.281431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.281431Z digest=sha256:ebdfc77d2a1c345de3bd86069943dcd3dde0705d4026914c37ef79e722cfd206

Observation 42419323-f0a6-4535-b20a-e3e7d48fd639 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.285395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.285395Z digest=sha256:346bdb493c246069ba9601ed7c0590915620e304adbf178a236a56a5643104a3

Observation bd24587f-f8b2-4f33-9e31-5baaa52f35f3 · outbound

This paper cites Neural Computation , volume =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Neural Computation , volume =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.289386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.289386Z digest=sha256:c3869066a6d243d9332cd085f9b776ed49be25419e7ed373d9344013f26fed38

Observation 39876256-6633-4d30-9909-20d932053f00 · outbound

This paper cites Learning Phrase Representations Using.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Learning Phrase Representations Using

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.292577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.292577Z digest=sha256:aeef1725ac2225d72a77f3768d49a2ac93b85f8c15f9ee8b5eddcfd0cfc0d798

Observation ad4b5534-944b-4956-a8e7-9dce4e4c1a7d · outbound

This paper cites AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.296565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.296565Z digest=sha256:c65e5cf9e0124b13e5ecceea889065c623af3a6f513af9993bc307afe244a632

Pith citing papers

No inbound Pith citation observations are available.