Pith. sign in

Paper Citation Record · LEDGER

Weak-to-Strong On-Policy Distillation

As of 22 August 2026, this Paper Citation Record lists 100 of 142 outbound references and 2 inbound Pith citation observations for arXiv:2607.26246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26246 v1

Coverage vector

measured 100 of 142 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:26:27.540391Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:03:18.871173Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T14:28:01.760078Z

Reference resolution

100 of 142 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved96
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a55e3b13-c499-41f3-a9df-7c98c08bb6a7 · outbound

This paper cites 2023 , url =.

Weak-to-Strong On-Policy Distillation 2023 , url =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.117357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.117357Z digest=sha256:e4bb3f29c6ac632abf962bd579f2457a1d98d3ff50ad9d3952e8f3d2b1686cf5

Observation 0b98eabb-a83e-43af-a504-59cfa30af337 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.190883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.190883Z digest=sha256:f34d0cff0fdcd66bebc12091e4bdf7346dcf94d25f9f05003d5e4c93bcb22947

Observation 1c864ef5-01f8-437b-acff-74d61ed00ec4 · outbound

This paper cites Reward Hacking in Rubric-Based Reinforcement Learning.

Weak-to-Strong On-Policy Distillation Reward Hacking in Rubric-Based Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.291280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.291280Z digest=sha256:d3e843fe5c82e5707a57730af1f1e2814063597216e40779abfa06ddf790dca1

Observation 1a908c5d-8948-4ca4-9a96-a873126f5d7a · outbound

This paper cites 2025 , note =.

Weak-to-Strong On-Policy Distillation 2025 , note =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.430305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.430305Z digest=sha256:50ec61e56c015e3dcd8fe4e81eb028036b60d3954bc0fed73b94bfb3449a2b01

Observation 677d4984-feb0-4f21-a993-5c8643224f4f · outbound

This paper cites arXiv preprint arXiv:2602.05125 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.05125 , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.528458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.528458Z digest=sha256:1e220ac0b74bc2721d522e83625ece746aa48b0aa17c934203c6e653ee8aff1d

Observation 006f6611-7d5a-4803-bf55-5ebf1a46730c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Weak-to-Strong On-Policy Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.656759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.656759Z digest=sha256:137a8abdb52e6482987fa32013b917cae7af0ee248d84a1bb411629d8ea0c811

Observation 639f0405-e45b-4a12-a8e3-6223f89ce915 · outbound

This paper cites arXiv preprint arXiv:2602.21628 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.21628 , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.773782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.773782Z digest=sha256:56f104ea4e896f3df52efc34aaa0370c94d24835f048998fe1393adb85c13b7c

Observation 9cda11fd-c187-4577-8691-28bea1c15d20 · outbound

This paper cites Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis.

Weak-to-Strong On-Policy Distillation Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.884060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.884060Z digest=sha256:1bd3cc942978f09554e374b66ffd6cdcd5cc4aa806569b9a925004e9f11f6f97

Observation 0871ec89-fb5b-463e-99cb-29e8befd5f31 · outbound

This paper cites arXiv preprint arXiv:2603.16600 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2603.16600 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.984931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.984931Z digest=sha256:63e2c3b296cdd0b1a37600ac94def708e35d7b6971a5018adf8b13d0bf9ddcb1

Observation da442d1f-fc9d-49ad-8ffc-c454aa514be0 · outbound

This paper cites Visual Preference Optimization with Rubric Rewards.

Weak-to-Strong On-Policy Distillation Visual Preference Optimization with Rubric Rewards

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.096591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.096591Z digest=sha256:1b0fed79103752ec2ee1aa4dc9ec7dbde17cdcbd325b599162c16050a6bd23e2

Observation d429081d-d486-4714-9c52-292a569175ca · outbound

This paper cites AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning.

Weak-to-Strong On-Policy Distillation AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.199370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.199370Z digest=sha256:76bac1a181dea6014fb2030dbb4c08d7c990275e99a18a8a5a158a106baa7990

Observation e4fe1920-783a-4f79-ab11-06a492ee180f · outbound

This paper cites arXiv preprint arXiv:2602.04649 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.04649 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.296021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.296021Z digest=sha256:41c8db6c4e7b84a3fd15a4ed587c1d2dcae286750b47f40eff48f5b8432063fd

Observation 7b15dc30-8edf-48be-a0b3-3e8fa9f599d3 · outbound

This paper cites arXiv preprint arXiv:2602.01511 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.01511 , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.367121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.367121Z digest=sha256:92a5554944272be24b678b55e1180a041575739217180d9042550d850d2b658a

Observation 76b69197-bb43-49de-8bb2-a74ca4509762 · outbound

This paper cites arXiv preprint arXiv:2510.07284 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2510.07284 , year=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.447882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.447882Z digest=sha256:361b7e1d98082263fe7a77cab0cf24012e66697ee9d7a7e5174d17bf108fb6bf

Observation 62685b6b-c711-40df-8f86-94c4f5657f9b · outbound

This paper cites arXiv preprint arXiv:2602.10885 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.10885 , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.497395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.497395Z digest=sha256:a2fbb2fe30b5bcb3aff08d4f0fef4b1329e8ee7b2a280617355e82050034ea51

Observation a03b8643-3e72-4672-ad6a-613268853c67 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Weak-to-Strong On-Policy Distillation DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.576394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.576394Z digest=sha256:92eb19efa7f2b9321c20aeee7a7a51a89fb8599a49d6a5753e7f65651ce36e83

Observation 299cef2e-f7a1-4c40-a823-698df32a7ad4 · outbound

This paper cites arXiv preprint arXiv:2508.16949 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2508.16949 , year=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.649501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.649501Z digest=sha256:f6221fe4183dc0f0d36ec155b39fd3280eddf3c057cb1ef538d3b2c75b1db796

Observation 39a7433e-7c50-4e95-97dd-de4d767ba954 · outbound

This paper cites arXiv e-prints , pages=.

Weak-to-Strong On-Policy Distillation arXiv e-prints , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.725825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.725825Z digest=sha256:3a74ab0278b42d4d8424c42fc2a65f6db914f7bf7d0aa6fba2b6861469c3ca29

Observation 8923774d-4bab-4dfa-a16f-ae65f4622edc · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Weak-to-Strong On-Policy Distillation Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.809547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.809547Z digest=sha256:cc807505574a3d45b21780b057b3972a6a6f0bc43dee2c33eb5e3498aad58639

Observation ec8c4125-4bb0-43ba-ae08-fcd8b4bf7ad1 · outbound

This paper cites arXiv preprint arXiv:2510.07743 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2510.07743 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.920466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.920466Z digest=sha256:528f6b33ad848383743801c7a1360828e17762ec5a39ccd952104659657b38a1

Observation e3000590-d0fd-4309-b7fb-988c6e303879 · outbound

This paper cites International Conference on Learning Representations , volume=.

Weak-to-Strong On-Policy Distillation International Conference on Learning Representations , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.997032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.997032Z digest=sha256:08e3d3d372e59fd08572bcd0eb10b69a8918e2ff31a1f1289a43933bdada1561

Observation ffec460e-c4e1-4ddc-bfb3-2b4b6d1f12cd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:17.072636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:17.072636Z digest=sha256:b45cdd6eb63f1dbb06b3d7f9bcd0463752578c323f7033016020e61002b8f6b2

Observation 65e970fa-7d0a-46fb-b353-40a2e1e26338 · outbound

This paper cites arXiv preprint arXiv:2512.20061 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2512.20061 , year=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:17.155022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:17.155022Z digest=sha256:bb517b73f20914a29bfe30a8308b97221bac2dd699bcfdf4b500bfa06db24143

Observation f6a70df5-0642-4629-8189-8893c9b177de · outbound

This paper cites Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation.

Weak-to-Strong On-Policy Distillation Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:17.251886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:17.251886Z digest=sha256:78a5f446a4f356257afa7e0987bfb32520cec71ae3894e4399f5d28c45518b44

Observation 1be3d6c8-d0e9-4690-9568-74779c25ffba · outbound

This paper cites Supervising strong learners by amplifying weak experts.

Weak-to-Strong On-Policy Distillation Supervising strong learners by amplifying weak experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:17.544081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:17.544081Z digest=sha256:972fc9ca6e26aa302525b75337e4e0479a4b04ff3a0580a1741fd3571f32c7de

Observation db40f63f-2d42-4e8f-93b3-12156ddc3a35 · outbound

This paper cites Qwen3 Technical Report.

Weak-to-Strong On-Policy Distillation Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:18.074258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:18.074258Z digest=sha256:261ea8b978bcf8ea49897631fea2e0ac08f3ffaba59da3c8cb7a5e70935f74a1

Observation 69862301-f537-479d-aada-98241c185954 · outbound

This paper cites International Conference on Learning Representations , volume=.

Weak-to-Strong On-Policy Distillation International Conference on Learning Representations , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:19.591846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:19.591846Z digest=sha256:6eff72eac5df5291c7e7df2f03686b656a15005f9194e0ab77733f93feef8019

Observation db04f71e-3553-4fa2-b3bc-94c389c63824 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:19.991996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:19.991996Z digest=sha256:4f8d543a5b6c9d99056b0af660775a4a0dd72ab0c64a152fe7a09559fc4630dc

Observation 9464bd17-9fdc-4dbd-9619-35e7de9a558e · outbound

This paper cites Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding.

Weak-to-Strong On-Policy Distillation Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.076885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.076885Z digest=sha256:4b7a2c845afc10e7cb28adc4b38f3c0cbbd0b52c7bd85dd02641d6f795272fa2

Observation 84301b1e-46b5-4694-b6b0-7c96d67735c7 · outbound

This paper cites Visual-Advantage On-Policy Distillation for Vision-Language Models.

Weak-to-Strong On-Policy Distillation Visual-Advantage On-Policy Distillation for Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.175720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.175720Z digest=sha256:1e0677cd204bad8de479f31034f615048abb824dc13ccea1fc09970b11be29e2

Observation 36264d60-5706-4ffc-bb4f-46b3862198b3 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Weak-to-Strong On-Policy Distillation Entropy-Aware On-Policy Distillation of Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.229686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.229686Z digest=sha256:df73b90543901a314ba36962496aa83bb382c90cf915d5a7acb6a09937fc1bd5

Observation b8ea40b1-dbf9-4830-9e99-001e47010cd4 · outbound

This paper cites arXiv preprint arXiv:2603.11137 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2603.11137 , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.302842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.302842Z digest=sha256:f2c6cfa0ad3e5aeea3dcff8e3fac7680d6e02775cbc41ab7fd63c53309d9fc42

Observation 632eaca8-dd22-4823-8081-9582477f93a5 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2026 , pages=.

Weak-to-Strong On-Policy Distillation Findings of the Association for Computational Linguistics: ACL 2026 , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.375260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.375260Z digest=sha256:0e44116ffde300de277e8098c8cc06e21882c2f6399c8c251108070ce1b703ee

Observation 0e964ad5-61da-4ce8-aee5-a9cf713c83b8 · outbound

This paper cites Trust Region On-Policy Distillation.

Weak-to-Strong On-Policy Distillation Trust Region On-Policy Distillation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.424327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.424327Z digest=sha256:b7ab5b1d6d06f40bb916b5666643b9ad6f3711196943faef768138774cd4c11b

Observation 51a7aace-0dae-4ea7-9aee-d81e65b0b1f9 · outbound

This paper cites On-Policy Context Distillation for Language Models.

Weak-to-Strong On-Policy Distillation On-Policy Context Distillation for Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.532339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.532339Z digest=sha256:8d583359842078b204c3e3f256ea6eaf51ee647faf3bff3bca64ffaaed616ee6

Observation 494cfeca-5b07-415c-a35c-f22e57f58bc9 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Weak-to-Strong On-Policy Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.620682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.620682Z digest=sha256:0fc8de8c38c90697ed63b533c8b4239b0dfcdffdfaca54e3724cc94cb14c478e

Observation 28370f0e-353e-4643-8770-11392d2ec6a7 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Weak-to-Strong On-Policy Distillation Self-Distillation Enables Continual Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.684529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.684529Z digest=sha256:e9e233291500da24ddb529210aec6338e38ab9be9fcc8d9e657951197257142a

Observation 2b7daaf2-3ecf-42ef-a436-68fdfaad4df4 · outbound

This paper cites Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation.

Weak-to-Strong On-Policy Distillation Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.783962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.783962Z digest=sha256:b6099052d53038f7957d40b08cf72ea7c6b654135c2737b4f9140434d7d32acd

Observation 209dfac0-a02f-4367-86b5-db31710fbe08 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Weak-to-Strong On-Policy Distillation Reinforcement Learning via Self-Distillation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.893238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.893238Z digest=sha256:7cc49776400904ab7591dd158b78610fdccda63a5828983b2447d8e02ffc9512

Observation 636c4765-8586-48d5-9811-753e16541890 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Weak-to-Strong On-Policy Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.970441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.970441Z digest=sha256:febd40c37d63c56192823104e22bb54fd03997d8d65cf3f482936ce010dc3b48

Observation 0f27cec9-87df-4c2f-afc4-bdaf1c935fa4 · outbound

This paper cites Hybrid Policy Distillation for LLMs.

Weak-to-Strong On-Policy Distillation Hybrid Policy Distillation for LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.077226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.077226Z digest=sha256:459f934ae35ba86def35882999bba138266b0a3e1b4920023a835bba2fa4728e

Observation e76c0b05-541f-4199-9896-e5336725ddf1 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

Weak-to-Strong On-Policy Distillation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.232534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.232534Z digest=sha256:0409b82158c466e6aef3eaedc084f06957b697cba857346e8fb76a2bc7112bfd

Observation 2acfd7d2-a822-474e-a4f4-f0410f2bead1 · outbound

This paper cites Weak-to-Strong Generalization via Direct On-Policy Distillation.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.345094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.345094Z digest=sha256:17d9bf07f817a6494906d8efb2001ae0006b629ef05a982bfc7650c2763729f7

Observation 18a9142f-aa88-46cb-b5be-814ee3dece4b · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Weak-to-Strong On-Policy Distillation RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.484742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.484742Z digest=sha256:8d7c92f8a2e9c1c47b0931716ce6fdf1da2d92b60fd2b71715ad6cb18e1aa5a4

Observation 78f458b4-ce1c-4dcb-a495-b0aa78de20cb · outbound

This paper cites Self-Distilled RLVR.

Weak-to-Strong On-Policy Distillation Self-Distilled RLVR

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.560326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.560326Z digest=sha256:8d30482b630e9e45d33e13a27246bb8024284f672b2b86f62bb3e7b4c09c16fd

Observation 0d2af395-005c-485c-b510-1ae3584bd838 · outbound

This paper cites 2014 , publisher=.

Weak-to-Strong On-Policy Distillation 2014 , publisher=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.657173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.657173Z digest=sha256:00f0c04d8d858e1172e994490b1d823a32cdbc002e807893d5460c442f40fb05

Observation 06afe223-7a1c-4bcb-8860-6598bedfd0d8 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.730787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.730787Z digest=sha256:45fa596d7751910a5d6ef83c63f1a3ed12adcbb3afd57d2c72aa868f9fffc2ee

Observation c85603ec-c901-468a-b644-c5089c74f37c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.839566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.839566Z digest=sha256:65bf87757b457f4a96ca451ce3e0c9698104707969e3ec699c72c862726f48f5

Observation 3a33cd26-26fc-420b-ba7f-df9a03ac50b0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.941002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.941002Z digest=sha256:2934c130aed455c30742f6a79a0ff1a928d8175083ed4a0b28d90abfc772d143

Observation 07e9f883-f5a4-431d-861a-15f32f0390ec · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=.

Weak-to-Strong On-Policy Distillation Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.036112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.036112Z digest=sha256:2aed349e555b9de7b35d2e02d2853cb167e117d94d96a020f1d6ae899160e01b

Observation 93ec4a7f-4515-41f3-8d9c-571389d6a9e9 · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.105162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.105162Z digest=sha256:74d5085aa88609d518bcd13ab089a918d29b3decbbb2a9bfd6eeebc101af7d0c

Observation 43156568-3c94-45cc-a0bd-a5ac98e111e7 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.214321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.214321Z digest=sha256:f279a67c98c5697aa897fe67313e678c725047bfe2c72b7c8cb08f2556b8a65b

Observation 253f5438-7de0-4c24-990d-9fdbf66d7a7f · outbound

This paper cites CoCon: A Self-Supervised Approach for Controlled Text Generation.

Weak-to-Strong On-Policy Distillation CoCon: A Self-Supervised Approach for Controlled Text Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.290413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.290413Z digest=sha256:33f8471cb9fdd07318909ddac738e11f452659db56dd0154ff5594edb4b92c2c

Observation e935d567-d00c-4e21-889d-501110686e46 · outbound

This paper cites CTRL: A Conditional Transformer Language Model for Controllable Generation.

Weak-to-Strong On-Policy Distillation CTRL: A Conditional Transformer Language Model for Controllable Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.360205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.360205Z digest=sha256:0a42f2651e8b6a64d66c99efb6eeb5aa3d9c389bb33f90f592b9d3bef186f9b1

Observation 5ee0fe25-c22e-4f53-bd3f-919500e7e411 · outbound

This paper cites Proceedings of the 28th International Conference on Computational Linguistics , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 28th International Conference on Computational Linguistics , pages=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.448551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.448551Z digest=sha256:2a7b9fd0d2498801c64153d190a3252ea4c271e0e4247a58f8f739484ea5c333

Observation bb56fba7-daac-4730-8568-69fe9a3d89f2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.563314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.563314Z digest=sha256:7b3449a83f33edc31a0c87e6c399ff40438fa76f2e303f4d0d4d97c9380fd39f

Observation dac88f3f-3967-4195-aa21-c0bc85aeca8a · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

Weak-to-Strong On-Policy Distillation MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.644318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.644318Z digest=sha256:deb810e4539dedc1f64e3480557a7b4a4563c20438ab89ab756b63e9db9c94ea

Observation 2f99256b-e251-4b69-a2ac-3f17cd92ca48 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Weak-to-Strong On-Policy Distillation GLM-5: from Vibe Coding to Agentic Engineering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.763779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.763779Z digest=sha256:5bf757b8fec6993cbaae56a058635760874655b611126a84aff15b36e6c53d14

Observation 157ed573-e4cd-4d1b-889f-721a05705655 · outbound

This paper cites arXiv preprint arXiv:2606.19348 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2606.19348 , year=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.869845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.869845Z digest=sha256:9e06e7edb4d190b5832ed4cfe364b884e28888ffde64ae5370afd847cb6ebed9

Observation 5827e367-6e79-401a-a52b-05cd9481865a · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

Weak-to-Strong On-Policy Distillation Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.018580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.018580Z digest=sha256:fd4c826c521e54cb9057ca3f360bc561209e03bd16e86730304eb3d1c1757925

Observation 1c61612d-ee26-4b4a-8e00-aa7796aa28de · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Weak-to-Strong On-Policy Distillation GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.123274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.123274Z digest=sha256:7621e7d1d1fa4e8d048c9624062ebae69f950cb60205fd5bb7d151227af3d1ed

Observation b2002792-b551-48b3-bd2e-d08f571dd581 · outbound

This paper cites an unresolved cited work.

Weak-to-Strong On-Policy Distillation Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.228289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.228289Z digest=sha256:730bbe38b5e53b8d809775828d78d1610662873a0454169572f3281ebd9f8340

Observation 8d2eae3d-df45-4cb0-b8b3-41f5bfc8df15 · outbound

This paper cites Tuning Language Models by Proxy.

Weak-to-Strong On-Policy Distillation Tuning Language Models by Proxy

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.349392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.349392Z digest=sha256:fb650c730481530cdd9c0e90684eca1e49c00d188b0bf2ec4f62ede3db600172

Observation 089f2637-e4c5-4abe-8244-3183a9e4b8ef · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.485869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.485869Z digest=sha256:0dcbcbe44db435e389d5244402ba60e779a0583b36aa72290a6076976e54c738

Observation 9d36fda2-fe7c-4062-a652-e9c3230f42a5 · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

Weak-to-Strong On-Policy Distillation A Survey of On-Policy Distillation for Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.639872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.639872Z digest=sha256:7419769b838096d7620402832242aa9fae4ce3347c84abbd9ec871e0c7ae408c

Observation 47529e3a-b68b-475c-a2c1-45808eaf61f4 · outbound

This paper cites Weak-to-Strong Jailbreaking on Large Language Models.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Jailbreaking on Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.761661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.761661Z digest=sha256:857aba76e9d425a6ba7aa10d017dec3719f4b96e8466dfb8f4694e12187727e5

Observation 758799fa-3759-42e2-924b-4af9fd0f8e79 · outbound

This paper cites MiMo-V2-Flash Technical Report.

Weak-to-Strong On-Policy Distillation MiMo-V2-Flash Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.854870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.854870Z digest=sha256:48a9a247eded21039850d7abdc67ec31c1b7d41b9d0c47ac307bf0e503e01c2f

Observation 0481d0ab-0c81-4ada-a26a-0a30a6c5ed41 · outbound

This paper cites Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.981367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.981367Z digest=sha256:be2f4c7611398e4b305c4cec26bfc7fd2b4d1ae95443ab30746f263d1656de36

Observation 6c4c09e2-dd65-4459-9260-675426845d9d · outbound

This paper cites On Weak-to-Strong Generalization and f-Divergence.

Weak-to-Strong On-Policy Distillation On Weak-to-Strong Generalization and f-Divergence

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.105193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.105193Z digest=sha256:a44db4ec5189ccb2de26f964a6bda106fa464c786bbe70506c1e7676a532350a

Observation a1f25405-7ce8-4f61-bde8-9050768d3c59 · outbound

This paper cites OPRD: On-Policy Representation Distillation.

Weak-to-Strong On-Policy Distillation OPRD: On-Policy Representation Distillation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.196739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.196739Z digest=sha256:cbf7e47ac7770b2905d7d3e031884c0b476ef886d5f0073ff0ff9e0921c2dbaf

Observation a09903bf-b5fe-4800-a220-d29b7586c4b5 · outbound

This paper cites International Conference on Learning Representations , volume=.

Weak-to-Strong On-Policy Distillation International Conference on Learning Representations , volume=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.281183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.281183Z digest=sha256:7c7caa9279344b2f4dad15de12fe3f5d692ca676d7157db93acc0f75215d0ea3

Observation c823c9a8-9ad6-4bfc-90ff-6d8e63403ce3 · outbound

This paper cites forward KL , author=.

Weak-to-Strong On-Policy Distillation forward KL , author=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.364454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.364454Z digest=sha256:63a93b398e6ca1ccdb3aac8f6a2e4625031ebdcdac00097dca9813393c4c4935

Observation 0520c4a3-0016-4277-8f82-b08a4d8420d7 · outbound

This paper cites Thinking Machines Lab: Connectionism , year =.

Weak-to-Strong On-Policy Distillation Thinking Machines Lab: Connectionism , year =

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.506888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.506888Z digest=sha256:0e23af4ac5c82565a68b53ec49bd8c199b1030767ce0a07cf48ed12dcab41918

Observation 16c884cc-ed3e-4ddd-82b3-af3c5fff7bfc · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Weak-to-Strong On-Policy Distillation Process Reinforcement through Implicit Rewards

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.653402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.653402Z digest=sha256:0ca83dc5033d6ee86e8dac37fece6bca0eb6220a8ca6e22d1d915e42a4f51bf3

Observation 88f7e6cc-8d1e-4b30-af61-8d381d919c6b · outbound

This paper cites an unresolved cited work.

Weak-to-Strong On-Policy Distillation Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.768775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.768775Z digest=sha256:2f140c41ea3666d5274cbc5460b00c363f923222594006db08ea5d0453822972

Observation 3c36225b-6de3-4a63-8363-09d4bb8883ad · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Weak-to-Strong On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.869605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.869605Z digest=sha256:fdd72c4359606c4c0cd69a0424fb380e112c1688b949b81edcba12cbd2d4565e

Observation a3545e34-3207-4116-a722-9fa81021fc4a · outbound

This paper cites an unresolved cited work.

Weak-to-Strong On-Policy Distillation Unresolved cited work

Reference 77

Resolution
parse uncertain
no resolver link, observed 2026-08-01T00:26:24.970768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.970768Z digest=sha256:a49c1e07dd6b16766df44bec0207402ce1174d390eab1954942af3ccd097ee9d

Observation 8ad81f8e-c470-4db2-ab72-2650a284f41f · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Weak-to-Strong On-Policy Distillation DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.081268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.081268Z digest=sha256:cffde4529912b559ed92ee425408fdde454ee9628dcd22c305986d5051fdfcef

Observation 4f4b306e-70d4-4ba6-b213-8a912aec2beb · outbound

This paper cites International Conference on Learning Representations , volume=.

Weak-to-Strong On-Policy Distillation International Conference on Learning Representations , volume=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.169140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.169140Z digest=sha256:11046456c1f1fce6250bedd71771feb174d65264fda692abc5082bad0c8c0272

Observation 8ec05144-8310-486f-87ca-4951390c544d · outbound

This paper cites URL https://matharena.

Weak-to-Strong On-Policy Distillation URL https://matharena

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.288499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.288499Z digest=sha256:12fc28ddf7363e6b7c3fc5ffd4af6b23a0f75ee3d53caa144e8bb4728758f793

Observation c7007b13-9e96-4e96-8c59-eaa8731a194d · outbound

This paper cites Advances in neural information processing systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in neural information processing systems , volume=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.385382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.385382Z digest=sha256:62e4dea0245ee128e779da25ee2ad3f4febfa4b77bd31f0afd4163179aff51c9

Observation fc9af86c-2049-46c8-8961-415335480d90 · outbound

This paper cites Kimi-Audio Technical Report.

Weak-to-Strong On-Policy Distillation Kimi-Audio Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.522384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.522384Z digest=sha256:736152d268d127791528d0bc677f87d2ff0ace1d221bebcaf150ef6d4f080a90

Observation 9083144a-06bf-4bbd-8be3-d5e6c2285e44 · outbound

This paper cites 2023 , url =.

Weak-to-Strong On-Policy Distillation 2023 , url =

Reference 83

Resolution
verified exact
doi, observed 2026-08-01T00:31:07.749637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-01T00:26:25.667153Z digest=sha256:e84488a8455ff99c6270112f3bf65ee08a761837478d3aae905dae2786c59fb0

Observation c6893a8c-bff4-45e4-9398-5ea127d52293 · outbound

This paper cites Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities.

Weak-to-Strong On-Policy Distillation Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.777254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.777254Z digest=sha256:2edd777305bc8243b40ed83eaabba1a02ff91b92f1aab7e2c0de06ce8e0bc392

Observation eb2323cc-1287-40c3-8a37-80c2e4629719 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Weak-to-Strong On-Policy Distillation Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.893734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.893734Z digest=sha256:9c87373e3c9c992027eab828728c55cc39d787cd88018a8580c271ef54dba5c7

Observation 79532ab5-4cd3-4c0d-bd19-67e062b0a829 · outbound

This paper cites Proceedings of the 30th ACM International Conference on Multimedia , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 30th ACM International Conference on Multimedia , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.023285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.023285Z digest=sha256:ed1db48b43c6434871b2fcee0e56b33265797e7bd65e8e744cf328fcb58f70c7

Observation eb25a9cb-433e-4cab-8feb-472ff4b34a52 · outbound

This paper cites GPT-4o System Card.

Weak-to-Strong On-Policy Distillation GPT-4o System Card

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.156814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.156814Z digest=sha256:1188fd8dc1f81d6b338ad85b24acbf88f7248f9a6387a6e9ca0164cdb0c8e0e3

Observation b5180e38-2847-4134-b0f5-909d126e32f3 · outbound

This paper cites GPT-4 Technical Report.

Weak-to-Strong On-Policy Distillation GPT-4 Technical Report

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.284433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.284433Z digest=sha256:0378072f21cfb5041c4d8681dbae4da171410876a2389430cd94c41c0293076c

Observation 3da63c6d-a74d-4676-b508-94c041a0e928 · outbound

This paper cites Wav2CLIP: Learning Robust Audio Representations from Clip , booktitle =.

Weak-to-Strong On-Policy Distillation Wav2CLIP: Learning Robust Audio Representations from Clip , booktitle =

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.410915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.410915Z digest=sha256:b11b6b04edb1b0f3eddd34ea6c089325de16ebbf37e658ff2dc6da7bffa0c241

Observation 5e3257b0-76a2-4728-9ec5-65a4a13295e1 · outbound

This paper cites Pengi: An Audio Language Model for Audio Tasks , booktitle =.

Weak-to-Strong On-Policy Distillation Pengi: An Audio Language Model for Audio Tasks , booktitle =

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.473503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.473503Z digest=sha256:ac2f8806c58f1bd03ffba49b74d0682b50150d2f9b3e8728da4a19836e4c06df

Observation b947c43a-0fcb-454a-8c9c-301e68ef6763 · outbound

This paper cites Liu and Leonid Karlinsky and James R.

Weak-to-Strong On-Policy Distillation Liu and Leonid Karlinsky and James R

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.552685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.552685Z digest=sha256:6910dc7825e6c1d09ba5ad036172c16d5dab024c994628bf39e644306f65bcb5

Observation 6e396538-fabf-40f1-81ca-84b80593d622 · outbound

This paper cites Optimal Transport for Treatment Effect Estimation , booktitle =.

Weak-to-Strong On-Policy Distillation Optimal Transport for Treatment Effect Estimation , booktitle =

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.585832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.585832Z digest=sha256:e547cdf52469d5dd7acf3475af1103fb187089a6423527967eb1e6970816baee

Observation 13963a24-32b3-43d2-8240-2305c8148ca8 · outbound

This paper cites A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions.

Weak-to-Strong On-Policy Distillation A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.646865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.646865Z digest=sha256:24878d432c623cdcc7d6f9e6ba62ab37438eebf151a98ed63523d95eba312054

Observation 7d5a76d6-8cb4-428a-b125-2269d5616a53 · outbound

This paper cites Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection , booktitle =.

Weak-to-Strong On-Policy Distillation Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection , booktitle =

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.784799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.784799Z digest=sha256:070a1e16e348d84c3ebf85dbc2308e6f810bb603165d549c030aa4a0a52e125f

Observation b78ee796-4202-4f8f-b32e-934422834d0b · outbound

This paper cites Generalized Data Distribution Iteration , booktitle =.

Weak-to-Strong On-Policy Distillation Generalized Data Distribution Iteration , booktitle =

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.902413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.902413Z digest=sha256:3a419861e48906c00ec9c62fb87455a4913b5461e502468d45616e9aee82c675

Observation 1d5146a5-1679-4da3-a8f8-e126d4c3fccd · outbound

This paper cites GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning.

Weak-to-Strong On-Policy Distillation GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:27.063018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:27.063018Z digest=sha256:74717aee3ba3903c53c1a02003f135b20e37e108847eae24174b39120c796d96

Observation aa07a1eb-f80c-46bc-8e96-57f92e94889b · outbound

This paper cites CoRR , volume =.

Weak-to-Strong On-Policy Distillation CoRR , volume =

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:27.202175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:27.202175Z digest=sha256:f931f1bc28d23d53981fedaf2991a09b842afb3ad1cb498a50b6b4bc35b29b95

Observation 687fa503-d69e-4af9-bc57-3db573a602a2 · outbound

This paper cites ConvFormer: Revisiting Transformer for Sequential User Modeling.

Weak-to-Strong On-Policy Distillation ConvFormer: Revisiting Transformer for Sequential User Modeling

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-08-01T00:31:07.611462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-01T00:26:27.325656Z digest=sha256:9103b7710a076739549863de4394209a398d9304614a216d37b7211f9920005e

Observation e8c97fdf-3f00-4a8d-8e9a-aa05912b2393 · outbound

This paper cites An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning.

Weak-to-Strong On-Policy Distillation An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:27.446030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:27.446030Z digest=sha256:33135f3afe96ab9774559e80e259873ae151cea7e7a1e8cc04773856ef71ba34

Observation 6e213f62-0170-4e5d-a5de-2178c5515c25 · outbound

This paper cites CoRR , volume =.

Weak-to-Strong On-Policy Distillation CoRR , volume =

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-08-01T00:31:07.518318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-01T00:26:27.540391Z digest=sha256:6f52c5ee50ddcd3e6d58e1b2d30aee7e9ddbc59014c1d68568c065c3335d7562

Pith citing papers

Observation 53f5c82c-5342-4939-976b-994ae60b48b5 · inbound

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning cites this paper.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Weak-to-Strong On-Policy Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:18.871173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:18.871173Z digest=sha256:b70908941f4b1935c88f7fab515674e3771dc17ea7f6222242a29d406686853a

Observation 5ce7d0d3-b555-4097-bb76-272c080c0820 · inbound

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training cites this paper.

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training Weak-to-Strong On-Policy Distillation

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:28:01.767121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T14:28:01.605540Z digest=sha256:8757cf743e975826eb4a13a25f69d64d6a26d552e1e15f20fc5a5278cf1fc042