Pith. sign in

Paper Citation Record · LEDGER

Weak-to-Strong On-Policy Distillation

As of 12 August 2026, this Paper Citation Record lists 100 of 142 outbound references and 0 inbound Pith citation observations for arXiv:2607.26246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26246 v1

Coverage vector

measured 100 of 142 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:26:27.540391Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 142 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved96
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a55e3b13-c499-41f3-a9df-7c98c08bb6a7 · outbound

This paper cites 2023 , url =.

Weak-to-Strong On-Policy Distillation 2023 , url =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.117357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.117357Z digest=sha256:e4bb3f29c6ac632abf962bd579f2457a1d98d3ff50ad9d3952e8f3d2b1686cf5

Observation 0b98eabb-a83e-43af-a504-59cfa30af337 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.190883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.190883Z digest=sha256:f34d0cff0fdcd66bebc12091e4bdf7346dcf94d25f9f05003d5e4c93bcb22947

Observation 1c864ef5-01f8-437b-acff-74d61ed00ec4 · outbound

This paper cites Reward Hacking in Rubric-Based Reinforcement Learning.

Weak-to-Strong On-Policy Distillation Reward Hacking in Rubric-Based Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.291280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.291280Z digest=sha256:d3e843fe5c82e5707a57730af1f1e2814063597216e40779abfa06ddf790dca1

Observation 1a908c5d-8948-4ca4-9a96-a873126f5d7a · outbound

This paper cites 2025 , note =.

Weak-to-Strong On-Policy Distillation 2025 , note =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.430305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.430305Z digest=sha256:50ec61e56c015e3dcd8fe4e81eb028036b60d3954bc0fed73b94bfb3449a2b01

Observation 677d4984-feb0-4f21-a993-5c8643224f4f · outbound

This paper cites arXiv preprint arXiv:2602.05125 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.05125 , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.528458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.528458Z digest=sha256:1e220ac0b74bc2721d522e83625ece746aa48b0aa17c934203c6e653ee8aff1d

Observation 006f6611-7d5a-4803-bf55-5ebf1a46730c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Weak-to-Strong On-Policy Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.656759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.656759Z digest=sha256:3ca6e3bc1c018af902858c8407b7d62341c31205885dc1985f568af299a89cbb

Observation 639f0405-e45b-4a12-a8e3-6223f89ce915 · outbound

This paper cites arXiv preprint arXiv:2602.21628 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.21628 , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.773782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.773782Z digest=sha256:56f104ea4e896f3df52efc34aaa0370c94d24835f048998fe1393adb85c13b7c

Observation 9cda11fd-c187-4577-8691-28bea1c15d20 · outbound

This paper cites Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis.

Weak-to-Strong On-Policy Distillation Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.884060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.884060Z digest=sha256:851b88669ca11db013be3ec8265bfbbea2bd9772d3f7176e67dcdee7d371dbfa

Observation 0871ec89-fb5b-463e-99cb-29e8befd5f31 · outbound

This paper cites arXiv preprint arXiv:2603.16600 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2603.16600 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:15.984931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:15.984931Z digest=sha256:63e2c3b296cdd0b1a37600ac94def708e35d7b6971a5018adf8b13d0bf9ddcb1

Observation da442d1f-fc9d-49ad-8ffc-c454aa514be0 · outbound

This paper cites Visual Preference Optimization with Rubric Rewards.

Weak-to-Strong On-Policy Distillation Visual Preference Optimization with Rubric Rewards

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.096591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.096591Z digest=sha256:b579c1b9197563b965f8877085e9692b50e0b205ef438215b2edc2b4b683d688

Observation d429081d-d486-4714-9c52-292a569175ca · outbound

This paper cites AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning.

Weak-to-Strong On-Policy Distillation AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.199370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.199370Z digest=sha256:5bce3780c997e4196f11da00cbed96e7a5a90918d1d9f75e8e4c2b9f65e27da2

Observation e4fe1920-783a-4f79-ab11-06a492ee180f · outbound

This paper cites arXiv preprint arXiv:2602.04649 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.04649 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.296021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.296021Z digest=sha256:41c8db6c4e7b84a3fd15a4ed587c1d2dcae286750b47f40eff48f5b8432063fd

Observation 7b15dc30-8edf-48be-a0b3-3e8fa9f599d3 · outbound

This paper cites arXiv preprint arXiv:2602.01511 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.01511 , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.367121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.367121Z digest=sha256:92a5554944272be24b678b55e1180a041575739217180d9042550d850d2b658a

Observation 76b69197-bb43-49de-8bb2-a74ca4509762 · outbound

This paper cites arXiv preprint arXiv:2510.07284 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2510.07284 , year=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.447882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.447882Z digest=sha256:361b7e1d98082263fe7a77cab0cf24012e66697ee9d7a7e5174d17bf108fb6bf

Observation 62685b6b-c711-40df-8f86-94c4f5657f9b · outbound

This paper cites arXiv preprint arXiv:2602.10885 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2602.10885 , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.497395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.497395Z digest=sha256:a2fbb2fe30b5bcb3aff08d4f0fef4b1329e8ee7b2a280617355e82050034ea51

Observation a03b8643-3e72-4672-ad6a-613268853c67 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Weak-to-Strong On-Policy Distillation DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.576394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.576394Z digest=sha256:84a125b52114f174119f20481d4e16123b05668160106a4fa579ff21c7fdaeb2

Observation 299cef2e-f7a1-4c40-a823-698df32a7ad4 · outbound

This paper cites arXiv preprint arXiv:2508.16949 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2508.16949 , year=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.649501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.649501Z digest=sha256:f6221fe4183dc0f0d36ec155b39fd3280eddf3c057cb1ef538d3b2c75b1db796

Observation 39a7433e-7c50-4e95-97dd-de4d767ba954 · outbound

This paper cites arXiv e-prints , pages=.

Weak-to-Strong On-Policy Distillation arXiv e-prints , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.725825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.725825Z digest=sha256:3a74ab0278b42d4d8424c42fc2a65f6db914f7bf7d0aa6fba2b6861469c3ca29

Observation 8923774d-4bab-4dfa-a16f-ae65f4622edc · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Weak-to-Strong On-Policy Distillation Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.809547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.809547Z digest=sha256:cc807505574a3d45b21780b057b3972a6a6f0bc43dee2c33eb5e3498aad58639

Observation ec8c4125-4bb0-43ba-ae08-fcd8b4bf7ad1 · outbound

This paper cites arXiv preprint arXiv:2510.07743 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2510.07743 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.920466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.920466Z digest=sha256:528f6b33ad848383743801c7a1360828e17762ec5a39ccd952104659657b38a1

Observation e3000590-d0fd-4309-b7fb-988c6e303879 · outbound

This paper cites International Conference on Learning Representations , volume=.

Weak-to-Strong On-Policy Distillation International Conference on Learning Representations , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:16.997032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:16.997032Z digest=sha256:08e3d3d372e59fd08572bcd0eb10b69a8918e2ff31a1f1289a43933bdada1561

Observation ffec460e-c4e1-4ddc-bfb3-2b4b6d1f12cd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:17.072636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:17.072636Z digest=sha256:b45cdd6eb63f1dbb06b3d7f9bcd0463752578c323f7033016020e61002b8f6b2

Observation 65e970fa-7d0a-46fb-b353-40a2e1e26338 · outbound

This paper cites arXiv preprint arXiv:2512.20061 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2512.20061 , year=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:17.155022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:17.155022Z digest=sha256:bb517b73f20914a29bfe30a8308b97221bac2dd699bcfdf4b500bfa06db24143

Observation f6a70df5-0642-4629-8189-8893c9b177de · outbound

This paper cites Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation.

Weak-to-Strong On-Policy Distillation Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:17.251886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:17.251886Z digest=sha256:37e35b03918c3b536050a9ebf6d33a2ad0734e844185e4dcdb9ba5d4de0c19d6

Observation 1be3d6c8-d0e9-4690-9568-74779c25ffba · outbound

This paper cites Supervising strong learners by amplifying weak experts.

Weak-to-Strong On-Policy Distillation Supervising strong learners by amplifying weak experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:17.544081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:17.544081Z digest=sha256:4a071b1ca967ff78a92ab5cf76690fc522caa7e42e745cd0c9cd68af1554a954

Observation db40f63f-2d42-4e8f-93b3-12156ddc3a35 · outbound

This paper cites Qwen3 Technical Report.

Weak-to-Strong On-Policy Distillation Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:18.074258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:18.074258Z digest=sha256:261ea8b978bcf8ea49897631fea2e0ac08f3ffaba59da3c8cb7a5e70935f74a1

Observation 69862301-f537-479d-aada-98241c185954 · outbound

This paper cites International Conference on Learning Representations , volume=.

Weak-to-Strong On-Policy Distillation International Conference on Learning Representations , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:19.591846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:19.591846Z digest=sha256:6eff72eac5df5291c7e7df2f03686b656a15005f9194e0ab77733f93feef8019

Observation db04f71e-3553-4fa2-b3bc-94c389c63824 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:19.991996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:19.991996Z digest=sha256:cf988ff4d3039e78e7a03ea54a5802753fa35ae63e9b64aa5cb74a5069c506f4

Observation 9464bd17-9fdc-4dbd-9619-35e7de9a558e · outbound

This paper cites Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding.

Weak-to-Strong On-Policy Distillation Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.076885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.076885Z digest=sha256:4b7a2c845afc10e7cb28adc4b38f3c0cbbd0b52c7bd85dd02641d6f795272fa2

Observation 84301b1e-46b5-4694-b6b0-7c96d67735c7 · outbound

This paper cites Visual-Advantage On-Policy Distillation for Vision-Language Models.

Weak-to-Strong On-Policy Distillation Visual-Advantage On-Policy Distillation for Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.175720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.175720Z digest=sha256:f97b62b6a227c066ae53a5c568418a50865d737b8fa60ad9e6aa78f20bb31623

Observation 36264d60-5706-4ffc-bb4f-46b3862198b3 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Weak-to-Strong On-Policy Distillation Entropy-Aware On-Policy Distillation of Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.229686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.229686Z digest=sha256:e0fad3cfa3a902f80f692f9d9b802ab46991dc707656aeec17aa74a443be64be

Observation b8ea40b1-dbf9-4830-9e99-001e47010cd4 · outbound

This paper cites arXiv preprint arXiv:2603.11137 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2603.11137 , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.302842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.302842Z digest=sha256:f2c6cfa0ad3e5aeea3dcff8e3fac7680d6e02775cbc41ab7fd63c53309d9fc42

Observation 632eaca8-dd22-4823-8081-9582477f93a5 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2026 , pages=.

Weak-to-Strong On-Policy Distillation Findings of the Association for Computational Linguistics: ACL 2026 , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.375260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.375260Z digest=sha256:0e44116ffde300de277e8098c8cc06e21882c2f6399c8c251108070ce1b703ee

Observation 0e964ad5-61da-4ce8-aee5-a9cf713c83b8 · outbound

This paper cites Trust Region On-Policy Distillation.

Weak-to-Strong On-Policy Distillation Trust Region On-Policy Distillation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.424327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.424327Z digest=sha256:b7ab5b1d6d06f40bb916b5666643b9ad6f3711196943faef768138774cd4c11b

Observation 51a7aace-0dae-4ea7-9aee-d81e65b0b1f9 · outbound

This paper cites On-Policy Context Distillation for Language Models.

Weak-to-Strong On-Policy Distillation On-Policy Context Distillation for Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.532339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.532339Z digest=sha256:c632e8f35c713dca432a0b45c8703e6ddedbfb7d2a2580114882df83055076ce

Observation 494cfeca-5b07-415c-a35c-f22e57f58bc9 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Weak-to-Strong On-Policy Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.620682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.620682Z digest=sha256:d0139ca5ad4d9277bfdcfdfc90e10e470304f34836402d0c31c78fa8f25bfe7c

Observation 28370f0e-353e-4643-8770-11392d2ec6a7 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Weak-to-Strong On-Policy Distillation Self-Distillation Enables Continual Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.684529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.684529Z digest=sha256:f76b16a1482af612ac5999c991e092a4b84dce1b9636fdd5ddd0c36f84408b4e

Observation 2b7daaf2-3ecf-42ef-a436-68fdfaad4df4 · outbound

This paper cites Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation.

Weak-to-Strong On-Policy Distillation Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.783962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.783962Z digest=sha256:0ef40e0da37ddd78f23ba1589fc49595b57315af0053c2f4f5e8bd71eb61b707

Observation 209dfac0-a02f-4367-86b5-db31710fbe08 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Weak-to-Strong On-Policy Distillation Reinforcement Learning via Self-Distillation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.893238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.893238Z digest=sha256:9d748619b78437dd031bf8f171ae377be765b7726bd31e1a224773b6111a0f6c

Observation 636c4765-8586-48d5-9811-753e16541890 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Weak-to-Strong On-Policy Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:20.970441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:20.970441Z digest=sha256:7a0a8979625913ced13ade92b4c49e8ce87e8141b462179fb30ab0c2e702b4a6

Observation 0f27cec9-87df-4c2f-afc4-bdaf1c935fa4 · outbound

This paper cites Hybrid Policy Distillation for LLMs.

Weak-to-Strong On-Policy Distillation Hybrid Policy Distillation for LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.077226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.077226Z digest=sha256:0aa1dc76bff2a4b6a534c5abfb1853078ee3070ab449b4500c246d799ba173ea

Observation e76c0b05-541f-4199-9896-e5336725ddf1 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

Weak-to-Strong On-Policy Distillation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.232534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.232534Z digest=sha256:3f49bdb7bf1a35684579c629e2161199ac8b6b2e95d8a7f4cbb69201f1bb2469

Observation 2acfd7d2-a822-474e-a4f4-f0410f2bead1 · outbound

This paper cites Weak-to-Strong Generalization via Direct On-Policy Distillation.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.345094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.345094Z digest=sha256:17d9bf07f817a6494906d8efb2001ae0006b629ef05a982bfc7650c2763729f7

Observation 18a9142f-aa88-46cb-b5be-814ee3dece4b · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Weak-to-Strong On-Policy Distillation RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.484742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.484742Z digest=sha256:e7f18139a95f8c5988c1b699870a4e9f754fb1c2172e342b4c1028b4176a77d2

Observation 78f458b4-ce1c-4dcb-a495-b0aa78de20cb · outbound

This paper cites Self-Distilled RLVR.

Weak-to-Strong On-Policy Distillation Self-Distilled RLVR

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.560326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.560326Z digest=sha256:8d30482b630e9e45d33e13a27246bb8024284f672b2b86f62bb3e7b4c09c16fd

Observation 0d2af395-005c-485c-b510-1ae3584bd838 · outbound

This paper cites 2014 , publisher=.

Weak-to-Strong On-Policy Distillation 2014 , publisher=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.657173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.657173Z digest=sha256:00f0c04d8d858e1172e994490b1d823a32cdbc002e807893d5460c442f40fb05

Observation 06afe223-7a1c-4bcb-8860-6598bedfd0d8 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.730787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.730787Z digest=sha256:45fa596d7751910a5d6ef83c63f1a3ed12adcbb3afd57d2c72aa868f9fffc2ee

Observation c85603ec-c901-468a-b644-c5089c74f37c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.839566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.839566Z digest=sha256:65bf87757b457f4a96ca451ce3e0c9698104707969e3ec699c72c862726f48f5

Observation 3a33cd26-26fc-420b-ba7f-df9a03ac50b0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.941002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.941002Z digest=sha256:2934c130aed455c30742f6a79a0ff1a928d8175083ed4a0b28d90abfc772d143

Observation 07e9f883-f5a4-431d-861a-15f32f0390ec · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=.

Weak-to-Strong On-Policy Distillation Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.036112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.036112Z digest=sha256:2aed349e555b9de7b35d2e02d2853cb167e117d94d96a020f1d6ae899160e01b

Observation 93ec4a7f-4515-41f3-8d9c-571389d6a9e9 · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.105162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.105162Z digest=sha256:74d5085aa88609d518bcd13ab089a918d29b3decbbb2a9bfd6eeebc101af7d0c

Observation 43156568-3c94-45cc-a0bd-a5ac98e111e7 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.214321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.214321Z digest=sha256:f279a67c98c5697aa897fe67313e678c725047bfe2c72b7c8cb08f2556b8a65b

Observation 253f5438-7de0-4c24-990d-9fdbf66d7a7f · outbound

This paper cites CoCon: A Self-Supervised Approach for Controlled Text Generation.

Weak-to-Strong On-Policy Distillation CoCon: A Self-Supervised Approach for Controlled Text Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.290413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.290413Z digest=sha256:fa563f116a18eaf0af66185ce983383a4e1f23d0e08dc47eace3b9589a540fef

Observation e935d567-d00c-4e21-889d-501110686e46 · outbound

This paper cites CTRL: A Conditional Transformer Language Model for Controllable Generation.

Weak-to-Strong On-Policy Distillation CTRL: A Conditional Transformer Language Model for Controllable Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.360205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.360205Z digest=sha256:9af6ccf3f93a79627cc41c03324729baa90ceae3b2ab597e6dd5d4c8e48a085b

Observation 5ee0fe25-c22e-4f53-bd3f-919500e7e411 · outbound

This paper cites Proceedings of the 28th International Conference on Computational Linguistics , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 28th International Conference on Computational Linguistics , pages=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.448551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.448551Z digest=sha256:2a7b9fd0d2498801c64153d190a3252ea4c271e0e4247a58f8f739484ea5c333

Observation bb56fba7-daac-4730-8568-69fe9a3d89f2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.563314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.563314Z digest=sha256:7b3449a83f33edc31a0c87e6c399ff40438fa76f2e303f4d0d4d97c9380fd39f

Observation dac88f3f-3967-4195-aa21-c0bc85aeca8a · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

Weak-to-Strong On-Policy Distillation MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.644318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.644318Z digest=sha256:deb810e4539dedc1f64e3480557a7b4a4563c20438ab89ab756b63e9db9c94ea

Observation 2f99256b-e251-4b69-a2ac-3f17cd92ca48 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Weak-to-Strong On-Policy Distillation GLM-5: from Vibe Coding to Agentic Engineering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.763779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.763779Z digest=sha256:5bf757b8fec6993cbaae56a058635760874655b611126a84aff15b36e6c53d14

Observation 157ed573-e4cd-4d1b-889f-721a05705655 · outbound

This paper cites arXiv preprint arXiv:2606.19348 , year=.

Weak-to-Strong On-Policy Distillation arXiv preprint arXiv:2606.19348 , year=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:22.869845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:22.869845Z digest=sha256:9e06e7edb4d190b5832ed4cfe364b884e28888ffde64ae5370afd847cb6ebed9

Observation 5827e367-6e79-401a-a52b-05cd9481865a · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

Weak-to-Strong On-Policy Distillation Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.018580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.018580Z digest=sha256:fd4c826c521e54cb9057ca3f360bc561209e03bd16e86730304eb3d1c1757925

Observation 1c61612d-ee26-4b4a-8e00-aa7796aa28de · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Weak-to-Strong On-Policy Distillation GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.123274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.123274Z digest=sha256:023484299f611a8343728d4f35333f87f424b3ae233a0101ce6d93047dc5b004

Observation b2002792-b551-48b3-bd2e-d08f571dd581 · outbound

This paper cites an unresolved cited work.

Weak-to-Strong On-Policy Distillation Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.228289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.228289Z digest=sha256:730bbe38b5e53b8d809775828d78d1610662873a0454169572f3281ebd9f8340

Observation 8d2eae3d-df45-4cb0-b8b3-41f5bfc8df15 · outbound

This paper cites Tuning Language Models by Proxy.

Weak-to-Strong On-Policy Distillation Tuning Language Models by Proxy

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.349392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.349392Z digest=sha256:1485a8ee27384ebcf878bb93230fd4eb7c6d852c242a651f239bc9816c733b17

Observation 089f2637-e4c5-4abe-8244-3183a9e4b8ef · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.485869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.485869Z digest=sha256:0dcbcbe44db435e389d5244402ba60e779a0583b36aa72290a6076976e54c738

Observation 9d36fda2-fe7c-4062-a652-e9c3230f42a5 · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

Weak-to-Strong On-Policy Distillation A Survey of On-Policy Distillation for Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.639872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.639872Z digest=sha256:4ea6b763328420799c2b67bb9fe4a5cfb816581b66497f6fb6c2ce159e47e292

Observation 47529e3a-b68b-475c-a2c1-45808eaf61f4 · outbound

This paper cites Weak-to-Strong Jailbreaking on Large Language Models.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Jailbreaking on Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.761661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.761661Z digest=sha256:c9b4c7d7c778e736d9c84ac56ac74333261bc1f2b84cc9b1a0a6da63a746b589

Observation 758799fa-3759-42e2-924b-4af9fd0f8e79 · outbound

This paper cites MiMo-V2-Flash Technical Report.

Weak-to-Strong On-Policy Distillation MiMo-V2-Flash Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.854870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.854870Z digest=sha256:dc059dcfdd6c8caebab064ba3287427d43bac1d815b4c01eb359ee49ffb280c3

Observation 0481d0ab-0c81-4ada-a26a-0a30a6c5ed41 · outbound

This paper cites Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.981367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.981367Z digest=sha256:be2f4c7611398e4b305c4cec26bfc7fd2b4d1ae95443ab30746f263d1656de36

Observation 6c4c09e2-dd65-4459-9260-675426845d9d · outbound

This paper cites On Weak-to-Strong Generalization and f-Divergence.

Weak-to-Strong On-Policy Distillation On Weak-to-Strong Generalization and f-Divergence

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.105193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.105193Z digest=sha256:79cfbc17703cfe4e7f077039dec80641fddaec19ba34aff284fa63ae0708998c

Observation a1f25405-7ce8-4f61-bde8-9050768d3c59 · outbound

This paper cites OPRD: On-Policy Representation Distillation.

Weak-to-Strong On-Policy Distillation OPRD: On-Policy Representation Distillation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.196739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.196739Z digest=sha256:f1a3eec482ee0eac3d4bc1f438ad66f7bc35652e8323f0214932b600f68eecdd

Observation a09903bf-b5fe-4800-a220-d29b7586c4b5 · outbound

This paper cites International Conference on Learning Representations , volume=.

Weak-to-Strong On-Policy Distillation International Conference on Learning Representations , volume=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.281183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.281183Z digest=sha256:7c7caa9279344b2f4dad15de12fe3f5d692ca676d7157db93acc0f75215d0ea3

Observation c823c9a8-9ad6-4bfc-90ff-6d8e63403ce3 · outbound

This paper cites forward KL , author=.

Weak-to-Strong On-Policy Distillation forward KL , author=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.364454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.364454Z digest=sha256:63a93b398e6ca1ccdb3aac8f6a2e4625031ebdcdac00097dca9813393c4c4935

Observation 0520c4a3-0016-4277-8f82-b08a4d8420d7 · outbound

This paper cites Thinking Machines Lab: Connectionism , year =.

Weak-to-Strong On-Policy Distillation Thinking Machines Lab: Connectionism , year =

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.506888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.506888Z digest=sha256:0e23af4ac5c82565a68b53ec49bd8c199b1030767ce0a07cf48ed12dcab41918

Observation 16c884cc-ed3e-4ddd-82b3-af3c5fff7bfc · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Weak-to-Strong On-Policy Distillation Process Reinforcement through Implicit Rewards

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.653402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.653402Z digest=sha256:18dfacb1d289bd0d6fe350c539d72e61f2948f4551fb2eeae1ee4b4fd51b3a39

Observation 88f7e6cc-8d1e-4b30-af61-8d381d919c6b · outbound

This paper cites an unresolved cited work.

Weak-to-Strong On-Policy Distillation Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.768775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.768775Z digest=sha256:2f140c41ea3666d5274cbc5460b00c363f923222594006db08ea5d0453822972

Observation 3c36225b-6de3-4a63-8363-09d4bb8883ad · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Weak-to-Strong On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:24.869605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.869605Z digest=sha256:4a4aaf814ca27cab65a1117eef3c2071c575a186354338313e2b4a85d778fb42

Observation a3545e34-3207-4116-a722-9fa81021fc4a · outbound

This paper cites an unresolved cited work.

Weak-to-Strong On-Policy Distillation Unresolved cited work

Reference 77

Resolution
parse uncertain
no resolver link, observed 2026-08-01T00:26:24.970768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:24.970768Z digest=sha256:a49c1e07dd6b16766df44bec0207402ce1174d390eab1954942af3ccd097ee9d

Observation 8ad81f8e-c470-4db2-ab72-2650a284f41f · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Weak-to-Strong On-Policy Distillation DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.081268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.081268Z digest=sha256:afa501b68bbb7792af3cba389076b9351cb335afa6381cab7d664fb5a24e2ccf

Observation 4f4b306e-70d4-4ba6-b213-8a912aec2beb · outbound

This paper cites International Conference on Learning Representations , volume=.

Weak-to-Strong On-Policy Distillation International Conference on Learning Representations , volume=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.169140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.169140Z digest=sha256:11046456c1f1fce6250bedd71771feb174d65264fda692abc5082bad0c8c0272

Observation 8ec05144-8310-486f-87ca-4951390c544d · outbound

This paper cites URL https://matharena.

Weak-to-Strong On-Policy Distillation URL https://matharena

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.288499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.288499Z digest=sha256:12fc28ddf7363e6b7c3fc5ffd4af6b23a0f75ee3d53caa144e8bb4728758f793

Observation c7007b13-9e96-4e96-8c59-eaa8731a194d · outbound

This paper cites Advances in neural information processing systems , volume=.

Weak-to-Strong On-Policy Distillation Advances in neural information processing systems , volume=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.385382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.385382Z digest=sha256:62e4dea0245ee128e779da25ee2ad3f4febfa4b77bd31f0afd4163179aff51c9

Observation fc9af86c-2049-46c8-8961-415335480d90 · outbound

This paper cites Kimi-Audio Technical Report.

Weak-to-Strong On-Policy Distillation Kimi-Audio Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.522384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.522384Z digest=sha256:736152d268d127791528d0bc677f87d2ff0ace1d221bebcaf150ef6d4f080a90

Observation 9083144a-06bf-4bbd-8be3-d5e6c2285e44 · outbound

This paper cites 2023 , url =.

Weak-to-Strong On-Policy Distillation 2023 , url =

Reference 83

Resolution
verified exact
doi, observed 2026-08-01T00:31:07.749637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-01T00:26:25.667153Z digest=sha256:ac16ba31cedb740a88408d1e5194e9c3773f60bb9f585f596161be7037c164e0

Observation c6893a8c-bff4-45e4-9398-5ea127d52293 · outbound

This paper cites Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities.

Weak-to-Strong On-Policy Distillation Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.777254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.777254Z digest=sha256:ba2d07d455846a607a1471131b7a8431918f6f71ca1c970f9a1e60a3b3ccec43

Observation eb2323cc-1287-40c3-8a37-80c2e4629719 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Weak-to-Strong On-Policy Distillation Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.893734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.893734Z digest=sha256:852134fbc3b27a0900468118dec427669701d67511cc43f72f620067dd948dac

Observation 79532ab5-4cd3-4c0d-bd19-67e062b0a829 · outbound

This paper cites Proceedings of the 30th ACM International Conference on Multimedia , pages=.

Weak-to-Strong On-Policy Distillation Proceedings of the 30th ACM International Conference on Multimedia , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.023285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.023285Z digest=sha256:ed1db48b43c6434871b2fcee0e56b33265797e7bd65e8e744cf328fcb58f70c7

Observation eb25a9cb-433e-4cab-8feb-472ff4b34a52 · outbound

This paper cites GPT-4o System Card.

Weak-to-Strong On-Policy Distillation GPT-4o System Card

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.156814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.156814Z digest=sha256:80937a884e07c278a4664da290f557647a8599415c0e9553fcb200ea5e7f672f

Observation b5180e38-2847-4134-b0f5-909d126e32f3 · outbound

This paper cites GPT-4 Technical Report.

Weak-to-Strong On-Policy Distillation GPT-4 Technical Report

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.284433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.284433Z digest=sha256:bf490868fbe2df76a0c61d9fff63a1d005e888217b50fe0765cc26d93015dcc5

Observation 3da63c6d-a74d-4676-b508-94c041a0e928 · outbound

This paper cites Wav2CLIP: Learning Robust Audio Representations from Clip , booktitle =.

Weak-to-Strong On-Policy Distillation Wav2CLIP: Learning Robust Audio Representations from Clip , booktitle =

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.410915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.410915Z digest=sha256:b11b6b04edb1b0f3eddd34ea6c089325de16ebbf37e658ff2dc6da7bffa0c241

Observation 5e3257b0-76a2-4728-9ec5-65a4a13295e1 · outbound

This paper cites Pengi: An Audio Language Model for Audio Tasks , booktitle =.

Weak-to-Strong On-Policy Distillation Pengi: An Audio Language Model for Audio Tasks , booktitle =

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.473503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.473503Z digest=sha256:ac2f8806c58f1bd03ffba49b74d0682b50150d2f9b3e8728da4a19836e4c06df

Observation b947c43a-0fcb-454a-8c9c-301e68ef6763 · outbound

This paper cites Liu and Leonid Karlinsky and James R.

Weak-to-Strong On-Policy Distillation Liu and Leonid Karlinsky and James R

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.552685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.552685Z digest=sha256:6910dc7825e6c1d09ba5ad036172c16d5dab024c994628bf39e644306f65bcb5

Observation 6e396538-fabf-40f1-81ca-84b80593d622 · outbound

This paper cites Optimal Transport for Treatment Effect Estimation , booktitle =.

Weak-to-Strong On-Policy Distillation Optimal Transport for Treatment Effect Estimation , booktitle =

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.585832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.585832Z digest=sha256:e547cdf52469d5dd7acf3475af1103fb187089a6423527967eb1e6970816baee

Observation 13963a24-32b3-43d2-8240-2305c8148ca8 · outbound

This paper cites A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions.

Weak-to-Strong On-Policy Distillation A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.646865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.646865Z digest=sha256:55f8ce5f8a5521317456dc72d16c956a7642f2c5c78b67d8f728f5b4e676f838

Observation 7d5a76d6-8cb4-428a-b125-2269d5616a53 · outbound

This paper cites Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection , booktitle =.

Weak-to-Strong On-Policy Distillation Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection , booktitle =

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.784799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.784799Z digest=sha256:070a1e16e348d84c3ebf85dbc2308e6f810bb603165d549c030aa4a0a52e125f

Observation b78ee796-4202-4f8f-b32e-934422834d0b · outbound

This paper cites Generalized Data Distribution Iteration , booktitle =.

Weak-to-Strong On-Policy Distillation Generalized Data Distribution Iteration , booktitle =

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:26.902413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:26.902413Z digest=sha256:3a419861e48906c00ec9c62fb87455a4913b5461e502468d45616e9aee82c675

Observation 1d5146a5-1679-4da3-a8f8-e126d4c3fccd · outbound

This paper cites GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning.

Weak-to-Strong On-Policy Distillation GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:27.063018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:27.063018Z digest=sha256:6f6203b61d6ce79ad990f4bc85b21ad2bad927d4f67d7f1a87c4e781e5b3028f

Observation aa07a1eb-f80c-46bc-8e96-57f92e94889b · outbound

This paper cites CoRR , volume =.

Weak-to-Strong On-Policy Distillation CoRR , volume =

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:27.202175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:27.202175Z digest=sha256:f931f1bc28d23d53981fedaf2991a09b842afb3ad1cb498a50b6b4bc35b29b95

Observation 687fa503-d69e-4af9-bc57-3db573a602a2 · outbound

This paper cites ConvFormer: Revisiting Transformer for Sequential User Modeling.

Weak-to-Strong On-Policy Distillation ConvFormer: Revisiting Transformer for Sequential User Modeling

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-08-01T00:31:07.611462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-01T00:26:27.325656Z digest=sha256:354e5f866499e4d1d79844645c65f3e427da9dd408100efe31c808048a5e118c

Observation e8c97fdf-3f00-4a8d-8e9a-aa05912b2393 · outbound

This paper cites An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning.

Weak-to-Strong On-Policy Distillation An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:27.446030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:27.446030Z digest=sha256:d07e6ad5c47d1f7f480918c2245c1b57148fe686dfb474cb6f1bbadb6eb4c62e

Observation 6e213f62-0170-4e5d-a5de-2178c5515c25 · outbound

This paper cites CoRR , volume =.

Weak-to-Strong On-Policy Distillation CoRR , volume =

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-08-01T00:31:07.518318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-01T00:26:27.540391Z digest=sha256:86df90deefa8d78552446700b9eae66046fbbf1271d21b36ec84f926cd3c8971

Pith citing papers

No inbound Pith citation observations are available.