Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T07:01:56.628017Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 4 inbound Pith citation observations for arXiv:2607.05394.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T07:01:56.628017Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:07:43.400511Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 105 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6384b32f-a969-43f8-a826-5270904df0ff · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781713bd-26e9-4bf9-a4c7-0a8c157c27ad · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation JustRL: Scaling a 1.5B LLM with a simple RL recipe
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d872834-608d-47e3-bd86-2822c9332ba4 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Qwen3 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa5a17f-54ca-4445-8b18-565ae2510235 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation POLARIS: A post-training recipe for scaling reinforcement learning on reasoning models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c1681c-4087-4f32-80df-36b75dd0da39 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36d5cf22-a34d-4ab9-91ff-2a758cec96f9 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9653133b-8622-4b08-8a47-917e10dfefc3 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation QuestA: Expanding reasoning capacity in LLMs via question augmentation.arXiv preprint arXiv:2507.13266, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e94af6ad-635d-4f35-9448-7b73035eb864 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Lyng, Sanjit Singh Batra, and Robert E
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50deb921-7e82-4e45-a46f-85698de79d12 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d3987e-c861-479b-96b7-fa783a730760 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Proximal Policy Optimization Algorithms
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94210e6c-4e52-4c8f-a8a9-bf63ba3344f3 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3eee08c-fcbf-4d05-9472-662838886b63 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation QwQ-32B: Embracing the power of reinforcement learning.https://qwenlm.github.io/blog/qw q-32b/, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5107ec-5293-491b-8077-284ae9ad5976 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Open-reasoner- zero: An open source approach to scaling reinforcement learning on the base model.https://github.com/Ope n-Reasoner-Zero/Open-Reasoner-Zero, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b534cb-3bb6-4c11-b93f-3ea3b04600c4 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ee0745-8e4e-40a9-a859-7f17b88c5fba · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Skywork Open Reasoner 1 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf49e12-3bfb-4301-9504-f6cef34cf187 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a95b12cb-86f7-4ad0-a343-52cd6930c9cd · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation How far can unsupervised rlvr scale llm training?arXiv preprintarXiv:2603.08660, 2026
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4c5fb1-4e19-45f4-aa99-5a4411c09190 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DeepSeek V4 preview release.https://api-docs.deepseek.com/news/news260424, 2026
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2f48df-f0f0-49ee-8791-83023a08cf92 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation GLM-5.2: Built for long-horizon tasks.https://z.ai/blog/glm-5.2, 2026
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e7e64b-cea9-457a-b420-8432ff36543f · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8404f6-722a-4ebb-a24f-a5ac0a817d8d · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation OpenThoughts: Data Recipes for Reasoning Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4581e08f-2d0a-40d5-9a7c-43a41ba0900a · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecc0a348-7280-492a-985b-560af29ffa29 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Distilling the Knowledge in a Neural Network
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2adf695e-99c8-4162-a617-959b9e5f835d · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Sequence-level knowledge distillation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6e95a5-65dc-4b9e-bde8-07b4ce1c8e55 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a69c8719-5814-4132-ace2-c08798ad8e0f · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Tinybert: Distilling bert for natural language understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1008496c-3766-4143-be47-59767fdbbf98 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b37e31-2aab-44df-b3a7-9cbacb74238c · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Improved knowledge distillation via teacher assistant
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6923504a-4d05-46ff-b1e0-b178db0f4a6f · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation On the efficacy of knowledge distillation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39e87fb7-d8c1-4488-a56e-e9092c417205 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Distillation Scaling Laws
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f424f5-2513-4320-b553-1367fbc27c14 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation MiniLLM: On-Policy Distillation of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d7ce90-7def-4c90-9ae3-971c097dea1f · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation f-Divergence Minimization for Sequence-Level Knowledge Distillation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06919646-9ff1-4cdd-b87a-59e50e152d61 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Revisiting Knowledge Distillation for Autoregressive Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cbe331a-0496-4e58-9684-cd0b1ebcb88c · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e193e4c4-d4be-4125-9b9c-a5b283fa6f3b · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DistiLLM: Towards Streamlined Distillation for Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d101de-8bd5-4513-a448-a79bafb8a2e5 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0c877e-fa71-480a-bd99-fb47ceb9b5e6 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation MiniPLM: Knowledge Distillation for Pre-Training Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7aab0c-affa-4812-afc1-8f3420495a68 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DistillSpec: Improving Speculative Decoding via Knowledge Distillation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77a8127-b322-463a-9b50-9174f09f7466 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 024037ba-1aa8-4095-a577-2a05dace3ead · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 728b120e-b2c7-4ebd-8f93-6ed3e04fcb5d · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation On-policy distillation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f79f2454-7766-4eb6-a227-81f81e99d52c · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation A Survey of On-Policy Distillation for Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e3d80cd-222f-4bf9-8239-ce5b1f2ac704 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e7219c-baa6-428c-ada3-027eeaf43a71 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Entropy-Aware On-Policy Distillation of Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 047a9d32-1acc-419c-a421-c27d9081f21f · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Stable On-Policy Distillation through Adaptive Target Reformulation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 634d9d0a-0f4b-4104-87d3-a211920d3e77 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Unifying group-relative and self-distillation policy optimization via sample routing
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590650dd-5a62-415d-917c-4b7c0c073182 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation On-Policy Context Distillation for Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c662a1-119b-458d-96fd-5c585807adb7 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Online Experiential Learning for Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edc387dc-4cc2-4c0f-b0e1-d663a455cabf · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f80e56c1-8224-4c47-a94f-165fd0dd53b4 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Self-distillation for multi-token prediction, 2026.arXiv preprint arXiv:2603.23911, 2026
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a91287-e81b-447f-ba55-06cc450123c2 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Reinforcement Learning via Self-Distillation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90474726-e136-48de-9e3b-6e4428f2baf5 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Distillation Enables Continual Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee600dd5-d64a-4e2f-b6f9-592d71db122c · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4348ce7b-e70f-47d8-80a7-4351fa53540f · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Scaling reasoning efficiently via relaxed on-policy distillation.arXiv preprint arXiv:2603.11137, 2026
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5cc2cab-e0d2-4d53-a4cd-ec2741f3039b · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e211856f-990b-4702-8110-4a7c6b1588a3 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Black-Box On-Policy Distillation of Large Language Models.arXiv preprint arXiv:2511.10643, 2025
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d55ba62-ba98-4143-9c40-df5269453ede · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation MiMo-V2-Flash Technical Report
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276f615d-744d-43a4-9f4a-e18d4b387fe2 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d41ed5-9b4c-4adb-a93e-dcf3a9891769 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468560f9-9e50-4064-8267-4f0333143a5f · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Reinforcement-aware Knowledge Distillation for LLM Reasoning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dae33936-2343-424a-a54f-d8b6cb275983 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199e62b2-4e82-4319-b42c-2c9f4923472e · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Distilled RLVR
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90af5db4-293a-48db-b8ee-8fcac00d6648 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Small models struggle to learn from strong reasoners
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f6fa9c3-f691-4d54-bab7-22445a3644f7 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4278c51-8044-46dd-8869-3e9a187f8195 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d9953a-31c0-4406-b7bc-bab71cdbf235 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Introducing Superalignment.OpenAI Blog, 2023
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fff07fb-add9-4446-8386-8f388a8a2e3b · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Semi-supervised learning by entropy minimization.Advances in neural information processing systems, 17, 2004
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bcd7216-249b-4b3a-a233-65f4d84875b7 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Semi-supervised sequence learning.Advances in neural information processing systems, 28, 2015
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d855c2-7bc0-4846-aeed-7732e5b6d353 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Temporal Ensembling for Semi-Supervised Learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3028b6e0-7e46-4822-97ff-68d65d743ec6 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation There Are Many Consistent Explanations of Unlabeled Data: Why You Should Average
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82abf378-ab71-4341-8a32-0c2805e6112f · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Co- teaching: Robust training of deep neural networks with extremely noisy labels.Advances in neural information processing systems, 31, 2018
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ead182-98ff-4c1c-88df-9ae548532948 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Mixmatch: A holistic approach to semi-supervised learning.Advancesin neural information processing systems, 32, 2019
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f07f7858-d323-44b4-b4d7-2cbf87c19d7a · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation DivideMix: Learning with Noisy Labels as Semi-supervised Learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950dff98-1d8e-443f-b012-6690f257a3cc · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Big self-supervised models are strong semi-supervised learners.Advancesin neural information processing systems, 33:22243–22255, 2020
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c27b717-4ded-4395-9db9-0e282e795c8c · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Self-training avoids using spurious features under domain shift
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64ef02a0-7147-45f0-8af1-8039bc4c3f2c · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Debiased self-training for semi-supervised learning.Advances in Neural Information Processing Systems, 35:32424–32437, 2022
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5da4bb-bd31-4da6-89f0-be1b9fc965bf · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Supervising strong learners by amplifying weak experts
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fbe11a8-925b-4d49-8365-804d4389dcb3 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Scalable agent alignment via reward modeling: a research direction
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c784bae3-20a2-41a5-974b-8598dea0a1d0 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation AI safety via debate
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efc4af27-2ee3-4a29-9127-9092046319d2 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Measuring Progress on Scalable Oversight for Large Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea91f91-1b4b-4cb1-a103-43bccd9bc4dc · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Artificial Sandwiching: When can we test scalable alignment protocols without humans?AI Alignment Forum, 2022
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99441c2c-ca2f-4586-bf71-7c1011ce5e6f · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Constitutional AI: Harmlessness from AI Feedback
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2a85736-198b-47e0-94f0-3c6478938af2 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Eliciting latent knowledge
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ddb4ac6-1bdc-4ba8-918d-16f06d406da1 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Discovering Latent Knowledge in Language Models Without Supervision
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59ddbb5-b5d1-453d-8626-7b68dba12af6 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e164f9-b7ce-4033-aa45-6d8ad297595e · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Datasets for Studying Generalization from Easy to Hard Examples
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e73740b-de68-4537-b9e3-29f49fd9d5fe · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Towards Scalable Automated Alignment of LLMs: A Survey
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e6647ff-7514-4774-80c3-60e33af2a2fc · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7fe6e15-85a8-4a3d-97e2-033207bc1c2d · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation On Weak-to-Strong Generalization and f-Divergence
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24e9d31-ee6b-4b32-8a90-579283d1a286 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05cc7a6e-dcbc-43d6-84f6-477e05ee8cee · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Incentivizing strong reasoning from weak supervision.arXiv preprint arXiv:2505.20072, 2025
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36f27840-916c-40e0-a746-c4ccd5e21f1b · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Rewarding Language Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30059df4-9f93-4e88-8475-7f9ccd21f731 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a17ee42-4f29-47d8-9461-8c67c3ee4b4c · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation A general theoretical paradigm to understand learning from human preferences.AISTATS, 2024
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6b0ef9-469d-4174-ba8c-81a441375b34 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation KTO: Model Alignment as Prospect Theoretic Optimization
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a573a63-9006-4612-9cca-b45b8d2aab0a · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation SimPO: Simple preference optimization with a reference-free reward
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61393b5f-49b2-49fd-bfbc-66143ca88918 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Let's Verify Step by Step
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6efae79e-95ab-44cf-9b7b-66b1143a3634 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9130062e-7c6d-431d-b775-576e47f93f5d · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Process Reinforcement through Implicit Rewards
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68e1632-e745-4550-a34c-fc003f8ded03 · outbound
Weak-to-Strong Generalization via Direct On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3bc2ed4-398d-4cbd-b8e9-451f2561bf49 · inbound
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals Weak-to-Strong Generalization via Direct On-Policy Distillation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc5e68fa-d86f-46ee-8b20-62d69c403240 · inbound
Visual Contrastive Self-Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e45157b-2721-467a-bed1-ad066595600d · inbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2acfd7d2-a822-474e-a4f4-f0410f2bead1 · inbound
Weak-to-Strong On-Policy Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.