Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:31:40.978324Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2608.08764.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:31:40.978324Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bbc3693d-ac26-4f0a-b217-c4088ef02f59 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast KTO: Model Alignment as Prospect Theoretic Optimization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9eec759-0224-4efe-a5f5-79d6cf0d0e7d · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Reinforcement Learning via Self-Distillation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa17ce90-6107-463f-8b99-06de3dd671f3 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 268847bd-388d-4ac4-94f9-051cea0e1ddd · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc29fc1-97b6-4435-9476-34db50f38de4 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast DistiLLM: Towards Streamlined Distillation for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d0201ab-a43f-49dd-9762-5644e0729f83 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Privileged Information Distillation for Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a0a7d45-0028-4ae4-b3f9-85f33a0685eb · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a03d18-c17c-4b11-8b7e-5ef7014ed2f8 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3718fcaf-3c77-4ba9-9e94-d4b1c4bd9cbd · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast arXiv preprint arXiv:2602.20574
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590da225-31ed-49f2-9a8f-67cbe11f1903 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd62bf11-6b5e-4479-bd97-dc5121eb6fc4 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Qwen3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f8995b-6e2c-4e50-b804-8984e7d8a593 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Self-Distilled RLVR
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71732ff7-32f2-43d4-899a-98146d2e9603 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast On-Policy Context Distillation for Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32a8d406-ca36-45cf-b120-8231effb7cd4 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Multi-Rollout On-Policy Distillation via Peer Successes and Failures
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44442b6a-9581-4299-9fb8-7788d0591050 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20003c9d-fca1-40f7-950f-319f121b8e79 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 953d7c51-62cd-4021-891c-c70fa83d153f · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Distilling the Knowledge in a Neural Network
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c24abb-978b-49d9-9c97-3a3325795f36 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Training language models to follow instructions with human feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c33c231-9f9d-49f4-b758-e1bc832e81cf · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast InFindings of the Association for Computational Linguistics: EMNLP 2023, 5687–5711
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 33d8db9a-da57-456a-a43f-2c4dda5b33f8 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast InInternational Conference on Learn- ing Representations, volume 2024, 21246–21263
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4ebc4a7-3eb0-4167-83fb-f072da754b68 · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast OpenThoughts: Data Recipes for Reasoning Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c723d0dc-7693-4866-8e27-24a711f934ce · outbound
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.