Pith. sign in

Paper Citation Record · LEDGER

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast

As of 18 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2608.08764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08764 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:31:40.978324Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bbc3693d-ac26-4f0a-b217-c4088ef02f59 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast KTO: Model Alignment as Prospect Theoretic Optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.893421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.893421Z digest=sha256:860c371099a8c85295358072e303740affa21171e594faa3fa06ecd649f0e5a4

Observation a9eec759-0224-4efe-a5f5-79d6cf0d0e7d · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Reinforcement Learning via Self-Distillation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.913039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.913039Z digest=sha256:e5de3816f05e1ed041f7644b5f6de9740b1dcbc609f1cd8ac64e9300515aa5af

Observation fa17ce90-6107-463f-8b99-06de3dd671f3 · outbound

This paper cites Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.917409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.917409Z digest=sha256:36015dfb7efc92b05ed36c6089b77198860626fa641ccf36828a7accaef0a3fd

Observation 268847bd-388d-4ac4-94f9-051cea0e1ddd · outbound

This paper cites UniSD: Towards a Unified Self-Distillation Framework for Large Language Models.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.921666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.921666Z digest=sha256:b04cb392695c5b919eef4345b1e3d7a8535c7b8d57cdedc4460d0fed7c1d121b

Observation 9dc29fc1-97b6-4435-9476-34db50f38de4 · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.926096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.926096Z digest=sha256:934e5f99a290180a47ffb3885416f81aea4af35a902e7fa5502cb70fdab36c4e

Observation 5d0201ab-a43f-49dd-9762-5644e0729f83 · outbound

This paper cites Privileged Information Distillation for Language Models.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Privileged Information Distillation for Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.934410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.934410Z digest=sha256:9e6984caf4895587cbc11137b95fe178399d96c2712cfe35fb80bff2c01a3b8a

Observation 4a0a7d45-0028-4ae4-b3f9-85f33a0685eb · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.942530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.942530Z digest=sha256:cde87308a46bbdb45208733750d7c457058e5e025113871909d9d0fca6aaa16d

Observation 95a03d18-c17c-4b11-8b7e-5ef7014ed2f8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.946297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.946297Z digest=sha256:af6778f6c6200f26c24ed5dcef491d37d06a31a0b8fd74d3931d752d7c30167b

Observation 3718fcaf-3c77-4ba9-9e94-d4b1c4bd9cbd · outbound

This paper cites arXiv preprint arXiv:2602.20574.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast arXiv preprint arXiv:2602.20574

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.950141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.950141Z digest=sha256:abe2619e970bc77d88f7c02bf126025ff234c2c93d5f7b75eb50300c8f505ef5

Observation 590da225-31ed-49f2-9a8f-67cbe11f1903 · outbound

This paper cites PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.953793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.953793Z digest=sha256:9dc138e9fd229ce8454054aaf124a55edb97c0b8ed9e36f1c38dfc60efa43155

Observation cd62bf11-6b5e-4479-bd97-dc5121eb6fc4 · outbound

This paper cites Qwen3 Technical Report.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Qwen3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.957549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.957549Z digest=sha256:6757057c7a75770490dfb99d9c472a3f27dfc77cd6d41ceb450c90f0ebbfb085

Observation 15f8995b-6e2c-4e50-b804-8984e7d8a593 · outbound

This paper cites Self-Distilled RLVR.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Self-Distilled RLVR

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.962082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.962082Z digest=sha256:86ac7d4decc6d202c332731af398190cfb32656dd52d381457f638a6e6699c2b

Observation 71732ff7-32f2-43d4-899a-98146d2e9603 · outbound

This paper cites On-Policy Context Distillation for Language Models.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast On-Policy Context Distillation for Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.965980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.965980Z digest=sha256:158d053199185a927e3d90c267d4c382ee10e6e0c820ce995cb82daec4534e98

Observation 32a8d406-ca36-45cf-b120-8231effb7cd4 · outbound

This paper cites Multi-Rollout On-Policy Distillation via Peer Successes and Failures.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Multi-Rollout On-Policy Distillation via Peer Successes and Failures

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.969905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.969905Z digest=sha256:e4f2dc65a246d42b91315f4d1f139180373836e9b102835f3fc3ae2ce2181840

Observation 44442b6a-9581-4299-9fb8-7788d0591050 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.974049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.974049Z digest=sha256:95211e709d3b4dbeb0f62fe10303bb31710805db24d2218a526851d4c405f7aa

Observation 20003c9d-fca1-40f7-950f-319f121b8e79 · outbound

This paper cites SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.978324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.978324Z digest=sha256:274c045d841888cdb3e16b019a2db9710bf3bcd8ae90200fd400ef2823ee1e2e

Observation 953d7c51-62cd-4021-891c-c70fa83d153f · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Distilling the Knowledge in a Neural Network

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.908467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.908467Z digest=sha256:a3efb7a1f8a9edfdd5e8e83b487b371acdb4c10ffc17c53210bdfeb371c679ac

Observation a6c24abb-978b-49d9-9c97-3a3325795f36 · outbound

This paper cites Training language models to follow instructions with human feedback.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.930257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.930257Z digest=sha256:c7bcddaae44a720b45df318d60c801995863c69350618cd3c437cb00bb72d456

Observation 4c33c231-9f9d-49f4-b758-e1bc832e81cf · outbound

This paper cites InFindings of the Association for Computational Linguistics: EMNLP 2023, 5687–5711.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast InFindings of the Association for Computational Linguistics: EMNLP 2023, 5687–5711

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:31:41.270942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:31:40.938535Z digest=sha256:8f934b7e18c0f16ba97220bcc5ea7f3d76808fb1563efa5b18378ecc2e4eb217

Observation 33d8db9a-da57-456a-a43f-2c4dda5b33f8 · outbound

This paper cites InInternational Conference on Learn- ing Representations, volume 2024, 21246–21263.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast InInternational Conference on Learn- ing Representations, volume 2024, 21246–21263

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:31:41.283534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:31:40.888392Z digest=sha256:d1cb2c1c4a59ec00d5a9d3b4e34c726daa217b49de7b33d7c5745533c0e3bfae

Observation c4ebc4a7-3eb0-4167-83fb-f072da754b68 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast OpenThoughts: Data Recipes for Reasoning Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.903633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.903633Z digest=sha256:fe0320f78cde266fe6185bfb18f8e5f79564350aee6835b317e8e0d101f19b78

Observation c723d0dc-7693-4866-8e27-24a711f934ce · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:40.898325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:40.898325Z digest=sha256:4ecf3ea9e27290078f0502b8a50fbdcdeba614cfc590228bca998def64f90168

Pith citing papers

No inbound Pith citation observations are available.