Towards on-policy sft: Distribution discriminant theory and its applications in llm training

Towards On-Policy SFT: Distribution Discriminant Theory, its Applications in LLM Training , author= · 2026 · arXiv 2602.12222

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

read on arXiv browse 3 citing papers

citation-role summary

other 1

citation-polarity summary

unclear 1

representative citing papers

ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design

cs.LG · 2026-05-11 · unverdicted · novelty 6.0

ProteinOPD uses token-level on-policy distillation from multiple preference-specific teacher models into a shared student to balance competing objectives in protein design, delivering gains on targets without losing designability and an 8x speedup over RL baselines.

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation

cs.LG · 2026-06-01 · unverdicted · novelty 5.0

FiRe-OPD introduces a two-stage filter-then-soft-reweight procedure for trajectory- and token-level supervision in on-policy distillation, claiming gains over prior token-level methods.

Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation

cs.CL · 2026-05-30 · unverdicted · novelty 4.0

DASD dynamically selects tokens in self-distillation to keep logical corrections while suppressing stylistic noise, improving robustness on math, code, and commonsense benchmarks.

citing papers explorer

Showing 3 of 3 citing papers.

ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design cs.LG · 2026-05-11 · unverdicted · none · ref 40
ProteinOPD uses token-level on-policy distillation from multiple preference-specific teacher models into a shared student to balance competing objectives in protein design, delivering gains on targets without losing designability and an 8x speedup over RL baselines.
Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation cs.LG · 2026-06-01 · unverdicted · none · ref 42
FiRe-OPD introduces a two-stage filter-then-soft-reweight procedure for trajectory- and token-level supervision in on-policy distillation, claiming gains over prior token-level methods.
Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation cs.CL · 2026-05-30 · unverdicted · none · ref 6
DASD dynamically selects tokens in self-distillation to keep logical corrections while suppressing stylistic noise, improving robustness on math, code, and commonsense benchmarks.

Towards on-policy sft: Distribution discriminant theory and its applications in llm training

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer