OPD-Evolver uses on-policy self-distillation in fast interaction and slow attribution loops to build agents with holistic memory competence, outperforming prior systems by up to 11.5% and allowing a 9B model to compete with much larger ones.
arXiv preprint arXiv:2111.10447 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Diagnoses attention dispersion in CTDG Transformers under temporal shift and introduces differential attention to suppress common signals and achieve SOTA on shifted benchmarks.
citing papers explorer
-
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation
OPD-Evolver uses on-policy self-distillation in fast interaction and slow attribution loops to build agents with holistic memory competence, outperforming prior systems by up to 11.5% and allowing a 9B model to compete with much larger ones.
-
Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix
Diagnoses attention dispersion in CTDG Transformers under temporal shift and introduces differential attention to suppress common signals and achieve SOTA on shifted benchmarks.