CoM-PT trains vision foundation models in ascending size order using inverse knowledge transfer, allowing larger models to achieve superior performance with significantly reduced overall computational cost compared to individual training.
arXiv preprint arXiv:2010.13002 , year=
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
A modified divergence decouples top-K teacher probabilities from the distribution tail during distillation, yielding competitive performance on decoder models with standard compute.
CANDLE applies CTC alignment to Arabic character deduplication, achieving 5.37% sentence error rate on clean text and up to 12.8% tokenizer fertility reduction.
citing papers explorer
-
Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
CoM-PT trains vision foundation models in ascending size order using inverse knowledge transfer, allowing larger models to achieve superior performance with significantly reduced overall computational cost compared to individual training.
-
Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
A modified divergence decouples top-K teacher probabilities from the distribution tail during distillation, yielding competitive performance on decoder models with standard compute.
-
CANDLE: CTC-based Arabic Noisy-character Deduplication using a Lightweight Encoder
CANDLE applies CTC alignment to Arabic character deduplication, achieving 5.37% sentence error rate on clean text and up to 12.8% tokenizer fertility reduction.