The paper fits power-law scaling laws for downstream vision tasks and claims a pretraining-data threshold where distilled models stop outperforming non-distilled ones, but the theory's assumptions come from the fitted exponents and the model-size scaling is contradicted by the paper's own figures.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Scaling Laws for Data-Efficient Visual Transfer Learning
The paper fits power-law scaling laws for downstream vision tasks and claims a pretraining-data threshold where distilled models stop outperforming non-distilled ones, but the theory's assumptions come from the fitted exponents and the model-size scaling is contradicted by the paper's own figures.