Pith. sign in

REVIEW 1 cited by

Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.18237 v3 pith:ZZ6WVMXJ submitted 2023-11-30 cs.CV cs.LG

classification cs.CVcs.LG
keywords pretrainingknowledgemodelstargettransferapproachcomputecost
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision Foundation Models (VFMs) pretrained on massive datasets exhibit impressive performance on various downstream tasks, especially with limited labeled target data. However, due to their high inference compute cost, these models cannot be deployed for many real-world applications. Motivated by this, we ask the following important question, "How can we leverage the knowledge from a large VFM to train a small task-specific model for a new target task with limited labeled training data?", and propose a simple task-oriented knowledge transfer approach as a highly effective solution to this problem. Our experimental results on five target tasks show that the proposed approach outperforms task-agnostic VFM distillation, web-scale CLIP pretraining, supervised ImageNet pretraining, and self-supervised DINO pretraining by up to 11.6%, 22.1%, 13.7%, and 29.8%, respectively. Furthermore, the proposed approach also demonstrates up to 9x, 4x and 15x reduction in pretraining compute cost when compared to task-agnostic VFM distillation, ImageNet pretraining and DINO pretraining, respectively, while outperforming them. We also show that the dataset used for transferring knowledge has a significant effect on the final target task performance, and introduce a retrieval-augmented knowledge transfer strategy that uses web-scale image retrieval to curate effective transfer sets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Contrasting local and global modeling with machine learning and satellite data: A case study estimating tree canopy height in African savannas

    cs.LG 2024-11 conditional novelty 6.0 of 10

    A locally trained five-layer convolutional network predicts tree canopy height in Karingani Game Reserve with 1.64 m RMSE, outperforming four global TCH maps (2.43 to 4.51 m) and locally fine-tuned global models (best...

Pith tools