REVIEW 6 cited by
CLIP model is an Efficient Continual Learner
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The continual learning setting aims to learn new tasks over time without forgetting the previous ones. The literature reports several significant efforts to tackle this problem with limited or no access to previous task data. Among such efforts, typical solutions offer sophisticated techniques involving memory replay, knowledge distillation, model regularization, and dynamic network expansion. The resulting methods have a retraining cost at each learning task, dedicated memory requirements, and setting-specific design choices. In this work, we show that a frozen CLIP (Contrastive Language-Image Pretraining) model offers astounding continual learning performance without any fine-tuning (zero-shot evaluation). We evaluate CLIP under a variety of settings including class-incremental, domain-incremental and task-agnostic incremental learning on five popular benchmarks (ImageNet-100 & 1K, CORe50, CIFAR-100, and TinyImageNet). Without any bells and whistles, the CLIP model outperforms the state-of-the-art continual learning approaches in the majority of the settings. We show the effect on the CLIP model's performance by varying text inputs with simple prompt templates. To the best of our knowledge, this is the first work to report the CLIP zero-shot performance in a continual setting. We advocate the use of this strong yet embarrassingly simple baseline for future comparisons in the continual learning tasks.
Forward citations
Cited by 6 Pith papers
-
Continual Learning with Vision-Language Models via Semantic-Geometry Preservation
SeGP-CL reduces catastrophic forgetting in CLIP-based continual learning by distilling cross-modal geometry around adversarial anchors at the old-new class boundary plus regularizing the text-space reference frame.
-
Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning
MG-CLIP preserves CLIP's modality gap by adaptively limiting fine-tuning epochs and compensates for its limits with a visual-space classifier, improving class-incremental learning without replay.
-
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
Proxy-FDA aligns the local neighborhood structure of pre-trained and fine-tuned feature spaces, generating synthetic proxies to reduce concept forgetting during fine-tuning.
-
Toward Verifiable Misinformation Detection: A Multi-Tool LLM Agent Framework
A multi-tool LLM agent for verifiable misinformation detection is claimed to outperform baselines on FakeNewsNet in accuracy, transparency, and rewriting resistance.
-
ChordPrompt: Orchestrating Cross-Modal Prompt Synergy for Multi-Domain Incremental Learning in CLIP
Cross-modal prompt sharing with domain-adaptive retrieval improves CLIP's continual learning performance on multi-domain image classification benchmarks.
-
DCFormer: Efficient 3D Vision-Language Modeling with Decomposed Convolutions
A decomposed 3D convolution encoder for chest CT reaches competitive pathology detection and image-text retrieval with far fewer parameters and FLOPs than transformer or full 3D convolution baselines.
Discussion (0). Continue with ORCID to comment.