{"work":{"id":"e9b1d13c-a8ca-43ed-af63-fd1158065caf","openalex_id":null,"doi":null,"arxiv_id":"2111.03930","raw_key":null,"title":"Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling","authors":null,"authors_text":"Zhang, R","year":2021,"venue":"cs.CV","abstract":"Contrastive Vision-Language Pre-training, known as CLIP, has provided a new paradigm for learning visual representations by using large-scale contrastive image-text pairs. It shows impressive performance on zero-shot knowledge transfer to downstream tasks. To further enhance CLIP's few-shot capability, CLIP-Adapter proposed to fine-tune a lightweight residual feature adapter and significantly improves the performance for few-shot classification. However, such a process still needs extra training and computational resources. In this paper, we propose \\textbf{T}raining-Free CL\\textbf{IP}-\\textbf{Adapter} (\\textbf{Tip-Adapter}), which not only inherits CLIP's training-free advantage but also performs comparably or even better than CLIP-Adapter. Tip-Adapter does not require any back propagation for training the adapter, but creates the weights by a key-value cache model constructed from the few-shot training set. In this non-parametric manner, Tip-Adapter acquires well-performed adapter weights without any training, which is both efficient and effective. Moreover, the performance of Tip-Adapter can be further boosted by fine-tuning such properly initialized adapter for only a few epochs with super-fast convergence speed. We conduct extensive experiments of few-shot classification on ImageNet and other 10 datasets to demonstrate the superiority of proposed Tip-Adapter. The code will be released at \\url{https://github.com/gaopengcuhk/Tip-Adapter}.","external_url":"https://arxiv.org/abs/2111.03930","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-08T10:54:49.125621+00:00","pith_arxiv_id":"2111.03930","created_at":"2026-05-11T16:36:09.388033+00:00","updated_at":"2026-07-08T10:54:49.125621+00:00","title_quality_ok":true,"display_title":"Deep sets","render_title":"Deep sets"},"hub":{"state":{"work_id":"e9b1d13c-a8ca-43ed-af63-fd1158065caf","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":16,"external_cited_by_count":null,"distinct_field_count":2,"first_pith_cited_at":"2023-02-10T23:12:37+00:00","last_pith_cited_at":"2026-07-07T14:11:54+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T10:29:57.898597+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":1},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":1},{"context_polarity":"use_method","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}