Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This survey claims to be the first to map CLIP's role in domain generalization and domain adaptation, organizing the field into a single two-branch taxonomy.

desk verdict A genuinely useful taxonomy of CLIP-based DG/DA, but the 'first comprehensive survey' claim needs a stated search protocol and a cross-check pass before the organization can be trusted. read the letter →

arxiv 2504.14280 v1 pith:JH3Z3UR7 submitted 2025-04-19 cs.CV cs.LG

classification cs.CVcs.LG
keywords domaingeneralizationadaptationCLIPvision-languagemodelspromptlearningsource-freezero-shotsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Domain generalization and domain adaptation both ask how a model trained on one or more source domains behaves under distribution shift, and this survey claims that CLIP, a vision-language model with contrastive and zero-shot capabilities, has become the common substrate for answering that question. The paper's central assertion is that no earlier review has focused on CLIP's role in these two tasks, and that a single taxonomy can organize the fast-growing literature. It sorts DG methods into prompt optimization versus using CLIP as a backbone, and DA methods into source-available versus source-free settings, with source-free further split into source-data-free and source-fully-free. A sympathetic reader would care because a reliable map of this territory tells a practitioner which method family fits their data budget and which cells of the field remain empty.

What carries the argument

The machinery that carries the argument is the taxonomy itself, built on CLIP's contrastive joint embedding space. CLIP trains image and text encoders so paired image-text embeddings are close and unpaired ones are pushed apart, which is what gives the model its zero-shot classification and robustness under domain shift; the survey's two decisive classification axes are then, for DG, prompt optimization versus CLIP as backbone or encoder, and for DA, source-available versus source-free, with source-free subdivided into source-data-free and source-fully-free and each cell further cut by closed-set, partial-set, open-set, and open-partial-set label relationships. This grid is what lets the paper place the cited methods, identify gaps, and connect benchmarks and metrics to scenario types.

What would settle it

Find the set of CLIP-based domain generalization or adaptation methods published before the survey's cutoff that are absent from the paper and cannot be slotted into its scenario table; if that set is substantial, say more than a handful of peer-reviewed methods, the completeness claim that anchors the survey collapses.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is organizational: the CLIP-powered DG and DA landscape is not a tangle of unrelated tricks but a structured space defined by two questions—how CLIP is used (its prompts are tuned, or its encoders are used as a feature extractor) and what information is available at test time (labeled source data, a source model, or nothing but target data). The survey claims that every current approach sits in one of these cells, that the closed-set, partial-set, open-set, and open-partial-set label-space relationships further refine each cell, and that the cells expose both mature method families and genuinely underexplored settings.

Load-bearing premise

The whole map stands on the assumption that the literature the survey chose to include is complete and that each method was placed correctly in its cell, but the paper does not state its search queries, inclusion rules, or cutoff date.

Editorial extensions

If this is right

  • A newcomer can choose their method family by answering two questions: can I touch source data, and is my target label space closed, partial, open, or mixed?
  • The taxonomy turns source-free into two precise settings—source-data-free (a source-trained model exists) and source-fully-free (no source domain exists)—so methods built for one setting are not mistakenly evaluated under the other.
  • The open-set and open-partial-set cells share a common evaluation protocol, the harmonic mean of known-class accuracy and unknown-class detection, which makes results across methods in those cells comparable.
  • CLIP's zero-shot ability shifts the default DG and DA recipe from training a task-specific classifier to choosing, prompting, or lightly adapting a frozen foundation model, which changes what counts as a parameter-efficient method.
  • Underexplored cells in the taxonomy, such as multi-source open-set domain generalization, become an explicit call for new work rather than an accident of the literature.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the survey does not report a search protocol or cutoff date, its comprehensive label is best treated as a working hypothesis; a reader who needs completeness should independently re-run a literature search before relying on the map.
  • The taxonomy suggests a testable diagnostic: if it is a true organizing structure, methods within one cell should share failure modes and design ingredients more than methods across cells, which quantitative benchmark analysis could verify.
  • The same two-axis scheme could plausibly extend to other contrastive vision-language models, which would show whether the structure is about CLIP specifically or about the general design space.
  • The scenario grid could be turned into a decision tool: given a data budget and a known label overlap, the table already names the method family to start from, which is more actionable than the survey explicitly claims.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript is a survey of methods that use CLIP for domain generalization (DG) and domain adaptation (DA). It proposes a taxonomy: for DG, prompt optimization versus CLIP as a backbone/encoder, with source-available and source-free branches; for DA, source-available versus source-free (further split into source-fully-free and source-data-free), with scenario labels closed/partial/open/open-partial. It also compiles common benchmarks, evaluation metrics, challenges, and future research directions. The paper claims to be the first comprehensive CLIP-focused review for DG/DA and maintains a living GitHub repository for updated references.

Significance. If the coverage and categorization are accurate, this survey would be a valuable reference for an active and rapidly growing area. Its strengths are the systematic scenario decomposition (notably the distinction between source-fully-free and source-data-free adaptation), a compact summary table of settings, a roadmap/timeline of the field, and a broad bibliography of recent CLIP-based work. Because the value of a survey of this type resides almost entirely in the reliability of its coverage and grouping, the load-bearing claim of comprehensiveness needs to be supported by a stated methodology, and several internal inconsistencies need to be resolved. These issues are substantive but fixable within the manuscript's scope.

major comments (4)
  1. [Section 1 (Contributions)] The central claim that this is 'the first comprehensive review' is not supported by a stated methodology: the manuscript reports no search databases, query strings, inclusion/exclusion criteria, or coverage cutoff date for the literature it surveys. Without such a protocol, the reader cannot verify that the cited set is complete or that its composition is not biased toward particular venues or time windows. Please add a methodology paragraph (in Section 1 or Section 5) specifying the search procedure, or alternatively soften the claim to 'a structured review' of the area.
  2. [Fig. 3 and Section 4.2.1.3] Fig. 3 places UOTA under 'DG → CLIP as Backbone or Encoder → Source-Available → Single-Source & Multi-Source', whereas Section 4.2.1.3 describes UOTA as a Source-Fully-Free Open-Set UDA method. These two locations are mutually exclusive under the paper's own taxonomy, so the roadmap and the body text cannot both be correct. Please determine UOTA's actual setting and correct the inconsistent component.
  3. [Table 2] In Table 2, the row labeled 'SFF-OSUDA a.k.a. OPS-UFT' duplicates the abbreviation SFF-OSUDA used in the immediately preceding row; the intended label is presumably 'SFF-OPSUDA a.k.a. OPS-UFT'. Since Table 2 is the compact reference for all scenario types, this duplication undermines confidence in the summary and must be corrected.
  4. [Section 4.1.2.1] Section 4.1.2.1, 'Multi-Source Closed-Set Unsupervised Domain Adaptation (MS-CSUDA)', opens with 'In SS-CSUDA, the label space...', which is a copy-paste from the single-source definition. As written, the formal definition for the multi-source case is not actually provided, weakening the definitional foundation of Section 4. The introductory sentence should be rewritten to define the MS-CSUDA setting explicitly.
minor comments (5)
  1. [Section 5.1 heading] The heading '5.1 Common Bechmarks' contains a typo; it should read 'Benchmarks'.
  2. [Table 3] Table 3 lists the Office-31 dataset with a URL that points to the Office-Home dataset page, and 'StandfordCars' should be 'StanfordCars'; please verify all dataset links and names.
  3. [Definition 10 (Section 2.2.4)] The contrastive loss formula in Definition 10 includes an extra denominator term exp(v_j^T t_i) that does not correspond to the standard CLIP loss; please check the equation against the cited source.
  4. [Section 2.2.4] The text states that CLIP 'is based on a transformer architecture', which is imprecise because CLIP encoders can be either Vision Transformers or ResNets; please qualify the statement.
  5. [Abstract] The GitHub URL in the abstract contains spaces ('Survey on CLIP-Powered...') and should be formatted as a single hyperlink.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this is a survey whose claims are organizational, not derived from fitted inputs or self-citation chains.

full rationale

The paper is a literature survey, not a derivation. It claims to provide the first comprehensive review of CLIP for domain generalization (DG) and domain adaptation (DA), and organizes existing methods into taxonomies (prompt optimization vs. backbone use for DG; source-available vs. source-free for DA). There is no equation in which an output is defined in terms of its own prediction, no parameter fitted to a subset of data and then reported as a prediction, and no load-bearing uniqueness theorem imported from the authors' prior work. The taxonomy is asserted from the surveyed literature rather than derived from first principles, so it cannot be circular in the sense of a derivation chain collapsing into its inputs. The reader's noted weaknesses—absence of a search protocol, misplacement of UOTA in Fig. 3 relative to Section 4.2.1.3, the duplicated SFF-OSUDA row in Table 2, and the 'In SS-CSUDA' typo in Section 4.1.2.1—are accuracy, consistency, and completeness concerns about the survey's coverage and organization, not circularity. They do not make the survey's central claim equivalent to its own inputs. No self-citation is used as evidence for any substantive claim, and no result is renamed as a novel finding when it is in fact the input. Accordingly, the honest finding is no significant circularity, with score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters or invented entities. It relies on standard DG/DA definitions and on CLIP's pretraining as an external backbone, and it imposes an organizational taxonomy that is the paper's contribution but is not independently verified.

assumptions (3)
  • domain assumption The definitions of DG and DA in Section 2.2 (Definitions 1 and 2) are standard and complete.
    These definitions are taken from prior literature and frame the entire survey.
  • domain assumption CLIP's zero-shot and robustness properties transfer to DG and DA benchmarks.
    The paper's motivation rests on this premise, stated in Section 1 and Figure 1.
  • ad hoc to paper The taxonomy divisions (prompt optimization vs backbone for DG; SA vs SF for DA) are exhaustive and mutually exclusive.
    This is the paper's own organizing structure; it is plausible but not proven.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey." pith.science (2026). https://pith.science/paper/JH3Z3UR7

@misc{pith2026250414280,
  author       = {Pith},
  title        = {Pith review of: CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JH3Z3UR7}},
  note         = {Machine review of arXiv:2504.14280}
}
read the original abstract

As machine learning evolves, domain generalization (DG) and domain adaptation (DA) have become crucial for enhancing model robustness across diverse environments. Contrastive Language-Image Pretraining (CLIP) plays a significant role in these tasks, offering powerful zero-shot capabilities that allow models to perform effectively in unseen domains. However, there remains a significant gap in the literature, as no comprehensive survey currently exists that systematically explores the applications of CLIP in DG and DA, highlighting the necessity for this review. This survey presents a comprehensive review of CLIP's applications in DG and DA. In DG, we categorize methods into optimizing prompt learning for task alignment and leveraging CLIP as a backbone for effective feature extraction, both enhancing model adaptability. For DA, we examine both source-available methods utilizing labeled source data and source-free approaches primarily based on target domain data, emphasizing knowledge transfer mechanisms and strategies for improved performance across diverse contexts. Key challenges, including overfitting, domain diversity, and computational efficiency, are addressed, alongside future research opportunities to advance robustness and efficiency in practical applications. By synthesizing existing literature and pinpointing critical gaps, this survey provides valuable insights for researchers and practitioners, proposing directions for effectively leveraging CLIP to enhance methodologies in domain generalization and adaptation. Ultimately, this work aims to foster innovation and collaboration in the quest for more resilient machine learning models that can perform reliably across diverse real-world scenarios. A more up-to-date version of the papers is maintained at: https://github.com/jindongli-Ai/Survey_on_CLIP-Powered_Domain_Generalization_and_Adaptation.

Figures

Figures reproduced from arXiv: 2504.14280 by the authors.

Figure 1
Figure 1. The characteristics of CLIP and its perfect fit for Domain Gener [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Structure of this paper with representative works. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 4
Figure 4. Comparison of domain generalization (DG) and domain adapta [PITH_FULL_IMAGE:figures/full_fig_p003_4.png] view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: Comparison of single-source (SS) and multi-source (MS) sce [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 8
Figure 8. Figure 8: The illustration of the (a) training process of CLIP (taking 4 class [PITH_FULL_IMAGE:figures/full_fig_p005_8.png]
Figure 9
Figure 9. Figure 9: The different ways about prompt learning optimization [ [PITH_FULL_IMAGE:figures/full_fig_p006_9.png]
Figure 11
Figure 11. Figure 11: The scenario of Multi-Source Open-Partial-Set Unsupervised [PITH_FULL_IMAGE:figures/full_fig_p009_11.png]
Figure 12
Figure 12. Figure 12: Percentage of cited papers in different sections. [PITH_FULL_IMAGE:figures/full_fig_p014_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hyperbolic Multimodal Continual Learning

    cs.LG 2026-08 conditional novelty 5.0 of 10

    The paper proposes projecting continual updates away from old-task spatial directions in hyperbolic multimodal models, claims this is theoretically required to prevent forgetting, but the necessity claim is unproven a...

Reference graph

Works this paper leans on

209 extracted references · 40 canonical work pages · cited by 1 Pith paper

  1. [1]

    Addepalli, A

    S. Addepalli, A. R. Asokan, L. Sharma, and R. V . Babu. Leveraging vision-language models for improving do- main generalization in image classification. In Proc. of CVPR, pages 23922–23932, 2024

  2. [2]

    E. L. Aleixo, J. G. Colonna, M. Cristo, and E. Fer- nandes. Catastrophic forgetting in deep learning: a comprehensive taxonomy. arXiv:2312.10549, 2023

  3. [3]

    E. Ali, S. Silva, and M. H. Khan. Dpa: Dual pro- totypes alignment for unsupervised adaptation of vision-language models. arXiv:2408.08855, 2024

  4. [4]

    Bahng, A

    H. Bahng, A. Jahanian, S. Sankaranarayanan, and P . Isola. Visual prompting: Modifying pixel space to adapt pre-trained models. arXiv:2203.17274, 3(11-12): 3, 2022

  5. [5]

    S. Bai, M. Zhang, W. Zhou, S. Huang, Z. Luan, D. Wang, and B. Chen. Prompt-based distribution alignment for unsupervised domain adaptation. In Proc. of AAAI, volume 38, pages 729–737, 2024

  6. [6]

    S. Bai, Y. Zhang, W. Zhou, Z. Luan, and B. Chen. Soft prompt generation for domain generalization. arXiv:2404.19286, 2024

  7. [7]

    Beery, G

    S. Beery, G. Van Horn, and P . Perona. Recognition in terra incognita. In Proc. of ECCV, pages 456–473, 2018

  8. [8]

    S. Bose, A. Jha, E. Fini, M. Singha, E. Ricci, and B. Banerjee. Stylip: Multi-scale style-conditioned prompt learning for clip-based domain generalization. In Proc. of WACV, pages 5542–5552, 2024

Show all 209 references
  1. [9]

    Bossard, M

    L. Bossard, M. Guillaumin, and L. Van Gool. Food- 101–mining discriminative components with random forests. In Proc. of ECCV , pages 446–461. Springer, 2014

  2. [10]

    Bucci, F

    S. Bucci, F. C. Borlino, B. Caputo, and T. Tommasi. Distance-based hyperspherical classification for multi- source open-set domain adaptation. In Proc. of WACV, pages 1119–1128, 2022

  3. [11]

    H. Cai, C. Gan, L. Zhu, and S. Han. Tinytl: Reduce memory, not parameters for efficient on-device learn- ing. Proc. of NeurIPS, 33:11285–11297, 2020

  4. [12]

    J. Cha, S. Chun, K. Lee, H.-C. Cho, S. Park, Y. Lee, and S. Park. Swad: Domain generalization by seeking flat minima. Proc. of NeurIPS, 34:22405–22418, 2021

  5. [13]

    X. Che, H. Zuo, J. Lu, and D. Chen. Fuzzy multioutput transfer learning for regression. IEEE Trans. on Fuzzy Systems, 30(7):2438–2451, 2021

  6. [14]

    G. Chen, W. Yao, X. Song, X. Li, Y. Rao, and K. Zhang. Prompt learning with optimal transport for vision- language models. OpenReview, 2022

  7. [15]

    G. Chen, W. Yao, X. Song, X. Li, Y. Rao, and K. Zhang. Plot: Prompt learning with optimal transport for vision-language models. arXiv:2210.01253, 2022

  8. [16]

    H. Chen, X. Han, Z. Wu, and Y.-G. Jiang. Multi- prompt alignment for multi-source unsupervised do- main adaptation. Proc. of NeurIPS , 36:74127–74139, 2023

  9. [17]

    S. Chen, C. Ge, Z. Tong, J. Wang, Y. Song, J. Wang, and P . Luo. Adaptformer: Adapting vision transformers for scalable visual recognition. Proc. of NeurIPS , 35: 16664–16678, 2022

  10. [18]

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In Proc. of ICML , pages 1597–1607. PMLR, 2020

  11. [19]

    Chen and B

    Z. Chen and B. Liu. Continual learning and catas- trophic forgetting. In Lifelong Machine Learning, pages 55–75. Springer, 2018. 15

  12. [20]

    Z. Chen, Y. Duan, W. Wang, J. He, T. Lu, J. Dai, and Y. Qiao. Vision transformer adapter for dense predictions. arXiv:2205.08534, 2022

  13. [21]

    Z. Chen, W. Wang, Z. Zhao, F. Su, A. Men, and Y. Dong. Instance paradigm contrastive learning for domain generalization. IEEE TCSVT, 34(2):1032–1042, 2023

  14. [22]

    Z. Chen, W. Wang, Z. Zhao, F. Su, A. Men, and H. Meng. Practicaldg: Perturbation distillation on vision-language models for hybrid domain general- ization. In Proc. of CVPR, pages 23501–23511, 2024

  15. [23]

    Cheng, Z

    D. Cheng, Z. Xu, X. Jiang, N. Wang, D. Li, and X. Gao. Disentangled prompt representation for domain gen- eralization. In Proc. of CVPR, pages 23595–23604, 2024

  16. [24]

    E. Cho, J. Kim, and H. J. Kim. Distribution-aware prompt tuning for vision-language models. In Proc. of ICCV, pages 22004–22013, 2023

  17. [25]

    J. Cho, G. Nam, S. Kim, H. Yang, and S. Kwak. Promptstyler: Prompt-driven style generation for source-free domain generalization. In Proc. of ICCV , pages 15702–15712, 2023

  18. [26]

    J. Choi, M. Jeong, T. Kim, and C. Kim. Pseudo- labeling curriculum for unsupervised domain adap- tation. arXiv:1908.00262, 2019

  19. [27]

    S. Choi, D. Das, S. Choi, S. Yang, H. Park, and S. Yun. Progressive random convolutions for single domain generalization. In Proc. of CVPR , pages 10312–10322, 2023

  20. [28]

    Cimpoi, S

    M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi. Describing textures in the wild. In Proc. of CVPR, pages 3606–3613, 2014

  21. [29]

    E. D. Cubuk, B. Zoph, J. Shlens, and Q. V . Le. Ran- daugment: Practical automated data augmentation with a reduced search space. In Proc. of CVPR Work- shop, pages 702–703, 2020

  22. [30]

    Das and P

    A. Das and P . Rad. Opportunities and challenges in explainable artificial intelligence (xai): A survey. arXiv:2006.11371, 2020

  23. [31]

    De Lange, R

    M. De Lange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, and T. Tuytelaars. A continual learning survey: Defying forgetting in clas- sification tasks. IEEE TP AMI, 44(7):3366–3385, 2021

  24. [32]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. of CVPR, pages 248–255. IEEE, 2009

  25. [33]

    Desai and J

    K. Desai and J. Johnson. Virtex: Learning visual representations from textual annotations. In Proc. of CVPR, pages 11162–11173, 2021

  26. [34]

    X. Dong, J. Bao, Y. Zheng, T. Zhang, D. Chen, H. Yang, M. Zeng, W. Zhang, L. Yuan, D. Chen, et al. Maskclip: Masked self-distillation advances contrastive language-image pretraining. In Proc. of CVPR, pages 10995–11005, 2023

  27. [35]

    Z. Du, X. Li, F. Li, K. Lu, L. Zhu, and J. Li. Domain- agnostic mutual prompting for unsupervised domain adaptation. In Proc. of CVPR, pages 23375–23384, 2024

  28. [36]

    Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias

    Chen Fang, Ye Xu, and Daniel N Rockmore. Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias. In Proc. of ICCV , pages 1657–1664, 2013

  29. [37]

    Fei-Fei, R

    L. Fei-Fei, R. Fergus, and P . Perona. Learning gener- ative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In Proc. of CVPR Workshop, pages 178–178. IEEE, 2004

  30. [38]

    R. Feng, T. Yu, X. Jin, X. Yu, L. Xiao, and Z. Chen. Rethinking domain adaptation and generalization in the era of clip. In Proc. of ICIP, pages 2585–2591. IEEE, 2024

  31. [39]

    F ¨urst, E

    A. F ¨urst, E. Rumetshofer, J. Lehner, V . T. Tran, F. Tang, H. Ramsauer, D. Kreil, M. Kopp, G. Klambauer, A. Bitto, et al. Cloob: Modern hopfield networks with infoloob outperform clip. Proc. of NeurIPS , 35:20450– 20468, 2022

  32. [40]

    R. Gal, O. Patashnik, H. Maron, A. H. Bermano, G. Chechik, and D. Cohen-Or. Stylegan-nada: Clip- guided domain adaptation of image generators. ACM TOG, 41(4):1–13, 2022

  33. [41]

    Gal and Z

    Y. Gal and Z. Ghahramani. Dropout as a bayesian ap- proximation: Representing model uncertainty in deep learning. In Proc. of ICML, pages 1050–1059, 2016

  34. [42]

    Z. Gan, L. Li, C. Li, L. Wang, Z. Liu, J. Gao, et al. Vision-language pre-training: Basics, recent advances, and future trends. FTCGV, 14(3–4):163–352, 2022

  35. [43]

    J. Gao, J. Ruan, S. Xiang, Z. Yu, K. Ji, M. Xie, T. Liu, and Y. Fu. Lamm: Label alignment for multi-modal prompt learning. In Proc. of AAAI , volume 38, pages 1815–1823, 2024

  36. [44]

    P . Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao. Clip-adapter: Better vision- language models with feature adapters. IJCV, 132(2): 581–595, 2024

  37. [45]

    T. Gao, A. Fisch, and D. Chen. Making pre- trained language models better few-shot learners. arXiv:2012.15723, 2020

  38. [46]

    C. Ge, R. Huang, M. Xie, Z. Lai, S. Song, S. Li, and G. Huang. Domain adaptation via prompt learning. IEEE TNNLS, 2023

  39. [47]

    J. Gu, Z. Han, S. Chen, A. Beirami, B. He, G. Zhang, R. Liao, Y. Qin, V . Tresp, and P . Torr. A system- atic survey of prompt engineering on vision-language foundation models. arXiv:2307.12980, 2023

  40. [48]

    J. Guo, L. Qi, and Y. Shi. Domaindrop: Suppressing domain-sensitive channels for domain generalization. In Proc. of CVPR, pages 19114–19124, 2023

  41. [49]

    K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick. Momen- tum contrast for unsupervised visual representation learning. In Proc. of CVPR, pages 9729–9738, 2020

  42. [50]

    Helber, B

    P . Helber, B. Bischke, A. Dengel, and D. Borth. Eu- rosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE J- STAR, 12(7):2217–2226, 2019

  43. [51]

    Hendrycks, S

    D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proc. of ICCV , pages 8340–8349, 2021

  44. [52]

    Hendrycks, K

    D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song. Natural adversarial examples. In Proc. of CVPR, pages 15262–15271, 2021

  45. [53]

    R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P . Bachman, A. Trischler, and Y. Bengio. 16 Learning deep representations by mutual information estimation and maximization. Proc. of ICLR, 2018

  46. [54]

    Houlsby, A

    N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly. Parameter-efficient transfer learning for nlp. In Proc. of ICML, pages 2790–2799. PMLR, 2019

  47. [55]

    E. J. Hu, Y. Shen, P . Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen. Lora: Low-rank adaptation of large language models. arXiv:2106.09685, 2021

  48. [56]

    X. Hu, K. Zhang, L. Xia, A. Chen, J. Luo, Y. Sun, K. Wang, N. Qiao, X. Zeng, M. Sun, et al. Reclip: Refine contrastive language image pre-training with source-free domain adaptation. In Proc. of WACV , pages 2994–3003, 2024

  49. [57]

    Huang, J

    T. Huang, J. Chu, and F. Wei. Unsuper- vised prompt learning for vision-language models. arXiv:2204.03649, 2022

  50. [58]

    Huang, C

    Y. Huang, C. Du, Z. Xue, X. Chen, H. Zhao, and L. Huang. What makes multi-modal learning better than single (provably). In Proc. NeurIPS, volume 34, pages 10944–10956, 2021

  51. [59]

    Huang, H

    Z. Huang, H. Wang, E. P . Xing, and D. Huang. Self- challenging improves cross-domain generalization. In Proc. of ECCV, pages 124–140. Springer, 2020

  52. [60]

    Huang, A

    Z. Huang, A. Zhou, Z. Ling, M. Cai, H. Wang, and Y. J. Lee. A sentence speaks a thousand images: Domain generalization through distilling clip with language guidance. In Proc. of ICCV, pages 11685–11695, 2023

  53. [61]

    Ilharco, M

    G. Ilharco, M. Wortsman, S. Y. Gadre, S. Song, H. Ha- jishirzi, S. Kornblith, A. Farhadi, and L. Schmidt. Patching open-vocabulary models by interpolating weights. Proc. of NeurIPS, 35:29262–29277, 2022

  54. [62]

    S. Ioffe. Batch normalization: Accelerating deep net- work training by reducing internal covariate shift. arXiv:1502.03167, 2015

  55. [63]

    C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Proc. of ICML , pages 4904–4916. PMLR, 2021

  56. [64]

    M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. Belongie, B. Hariharan, and S.-N. Lim. Visual prompt tuning. In Proc. of ECCV, pages 709–727. Springer, 2022

  57. [65]

    Jiang, F

    Z. Jiang, F. F. Xu, J. Araki, and G. Neubig. How can we know what language models know? Trans. of ACL, 8:423–438, 2020

  58. [66]

    C. Ju, T. Han, K. Zheng, Y. Zhang, and W. Xie. Prompt- ing visual-language models for efficient video under- standing. In Proc. of ECCV , pages 105–124. Springer, 2022

  59. [67]

    J. Kang, S. Lee, N. Kim, and S. Kwak. Style neophile: Constantly seeking novel styles for domain general- ization. In Proc. of CVPR, pages 7130–7140, 2022

  60. [68]

    Katsumata, I

    K. Katsumata, I. Kishida, A. Amma, and H. Nakayama. Open-set domain generalization via metric learning. In Proc. of ICIP , pages 459–463, 2021

  61. [69]

    Kemker, M

    R. Kemker, M. McClure, A. Abitino, T. Hayes, and C. Kanan. Measuring catastrophic forgetting in neural networks. In Proc. AAAI, volume 32, 2018

  62. [70]

    M. U. Khattak, H. Rasheed, M. Maaz, S. Khan, and F. S. Khan. Maple: Multi-modal prompt learning. In Proc. of CVPR, pages 19113–19122, 2023

  63. [71]

    M. U. Khattak, S. T. Wasim, M. Naseer, S. Khan, M.- H. Yang, and F. S. Khan. Self-regulating prompts: Foundational model adaptation without forgetting. In Proc. of ICCV, pages 15190–15200, 2023

  64. [72]

    G. Kim, T. Kwon, and J. C. Ye. Diffusionclip: Text- guided diffusion models for robust image manipula- tion. In Proc. of CVPR, pages 2426–2435, 2022

  65. [73]

    J. Kim, K. Ryoo, J. Seo, G. Lee, D. Kim, H. Cho, and S. Kim. Semi-supervised learning of semantic correspondence with pseudo-labels. In Proc. of CVPR, pages 19699–19709, 2022

  66. [74]

    Krause, M

    J. Krause, M. Stark, J. Deng, and L. Fei-Fei. 3d object representations for fine-grained categorization. In Proc. of ICCV Workshop, pages 554–561, 2013

  67. [75]

    Krizhevsky, G

    A. Krizhevsky, G. Hinton, et al. Learning multiple layers of features from tiny images. 2009

  68. [76]

    Kwon and J

    G. Kwon and J. C. Ye. Clipstyler: Image style transfer with a single text condition. In Proc. of CVPR , pages 18062–18071, 2022

  69. [77]

    Lafon, E

    M. Lafon, E. Ramzi, C. Rambour, N. Audebert, and N. Thome. Gallop: Learning global and local prompts for vision-language models. arXiv:2407.01400, 2024

  70. [78]

    Z. Lai, N. Vesdapunt, N. Zhou, J. Wu, C. P . Huynh, X. Li, K. K. Fu, and C.-N. Chuah. Padclip: Pseudo- labeling with adaptive debiasing in clip for unsuper- vised domain adaptation. In Proc. of ICCV , pages 16155–16165, 2023

  71. [79]

    Z. Lai, H. Bai, H. Zhang, X. Du, J. Shan, Y. Yang, C.-N. Chuah, and M. Cao. Empowering unsupervised do- main adaptation with large-scale pre-trained vision- language models. In Proc. of WACV, pages 2691–2701, 2024

  72. [80]

    K. Lee, S. Kim, and S. Kwak. Cross-domain ensemble distillation for domain generalization. In Proc. of ECCV, pages 1–20. Springer, 2022

  73. [81]

    S. Lee, J. Bae, and H. Y. Kim. Decompose, adjust, compose: Effective normalization by playing with fre- quency for domain generalization. In Proc. of CVPR , pages 11776–11785, 2023

  74. [82]

    Lester, R

    B. Lester, R. Al-Rfou, and N. Constant. The power of scale for parameter-efficient prompt tuning. arXiv:2104.08691, 2021

  75. [83]

    Deeper, broader and artier domain gen- eralization

    Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M Hospedales. Deeper, broader and artier domain gen- eralization. In Proc. of ICCV, pages 5542–5550, 2017

  76. [84]

    H. Li, S. J. Pan, S. Wang, and A. C. Kot. Domain generalization with adversarial feature learning. In Proc. of CVPR, pages 5400–5409, 2018

  77. [85]

    J. Li, Z. Yu, Z. Du, L. Zhu, and H. T. Shen. A com- prehensive survey on source-free domain adaptation. IEEE TP AMI, 2024

  78. [86]

    K. Li, J. Lu, H. Zuo, and G. Zhang. Attention-bridging ts fuzzy rules for universal multi-domain adaptation without source data. In Proc. of FUZZ-IEEE, pages 1–6. IEEE, 2023

  79. [87]

    K. Li, J. Lu, H. Zuo, and G. Zhang. Source-free multidomain adaptation with fuzzy rule-based deep neural networks. IEEE Trans. on Fuzzy Systems, 31(12): 4180–4194, 2023. 17

  80. [88]

    K. Li, J. Lu, H. Zuo, and G. Zhang. Source-free multidomain adaptation with fuzzy rule-based deep neural networks. IEEE Trans. on Fuzzy Systems, 31(12): 4180–4194, 2023

  81. [89]

    X. Li, Y. Li, Z. Du, F. Li, K. Lu, and J. Li. Split to merge: Unifying separated modalities for unsupervised do- main adaptation. In Proc. of CVPR, pages 23364–23374, 2024

  82. [90]

    X. L. Li and P . Liang. Prefix-tuning: Optimizing continuous prompts for generation. arXiv:2101.00190, 2021

  83. [91]

    Y. Li, Y. Cao, J. Li, Q. Wang, and S. Wang. Data-efficient clip-powered dual-branch networks for source-free unsupervised domain adaptation. arXiv:2410.15811, 2024

  84. [92]

    Y.-J. Li, X. Dai, C.-Y. Ma, Y.-C. Liu, K. Chen, B. Wu, Z. He, K. Kitani, and P . Vajda. Cross-domain adaptive teacher for object detection. In Proc. of CVPR , pages 7581–7590, 2022

  85. [93]

    Liang, R

    J. Liang, R. He, and T. Tan. A comprehensive sur- vey on test-time adaptation under distribution shifts. IJCV, pages 1–34, 2024

  86. [94]

    Liang, L

    J. Liang, L. Sheng, Z. Wang, R. He, and T. Tan. Realistic unsupervised clip fine-tuning with universal entropy optimization. In Proc. of ICML, 2024

  87. [95]

    F. Liu, J. Lu, and G. Zhang. Unsupervised hetero- geneous domain adaptation via shared fuzzy equiv- alence relations. IEEE Trans. on Fuzzy Systems , 26(6): 3555–3568, 2018

  88. [96]

    F. Liu, G. Zhang, and J. Lu. Multisource heteroge- neous unsupervised domain adaptation via fuzzy re- lation neural networks. IEEE Trans. on Fuzzy Systems, 29(11):3308–3322, 2020

  89. [97]

    H. Liu, J. Wang, and M. Long. Cycle self-training for domain adaptation. Proc. of NeurIPS, 34:22968–22981, 2021

  90. [98]

    Pre-train, prompt, and predict: A systematic survey of prompt- ing methods in natural language processing

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompt- ing methods in natural language processing. ACM Computing Surveys, 55(9):1–35, 2023

  91. [99]

    M. Long, H. Zhu, J. Wang, and M. I. Jordan. Deep transfer learning with joint adaptation networks. In Proc. of ICML, pages 2208–2217. PMLR, 2017

  92. [100]

    M. Long, Z. Cao, J. Wang, and M. I. Jordan. Con- ditional adversarial domain adaptation. Proc. of NeurIPS, 31, 2018

  93. [101]

    S. Long, L. Wang, Z. Zhao, Z. Tan, Y. Wu, S. Wang, and J. Wang. Training-free unsupervised prompt for vision-language models. arXiv:2404.16339, 2024

  94. [102]

    Loshchilov

    I. Loshchilov. Decoupled weight decay regularization. arXiv:1711.05101, 2017

  95. [103]

    J. Lu, H. Zuo, and G. Zhang. Fuzzy multiple-source transfer learning. IEEE Trans. on Fuzzy Systems, 28(12): 3418–3431, 2019

  96. [104]

    Y. Lu, J. Liu, Y. Zhang, Y. Liu, and X. Tian. Prompt distribution learning. In Proc. of CVPR , pages 5206– 5215, 2022

  97. [105]

    F. Lv, J. Liang, S. Li, B. Zang, C. H. Liu, Z. Wang, and D. Liu. Causality inspired representation learning for domain generalization. In Proc. of CVPR, pages 8046– 8056, 2022

  98. [106]

    W. Ma, S. Li, J. Zhang, C. H. Liu, J. Kang, Y. Wang, and G. Huang. Borrowing knowledge from pre-trained language model: A new data-efficient visual learning paradigm. In Proc. of ICCV, pages 18786–18797, 2023

  99. [107]

    S. Maji, E. Rahtu, J. Kannala, M. Blaschko, and A. Vedaldi. Fine-grained visual classification of air- craft. arXiv:1306.5151, 2013

  100. [108]

    K. Mei, C. Zhu, J. Zou, and S. Zhang. Instance adaptive self-training for unsupervised domain adap- tation. In Proc. of ECCV, pages 415–430. Springer, 2020

  101. [109]

    Y. Min, K. Ryoo, B. Kim, and T. Kim. Uota: Un- supervised open-set task adaptation using a vision- language foundation model. In Proc. of ICML Work- shop, 2023

  102. [110]

    M. J. Mirza, L. Karlinsky, W. Lin, H. Possegger, M. Kozinski, R. Feris, and H. Bischof. Lafter: Label- free tuning of zero-shot classifier using language and unlabeled image collections. Proc. of NeurIPS, 36, 2024

  103. [111]

    Miyai, Q

    A. Miyai, Q. Yu, G. Irie, and K. Aizawa. Locoop: Few- shot out-of-distribution detection via prompt learn- ing. Proc. of NeurIPS, 36:76298–76310, 2023

  104. [112]

    Miyai, Q

    A. Miyai, Q. Yu, G. Irie, and K. Aizawa. Zero-shot in-distribution detection in multi-object settings using vision-language foundation models. arXiv:2304.04521, 2023

  105. [113]

    Monga, S

    M. Monga, S. K. Giroh, A. Jha, M. Singha, B. Baner- jee, and J. Chanussot. Cosmo: Clip talks on open- set multi-target domain adaptation. arXiv:2409.00397, 2024

  106. [114]

    C. Mou, X. Wang, L. Xie, Y. Wu, J. Zhang, Z. Qi, and Y. Shan. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Proc. of AAAI, volume 38, pages 4296–4304, 2024

  107. [115]

    Nilsback and A

    M.-E. Nilsback and A. Zisserman. Automated flower classification over a large number of classes. In 2008 Sixth Indian Conf. on CVGIP , pages 722–729. IEEE, 2008

  108. [116]

    H. Niu, H. Li, F. Zhao, and B. Li. Domain-unified prompt representations for source-free domain gener- alization. arXiv:2209.14926, 2022

  109. [117]

    Noguchi and S

    M. Noguchi and S. Shirakawa. Simple domain gen- eralization methods are strong baselines for open domain generalization. In Proc. of IJCNN , pages 1–8, 2024

  110. [118]

    Oppenlaender

    J. Oppenlaender. The creativity of text-to-image gen- eration. In Proc. of MindTrek, pages 192–202, 2022

  111. [119]

    O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V . Jawahar. Cats and dogs. In Proc. of CVPR, pages 3498–

  112. [120]

    Patashnik, Z

    O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischinski. Styleclip: Text-driven manipulation of stylegan imagery. In Proc. of ICCV , pages 2085–2094, 2021

  113. [121]

    Visda: The visual domain adaptation challenge

    Xingchao Peng, Ben Usman, Neela Kaushik, Judy Hoffman, Dequan Wang, and Kate Saenko. Visda: The visual domain adaptation challenge. arXiv:1710.06924, 2017

  114. [122]

    Moment matching for 18 multi-source domain adaptation

    Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for 18 multi-source domain adaptation. In Proc. of ICCV , pages 1406–1415, 2019

  115. [123]

    Petroni, T

    F. Petroni, T. Rockt ¨aschel, P . Lewis, A. Bakhtin, Y. Wu, A. H. Miller, and S. Riedel. Language models as knowledge bases? arXiv:1909.01066, 2019

  116. [124]

    Y. Qiao, K. Li, J. Lin, R. Wei, C. Jiang, Y. Luo, and H. Yang. Robust domain generalization for multi- modal object recognition. In Proc. of AIEA, pages 392–

  117. [125]

    Lfpt5: A unified frame- work for lifelong few-shot language learning based on prompt tuning of t5

    Chengwei Qin and Shafiq Joty. Lfpt5: A unified frame- work for lifelong few-shot language learning based on prompt tuning of t5. arXiv:2110.07298, 2021

  118. [126]

    S. Qu, Y. Pan, G. Chen, T. Yao, C. Jiang, and T. Mei. Modality-agnostic debiasing for single domain gener- alization. In Proc. of CVPR, pages 24142–24151, 2023

  119. [127]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P . Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In Proc. of ICML , pages 8748–

  120. [128]

    Ramesh, M

    A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever. Zero-shot text- to-image generation. In Proc. of ICML , pages 8821–

  121. [129]

    Y. Rao, W. Zhao, G. Chen, Y. Tang, Z. Zhu, G. Huang, J. Zhou, and J. Lu. Denseclip: Language-guided dense prediction with context-aware prompting. In Proc. of CVPR, pages 18082–18091, 2022

  122. [130]

    Recht, R

    B. Recht, R. Roelofs, L. Schmidt, and V . Shankar. Do imagenet classifiers generalize to imagenet? In Proc. of ICML, pages 5389–5400. PMLR, 2019

  123. [131]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P . Esser, and B. Ommer. High-resolution image synthesis with la- tent diffusion models. In Proc. of CVPR, pages 10684– 10695, 2022

  124. [132]

    Adapting visual category models to new do- mains

    Kate Saenko, Brian Kulis, Mario Fritz, and Trevor Darrell. Adapting visual category models to new do- mains. In Proc. of ECCV 2010, pages 213–226. Springer, 2010

  125. [133]

    Saharia, W

    C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Proc. of NeurIPS, 35:36479–36494, 2022

  126. [134]

    M. B. Sariyildiz, J. Perez, and D. Larlus. Learning visual representations with caption annotations. In Proc. of ECCV, pages 153–170. Springer, 2020

  127. [135]

    Schulhoff, M

    S. Schulhoff, M. Ilie, N. Balepur, K. Kahadze, A. Liu, C. Si, Y. Li, A. Gupta, H. Han, S. Schulhoff, et al. The prompt report: A systematic survey of prompting techniques. arXiv:2406.06608, 2024

  128. [136]

    Schwalbe and B

    G. Schwalbe and B. Finzel. A comprehensive taxon- omy for explainable artificial intelligence: a systematic survey of surveys on methods and concepts. DMKD, 38(5):3043–3101, 2024

  129. [137]

    Shankar, V

    S. Shankar, V . Piratla, S. Chakrabarti, S. Chaudhuri, P . Jyothi, and S. Sarawagi. Generalizing across do- mains via cross-gradient training. arXiv:1804.10745, 2018

  130. [138]

    S. Shen, S. Yang, T. Zhang, B. Zhai, J. E. Gonzalez, K. Keutzer, and T. Darrell. Multitask vision-language prompt tuning. In Proc. of CVPR , pages 5656–5667, 2024

  131. [139]

    K. Shi, J. Lu, Z. Fang, and G. Zhang. Enhancing vision- language models incorporating tsk fuzzy system for domain adaptation. In Proc. of FUZZ-IEEE, pages 1–8, 2024

  132. [140]

    K. Shi, J. Lu, Z. Fang, and G. Zhang. Clip-enhanced unsupervised domain adaptation with consistency regularization. In Proc. of IJCNN, pages 1–8, 2024

  133. [141]

    K. Shi, J. Lu, Z. Fang, and G. Zhang. Unsupervised do- main adaptation enhanced by fuzzy prompt learning. IEEE TFS, 2024

  134. [142]

    T. Shin, Y. Razeghi, R. L. Logan IV , E. Wallace, and S. Singh. Autoprompt: Eliciting knowledge from lan- guage models with automatically generated prompts. arXiv:2010.15980, 2020

  135. [143]

    Y. Shu, Z. Cao, C. Wang, J. Wang, and M. Long. Open domain generalization with domain-augmented meta-learning. In Proc. of CVPR , pages 9624–9633, 2021

  136. [144]

    Y. Shu, X. Guo, J. Wu, X. Wang, J. Wang, and M. Long. Clipood: Generalizing clip to out-of-distributions. In Proc. of ICML, pages 31716–31731. PMLR, 2023

  137. [145]

    Singh, R

    A. Singh, R. Hu, V . Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela. Flava: A foundational language and vision alignment model. In Proc. of CVPR, pages 15638–15650, 2022

  138. [146]

    Singha, H

    M. Singha, H. Pal, A. Jha, and B. Banerjee. Ad-clip: Adapting domains in prompt space using clip. InProc. of ICCV Workshop, pages 4355–4364, 2023

  139. [147]

    Singha, A

    M. Singha, A. Jha, S. Bose, A. Nair, M. Abdar, and B. Banerjee. Unknown prompt the only lacuna: Un- veiling clip’s potential for open domain generaliza- tion. In Proc. of CVPR, pages 13309–13319, 2024

  140. [148]

    K. Sohn, D. Berthelot, N. Carlini, Z. Zhang, H. Zhang, C. A. Raffel, E. D. Cubuk, A. Kurakin, and C.-L. Li. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. Proc. of NeurIPS , 33:596– 608, 2020

  141. [149]

    K. Soomro. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv:1212.0402, 2012

  142. [150]

    Srivastava, G

    N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. JMLR, 15 (1):1929–1958, 2014

  143. [151]

    A. C. Stickland and I. Murray. Bert and pals: Projected attention layers for efficient adaptation in multi-task learning. In Proc. of ICML , pages 5986–5995. PMLR, 2019

  144. [152]

    Sun and K

    B. Sun and K. Saenko. Deep coral: Correlation align- ment for deep domain adaptation. In Proc. of ECCV Workshop, pages 443–450. Springer, 2016

  145. [153]

    X. Sun, P . Hu, and K. Saenko. Dualcoop: Fast adap- tation to multi-label recognition with limited annota- tions. Proc. of NeurIPS, 35:30569–30582, 2022

  146. [154]

    Szegedy, V

    C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In Proc. of CVPR , pages 2818–2826, 2016

  147. [155]

    S. Tang, W. Su, M. Ye, and X. Zhu. Source-free do- main adaptation with frozen multimodal foundation 19 model. In Proc. of CVPR, pages 23711–23720, 2024

  148. [156]

    Y. Tang, Y. Wan, L. Qi, and X. Geng. Dpstyler: Dynamic promptstyler for source-free domain gener- alization. arXiv:2403.16697, 2024

  149. [157]

    Tanwisuth, S

    K. Tanwisuth, S. Zhang, H. Zheng, P . He, and M. Zhou. Pouf: Prompt-oriented unsupervised fine- tuning for large pre-trained models. In Proc. of ICML, pages 33816–33832, 2023

  150. [158]

    L. Tian, M. Ye, L. Zhou, and Q. He. Clip-guided black- box domain adaptation of image classification. Signal, Image and Video Processing, 18(5):4637–4646, 2024

  151. [159]

    Tsimpoukelli, J

    M. Tsimpoukelli, J. L. Menick, S. Cabi, S. M. Eslami, O. Vinyals, and F. Hill. Multimodal few-shot learning with frozen language models. Proc. of NeurIPS , 34: 200–212, 2021

  152. [160]

    Continual learning and catastrophic forgetting

    Gido M van de Ven, Nicholas Soures, and Dhireesha Kudithipudi. Continual learning and catastrophic forgetting. arXiv:2403.05175, 2024

  153. [161]

    Deep hashing network for unsupervised domain adaptation

    Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty, and Sethuraman Panchanathan. Deep hashing network for unsupervised domain adaptation. In Proc. of CVPR, pages 5018–5027, 2017

  154. [162]

    H. Wang, S. Ge, Z. Lipton, and E. P . Xing. Learning robust global representations by penalizing local pre- dictive power. Proc. of NeurIPS, 32, 2019

  155. [163]

    P . Wang, Z. Zhang, Z. Lei, and L. Zhang. Sharpness- aware gradient matching for domain generalization. In Proc. of CVPR, pages 3769–3778, 2023

  156. [164]

    X. Wang, J. Zhang, L. Qi, and Y. Shi. Generalizable decision boundaries: Dualistic meta-learning for open set domain generalization. In Proc. of ICCV , pages 11564–11573, 2023

  157. [165]

    Y. Wang. Survey on deep multi-modal data analytics: Collaboration, rivalry, and fusion. ACM TOMM , 17 (1s):1–25, 2021

  158. [166]

    Z. Wang, L. Zhang, L. Wang, and M. Zhu. Landa: Language-guided multi-source domain adaptation. arXiv:2401.14148, 2024

  159. [167]

    H. Wei, L. Chen, K. Ruan, and L. Li. Low-rank tensor regularized fuzzy clustering for multiview data. IEEE Trans. on Fuzzy Systems, 28(12):3087–3099, 2020

  160. [168]

    Wortsman, G

    M. Wortsman, G. Ilharco, J. W. Kim, M. Li, S. Korn- blith, R. Roelofs, R. G. Lopes, H. Hajishirzi, A. Farhadi, H. Namkoong, et al. Robust fine-tuning of zero-shot models. In Proc. of CVPR, pages 7959–7971, 2022

  161. [169]

    J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Tor- ralba. Sun database: Large-scale scene recognition from abbey to zoo. In Proc. of CVPR, pages 3485–3492. IEEE, 2010

  162. [170]

    Z. Xiao, J. Shen, M. M. Derakhshani, S. Liao, and C. G. M. Snoek. Any-shift prompting for generaliza- tion over distributions. In Proc. of CVPR, pages 13849– 13860, 2024

  163. [171]

    P . Xu, Z. Deng, J. Wang, Q. Zhang, K.-S. Choi, and S. Wang. Transfer representation learning with tsk fuzzy system. IEEE Trans. on Fuzzy Systems, 29(3):649– 663, 2019

  164. [172]

    Q. Xu, R. Zhang, Y. Zhang, Y. Wang, and Q. Tian. A fourier-based framework for domain generalization. In Proc. of CVPR, pages 14383–14392, 2021

  165. [173]

    Q. Xuan, T. Yu, L. Bai, and Y. Ruan. Consistent augmentation learning for generalizing clip to unseen domains. IEEE Access, 2024

  166. [174]

    S. Yan, C. Luo, Z. Yu, and Z. Ge. Generalizing clip to unseen domain via text-guided diverse novel feature synthesis. arXiv:2405.02586, 2024

  167. [175]

    L. Yang, R. Y. Zhang, Y. Wang, and X. Xie. Mma: Multi- modal adapter for vision-language models. In Proc. of CVPR, pages 23826–23837, 2024

  168. [176]

    Y. Yang, Y. Hou, L. Wen, P . Zeng, and Y. Wang. Semantic-aware adaptive prompt learning for uni- versal multi-source domain adaptation. IEEE Signal Processing Letters, 2024

  169. [177]

    H. Yao, R. Zhang, and C. Xu. Visual-language prompt tuning with knowledge-guided context optimization. In Proc. of CVPR, pages 6757–6767, 2023

  170. [178]

    Y. Yao, A. Zhang, Z. Zhang, Z. Liu, T.-S. Chua, and M. Sun. Cpt: Colorful prompt tuning for pre-trained vision-language models. arXiv e-prints, pages arXiv– 2109, 2021

  171. [179]

    M. Yi, L. Hou, J. Sun, L. Shang, X. Jiang, Q. Liu, and Z. Ma. Improved ood generalization via adversarial training and pretraining. In Proc. of ICML , pages 11987–11997. PMLR, 2021

  172. [180]

    Y. Yin, Z. Yang, H. Hu, and X. Wu. Universal multi- source domain adaptation for image classification.PR, 121:108238, 2022

  173. [181]

    H. Yu, C. Jin, Y. Zhang, X. Cao, and Z. Fang. Domain prompt matters a lot in multi-source few-shot domain adaptation. openreview.net, 2024

  174. [182]

    Q. Yu, G. Irie, and K. Aizawa. Open-set domain adaptation with visual-language foundation models. arXiv:2307.16204, 2023

  175. [183]

    L. Yuan, D. Chen, Y.-L. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li, et al. Flo- rence: A new foundation model for computer vision. arXiv:2111.11432, 2021

  176. [184]

    S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proc. of ICCV , pages 6023–6032, 2019

  177. [185]

    Y. Zang, W. Li, K. Zhou, C. Huang, and C. C. Loy. Unified vision and language prompt learning. arXiv:2210.07225, 2022

  178. [186]

    S. Zeng, X. Liu, and Y. Zhou. Decoupling domain invariance and variance with tailored prompts for open-set domain adaptation. In Proc. of ICIP , pages 645–651, 2024

  179. [187]

    Zhang, Y

    B. Zhang, Y. Wang, W. Hou, H. Wu, J. Wang, M. Oku- mura, and T. Shinozaki. Flexmatch: Boosting semi- supervised learning with curriculum pseudo labeling. Proc. of NeurIPS, 34:18408–18419, 2021

  180. [188]

    H. Zhang. Mixup: Beyond empirical risk minimiza- tion. arXiv:1710.09412, 2017

  181. [189]

    Zhang, S

    H. Zhang, S. Bai, W. Zhou, J. Fu, and B. Chen. Promptta: Prompt-driven text adapter for source-free domain generalization. arXiv:2409.14163, 2024

  182. [190]

    Zhang, Q

    J. Zhang, Q. Wei, F. Liu, and L. Feng. Candidate pseu- dolabel learning: Enhancing vision-language models by prompt tuning with unlabeled data. In Proc. of ICML, page 235, 2024

  183. [191]

    Zhang, Z

    R. Zhang, Z. Guo, W. Zhang, K. Li, X. Miao, B. Cui, 20 Y. Qiao, P . Gao, and H. Li. Pointclip: Point cloud understanding by clip. In Proc. of CVPR, pages 8552– 8562, 2022

  184. [192]

    Zhang, W

    W. Zhang, W. Ouyang, W. Li, and D. Xu. Collaborative and adversarial network for unsupervised domain adaptation. In Proc. of CVPR, pages 3801–3809, 2018

  185. [193]

    Zhang, L

    W. Zhang, L. Shen, and C.-S. Foo. Source-free domain adaptation guided by vision and vision-language pre- training. IJCV, pages 1–23, 2024

  186. [194]

    Zhang, S

    X. Zhang, S. S. Gu, Y. Matsuo, and Y. Iwasawa. Do- main prompt learning for efficiently adapting clip to unseen domains. Transactions of the Japanese Society for Artificial Intelligence, 38(6):B–MC2 1, 2023

  187. [195]

    Zhang, Y

    X. Zhang, Y. He, R. Xu, H. Yu, Z. Shen, and P . Cui. Nico++: Towards better benchmarking for domain generalization. In Proc. of CVPR , pages 16036–16047, 2023

  188. [196]

    Zhang, R

    X. Zhang, R. Xu, H. Yu, Y. Dong, P . Tian, and P . Cui. Flatness-aware minimization for domain generaliza- tion. In Proc. of ICCV, pages 5189–5202, 2023

  189. [197]

    Zhang, H

    Y. Zhang, H. Jiang, Y. Miura, C. D. Manning, and C. P . Langlotz. Contrastive learning of medical visual representations from paired images and text. In Proc. of MLHC, pages 2–25. PMLR, 2022

  190. [198]

    H. Zhao, H. Chen, F. Yang, N. Liu, H. Deng, H. Cai, S. Wang, D. Yin, and M. Du. Explainability for large language models: A survey. ACM TIST , 15(2):1–38, 2024

  191. [199]

    Zhong, D

    Z. Zhong, D. Friedman, and D. Chen. Factual probing is [mask]: Learning vs. learning to recall. arXiv:2104.05240, 2021

  192. [200]

    C. Zhou, C. C. Loy, and B. Dai. Extract free dense labels from clip. In Proc. of ECCV , pages 696–712. Springer, 2022

  193. [201]

    K. Zhou, Y. Yang, T. Hospedales, and T. Xiang. Learn- ing to generate novel domains for domain generaliza- tion. In Proc. of ECCV, pages 561–578. Springer, 2020

  194. [202]

    K. Zhou, Y. Yang, Y. Qiao, and T. Xiang. Domain generalization with mixstyle. arXiv:2104.02008, 2021

  195. [203]

    K. Zhou, J. Yang, C. C. Loy, and Z. Liu. Conditional prompt learning for vision-language models. In Proc. of CVPR, pages 16816–16825, 2022

  196. [204]

    K. Zhou, J. Yang, C. C. Loy, and Z. Liu. Learning to prompt for vision-language models. IJCV, 130(9): 2337–2348, 2022

  197. [205]

    Zhou and Z

    W. Zhou and Z. Zhou. Unsupervised domain adap- tation harnessing vision-language pre-training. IEEE TCSVT, 2024

  198. [206]

    B. Zhu, Y. Niu, Y. Han, Y. Wu, and H. Zhang. Prompt- aligned gradient for prompt tuning. In Proc. of ICCV, pages 15659–15669, 2023

  199. [207]

    J. Zhu, Y. Chen, and L. Wang. Clip the divergence: Language-guided unsupervised domain adaptation. arXiv:2407.01842, 2024

  200. [208]

    Y. Zou, Z. Yu, B. V . K. Kumar, and J. Wang. Unsuper- vised domain adaptation for semantic segmentation via class-balanced self-training. In Proc. of ECCV , pages 289–305, 2018

  201. [209]

    H. Zuo, J. Lu, G. Zhang, and W. Pedrycz. Fuzzy rule-based domain adaptation in homogeneous and heterogeneous spaces. IEEE Trans. on Fuzzy Systems , 27(2):348–361, 2018

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.