REVIEW 3 major objections 8 minor 71 references
VLMs meet UDA: Boosting Transferability of Open Vocabulary Segmentation with Unsupervised Domain Adaptation
T0 review · 3 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Unsupervised domain adaptation can boost open-vocabulary segmentation across domains that share no categories.
desk verdict First true UDA-OVSS integration with real target-private segmentation, but the 'no shared categories' claim is overstated: the benchmark still shares 16/19 classes and the SOTA margin is entirely target-private. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The cost volume is the central object: for each patch and category prompt, the cosine similarity between dense CLIP visual features and text features, which the decoder refines into pixel-wise maps. Three supporting mechanisms carry the argument: robust text embeddings built by averaging many LLM-generated prompt variations per category; a fine-tuning scheme that freezes MLP layers, tunes spatial layers, and decays learning rates by a factor of $\beta$ from the last encoder layer backward; and a UDA loop in which a teacher with a frozen encoder and EMA-updated decoder produces target pseudo-labels, a student trains on source labels and cross-domain mixed samples, and a gamma schedule shifts pseudo-label trust from teacher to student over time.
What would settle it
Use the paper's own full-EMA variant as the control: updating the teacher encoder makes the model forget the target-private train class, dropping it to 0.1 mIoU in the paper's comparison while UDA-FROVSS reaches 60.2; any reproduction that preserves the frozen-encoder/EMA-decoder design yet still leaves train and truck near zero would falsify the claim that this design is what enables open-vocabulary transfer.
Extended reading notes
Core claim
The paper's central discovery, in the authors' own framing, is that UDA and VLM-based open-vocabulary segmentation are mutually reinforcing and can be combined into the first UDA framework that works without shared categories between source and target. The evidence is four-fold: the FROVSS decoder and prompt augmentation improve open-vocabulary segmentation on every benchmark tested; the layer-wise fine-tuning preserves CLIP's generalization; the teacher-student design with a frozen teacher encoder and EMA-updated teacher decoder lets the model learn target-private labels; and the full UDA-FROVSS pipeline sets a new state of the art on the Synthia-to-Cityscapes UDA benchmark while remaining open vocabulary. In the authors' words, the framework removes the need for shared categories; per-class results show that classes absent from the source (truck, train) are segmented at 80.3 and 60.2 mIoU respectively, whereas closed-set UDA baselines score zero on them.
Load-bearing premise
The load-bearing premise is that the teacher's pseudo-labels on unlabeled target images are reliable enough to supervise the student, especially for target-private categories that have no labeled examples anywhere in training.
Editorial extensions
If this is right
- Open-vocabulary segmentation models can use large unlabeled target image collections to specialize to a domain without giving up their ability to name novel categories.
- UDA systems no longer need to regenerate synthetic data or retrain when a new category appears; the open-vocabulary head can recognize it from the prompt alone.
- The Synthia-to-Cityscapes setting becomes a usable UDA benchmark for open-vocabulary models, with target-private classes scored explicitly rather than ignored.
- The FROVSS components (prompt augmentation and layer-wise fine-tuning) account for most of the observed UDA gains in the paper's ablation, so improving open-vocabulary capability is itself a route to better domain transfer.
Reading between the lines
- A general recipe may follow: keep the pretrained vision-language geometry frozen and adapt a small spatially-aware head; the same split could be tested on open-vocabulary detection, panoptic segmentation, or monocular depth estimation.
- The paper's Synthia-to-Cityscapes experiment still shares 16 of 19 classes between source and target, so 'no shared categories' is demonstrated only partially; a stronger test would use a source-target pair with zero overlap and check whether private classes emerge purely from teacher pseudo-labels.
- The gamma decay schedule is an acknowledged patch for teacher encoder-decoder misalignment; re-aligning the teacher (for example by periodic reset or low-rank adapters) could remove the need for the schedule and is a natural next step.
- Photometry-based prompt augmentations hurt performance in the paper's ablations, suggesting that the prompt-augmentation recipe is not uniformly beneficial and may need to be tuned per domain.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FROVSS, an open-vocabulary semantic segmentation (OVSS) model that combines a CLIP image encoder with a hybrid convolutional-transformer decoder, LLM-based prompt augmentation, and layer-wise learning-rate decay. It then introduces UDA-FROVSS, a UDA extension with a teacher-student setup in which the teacher's encoder is frozen while the decoder is updated by EMA, cross-domain image mixup, and confidence-weighted pseudo-label blending controlled by a gamma decay schedule. The authors report improvements on several OVSS benchmarks and a new state of the art of 61.5 mIoU on Synthia-to-Cityscapes, claiming this is the first UDA framework that adapts across domains without requiring shared categories. The central contribution is the integration of VLM-based open-vocabulary reasoning with standard UDA techniques to recognize target-private classes during adaptation.
Significance. If the central claim is accepted, the paper opens a valuable direction: adapting open-vocabulary models to unlabeled target domains while retaining the ability to segment categories absent from the source label set. The paper is honest about the teacher encoder-decoder misalignment (Sec. 4.4) and provides detailed ablations (Tables 10, 11). The cross-dataset OVSS experiments are extensive, and the prompt-augmentation analysis is a useful empirical contribution. However, the manuscript currently does not provide code, error bars, or multiple-seed runs, and the headline claim about disjoint label sets rests on only three target-private classes in a benchmark that shares 16 of 19 classes. The shared-class average on Synthia-to-Cityscapes is below prior closed-set UDA methods, so the SOTA margin is entirely attributable to target-private categories whose pseudo-label quality is not measured. These are fixable but load-bearing issues.
major comments (3)
- [Sec. 4.4 and Table 11] The central claim that UDA-FROVSS adapts "without requiring shared categories" is not directly demonstrated: the only UDA benchmark with target-private classes, Synthia-to-Cityscapes, shares 16 of 19 classes, and the three private classes (terrain, truck, train) are never separated out in the evaluation. Recomputing the shared-class average from Table 11, UDA-FROVSS scores 64.3 mIoU on the 16 shared classes, while DCF in Table 13 scores 69.3; thus the reported advantage over prior SOTA comes entirely from truck/train (terrain is 0.0 for both). The paper must report the shared-class breakdown and, ideally, run a fully disjoint label split before claiming the advertised generalization.
- [Sec. 3.3, Eqs. (10)-(12)] The pseudo-label schedule that enables learning of target-private categories is introduced as a patch for the acknowledged teacher encoder-decoder misalignment (Sec. 4.4), but no sensitivity analysis, no ablation on gamma_0, and no pseudo-label quality measure (e.g., precision/recall on a held-out target set) are provided. Since the method's headline capability depends on this schedule, its behavior under different gamma_0 values and different domain gaps should be characterized. As written, Eq. (12) is ambiguous and possibly missing parentheses: "gamma_{delta+1} <- (1/delta gamma_delta + 1)" should be clarified.
- [Sec. 4.1, Tables 10 and 13] The experimental evidence for the claimed SOTA margin lacks error bars, multiple seeds, and code release. Hyperparameters beta and mu are set via "initial exploration" (Sec. 4.1), and the ablation in Table 10 reports single-run numbers. The reported Synthia-to-Cityscapes advantage over DCF is 61.5 vs 58.4 mIoU, which is an absolute margin of 3.1; the paper's phrase "by over 8%" should specify whether this refers to relative improvement, and the reader needs to know the variance of these numbers before the superiority claim can be assessed.
minor comments (8)
- [Sec. 1] There are several typos in the introduction, including "class dviersity" and "Alltogether"; the text should be proofread.
- [Sec. 3.3, Eq. (12)] The gamma update rule as printed is notationally ambiguous; please add parentheses and define the intended recurrence clearly.
- [Sec. 4.4] The sentence "hence defining the combination of teacher and student labels described in equation 6 of the paper" should refer to the correct equation number, which appears to be Eq. (11) or (12), not Eq. (6).
- [Table 6] The row labels "Spatial[13]" and "Proposed" are unclear; please specify the exact fine-tuning protocol for each row (which layers are updated and how).
- [Table 7] The "OV" column uses checkmarks and crosses without explaining the criterion; clarify whether it denotes the ability to handle unseen categories at inference.
- [Sec. 2] "Covariant distribution shift" should be "covariate shift" in the related-work discussion.
- [Sec. 4.2 and Table 5] The photometry prompt augmentation is reported to not improve performance, yet the contributions section presents prompt augmentation as a generally beneficial strategy; the paper should temper the claim or explain why the photometry variant failed.
- [Sec. 4.2 and Table 4] The statement that prompt augmentation "exclusively during testing enhances the model's generality" is not uniformly supported: for the ADE-20-trained model, same-dataset performance drops from 53.4 to 53.0 and cross-dataset performance is mixed; the discussion should be more balanced.
Circularity Check
No significant circularity: target-private recognition rests on frozen CLIP embeddings and external benchmarks, not on fitted inputs or load-bearing self-citation.
full rationale
The paper's derivation chain is self-contained and externally validated. FROVSS is benchmarked against CAT-Seg and other CLIP-based segmentators on COCO, ADE-20, Pascal-Context, and Pascal VOC using held-out validation sets, and UDA-FROVSS is evaluated on Synthia-to-Cityscapes against prior UDA methods (MM, DIGA, MIC, DCF). The target-private classes (truck, train) are obtained through the teacher's frozen CLIP image encoder and the always-frozen text encoder, so this capability is not a fitted parameter renamed as a prediction. The teacher decoder EMA (Eq. 10) and gamma-decay pseudo-label combination (Eqs. 11-12) are training heuristics, not re-statements of the evaluation metric. The self-citations ([2], [7], [33]) appear in background discussion of UDA and out-of-distribution detection and are not load-bearing for the central claim. The reviewer concern that the 'without shared categories' claim is only demonstrated on a 16/19 shared-class benchmark, and that the reported SOTA margin is concentrated in target-private classes, is a scope and benchmark-design objection rather than a circularity: no claimed result is shown to reduce by construction to its own inputs.
Assumptions & free parameters
free parameters (5)
- beta (layer-wise learning rate decay) =
0.95
- mu (pixel confidence threshold) =
0.96
- gamma_0 (initial teacher weight) =
not reported
- alpha (EMA decay) =
not reported
- delta (EMA time step) =
not reported
assumptions (6)
- domain assumption CLIP image and text encoders pretrained on web-scale data contain sufficient semantic knowledge for open-vocabulary pixel-level segmentation.
- domain assumption Tuning only spatial-interaction layers (attention, positional embeddings) while freezing MLP layers suffices to transfer CLIP from image-level to pixel-level prediction.
- domain assumption Layer-wise decayed learning rate (Eq. 1) preserves pre-trained knowledge and improves fine-tuning.
- domain assumption A teacher with frozen encoder and EMA-updated decoder generates reliable pseudo-labels for the target domain, even as the decoder drifts from its encoder.
- domain assumption Cross-domain mixed sampling (DACS-style overlaying source instances on target images) enforces domain-invariant features.
- domain assumption LLM-generated prompt augmentations and average text embeddings improve robustness of category recognition.
Cite this review
Pith. "Pith review of VLMs meet UDA: Boosting Transferability of Open Vocabulary Segmentation with Unsupervised Domain Adaptation." pith.science (2026). https://pith.science/paper/QBRTBJSH
@misc{pith2026241209240,
author = {Pith},
title = {Pith review of: VLMs meet UDA: Boosting Transferability of Open Vocabulary Segmentation with Unsupervised Domain Adaptation},
year = {2026},
howpublished = {\url{https://pith.science/paper/QBRTBJSH}},
note = {Machine review of arXiv:2412.09240}
}
read the original abstract
Segmentation models are typically constrained by the categories defined during training. To address this, researchers have explored two independent approaches: adapting Vision-Language Models (VLMs) and leveraging synthetic data. However, VLMs often struggle with granularity, failing to disentangle fine-grained concepts, while synthetic data-based methods remain limited by the scope of available datasets. This paper proposes enhancing segmentation accuracy across diverse domains by integrating Vision-Language reasoning with key strategies for Unsupervised Domain Adaptation (UDA). First, we improve the fine-grained segmentation capabilities of VLMs through multi-scale contextual data, robust text embeddings with prompt augmentation, and layer-wise fine-tuning in our proposed Foundational-Retaining Open Vocabulary Semantic Segmentation (FROVSS) framework. Next, we incorporate these enhancements into a UDA framework by employing distillation to stabilize training and cross-domain mixed sampling to boost adaptability without compromising generalization. The resulting UDA-FROVSS framework is the first UDA approach to effectively adapt across domains without requiring shared categories.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Class- conditional domain adaptation for semantic seg- mentation
Wang Y, Li Y, Elder JH, Wu R, Lu H. Class- conditional domain adaptation for semantic seg- mentation. Computational Visual Media , 2024: 1–18
work page 2024
-
[2]
Alcover-Couso R, SanMiguel JC, Escudero- Vinolo M, Garcia-Martin A. On exploring weakly supervised domain adaptation strate- gies for semantic segmentation using synthetic data. Multimedia Tools and Applications , 2023: 35879–35911. 17
work page 2023
-
[3]
Taming diffusion model for exemplar-based image translation
Ma H, Yang J, Huang H. Taming diffusion model for exemplar-based image translation. Computa- tional Visual Media , 2024: 1–13
work page 2024
-
[4]
Learning layout generation for virtual worlds
Cheng W, Shan Y. Learning layout generation for virtual worlds. Computational Visual Media , 2024: 1–16
work page 2024
-
[5]
Adap- tive sampling and reconstruction for gradient- domain rendering
Liang Y, Liu T, Huo Y, Wang R, Bao H. Adap- tive sampling and reconstruction for gradient- domain rendering. Computational Visual Media, 2024: 1–18
work page 2024
-
[6]
Multi3D: 3D-aware multimodal image synthesis
Zhou W, Yuan L, Mu T. Multi3D: 3D-aware multimodal image synthesis. Computational Vi- sual Media, 2024: 1–13
work page 2024
-
[7]
Alcover-Couso R, SanMiguel JC, Escudero- Vi˜ nolo M. Biased Class disagreement: detection of out of distribution instances by using differ- ently biased semantic segmentation models. In Int. Conf. Comput. Vis. (ICCVW) , 2023, 4580– 4588
work page 2023
-
[8]
Cross-modal learning using privileged informa- tion for long-tailed image classification
Li X, Zheng Y, Ma H, Qi Z, Meng X, Meng L. Cross-modal learning using privileged informa- tion for long-tailed image classification. Compu- tational Visual Media , 2024: 1–12
work page 2024
Show all 71 references
-
[9]
Don’t Stop Learning: Towards Continual Learning for the CLIP Model
Ding Y, Liu L, Tian C, Yang J, Ding H. Don’t Stop Learning: Towards Continual Learning for the CLIP Model. ArXiv, 2022, abs/2207.09248
2022 arXiv
-
[10]
Generative Negative Text Replay for Continual Vision-Language Pretraining
Yan S, Hong L, Xu H, Han J, Tuytelaars T, Li Z, He X. Generative Negative Text Replay for Continual Vision-Language Pretraining. ArXiv, 2022, abs/2210.17322
2022 arXiv
-
[11]
Extract Free Dense Labels from CLIP
Zhou C, Loy CC, Dai B. Extract Free Dense Labels from CLIP. In Eur. Conf. Comput. Vis. (ECCV), 2022
2022
- [12]
-
[13]
CAT-Seg: Cost Aggre- gation for Open-Vocabulary Semantic Segmen- tation
Cho S, Shin H, Hong S, An S, Lee S, Arnab A, Seo PH, Kim S. CAT-Seg: Cost Aggre- gation for Open-Vocabulary Semantic Segmen- tation. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), 2024
2024
-
[14]
COCO-Stuff: Thing and Stuff Classes in Context
Holger Caesar VF Jasper Uijlings. COCO-Stuff: Thing and Stuff Classes in Context. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR) , 2018
2018
-
[15]
Early Convolutions Help Transform- ers See Better
Xiao T, Singh M, Mintun E, Darrell T, Dollar P, Girshick R. Early Convolutions Help Transform- ers See Better. In Adv. Neural Inform. Process. Syst. (NeurIPS), volume 34, 2021, 30392–30400
2021
-
[16]
Incorporating Convolution Designs Into Visual Transformers
Yuan K, Guo S, Liu Z, Zhou A, Yu F, Wu W. Incorporating Convolution Designs Into Visual Transformers. In IEEE Int. Conf. Comput. Vis. (ICCV), 2021, 579–588
2021
-
[17]
Pyramid Geometric Consistency Learning For Seman- tic Segmentation
Zhang X, Li Q, Quan Z, Yang W. Pyramid Geometric Consistency Learning For Seman- tic Segmentation. Pattern Recognition , 2023, 133: 109020, doi:https://doi.org/10.1016/j. patcog.2022.109020
2023
-
[18]
Learning Trans- ferable Visual Models From Natural Language Supervision
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I. Learning Trans- ferable Visual Models From Natural Language Supervision. In Int. Conf. Mach. Lear. (ICML) , volume 139, 2021, 8748–8763
2021
-
[19]
SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation
Luo H, Bao J, Wu Y, He X, Li T. SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation. Int. Conf. Mach. Lear. (ICML) , 2023
2023
-
[20]
Open-Vocabulary Panop- tic Segmentation with MaskCLIP
Ding Z, Wang J, Tu Z. Open-Vocabulary Panop- tic Segmentation with MaskCLIP. Int. Conf. Comput. Vis. (ICCV) , 2023
2023
-
[21]
Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling
Huynh DT, Kuen J, nan Lin Z, Gu J, Elhami- far E. Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling. IEEE Conf. Comput. Vis. Pattern Recog. (CVPR) , 2021: 7010–7021. 18
2021
-
[22]
Collaborating Foundation models for Domain Generalized Semantic Segmentation
Benigmim Y, Roy S, Essid S, Kalogeiton V, Lathuili` ere S. Collaborating Foundation models for Domain Generalized Semantic Segmentation. arXiv:2312.09788, 2023
2023 arXiv
-
[23]
Side Adapter Network for Open-Vocabulary Seman- tic Segmentation
Xu M, Zhang Z, Wei F, Hu H, Bai X. Side Adapter Network for Open-Vocabulary Seman- tic Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), 2023
2023
-
[24]
CLIP-SP: Vision-language model with adap- tive prompting for scene parsing
Li J, Huang Y, Wu M, Zhang B, Ji X, Zhang C. CLIP-SP: Vision-language model with adap- tive prompting for scene parsing. Computational Visual Media, 2024: 741–752
2024
-
[25]
Ex- ploring Visual Interpretability for Con- trastive Language-Image Pre-training
Li Y, Wang H, Duan Y, Xu H, Li X. Ex- ploring Visual Interpretability for Con- trastive Language-Image Pre-training. arXiv:2209.07046, 2022
2022 arXiv
-
[26]
CLIP Surgery for Better Explainability with Enhancement in Open-Vocabulary Tasks
Li Y, Wang H, Duan Y, Li X. CLIP Surgery for Better Explainability with Enhancement in Open-Vocabulary Tasks. arXiv:2304.05653, 2023
2023 arXiv
-
[27]
SemiVL: Semi-Supervised Semantic Segmen- tation with Vision-Language Guidance
Hoyer L, Tan DJ, Naeem MF, Gool L V, Tombari F. SemiVL: Semi-Supervised Semantic Segmen- tation with Vision-Language Guidance. In Eur. Conf. Comput. Vis. (ECCV) , 2023
2023
-
[28]
Open- vocabulary semantic segmentation with mask- adapted clip
Liang F, Wu B, Dai X, Li K, Zhao Y, Zhang H, Zhang P, Vajda P, Marculescu D. Open- vocabulary semantic segmentation with mask- adapted clip. In IEEE Conf. Comput. Vis. Pat- tern Recog. (CVPR), 2023, 7061–7070
2023
-
[29]
Decoupling Zero- Shot Semantic Segmentation
Ding J, Xue N, Xia GS, Dai D. Decoupling Zero- Shot Semantic Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR) , 2022, 11573–11582
2022
-
[30]
A Simple Baseline for Open Vocab- ulary Semantic Segmentation with Pre-trained Vision-language Model
Xu M, Zhang Z, Wei F, Lin Y, Cao Y, Hu H, Bai X. A Simple Baseline for Open Vocab- ulary Semantic Segmentation with Pre-trained Vision-language Model. Eur. Conf. Comput. Vis. (ECCV) , 2022
2022
-
[31]
Discovering latent target subdomains for domain adaptive seman- tic segmentation via style clustering
Wang S, Zhao X, Chen J. Discovering latent target subdomains for domain adaptive seman- tic segmentation via style clustering. Multimedia Tools and Applications , 2023: 3234–3243, doi: 10.1007/s11042-023-15620-6
2023 doi
-
[32]
Survey on Unsupervised Domain Adaptation for Semantic Segmentation for Vi- sual Perception in Automated Driving
Schwonberg M, Niemeijer J, Term¨ ohlen JA, sch¨ afer JP, Schmidt NM, Gottschalk H, Fin- gscheidt T. Survey on Unsupervised Domain Adaptation for Semantic Segmentation for Vi- sual Perception in Automated Driving. IEEE Access, 2023, 11: 54296–54336
2023
-
[33]
Per-Class Curriculum for Unsupervised Domain Adaptation in Se- mantic Segmentation
Alcover-Couso R, SanMiguel JC, Escudero- Vi˜ nolo M, Caballeira P. Per-Class Curriculum for Unsupervised Domain Adaptation in Se- mantic Segmentation. In The Visual Computer , 2023, 1–19
2023
-
[34]
Pseudo-Label : The Simple and Ef- ficient Semi-Supervised Learning Method for Deep Neural Networks
Lee DH. Pseudo-Label : The Simple and Ef- ficient Semi-Supervised Learning Method for Deep Neural Networks. In Int. Conf. Mach. Lear. (ICML W), 2013
2013
-
[35]
DAFormer: Im- proving Network Architectures and Training Strategies for Domain-Adaptive Semantic Seg- mentation
Hoyer L, Dai D, Van Gool L. DAFormer: Im- proving Network Architectures and Training Strategies for Domain-Adaptive Semantic Seg- mentation. In IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022, 9924–9935
2022
-
[36]
HRDA: Context- Aware High-Resolution Domain-Adaptive Se- mantic Segmentation
Hoyer L, Dai D, Van Gool L. HRDA: Context- Aware High-Resolution Domain-Adaptive Se- mantic Segmentation. In IEEE Eur. Conf. Com- put. Vis. (ECCV) , 2022, 372–391
2022
-
[37]
CDAC: Cross-domain Attention Consistency in Transformer for Domain Adaptive Seman- tic Segmentation
Wang K, Kim D, Feris R, Saenko K, Betke M. CDAC: Cross-domain Attention Consistency in Transformer for Domain Adaptive Seman- tic Segmentation. In IEEE Conf. Comput. Vis. (ICCV), 2023
2023
-
[38]
CoN- Mix for Source-free Single and Multi-target Do- main Adaptation
Kumar V, Lal R, Patil H, Chakraborty A. CoN- Mix for Source-free Single and Multi-target Do- main Adaptation. In Wint. App. Comp. Vis. (WACV), 2023, 4178–4188
2023
-
[39]
MIC: Masked Image Consistency for Context- Enhanced Domain Adaptation
Hoyer L, Dai D, Wang H, Van Gool L. MIC: Masked Image Consistency for Context- Enhanced Domain Adaptation. In IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2023. 19
2023
-
[40]
Mean teachers are bet- ter role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Tarvainen A, Valpola H. Mean teachers are bet- ter role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Adv. Neural Inform. Process. Syst. (NeurIPS), 2017
2017
-
[41]
Research On Data Model Migration In Image Semantic Segmenta- tion Based On Deep Learning
Guo W, Liu F, Song Y, Qin C. Research On Data Model Migration In Image Semantic Segmenta- tion Based On Deep Learning. Int. Conf. Mea- suring Technology and Mechatronics Automa- tion, 2022: 417–420
2022
-
[42]
Rectifying Pseudo Label Learning via Uncertainty Estimation for Domain Adaptive Semantic Segmentation
Zheng Z, Yang Y. Rectifying Pseudo Label Learning via Uncertainty Estimation for Domain Adaptive Semantic Segmentation. Int. J. Com- put. Vis. (IJCV) , 2020: 1–15
2020
-
[43]
Characterizations of semantic do- mains for randomized algorithms
Yamada S. Characterizations of semantic do- mains for randomized algorithms. Japan Journal of Applied Mathematics , 1989, 6: 111–146
1989
-
[44]
Training Deep Networks with Synthetic Data: Bridging the Reality Gap by Domain Randomization
Tremblay J, Prakash A, Acuna D, Brophy M, Jampani V, Anil C, To T, Cameracci E, Boo- choon S, Birchfield S. Training Deep Networks with Synthetic Data: Bridging the Reality Gap by Domain Randomization. IEEE Conf. Com- put. Vis. Pattern Recog. (CVPR W), 2018: 1082– 10828
2018
-
[45]
Structured Domain Randomization: Bridg- ing the Reality Gap by Context-Aware Synthetic Data
Prakash A, Boochoon S, Brophy M, Acuna D, Cameracci E, State G, Shapira O, Birchfield S. Structured Domain Randomization: Bridg- ing the Reality Gap by Context-Aware Synthetic Data. IEEE Int. Conf. Rob. Aut. (ICRA) , 2018: 7249–7255
2018
-
[46]
Domain randomization for neural network classification
Valtchev SZ, Wu J. Domain randomization for neural network classification. Journal of Big Data, 2020, 8
2020
-
[47]
DACS: Domain Adaptation via Cross-domain Mixed Sampling
Tranheden W, Olsson V, Pinto J, Svensson L. DACS: Domain Adaptation via Cross-domain Mixed Sampling. IEEE Winter Conf. App. Comp. Vis. (WACV) , 2020: 1378–1388
2020
-
[48]
CLIP-Flow: Decoding images encoded in CLIP space
Ma H, Li M, Yang J, Patashnik O, Lischinski D, Cohen-Or D, Huang H. CLIP-Flow: Decoding images encoded in CLIP space. Computational Visual Media, 2024: 1–12
2024
-
[49]
Attention is All you Need
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Lu, Polosukhin I. Attention is All you Need. In Adv. Neural Inform. Process. Syst. (NeurIPS) , volume 30, 2017
2017
-
[50]
Universal Language Model Fine-tuning for Text Classification
Howard J, Ruder S. Universal Language Model Fine-tuning for Text Classification. In ACL, 2018
2018
-
[51]
Fast End-to- End Trainable Guided Filter
Wu H, Zheng S, Zhang J, Huang K. Fast End-to- End Trainable Guided Filter. IEEE Conf. Com- put. Vis. Pattern Recog. (CVPR) , 2018: 1838– 1847
2018
-
[52]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy A, Beyer L, Kolesnikov A, Weis- senborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, Uszkoreit J, Houlsby N. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Int. Conf. Learn. Rep. (ICLR) , 2021
2021
-
[53]
Convolu- tional neural network architecture for geometric matching
Rocco I, Arandjelovi´ c R, Sivic J. Convolu- tional neural network architecture for geometric matching. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), 2017
2017
-
[54]
Cost Ag- gregation with 4D Convolutional Swin Trans- former for Few-Shot Segmentation
Hong S, Cho S, Nam J, Lin S, Kim S. Cost Ag- gregation with 4D Convolutional Swin Trans- former for Few-Shot Segmentation. In ECCV, 2022, 108–126
2022
-
[55]
The Cityscapes Dataset for Seman- tic Urban Scene Understanding
Cordts M, Omran M, Ramos S, Rehfeld T, Enzweiler M, Benenson R, Franke U, Roth S, Schiele B. The Cityscapes Dataset for Seman- tic Urban Scene Understanding. In IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2016, 3212–3223
2016
-
[56]
Language Models are Few-Shot Learners
Brown TB, Mann B, Ryder N, Subbiah M, Ka- plan J, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, et al.. Language Models are Few-Shot Learners. Adv. Neural Inform. Pro- cess. Syst. (NeurIPS) , 2020: 1877–1901
2020
-
[57]
The Mapillary Vistas Dataset for Semantic Understanding of Street Scenes
Neuhold G, Ollmann T, Rota Bul` o S, Kontschieder P. The Mapillary Vistas Dataset for Semantic Understanding of Street Scenes. In 20 IEEE Int. Conf. Comput. Vis. (ICCV) , 2017, 5000–5009
2017
-
[58]
Domain randomization for trans- ferring deep neural networks from simulation to the real world
Tobin J, Fong R, Ray A, Schneider J, Zaremba W, Abbeel P. Domain randomization for trans- ferring deep neural networks from simulation to the real world. IEEE Conf. Intell. Rob. Sys. (IROS), 2017: 23–30
2017
-
[59]
Semantic understanding of scenes through the ade20k dataset
Zhou B, Zhao H, Puig X, Xiao T, Fidler S, Bar- riuso A, Torralba A. Semantic understanding of scenes through the ade20k dataset. Int. Journal of Computer Vision , 2019, 127: 302–321
2019
-
[60]
The Role of Con- text for Object Detection and Semantic Segmen- tation in the Wild
Mottaghi R, Chen X, Liu X, Cho NG, Lee SW, Fidler S, Urtasun R, Yuille A. The Role of Con- text for Object Detection and Semantic Segmen- tation in the Wild. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), 2014
2014
-
[61]
The Pascal Visual Object Classes Challenge: A Retrospec- tive
Everingham M, Eslami SMA, Van Gool L, Williams CKI, Winn J, Zisserman A. The Pascal Visual Object Classes Challenge: A Retrospec- tive. IJCV, 2015, 111(1): 98–136
2015
-
[62]
The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Seg- mentation of Urban Scenes
Ros G, Sellart L, Materzynska J, Vazquez D, Lopez AM. The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Seg- mentation of Urban Scenes. IEEE Conf. Com- put. Vis. Pattern Recog. (CVPR) , 2016: 3234– 3243
2016
-
[63]
Swin Transformer V2: Scaling Up Capacity and Res- olution
Liu Z, Hu H, Lin Y, Yao Z, Xie Z, Wei Y, Ning J, Cao Y, Zhang Z, Dong L, Wei F, Guo B. Swin Transformer V2: Scaling Up Capacity and Res- olution. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), 2022
2022
-
[64]
Segment Anything
Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L, Xiao T, Whitehead S, Berg AC, Lo WY, Dollar P, Girshick R. Segment Anything. In Int. Conf. Comput. Vis. (ICCV) , 2023, 4015– 4026
2023
-
[65]
FreeSeg: Unified, Universal and Open-Vocabulary Image Segmentation
Qin J, Wu J, Yan P, Li M, Yuxi R, Xiao X, Wang Y, Wang R, Wen S, Pan X, et al.. FreeSeg: Unified, Universal and Open-Vocabulary Image Segmentation. IEEE Conf. Comput. Vis. Pat- tern Recog. (CVPR), 2023
2023
-
[66]
MasQCLIP for Open-Vocabulary Universal Image Segmenta- tion
Xu X, Xiong T, Ding Z, Tu Z. MasQCLIP for Open-Vocabulary Universal Image Segmenta- tion. In Int. Conf. Comput. Vis. (ICCV) , 2023, 887–898
2023
-
[67]
ZegCLIP: Towards Adapting CLIP for Zero-Shot Seman- tic Segmentation
Zhou Z, Lei Y, Zhang B, Liu L, Liu Y. ZegCLIP: Towards Adapting CLIP for Zero-Shot Seman- tic Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), 2023, 11175–11185
2023
-
[68]
Hierarchical Open-vocabulary Uni- versal Image Segmentation
Wang X, Li S, Kallidromitis K, Kato Y, Kozuka K, Darrell T. Hierarchical Open-vocabulary Uni- versal Image Segmentation. In Adv. Neural In- form. Process. Syst. (NeurIPS) , 2023
2023
-
[69]
Transferring Multi-Modal Domain Knowledge to Uni-Modal Domain for Urban Scene Segmenta- tion
Liu P, Ge Y, Duan L, Li W, Luo H, Lv F. Transferring Multi-Modal Domain Knowledge to Uni-Modal Domain for Urban Scene Segmenta- tion. IEEE Transactions on Intelligent Trans- portation Systems, 2024: 11576–11589
2024
-
[70]
DiGA: Distil To Generalize and Then Adapt for Domain Adaptive Semantic Segmentation
Shen F, Gurram A, Liu Z, Wang H, Knoll A. DiGA: Distil To Generalize and Then Adapt for Domain Adaptive Semantic Segmentation. In IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, 15866–15877
2023
-
[71]
Transferring to Real- World Layouts: A Depth-aware Framework for Scene Adaptation
Chen M, Zheng Z, Yang Y. Transferring to Real- World Layouts: A Depth-aware Framework for Scene Adaptation. In ACM Multimedia, 2024. 21
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.