REVIEW 3 major objections 5 minor 77 references
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Two-stage open-vocabulary segmentation improves when mask proposals read the text prompt.
desk verdict Consistent 1–3 mIOU gains for a clean plug-in, but the arbitrary single-prompt claim is untested and likely distribution-shifted (M=1 vs. M≈171 training). read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a text-query cross-attention block inserted before the standard cross-attention in the transformer decoder. Given query features $Q_l$, the text tokens are projected to keys and values $K_t, V_t$; the queries first attend to the text to produce $Q'_l = \mathrm{softmax}(Q_l K_t^\top)V_t$, and this prompt-conditioned query then attends to the image features. This is what makes the $N$ mask embeddings prompt-specific without changing the number of queries, and it can be dropped into Mask2Former-style decoders used by several existing two-stage models.
What would settle it
Run the trained PMP pipeline on a set of abstract and proprietary prompts (for example "love," "parking," "Washington") with human-annotated ground-truth masks; if first-stage recall or final mIOU is no better than the class-agnostic baseline on that set, the central claim of prompt-guided transfer fails.
Extended reading notes
Core claim
The paper's central claim is that the missing-mask failure of two-stage open-vocabulary segmentation is fixable at the proposal stage: if the transformer decoder conditions its learned queries on the input text, the first-stage mask proposals become prompt-specific, and the frozen CLIP classifier can then retrieve regions that class-agnostic proposals never contained. On its own terms, this is a lightweight plug-in that yields absolute mIOU gains of roughly 1-3 points over OVSeg, FC-CLIP, SAN, and ODISE across ADE-847, PC-459, ADE-150, PC-59, and VOC, and enables qualitative segmentation of prompts such as "Yellowstone" and "love."
Load-bearing premise
The paper assumes that a decoder trained on COCO-Stuff class names and captions learns to condition masks on any CLIP text embedding, so the first-stage recall gains seen on standard benchmarks also hold for arbitrary test-time words like "love" or "Times Square"; the evidence for that transfer is qualitative only.
Editorial extensions
If this is right
- Because PMP is a decoder-level change, any two-stage model built on Mask2Former can absorb it without touching its frozen CLIP classifier or retraining the second stage.
- The reported first-stage recall gains are larger than the final mIOU gains (for example, OVSeg recall mIOU rises by 4.3-8.5 points across the five benchmarks), so improvements in proposal generation are the main driver of the final result.
- Prompt-specific proposals make the pipeline usable for open-ended prompts such as proper nouns, adjectives, and full captions, not just the class names used in training.
- The added cost is small: the paper reports roughly 0.01 seconds of extra first-stage latency per image while keeping the rest of the pipeline unchanged.
- The same plug-in also improves panoptic segmentation on ADE20K and COCO when added to ODISE and FC-CLIP, so the benefit is not limited to semantic segmentation benchmarks.
Reading between the lines
- Beyond the paper's experiments, the qualitative evidence suggests the mechanism transfers to prompts far outside COCO-Stuff vocabulary, but that transfer is not quantified; a natural next step is to build a benchmark of abstract and proprietary prompts with human-annotated masks and measure first-stage recall there.
- Because the improvement concentrates in first-stage recall, pairing PMP with a better proposal-to-prompt matching or ranking step in the second stage could convert more of that recall into final mIOU than the current geometric-mean classifier does.
- Training the same text-query cross-attention on a larger, more diverse set of captions and prompt pairs would likely strengthen the prompt-conditioning transfer, a testable extension the paper's framing suggests but does not run.
- The cross-attention design is generic enough that it could be applied to other query-based architectures beyond the four baselines tested, though the paper only demonstrates it on those models.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Prompt-Guided Mask Proposal (PMP), a lightweight modification to query-based mask-generator decoders (Mask2Former/MaskFormer style) in two-stage open-vocabulary segmentation. The method inserts a text-query cross-attention step before the standard image cross-attention in each decoder layer, so that mask proposals are conditioned on the input text prompt. The authors combine PMP with OVSeg, ODISE, SAN, and FC-CLIP, and report consistent mIOU gains of roughly 0.6–3.9 points on ADE-847, PC-459, ADE-150, PC-59, and Pascal VOC, plus panoptic segmentation gains on ADE20K and COCO. The paper also presents qualitative examples for prompts such as "Yellowstone," "love," and "Times Square," arguing that PMP enables mask retrieval for arbitrary single prompts rather than only benchmark class lists. The main claimed contribution is improved first-stage mask proposal quality conditioned on text, supported by first-stage recall ablations in Table 4.
Significance. If the reported gains hold, PMP is a useful and simple adapter for two-stage open-vocabulary segmentation models, and the breadth of the evaluation—four different methods/backbones across five benchmarks, plus panoptic results—is a genuine strength. The first-stage recall improvements in Table 4 are particularly informative because they show the effect is in proposal generation rather than only in the second-stage classifier. However, the paper's most distinctive claim, that the method supports arbitrary single text prompts in the style of the qualitative figures, is not quantitatively evaluated, and the architectural change has an untested distribution-shift issue at M=1. The empirical contribution is solid but currently supported without variance estimates or a controlled single-prompt benchmark.
major comments (3)
- [Sec. 4.3, Figs. 3 and 7; Eq. (5)] The central claim that PMP enables mask retrieval for arbitrary single prompts such as "love," "Times Square," or "MIT CSAIL" is supported only by qualitative examples. All quantitative tables evaluate full benchmark class lists (M=150–847 text tokens). For a single prompt, M=1 in Eq. (5): the softmax over text tokens is degenerate, every query receives the same projected text vector, and the subsequent image cross-attention produces the same attention weights across queries except through the residual X_{l-1}. This is a structurally different regime from the M≈171 training distribution, and the manuscript neither trains nor quantitatively evaluates in this regime. Appendix A.2 reports failure cases only anecdotally, and the limitation statement in Appendix A.5 addresses mask precision rather than vocabulary transfer. Please add a quantitative single-prompt/arbitrary-prompt evaluation (for example, mask IoU or recall against manually annotated regions for a set of proper nouns and abstract words), and consider whether training should sample variable M to cover the M=1 case.
- [Table 1] The central quantitative claim of consistent gains is presented without error bars, multiple seeds, or significance tests. Some reported gains are small (for example, 0.6 mIOU for FC-CLIP + PMP on VOC, and 0.6 mIOU for ODISE + PMP on PC-459), and the manuscript does not state whether the baseline numbers were re-run in the same codebase or taken from the original papers. Without variance information or a statement about experimental control, it is hard to assess whether the smaller gains exceed noise. Please provide per-seed results or confidence intervals and clarify the origin of the baseline numbers.
- [Sec. 4.2 and A.3] The training-time construction of the text tokens is unspecified. The paper does not state whether the text input during COCO-Stuff training is the full set of 171 class names, extracted nouns from captions, or a sampled subset, nor how M is distributed across training examples. Since the behavior of Eq. (5) depends on M, and since the paper argues for transfer to arbitrary prompts, this detail is necessary both for reproducibility and for assessing the M=1 distribution shift described above.
minor comments (5)
- [Sec. 4.4] The text says the four candidate decoding strategies are compared in Table 5, but the strategy comparison is presented in Table 3; Table 5 contains the backbone and hyperparameter ablations.
- [Eq. (5)] The index notation "{t_j}_{N}^{M}=1" is malformed; it should indicate text tokens indexed j=1,...,M.
- [Appendix A.4] The text states that PMP brings "0.1s" of added latency, but the reported inference times (1.03s vs 1.02s) imply 0.01s of total added latency, and the first-stage difference is also 0.01s.
- [Sec. 4.2] In the OVSeg paragraph, "to combine PMP with ODISE" should read "to combine PMP with OVSeg."
- [Throughout] There are numerous typographical and grammatical errors, including "the in the second stage," "summmer," "ensambling," and "demonstracts"; the manuscript should be carefully proofread before resubmission.
Circularity Check
No significant circularity: the paper reports external benchmark results of a trained cross-attention module, not a derivation that reduces to its inputs.
full rationale
The paper's derivation chain is empirical rather than definitional. The PMP module inserts a text-query cross-attention (Eq. 5) into existing query-based mask decoders, trains on COCO-Stuff, and evaluates on held-out benchmarks (ADE20K, Pascal Context, Pascal VOC) against the unmodified baselines. No parameter is fitted to the benchmark numbers and then reported as a prediction; the reported mIOU gains are external measurements of the trained system. CLIP is used as a fixed feature extractor in both training and evaluation, which is a standard component rather than a tautological reduction. The qualitative claims about arbitrary prompts such as 'love' or 'Times Square' are supported only by examples and may involve an untested distribution shift from many-text-token training to single-prompt inference, but that is an evidence/completeness concern and not a circularity. The paper contains no load-bearing self-citation chain, no uniqueness theorem imported from the authors' prior work, and no renamed known result presented as derivation. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (3)
- lambda (blending factor) =
0.65
- number of decoder layers L =
3
- loss weights lambda_ce, lambda_dice =
5.0, 5.0
assumptions (4)
- domain assumption CLIP text-image embedding space supports zero-shot classification across datasets and prompt types.
- domain assumption Training on COCO-Stuff with class-name and caption text prompts is sufficient for the decoder to learn prompt-conditioned mask generation that generalizes to new datasets.
- domain assumption Adding a text-query cross-attention layer before each standard cross-attention in Mask2Former and SAN decoders does not destabilize the bipartite matching training objective.
- domain assumption Mask proposal recall is the bottleneck in two-stage open-vocab segmentation for arbitrary prompts.
Cite this review
Pith. "Pith review of Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation." pith.science (2026). https://pith.science/paper/SF7J6EAX
@misc{pith2026241210292,
author = {Pith},
title = {Pith review of: Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/SF7J6EAX}},
note = {Machine review of arXiv:2412.10292}
}
read the original abstract
We tackle the challenge of open-vocabulary segmentation, where we need to identify objects from a wide range of categories in different environments, using text prompts as our input. To overcome this challenge, existing methods often use multi-modal models like CLIP, which combine image and text features in a shared embedding space to bridge the gap between limited and extensive vocabulary recognition, resulting in a two-stage approach: In the first stage, a mask generator takes an input image to generate mask proposals, and the in the second stage the target mask is picked based on the query. However, the expected target mask may not exist in the generated mask proposals, which leads to an unexpected output mask. In our work, we propose a novel approach named Prompt-guided Mask Proposal (PMP) where the mask generator takes the input text prompts and generates masks guided by these prompts. Compared with mask proposals generated without input prompts, masks generated by PMP are better aligned with the input prompts. To realize PMP, we designed a cross-attention mechanism between text tokens and query tokens which is capable of generating prompt-guided mask proposals after each decoding. We combined our PMP with several existing works employing a query-based segmentation backbone and the experiments on five benchmark datasets demonstrate the effectiveness of this approach, showcasing significant improvements over the current two-stage models (1% ~ 3% absolute performance gain in terms of mIOU). The steady improvement in performance across these benchmarks indicates the effective generalization of our proposed lightweight prompt-aware method.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems , 35:23716–23736,
-
[2]
Yolact: Real-time instance segmentation
Daniel Bolya, Chong Zhou, Fanyi Xiao, and Yong Jae Lee. Yolact: Real-time instance segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9157–9166, 2019. 2
work page 2019
-
[3]
Lan- guage models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Lan- guage models are few-shot learners. Advances in neural in- formation processing systems, 33:1877–1901, 2020. 2
1901
-
[4]
Coco- stuff: Thing and stuff classes in context
Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco- stuff: Thing and stuff classes in context. In Proceedings of the IEEE conference on computer vision and pattern recog- nition, pages 1209–1218, 2018. 4, 5
work page 2018
-
[5]
Cascade r-cnn: Delv- ing into high quality object detection
Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delv- ing into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 6154–6162, 2018. 2
2018
-
[6]
End-to- end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to- end object detection with transformers. In European confer- ence on computer vision, pages 213–229. Springer, 2020. 2, 4
work page 2020
-
[7]
Hybrid task cascade for instance seg- mentation
Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaox- iao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, et al. Hybrid task cascade for instance seg- mentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4974–4983,
-
[8]
Semantic image segmen- tation with deep convolutional nets and fully connected crfs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmen- tation with deep convolutional nets and fully connected crfs. arXiv preprint arXiv:1412.7062, 2014. 2
arXiv 2014
Show all 77 references
-
[9]
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolu- tion, and fully connected crfs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolu- tion, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–...
2017
-
[10]
Rethinking atrous convolution for seman- tic image segmentation
Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for seman- tic image segmentation. arXiv preprint arXiv:1706.05587 , 2017
2017 arXiv
-
[11]
Encoder-decoder with atrous separable convolution for semantic image segmentation
Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer vision (ECCV), pages 801–818, 2018. 2, 14
2018
-
[12]
Scal- ing wide residual networks for panoptic segmentation
Liang-Chieh Chen, Huiyu Wang, and Siyuan Qiao. Scal- ing wide residual networks for panoptic segmentation. arXiv preprint arXiv:2011.11675, 2020. 2
2011 arXiv
-
[13]
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision , pages 104–120. Springer,
-
[14]
Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation
Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen. Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognit...
2020
-
[15]
Per- pixel classification is not all you need for semantic segmen- tation
Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per- pixel classification is not all you need for semantic segmen- tation. Advances in Neural Information Processing Systems, 34:17864–17875, 2021. 2, 4, 5, 13
2021
-
[16]
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1290–1299, 2022. 2, 4, 5, 6, 13, 14
2022
-
[17]
Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation
Seokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab, Paul Hongsuck Seo, and Seungryong Kim. Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4113– 4123, 2024. 3
2024
-
[18]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 1, 2
2018 arXiv
-
[19]
De- coupling zero-shot semantic segmentation
Jian Ding, Nan Xue, Gui-Song Xia, and Dengxin Dai. De- coupling zero-shot semantic segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11583–11592, 2022. 2, 5
2022
-
[20]
Open- vocabulary universal image segmentation with maskclip
Zheng Ding, Jieke Wang, and Zhuowen Tu. Open- vocabulary universal image segmentation with maskclip
-
[21]
The pascal visual object classes challenge: A retrospective
Mark Everingham, SM Ali Eslami, Luc Van Gool, Christo- pher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. In- ternational journal of computer vision , 111:98–136, 2015. 5
2015
-
[22]
Dual attention network for scene seg- mentation
Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. Dual attention network for scene seg- mentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3146–3154,
-
[23]
Scal- ing open-vocabulary image segmentation with image-level labels
Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Scal- ing open-vocabulary image segmentation with image-level labels. In European Conference on Computer Vision, pages 540–557. Springer, 2022. 1, 2, 5
2022
-
[24]
Multi-scale high-resolution vision transformer for se- mantic segmentation
Jiaqi Gu, Hyoukjun Kwon, Dilin Wang, Wei Ye, Meng Li, Yu-Hsin Chen, Liangzhen Lai, Vikas Chandra, and David Z Pan. Multi-scale high-resolution vision transformer for se- mantic segmentation. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition...
2022
-
[25]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 13
2016
-
[26]
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross Gir- shick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 2
2017
-
[27]
Oneformer: One transformer to rule universal image segmentation
Jitesh Jain, Jiachen Li, Mang Tik Chiu, Ali Hassani, Nikita Orlov, and Humphrey Shi. Oneformer: One transformer to rule universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2989–2998, 2023. 2
2023
-
[28]
Scaling up visual and vision-language representa- tion learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representa- tion learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR,
-
[29]
Instancecut: from edges to instances with multicut
Alexander Kirillov, Evgeny Levinkov, Bjoern Andres, Bog- dan Savchynskyy, and Carsten Rother. Instancecut: from edges to instances with multicut. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 5008–5017, 2017. 2
2017
-
[30]
Panoptic segmentation
Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Doll ´ar. Panoptic segmentation. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9404–9413, 2019. 2
2019
-
[31]
Segment any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. arXiv preprint arXiv:2304.02643, 2023. 7, 8
2023 arXiv
-
[32]
Language-driven semantic seg- mentation
Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic seg- mentation. In International Conference on Learning Rep- resentations, 2022. 1, 2, 5
2022
-
[33]
Mask dino: Towards a unified transformer-based framework for object detection and segmentation
Feng Li, Hao Zhang, Huaizhe Xu, Shilong Liu, Lei Zhang, Lionel M Ni, and Heung-Yeung Shum. Mask dino: Towards a unified transformer-based framework for object detection and segmentation. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pa...
2023
-
[34]
Unifying training and inference for panoptic segmentation
Qizhu Li, Xiaojuan Qi, and Philip HS Torr. Unifying training and inference for panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13320–13328, 2020. 2
2020
-
[35]
Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers
Zhiqi Li, Wenhai Wang, Enze Xie, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, Ping Luo, and Tong Lu. Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages...
2022
-
[36]
Open-vocabulary semantic segmentation with mask-adapted clip
Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pag...
2023
-
[37]
Feature pyra- mid networks for object detection
Tsung-Yi Lin, Piotr Doll ´ar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyra- mid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 2117–2125, 2017. 14
2017
-
[38]
An end-to-end network for panoptic segmentation
Huanyu Liu, Chao Peng, Changqian Yu, Jingbo Wang, Xu Liu, Gang Yu, and Wei Jiang. An end-to-end network for panoptic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 6172–6181, 2019. 2
2019
-
[39]
Path aggregation network for instance segmentation
Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. Path aggregation network for instance segmentation. In Pro- ceedings of the IEEE conference on computer vision and pat- tern recognition, pages 8759–8768, 2018. 2
2018
-
[40]
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettle- moyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692,
1907 arXiv
-
[41]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021. 13
2021
-
[42]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , pages 11976–11986,
-
[43]
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV), pages 565–571. Ieee, 2016. 14
2016
-
[44]
The role of context for object detection and semantic segmentation in the wild
Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. The role of context for object detection and semantic segmentation in the wild. In Proceedings of the IEEE conference on computer vision and pattern recogni...
2014
-
[45]
Detectors: Detecting objects with recursive feature pyramid and switch- able atrous convolution
Siyuan Qiao, Liang-Chieh Chen, and Alan Yuille. Detectors: Detecting objects with recursive feature pyramid and switch- able atrous convolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 10213–10224, 2021. 2
2021
-
[46]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[47]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020. 1
2020
-
[48]
High-resolution image 10 synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image 10 synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3
2022
-
[49]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Pa...
2015
-
[50]
Segmenter: Transformer for semantic segmenta- tion
Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmenta- tion. In Proceedings of the IEEE/CVF international confer- ence on computer vision, pages 7262–7272, 2021. 2
2021
-
[51]
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. Lxmert: Learning cross-modality encoder representations from transformers. EMNLP, 2019. 2
2019
-
[52]
Conditional con- volutions for instance segmentation
Zhi Tian, Chunhua Shen, and Hao Chen. Conditional con- volutions for instance segmentation. In Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, Au- gust 23–28, 2020, Proceedings, Part I 16 , pages 282–298. Springer, 2020. 2
2020
-
[53]
Axial-deeplab: Stand- alone axial-attention for panoptic segmentation
Huiyu Wang, Yukun Zhu, Bradley Green, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Axial-deeplab: Stand- alone axial-attention for panoptic segmentation. InEuropean conference on computer vision , pages 108–126. Springer,
-
[54]
Max-deeplab: End-to-end panoptic segmentation with mask transformers
Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Max-deeplab: End-to-end panoptic segmentation with mask transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5463–5474, 2021. 2
2021
-
[55]
Solov2: Dynamic and fast instance segmenta- tion
Xinlong Wang, Rufeng Zhang, Tao Kong, Lei Li, and Chun- hua Shen. Solov2: Dynamic and fast instance segmenta- tion. Advances in Neural information processing systems , 33:17721–17732, 2020. 2
2020
-
[56]
Segformer: Simple and efficient design for semantic segmentation with transform- ers
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transform- ers. Advances in Neural Information Processing Systems , 34:12077–12090, 2021. 2
2021
-
[57]
Upsnet: A unified panoptic segmentation network
Yuwen Xiong, Renjie Liao, Hengshuang Zhao, Rui Hu, Min Bai, Ersin Yumer, and Raquel Urtasun. Upsnet: A unified panoptic segmentation network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8818–8826, 2019. 2
2019
-
[58]
Groupvit: Semantic segmentation emerges from text supervision
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 18134–18144, 2022. 2, 5
2022
-
[59]
Open-vocabulary panop- tic segmentation with text-to-image diffusion models
Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiao- long Wang, and Shalini De Mello. Open-vocabulary panop- tic segmentation with text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 2955–2966, 2023. ...
2023
-
[60]
A simple baseline for zero- shot semantic segmentation with pre-trained vision-language model
Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, and Xiang Bai. A simple baseline for zero- shot semantic segmentation with pre-trained vision-language model. ECCV, 3, 2021. 1, 2, 5
2021
-
[61]
A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model
Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, and Xiang Bai. A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model. In European Conference on Computer Vi- sion, pages 736–753. Springer, 2022. 5
2022
-
[62]
Side adapter network for open-vocabulary semantic segmentation
Mengde Xu, Zheng Zhang, Fangyun Wei, Han Hu, and Xi- ang Bai. Side adapter network for open-vocabulary semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2945– 2954, 2023. 2, 5, 6, 13, 15
2023
-
[63]
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mo- jtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022. 2
2022 arXiv
-
[64]
Cmt-deeplab: Clustering mask transformers for panoptic segmentation
Qihang Yu, Huiyu Wang, Dahun Kim, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Cmt-deeplab: Clustering mask transformers for panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition...
2022
-
[65]
k-means mask transformer
Qihang Yu, Huiyu Wang, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. k-means mask transformer. In European Conference on Computer Vision, pages 288–307. Springer, 2022. 2
2022
-
[66]
Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip
Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, and Liang- Chieh Chen. Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip. arXiv preprint arXiv:2308.02487, 2023. 3, 6, 7, 13, 14, 15
2023 arXiv
-
[67]
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021. 2
2021 arXiv
-
[68]
Object- contextual representations for semantic segmentation
Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object- contextual representations for semantic segmentation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16, pages 173–190. Springer, 2020. 2
2020
-
[69]
Open-vocabulary object detection using captions
Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih- Fu Chang. Open-vocabulary object detection using captions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14393–14402, 2021. 1
2021
-
[70]
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5579–5588,
-
[71]
Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers. In Proceedings of the IEEE/CVF conference...
-
[72]
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 633–641,
-
[73]
Extract free dense labels from clip
Chong Zhou, Chen Change Loy, and Bo Dai. Extract free dense labels from clip. In European Conference on Com- puter Vision, pages 696–712. Springer, 2022. 2, 5
2022
-
[74]
Zegclip: Towards adapting clip for zero-shot se- mantic segmentation
Ziqin Zhou, Yinjie Lei, Bowen Zhang, Lingqiao Liu, and Yifan Liu. Zegclip: Towards adapting clip for zero-shot se- mantic segmentation. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 11175–11185, 2023. 2, 5
2023
-
[75]
Deformable detr: Deformable trans- formers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable trans- formers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 14
2010 arXiv
-
[76]
Generalized decoding for pixel, image, and lan- guage
Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, et al. Generalized decoding for pixel, image, and lan- guage. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages...
-
[2023]
MIT CSAIL
2, 5 12 A. Appendix A.1. More ablation studies In this section we provide more ablation studies on either each stage of pipeline, the backbones, or the hyperparam- eters to further analyze the sensitivity of them in Table 4 and Table 5. We conduct the ablation studies on top o...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.