REVIEW 5 major objections 6 minor 221 references
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that the entire video scene parsing literature can be organized as five related tasks on one architectural arc, with four cross-cutting failure modes.
desk verdict A useful organizational survey whose benchmark tables currently are not trustworthy enough to serve as the reference it aims to be. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is the five-task taxonomy crossed with the architectural arc. Each task is given a formal input-output definition: video semantic segmentation maps a clip $V \in \mathbb{R}^{T \times H \times W \times 3}$ to per-pixel semantic labels, video instance segmentation adds instance IDs, video panoptic segmentation adds "stuff" classes plus tracked "thing" instances, video tracking and segmentation adds persistent identities, and open-vocabulary video segmentation replaces the fixed label set with an open vocabulary drawn from a vision-language model. The architectural arc is the historical axis: over-segmentation and hand-crafted spatiotemporal cues, frame-wise fully convolutional parsing with flow or conditional-random-field post-processing, attention-based temporal aggregation, end-to-end query-based transformers, and foundation-model components such as CLIP, SAM, and diffusion backbones. Placing every method in this task-by-architecture grid lets the paper read across rows for shared trade-offs, like accuracy versus latency and identity stability versus flexibility, and down columns for the evolution of temporal modeling, which is how it arrives at the four cross-cutting failure modes.
What would settle it
Run a matched-protocol re-evaluation: retrain the leading VSS, VIS, and VPS methods with the same backbone, pre-training, and epochs, and measure mIoU, VPQ, STQ, and IDS alongside the values in Tables 6-10; if the current leaders change or the reported gaps narrow, the survey's comparative claims are not robust.
Extended reading notes
Core claim
The paper's central discovery is organizational: despite separate benchmarks and communities, video semantic segmentation, video instance segmentation, video panoptic segmentation, video tracking and segmentation, and open-vocabulary video segmentation are facets of one problem, assigning every pixel a semantic name while keeping object identity coherent over time. The survey arranges the methods into task-specific families, such as flow-based, attention-based, real-time, and semi-supervised methods for semantic segmentation, and query-based, depth-aware, and dual-branch methods for panoptic segmentation, and threads them together with one architectural arc from hand-crafted features to transformers and foundation models. It claims that four failure modes, temporal flicker, occlusion-induced identity switches, long-tail categories, and the annotation-capacity-latency tension, cut across all five tasks, and it uses benchmark tables to name current leaders in each setting. The stated conclusion is that the field is moving from closed-set, frame-wise parsers toward open-world, unified, efficient systems, with multimodal fusion, visual reasoning, generative segmentation, and large-language-model-based segmentation as the active frontiers.
Load-bearing premise
The comparisons in Section 5 assume that numbers reported by different papers with different backbones and training protocols are directly comparable; if that assumption fails, the claimed leaders and accuracy rankings are not established.
Editorial extensions
If this is right
- If the survey's benchmark tables are taken at face value, VPSeg is the current semantic-segmentation accuracy leader on Cityscapes, CTVIS on YouTube-VIS, PolyphonicFormer on Cityscapes-VPS, Video K-Net on KITTI-STEP, and TubeLink on VIPSeg.
- The recurring failure modes give a reporting checklist: new video scene parsing methods should report temporal consistency (mVC, VPQ), identity stability (IDS, STQ), long-tail behavior, and latency, not only mean accuracy.
- Open-vocabulary video segmentation remains far behind fully supervised methods on the same benchmarks, so the open direction is to close that gap while preserving the ability to name novel categories.
- The paper's forward-looking claim is that the next generation of video scene parsing systems will be open-world, unified across tasks, multimodal, reasoning-capable, generative, efficient, and built on large language and foundation models.
Reading between the lines
- The taxonomy's usefulness could be tested by trying to place recent hybrid methods, such as state-space or language-model-prompted segmenters, in exactly one task family; the survey's own inclusion of TV3S and SAM-based trackers suggests several will straddle two families.
- If the four failure modes are truly cross-cutting, a single diagnostic suite using videos with forced occlusions, rare categories, and variable frame rates could replace the current per-dataset metrics for comparing methods across tasks.
- The benchmark ranking claims are sensitive to backbone choice and training protocol, so a matched-protocol re-run would likely reorder the leaders without overturning the qualitative division of methods into accuracy-oriented and efficiency-oriented families.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript surveys Video Scene Parsing (VSP), organizing the literature across five tasks (Video Semantic Segmentation, Video Instance Segmentation, Video Panoptic Segmentation, Video Tracking & Segmentation, and Open-Vocabulary Video Segmentation) and along one architectural arc from hand-crafted cues through fully convolutional, attention-based, query-based, and foundation-model approaches. It provides task definitions, a method taxonomy, dataset statistics, metric formulations, benchmark tables, and a set of future directions. The stated goal is to serve as a holistic and reliable integrative reference for the field.
Significance. If the survey were fully accurate, it would fill a useful niche by covering five related video segmentation tasks in one place, with a unified architectural narrative and cross-task discussions of temporal consistency, identity preservation, and efficiency. The paper is commendably broad: it includes recent open-vocabulary methods, emerging state-space models, and a dedicated metrics section. However, the survey's central value is as a reference, and the current reliability problems in the benchmark tables and dataset statistics—especially the inclusion of a non-video method in the VSS efficiency comparison and the contradictory statements about the best mVC8 method—undermine that central claim until corrected.
major comments (5)
- [Section 5.1.2, Table 6] The VSS efficiency and accuracy comparison in Table 6 is not valid as presented because BLO [107] is an image semantic segmentation pruning method with no video experiments, yet it is listed with 'Keyframe Selection' and trained on image datasets (Cityscapes, ADE20K, Pascal VOC). Using BLO to support the claim that it 'obtains the best inference speed with 30.8 FPS' is therefore unsupported and makes the ranking in Section 5.1.2 misleading. This table should either exclude non-video methods or clearly separate them with a statement that the numbers are not directly comparable.
- [Section 5.1.2 and Table 7] The text contains a direct contradiction: it states that 'TubeFormer achieves the best performance on all mIoU (63.2) and mVC8 (92.1),' then immediately says 'TV3S achieves the best performance on mVC8,' even though the same table reports mVC8 = 92.1 for TubeFormer and mVC8 = 91.7 for TV3S. This needs to be corrected to identify TubeFormer as the best on both mIoU and mVC8 (or the table numbers revised if they are wrong).
- [Section 4.1.1 and Table 5] The CamVid statistics are inconsistent between the text and the table: Section 4.1.1 states 'five continuous videos' with a training/validation/test frame split of 367/101/233, while Table 5 reports 4 videos with a split of 467/100/233. Similarly, YouTube-VIS is given as 3,859 videos with a 2,985/421/453 split in Table 5, but Section 4.1.2 states 2,883 videos with 2,238/302/343. Since Section 4 is the authoritative dataset reference for the rest of the paper, these discrepancies cast doubt on the reliability of the survey as a reference and must be resolved against the original dataset papers.
- [Section 5.1.2, 5.2.2, and Tables 6-10] The benchmark comparisons mix methods with different backbone architectures (ResNet-101, MiT-B1, Swin-L, MiT-B3, searched backbones), different training sets, and different evaluation protocols, yet the text draws unqualified conclusions such as 'VPSeg achieves the highest segmentation accuracy' and 'CTVIS ... setting the new state-of-the-art.' Without controlled conditions or explicit caveats about these confounds, these numerical rankings are not supported. The paper should either add a clear limitations paragraph explaining the non-comparability or restructure the tables to group methods by comparable settings.
- [Section 5.5 and Tables 7/8] The open-vocabulary comparison is presented in the same tables as fully supervised methods without a clear protocol distinction: OV2VSS, OVFormer, OV2Seg+, and CLIP-VIS are evaluated under open-vocabulary training regimes and report low AP/mIoU values, while the surrounding rows are closed-set methods. The caption notes 'Methods in gray use open-vocabulary supervision,' but the text does not explain how these numbers were obtained or whether the comparison is intended to be quantitative at all. This should be clarified, and the comparison framed as illustrative rather than a direct ranking.
minor comments (6)
- [Table 6] There is a comma-decimal typo in the CFFM row ('75,1' instead of '75.1'), and the BLO row shows '74.730.8' with no separator between mIoU and FPS.
- [Equation (3)] The Video Consistency formula appears to contain indexing errors: the two intersection terms use inconsistent subscripts (i+j and i−j), and the intended intersection over n consecutive frames is not clearly expressed. Please revise the notation.
- [Equation (14)] The IDS formula uses c−1(m) as a subscript in 'id_{c−1(m)}' and 'id_{c−1(pred(m))}', which is confusing because c−1(m) is already defined as a matched prediction. The notation should be made clearer.
- [Section 5.3.2 and 5.4.2] Section 5.3.2 says 'nine VIS approaches' when reporting VPS results, and Section 5.4.2 says 'six VIS methods' when reporting VTS results on KITTI-MOTS; both should say VPS and VTS respectively.
- [Section 4.1.1 (NYUDv2)] The description of NYUDv2 as '464 novel scenes from three cities' needs a source check; the original dataset is a single indoor environment dataset and the phrase 'from three cities' is not standard in the dataset's official description.
- [References] Reference [161] for HiEve is listed as an arXiv preprint even though Table 4 says IJCV 2023; please update the reference to the published version if it exists.
Circularity Check
No significant circularity: the survey's taxonomy and benchmark comparisons are external descriptive syntheses, not derivations from fitted inputs or self-cited theorems.
full rationale
This manuscript is a literature survey rather than a derivation chain: it organizes published methods into a taxonomy and compiles benchmark numbers from the cited papers. No equation in the paper is derived from a fitted parameter, and no prediction is produced from a model whose inputs are the outcome being asserted. The authors' own methods (CFFM, MRCFA, CFFM+, TV3S, OV2VSS) receive prominent coverage and are cited from the authors' prior work, but the survey's central claims—that VSP spans five tasks and an architectural arc, and that certain methods lead the benchmarks—do not rest on those self-citations as evidence for the claims themselves; the cited methods are independently published and externally benchmarked. The internal contradictions and possible benchmark-table errors noted by a careful reader (e.g., the TubeFormer/TV3S mVC8 statement in Sec. 5.1.2, the inclusion of the image-only BLO method in Table 6, and dataset-statistic mismatches in Table 5 and the text) are correctness and reliability concerns about the survey's synthesis, not circularity: they do not make any claim equivalent to its own input by construction. Accordingly, the circularity score is 0, while the survey's accuracy as a reference is a separate question.
Assumptions & free parameters
assumptions (2)
- domain assumption The papers selected for review are representative of the field and their reported results are accurately transcribed.
- domain assumption The five-task taxonomy (VSS, VIS, VPS, VTS, OVVS) is a meaningful organizing principle for the field.
Cite this review
Pith. "Pith review of A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects." pith.science (2026). https://pith.science/paper/OFZU2NNG
@misc{pith2026250613552,
author = {Pith},
title = {Pith review of: A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects},
year = {2026},
howpublished = {\url{https://pith.science/paper/OFZU2NNG}},
note = {Machine review of arXiv:2506.13552}
}
read the original abstract
Video Scene Parsing (VSP) studies dense video understanding, where every pixel in each frame must be segmented, each region must be named, and each object identity must remain coherent over time. This survey reviews recent progress in VSP across five tasks, spanning Video Semantic Segmentation (VSS), Video Instance Segmentation (VIS), Video Panoptic Segmentation (VPS), Video Tracking \& Segmentation (VTS), and Open-Vocabulary Video Segmentation (OVVS). We organize the literature as one architectural arc, running from hand-crafted motion and appearance cues through fully convolutional, attention-based and query-based designs to recent foundation-model approaches, and we trace how each family models temporal context, preserves identity, and balances accuracy against efficiency. We then compare the datasets, metrics and benchmark trends that shape current evaluation. Beyond cataloguing methods, we foreground the design trade-offs and recurring failure modes that cut across the field, namely temporal flicker, occlusion-induced identity switches, long-tail categories, and the annotation--capacity--latency tension. We close with open directions towards robust, efficient and open-world VSP systems.
Figures
Reference graph
Works this paper leans on
-
[107]
Pruning parameterization with bi-level optimiza- tion for efficient semantic segmentation on the edge,
C. Yang, P. Zhao, Y . Li, W. Niu, J. Guan, H. Tang, M. Qin, B. Ren, X. Lin, and Y . Wang, “Pruning parameterization with bi-level optimiza- tion for efficient semantic segmentation on the edge,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 15 402–15 412
2023
-
[1]
Fully convolutional networks for semantic segmentation,
E. Shelhamer, J. Long, and T. Darrell, “Fully convolutional networks for semantic segmentation,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 4, pp. 640–651, 2017
2017
-
[2]
Deep feature flow for video recognition,
X. Zhu, Y . Xiong, J. Dai, L. Yuan, and Y . Wei, “Deep feature flow for video recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 2349–2358
2017
-
[3]
Semantic video CNNs through representation warping,
R. Gadde, V . Jampani, and P. Gehler, “Semantic video CNNs through representation warping,” inInt. Conf. Comput. Vis., 2017, pp. 4463– 4472
2017
-
[4]
Accel: A corrective fusion network for efficient semantic segmentation on video,
S. Jain, X. Wang, and J. E. Gonzalez, “Accel: A corrective fusion network for efficient semantic segmentation on video,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 8858–8867
2018
-
[5]
Feature space optimization for semantic video segmentation,
A. Kundu, V . Vineet, and V . Koltun, “Feature space optimization for semantic video segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 3168–3175
2016
-
[6]
GSVNet: Guided spatially- varying convolution for fast semantic segmentation on video,
S.-P. Lee, S.-C. Chen, and W.-H. Peng, “GSVNet: Guided spatially- varying convolution for fast semantic segmentation on video,” inIEEE Int. Conf. Multimedia Expo, 2021, pp. 1–6
2021
-
[7]
Faster R-CNN: Towards real-time object detection with region proposal networks,
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 6, pp. 1137–1149, 2015
2015
Show all 221 references
-
[8]
The 2018 DA VIS challenge on video object segmentation,
S. Caelles, A. Montes, K.-K. Maninis, Y . Chen, L. Van Gool, F. Per- azzi, and J. Pont-Tuset, “The 2018 DA VIS challenge on video object segmentation,”arXiv preprint arXiv:1803.00557, 2018
2018 arXiv
-
[9]
Video segmentation using color difference histogram,
C. Lam and M.-C. Lee, “Video segmentation using color difference histogram,” inIAPR International Workshop on Multimedia Information Analysis and Retrieval, 1998, pp. 159–174
1998
-
[10]
Dynamic texture detection based on motion analysis,
S. Fazekas, T. Amiaz, D. Chetverikov, and N. Kiryati, “Dynamic texture detection based on motion analysis,”Int. J. Comput. Vis., vol. 82, pp. 48–63, 2009
2009
-
[11]
Illumination robust optical flow model based on histogram of oriented gradients,
H. A. Rashwan, M. A. Mohamed, M. A. Garc ´ıa, B. Mertsching, and D. Puig, “Illumination robust optical flow model based on histogram of oriented gradients,” inPattern Recogn.Springer, 2013, pp. 354–363
2013
-
[12]
Efficient video seg- mentation using parametric graph partitioning,
C.-P. Yu, H. M. Le, G. J. Zelinsky, and D. Samaras, “Efficient video seg- mentation using parametric graph partitioning,” inInt. Conf. Comput. Vis., 2015, pp. 3155–3163
2015
-
[13]
Efficient hierarchical graph-based video segmentation,
M. Grundmann, V . Kwatra, M. Han, and I. Essa, “Efficient hierarchical graph-based video segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2010, pp. 2141–2148
2010
-
[14]
Fully connected object proposals for video segmentation,
F. Perazzi, O. Wang, M. H. Gross, and A. Sorkine-Hornung, “Fully connected object proposals for video segmentation,” inInt. Conf. Comput. Vis., 2015, pp. 3227–3234
2015
-
[15]
Semi-supervised video segmentation using tree structured graphical models,
V . Badrinarayanan, I. Budvytis, and R. Cipolla, “Semi-supervised video segmentation using tree structured graphical models,” inIEEE Trans. Pattern Anal. Mach. Intell., vol. 35, no. 11. IEEE, 2013, pp. 2751– 2764
2013
-
[16]
Streaming video segmentation via short- term hierarchical segmentation and frame-by-frame markov random field optimization,
W.-D. Jang and C.-S. Kim, “Streaming video segmentation via short- term hierarchical segmentation and frame-by-frame markov random field optimization,” inEur . Conf. Comput. Vis., 2016, pp. 599–615
2016
-
[17]
Multiclass semantic video segmentation with object- level active inference,
B. Liu and X. He, “Multiclass semantic video segmentation with object- level active inference,” inIEEE Conf. Comput. Vis. Pattern Recog., 2015, pp. 4286–4294
2015
-
[18]
DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs,
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 4, pp. 834–848, 2018
2018
-
[19]
Pyramid scene parsing network,
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 2881– 2890. 16
2017
-
[20]
Context contrasted feature and gated multi-scale aggregation for scene segmen- tation,
H. Ding, X. Jiang, B. Shuai, A. Q. Liu, and G. Wang, “Context contrasted feature and gated multi-scale aggregation for scene segmen- tation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 2393– 2402
2018
-
[21]
Video segmentation via object flow,
Y .-H. Tsai, M.-H. Yang, and M. J. Black, “Video segmentation via object flow,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 3899–3908
2016
-
[22]
Improving semantic segmentation via video propagation and label relaxation,
Y . Zhu, K. Sapra, F. A. Reda, K. J. Shih, S. Newsam, A. Tao, and B. Catanzaro, “Improving semantic segmentation via video propagation and label relaxation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 8856–8865
2019
-
[23]
Efficient uncertainty estimation for semantic segmentation in videos,
P.-Y . Huang, W.-T. Hsu, C.-Y . Chiu, T.-F. Wu, and M. Sun, “Efficient uncertainty estimation for semantic segmentation in videos,” inEur . Conf. Comput. Vis., 2018, pp. 520–535
2018
-
[24]
Efficient video semantic segmentation with labels propagation and refinement,
M. Paul, C. Mayer, L. V . Gool, and R. Timofte, “Efficient video semantic segmentation with labels propagation and refinement,” inIEEE Winter Conf. App. Comput. Vis., 2020, pp. 2873–2882
2020
-
[25]
ICNet for real-time semantic segmentation on high-resolution images,
H. Zhao, X. Qi, X. Shen, J. Shi, and J. Jia, “ICNet for real-time semantic segmentation on high-resolution images,” inEur . Conf. Comput. Vis., 2018, pp. 405–420
2018
-
[26]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” inAnnu. Conf. Neur . Inform. Process. Syst., 2017, pp. 6000–6010
2017
-
[27]
Scaling vision transformers,
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, “Scaling vision transformers,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 12 104–12 113
2022
-
[28]
AdaViT: Adaptive vision transformers for efficient image recognition,
L. Meng, H. Li, B.-C. Chen, S. Lan, Z. Wu, Y .-G. Jiang, and S.-N. Lim, “AdaViT: Adaptive vision transformers for efficient image recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 12 309–12 318
2022
-
[29]
Vision transformers with hierarchical attention,
Y . Liu, Y .-H. Wu, G. Sun, L. Zhang, A. Chhatkuli, and L. Van Gool, “Vision transformers with hierarchical attention,”Machine Intelligence Research, vol. 21, no. 4, pp. 670–683, 2024
2024
-
[30]
Swin transformer v2: Scaling up capacity and resolution,
Z. Liu, H. Hu, Y . Lin, Z. Yao, Z. Xie, Y . Wei, J. Ning, Y . Cao, Z. Zhang, L. Donget al., “Swin transformer v2: Scaling up capacity and resolution,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 12 009–12 019
2022
-
[31]
Scaling vision transformers to gigapixel images via hierarchical self-supervised learning,
R. J. Chen, C. Chen, Y . Li, T. Y . Chen, A. D. Trister, R. G. Krishnan, and F. Mahmood, “Scaling vision transformers to gigapixel images via hierarchical self-supervised learning,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 16 144–16 155
2022
-
[32]
Multiview transformers for video recognition,
S. Yan, X. Xiong, A. Arnab, Z. Lu, M. Zhang, C. Sun, and C. Schmid, “Multiview transformers for video recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 3333–3343
2022
-
[33]
Multi-scale high-resolution vision transformer for semantic segmentation,
J. Gu, H. Kwon, D. Wang, W. Ye, M. Li, Y .-H. Chen, L. Lai, V . Chandra, and D. Z. Pan, “Multi-scale high-resolution vision transformer for semantic segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 12 094–12 103
2022
-
[34]
UniRepLKNet: A universal perception large-kernel ConvNet for audio video point cloud time-series and image recognition,
X. Ding, Y . Zhang, Y . Ge, S. Zhao, L. Song, X. Yue, and Y . Shan, “UniRepLKNet: A universal perception large-kernel ConvNet for audio video point cloud time-series and image recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 5513–5524
2024
-
[35]
P2T: Pyramid pooling transformer for scene understanding,
Y .-H. Wu, Y . Liu, X. Zhan, and M.-M. Cheng, “P2T: Pyramid pooling transformer for scene understanding,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 11, pp. 12 760–12 771, 2022
2022
-
[36]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” inInt. Conf. Learn. Represent., 2021
2021
-
[37]
End-to-end object detection with transformers,
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” inEur . Conf. Comput. Vis., 2020, pp. 213—-229
2020
-
[38]
MOTS: Multi-object tracking and segmenta- tion,
P. V oigtlaender, M. Krause, A. Osep, J. Luiten, B. B. G. Sekar, A. Geiger, and B. Leibe, “MOTS: Multi-object tracking and segmenta- tion,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 7942–7951
2019
-
[39]
Tracking anything with decoupled video segmentation,
H. K. Cheng, S. W. Oh, B. Price, A. Schwing, and J.-Y . Lee, “Tracking anything with decoupled video segmentation,” inInt. Conf. Comput. Vis., 2023, pp. 1316–1326
2023
-
[40]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inInt. Conf. Mach. Learn., 2021, pp. 8748–8763
2021
-
[41]
Global knowledge calibration for fast open- vocabulary segmentation,
K. Han, Y . Liu, J. H. Liew, H. Ding, J. Liu, Y . Wang, Y . Tang, Y . Yang, J. Feng, Y . Zhaoet al., “Global knowledge calibration for fast open- vocabulary segmentation,” inInt. Conf. Comput. Vis., 2023, pp. 797– 807
2023
-
[42]
A simple framework for open-vocabulary segmentation and detection,
H. Zhang, F. Li, X. Zou, S. Liu, C. Li, J. Yang, and L. Zhang, “A simple framework for open-vocabulary segmentation and detection,” in Int. Conf. Comput. Vis., 2023, pp. 1020–1031
2023
-
[43]
Learning open-vocabulary semantic segmentation models from natural language supervision,
J. Xu, J. Hou, Y . Zhang, R. Feng, Y . Wang, Y . Qiao, and W. Xie, “Learning open-vocabulary semantic segmentation models from natural language supervision,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 2935–2944
2023
-
[44]
Open-vocabulary semantic segmentation with decoupled one-pass network,
C. Han, Y . Zhong, D. Li, K. Han, and L. Ma, “Open-vocabulary semantic segmentation with decoupled one-pass network,” inInt. Conf. Comput. Vis., 2023, pp. 1086–1096
2023
-
[45]
Open-vocabulary panoptic segmentation with embedding modulation,
X. Chen, S. Li, S.-N. Lim, A. Torralba, and H. Zhao, “Open-vocabulary panoptic segmentation with embedding modulation,” inInt. Conf. Com- put. Vis., 2023, pp. 1141–1150
2023
-
[46]
A survey on deep learning technique for video segmentation,
W. Wang, T. Zhou, F. M. Porikli, D. J. Crandall, and L. V . Gool, “A survey on deep learning technique for video segmentation,” inIEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 6. IEEE, 2022, pp. 7099–7122
2022
-
[47]
Transformer-based visual segmentation: A survey,
X. Li, H. Ding, W. Zhang, H. Yuan, J. Pang, G. Cheng, K. Chen, Z. Liu, and C. C. Loy, “Transformer-based visual segmentation: A survey,” in IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 12. IEEE, 2024, pp. 10 138–10 163
2024
-
[48]
Large- scale video panoptic segmentation in the wild: A benchmark,
J. Miao, X. Wang, Y . Wu, W. Li, X. Zhang, Y . Wei, and Y . Yang, “Large- scale video panoptic segmentation in the wild: A benchmark,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 21 001–21 011
2022
-
[49]
Machine perception of three-dimensional, so lids,
P. E. DO CT OR OF, “Machine perception of three-dimensional, so lids,” Ph.D. dissertation, PhD thesis, MASSACHUSETTS INSTITUTE OF TECHNOLOGY , 1961
1961
-
[50]
Evaluation of super-voxel methods for early video processing,
C. Xu and J. J. Corso, “Evaluation of super-voxel methods for early video processing,” inIEEE Conf. Comput. Vis. Pattern Recog., 2012, pp. 1202–1209
2012
-
[51]
A video representation using temporal superpixels,
J. Chang, D. Wei, and J. W. Fisher, “A video representation using temporal superpixels,” inIEEE Conf. Comput. Vis. Pattern Recog., 2013, pp. 2051–2058
2013
-
[52]
Segmentation of moving objects by long term video analysis,
P. Ochs, J. Malik, and T. Brox, “Segmentation of moving objects by long term video analysis,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 36, no. 6, pp. 1187–1200, 2013
2013
-
[53]
Fast object segmentation in uncon- strained video,
A. Papazoglou and V . Ferrari, “Fast object segmentation in uncon- strained video,” inInt. Conf. Comput. Vis., 2013, pp. 1777–1784
2013
-
[54]
Coherent motion segmentation in moving camera videos using optical flow orientations,
M. Narayana, A. Hanson, and E. Learned-Miller, “Coherent motion segmentation in moving camera videos using optical flow orientations,” inInt. Conf. Comput. Vis., 2013, pp. 1577–1584
2013
-
[55]
Layered segmentation and optical flow estimation over time,
D. Sun, E. B. Sudderth, and M. J. Black, “Layered segmentation and optical flow estimation over time,” inIEEE Conf. Comput. Vis. Pattern Recog.IEEE, 2012, pp. 1768–1775
2012
-
[56]
Dynamic color flow: A motion- adaptive color model for object segmentation in video,
X. Bai, J. Wang, and G. Sapiro, “Dynamic color flow: A motion- adaptive color model for object segmentation in video,” inEur . Conf. Comput. Vis., 2010, pp. 617–630
2010
-
[57]
Steadyflow: Spatially smooth optical flow for video stabilization,
S. Liu, L. Yuan, P. Tan, and J. Sun, “Steadyflow: Spatially smooth optical flow for video stabilization,” inIEEE Conf. Comput. Vis. Pattern Recog., 2014, pp. 4209–4216
2014
-
[58]
Classifier based graph construction for video segmentation,
A. Khoreva, F. Galasso, M. Hein, and B. Schiele, “Classifier based graph construction for video segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2015, pp. 951–960
2015
-
[59]
Jf-cut: A parallel graph cut approach for large-scale image and video,
Y . Peng, L. Chen, F.-X. Ou-Yang, W. Chen, and J.-H. Yong, “Jf-cut: A parallel graph cut approach for large-scale image and video,”IEEE Trans. Image Process., vol. 24, no. 2, pp. 655–666, 2014
2014
-
[60]
Video scene parsing with predictive feature learning,
X. Jin, X. Li, H. Xiao, X. Shen, Z. Lin, J. Yang, Y . Chen, J. Dong, L. Liu, Z. Jieet al., “Video scene parsing with predictive feature learning,” in Int. Conf. Comput. Vis., 2017, pp. 5580–5588
2017
-
[61]
Neural window fully- connected CRFs for monocular depth estimation,
W. Yuan, X. Gu, Z. Dai, S. Zhu, and P. Tan, “Neural window fully- connected CRFs for monocular depth estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 3916–3925
2022
-
[62]
Video instance segmentation,
L. Yang, Y . Fan, and N. Xu, “Video instance segmentation,” inInt. Conf. Comput. Vis., 2019, pp. 5187–5196
2019
-
[63]
Naive-Student: Leveraging semi- supervised learning in video sequences for urban scene segmentation,
L.-C. Chen, R. G. Lopes, B. Cheng, M. D. Collins, E. D. Cubuk, B. Zoph, H. Adam, and J. Shlens, “Naive-Student: Leveraging semi- supervised learning in video sequences for urban scene segmentation,” inEur . Conf. Comput. Vis., 2020, pp. 695–714
2020
-
[64]
Three ways to improve semantic segmentation with self-supervised depth estimation,
L. Hoyer, D. Dai, Y . Chen, A. K ¨oring, S. Saha, and L. V . Gool, “Three ways to improve semantic segmentation with self-supervised depth estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 11 125–11 135
2020
-
[65]
Simultaneously short- and long-term temporal modeling for semi- supervised video semantic segmentation,
J. Lao, W. Hong, X. Guo, Y . Zhang, J. Wang, J. Chen, and W. Chu, “Simultaneously short- and long-term temporal modeling for semi- supervised video semantic segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 14 763–14 772. 17
2023
-
[66]
Clockwork convnets for video semantic segmentation,
E. Shelhamer, K. Rakelly, J. Hoffman, and T. Darrell, “Clockwork convnets for video semantic segmentation,” inEur . Conf. Comput. Vis., 2016, pp. 852–868
2016
-
[67]
Real-time, accurate, and consistent video semantic segmentation via unsupervised adaptation and cross-unit deployment on mobile device,
H. Park, A. Yessenbayev, T. Singhal, N. K. Adhikari, Y . Zhang, S. M. Borse, H. Cai, N. P. Pandey, F. Yin, F. Mayeret al., “Real-time, accurate, and consistent video semantic segmentation via unsupervised adaptation and cross-unit deployment on mobile device,” inIEEE Conf. Com...
2022
-
[68]
Dynamic video segmen- tation network,
Y .-S. Xu, T.-J. Fu, H.-K. Yang, and C.-Y . Lee, “Dynamic video segmen- tation network,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 6556–6565
2018
-
[69]
Low-latency video semantic segmentation,
Y . Li, J. Shi, and D. Lin, “Low-latency video semantic segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 5997–6005
2018
-
[70]
Coarse-to- fine feature mining for video semantic segmentation,
G. Sun, Y . Liu, H. Ding, T. Probst, and L. van Gool, “Coarse-to- fine feature mining for video semantic segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 3116–3127
2022
-
[71]
Mining relations among cross-frame affinities for video semantic segmentation,
G. Sun, Y . Liu, H. Tang, A. Chhatkuli, L. Zhang, and L. Van Gool, “Mining relations among cross-frame affinities for video semantic segmentation,” inEur . Conf. Comput. Vis., 2022, pp. 522–539
2022
-
[72]
Deep dual learning for semantic image segmentation,
P. Luo, G. Wang, L. Lin, and X. Wang, “Deep dual learning for semantic image segmentation,” inInt. Conf. Comput. Vis., 2017, pp. 2718–2726
2017
-
[73]
Learning pixel-level semantic affinity with image- level supervision for weakly supervised semantic segmentation,
J. Ahn and S. Kwak, “Learning pixel-level semantic affinity with image- level supervision for weakly supervised semantic segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 4981–4990
2018
-
[74]
PANet: Few-shot image semantic segmentation with prototype alignment,
K. Wang, J. H. Liew, Y . Zou, D. Zhou, and J. Feng, “PANet: Few-shot image semantic segmentation with prototype alignment,” inInt. Conf. Comput. Vis., 2019, pp. 9197–9206
2019
-
[75]
Single-stage semantic segmentation from image labels,
N. Araslanov and S. Roth, “Single-stage semantic segmentation from image labels,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 4253–4262
2020
-
[76]
CIAN: Cross-image affinity net for weakly supervised semantic segmentation,
J. Fan, Z. Zhang, T. Tan, C. Song, and J. Xiao, “CIAN: Cross-image affinity net for weakly supervised semantic segmentation,” inAAAI Conf. Artif. Intell., 2020, pp. 10 762–10 769
2020
-
[77]
Budget-aware deep semantic video segmentation,
B. Mahasseni, S. Todorovic, and A. Fern, “Budget-aware deep semantic video segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 1029–1038
2017
-
[78]
Convolutional gated recurrent networks for video segmentation,
M. Siam, S. Valipour, M. Jagersand, and N. Ray, “Convolutional gated recurrent networks for video segmentation,” inInt. Conf. Image Process., 2017, pp. 3090–3094
2017
-
[79]
One-shot video object segmentation,
S. Caelles, K.-K. Maninis, J. Pont-Tuset, L. Leal-Taix ´e, D. Cremers, and L. Van Gool, “One-shot video object segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 221–230
2017
-
[80]
Space-time memory networks for video object segmentation with user guidance,
S. W. Oh, J.-Y . Lee, N. Xu, and S. J. Kim, “Space-time memory networks for video object segmentation with user guidance,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 1, pp. 442–455, 2020
2020
-
[81]
Learning video object segmentation from static images,
F. Perazzi, A. Khoreva, R. Benenson, B. Schiele, and A. Sorkine- Hornung, “Learning video object segmentation from static images,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 2663–2672
2017
-
[82]
RVOS: End-to-end recurrent network for video object segmentation,
C. Ventura, M. Bellver, A. Girbau, A. Salvador, F. Marques, and X. Giro-i Nieto, “RVOS: End-to-end recurrent network for video object segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 5277–5286
2019
-
[83]
Putting the object back into video object segmentation,
H. K. Cheng, S. W. Oh, B. Price, J.-Y . Lee, and A. Schwing, “Putting the object back into video object segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 3151–3161
2024
-
[84]
Towards high performance video object detection,
X. Zhu, J. Dai, L. Yuan, and Y . Wei, “Towards high performance video object detection,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 7210–7218
2018
-
[85]
Flow-guided feature aggregation for video object detection,
X. Zhu, Y . Wang, J. Dai, L. Yuan, and Y . Wei, “Flow-guided feature aggregation for video object detection,” inInt. Conf. Comput. Vis., 2017, pp. 408–417
2017
-
[86]
Fully motion-aware network for video object detection,
S. Wang, Y . Zhou, J. Yan, and Z. Deng, “Fully motion-aware network for video object detection,” inEur . Conf. Comput. Vis., 2018, pp. 542– 557
2018
-
[87]
Mamba: Multi-level aggregation via memory bank for video object detection,
G. Sun, Y . Hua, G. Hu, and N. Robertson, “Mamba: Multi-level aggregation via memory bank for video object detection,” inAAAI Conf. Artif. Intell., 2021, pp. 2620–2627
2021
-
[88]
Dynamic context-sensitive filtering network for video salient object detection,
M. Zhang, J. Liu, Y . Wang, Y . Piao, S. Yao, W. Ji, J. Li, H. Lu, and Z. Luo, “Dynamic context-sensitive filtering network for video salient object detection,” inInt. Conf. Comput. Vis., 2021, pp. 1553–1563
2021
-
[89]
SAM-PM: Enhancing video camouflaged object detection using spatio-temporal attention,
M. N. Meeran, B. P. Manthaet al., “SAM-PM: Enhancing video camouflaged object detection using spatio-temporal attention,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 1857–1866
2024
-
[90]
Deep spatio-temporal random fields for efficient video segmentation,
S. Chandra, C. Couprie, and I. Kokkinos, “Deep spatio-temporal random fields for efficient video segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 8915–8924
2018
-
[91]
A benchmark dataset and evaluation methodol- ogy for video object segmentation,
F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “A benchmark dataset and evaluation methodol- ogy for video object segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 724–732
2016
-
[92]
The 2017 DA VIS challenge on video object segmen- tation,
J. Pont-Tuset, F. Perazzi, S. Caelles, P. Arbel ´aez, A. Sorkine-Hornung, and L. Van Gool, “The 2017 DA VIS challenge on video object segmen- tation,”arXiv preprint arXiv:1704.00675, 2017
2017 arXiv
-
[93]
Segmentation and recognition using structure from motion point clouds,
G. J. Brostow, J. Shotton, J. Fauqueur, and R. Cipolla, “Segmentation and recognition using structure from motion point clouds,” inEur . Conf. Comput. Vis., 2008, pp. 44–57
2008
-
[94]
The Cityscapes dataset for semantic urban scene understanding,
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Be- nenson, U. Franke, S. Roth, and B. Schiele, “The Cityscapes dataset for semantic urban scene understanding,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 3213–3223
2016
-
[95]
Semantic video segmentation by gated recurrent flow propagation,
D. Nilsson and C. Sminchisescu, “Semantic video segmentation by gated recurrent flow propagation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 6819–6828
2018
-
[96]
Are we ready for autonomous driving? the KITTI vision benchmark suite,
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the KITTI vision benchmark suite,” inIEEE Conf. Comput. Vis. Pattern Recog., 2012, pp. 3354–3361
2012
-
[97]
Every frame counts: Joint learning of video segmentation and optical flow,
M. Ding, Z. Wang, B. Zhou, J. Shi, Z. Lu, and P. Luo, “Every frame counts: Joint learning of video segmentation and optical flow,” inAAAI Conf. Artif. Intell., 2020, pp. 10 713–10 720
2020
-
[98]
Temporally distributed networks for fast video semantic segmenta- tion,
P. Hu, F. C. Heilbron, O. Wang, Z. L. Lin, S. Sclaroff, and F. Perazzi, “Temporally distributed networks for fast video semantic segmenta- tion,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 8815–8824
2020
-
[99]
Indoor segmentation and support inference from RGBD images,
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmentation and support inference from RGBD images,” inEur . Conf. Comput. Vis., 2012, pp. 746–760
2012
-
[100]
Efficient semantic video segmentation with per-frame inference,
Y . Liu, C. Shen, C. Yu, and J. Wang, “Efficient semantic video segmentation with per-frame inference,” inEur . Conf. Comput. Vis., 2020, pp. 352–368
2020
-
[101]
VSPW: A large-scale dataset for video scene parsing in the wild,
J. Miao, Y . Wei, Y . Wu, C. Liang, G. Li, and Y . Yang, “VSPW: A large-scale dataset for video scene parsing in the wild,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 4133–4143
2021
-
[102]
Cross-image relational knowledge distillation for semantic segmentation,
C. Yang, H. Zhou, Z. An, X. Jiang, Y . Xu, and Q. Zhang, “Cross-image relational knowledge distillation for semantic segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 12 309–12 318
2022
-
[103]
The pascal visual object classes challenge: A retrospective,
M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,”Int. J. Comput. Vis., vol. 111, no. 1, pp. 98–136, 2015
2015
-
[104]
Semi-supervised video semantic segmentation with inter-frame feature reconstruction,
J. Zhuang, Z. Wang, and Y . Gao, “Semi-supervised video semantic segmentation with inter-frame feature reconstruction,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 3263–3271
2022
-
[105]
Multispectral video semantic segmentation: A benchmark dataset and baseline,
W. Ji, J. Li, C. Bian, Z. Zhou, J. Zhao, A. L. Yuille, and L. Cheng, “Multispectral video semantic segmentation: A benchmark dataset and baseline,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 1094– 1104
2023
-
[106]
Mask propagation for efficient video semantic segmentation,
Y . Weng, M. Han, H. He, M. Li, L. Yao, X. Chang, and B. Zhuang, “Mask propagation for efficient video semantic segmentation,” inAnnu. Conf. Neur . Inform. Process. Syst., 2023, pp. 7170–7183
2023
-
[108]
Semantic understanding of scenes through the ADE20K dataset,
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ADE20K dataset,”Int. J. Comput. Vis., vol. 127, no. 3, pp. 302–321, 2019
2019
-
[109]
Vanishing- point-guided video semantic segmentation of driving scenes,
D. Guo, D.-P. Fan, T. Lu, C. Sakaridis, and L. V . Gool, “Vanishing- point-guided video semantic segmentation of driving scenes,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 3544–3553
2024
-
[110]
ACDC: The adverse conditions dataset with correspondences for semantic driving scene understand- ing,
C. Sakaridis, D. Dai, and L. Van Gool, “ACDC: The adverse conditions dataset with correspondences for semantic driving scene understand- ing,” inInt. Conf. Comput. Vis., 2021, pp. 10 745–10 755
2021
-
[111]
Infer from what you have seen before: Temporally-dependent classifier for semi-supervised video segmentation,
J. Zhuang, Z. Wang, Y . Zhang, and Z. Fan, “Infer from what you have seen before: Temporally-dependent classifier for semi-supervised video segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 3575–3584
2024
-
[112]
Learning local and global temporal contexts for video semantic segmentation,
G. Sun, Y . Liu, H. Ding, M. Wu, and L. V . Gool, “Learning local and global temporal contexts for video semantic segmentation,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 10, pp. 6919–6934, 2024
2024
-
[113]
Exploiting temporal state space sharing for video semantic segmentation,
S. A. S. Hesham, Y . Liu, G. Sun, H. Ding, J. Yang, E. Konukoglu, X. Geng, and X. Jiang, “Exploiting temporal state space sharing for video semantic segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2025. 18
2025
-
[114]
New generation deep learning for video object detection: A survey,
L. Jiao, R. Zhang, F. Liu, S. Yang, B. Hou, L. Li, and X. Tang, “New generation deep learning for video object detection: A survey,”IEEE Trans. Neur . Net. Learn. Syst., vol. 33, no. 8, pp. 3195–3215, 2021
2021
-
[115]
STEm- Seg: Spatio-temporal embeddings for instance segmentation in videos,
A. Athar, S. Mahadevan, A. Osep, L. Leal-Taix ´e, and B. Leibe, “STEm- Seg: Spatio-temporal embeddings for instance segmentation in videos,” inEur . Conf. Comput. Vis., 2020, pp. 158–177
2020
-
[116]
Classifying, segmenting, and tracking object instances in video with mask propagation,
G. Bertasius and L. Torresani, “Classifying, segmenting, and tracking object instances in video with mask propagation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 9739–9748
2020
-
[117]
Video instance segmentation tracking with a modified V AE architecture,
C.-C. Lin, Y . Hung, R. Feris, and L. He, “Video instance segmentation tracking with a modified V AE architecture,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 13 147–13 157
2020
-
[118]
Learning multi-object tracking and segmentation from automatic an- notations,
L. Porzi, M. Hofinger, I. Ruiz, J. Serrat, S. R. Bul `o, and P. Kontschieder, “Learning multi-object tracking and segmentation from automatic an- notations,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 6845– 6854
2019
-
[119]
BDD100K: A diverse driving dataset for heterogeneous multitask learning,
F. Yu, H. Chen, X. Wang, W. Xian, Y . Chen, F. Liu, V . Madhavan, and T. Darrell, “BDD100K: A diverse driving dataset for heterogeneous multitask learning,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 2636–2645
2020
-
[120]
Learning to track instances without video annotations,
Y . Fu, S. Liu, U. Iqbal, S. D. Mello, H. Shi, and J. Kautz, “Learning to track instances without video annotations,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 8676–8685
2021
-
[121]
CompFeat: Compre- hensive feature aggregation for video instance segmentation,
Y . Fu, L. Yang, D. Liu, T. S. Huang, and H. Shi, “CompFeat: Compre- hensive feature aggregation for video instance segmentation,” inAAAI Conf. Artif. Intell., vol. 35, no. 2, 2021, pp. 1361–1369
2021
-
[122]
Video instance segmen- tation using inter-frame communication transformers,
S. Hwang, M. Heo, S. W. Oh, and S. J. Kim, “Video instance segmen- tation using inter-frame communication transformers,” inAnnu. Conf. Neur . Inform. Process. Syst., 2021, pp. 13 352–13 363
2021
-
[123]
Video instance segmentation with a propose-reduce paradigm,
H. Lin, R. Wu, S. Liu, J. Lu, and J. Jia, “Video instance segmentation with a propose-reduce paradigm,” inInt. Conf. Comput. Vis., 2021, pp. 1739–1748
2021
-
[124]
SG-Net: Spatial granularity network for one-stage video instance segmentation,
D. Liu, Y . Cui, W. Tan, and Y . Chen, “SG-Net: Spatial granularity network for one-stage video instance segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 9816–9825
2021
-
[125]
Weakly supervised instance segmentation for videos with temporal mask consistency,
Q. Liu, V . Ramanathan, D. K. Mahajan, A. L. Yuille, and Z. Yang, “Weakly supervised instance segmentation for videos with temporal mask consistency,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 13 963–13 973
2021
-
[126]
End-to-end video instance segmentation with transformers,
Y . Wang, Z. Xu, X. Wang, C. Shen, B. Cheng, H. Shen, and H. Xia, “End-to-end video instance segmentation with transformers,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 8741–8750
2021
-
[127]
Track to detect and segment: An online multi-object tracker,
J. Wu, J. Cao, L. Song, Y . Wang, M. Yang, and J. Yuan, “Track to detect and segment: An online multi-object tracker,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 12 347–12 356
2021
-
[128]
Crossover learning for fast online video instance segmentation,
S. Yang, Y . Fang, X. Wang, Y . Li, C. Fang, Y . Shan, B. Feng, and W. Liu, “Crossover learning for fast online video instance segmentation,” inInt. Conf. Comput. Vis., 2021, pp. 8043–8052
2021
-
[129]
Occluded video instance segmentation: A benchmark,
J. Qi, Y . Gao, Y . Hu, X. Wang, X. Liu, X. Bai, S. J. Belongie, A. L. Yuille, P. H. S. Torr, and S. Bai, “Occluded video instance segmentation: A benchmark,” inInt. J. Comput. Vis., vol. 130, no. 8. Springer, 2022, pp. 2022–2039
2022
-
[130]
Improving video instance segmentation via temporal pyramid routing,
X. Li, H. He, Y . Yang, H. Ding, K. Yang, G. Cheng, Y . Tong, and D. Tao, “Improving video instance segmentation via temporal pyramid routing,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 5, pp. 6594–6601, 2022
2022
-
[131]
VISOLO: Grid-based space-time aggregation for efficient online video instance segmentation,
S. H. Han, S. Hwang, S. W. Oh, Y . Park, H. Kim, M.-J. Kim, and S. J. Kim, “VISOLO: Grid-based space-time aggregation for efficient online video instance segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 2896–2905
2022
-
[132]
Microsoft coco: Common objects in context,
T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” inEur . Conf. Comput. Vis., 2014, pp. 740–755
2014
-
[133]
MinVIS: A minimal video instance segmentation framework without video-based training,
D.-A. Huang, Z. Yu, and A. Anandkumar, “MinVIS: A minimal video instance segmentation framework without video-based training,” in Annu. Conf. Neur . Inform. Process. Syst., 2022, pp. 31 265–31 277
2022
-
[134]
STC: Spatio-temporal contrastive learning for video instance segmentation,
Z. Jiang, Z. Gu, J. Peng, H. Zhou, L. Liu, Y . Wang, Y . Tai, C. Wang, and L. Zhang, “STC: Spatio-temporal contrastive learning for video instance segmentation,” inEur . Conf. Comput. Vis., 2022, pp. 539–556
2022
-
[135]
Track- Former: Multi-object tracking with transformers,
T. Meinhardt, A. Kirillov, L. Leal-Taixe, and C. Feichtenhofer, “Track- Former: Multi-object tracking with transformers,” inIEEE Conf. Com- put. Vis. Pattern Recog., 2022, pp. 8844–8854
2022
-
[136]
MOT16: A benchmark for multi-object tracking,
A. Milan, L. Leal-Taix ´e, I. Reid, S. Roth, and K. Schindler, “MOT16: A benchmark for multi-object tracking,”arXiv preprint arXiv:1603.00831, 2016
2016 arXiv
-
[137]
SeqFormer: Sequential transformer for video instance segmentation,
J. Wu, Y . Jiang, S. Bai, W. Zhang, and X. Bai, “SeqFormer: Sequential transformer for video instance segmentation,” inEur . Conf. Comput. Vis., 2022, pp. 553–569
2022
-
[138]
In defense of online models for video instance segmentation,
J. Wu, Q. Liu, Y . Jiang, S. Bai, A. Yuille, and X. Bai, “In defense of online models for video instance segmentation,” inEur . Conf. Comput. Vis., 2022, pp. 588–605
2022
-
[139]
Temporally efficient vision transformer for video instance segmentation,
S. Yang, X. Wang, Y . Li, Y . Fang, J. Fang, W. Liu, X. Zhao, and Y . Shan, “Temporally efficient vision transformer for video instance segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 2885–2895
2022
-
[140]
A generalized framework for video instance segmentation,
M. Heo, S. Hwang, J. Hyun, H. Kim, S. W. Oh, J.-Y . Lee, and S. J. Kim, “A generalized framework for video instance segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 14 623–14 632
2023
-
[141]
VideoCut- LER: Surprisingly simple unsupervised video instance segmentation,
X. Wang, I. Misra, Z. Zeng, R. Girdhar, and T. Darrell, “VideoCut- LER: Surprisingly simple unsupervised video instance segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 22 755–22 764
2023
-
[142]
CTVIS: Consistent training for online video instance segmentation,
K. Ying, Q. Zhong, W. Mao, Z. Wang, H. Chen, L. Y . Wu, Y . Liu, C. Fan, Y . Zhuge, and C. Shen, “CTVIS: Consistent training for online video instance segmentation,” inInt. Conf. Comput. Vis., 2023, pp. 899– 908
2023
-
[143]
Dvis: Decoupled video instance segmentation framework,
T. Zhang, X. Tian, Y . Wu, S. Ji, X. Wang, Y . Zhang, and P. Wan, “Dvis: Decoupled video instance segmentation framework,” inInt. Conf. Comput. Vis., 2023, pp. 1282–1291
2023
-
[144]
Unified embedding alignment for open-vocabulary video instance segmentation,
H. Fang, P. Wu, Y . Li, X. Zhang, and X. Lu, “Unified embedding alignment for open-vocabulary video instance segmentation,” inEur . Conf. Comput. Vis., 2024, pp. 225–241
2024
-
[145]
Towards open-vocabulary video instance segmentation,
H. Wang, C. Yan, S. Wang, X. Jiang, X. Tang, Y . Hu, W. Xie, and E. Gavves, “Towards open-vocabulary video instance segmentation,” in Int. Conf. Comput. Vis., 2023, pp. 4057–4066
2023
-
[146]
OV-VIS: Open-vocabulary video instance segmentation,
H. Wang, C. Yan, K. Chen, X. Jiang, X. Tang, Y . Hu, G. Kang, W. Xie, and E. Gavves, “OV-VIS: Open-vocabulary video instance segmentation,”Int. J. Comput. Vis., vol. 132, no. 11, pp. 5048–5065, 2024
2024
-
[147]
LVIS: A dataset for large vocabulary instance segmentation,
A. Gupta, P. Dollar, and R. Girshick, “LVIS: A dataset for large vocabulary instance segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 5356–5364
2019
-
[148]
Video panoptic segmenta- tion,
D. Kim, S. Woo, J.-Y . Lee, and I.-S. Kweon, “Video panoptic segmenta- tion,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 9856–9865
2020
-
[149]
ViP- DeepLab: Learning visual perception with depth-aware video panoptic segmentation,
S. Qiao, Y . Zhu, H. Adam, A. L. Yuille, and L.-C. Chen, “ViP- DeepLab: Learning visual perception with depth-aware video panoptic segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 3996–4007
2020
-
[150]
Learning to associate every segment for video panoptic segmentation,
S. Woo, D. Kim, J.-Y . Lee, and I. S. Kweon, “Learning to associate every segment for video panoptic segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 2705–2714
2021
-
[151]
Waymo open dataset: Panoramic video panoptic segmentation,
J. Mei, A. Z. Zhu, X. Yan, H. Yan, S. Qiao, Y . Zhu, L.-C. Chen, H. Kretzschmar, and D. Anguelov, “Waymo open dataset: Panoramic video panoptic segmentation,” inEur . Conf. Comput. Vis., 2022, pp. 53–72
2022
-
[152]
Video K-Net: A simple, strong, and unified baseline for video segmentation,
X. Li, W. Zhang, J. Pang, K. Chen, G. Cheng, Y . Tong, and C. C. Loy, “Video K-Net: A simple, strong, and unified baseline for video segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 18 847–18 857
2022
-
[153]
Tube- Link: A flexible cross tube framework for universal video segmenta- tion,
X. Li, H. Yuan, W. Zhang, G. Cheng, J. Pang, and C. C. Loy, “Tube- Link: A flexible cross tube framework for universal video segmenta- tion,” inInt. Conf. Comput. Vis., 2023, pp. 13 923–13 933
2023
-
[154]
PolyphonicFormer: Unified query learning for depth-aware video panoptic segmentation,
H. Yuan, X. Li, Y . Yang, G. Cheng, J. Zhang, Y . Tong, L. Zhang, and D. Tao, “PolyphonicFormer: Unified query learning for depth-aware video panoptic segmentation,” inEur . Conf. Comput. Vis., 2022, pp. 582–599
2022
-
[155]
Open- vocabulary panoptic segmentation with text-to-image diffusion models,
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello, “Open- vocabulary panoptic segmentation with text-to-image diffusion models,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 2955–2966
2023
-
[156]
COCO-Stuff: Thing and stuff classes in context,
H. Caesar, J. Uijlings, and V . Ferrari, “COCO-Stuff: Thing and stuff classes in context,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 1209–1218
2018
-
[157]
Segment as points for efficient online multi-object tracking and segmentation,
Z. Xu, W. Zhang, X. Tan, W. Yang, H. Huang, S. Wen, E. Ding, and L. Huang, “Segment as points for efficient online multi-object tracking and segmentation,” inEur . Conf. Comput. Vis., 2020, pp. 264–281
2020
-
[158]
Assignment-space- based multi-object tracking and segmentation,
A. Choudhuri, G. Chowdhary, and A. G. Schwing, “Assignment-space- based multi-object tracking and segmentation,” inInt. Conf. Comput. Vis., 2021, pp. 13 598–13 607
2021
-
[159]
Multi-object tracking and segmentation via neural message passing,
G. Bras ´o, O. Cetintas, and L. Leal-Taix ´e, “Multi-object tracking and segmentation via neural message passing,”Int. J. Comput. Vis., vol. 130, no. 12, pp. 3035–3053, 2022
2022
-
[160]
MOT20: A bench- 19 mark for multi object tracking in crowded scenes,
P. Dendorfer, H. Rezatofighi, A. Milan, J. Shi, D. Cremers, I. Reid, S. Roth, K. Schindler, and L. Leal-Taix ´e, “MOT20: A bench- 19 mark for multi object tracking in crowded scenes,”arXiv preprint arXiv:2003.09003, 2020
2003 arXiv
-
[161]
Human in events: A large-scale benchmark for human-centric video analysis in complex events,
W. Lin, H. Liu, S. Liu, Y . Li, R. Qian, T. Wang, N. Xu, H. Xiong, G.-J. Qi, and N. Sebe, “Human in events: A large-scale benchmark for human-centric video analysis in complex events,”arXiv preprint arXiv:2005.04490, 2020
2005 arXiv
-
[162]
Integrating boxes and masks: A multi- object framework for unified visual tracking and segmentation,
Y . Xu, Z. Yang, and Y . Yang, “Integrating boxes and masks: A multi- object framework for unified visual tracking and segmentation,” inInt. Conf. Comput. Vis., 2023, pp. 9738–9751
2023
-
[163]
YouTube-VOS: A large-scale video object segmentation benchmark,
N. Xu, L. Yang, Y . Fan, D. Yue, Y . Liang, J. Yang, and T. Huang, “YouTube-VOS: A large-scale video object segmentation benchmark,” arXiv preprint arXiv:1809.03327, 2018
2018 arXiv
-
[164]
Segment and track anything,
Y . Cheng, L. Li, Y . Xu, X. Li, Z. Yang, W. Wang, and Y . Yang, “Segment and track anything,”arXiv preprint arXiv:2305.06558, 2023
2023 arXiv
-
[165]
Seg- ment anything meets point tracking,
F. Raji ˇc, L. Ke, Y .-W. Tai, C.-K. Tang, M. Danelljan, and F. Yu, “Seg- ment anything meets point tracking,”arXiv preprint arXiv:2307.01197, 2023
2023 arXiv
-
[166]
SAM 2: Segment anything in images and videos,
N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R ¨adle, C. Rolland, L. Gustafsonet al., “SAM 2: Segment anything in images and videos,”arXiv preprint arXiv:2408.00714, 2024
2024 arXiv
-
[167]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Loet al., “Segment anything,” inInt. Conf. Comput. Vis., 2023, pp. 4015–4026
2023
-
[168]
Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection,
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Suet al., “Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection,” inEur . Conf. Comput. Vis., 2024, pp. 38–55
2024
-
[169]
CLIP-VIS: Adapting CLIP for open-vocabulary video instance segmentation,
W. Zhu, J. Cao, J. Xie, S. Yang, and Y . Pang, “CLIP-VIS: Adapting CLIP for open-vocabulary video instance segmentation,”IEEE Trans. Circ. Syst. Video Technol., vol. 35, no. 2, pp. 1098–1110, 2025
2025
-
[170]
Towards open- vocabulary video semantic segmentation,
X. Li, Y . Liu, G. Sun, M. Wu, L. Zhang, and C. Zhu, “Towards open- vocabulary video semantic segmentation,”IEEE Trans. Multimedia, 2025
2025
-
[171]
STEP: Segmenting and tracking every pixel,
M. Weber, J. Xie, M. D. Collins, Y . Zhu, P. V oigtlaender, H. Adam, B. Green, A. Geiger, B. Leibe, D. Cremers, A. Osep, L. Leal-Taix´e, and L.-C. Chen, “STEP: Segmenting and tracking every pixel,” inNeurIPS Datasets and Benchmarks, 2021
2021
-
[172]
Tubeformer-DeepLab: Video mask transformer,
D. Kim, J. Xie, H. Wang, S. Qiao, Q. Yu, H.-S. Kim, H. Adam, I. S. Kweon, and L.-C. Chen, “Tubeformer-DeepLab: Video mask transformer,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 13 914–13 924
2022
-
[173]
Slot-VPS: Object-centric representation learning for video panoptic segmentation,
Y . Zhou, H. Zhang, H. Lee, S. Sun, P. Li, Y . Zhu, B. Yoo, X. Qi, and J.-J. Han, “Slot-VPS: Object-centric representation learning for video panoptic segmentation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 3093–3103
2022
-
[174]
Combined image- and world-space tracking in traffic scenes,
A. Osep, W. Mehner, M. Mathias, and B. Leibe, “Combined image- and world-space tracking in traffic scenes,” inIEEE Int. Conf. Robot. Autom., 2017, pp. 1988–1995
2017
-
[175]
Track to reconstruct and reconstruct to track,
J. Luiten, T. Fischer, and B. Leibe, “Track to reconstruct and reconstruct to track,”IEEE Trans. Robot. Autom. Let., vol. 5, no. 2, pp. 1803–1810, 2020
2020
-
[176]
Track, then decide: Category-agnostic vision-based multi-object tracking,
A. O ˇsep, W. Mehner, P. V oigtlaender, and B. Leibe, “Track, then decide: Category-agnostic vision-based multi-object tracking,” inIEEE Int. Conf. Robot. Autom., 2018, pp. 3494–3501
2018
-
[177]
Unidentified video objects: A benchmark for dense, open-world segmentation,
W. Wang, M. Feiszli, H. Wang, and D. Tran, “Unidentified video objects: A benchmark for dense, open-world segmentation,” inInt. Conf. Comput. Vis., 2021, pp. 10 776–10 785
2021
-
[178]
Towards open vocabulary learning: A survey,
J. Wu, X. Li, S. Xu, H. Yuan, H. Ding, Y . Yang, X. Li, J. Zhang, Y . Tong, X. Jianget al., “Towards open vocabulary learning: A survey,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 7, pp. 5092–5113, 2024
2024
-
[179]
VideoSAM: Open-world video segmentation,
P. Guo, Z. Zhao, J. Gao, C. Wu, T. He, Z. Zhang, T. Xiao, and W. Zhang, “VideoSAM: Open-world video segmentation,”arXiv preprint arXiv:2410.08781, 2024
2024 arXiv
-
[180]
Video instance segmentation in an open-world,
O. Thawakar, S. Narayan, H. Cholakkal, R. M. Anwer, S. Khan, J. Laaksonen, M. Shah, and F. S. Khan, “Video instance segmentation in an open-world,”Int. J. Comput. Vis., vol. 133, no. 1, pp. 398–409, 2025
2025
-
[181]
Learning object state changes in videos: An open-world perspective,
Z. Xue, K. Ashutosh, and K. Grauman, “Learning object state changes in videos: An open-world perspective,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 18 493–18 503
2024
-
[182]
Towards training-free open- world segmentation via image prompt foundation models,
L. Tang, P.-T. Jiang, H. Xiao, and B. Li, “Towards training-free open- world segmentation via image prompt foundation models,”Int. J. Comput. Vis., vol. 133, no. 1, pp. 1–15, 2025
2025
-
[183]
Open-world instance segmentation: Top-down learning with bottom-up supervision,
T. Kalluri, W. Wang, H. Wang, M. Chandraker, L. Torresani, and D. Tran, “Open-world instance segmentation: Top-down learning with bottom-up supervision,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 2693–2703
2024
-
[184]
BURST: A benchmark for unifying object recogni- tion, segmentation and tracking in video,
A. Athar, J. Luiten, P. V oigtlaender, T. Khurana, A. Dave, B. Leibe, and D. Ramanan, “BURST: A benchmark for unifying object recogni- tion, segmentation and tracking in video,” inIEEE Winter Conf. App. Comput. Vis., 2023, pp. 1674–1683
2023
-
[185]
OMG-Seg: Is one model good enough for all segmenta- tion?
X. Li, H. Yuan, W. Li, H. Ding, S. Wu, W. Zhang, Y . Li, K. Chen, and C. C. Loy, “OMG-Seg: Is one model good enough for all segmenta- tion?” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 27 948– 27 959
2024
-
[186]
General and task- oriented video segmentation,
M. Chen, L. Li, W. Wang, R. Quan, and Y . Yang, “General and task- oriented video segmentation,” inEur . Conf. Comput. Vis., 2024, pp. 72–92
2024
-
[187]
OMG-LLaV A: Bridging image-level, object-level, pixel-level reasoning and understanding,
T. Zhang, X. Li, H. Fei, H. Yuan, S. Wu, S. Ji, C. C. Loy, and S. Yan, “OMG-LLaV A: Bridging image-level, object-level, pixel-level reasoning and understanding,” inAnnu. Conf. Neur . Inform. Process. Syst., 2024, pp. 71 737–71 767
2024
-
[188]
BA-SAM: Scalable bias-mode attention mask for segment anything model,
Y . Song, Q. Zhou, X. Li, D.-P. Fan, X. Lu, and L. Ma, “BA-SAM: Scalable bias-mode attention mask for segment anything model,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 3162–3173
2024
-
[189]
Instruction-guided multi-granularity segmentation and captioning with large multimodal model,
X. Yuan, L. Zhou, Z. Sun, Z. Zhou, and J. Lan, “Instruction-guided multi-granularity segmentation and captioning with large multimodal model,” inAAAI Conf. Artif. Intell., 2025, pp. 9725–9733
2025
-
[190]
DVIS++: Improved decoupled framework for universal video segmentation,
T. Zhang, X. Tian, Y . Zhou, S. Ji, X. Wang, X. Tao, Y . Zhang, P. Wan, Z. Wang, and Y . Wu, “DVIS++: Improved decoupled framework for universal video segmentation,”IEEE Trans. Pattern Anal. Mach. Intell., 2025
2025
-
[191]
Referred by multi-modality: A unified temporal transformer for video object segmentation,
S. Yan, R. Zhang, Z. Guo, W. Chen, W. Zhang, H. Li, Y . Qiao, H. Dong, Z. He, and P. Gao, “Referred by multi-modality: A unified temporal transformer for video object segmentation,” inAAAI Conf. Artif. Intell., 2024, pp. 6449–6457
2024
-
[192]
Connecting vision and language with video localized narratives,
P. V oigtlaender, S. Changpinyo, J. Pont-Tuset, R. Soricut, and V . Ferrari, “Connecting vision and language with video localized narratives,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 2461–2471
2023
-
[193]
VisionLLM: Large language model is also an open-ended decoder for vision-centric tasks,
W. Wang, Z. Chen, X. Chen, J. Wu, X. Zhu, G. Zeng, P. Luo, T. Lu, J. Zhou, Y . Qiaoet al., “VisionLLM: Large language model is also an open-ended decoder for vision-centric tasks,” inAnnu. Conf. Neur . Inform. Process. Syst., 2023, pp. 61 501–61 513
2023
-
[194]
Monkey: Image resolution and text label are important things for large multi-modal models,
Z. Li, B. Yang, Q. Liu, Z. Ma, S. Zhang, J. Yang, Y . Sun, Y . Liu, and X. Bai, “Monkey: Image resolution and text label are important things for large multi-modal models,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 26 763–26 773
2024
-
[195]
Unified-IO 2: Scaling autoregressive multimodal models with vision language audio and action,
J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi, “Unified-IO 2: Scaling autoregressive multimodal models with vision language audio and action,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 26 439–26 455
2024
-
[196]
Generalized decoding for pixel, image, and language,
X. Zou, Z.-Y . Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, H. Behl, J. Wang, L. Yuanet al., “Generalized decoding for pixel, image, and language,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 15 116–15 127
2023
-
[197]
Meta-Transformer: A unified framework for multimodal learning,
Y . Zhang, K. Gong, K. Zhang, H. Li, Y . Qiao, W. Ouyang, and X. Yue, “Meta-Transformer: A unified framework for multimodal learning,” arXiv preprint arXiv:2307.10802, 2023
2023 arXiv
-
[198]
What makes multimodal in-context learning work?
F. B. Baldassini, M. Shukor, M. Cord, L. Soulier, and B. Piwowarski, “What makes multimodal in-context learning work?” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 1539–1550
2024
-
[199]
CLiMB: A continual learning benchmark for vision- and-language tasks,
T. Srinivasan, T.-Y . Chang, L. Pinto Alva, G. Chochlakis, M. Rostami, and J. Thomason, “CLiMB: A continual learning benchmark for vision- and-language tasks,” inAnnu. Conf. Neur . Inform. Process. Syst., 2022, pp. 29 440–29 453
2022
-
[200]
ViLLa: Video reasoning segmentation with large language model,
R. Zheng, L. Qi, X. Chen, Y . Wang, K. Wang, Y . Qiao, and H. Zhao, “ViLLa: Video reasoning segmentation with large language model,” arXiv preprint arXiv:2407.14500, 2024
2024 arXiv
-
[201]
VISA: Reasoning video object segmentation via large language models,
C. Yan, H. Wang, S. Yan, X. Jiang, Y . Hu, G. Kang, W. Xie, and E. Gavves, “VISA: Reasoning video object segmentation via large language models,” inEur . Conf. Comput. Vis., 2024, pp. 98–115
2024
-
[202]
One token to seg them all: Language instructed reasoning segmentation in videos,
Z. Bai, T. He, H. Mei, P. Wang, Z. Gao, J. Chen, Z. Zhang, and M. Z. Shou, “One token to seg them all: Language instructed reasoning segmentation in videos,” inAnnu. Conf. Neur . Inform. Process. Syst., 2024, pp. 6833–6859
2024
-
[203]
The devil is in temporal token: High quality video reasoning segmentation,
S. Gong, Y . Zhuge, L. Zhang, Z. Yang, P. Zhang, and H. Lu, “The devil is in temporal token: High quality video reasoning segmentation,”arXiv preprint arXiv:2501.08549, 2025
2025 arXiv
-
[204]
A generalist framework for panoptic segmentation of images and videos,
T. Chen, L. Li, S. Saxena, G. Hinton, and D. J. Fleet, “A generalist framework for panoptic segmentation of images and videos,” inInt. Conf. Comput. Vis., 2023, pp. 909–919. 20
2023
-
[205]
DiffusionInst: Diffusion model for instance segmentation,
Z. Gu, H. Chen, and Z. Xu, “DiffusionInst: Diffusion model for instance segmentation,” inIEEE Int. Conf. Acoust. Speech SP, 2024, pp. 2730– 2734
2024
-
[206]
UniGS: Unified representation for image generation and segmenta- tion,
L. Qi, L. Yang, W. Guo, Y . Xu, B. Du, V . Jampani, and M.-H. Yang, “UniGS: Unified representation for image generation and segmenta- tion,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 6305–6315
2024
-
[207]
MomentDiff: Generative video moment retrieval from random to real,
P. Li, C.-W. Xie, H. Xie, L. Zhao, L. Zhang, Y . Zheng, D. Zhao, and Y . Zhang, “MomentDiff: Generative video moment retrieval from random to real,” inAnnu. Conf. Neur . Inform. Process. Syst., 2023, pp. 65 948–65 966
2023
-
[208]
TSM: Temporal shift module for efficient video understanding,
J. Lin, C. Gan, and S. Han, “TSM: Temporal shift module for efficient video understanding,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 7083–7093
2019
-
[209]
MobileViT: Light-weight, general- purpose, and mobile-friendly vision transformer,
S. Mehta and M. Rastegari, “MobileViT: Light-weight, general- purpose, and mobile-friendly vision transformer,”arXiv preprint arXiv:2110.02178, 2021
2021 arXiv
-
[210]
MeMViT: Memory-augmented multiscale vision transformer for efficient long-term video recognition,
C.-Y . Wu, Y . Li, K. Mangalam, H. Fan, B. Xiong, J. Malik, and C. Feichtenhofer, “MeMViT: Memory-augmented multiscale vision transformer for efficient long-term video recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 13 587–13 597
2022
-
[211]
Faster segment anything: Towards lightweight SAM for mobile applications,
C. Zhang, D. Han, Y . Qiao, J. U. Kim, S.-H. Bae, S. Lee, and C. S. Hong, “Faster segment anything: Towards lightweight SAM for mobile applications,”arXiv preprint arXiv:2306.14289, 2023
2023 arXiv
-
[212]
PIDNet: A real-time semantic segmentation network inspired by PID controllers,
J. Xu, Z. Xiong, and S. P. Bhattacharyya, “PIDNet: A real-time semantic segmentation network inspired by PID controllers,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 19 529–19 539
2023
-
[213]
Rethinking dilated convolution for real-time semantic segmen- tation,
R. Gao, “Rethinking dilated convolution for real-time semantic segmen- tation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 4675– 4684
2023
-
[214]
LLM-Seg: Bridging image segmentation and large language model reasoning,
J. Wang and L. Ke, “LLM-Seg: Bridging image segmentation and large language model reasoning,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 1765–1774
2024
-
[215]
Visual large language models for generalized and specialized applications,
Y . Li, Z. Lai, W. Bao, Z. Tan, A. Dao, K. Sui, J. Shen, D. Liu, H. Liu, and Y . Kong, “Visual large language models for generalized and specialized applications,”arXiv preprint arXiv:2501.02765, 2025
2025 arXiv
-
[216]
LLMFormer: Large language model for open-vocabulary semantic segmentation,
H. Shi, S. D. Dao, and J. Cai, “LLMFormer: Large language model for open-vocabulary semantic segmentation,”Int. J. Comput. Vis., vol. 133, no. 2, pp. 742–759, 2025
2025
-
[217]
LISA: Reasoning segmentation via large language model,
X. Lai, Z. Tian, Y . Chen, Y . Li, Y . Yuan, S. Liu, and J. Jia, “LISA: Reasoning segmentation via large language model,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 9579–9589
2024
-
[218]
GSV A: Generalized segmentation via multimodal large language models,
Z. Xia, D. Han, Y . Han, X. Pan, S. Song, and G. Huang, “GSV A: Generalized segmentation via multimodal large language models,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 3858–3869
2024
-
[219]
InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Luet al., “InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 24 185–24 198
2024
-
[220]
Groma: Localized visual tokenization for grounding multimodal large language models,
C. Ma, Y . Jiang, J. Wu, Z. Yuan, and X. Qi, “Groma: Localized visual tokenization for grounding multimodal large language models,” inEur . Conf. Comput. Vis., 2024, pp. 417–435
2024
-
[221]
ShareGPT4Video: Improving video understanding and generation with better captions,
L. Chen, X. Wei, J. Li, X. Dong, P. Zhang, Y . Zang, Z. Chen, H. Duan, Z. Tang, L. Yuanet al., “ShareGPT4Video: Improving video understanding and generation with better captions,” inAnnu. Conf. Neur . Inform. Process. Syst., 2024, pp. 19 472–19 495. Guohuan Xieis currently pur...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.