REVIEW 4 major objections 5 minor 72 references
MOVE: Motion-Guided Few-Shot Video Object Segmentation
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper introduces MOVE, a dataset and task in which a few support videos of a motion specify what to segment, and reports that its decoupled motion-appearance network outperforms all six compared methods on it.
desk verdict The MOVE dataset is a genuine new resource for motion-guided few-shot video segmentation, but the paper's central claim that it isolates motion from object category is not yet backed by dataset statistics. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Decoupled Motion-Appearance module (DMA), which converts a support video and its mask into two separate prototype sequences. The appearance prototype is produced by mask-pooling per-frame features, so it encodes what the object looks like, while the motion prototype is produced by temporally differencing adjacent-frame features and enhancing them with 3D convolutions, so it encodes how the object moves. Two auxiliary classification heads, one for object categories and one for motion categories, keep the prototypes separated, and a transformer with cross-attention refines them before the mask decoder. This decoupling is the mechanism that lets a query object be matched to a support motion even when the two objects belong to different categories.
What would settle it
Compute the per-motion distribution of object categories in MOVE; if most motion classes involve a single object category, and a baseline that simply segments that category reaches a J&F comparable to DMA's, the benchmark would be measuring category matching rather than motion understanding.
Extended reading notes
Core claim
The central claim is that temporal motion, not object category, can serve as the supervised signal for few-shot video object segmentation, and that a benchmark built on this idea exposes a real weakness in existing methods. MOVE changes the support set from static images to video clips with mask sequences, so the model must infer a motion prototype across frames, and the query set includes objects performing the support motion among distractors and empty frames. The paper's DMA network computes an appearance prototype by mask-pooling frame features and a motion prototype by differencing adjacent-frame features, keeps the two prototypes decoupled through separate auxiliary classification heads, and fuses them with transformer attention before decoding masks. On MOVE, DMA is reported to outperform all baselines in both 2-way-1-shot and 5-way-1-shot settings, on overlapping and non-overlapping motion splits, and with ResNet50 and VideoSwin-T backbones. The paper also shows that adding frame-level temporal modeling to a strong category-based baseline raises its J&F from 44.4% to 46.3%, supporting the claim that motion is the active ingredient.
Load-bearing premise
The benchmark assumes that in MOVE the same motion is performed by many different kinds of objects, so a model cannot solve the task by recognizing the dominant object category, such as always segmenting the human, instead of understanding the motion.
Editorial extensions
If this is right
- Existing few-shot video segmentation methods will need explicit temporal modeling to perform well on MOVE; the paper reports that adding a simple self-attention step across frames raises a category-based baseline from 44.4% to 46.3% J&F.
- Motion prototypes alone outperform appearance prototypes alone on MOVE (43.8% versus 36.5% J&F), so dynamic cues appear to carry most of the discriminative signal in this task.
- Because DMA beats all six baselines under both backbones and both data splits, it can serve as a stable reference baseline for future work on MOVE.
- The non-overlapping split is harder than the overlapping split for every method, meaning generalization to motion families that share no parent class with training is a distinct open challenge.
- Low N-Acc values across all methods show that rejecting empty query frames and suppressing false positives is a common weakness, pointing to background modeling as a needed direction.
Reading between the lines
- A check the paper leaves implicit is to report the conditional distribution of object categories within each motion class; if most motions are performed by one object type, a model could solve the task by category matching alone.
- The same episode construction could extend to other dense prediction settings, such as few-shot video instance segmentation or motion-guided video object detection, where the support videos would define a motion rather than a class.
- The paper lists decomposing motions into primitive units as future work; if supported, that direction would test whether motion prototypes learned on known motions transfer to unseen motion combinations.
- The matching-score head could be evaluated as a stand-alone few-shot action retrieval signal, since the paper motivates motion-based retrieval but reports no retrieval experiment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MOVE, a large-scale benchmark for motion-guided few-shot video object segmentation (FSVOS), where support and query videos are matched by motion pattern rather than object category, and introduces DMA (Decoupled Motion-Appearance Network), a baseline that extracts separate appearance and motion prototypes with auxiliary classification supervision, prototype attention, and a mask decoder. The authors evaluate six prior methods from FSVOS, few-shot image segmentation, and referring video object segmentation on MOVE under overlapping and non-overlapping splits, in 2-way-1-shot and 5-way-1-shot settings with ResNet50 and VideoSwin-T backbones, and report that DMA outperforms all baselines on J&F and T-Acc, though all methods achieve very low N-Acc. The paper also reports a necessity study, ablations of the motion extractor and prototype decoupling, oracle experiments, and qualitative examples.
Significance. If the benchmark's motion-isolation claim holds, MOVE would be a valuable new resource: it is larger and more motion-focused than existing FSVOS datasets, uses video-level support sets, and provides a concrete task definition that could catalyze research on motion-centric few-shot segmentation. The DMA method is a reasonable first baseline, and the oracle experiments give useful upper bounds. The internal ablations are consistent and the benchmarking effort is broad. However, the central claim that MOVE isolates motion as the discriminative signal is currently under-supported because the dataset statistics do not rule out strong correlations between motion categories and object categories, and the baseline adaptations are not described in sufficient detail to allow reproduction or to rule out benchmarking bias. The low N-Acc values also temper the claim of consistent superiority across all metrics.
major comments (4)
- [§3.3 and Table 2]
- [§5.2, Tables 3-4]
- [§5, Implementation Details and baseline adaptation]
- [§5.1, Table 2, HPAN*]
minor comments (5)
- [§5.2, paragraph 2]
- [§4.3, Eq. (3)]
- [§5.3, Table 7]
- [§3.3 and NS split description]
- [§5.4, Figure 6]
Circularity Check
No circularity: the MOVE dataset and DMA method are independently derived, and the unverified motion/category decorrelation is a validity concern, not a circular reduction.
full rationale
The paper's derivation chain contains no step in which an output is equivalent to an input by construction. The MOVE task is defined by a test protocol in Section 3.1 and a newly constructed dataset, while the DMA method in Section 4 is an independent network whose ablations in Tables 5-7 vary concrete components; none of the reported quantities is a fitted parameter renamed as a prediction. The benchmark comparisons in Tables 3-4 use externally published methods (DANet, HPAN, TTI, LMPM, SCCAN, CyCTR) with stated venues, and the gains of DMA over these baselines are empirical measurements rather than consequences of the task definition. Self-citations, such as the MeViS LMPM baseline [11] and the GRES metrics [35], are used as tooling or as one baseline among several, and no load-bearing uniqueness claim is imported from the authors' prior work. The skeptic's concern that MOVE may not decorrelate object categories from motion categories is a dataset-validity limitation, because the paper does not report P(object|motion); even if true, it would weaken the interpretation of the benchmark, but it would not make the DMA-versus-baseline results true by definition. Accordingly, no circular step is identified.
Assumptions & free parameters
free parameters (3)
- DMA transformer depth (number of layers)
- Auxiliary classification loss weights
- Training episode count =
240,000 main; 150,000 ablations
assumptions (4)
- domain assumption Adjacent-frame feature differencing captures the motion of the target object (Eq. 3).
- domain assumption The MOVE vocabulary categories are mutually exclusive and semantically unambiguous (Sec 3.2).
- domain assumption Support masks are accurate and correspond to the object performing the target motion (Sec 3.1).
- domain assumption Pre-training on ImageNet and Kinetics-400 transfers to motion-guided segmentation (Sec 5 Implementation Details).
Cite this review
Pith. "Pith review of MOVE: Motion-Guided Few-Shot Video Object Segmentation." pith.science (2026). https://pith.science/paper/VYEQSFDE
@misc{pith2026250722061,
author = {Pith},
title = {Pith review of: MOVE: Motion-Guided Few-Shot Video Object Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/VYEQSFDE}},
note = {Machine review of arXiv:2507.22061}
}
read the original abstract
This work addresses motion-guided few-shot video object segmentation (FSVOS), which aims to segment dynamic objects in videos based on a few annotated examples with the same motion patterns. Existing FSVOS datasets and methods typically focus on object categories, which are static attributes that ignore the rich temporal dynamics in videos, limiting their application in scenarios requiring motion understanding. To fill this gap, we introduce MOVE, a large-scale dataset specifically designed for motion-guided FSVOS. Based on MOVE, we comprehensively evaluate 6 state-of-the-art methods from 3 different related tasks across 2 experimental settings. Our results reveal that current methods struggle to address motion-guided FSVOS, prompting us to analyze the associated challenges and propose a baseline method, Decoupled Motion Appearance Network (DMA). Experiments demonstrate that our approach achieves superior performance in few shot motion understanding, establishing a solid foundation for future research in this direction.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Learning what to learn for video object segmentation
Goutam Bhat, Felix J ¨aremo Lawin, Martin Danelljan, Andreas Robinson, Michael Felsberg, Luc Van Gool, and Radu Timofte. Learning what to learn for video object segmentation. In Eur. Conf. Comput. Vis., 2020. 2
work page 2020
-
[2]
One- shot video object segmentation
Sergi Caelles, Kevis-Kokitsi Maninis, Jordi Pont-Tuset, Laura Leal-Taix´e, Daniel Cremers, and Luc Van Gool. One- shot video object segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2017. 2
work page 2017
-
[3]
Delving Deep Into Many-to-Many Attention for Few-Shot Video Object Segmentation
Haoxin Chen, Hanjie Wu, Nanxuan Zhao, Sucheng Ren, and Shengfeng He. Delving Deep Into Many-to-Many Attention for Few-Shot Video Object Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2021. 1, 2, 3, 5, 6, 7, 8
work page 2021
-
[4]
Mammalnet: A large-scale video benchmark for mammal recognition and behavior understanding
Jun Chen, Ming Hu, Darren J Coker, Michael L Berumen, Blair Costelloe, Sara Beery, Anna Rohrbach, and Mohamed Elhoseiny. Mammalnet: A large-scale video benchmark for mammal recognition and behavior understanding. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 3
work page 2023
-
[5]
Efficient video action detection with token dropout and context refinement
Lei Chen, Zhan Tong, Yibing Song, Gangshan Wu, and Limin Wang. Efficient video action detection with token dropout and context refinement. In Int. Conf. Comput. Vis.,
-
[6]
Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model
Ho Kei Cheng and Alexander G Schwing. Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model. In Eur. Conf. Comput. Vis., 2022. 2
work page 2022
-
[7]
Ho Kei Cheng, Yu-Wing Tai, and Chi-Keung Tang. Rethinking space-time networks with improved memory coverage for efficient video object segmentation.Adv. Neural Inform. Process. Syst., 2021. 2
work page 2021
-
[8]
Putting the object back into video object segmentation
Ho Kei Cheng, Seoung Wug Oh, Brian Price, Joon-Young Lee, and Alexander Schwing. Putting the object back into video object segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 2
work page 2024
Show all 72 references
-
[9]
Haa500: Human-centric atomic action dataset with curated videos
Jihoon Chung, Cheng-hsin Wuu, Hsuan-ru Yang, Yu-Wing Tai, and Chi-Keung Tang. Haa500: Human-centric atomic action dataset with curated videos. In Int. Conf. Comput. Vis., 2021. 3
2021
-
[10]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In IEEE Conf. Comput. Vis. Pattern Recog., 2009. 5
2009
-
[11]
MeViS: A large-scale benchmark for video segmentation with motion expressions
Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, and Chen Change Loy. MeViS: A large-scale benchmark for video segmentation with motion expressions. In Int. Conf. Comput. Vis., 2023. 2, 3, 5, 6
2023
-
[12]
MOSE: A new dataset for video object segmentation in complex scenes
Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, Philip HS Torr, and Song Bai. MOSE: A new dataset for video object segmentation in complex scenes. In Int. Conf. Comput. Vis., 2023. 2, 5
2023
-
[13]
Multimodal referring segmentation: A survey
Henghui Ding, Song Tang, Shuting He, Chang Liu, Zuxuan Wu, and Yu-Gang Jiang. Multimodal referring segmentation: A survey. arXiv, 2025. 2
2025
-
[14]
MOSEv2: A more challenging dataset for video object segmentation in complex scenes
Henghui Ding, Kaining Ying, Chang Liu, Shuting He, Yu- Gang Jiang, Philip HS Torr, and Song Bai. MOSEv2: A more challenging dataset for video object segmentation in complex scenes. arXiv, 2025. 2
2025
-
[15]
Wlasl (world level american sign language) video, 2022
Hongdong Li Dongxu Li. Wlasl (world level american sign language) video, 2022. 3
2022
-
[16]
Few-Shot Video Object Detection
Fan, Qi, Tang, Chi-Keung, Tai, and Yu-Wing. Few-Shot Video Object Detection. In Eur. Conf. Comput. Vis., 2022. 3
2022
-
[17]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 4, 5, 6
2016
-
[18]
Decoupling static and hier- archical motion perception for referring video segmentation
Shuting He and Henghui Ding. Decoupling static and hier- archical motion perception for referring video segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 2, 3
2024
-
[19]
Asap: Aligning simulation and real-world physics for learning agile humanoid whole-body skills
Tairan He, Jiawei Gao, Wenli Xiao, Yuanhang Zhang, Zi Wang, Jiashun Wang, Zhengyi Luo, Guanqi He, Nikhil Sobanbab, Chaoyi Pan, et al. Asap: Aligning simulation and real-world physics for learning agile humanoid whole-body skills. arXiv preprint arXiv:2502.01143, 2025. 1
2025 arXiv
-
[20]
Lvos: A benchmark for long-term video object segmentation
Lingyi Hong, Wenchao Chen, Zhongying Liu, Wei Zhang, Pinxue Guo, Zhaoyu Chen, and Wenqiang Zhang. Lvos: A benchmark for long-term video object segmentation. In Int. Conf. Comput. Vis., 2023. 2
2023
-
[21]
Lvos: A benchmark for large- scale long-term video object segmentation
Lingyi Hong, Zhongying Liu, Wenchao Chen, Chenzhi Tan, Yuang Feng, Xinyu Zhou, Pinxue Guo, Jinglun Li, Zhaoyu Chen, Shuyong Gao, et al. Lvos: A benchmark for large- scale long-term video object segmentation. arXiv preprint arXiv:2404.19326, 2024. 2
2024 arXiv
-
[22]
Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models
Wenyi Hong, Yean Cheng, Zhuoyi Yang, Weihan Wang, Lefan Wang, Xiaotao Gu, Shiyu Huang, Yuxiao Dong, and Jie Tang. Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models. In IEEE Conf. Comput. Vis. Pattern Recog., 2025. 3
2025
-
[23]
Visual recognition of chinese traffic police gestures based on spatial context and temporal features
HE Jian and W ANG Weidong. Visual recognition of chinese traffic police gestures based on spatial context and temporal features. Acta Electronica Sinica, 2020. 3
2020
-
[24]
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017. 3, 5
2017 arXiv
-
[25]
Referring video object segmentation via language-aligned track selection
Seongchan Kim, Woojeong Jin, Sangbeom Lim, Heeji Yoon, Hyunwook Choi, and Seungryong Kim. Referring video object segmentation via language-aligned track selection. arXiv preprint arXiv:2412.01136, 2024. 3
2024 arXiv
-
[26]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Int. Conf. Comput. Vis., 2023. 2
2023
-
[27]
Learning human interaction by interactive phrases
Yu Kong, Yunde Jia, and Yun Fu. Learning human interaction by interactive phrases. InEur. Conf. Comput. Vis.,
-
[28]
You only watch once: A unified cnn architecture for real- time spatiotemporal action localization
Okan K ¨op¨ukl¨u, Xiangyu Wei, and Gerhard Rigoll. You only watch once: A unified cnn architecture for real- time spatiotemporal action localization. arXiv preprint arXiv:1911.06644, 2019. 3
1911 arXiv
-
[29]
Hmdb: a large video 9 database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Est ´ıbaliz Garrote, Tomaso Poggio, and Thomas Serre. Hmdb: a large video 9 database for human motion recognition. In Int. Conf. Comput. Vis., 2011. 3
2011
-
[30]
Motion expressions guided video segmentation via effective motion information mining
Ge Li, Hanqing Sun, Aiping Yang, Jiale Cao, and Yanwei Pang. Motion expressions guided video segmentation via effective motion information mining. IEEE Trans. Emerg. Topics Comput. Intell., 2025. 2, 3
2025
-
[31]
Ai choreographer: Music conditioned 3d dance generation with aist++
Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa. Ai choreographer: Music conditioned 3d dance generation with aist++. In Int. Conf. Comput. Vis., 2021. 3
2021
-
[32]
Multisports: A multi-person video dataset of spatio-temporally localized sports actions
Yixuan Li, Lei Chen, Runyu He, Zhenzhi Wang, Gangshan Wu, and Limin Wang. Multisports: A multi-person video dataset of spatio-temporally localized sports actions. In Int. Conf. Comput. Vis., 2021. 3
2021
-
[33]
Mask-adapter: The devil is in the masks for open- vocabulary segmentation
Yongkang Li, Tianheng Cheng, Wenyu Liu, and Xinggang Wang. Mask-adapter: The devil is in the masks for open- vocabulary segmentation. arXiv preprint arXiv:2412.04533,
-
[34]
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Doll ´ar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In IEEE Conf. Comput. Vis. Pattern Recog., 2017. 4
2017
-
[35]
GRES: Generalized referring expression segmentation
Chang Liu, Henghui Ding, and Xudong Jiang. GRES: Generalized referring expression segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 5
2023
-
[36]
Primitivenet: decomposing the global constraints for referring segmenta- tion
Chang Liu, Xudong Jiang, and Henghui Ding. Primitivenet: decomposing the global constraints for referring segmenta- tion. Visual Intelligence, 2024. 2
2024
-
[37]
Multi- grained Temporal Prototype Learning for Few-shot Video Object Segmentation
Nian Liu, Kepan Nan, Wangbo Zhao, Yuanwei Liu, Xiwen Yao, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Junwei Han, and Fahad Shahbaz Khan. Multi- grained Temporal Prototype Learning for Few-shot Video Object Segmentation. In Int. Conf. Comput. Vis. , 2023. 1, 2, 3
2023
-
[38]
Convbench: A multi-turn conversation evaluation benchmark with hierarchical ablation capability for large vision-language models
Shuo Liu, Kaining Ying, Hao Zhang, Yue Yang, Yuqi Lin, Tianle Zhang, Chuanhao Li, Yu Qiao, Ping Luo, Wenqi Shao, et al. Convbench: A multi-turn conversation evaluation benchmark with hierarchical ablation capability for large vision-language models. In Adv. Neural Inform. Proc...
2024
-
[39]
Video swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 4, 5, 6
2022
-
[40]
Exploring the Better Correlation for Few-Shot Video Object Segmentation
Naisong Luo, Yuan Wang, Rui Sun, Guoxin Xiong, Tianzhu Zhang, and Feng Wu. Exploring the Better Correlation for Few-Shot Video Object Segmentation. IEEE Trans. Circuits Syst. Video Technol., 2024. 1, 2, 3
2024
-
[41]
Holistic prototype attention network for few-shot video object segmentation.IEEE Trans
Naisong Luo, Yuan Wang, Rui Sun, Guoxin Xiong, Tianzhu Zhang, and Feng Wu. Holistic prototype attention network for few-shot video object segmentation.IEEE Trans. Circuits Syst. Video Technol., 2024. 3
2024
-
[42]
Few-shot video object segmentation with prototype evolution
Binjie Mao, Xiyan Liu, Linsu Shi, Jiazhong Yu, Fei Li, and Shiming Xiang. Few-shot video object segmentation with prototype evolution. Neural Computing and Applications ,
-
[43]
Video object segmentation using space-time memory networks
Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim. Video object segmentation using space-time memory networks. In Int. Conf. Comput. Vis., 2019. 2
2019
-
[44]
Actor-context-actor relation network for spatio-temporal action localization
Junting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu, Jing Shao, and Hongsheng Li. Actor-context-actor relation network for spatio-temporal action localization. In Int. Conf. Comput. Vis., 2021. 3
2021
-
[45]
Idd-x: A multi-view dataset for ego- relative important object localization and explanation in dense and unstructured traffic
Chirag Parikh, Rohit Saluja, CV Jawahar, and Ravi Kiran Sarvadevabhatla. Idd-x: A multi-view dataset for ego- relative important object localization and explanation in dense and unstructured traffic. In IEEE Int. Conf. Robot. Autom., 2024. 3
2024
-
[46]
A benchmark dataset and evaluation methodology for video object segmentation
Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 2
2016
-
[48]
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 3
2024 arXiv
-
[49]
A 3- dimensional sift descriptor and its application to action recognition
Paul Scovanner, Saad Ali, and Mubarak Shah. A 3- dimensional sift descriptor and its application to action recognition. In ACM Int. Conf. Multimedia, 2007. 3
2007
-
[50]
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang. Temporal action localization in untrimmed videos via multi-stage cnns. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 3
2016
-
[51]
Temporal transductive inference for few- shot video object segmentation
Mennatullah Siam. Temporal transductive inference for few- shot video object segmentation. Int. J. Comput. Vis., 2025. 1, 3, 6
2025
-
[52]
Temporal transductive inference for few-shot video object segmentation
Mennatullah Siam, Konstantinos G Derpanis, and Richard P Wildes. Temporal transductive inference for few-shot video object segmentation. arXiv preprint arXiv:2203.14308 ,
-
[53]
Ucf101: A dataset of 101 human actions classes from videos in the wild
K Soomro. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 ,
-
[54]
Holistic prototype attention network for few-shot video object segmentation
Yin Tang, Tao Chen, Xiruo Jiang, Yazhou Yao, Guo-Sen Xie, and Heng-Tao Shen. Holistic prototype attention network for few-shot video object segmentation. IEEE Trans. Circuit Syst. Video Technol., 2023. 1, 2, 3, 5, 6, 7, 8
2023
-
[55]
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. J. Mach. Learn. Res., 2008. 7
2008
-
[56]
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid. Action recognition with improved trajectories. In Int. Conf. Comput. Vis. , pages 3551–3558, 2013. 3
2013
-
[57]
Panet: Few-shot image semantic segmentation with prototype alignment
Kaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou, and Jiashi Feng. Panet: Few-shot image semantic segmentation with prototype alignment. In Int. Conf. Comput. Vis., 2019. 1
2019
-
[58]
Monet: Deep motion exploitation for video 10 object segmentation
Huaxin Xiao, Jiashi Feng, Guosheng Lin, Yu Liu, and Maojun Zhang. Monet: Deep motion exploitation for video 10 object segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2018. 2
2018
-
[59]
Self-calibrated cross attention network for few-shot segmentation
Qianxiong Xu, Wenting Zhao, Guosheng Lin, and Cheng Long. Self-calibrated cross attention network for few-shot segmentation. In Int. Conf. Comput. Vis., 2023. 1, 5, 6
2023
-
[60]
Eliminating feature ambiguity for few-shot segmentation
Qianxiong Xu, Guosheng Lin, Chen Change Loy, Cheng Long, Ziyue Li, and Rui Zhao. Eliminating feature ambiguity for few-shot segmentation. In Eur. Conf. Comput. Vis., 2024. 1
2024
-
[61]
Referred by multi-modality: A unified temporal transformer for video object segmentation
Shilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen, Wei Zhang, Hongyang Li, Yu Qiao, Hao Dong, Zhongjiang He, and Peng Gao. Referred by multi-modality: A unified temporal transformer for video object segmentation. In AAAI, 2024. 2
2024
-
[62]
Efficient video object segmentation via network modulation
Linjie Yang, Yanran Wang, Xuehan Xiong, Jianchao Yang, and Aggelos K Katsaggelos. Efficient video object segmentation via network modulation. In IEEE Conf. Comput. Vis. Pattern Recog., 2018. 2
2018
-
[63]
Video instance segmentation
Linjie Yang, Yuchen Fan, and Ning Xu. Video instance segmentation. In Int. Conf. Comput. Vis., 2019. 3, 5
2019
-
[64]
Revisiting anchor mechanisms for temporal action localization
Le Yang, Houwen Peng, Dingwen Zhang, Jianlong Fu, and Junwei Han. Revisiting anchor mechanisms for temporal action localization. IEEE Trans. Image Process., 2020. 3
2020
-
[65]
Isda: Position-aware instance segmentation with deformable attention
Kaining Ying, Zhenhua Wang, Cong Bai, and Pengfei Zhou. Isda: Position-aware instance segmentation with deformable attention. In IEEE Int. Conf. Acoust. Speech Signal Process.,
-
[66]
CTVIS: Consistent training for online video instance segmentation
Kaining Ying, Qing Zhong, Weian Mao, Zhenhua Wang, Hao Chen, Lin Yuanbo Wu, Yifan Liu, Chengxiang Fan, Yunzhi Zhuge, and Chunhua Shen. CTVIS: Consistent training for online video instance segmentation. In Int. Conf. Comput. Vis., 2023. 5
2023
-
[67]
MMT-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask AGI
Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu, Renrui Zhang, Haozhe Zhang, Peng Gao, Yali Wang, Yu Qiao, Ping Luo, Kaipeng Zhang, and Wenqi Shao. MMT-bench: A compr...
2024
-
[68]
Towards omnimodal expressions and reasoning in referring audio-visual segmentation
Kaining Ying, Henghui Ding, Guangquan Jie, and Yu-Gang Jiang. Towards omnimodal expressions and reasoning in referring audio-visual segmentation. In Int. Conf. Comput. Vis., 2025. 2
2025
-
[69]
Few-shot segmentation via cycle-consistent trans- former
Gengwei Zhang, Guoliang Kang, Yi Yang, and Yunchao Wei. Few-shot segmentation via cycle-consistent trans- former. Adv. Neural Inform. Process. Syst., 2021. 1, 6
2021
-
[70]
Egogesture: a new dataset and benchmark for egocentric hand gesture recognition
Yifan Zhang, Congqi Cao, Jian Cheng, and Hanqing Lu. Egogesture: a new dataset and benchmark for egocentric hand gesture recognition. IEEE Trans. Multimedia , 2018. 3
2018
-
[71]
Video self- stitching graph network for temporal action localization
Chen Zhao, Ali K Thabet, and Bernard Ghanem. Video self- stitching graph network for temporal action localization. In Int. Conf. Comput. Vis., 2021. 3
2021
-
[72]
A survey on deep learning technique for video segmentation
Tianfei Zhou, Fatih Porikli, David J Crandall, Luc Van Gool, and Wenguan Wang. A survey on deep learning technique for video segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 2022. 1
2022
-
[73]
A key volume mining deep framework for action recognition
Wangjiang Zhu, Jie Hu, Gang Sun, Xudong Cao, and Yu Qiao. A key volume mining deep framework for action recognition. In IEEE Conf. Comput. Vis. Pattern Recog. ,
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.