REVIEW 4 major objections 5 minor 1 cited by
On Moving Object Segmentation from Monocular Video with Transformers
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read M3Former, a two-stream transformer that fuses appearance and motion features through attention, claims state-of-the-art moving-object segmentation on KITTI and DAVIS when trained on a diverse mix of datasets.
desk verdict Useful transformer fusion and data-mix analysis for motion segmentation, but the SOTA claim rests on a training-mix mismatch and fails on DAVIS. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the two-stream attention fusion inside M3Former, a Multi-Modal Mask2Former built on the Mask2Former architecture. An appearance branch and a motion branch each use a ResNet-50 backbone, a multi-scale deformable-attention encoder, and a masked-attention decoder with 100 object queries; fusion happens by concatenating feature maps, query embeddings, and learned attention masks so that attention can flow between the streams, and a $1\times1$ convolution merges the mask and class outputs of all queries. A second load-bearing mechanism is the negative-example augmentation: with probability $p_{\mathrm{neg}}$, the motion input is replaced by a constant flow field, which forces the network not to over-rely on appearance data. The input motion representations are computed by frozen expert models—RAFT optical flow, RAFT-3D scene flow, DPT depth—so that the architecture consumes pseudo-modalities rather than raw video.
What would settle it
Retrain M3Former under the identical zero-shot protocol of the prior work—only the SceneFlow-based Mix 1 training set—and compare KITTI average precision against Raptor's reported 40.07; the paper already reports 25.67 for this setting, so a gap in either direction would settle whether the architecture itself is state of the art or whether the diverse training mix is doing the work.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that motion segmentation should be formulated as a fusion problem and that a two-stream transformer with dedicated parameters for appearance and motion—M3Former—solves it better than prior CNN-based fusion. The two branches produce multi-scale features and separate sets of object queries; at fusion points the feature maps, queries, and learned attention masks are concatenated, so attention can decide freely which modality to trust, and a $1\times1$ convolution merges the final predictions. Systematic ablations show that no single fusion strategy wins on all datasets or modalities, that scene flow—a per-pixel rigid-body motion field—carries the most information when depth quality is sufficient, and that the largest improvement comes from training on several datasets with diverse motion patterns and semantic classes. The paper's stated result is state-of-the-art average precision on KITTI and DAVIS for the fusion model trained on its most diverse data mix.
Load-bearing premise
The state-of-the-art comparison is fair even though the new model's training data includes the training splits of the very KITTI and DAVIS sets it is evaluated on, while the prior methods it is compared with (Raptor, Generic MoSeg) were evaluated without seeing those datasets.
Editorial extensions
If this is right
- A practical motion-segmentation system can be assembled from an off-the-shelf appearance transformer, a motion branch, and a diverse training mix, without handcrafted geometric criteria.
- Training-data construction and balancing become part of the method itself: adding real driving data and casual video with non-rigid motion removes failure modes that architecture changes cannot.
- The negative-example augmentation is a cheap, transferable regularizer for any fusion setting where one modality is noisy or sometimes absent.
- With enough diverse data, 2D optical flow suffices for competitive results on driving scenes; 3D scene flow only pays off when depth maps are reliable and scale-accurate.
- Because fusion happens through attention, the same two-stream design extends to longer temporal windows or additional modalities by adding streams.
Reading between the lines
- The headline state-of-the-art claim is protocol-dependent: under the zero-shot Mix 1 protocol the paper reports KITTI AP 25.67 against Raptor's 40.07, so the published comparison mixes training protocols; a matched-protocol benchmark would settle how much of the gain is architecture versus training data.
- The dataset-mix study suggests a practical recipe many practitioners could adopt: combine synthetic scene-flow data, real driving data, and casual video, sample datasets roughly equally, and set the negative-example probability near 0.3 for small mixes and near 0.05 for large ones.
- If segmentation is indeed driven mostly by the appearance stream, a much lighter motion branch would likely suffice; the paper hints at this but keeps both streams equal in size.
- A quantitative evaluation on held-out camouflage datasets (the paper shows only qualitative MoCA examples) would directly test the claim that the model discovers truly generic moving objects, not just familiar classes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes M3Former, a two-stream transformer architecture for monocular motion segmentation that fuses appearance features (RGB) with motion features obtained from frozen expert models (optical flow, scene flow, and higher-dimensional motion costs). The authors systematically analyze different motion representations, fusion mechanisms (encoder, decoder, and bottleneck-based), and a negative-example augmentation that prevents over-reliance on appearance. They also study the effect of training-data diversity by designing several dataset mixes and report instance-segmentation metrics on KITTI and DAVIS. The paper claims state-of-the-art (SotA) performance on both benchmarks, attributing this to the flexible attention mechanism and diverse training data.
Significance. If the SotA claim were fully supported, the paper would be a valuable contribution to monocular motion segmentation, particularly because it provides a clean comparison of 2D vs. 3D motion representations and a simple, effective negative-example augmentation. The systematic ablation of fusion locations and the analysis of training-data composition are useful and reproducible extensions of Mask2Former. However, the central empirical claim is undermined by a non-comparable evaluation protocol: the best M3Former results are obtained with a training mix that includes the target-domain training splits, whereas the comparison methods were evaluated zero-shot after training only on SceneFlow. The paper's own admission that it 'partially trained on the target domain' (§4.4) makes this limitation explicit. Because the headline claim is load-bearing and is not supported by an apples-to-apples comparison, the contribution reduces to an architecture study plus a data-diversity analysis unless the comparison is redone.
major comments (4)
- [Table 7 / §4.4] The state-of-the-art claim rests on a non-comparable training protocol. Table 7 reports Mix 3 results for M3Former (e.g., KITTI AP 52.27 with RGB+SF), but the comparison methods Raptor [44] and Generic MoSeg [18] were trained only on Mix 1 (SceneFlow) and evaluated zero-shot. The text at §4.4 explicitly admits: 'Our results are not necessarily surprising, as we partially trained on the target domain.' Under the zero-shot Mix 1 rows of Table 7, M3Former scores KITTI AP 25.67 (RGB+OF), 26.70 (RGB+SF), and 32.40 (RGB+Cost), all below Raptor's 40.07. To substantiate the SotA claim, the authors must either retrain the baselines under the Mix 3 protocol or present Mix 1 as the primary comparison; otherwise the improvement cannot be attributed to the architecture.
- [Table 7] Even taking the reported Mix 3 numbers at face value, the abstract's claim of 'SotA performance on Kitti and Davis' is not supported on DAVIS. In Table 7, Raptor attains DAVIS AP 40.9, while the best M3Former entry (RGB+SF, Mix 3) reaches only 37.07; no M3Former row exceeds Raptor's DAVIS AP. The claim should be restricted to KITTI, or the DAVIS comparison should be reworked with baselines retrained under the same protocol.
- [§4.4, Mix 1 rows of Table 7] The statement that 'we could not replicate the performance of [44] with the training setting of Mix 1' is a self-reported negative result that is not adequately investigated. Under identical Mix 1 data, Raptor's RGB + Cost input reaches KITTI AP 40.07, whereas the M3Former variant with the same input (RGB + Cost [44]) reaches only 32.40. The paper does not analyze whether this gap stems from the motion-cost computation, the fusion architecture, or training details. Since this is the only controlled comparison where both methods use the same modality and same training set, this gap is decisive for the claim that M3Former itself is competitive.
- [Table 2 and §4.4] The conclusion that 'diverse datasets are required to achieve SotA performance' is confounded with target-domain supervision. Mix 2 includes DAVIS train frames and Mix 3 includes both KITTI and DAVIS train frames (Table 2), so the large gains reported on those benchmarks from Mix 1 to Mix 2/3 may reflect label-space familiarity rather than diversity per se. To support the diversity claim, the authors should isolate the effect of adding datasets that are neither KITTI nor DAVIS (e.g., YTVOS, which they explicitly exclude), or ablate Mix 3 with the target-domain splits removed.
minor comments (5)
- [§3.2] There is an incomplete sentence in the paragraph on Multi-modal Alignment: '...the model needs to figure out to rely on the' — the sentence ends mid-thought. It should be completed accordingly.
- [§2 and §3.1] Typographical errors: 'adress' should be 'address' (§2) and 'indepent' should be 'independent' (§3.1).
- [Table 5 vs. Tables 6–7] Table 5 reports raw FP/FN counts and notes they are not normalized, while Tables 6 and 7 report per-frame normalized values. Using different conventions without a clear cross-reference is confusing; state the convention explicitly in each caption or use a single convention throughout.
- [Table 2] The caption refers to 'Colored datasets are used for evaluation,' but in monochrome the table does not indicate which datasets are colored. Name the evaluation datasets explicitly (KITTI and DAVIS) in the caption.
- [Abstract / Conclusion] The abstract states 'SotA performance on Kitti and Davis,' while the conclusion words the claim as 'our approach achieves SotA performance by leveraging the flexible attention mechanism and diverse training data.' The discrepancy in specificity should be resolved, especially in light of the protocol issues discussed above.
Circularity Check
No circularity: the paper is an empirical benchmark study whose claims are supported by test-split evaluations and whose components are cited from external prior work.
full rationale
M3Former is presented as a new fusion architecture, and its headline claims ('SotA performance on Kitti and Davis', 'diverse datasets are required') are empirical statements backed by measured metrics in Tables 3-7 rather than by a formal derivation. There is no equation-level chain in which a predicted quantity is defined in terms of the same quantity or in which a fitted parameter is renamed as a prediction. The only potentially problematic aspect is evaluation protocol: the paper's best Mix 3 results include the KITTI and DAVIS train splits, while prior methods such as Raptor and Generic MoSeg are reported after Mix 1 training only; the paper itself acknowledges this ('Our results are not necessarily surprising, as we partially trained on the target domain'). That is a comparison-fairness or soundness concern, not a circularity of definitions or arguments. The paper's citations of Mask2Former, RAFT, RAFT-3D, DPT, Raptor, and Learning Rigid Motions are standard external references; none of the load-bearing premises is justified solely by a self-citation to a prior work of the same authors, and no uniqueness theorem or ansatz is smuggled in via citation. Therefore, applying the hard rule that circularity may only be claimed when a specific reduction can be exhibited, no circular step can be identified, and the honest finding is a score of 0.
Assumptions & free parameters
free parameters (3)
- negative-example probability p_neg =
0.3 for Mix 1/2, 0.05 for Mix 3
- fusion mechanism =
decoder (D) or encoder+decoder (E+D), chosen per dataset and modality
- training data mix =
Mix 3 (FT3D, Monkaa, Driving, Vkitti, Kitti, Davis)
assumptions (4)
- domain assumption Motion segmentation labels are merged into a single 'object' class, and the task is formulated as supervised instance segmentation over binary motion state.
- domain assumption Frozen expert models (RAFT, RAFT-3D, DPT, etc.) provide motion pseudo-modalities that are useful inputs, and their errors can be compensated by fusion with appearance.
- domain assumption Transformer attention can effectively fuse multi-scale features from two modalities, and Mask2Former's masked-attention decoder transfers to a two-stream design.
- domain assumption A COCO-pretrained appearance branch provides a semantic prior that should be frozen, and this prior helps generalization to moving objects.
Cite this review
Pith. "Pith review of On Moving Object Segmentation from Monocular Video with Transformers." pith.science (2026). https://pith.science/paper/SXYPGQ5N
@misc{pith2026241119141,
author = {Pith},
title = {Pith review of: On Moving Object Segmentation from Monocular Video with Transformers},
year = {2026},
howpublished = {\url{https://pith.science/paper/SXYPGQ5N}},
note = {Machine review of arXiv:2411.19141}
}
read the original abstract
Moving object detection and segmentation from a single moving camera is a challenging task, requiring an understanding of recognition, motion and 3D geometry. Combining both recognition and reconstruction boils down to a fusion problem, where appearance and motion features need to be combined for classification and segmentation. In this paper, we present a novel fusion architecture for monocular motion segmentation - M3Former, which leverages the strong performance of transformers for segmentation and multi-modal fusion. As reconstructing motion from monocular video is ill-posed, we systematically analyze different 2D and 3D motion representations for this problem and their importance for segmentation performance. Finally, we analyze the effect of training data and show that diverse datasets are required to achieve SotA performance on Kitti and Davis.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 1 Pith paper
-
STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision
Using automotive radar Doppler as supervision, STAR-VLM enables a vision-language model to estimate metric radial velocity and motion state of objects from video, outperforming zero-shot task-specific baselines on a n...
Reference graph
Works this paper leans on
-
[44]
Monocular arbitrary moving object discovery and segmentation
Michal Neoral, Jan ˇSochman, and Jir ´ı Matas. Monocular arbitrary moving object discovery and segmentation. 2021. 1, 2, 3, 4, 5, 6, 7, 8, 13, 14, 18, 19
work page 2021
-
[18]
To- wards segmenting anything that moves
Achal Dave, Pavel Tokmakov, and Deva Ramanan. To- wards segmenting anything that moves. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 1, 2, 3, 5, 6, 7, 8, 13, 14
2019
-
[1]
A database and eval- uation methodology for optical flow
Simon Baker, Daniel Scharstein, JP Lewis, Stefan Roth, Michael J Black, and Richard Szeliski. A database and eval- uation methodology for optical flow. International journal of computer vision, 92:1–31, 2011. 4
2011
-
[2]
Multimodal machine learning: A survey and tax- onomy
Tadas Baltru ˇsaitis, Chaitanya Ahuja, and Louis-Philippe Morency. Multimodal machine learning: A survey and tax- onomy. IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443, 2018. 5
2018
-
[3]
Discovering objects that can move
Zhipeng Bao, Pavel Tokmakov, Allan Jabri, Yu-Xiong Wang, Adrien Gaidon, and Martial Hebert. Discovering objects that can move. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11789– 11798, 2022. 2, 13
2022
-
[4]
Object Discovery from Motion-Guided Tokens
Zhipeng Bao, Pavel Tokmakov, Yu-Xiong Wang, Adrien Gaidon, and Martial Hebert. Object discovery from motion- guided tokens. arXiv preprint arXiv:2303.15555, 2023. 2, 13
work page Pith review arXiv 2023
-
[5]
It’s moving! a prob- abilistic model for causal motion segmentation in moving camera videos
Pia Bideau and Erik Learned-Miller. It’s moving! a prob- abilistic model for causal motion segmentation in moving camera videos. In Computer Vision–ECCV 2016: 14th Eu- ropean Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part VIII 14 , pages 433–449. Springer, 2016. 2, 13
2016
-
[6]
Moa-net: self-supervised motion segmentation
Pia Bideau, Rakesh R Menon, and Erik Learned-Miller. Moa-net: self-supervised motion segmentation. In Pro- ceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0–0, 2018. 2, 13
2018
Show all 110 references
-
[7]
The best of both worlds: Combining cnns and geometric constraints for hierarchical motion seg- mentation
Pia Bideau, Aruni RoyChowdhury, Rakesh R Menon, and Erik Learned-Miller. The best of both worlds: Combining cnns and geometric constraints for hierarchical motion seg- mentation. In Proceedings of the IEEE conference on com- puter vision and pattern recognition , pages 508–517...
2018
-
[8]
Neural-guided ransac: Learning where to sample model hypotheses
Eric Brachmann and Carsten Rother. Neural-guided ransac: Learning where to sample model hypotheses. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 4322–4331, 2019. 2
2019
-
[9]
Object segmentation by long term analysis of point trajectories
Thomas Brox and Jitendra Malik. Object segmentation by long term analysis of point trajectories. In Computer Vision– ECCV 2010: 11th European Conference on Computer Vi- sion, Heraklion, Crete, Greece, September 5-11, 2010, Pro- ceedings, Part V 11, pages 282–295. Springer, 2010. 2
2010
-
[10]
A naturalistic open source movie for opti- cal flow evaluation
Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black. A naturalistic open source movie for opti- cal flow evaluation. In Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part VI 12 , pages 611–
2012
-
[11]
Vir- tual kitti 2
Yohann Cabon, Naila Murray, and Martin Humenberger. Vir- tual kitti 2. arXiv preprint arXiv:2001.10773, 2020. 6
2001 arXiv
-
[12]
End-to- end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to- end object detection with transformers. In Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, Au- gust 23–28, 2020, Proceedings, Part I 16 , pa...
2020
-
[13]
Mask2former for video instance segmentation
Bowen Cheng, Anwesa Choudhuri, Ishan Misra, Alexan- der Kirillov, Rohit Girdhar, and Alexander G Schwing. Mask2former for video instance segmentation. arXiv preprint arXiv:2112.10764, 2021. 1, 3, 4, 5, 8, 19
2021 arXiv
-
[14]
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1290–1299, 2022. 4, 5, 6, 13, 14
2022
-
[15]
Per- pixel classification is not all you need for semantic segmen- tation
Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per- pixel classification is not all you need for semantic segmen- tation. Advances in Neural Information Processing Systems, 34:17864–17875, 2021. 4
2021
-
[16]
Guess what moves: unsupervised video and image segmentation by anticipating motion
Subhabrata Choudhury, Laurynas Karazija, Iro Laina, An- drea Vedaldi, and Christian Rupprecht. Guess what moves: unsupervised video and image segmentation by anticipating motion. arXiv preprint arXiv:2205.07844, 2022. 2, 13
2022 arXiv
-
[17]
Robust estimation of a multi-layered motion representation
Trevor Darrell and Alexander Pentland. Robust estimation of a multi-layered motion representation. In Proceedings of the IEEE Workshop on Visual Motion, pages 173–174. IEEE Computer Society, 1991. 2
1991
-
[19]
Fusion- seg: Learning to combine motion and appearance for fully automatic segmentation of generic objects in videos
Suyog Dutt Jain, Bo Xiong, and Kristen Grauman. Fusion- seg: Learning to combine motion and appearance for fully automatic segmentation of generic objects in videos. In Pro- ceedings of the IEEE conference on computer vision and pat- tern recognition, pages 3664–3673, 2017. 2, 3
2017
-
[20]
Savi++: Towards end-to-end object-centric learning from real-world videos
Gamaleldin Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Michael C Mozer, and Thomas Kipf. Savi++: Towards end-to-end object-centric learning from real-world videos. Advances in Neural Information Processing Systems, 35:28940–28954, 2022. 2, 13
2022
-
[21]
Detection free track- ing: Exploiting motion and topology for segmenting and tracking under entanglement
Katerina Fragkiadaki and Jianbo Shi. Detection free track- ing: Exploiting motion and topology for segmenting and tracking under entanglement. In CVPR 2011, pages 2073–
2011
-
[22]
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. The Inter- national Journal of Robotics Research , 32(11):1231–1237,
-
[23]
Separate visual pathways for perception and action.Trends in neurosciences, 15(1):20–25, 1992
Melvyn A Goodale and A David Milner. Separate visual pathways for perception and action.Trends in neurosciences, 15(1):20–25, 1992. 1
1992
-
[24]
In defense of the eight-point algorithm
Richard I Hartley. In defense of the eight-point algorithm. IEEE Transactions on pattern analysis and machine intelli- gence, 19(6):580–593, 1997. 4
1997
-
[25]
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross Gir- shick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 14
2017
-
[26]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 5, 13
2016
-
[27]
Vita: Video instance seg- mentation via object token association
Miran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee, and Seon Joo Kim. Vita: Video instance seg- mentation via object token association. arXiv preprint arXiv:2206.04403, 2022. 3
2022 arXiv
-
[28]
Detect- ing and tracking multiple moving objects using temporal in- tegration
Michal Irani, Benny Rousso, and Shmuel Peleg. Detect- ing and tracking multiple moving objects using temporal in- tegration. In Computer Vision—ECCV’92: Second Euro- pean Conference on Computer Vision Santa Margherita Lig- ure, Italy, May 19–22, 1992 Proceedings 2, pages 282–2...
1992
-
[29]
Pixel objectness
Suyog Dutt Jain, Bo Xiong, and Kristen Grauman. Pixel objectness. arXiv preprint arXiv:1701.05349, 2017. 2
2017 arXiv
-
[30]
Unsupervised multi- object segmentation by predicting probable motion patterns
Laurynas Karazija, Subhabrata Choudhury, Iro Laina, Chris- tian Rupprecht, and Andrea Vedaldi. Unsupervised multi- object segmentation by predicting probable motion patterns. arXiv preprint arXiv:2210.12148, 2022. 2, 13
2022 arXiv
-
[31]
Segment any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. arXiv preprint arXiv:2304.02643, 2023. 3
2023 arXiv
-
[32]
Ro- bust consistent video depth estimation
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang. Ro- bust consistent video depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1611–1621, 2021. 6
2021
-
[33]
Betrayed by motion: Camouflaged object discovery via motion segmentation
Hala Lamdouar, Charig Yang, Weidi Xie, and Andrew Zis- serman. Betrayed by motion: Camouflaged object discovery via motion segmentation. In Proceedings of the Asian Con- ference on Computer Vision, 2020. 5, 16, 19
2020
-
[34]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceeding...
2014
-
[35]
The emergence of objectness: Learning zero-shot segmentation from videos
Runtao Liu, Zhirong Wu, Stella Yu, and Stephen Lin. The emergence of objectness: Learning zero-shot segmentation from videos. Advances in Neural Information Processing Systems, 34:13137–13152, 2021. 2, 13
2021
-
[36]
Prismer: A vision- language model with an ensemble of experts
Shikun Liu, Linxi Fan, Edward Johns, Zhiding Yu, Chaowei Xiao, and Anima Anandkumar. Prismer: A vision- language model with an ensemble of experts. arXiv preprint arXiv:2303.02506, 2023. 3
2023 arXiv
-
[37]
Robust dynamic radiance fields
Yu-Lun Liu, Chen Gao, Andreas Meuleman, Hung-Yu Tseng, Ayush Saraf, Changil Kim, Yung-Yu Chuang, Jo- hannes Kopf, and Jia-Bin Huang. Robust dynamic radiance fields. arXiv preprint arXiv:2301.02239, 2023. 1, 2, 6, 15
2023 arXiv
-
[38]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 13
2017 arXiv
-
[39]
See more, know more: Unsuper- vised video object segmentation with co-attention siamese networks
Xiankai Lu, Wenguan Wang, Chao Ma, Jianbing Shen, Ling Shao, and Fatih Porikli. See more, know more: Unsuper- vised video object segmentation with co-attention siamese networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3623–3632,
-
[40]
Learning rigidity in dynamic scenes with a moving camera for 3d motion field estimation
Zhaoyang Lv, Kihwan Kim, Alejandro Troccoli, Deqing Sun, James M Rehg, and Jan Kautz. Learning rigidity in dynamic scenes with a moving camera for 3d motion field estimation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 468–484, 2018. 2, 13
2018
-
[41]
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE conference on computer vision and ...
2016
-
[42]
Modetr: Mov- ing object detection with transformers
Eslam Mohamed and Ahmad El-Sallab. Modetr: Mov- ing object detection with transformers. arXiv preprint arXiv:2106.11422, 2021. 3, 13
2021 arXiv
-
[43]
Attention bottlenecks for multimodal fusion
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun. Attention bottlenecks for multimodal fusion. Advances in Neural Information Pro- cessing Systems, 34:14200–14213, 2021. 3, 5, 7, 14, 15, 16
2021
-
[45]
Segmenta- tion of moving objects by long term video analysis
Peter Ochs, Jitendra Malik, and Thomas Brox. Segmenta- tion of moving objects by long term video analysis. IEEE transactions on pattern analysis and machine intelligence , 36(6):1187–1200, 2013. 2
2013
-
[46]
The 2017 davis challenge on video object segmentation.arXiv preprint arXiv:1704.00675, 2017
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Ar- bel´aez, Alex Sorkine-Hornung, and Luc Van Gool. The 2017 davis challenge on video object segmentation.arXiv preprint arXiv:1704.00675, 2017. 6
2017 arXiv
-
[47]
Vi- sion transformers for dense prediction
Ren ´e Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vi- sion transformers for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 12179–12188, 2021. 1, 2, 3, 4, 6, 8, 15, 17
2021
-
[48]
Optical flow estima- tion using a spatial pyramid network
Anurag Ranjan and Michael J Black. Optical flow estima- tion using a spatial pyramid network. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 4161–4170, 2017. 4
2017
-
[49]
Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation
Anurag Ranjan, Varun Jampani, Lukas Balles, Kihwan Kim, Deqing Sun, Jonas Wulff, and Michael J Black. Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation. In Proceed- ings of the IEEE/CVF conference on computer v...
2019
-
[50]
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San- jeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115:211–252, 2015. 13
2015
-
[51]
3d geometry from planar parallax
Harpreet S Sawhney. 3d geometry from planar parallax. In CVPR, volume 94, pages 929–934, 1994. 2
1994
-
[52]
Motion segmentation and tracking using normalized cuts
Jianbo Shi and Jitendra Malik. Motion segmentation and tracking using normalized cuts. InSixth international confer- ence on computer vision (IEEE Cat. No. 98CH36271), pages 1154–1160. IEEE, 1998. 2
1998
-
[53]
Simple unsu- pervised object-centric learning for complex and naturalistic videos
Gautam Singh, Yi-Fu Wu, and Sungjin Ahn. Simple unsu- pervised object-centric learning for complex and naturalistic videos. arXiv preprint arXiv:2205.14065, 2022. 2, 13
2022 arXiv
-
[54]
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part II 16, pages 402–419. Springer,
2020
-
[55]
Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
Zachary Teed and Jia Deng. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in neu- ral information processing systems, 34:16558–16569, 2021. 7, 15
2021
-
[56]
Raft-3d: Scene flow using rigid- motion embeddings
Zachary Teed and Jia Deng. Raft-3d: Scene flow using rigid- motion embeddings. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 8375–8384, 2021. 1, 2, 3, 4, 6
2021
-
[57]
Learning video object segmentation with visual memory
Pavel Tokmakov, Karteek Alahari, and Cordelia Schmid. Learning video object segmentation with visual memory. In Proceedings of the IEEE International Conference on Com- puter Vision, pages 4481–4490, 2017. 2, 3, 13
2017
-
[58]
Geometric motion segmentation and model selection
Philip HS Torr. Geometric motion segmentation and model selection. Philosophical Transactions of the Royal Society of London. Series A: Mathematical, Physical and Engineering Sciences, 356(1740):1321–1340, 1998. 2
1998
-
[59]
The problem of degeneracy in structure and motion recovery from uncalibrated image sequences
Philip HS Torr, Andrew W Fitzgibbon, and Andrew Zisser- man. The problem of degeneracy in structure and motion recovery from uncalibrated image sequences. International Journal of Computer Vision, 32:27–44, 1999. 2
1999
-
[60]
Robust detection of degenerate configurations while estimat- ing the fundamental matrix
Philip HS Torr, Andrew Zisserman, and Stephen J Maybank. Robust detection of degenerate configurations while estimat- ing the fundamental matrix. Computer vision and image un- derstanding, 71(3):312–333, 1998. 2
1998
-
[61]
A benchmark for the com- parison of 3-d motion segmentation algorithms
Roberto Tron and Ren ´e Vidal. A benchmark for the com- parison of 3-d motion segmentation algorithms. In 2007 IEEE conference on computer vision and pattern recogni- tion, pages 1–8. IEEE, 2007. 2
2007
-
[62]
Video segmentation via object flow
Yi-Hsuan Tsai, Ming-Hsuan Yang, and Michael J Black. Video segmentation via object flow. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 3899–3908, 2016. 2
2016
-
[63]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 3, 4
2017
-
[64]
Motion segmentation with missing data using powerfactorization and gpca
Ren ´e Vidal and Richard Hartley. Motion segmentation with missing data using powerfactorization and gpca. In Pro- ceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004., volume 2, pages II–II. IEEE, 2004. 2
2004
-
[65]
Optimal segmentation of dynamic scenes from two perspective views
Ren ´e Vidal and Shankar Sastry. Optimal segmentation of dynamic scenes from two perspective views. In 2003 IEEE Computer Society Conference on Computer Vision and Pat- tern Recognition, 2003. Proceedings., volume 2, pages II–II. IEEE, 2003. 2
2003
-
[66]
Sfm- net: Learning of structure and motion from video
Sudheendra Vijayanarasimhan, Susanna Ricco, Cordelia Schmid, Rahul Sukthankar, and Katerina Fragkiadaki. Sfm- net: Learning of structure and motion from video. arXiv preprint arXiv:1704.07804, 2017. 2
2017 arXiv
-
[67]
Opti- cal flow in mostly rigid scenes
Jonas Wulff, Laura Sevilla-Lara, and Michael J Black. Opti- cal flow in mostly rigid scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 4671–4680, 2017. 2
2017
-
[68]
Object discovery in videos as foreground motion clustering
Christopher Xie, Yu Xiang, Zaid Harchaoui, and Dieter Fox. Object discovery in videos as foreground motion clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 9994–10003, 2019. 2, 3, 13
2019
-
[69]
Segment- ing moving objects via an object-centric layered representa- tion
Junyu Xie, Weidi Xie, and Andrew Zisserman. Segment- ing moving objects via an object-centric layered representa- tion. In Advances in Neural Information Processing Systems,
-
[70]
Unify- ing flow, stereo and depth estimation
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. Unify- ing flow, stereo and depth estimation. arXiv preprint arXiv:2211.05783, 2022. 6
2022 arXiv
-
[71]
Youtube-vos: A large-scale video object segmentation benchmark
Ning Xu, Linjie Yang, Yuchen Fan, Dingcheng Yue, Yuchen Liang, Jianchao Yang, and Thomas Huang. Youtube-vos: A large-scale video object segmentation benchmark. arXiv preprint arXiv:1809.03327, 2018. 6, 17
2018 arXiv
-
[72]
3d rigid mo- tion segmentation with mixed and unknown number of mod- els
Xun Xu, Loong-Fah Cheong, and Zhuwen Li. 3d rigid mo- tion segmentation with mixed and unknown number of mod- els. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(1):1–16, 2019. 2
2019
-
[73]
A general framework for motion segmentation: Independent, articulated, rigid, non- rigid, degenerate and non-degenerate
Jingyu Yan and Marc Pollefeys. A general framework for motion segmentation: Independent, articulated, rigid, non- rigid, degenerate and non-degenerate. In Computer Vision– ECCV 2006: 9th European Conference on Computer Vi- sion, Graz, Austria, May 7-13, 2006, Proceedings, Part...
2006
-
[74]
Self-supervised video object segmentation by motion grouping
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, and Weidi Xie. Self-supervised video object segmentation by motion grouping. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 7177–7188,
-
[75]
Upgrading optical flow to 3d scene flow through optical expansion
Gengshan Yang and Deva Ramanan. Upgrading optical flow to 3d scene flow through optical expansion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 1334–1343, 2020. 2, 4
2020
-
[76]
Learning to seg- ment rigid motions from two frames
Gengshan Yang and Deva Ramanan. Learning to seg- ment rigid motions from two frames. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1266–1275, 2021. 2, 3, 4, 5, 6, 7, 8, 13, 14 11
2021
-
[77]
Video instance seg- mentation
Linjie Yang, Yuchen Fan, and Ning Xu. Video instance seg- mentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5188–5197, 2019. 3
2019
-
[78]
Unsupervised moving object detection via contextual information separation
Yanchao Yang, Antonio Loquercio, Davide Scaramuzza, and Stefano Soatto. Unsupervised moving object detection via contextual information separation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 879–888, 2019. 2
2019
-
[79]
Every pixel counts: Unsupervised geometry learn- ing with holistic 3d motion understanding
Zhenheng Yang, Peng Wang, Yang Wang, Wei Xu, and Ram Nevatia. Every pixel counts: Unsupervised geometry learn- ing with holistic 3d motion understanding. InProceedings of the European conference on computer vision (ECCV) work- shops, pages 0–0, 2018. 2
2018
-
[80]
Detecting motion regions in the presence of a strong parallax from a moving camera by multiview geometric con- straints
Chang Yuan, Gerard Medioni, Jinman Kang, and Isaac Co- hen. Detecting motion regions in the presence of a strong parallax from a moving camera by multiview geometric con- straints. IEEE transactions on pattern analysis and machine intelligence, 29(9):1627–1641, 2007. 2
2007
-
[81]
Consistent depth of moving objects in video
Zhoutong Zhang, Forrester Cole, Richard Tucker, William T Freeman, and Tali Dekel. Consistent depth of moving objects in video. ACM Transactions on Graphics (TOG) , 40(4):1– 12, 2021. 6
2021
-
[82]
Particlesfm: Exploiting dense point trajecto- ries for localizing moving cameras in the wild
Wang Zhao, Shaohui Liu, Hengkai Guo, Wenping Wang, and Yong-Jin Liu. Particlesfm: Exploiting dense point trajecto- ries for localizing moving cameras in the wild. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Is- rael, October 23–27, 2022, Proceedings, Part...
2022
-
[83]
Motion-attentive transition for zero-shot video object segmentation
Tianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao, Jianwu Li, and Ling Shao. Motion-attentive transition for zero-shot video object segmentation. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 13066–13073, 2020. 3, 13
2020
-
[84]
Deformable detr: Deformable trans- formers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable trans- formers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 3, 5, 13
2010 arXiv
-
[85]
Segment everything every- where all at once
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Gao, and Yong Jae Lee. Segment everything every- where all at once. arXiv preprint arXiv:2304.06718, 2023. 3 12 Supplementary Material In this supplementary material, we provide further de- tails on our approach...
2023 arXiv
-
[86]
Placement in the Literature There is a vast amount of related literature on segmen- tation, motion segmentation, moving object discovery and unsupervised feature learning
Our Approach 6.1. Placement in the Literature There is a vast amount of related literature on segmen- tation, motion segmentation, moving object discovery and unsupervised feature learning. Table 8 summarizes the de- velopment in this field and where our approach fits in. We c...
-
[87]
Training Details We follow a similar training setup as [14]
Experimental Setup 7.1. Training Details We follow a similar training setup as [14]. Our net- works are optimized using AdamW [38] with a learning rate of 1.0 × 10−4 and a weight decay of 0.05 for all back- bones. A learning rate multiplier of 0.1 is applied to the backbone. W...
-
[88]
ECCV 2016 ✗ 2 ✗ Optical Flow BS ✗
2016
-
[89]
ICCV 2017 ✓ ≥3 ✓ RGB + Optica l Flow BS GRU
2017
-
[90]
ECCV 2018 ✓ 2 ✗ Optical Flow BS ✗
2018
-
[91]
ECCV 2018 ✓ 2 ✓ RGB-D BS, Odometry ✗
2018
-
[92]
CVPR 2018 ( ✓) 2 ✓ RGB + Optical Flow BS ✗
2018
-
[93]
CVPR 2019 ✓ 2 ✗ RGB BS Attention
2019
-
[94]
CVPR 2019 ✓ >2 ✗ RGB + Optical Flow IS / (VIS) Convolution
2019
-
[95]
ICCV 2019 ✓ 2 ✓ RGB + Optical Flow IS / (VIS) Convolution
2019
-
[96]
NeurIPS 2020 ✓ 2 ✓ RGB + Optical Flow Detection Attention
2020
-
[97]
AAAI 2020 ✓ 2 ✓ RGB + Optical Flow BS Attention
2020
-
[98]
NeurIPS 2021 ✓ 2 ✗ Optical Flow BS ✗
2021
-
[99]
ICCV 2021 ✓ 2 ✗ Optical Flow BS ✗
2021
-
[100]
Costs BS / IS ✗
CVPR 2021 ✓ 2 ✓ Geom. Costs BS / IS ✗
2021
-
[101]
Costs IS Convolution
BMVC 2021 ✓ 3 ✓ RGB + Geom. Costs IS Convolution
2021
-
[102]
CVPR 2022 ✓ 3 ✗ RGB Object Discovery GRU
2022
-
[103]
NeurIPS 2022 ✓ 1 ✗ RGB BS ✗
2022
-
[105]
NeurIPS 2022 ✓ ≤2 ✗ RGB VOS / Object Discovery RNN
2022
-
[106]
BMVC 2022 ✓ ≤2 ✗ RGB + Optical Flow Binary Segmentation Spectral Clustering
2022
-
[107]
no-object
CVPR 2023 ✓ 3 ✗ RGB Object Discovery Attention Ours 2023 ✓ ≥2 ✓ RGB + 3D scene flow IS Attention Table 8: Taxonomy of related segmentation literature. We distinguish Binary Segmentation (BS), Instance Segmentation (IS) and Object Segmentation (OS) as task acronyms. 13 matched ...
2023
-
[108]
Similar to [44], we average over different IoU’s [0.01, 0.1, 0.3, 0.5, 0.75, 0.9, 0.95]
metrics, foreground/background precision [76], Preci- sion (Pu), Recall (Ru) and F-score (Fu) [18] and the num- ber of false positives and false negatives over the whole split [44]. Similar to [44], we average over different IoU’s [0.01, 0.1, 0.3, 0.5, 0.75, 0.9, 0.95]. Since ...
-
[109]
Note how scal- ing depends on the individual complexities (of the atten- tion mechanism)
We measure with a batch size of 1. Note how scal- ing depends on the individual complexities (of the atten- tion mechanism). Since computation of pseudo-modalities is dependent on specific off-the-shelf expert models and in- put resolution, we omit a total runtime comparison. ...
-
[110]
Ablation Fusion Strategies We ablate different fusion mechanisms for image and op- tical flow input data and measure the effect of using differ- ent training data
Additional Results 8.1. Ablation Fusion Strategies We ablate different fusion mechanisms for image and op- tical flow input data and measure the effect of using differ- ent training data. Results can be seen in Table 11. During our initial experiments we did not find consisten...
-
[111]
More Visualizations In this section we add more visualizations to better ex- plain our model behavior. We give further examples of the attention in both streams, failure cases, differences between training data mixes and generalization on the Moving Cam- ouflaged Animal (MoCA)...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.