REVIEW 3 major objections 6 minor 60 references
Reprojecting frozen vision-language features across views is a free training signal that improves feed-forward 3D reconstruction and open-vocabulary 3D semantics.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 07:34 UTC pith:EP67DIMZ
load-bearing objection Clean auxiliary loss that measurably helps both geometry and open-vocab 3D fusion; the proxy assumption is untested but the empirical package is solid enough to engage. the 3 major comments →
VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Multi-view inconsistency of frozen dense vision-language features, when induced by a model’s own predicted geometry, is a usable training signal for feed-forward 3D estimators. Adding Vision-Language Reprojection Consistency as an auxiliary loss improves depth, camera pose, and intrinsics for self-supervised monocular reconstruction and for unlabeled adaptation of supervised-pretrained models, without new 3D annotations, and yields geometry better aligned for open-vocabulary 3D semantic fusion.
What carries the argument
Vision-Language Reprojection Consistency (VLRC): predicted depth, pose, and intrinsics warp dense frozen vision-language feature maps across views; a cosine dissimilarity loss (with a validity mask for unreliable correspondences) is applied between reprojected and target features. Gradients update the 3D model while the vision-language encoder stays frozen, forcing geometry to explain multi-view language-aligned consistency.
Load-bearing premise
If the predicted geometry is correct, frozen language-aligned features will agree across views; if it is wrong, they will disagree enough to give a useful training gradient.
What would settle it
Train matched models with and without VLRC on the same unlabeled video, then check whether reprojected CLIP cosine similarity at known true correspondences rises exactly when measured depth and pose error fall; if geometry improves while feature consistency does not (or the reverse), the claimed mechanism fails.
If this is right
- Self-supervised monocular 3D models can improve depth, pose, and intrinsics by adding a frozen VLM feature-reprojection term without collecting geometric labels.
- Unlabeled post-training of large supervised feed-forward 3D models can be strengthened by the same auxiliary signal.
- Geometry trained with VLRC supports more coherent multi-view aggregation of dense VLM features, raising zero-shot open-vocabulary 3D segmentation scores.
- Casual monocular video becomes more usable for both metric reconstruction and free-form text localization in 3D.
- The same fused 3D VLM representation can serve both fixed category sets and arbitrary text queries.
Where Pith is reading between the lines
- If the vision-language model were unfrozen and co-adapted, the same consistency loop could make features more geometry-aware rather than treating them as a fixed teacher.
- The same reprojection idea may transfer to other frozen foundation encoders (audio-visual, event, or region-level descriptors) wherever multi-view semantic structure already exists.
- Photometric failure modes on specular or textureless surfaces may systematically shrink when language-aligned features remain stable across those regions.
- How far VLRC can push reconstruction may be bounded by the domain gap between the VLM’s pretraining distribution and extreme in-the-wild video.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Vision-Language Reprojection Consistency (VLRC), an auxiliary multi-view loss that reprojects frozen dense vision-language features (primarily CLIP-Seg) using a feed-forward model’s own predicted depth, pose, and intrinsics and penalizes cosine inconsistency (Eq. 5). VLRC is added to photometric self-supervised training (SS3D) and to SelfEvo-style unlabeled adaptation of a supervised-pretrained model (VGGT). Across KITTI, NYUv2, Sintel, TUM-RGBD, and ScanNet200, the authors report consistent improvements in depth, camera motion, and intrinsics, plus stronger zero-shot open-vocabulary 3D semantic segmentation under a Casper3D protocol and a new KITTI geometry-to-semantics protocol. No extra 3D annotations are required; the VLM encoder remains frozen.
Significance. If the empirical gains hold under stronger controls, VLRC is a practical, annotation-free auxiliary signal that can be dropped into both self-supervised monocular reconstruction and unlabeled post-training of large feed-forward 3D models. The dual benefit—better core geometry and more coherent multi-view VLM fusion for open-vocabulary 3D understanding—is useful for scalable in-the-wild 3D pretraining. Strengths include a clean formulation, evaluation under two training regimes, a backbone ablation (Table 3), a new KITTI open-vocab protocol with released evaluation code, and qualitative free-form localization on web video. The contribution is incremental relative to prior feature-metric and multi-view consistency losses, but the specific use of frozen language-aligned dense features for feed-forward 3D pretraining is a clear and reusable idea.
major comments (3)
- [Sec. 3.2, Eq. (5)] Sec. 3.2, Eq. (5): The central mechanism assumes that multi-view inconsistency of frozen dense VLM features is a reliable proxy for geometric error. The manuscript never reports L_VLRC (or residual multi-view feature variance) under ground-truth depth/pose/intrinsics versus under baseline predicted geometry on any calibrated multi-view set. Without this diagnostic, it is unclear how much of the signal is geometric versus residual viewpoint/lighting/occlusion sensitivity of CLIP-style features. A short GT-geometry vs. predicted-geometry comparison (even on a subset of KITTI or ScanNet) is load-bearing for interpreting the method and should be added.
- [Sec. 4.1, Eq. (5)] Sec. 4.1 and Eq. (5): In the SS3D regime the validity mask m_t,s is built from optical-flow vs. depth-pose discrepancy and excludes dynamic objects, occlusions, and unreliable geometry—precisely the ambiguous regions the introduction claims VLRC should help. The paper should quantify how much of the reported depth/pose gains remain when the mask is removed or when the loss is restricted to hard pixels, and should reconcile this with the SelfEvo setting (no mask, smaller gains in Table 4). Otherwise the claim that VLRC supplies useful gradients where photometric consistency fails is not supported by the experimental design.
- [Tables 1, 2, 4, 5] Tables 1, 2, 4, 5: All main results are single-run point estimates with no error bars, seeds, or significance tests. Several improvements are small (e.g., Table 4 Sintel Abs Rel 0.212→0.209; Table 1 Abs Rel 0.064→0.060). Given that λ_VLRC, crop, and schedule are free hyperparameters, at least multi-seed means/std or a short sensitivity sweep on λ_VLRC is needed before the “consistent gains” claim can be treated as robust, especially for the SelfEvo adaptation numbers.
minor comments (6)
- [Fig. 1, Abstract] Fig. 1 caption and abstract: “Y ouTube8M” has a spurious space; fix throughout.
- [Fig. 3] Fig. 3 prompt text: “Eifel-tower” should be “Eiffel Tower” for consistency with the figure caption.
- [Sec. 3.3] Sec. 3.3: The fusion weight α_ij is introduced then set to uniform averaging “for clarity”; state explicitly which fusion is used in each reported open-vocab number (Table 5 vs. qualitative figures).
- [Sec. 2] Related work: Feature-metric loss (Shu et al., ECCV 2020) and other non-RGB reprojection cues are cited briefly; a short paragraph contrasting VLRC with those feature-space photometric alternatives would clarify novelty.
- [Appendix A] Appendix A / Table A.1: Clarify whether SegFormer pseudo-labels and CLIP Cityscapes prompts use identical class sets and whether mIoU is computed only on LiDAR-projected points that are valid in all compared methods.
- [Sec. 4.1] Implementation: λ_VLRC = 0.01 (SS3D) vs. 0.1 (SelfEvo) is stated without justification; a one-sentence note on how these values were chosen would help reproducibility.
Circularity Check
No significant circularity: VLRC is an empirical auxiliary loss whose gains are measured on external held-out benchmarks; self-citations supply only baselines and base models.
full rationale
The paper defines VLRC (Eq. 5) as a cosine-dissimilarity reprojection loss on frozen dense VLM features induced by the model's own predicted depth/pose/intrinsics, then adds it with a fixed hyper-parameter λ_VLRC to either the photometric self-supervised objective (SS3D) or a SelfEvo-style distillation objective (VGGT). All reported improvements (Abs Rel, δ1, ATE, mIoU, etc.) are measured on standard external datasets (KITTI, NYUv2, Sintel, TUM-RGBD, ScanNet200) that are not used to fit free parameters of the claimed result. Self-citations to the authors' prior SS3D and Casper3D works merely provide the base architectures and the downstream fusion protocol; they do not supply the target metrics or force the observed gains by construction. No equation reduces a reported quantity to a fitted input, no uniqueness theorem is imported, and no known empirical pattern is merely renamed. The untested proxy assumption (that multi-view VLM inconsistency reliably tracks geometric error) is a validity concern, not circularity. Score 1 reflects only the presence of non-load-bearing self-citations for baselines.
Axiom & Free-Parameter Ledger
free parameters (3)
- λ_VLRC =
0.01 / 0.1
- learning-rate schedule =
1e-5 → 1e-8
- similarity threshold τ =
0.5
axioms (4)
- domain assumption Correct 3D geometry induces multi-view consistency of dense frozen VLM features; incorrect geometry produces detectable inconsistency that can be used as a training signal.
- standard math Standard multi-view geometry (back-projection, SE(3) relative pose, perspective projection, bilinear sampling) correctly maps corresponding pixels when depth/pose/intrinsics are accurate.
- domain assumption Optical-flow vs. depth-pose discrepancy reliably identifies dynamic objects and occlusions for the validity mask m_t,s.
- domain assumption Frozen dense CLIP-derived features (upsampled by FeatUp) remain informative for both geometric supervision and open-vocabulary fusion.
invented entities (1)
-
Vision-Language Reprojection Consistency (VLRC) loss
independent evidence
read the original abstract
Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojection Consistency (VLRC), a scalable auxiliary objective that exploits frozen vision-language representations as semantic multi-view supervision. Given a predicted 3D reconstruction, VLRC reprojects dense vision-language features across views and enforces feature consistency between corresponding image locations, requiring no additional 3D annotations. The objective integrates seamlessly with both self-supervised monocular reconstruction and supervised-pretrained feed-forward 3D models during unlabeled adaptation. By aligning geometry with language-grounded features, VLRC not only improves depth and camera estimation but also enables more coherent multi-view semantic fusion for open-vocabulary 3D scene understanding. Experiments on indoor and outdoor benchmarks demonstrate consistent gains in 3D reconstruction accuracy and zero-shot open-vocabulary 3D semantic segmentation.
Figures
Reference graph
Works this paper leans on
-
[1]
Youtube-8m: A large- scale video classification benchmark.arXiv preprint arXiv:1609.08675, 2016
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large- scale video classification benchmark.arXiv preprint arXiv:1609.08675, 2016. 6, 1, 3
Pith/arXiv arXiv 2016
-
[2]
Met3r: Measuring multi-view consistency in generated images
Mohammad Asim, Christopher Wewer, Thomas Wimmer, Bernt Schiele, and Jan Eric Lenssen. Met3r: Measuring multi-view consistency in generated images. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 6034–6044, 2025. 3
2025
-
[3]
A naturalistic open source movie for op- tical flow evaluation
Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black. A naturalistic open source movie for op- tical flow evaluation. InEuropean conference on computer vision, pages 611–625. Springer, 2012. 6, 1
2012
-
[4]
Tipsv2: Advancing vision-language pretraining with enhanced patch-text align- ment
Bingyi Cao, Koert Chen, Kevis-Kokitsi Maninis, Kaifeng Chen, Arjun Karpur, Ye Xia, Sahil Dua, Tanmaya Dabral, Guangxing Han, Bohyung Han, et al. Tipsv2: Advancing vision-language pretraining with enhanced patch-text align- ment. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 29325–29335,
-
[5]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. InPro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 3
2021
-
[6]
Vl-jepa: Joint embedding predictive architecture for vision-language
Delong Chen, Mustafa Shukor, Theo Moutakanni, Willy Chung, Jade Yu, Tejaswi Kasarla, Yejin Bang, Allen Bolourchi, Yann LeCun, and Pascale Fung. Vl-jepa: Joint embedding predictive architecture for vision-language. arXiv preprint arXiv:2512.10942, 2025. 2
arXiv 2025
-
[7]
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017. 6
2017
-
[8]
Pla: Language-driven open- vocabulary 3d scene understanding
Runyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang, Song Bai, and Xiaojuan Qi. Pla: Language-driven open- vocabulary 3d scene understanding. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7010–7019, 2023. 3
2023
-
[9]
Deeper into self-supervised monocular indoor depth estimation
Chao Fan, Zhenyu Yin, Yue Li, and Feiqing Zhang. Deeper into self-supervised monocular indoor depth estimation. arXiv preprint arXiv:2312.01283, 2023. 7
Pith/arXiv arXiv 2023
-
[10]
Self- supervised models are continual learners
Enrico Fini, Victor G Turrisi Da Costa, Xavier Alameda- Pineda, Elisa Ricci, Karteek Alahari, and Julien Mairal. Self- supervised models are continual learners. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9621–9630, 2022. 3
2022
-
[11]
Stephanie Fu, Mark Hamilton, Laura Brandt, Axel Feldman, Zhoutong Zhang, and William T Freeman. Featup: A model- agnostic framework for features at any resolution.arXiv preprint arXiv:2403.10516, 2024. 6
Pith/arXiv arXiv 2024
-
[12]
Vision meets robotics: The kitti dataset.The in- ternational journal of robotics research, 32(11):1231–1237,
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset.The in- ternational journal of robotics research, 32(11):1231–1237,
-
[13]
Digging into self-supervised monocular depth estimation
Cl ´ement Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J Brostow. Digging into self-supervised monocular depth estimation. InProceedings of the IEEE/CVF inter- national conference on computer vision, pages 3828–3838,
-
[14]
Re- balancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictions
Marwane Hariat, Antoine Manzanera, and David Filliat. Re- balancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictions. InProceed- ings of the IEEE/CVF winter conference on applications of computer vision, pages 1267–1276, 2023. 3, 6, 2
2023
-
[15]
Im- proved monocular depth prediction using distance transform over pre-semantic contours with self-supervised neural net- works
Marwane Hariat, Antoine Manzanera, and David Filliat. Im- proved monocular depth prediction using distance transform over pre-semantic contours with self-supervised neural net- works. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 21868–21879, 2025. 3, 7
2025
-
[16]
Lightweight 3d feature pretraining by bayesian inversion of 2d foundation models
Marwane Hariat, Gianni Franchi, David Filliat, and Antoine Manzanera. Lightweight 3d feature pretraining by bayesian inversion of 2d foundation models. 2026. 3, 6, 8
2026
-
[17]
Ss3d: End2end self-supervised 3d from web videos.arXiv preprint arXiv:2604.22686, 2026
Marwane Hariat, Gianni Franchi, David Filliat, and Antoine Manzanera. Ss3d: End2end self-supervised 3d from web videos.arXiv preprint arXiv:2604.22686, 2026. 2, 4, 6
Pith/arXiv arXiv 2026
-
[18]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000– 16009, 2022. 3
2022
-
[19]
Ra-depth: Resolution adaptive self-supervised monocular depth estimation
Mu He, Le Hui, Yikai Bian, Jian Ren, Jin Xie, and Jian Yang. Ra-depth: Resolution adaptive self-supervised monocular depth estimation. InEuropean Conference on Computer Vi- sion, pages 565–581. Springer, 2022. 7
2022
-
[20]
Self- improving 4d perception via self-distillation.arXiv preprint arXiv:2604.08532, 2026
Nan Huang, Pengcheng Yu, Weijia Zeng, James M Rehg, Angjoo Kanazawa, Haiwen Feng, and Qianqian Wang. Self- improving 4d perception via self-distillation.arXiv preprint arXiv:2604.08532, 2026. 2, 3, 6
Pith/arXiv arXiv 2026
-
[21]
Open-vocabulary 3d semantic segmentation with foundation models
Li Jiang, Shaoshuai Shi, and Bernt Schiele. Open-vocabulary 3d semantic segmentation with foundation models. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21284–21294, 2024. 3, 6
2024
-
[22]
Mapanything: Universal feed-forward metric 3d re- construction.arXiv preprint arXiv:2509.13414, 2025
Nikhil Keetha, Norman M ¨uller, Johannes Sch ¨onberger, Lorenzo Porzi, Yuchen Zhang, Tobias Fischer, Arno Knapitsch, Duncan Zauss, Ethan Weber, Nelson Antunes, et al. Mapanything: Universal feed-forward metric 3d re- construction.arXiv preprint arXiv:2509.13414, 2025. 1, 2
Pith/arXiv arXiv 2025
-
[23]
Adam: A method for stochastic opti- mization.arXiv preprint arXiv:1412.6980, 2014
Diederik P Kingma. Adam: A method for stochastic opti- mization.arXiv preprint arXiv:1412.6980, 2014. 6
Pith/arXiv arXiv 2014
-
[24]
Ground- ing image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and Jerome Revaud. Ground- ing image matching in 3d with mast3r. InComputer Vi- sion – ECCV 2024: 18th European Conference, page 71–91. Springer-Verlag, 2024. 2
2024
-
[25]
Structdepth: Leveraging the structural regularities 9 for self-supervised indoor depth estimation
Boying Li, Yuan Huang, Zeyu Liu, Danping Zou, and Wenx- ian Yu. Structdepth: Leveraging the structural regularities 9 for self-supervised indoor depth estimation. InProceedings of the IEEE/CVF international conference on computer vi- sion, pages 12663–12673, 2021. 7
2021
-
[26]
Unsupervised monocular depth learning in dynamic scenes
Hanhan Li, Ariel Gordon, Hang Zhao, Vincent Casser, and Anelia Angelova. Unsupervised monocular depth learning in dynamic scenes. InConference on Robot Learning, pages 1908–1917. PMLR, 2021. 3
1908
-
[27]
Runze Li, Pan Ji, Yi Xu, and Bir Bhanu. Monoindoor++: Towards better practice of self-supervised monocular depth estimation for indoor environments.IEEE Transactions on Circuits and Systems for Video Technology, 33(2):830–846,
-
[28]
Depth anything 3: Recovering the visual space from any views
Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647, 2025. 1, 2
Pith/arXiv arXiv 2025
-
[29]
Partslip: Low-shot part segmentation for 3d point clouds via pretrained image- language models
Minghua Liu, Yinhao Zhu, Hong Cai, Shizhong Han, Zhan Ling, Fatih Porikli, and Hao Su. Partslip: Low-shot part segmentation for 3d point clouds via pretrained image- language models. InProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 21736–21746, 2023. 3
2023
-
[30]
Image segmenta- tion using text and image prompts
Timo L ¨uddecke and Alexander Ecker. Image segmenta- tion using text and image prompts. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7086–7096, 2022. 2, 6, 7
2022
-
[31]
Hr-depth: High resolution self-supervised monocular depth estimation
Xiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong, Lina Liu, Yong Liu, Xinxin Chen, and Yi Yuan. Hr-depth: High resolution self-supervised monocular depth estimation. InProceedings of the AAAI conference on artificial intelli- gence, pages 2294–2301, 2021. 7
2021
-
[32]
Brightflow: Brightness-change-aware un- supervised learning of optical flow
R ´emi Marsal, Florian Chabot, Ang ´elique Loesch, and Hichem Sahbi. Brightflow: Brightness-change-aware un- supervised learning of optical flow. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2061–2070, 2023. 3
2061
-
[33]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 2, 7
Pith/arXiv arXiv 2023
-
[34]
Pytorch: An im- perative style, high-performance deep learning library.Ad- vances in neural information processing systems, 32, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An im- perative style, high-performance deep learning library.Ad- vances in neural information processing systems, 32, 2019. 6
2019
-
[35]
Openscene: 3d scene understanding with open vocabularies
Songyou Peng, Kyle Genova, Chiyu Jiang, Andrea Tagliasacchi, Marc Pollefeys, Thomas Funkhouser, et al. Openscene: 3d scene understanding with open vocabularies. InProceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 815–824, 2023. 3
2023
-
[36]
Deep hough voting for 3d object detection in point clouds
Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. Inproceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 9277–9286, 2019. 3
2019
-
[37]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. InInternational conference on machine learning, pages 8748–8763. PmLR, 2021. 2
2021
-
[38]
Ren ´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637, 2020. 2
2020
-
[39]
Multi-task learning as multi-objective optimization.Advances in neural informa- tion processing systems, 31, 2018
Ozan Sener and Vladlen Koltun. Multi-task learning as multi-objective optimization.Advances in neural informa- tion processing systems, 31, 2018. 2
2018
-
[40]
Feature-metric loss for self-supervised learning of depth and egomotion
Chang Shu, Kun Yu, Zhixiang Duan, and Kuiyuan Yang. Feature-metric loss for self-supervised learning of depth and egomotion. InEuropean Conference on Computer Vision, pages 572–588. Springer, 2020. 3
2020
-
[41]
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgbd images. InEuropean conference on computer vision, pages 746–760. Springer, 2012. 6, 7
2012
-
[42]
Mv-clip: Multi- view clip for zero-shot 3d shape recognition.IEEE Transac- tions on Circuits and Systems for Video Technology, 35(9): 8767–8779, 2025
Dan Song, Xinwei Fu, Ning Liu, Wei-Zhi Nie, Wen-Hui Li, Lan-Jun Wang, You Yang, and An-An Liu. Mv-clip: Multi- view clip for zero-shot 3d shape recognition.IEEE Transac- tions on Circuits and Systems for Video Technology, 35(9): 8767–8779, 2025. 3
2025
-
[43]
A benchmark for the evalua- tion of rgb-d slam systems
J ¨urgen Sturm, Nikolas Engelhard, Felix Endres, Wolfram Burgard, and Daniel Cremers. A benchmark for the evalua- tion of rgb-d slam systems. In2012 IEEE/RSJ international conference on intelligent robots and systems, pages 573–580. IEEE, 2012. 6, 1
2012
-
[44]
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muham- mad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. Siglip 2: Multilingual vision-language en- coders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025. 2
Pith/arXiv arXiv 2025
-
[45]
Vggt: Vi- sual geometry grounded transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Vi- sual geometry grounded transformer. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 5294–5306, 2025. 1, 2, 4, 6, 7
2025
-
[46]
Continuous 3d per- ception model with persistent state
Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A Efros, and Angjoo Kanazawa. Continuous 3d per- ception model with persistent state. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 10510–10522, 2025. 1, 3
2025
-
[47]
Dust3r: Geometric 3d vi- sion made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vi- sion made easy. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20697– 20709, 2024. 1, 2
2024
-
[48]
Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chun- hua Shen, and Tong He.π 3: Permutation-equivariant visual geometry learning.arXiv preprint arXiv:2507.13347, 2025. 1 10
Pith/arXiv arXiv 2025
-
[49]
Croco: self-supervised pre-training for 3d vision tasks by cross-view completion
Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Ro- main Br´egier, Yohann Cabon, Vaibhav Arora, Leonid Ants- feld, Boris Chidlovskii, Gabriela Csurka, and J ´erˆome Re- vaud. Croco: self-supervised pre-training for 3d vision tasks by cross-view completion. InProceedings of the 36th Inter- national Conference on Neural Information Processing Sys- tems (...
2022
-
[50]
Anycam: Learning to re- cover camera poses and intrinsics from casual videos
Felix Wimbauer, Weirong Chen, Dominik Muhle, Christian Rupprecht, and Daniel Cremers. Anycam: Learning to re- cover camera poses and intrinsics from casual videos. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 16717–16727, 2025. 1, 2
2025
-
[51]
Segformer: Simple and efficient design for semantic segmentation with transform- ers.Advances in neural information processing systems, 34: 12077–12090, 2021
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transform- ers.Advances in neural information processing systems, 34: 12077–12090, 2021. 1
2021
-
[52]
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10371–10381, 2024. 2
2024
-
[53]
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. InProceedings of the IEEE/CVF international conference on computer vision, pages 11975–11986, 2023. 2
2023
-
[54]
Pointclip: Point cloud understanding by clip
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xu- peng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8552–8562, 2022. 3
2022
-
[55]
Monovit: Self-supervised monocular depth estimation with a vision transformer
Chaoqiang Zhao, Youmin Zhang, Matteo Poggi, Fabio Tosi, Xianda Guo, Zheng Zhu, Guan Huang, Yang Tang, and Ste- fano Mattoccia. Monovit: Self-supervised monocular depth estimation with a vision transformer. In2022 international conference on 3D vision (3DV), pages 668–678. IEEE, 2022. 7
2022
-
[56]
Extract free dense labels from clip
Chong Zhou, Chen Change Loy, and Bo Dai. Extract free dense labels from clip. InEuropean conference on computer vision, pages 696–712. Springer, 2022. 2
2022
-
[57]
Self- supervised monocular depth estimation with internal feature fusion
Hang Zhou, David Greenwood, and Sarah Taylor. Self- supervised monocular depth estimation with internal feature fusion. InBritish Machine Vision Conference (BMVC), 2021. 7
2021
-
[58]
Moving indoor: Unsupervised video depth learning in challenging environments
Junsheng Zhou, Yuwang Wang, Kaihuai Qin, and Wenjun Zeng. Moving indoor: Unsupervised video depth learning in challenging environments. InProceedings of the IEEE/CVF international conference on computer vision, pages 8618– 8627, 2019. 7
2019
-
[59]
Ov3d-cg: Open-vocabulary 3d instance segmentation with contextual guidance
Mingquan Zhou, Chen He, Ruiping Wang, and Xilin Chen. Ov3d-cg: Open-vocabulary 3d instance segmentation with contextual guidance. InProceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 5305–5314,
-
[60]
Unsupervised learning of depth and ego-motion from video
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe. Unsupervised learning of depth and ego-motion from video. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 1851–1858, 2017. 3 11 VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining Supplementary Material...
arXiv 2017
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.