Pith. sign in

REVIEW 1 minor 181 references

A data-centric taxonomy connects 3D geometric representations, datasets, and learning paradigms into one map.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-28 10:13 UTC pith:LWRYQQNG

load-bearing objection This is a survey paper that builds a taxonomy of 3D vision around representations like point clouds and Gaussians plus supervision types, but adds no new results or derivations.

arxiv 2606.04291 v1 pith:LWRYQQNG submitted 2026-06-02 cs.CV

A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

classification cs.CV
keywords 3D visiondata-centric taxonomygeometric representationspoint cloudsmeshesvoxels3D Gaussiansimplicit neural representations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper establishes a unified conceptual map for 3D vision by organizing it around geometric data representations and how they pair with different learning approaches. A reader would care because the field is currently split across many benchmarks and methods, making it hard to see how to build scalable systems. The taxonomy starts with core representations including point clouds, meshes, voxels, and 3D Gaussians and their data collection methods. It then connects these to dataset designs, supervision types like 2D-supervised learning and implicit neural representations, and applications in reconstruction, generation, and 4D modeling. The result is a clearer picture of trends that aim to improve both efficiency and accuracy in 3D tasks.

Core claim

We provide a data-centric taxonomy of 3D vision that connects geometric representations, datasets, learning frameworks, and applications within a single conceptual map. We begin by analysing the principal structural representations of 3D data--point clouds, meshes, voxels, and 3D Gaussians--along with their acquisition pipelines. We then examine how dataset design, benchmark construction, and supervision regimes shape recent advances, spanning 2D-supervised 3D learning, implicit neural representations, and 4D world modeling. Through this integrative lens, we clarify the relationships among representations, learning paradigms, and downstream tasks in reconstruction, generation, and video mode

What carries the argument

The data-centric taxonomy that integrates principal structural representations such as point clouds, meshes, voxels and 3D Gaussians with dataset design and supervision regimes to map connections to downstream tasks.

Load-bearing premise

The relationships among representations, learning paradigms, and downstream tasks can be clarified by analyzing principal structural representations and supervision regimes as described.

What would settle it

A survey or experiment finding no consistent patterns linking specific representation choices to measurable differences in task efficiency or fidelity would show that the taxonomy does not clarify the claimed relationships.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • The taxonomy reveals how choices among point clouds, meshes, voxels, and 3D Gaussians interact with 2D-supervised learning to affect reconstruction quality.
  • Dataset design and benchmark construction directly shape advances in implicit neural representations and 4D world modeling.
  • Clearer links between supervision regimes and applications support trends that balance efficiency against fidelity in generation and video modeling.
  • The map points to multimodal geometric grounding as a direction that ties representations to new task types.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy could be used to identify gaps where certain representation-supervision pairs lack dedicated benchmarks.
  • Developers might apply the map to choose representations that match hardware or latency constraints in new applications.
  • The same data-centric approach could be tested on related domains such as dynamic scene understanding to check if similar connections appear.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 1 minor

Summary. The manuscript presents a data-centric taxonomy of 3D vision that organizes geometric representations (point clouds, meshes, voxels, 3D Gaussians) and their acquisition pipelines, then connects these to dataset design, benchmark construction, and supervision regimes (2D-supervised 3D learning, implicit neural representations, 4D world modeling), ultimately mapping the elements to applications in reconstruction, generation, and video modeling.

Significance. If the taxonomy accurately and comprehensively links representations, supervision regimes, and tasks, the work would supply a useful integrative map for a fragmented field, clarifying relationships and trends toward efficiency-fidelity trade-offs and multimodal grounding without advancing new theorems or empirical results.

minor comments (1)
  1. [Abstract] The abstract is dense with terminology; a short overview paragraph or figure in the introduction that visually summarizes the taxonomy axes would improve accessibility for readers new to the subfield.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their positive review and recommendation to accept the manuscript. We appreciate the recognition that the data-centric taxonomy can serve as an integrative map linking representations, supervision regimes, and applications in 3D vision.

Circularity Check

0 steps flagged

No significant circularity; purely descriptive survey

full rationale

The paper constructs a data-centric taxonomy of 3D vision by surveying existing representations (point clouds, meshes, voxels, 3D Gaussians), acquisition methods, datasets, supervision regimes (2D-supervised, implicit, 4D), and applications. No equations, derivations, predictions, fitted parameters, or uniqueness theorems are present. The central contribution is an organizational map of the literature rather than any claim that reduces to its own inputs by construction. Self-citations, if present, are not load-bearing for any technical result because no technical results are derived. This matches the default expectation for a review paper with no derivation chain.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

This is a survey paper. No free parameters, axioms, or invented entities are introduced because the work does not advance new technical claims.

pith-pipeline@v0.9.1-grok · 5734 in / 1051 out tokens · 20804 ms · 2026-06-28T10:13:04.479521+00:00 · methodology

0 comments
read the original abstract

3D vision has rapidly evolved, driven by increasingly diverse data representations, learning paradigms, and modeling strategies. Yet the field remains fragmented across representations and benchmarks, making it difficult to develop unified perspectives on efficiency, fidelity, and scalability. This work provides a data-centric taxonomy of 3D vision that connects geometric representations, datasets, learning frameworks, and applications within a single conceptual map. We begin by analysing the principal structural representations of 3D data--point clouds, meshes, voxels, and 3D Gaussians--along with their acquisition pipelines. We then examine how dataset design, benchmark construction, and supervision regimes shape recent advances, spanning 2D-supervised 3D learning, implicit neural representations, and 4D world modeling. Through this integrative lens, we clarify the relationships among representations, learning paradigms, and downstream tasks in reconstruction, generation, and video modeling, offering a consolidated view of emerging trends toward balancing efficiency and fidelity and toward multimodal geometric grounding.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

181 extracted references · 13 canonical work pages · 11 internal anchors

  1. [1]

    Pointnetlk: Robust & efficient point cloud registration using pointnet

    Yasuhiro Aoki, Hunter Goforth, Rangaprasad Arun Srivatsan, and Simon Lucey. Pointnetlk: Robust & efficient point cloud registration using pointnet. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  2. [2]

    Navigation world models, 2025

    Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation world models, 2025

  3. [3]

    A dataset for semantic scene understanding of lidar sequences

    J Behley, M Garbade, A Milioto, J Quenzel, S Behnke, C Stachniss, J Gall, and Semantickitti. A dataset for semantic scene understanding of lidar sequences. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9297–9307

  4. [4]

    Virtual kitti 2, 2020

    Yohann Cabon, Naila Murray, and Martin Humenberger. Virtual kitti 2, 2020

  5. [5]

    Monoscene: Monocular 3d semantic scene completion, 2022

    Anh-Quan Cao and Raoul de Charette. Monoscene: Monocular 3d semantic scene completion, 2022

  6. [6]

    Matterport3d: Learning from rgb-d data in indoor environments, 2017

    Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Nießner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments, 2017

  7. [7]

    Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu

    Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. Shapenet: An information-rich 3d model repository, 2015

  8. [8]

    Tensorf: Tensorial radiance fields

    Anpei Chen, Zexiang Xu, Matthew Tancik, Jingyi Xu, Xiuming Zhang, Hiroharu Kato, and Jingyi Yu. Tensorf: Tensorial radiance fields. InECCV, pages 333–350, 2022

  9. [9]

    A survey on 3d gaussian splatting, 2025

    Guikun Chen and Wenguan Wang. A survey on 3d gaussian splatting, 2025

  10. [10]

    SAM 3D: 3Dfy Anything in Images

    Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, et al. Sam 3d: 3dfy anything in images.arXiv preprint arXiv:2511.16624, 2025

  11. [11]

    Parametric 20000

    Xi Cheng. Parametric 20000. Mendeley Data, V1, 2024

  12. [12]

    Robust reconstruction of indoor scenes

    Sungjoon Choi, Qian-Yi Zhou, and Vladlen Koltun. Robust reconstruction of indoor scenes. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5556–5565, 2015

  13. [13]

    A Large Dataset of Object Scans

    Sungjoon Choi, Qian-Yi Zhou, Stephen Miller, and Vladlen Koltun. A large dataset of object scans. arXiv:1602.02481, 2016

  14. [14]

    4d spatio-temporal convnets: Minkowski convolutional neural networks

    Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4d spatio-temporal convnets: Minkowski convolutional neural networks. InCVPR, pages 3075–3084, 2019

  15. [15]

    Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese

    Christopher B. Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction, 2016

  16. [16]

    3d u-net: Learning dense volumetric segmentation from sparse annotation

    Ozgun Cicek, Ahmed Abdulkadir, Soeren S Lienkamp, Thomas Brox, and Olaf Ronneberger. 3d u-net: Learning dense volumetric segmentation from sparse annotation. InMICCAI, pages 424–432, 2016

  17. [17]

    Abo: Dataset and benchmarks for real-world 3d object understanding.CVPR, 2022

    Jasmine Collins, Shubham Goel, Kenan Deng, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Xi Zhang, Tomas F Yago Vicente, Thomas Dideriksen, Himanshu Arora, Matthieu Guillaumin, and Jitendra Malik. Abo: Dataset and benchmarks for real-world 3d object understanding.CVPR, 2022

  18. [18]

    M. G. Cox. The numerical evaluation of b-splines.IMA Journal of Applied Mathematics, 10(2):134–149, 1972

  19. [19]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes

    Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2432–2443, 2017

  20. [20]

    Bundlefusion: Real- time globally consistent 3d reconstruction using on-the-fly surface reintegration.ACM Transactions on Graphics (TOG), 36(4):1–18, 2017

    Angela Dai, Matthias Nießner, Michael Zollhöfer, Shahram Izadi, and Christian Theobalt. Bundlefusion: Real- time globally consistent 3d reconstruction using on-the-fly surface reintegration.ACM Transactions on Graphics (TOG), 36(4):1–18, 2017

  21. [21]

    Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans

    Angela Dai, Maximilian Dahnert, and Matthias Nießner. Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans. InCVPR, pages 4578–4587, 2018

  22. [22]

    Brepformer: Transformer-based b-rep geometric feature recognition

    Yongkang Dai, Xiaoshui Huang, Yunpeng Bai, Hao Guo, Hongping Gan, Ling Yang, and Yilei Shi. Brepformer: Transformer-based b-rep geometric feature recognition. InProceedings of the 2025 International Conference on Multimedia Retrieval, page 155–163, New York, NY, USA, 2025. Association for Computing Machinery. 11

  23. [23]

    Springer, New York, revised 2001 edition, 1978

    Carl de Boor.A Practical Guide to Splines. Springer, New York, revised 2001 edition, 1978

  24. [24]

    Objaverse-XL: A Universe of 10M+ 3D Objects

    Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Objaverse-xl: A universe of 10m+ 3d objects.arXiv preprint arXiv:2307.05663, 2023

  25. [25]

    Videogpa: Distilling geometry priors for 3d-consistent video generation, 2026

    Hongyang Du, Junjie Ye, Xiaoyan Cong, Runhao Li, Jingcheng Ni, Aman Agarwal, Zeqi Zhou, Zekun Li, Randall Balestriero, and Yue Wang. Videogpa: Distilling geometry priors for 3d-consistent video generation, 2026

  26. [26]

    The mapillary traffic sign dataset for detection and classification on a global scale, 2020

    Christian Ertler, Jerneja Mislej, Tobias Ollmann, Lorenzo Porzi, Gerhard Neuhold, and Yubin Kuang. The mapillary traffic sign dataset for detection and classification on a global scale, 2020

  27. [27]

    A point set generation network for 3d object reconstruction from a single image, 2016

    Haoqiang Fan, Hao Su, and Leonidas Guibas. A point set generation network for 3d object reconstruction from a single image, 2016

  28. [28]

    A history-based parametric cad sketch dataset with advanced engineering commands.Computer-Aided Design, 182:103848, 2025

    Rubin Fan, Fazhi He, Yuxin Liu, and Jing Lin. A history-based parametric cad sketch dataset with advanced engineering commands.Computer-Aided Design, 182:103848, 2025

  29. [29]

    Morgan Kaufmann, San Diego, 5 edition, 2002

    Gerald Farin.Curves and Surfaces for CAGD: A Practical Guide. Morgan Kaufmann, San Diego, 5 edition, 2002

  30. [30]

    Alex Fisher, Ricardo Cannizzaro, Madeleine Cochrane, Chatura Nagahawatte, and Jennifer L. Palmer. Colmap: A memory-efficient occupancy grid mapping framework.Robotics and Autonomous Systems, 142:103755, 2021

  31. [31]

    3d-future: 3d furniture shape with texture, 2020

    Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang Zhao, Steve Maybank, and Dacheng Tao. 3d-future: 3d furniture shape with texture, 2020

  32. [32]

    3d-front: 3d furnished rooms with layouts and semantics, 2021

    Huan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang, Jiaming Wang Cao Li, Zengqi Xun, Chengyue Sun, Rongfei Jia, Binqiang Zhao, and Hao Zhang. 3d-front: 3d furnished rooms with layouts and semantics, 2021

  33. [33]

    Anyhome: Open-vocabulary generation of structured and textured 3d homes, 2024

    Rao Fu, Zehao Wen, Zichen Liu, and Srinath Sridhar. Anyhome: Open-vocabulary generation of structured and textured 3d homes, 2024

  34. [34]

    Gigahands: A massive annotated dataset of bimanual hand activities

    Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu, Austin Fund, Daniel Ritchie, and Srinath Sridhar. Gigahands: A massive annotated dataset of bimanual hand activities. 2025

  35. [35]

    Efros, and Xiaolong Wang

    Yang Fu, Sifei Liu, Amey Kulkarni, Jan Kautz, Alexei A. Efros, and Xiaolong Wang. Colmap-free 3d gaussian splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20796–20805, 2024

  36. [36]

    Accurate, dense, and robust multi-view stereopsis.IEEE Trans

    Yasutaka Furukawa and Jean Ponce. Accurate, dense, and robust multi-view stereopsis.IEEE Trans. on Pattern Analysis and Machine Intelligence, 32(8):1362–1376, 2010

  37. [37]

    Virtual worlds as proxy for multi-object tracking analysis, 2016

    Adrien Gaidon, Qiao Wang, Yohann Cabon, and Eleonora Vig. Virtual worlds as proxy for multi-object tracking analysis, 2016

  38. [38]

    Nerf: Neural radiance field in 3d vision: A comprehensive review (updated post-gaussian splatting), 2025

    Kyle Gao, Yina Gao, Hongjie He, Dening Lu, Linlin Xu, and Jonathan Li. Nerf: Neural radiance field in 3d vision: A comprehensive review (updated post-gaussian splatting), 2025

  39. [39]

    Submanifold sparse convolutional networks

    Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. Submanifold sparse convolutional networks. InCVPR, pages 9224–9232, 2018

  40. [40]

    Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

    Kristen Grauman, Andrew Westbury, Eugene Patterson, Tsung-Yi Fu, Gijsbert Halbertsma, Lijun Zhao, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 23533–23545, 2024

  41. [41]

    Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam Laradji, Hsueh-Ti (Derek) Liu, Henning Meyer, Yishu Miao, Derek Nowrouzezahrai, Cengiz Oztireli, Etienne Pot, Noha Radwan, Daniel Rebain, Sara Sabour, Meh...

  42. [42]

    Diffusion as shader: 3d-aware video diffusion for versatile video generation control, 2025

    Zekai Gu, Rui Yan, Jiahao Lu, Peng Li, Zhiyang Dou, Chenyang Si, Zhen Dong, Qifeng Liu, Cheng Lin, Ziwei Liu, Wenping Wang, and Yuan Liu. Diffusion as shader: 3d-aware video diffusion for versatile video generation control, 2025. 12

  43. [43]

    Roca: Robust cad model retrieval and alignment from a single image

    Can Gumeli, Angela Dai, and Matthias Niebner. Roca: Robust cad model retrieval and alignment from a single image. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 4012–4021. IEEE, 2022

  44. [44]

    Martin, and Shi-Min Hu

    Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R. Martin, and Shi-Min Hu. Pct: Point cloud transformer.Computational Visual Media, 7(2):187–199, 2021

  45. [45]

    Learning rich features from rgb-d images for object detection and segmentation

    Saurabh Gupta, Ross Girshick, Pablo Arbeláez, and Jitendra Malik. Learning rich features from rgb-d images for object detection and segmentation. InEuropean Conference on Computer Vision (ECCV), pages 345–360, 2014

  46. [46]

    Savinov, L

    Timo Hackel, N. Savinov, L. Ladicky, Jan D. Wegner, K. Schindler, and M. Pollefeys. SEMANTIC3D.NET: A new large-scale point cloud classification benchmark. InISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, pages 91–98, 2017

  47. [47]

    Dual transformer for point cloud analysis

    Xian-Feng Han, Yi-Fei Jin, Hui-Xian Cheng, and Guo-Qiang Xiao. Dual transformer for point cloud analysis. IEEE Transactions on Multimedia, 25:5638–5648, 2023

  48. [48]

    Meshcnn: A network with an edge

    Rana Hanocka, Amir Hertz, Noa Fish, Raja Giryes, Shachar Fleishman, and Daniel Cohen-Or. Meshcnn: A network with an edge. InACM Transactions on Graphics (TOG), pages 1–12, 2019

  49. [49]

    Deep learning based 3d segmentation: A survey, 2024

    Yong He, Hongshan Yu, Xiaoyan Liu, Zhengeng Yang, Wei Sun, Saeed Anwar, and Ajmal Mian. Deep learning based 3d segmentation: A survey, 2024

  50. [50]

    Rgb-d mapping: Using depth cameras for dense 3d modeling of indoor environments.The International Journal of Robotics Research, 31(5):647–663, 2012

    Peter Henry, Michael Krainin, Evan Herbst, Xiaofeng Ren, and Dieter Fox. Rgb-d mapping: Using depth cameras for dense 3d modeling of indoor environments.The International Journal of Robotics Research, 31(5):647–663, 2012

  51. [51]

    Hoffmann.Geometric and Solid Modeling: An Introduction

    Christoph M. Hoffmann.Geometric and Solid Modeling: An Introduction. Morgan Kaufmann, San Mateo, CA, 1989

  52. [52]

    Pf3plat: Pose-free feed-forward 3d gaussian splatting, 2025

    Sunghwan Hong, Jaewoo Jung, Heeseong Shin, Jisang Han, Jiaolong Yang, Chong Luo, and Seungryong Kim. Pf3plat: Pose-free feed-forward 3d gaussian splatting, 2025

  53. [53]

    CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers.arXiv preprint arXiv:2205.15868, 2022

  54. [54]

    LRM: Large Reconstruction Model for Single Image to 3D

    Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d.arXiv preprint arXiv:2311.04400, 2023

  55. [55]

    3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36, 2023

    Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36, 2023

  56. [56]

    Scenenn: A scene meshes dataset with annotations

    Binh-Son Hua, Quang-Hieu Pham, Duc Thanh Nguyen, Minh-Khoi Tran, Lap-Fai Yu, and Sai-Kit Yeung. Scenenn: A scene meshes dataset with annotations. InInternational Conference on 3D Vision (3DV), 2016

  57. [57]

    An embodied generalist agent in 3d world

    Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An embodied generalist agent in 3d world. InProceedings of the 41st International Conference on Machine Learning. JMLR.org, 2024

  58. [58]

    Deepmvs: Learning multi-view stereopsis

    Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang. Deepmvs: Learning multi-view stereopsis. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  59. [59]

    Particleformer: A 3d point cloud world model for multi-object, multi-material robotic manipulation, 2025

    Suning Huang, Qianzhong Chen, Xiaohan Zhang, Jiankai Sun, and Mac Schwager. Particleformer: A 3d point cloud world model for multi-object, multi-material robotic manipulation, 2025

  60. [60]

    Pointworld: Scaling 3d world models for in-the-wild robotic manipulation, 2026

    Wenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu, Dieter Fox, Kaichun Mo, and Li Fei-Fei. Pointworld: Scaling 3d world models for in-the-wild robotic manipulation, 2026

  61. [61]

    Kinectfusion: Real-time 3d reconstruction and interaction using a moving depth camera

    Shahram Izadi, David Kim, Otmar Hilliges, David Molyneaux, Richard Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Daniel Freeman, Andrew Davison, et al. Kinectfusion: Real-time 3d reconstruction and interaction using a moving depth camera. InProceedings of the 24th Annual ACM Symposium on User Interface Software and Technology (UIST), pages 559–568, 2011

  62. [62]

    Rayzer: A self-supervised large view synthesis model, 2025

    Hanwen Jiang, Hao Tan, Peng Wang, Haian Jin, Yue Zhao, Sai Bi, Kai Zhang, Fujun Luan, Kalyan Sunkavalli, Qixing Huang, and Georgios Pavlakos. Rayzer: A self-supervised large view synthesis model, 2025

  63. [63]

    Megasynth: Scaling up 3d scene 13 reconstruction with synthesized data

    Hanwen Jiang, Zexiang Xu, Desai Xie, Ziwen Chen, Haian Jin, Fujun Luan, Zhixin Shu, Kai Zhang, Sai Bi, Xin Sun, Jiuxiang Gu, Qixing Huang, Georgios Pavlakos, and Hao Tan. Megasynth: Scaling up 3d scene 13 reconstruction with synthesized data. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16441–16452, 2025

  64. [64]

    Rellis-3d dataset: Data, benchmarks and analysis, 2020

    Peng Jiang, Philip Osteen, Maggie Wigness, and Srikanth Saripalli. Rellis-3d dataset: Data, benchmarks and analysis, 2020

  65. [65]

    Tensoir: Tensorial inverse rendering, 2024

    Haian Jin, Isabella Liu, Peijia Xu, Xiaoshuai Zhang, Songfang Han, Sai Bi, Xiaowei Zhou, Zexiang Xu, and Hao Su. Tensoir: Tensorial inverse rendering, 2024

  66. [66]

    Neural 3d mesh renderer, 2017

    Hiroharu Kato, Yoshitaka Ushiku, and Tatsuya Harada. Neural 3d mesh renderer, 2017

  67. [67]

    Differentiable rendering: A survey, 2020

    Hiroharu Kato, Deniz Beker, Mihai Morariu, Takahiro Ando, Toru Matsuoka, Wadim Kehl, and Adrien Gaidon. Differentiable rendering: A survey, 2020

  68. [68]

    Screened poisson surface reconstruction.ACM Trans

    Michael Kazhdan and Hugues Hoppe. Screened poisson surface reconstruction.ACM Trans. Graph., 32(3), 2013

  69. [69]

    Poisson surface reconstruction.Proceedings of the Fourth Eurographics Symposium on Geometry Processing, 7:61–70, 2006

    Michael Kazhdan, Michael Bolitho, and Hugues Hoppe. Poisson surface reconstruction.Proceedings of the Fourth Eurographics Symposium on Geometry Processing, 7:61–70, 2006

  70. [70]

    Posenet: A convolutional network for real-time 6-dof camera relocalization

    Alex Kendall, Matthew Grimes, and Roberto Cipolla. Posenet: A convolutional network for real-time 6-dof camera relocalization. InProceedings of the IEEE International Conference on Computer Vision (ICCV), pages 2938–2946, 2015

  71. [71]

    3d gaussian splatting for real-time radiance field rendering, 2023

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering, 2023

  72. [72]

    Egohumans: An egocentric 3d multi-human benchmark, 2023

    Rawal Khirodkar, Aayush Bansal, Lingni Ma, Richard Newcombe, Minh Vo, and Kris Kitani. Egohumans: An egocentric 3d multi-human benchmark, 2023

  73. [73]

    Parallel tracking and mapping for small ar workspaces

    Georg Klein and David Murray. Parallel tracking and mapping for small ar workspaces. In2007 6th IEEE and ACM International Symposium on Mixed and Augmented Reality, pages 225–234, 2007

  74. [74]

    Abc: A big cad model dataset for geometric deep learning

    Sebastian Koch, Albert Matveev, Zhongshi Jiang, Francis Williams, Alexey Artemov, Evgeny Burnaev, Marc Alexa, Denis Zorin, and Daniele Panozzo. Abc: A big cad model dataset for geometric deep learning. In2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9593–9603, 2019

  75. [75]

    Oneformer3d: One transformer for unified point cloud segmentation

    Maxim Kolodiazhnyi, Anna Vorontsova, Anton Konushin, and Danila Rukhovich. Oneformer3d: One transformer for unified point cloud segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20943–20953, 2024

  76. [76]

    Epipolar geometry improves video generation models, 2025

    Orest Kupyn, Fabian Manhardt, Federico Tombari, and Christian Rupprecht. Epipolar geometry improves video generation models, 2025

  77. [77]

    3d vision with transformers: A survey, 2022

    Jean Lahoud, Jiale Cao, Fahad Shahbaz Khan, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, and Ming-Hsuan Yang. 3d vision with transformers: A survey, 2022

  78. [78]

    Stratified transformer for 3d point cloud segmentation

    Xin Lai, Jianhui Liu, Li Jiang, Liwei Wang, Hengshuang Zhao, Shu Liu, Xiaojuan Qi, and Jiaya Jia. Stratified transformer for 3d point cloud segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8500–8509, 2022

  79. [79]

    Advances in 3d generation: A survey, 2024

    Xiaoyu Li, Qi Zhang, Di Kang, Weihao Cheng, Yiming Gao, Jingbo Zhang, Zhihao Liang, Jing Liao, Yan-Pei Cao, and Ying Shan. Advances in 3d generation: A survey, 2024

  80. [80]

    Megadepth: Learning single-view depth prediction from internet photos

    Zhengqi Li and Noah Snavely. Megadepth: Learning single-view depth prediction from internet photos. In Computer Vision and Pattern Recognition (CVPR), 2018

Showing first 80 references.