REVIEW 3 major objections 5 minor 9 cited by
Reconstructing 4D Spatial Intelligence: A Survey
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Five levels map how machines build 4D scenes from video
desk verdict A current, well-organized survey whose five-level taxonomy is useful as a thematic map, but the 'progressive' hierarchy is asserted, not demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the five-level taxonomy itself, defined in the introduction and used to organize every section: Level 1 low-level 3D cues, Level 2 3D scene components, Level 3 4D dynamic scenes, Level 4 interactions among components, and Level 5 physical laws and constraints. The taxonomy does the argument's work by assigning each surveyed method to a rung and by framing each section's open challenges as what must be solved before the next rung becomes reachable. Supporting machinery includes the 3D representations that appear across levels, such as neural radiance fields, 3D Gaussian splatting, signed distance functions, and parametric body models like SMPL, but these are tools rather than the survey's contribution.
What would settle it
A dependency analysis of the surveyed methods would falsify the progression claim if it found that a large share of methods at Levels 4 and 5 were developed without using outputs from Levels 1 and 2, or if the levels did not cluster in citation space.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that achieving full 4D spatial intelligence from video, capturing geometry, objects, motion, interaction, and physical behavior, is not a single problem but a layered one, and that the literature already reflects this layering. The five levels are: (1) low-level 3D cues such as depth, camera pose, point maps, and 3D tracking; (2) reconstruction of 3D scene components such as objects, humans, and structures; (3) reconstruction of dynamic 4D scenes, typically by canonical-space deformation or by adding time to the representation; (4) modeling of interactions among scene components, mostly human-centric; and (5) incorporation of physical laws and constraints so reconstructions behave plausibly under gravity, friction, and contact. The paper further claims that each level supports the next, and it ends each section by listing the challenges that block progress to the next level.
Load-bearing premise
The load-bearing premise is that the five levels really are a progression in which higher levels depend on lower ones, but the survey offers no formal dependency criterion, so if the ordering is only a convenient grouping the structural claim weakens.
Editorial extensions
If this is right
- Researchers can use the five levels as a shared coordinate system for placing new methods and for spotting which rung a paper actually advances.
- The survey identifies specific open challenges per level, including occlusions and dynamic motion at Level 1, fluids and topological change at Level 3, and physical contact at Level 4, so the gaps constitute a de facto research agenda.
- Because each level is claimed to build on the previous one, progress at lower levels, such as unified feed-forward estimation of depth, pose, and tracking, should directly accelerate the higher levels.
- The concluding discussion of a possible Level 6 implies that the hierarchy is intended to be extensible, with richer spatial intelligence beyond physics-grounded reconstruction.
Reading between the lines
- I would read the taxonomy as a claim about research dependencies rather than just a classification; if that is right, work at Level 1 has outsized downstream leverage on everything above it.
- The taxonomy predicts that near-term breakthroughs will come at the boundaries between levels, for example feed-forward systems that jump from raw video to interaction modeling without explicitly reconstructing every intermediate representation.
- A testable extension would be to derive the five levels automatically from citation or method-dependency data and compare the empirical clusters with the paper's assignments.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey organizes 4D scene reconstruction from video into five progressive levels: low-level 3D cues, 3D scene components, 4D dynamic scenes, interactions among components, and physical laws. For each level it reviews representative methods, describes paradigm architectures, and closes with challenges and future directions. The paper's central claim, stated in the abstract and Section 1, is that the five levels form a progressive hierarchy of 4D spatial intelligence, with each level building on the previous one.
Significance. The survey is timely and unusually broad: it covers roughly five hundred references, including many 2024–2025 preprints, and the level-by-level structure makes it a convenient entry point for newcomers. The paradigm figures (Figs. 2, 4, 5, 7, 8) are informative, and the maintained project page is a practical asset. The main conceptual contribution is the five-level taxonomy. Its value depends on whether 'progressive levels' is a genuine structural claim about the field; as submitted, that claim is not established, and some method assignments contradict it. With a clearly stated ordering criterion and a consistent assignment rule, the taxonomy could be a useful map of the field; without them, the contribution reduces to a topical grouping.
major comments (3)
- [Abstract and Section 1] The central claim that the five levels are 'progressive' is asserted but never defined, and the method assignments contradict a dependency reading. Level 5 entries such as DeepMimic [491], AMP [496], CALM [499], PULSE [88], ASAP [503], and UniPhys [504] in Section 6.1 are physics-based character animation and control methods trained on motion capture or reinforcement learning; they do not consume a Level 4 reconstructed interaction or a Level 3 reconstructed 4D scene. Likewise, HDM [452] and InterTrack [79] in Section 5.1 estimate human-object interaction directly from video frames without first solving Level 3 dynamic-scene reconstruction. If 'progressive' means that higher levels build on the outputs of lower levels, the taxonomy is internally inconsistent. Please either define an explicit ordering criterion (e.g., dependency, representational complexity, or task semantics) and justify each assignment against it, or reframe the contribution as a thematic organization into five topics of increasing complexity.
- [Section 6.1 and Scope] Section 6.1, 'Dynamic 4D human simulation with physics,' contains numerous methods that are not reconstructing 4D scenes from video, despite the Scope statement that the survey focuses on approaches for reconstructing 4D scenes from video inputs. DeepMimic, AMP, ASE, CALM, ControlVAE, PULSE, OmniGrasp, HOVER, ASAP, UniPhys, MaskedMimic, SuperPADL, PDP, and CLoSD learn policies from MoCap data, reinforcement learning, or text commands, not from video observations of a scene to be reconstructed. These methods may be relevant as downstream consumers of reconstructions, but as presented they address a different task (character control/synthesis). Please either narrow Level 5 to methods that operate on reconstructed 4D representations as input (e.g., PhysHOI [86], SkillMimic [481], PhysicsNeRF [93], PhyRecon [482]), or add a subsection that explicitly distinguishes reconstruction from control and justifies the inclusion of the control methods.
- [Section 1 and throughout] The paper does not provide a systematic, operational criterion for assigning a method to a level, and some assignments are hard to reconcile. For example, 3D tracking is presented in Level 1 as a low-level cue (Section 2.3), but the same capability reappears in Level 3 dynamic-reconstruction methods such as st4rtrack (Section 4.1); the boundary between Level 3 human-centric dynamic modeling (Section 4.2) and Level 4 interactions (Section 5.1) is not drawn in terms of what a method consumes or outputs. A survey taxonomy does not need a formal algorithm, but the central claim of progressivity requires a stated criterion. Please add a short 'taxonomy criteria' paragraph in Section 1 defining the level-assignment rule, and include a table that lists representative methods with their assigned levels.
minor comments (5)
- [Scope paragraph] Typo: '4D sptial intelligence' should be '4D spatial intelligence'.
- [Section 3.2] The sentence 'An overview of representative approaches in this category is shown in Fig. 2' should refer to Fig. 3; Fig. 2 is the low-level-cues paradigm figure from Section 2.
- [Section 5.1] The caption of Fig. 6 names InterDreamer, CIRCLE, and BUDDI, but InterDreamer is not discussed in the body text; please add a cross-reference or remove the name from the caption.
- [Section 2.3] The sentence about EgoPoints, 'It opens the door for future works,' is vague; please specify what future directions the new benchmark enables.
- [References] References [17] and [55] cite the same NeRF paper twice; please consolidate them into one canonical citation.
Circularity Check
No circular derivation: the paper is a survey whose five-level taxonomy organizes existing methods and is not derived from fitted inputs or load-bearing self-citations.
full rationale
This is a survey paper; its central contribution is a taxonomy (five levels of 4D spatial intelligence) imposed on existing methods. There is no quantity derived from inputs, no fitted parameter that is later 'predicted,' and no uniqueness theorem or ansatz imported from the authors' prior work. The level definitions in the Introduction (Levels 1 through 5) are organizing criteria, not outputs of a calculation; whether the levels are 'progressive' is an interpretive claim justified (or not) by the survey's coverage, not by construction. The paper cites several works by its own authors (e.g., AvatrGo [80], DreamAvatar [98], CrowdMoGen [103], EgoLM [423]), but these are ordinary survey entries describing external methods; none supplies a premise on which the taxonomy depends, and removing them would not change the five-level structure. The reader's 'weakest assumption' and the skeptic note that some methods skip levels are correctness or conceptual-coherence concerns about the taxonomy's ordering claim; they are not circularity because the paper never reduces a conclusion to its own definition or to a self-citation chain. No equations are used to derive a result from an input; equations (1) through (5) describe existing representations (3DGS parameters, SMPL kinematics, HOSNeRF field) and do not feed into the taxonomy. Accordingly, no circular steps are identified.
Assumptions & free parameters
assumptions (3)
- domain assumption The five levels form a progressive hierarchy where each level builds on the previous one.
- domain assumption The surveyed methods are representative and each method is correctly assigned to one level.
- domain assumption Existing surveys do not provide a comprehensive hierarchical analysis of 4D scene reconstruction.
Cite this review
Pith. "Pith review of Reconstructing 4D Spatial Intelligence: A Survey." pith.science (2026). https://pith.science/paper/IB3TXDV4
@misc{pith2026250721045,
author = {Pith},
title = {Pith review of: Reconstructing 4D Spatial Intelligence: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/IB3TXDV4}},
note = {Machine review of arXiv:2507.21045}
}
read the original abstract
Reconstructing 4D spatial intelligence from visual observations has long been a central yet challenging task in computer vision, with broad real-world applications. These range from entertainment domains like movies, where the focus is often on reconstructing fundamental visual elements, to embodied AI, which emphasizes interaction modeling and physical realism. Fueled by rapid advances in 3D representations and deep learning architectures, the field has evolved quickly, outpacing the scope of previous surveys. Additionally, existing surveys rarely offer a comprehensive analysis of the hierarchical structure of 4D scene reconstruction. To address this gap, we present a new perspective that organizes existing methods into five progressive levels of 4D spatial intelligence: (1) Level 1 -- reconstruction of low-level 3D attributes (e.g., depth, pose, and point maps); (2) Level 2 -- reconstruction of 3D scene components (e.g., objects, humans, structures); (3) Level 3 -- reconstruction of 4D dynamic scenes; (4) Level 4 -- modeling of interactions among scene components; and (5) Level 5 -- incorporation of physical laws and constraints. We conclude the survey by discussing the key challenges at each level and highlighting promising directions for advancing toward even richer levels of 4D spatial intelligence. To track ongoing developments, we maintain an up-to-date project page: https://github.com/yukangcao/Awesome-4D-Spatial-Intelligence.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 9 Pith papers
-
One Video, One World: Turning Monocular Video into Physical 4D Scenes
OVOW reconstructs instance-level, simulation-ready 4D mesh scenes from monocular video via a four-stage training-free pipeline and introduces a new benchmark for structured Video-to-4D evaluation.
-
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
ACE-Data-0 is a 150-hour home HOI dataset with millisecond-synced ego/exo video, mocap body/hands, object 6-DoF, audio, and tactile signals, plus a three-level benchmark exposing large SOTA gaps.
-
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.
-
Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos
HA-HOI produces physically plausible 4D HOI animations from monocular videos by anchoring object reconstruction to human motion and refining the result in a physics-based humanoid-object simulator.
-
Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation
Stitch4D reconstructs coherent 4D urban scenes from sparse non-overlapping camera placements by synthesizing bridge views and enforcing inter-location spatio-temporal consistency.
-
PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception
PAGE-4D is a feedforward extension of VGGT that uses a dynamics-aware aggregator and mask to disentangle pose estimation from geometry reconstruction in videos with moving objects.
-
PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception
A fine-tuned VGGT with a learned dynamics mask improves camera pose, depth, and point-cloud reconstruction on dynamic-scene benchmarks over the original static-scene model.
-
Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation
Synthesizing intermediate bridge views between sparse, non-overlapping urban cameras and jointly optimizing them stabilizes 4D reconstruction where dense-view methods collapse.
-
Advances in 4D Representation: Geometry, Motion, and Interaction
A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.
Reference graph
Works this paper leans on
-
[88]
Universal humanoid motion representations for physics-based control,
Z. Luo, J. Cao, J. Merel, A. Winkler, J. Huang, K. M. Kitani, and W. Xu, “Universal humanoid motion representations for physics-based control,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https: //openreview.net/forum?id=OrOd8PxOO2
2024
-
[79]
Intertrack: Tracking hu- man object interaction without object templates,
X. Xie, J. E. Lenssen, and G. Pons-Moll, “Intertrack: Tracking hu- man object interaction without object templates,” arXiv preprint arXiv:2408.13953, 2024
arXiv 2024
-
[86]
Physhoi: Physics-based imitation of dynamic human-object in- teraction,
Y. Wang, J. Lin, A. Zeng, Z. Luo, J. Zhang, and L. Zhang, “Physhoi: Physics-based imitation of dynamic human-object in- teraction,” arXiv preprint arXiv:2312.04393, 2023
arXiv 2023
-
[93]
Physicsnerf: Physics-guided 3d reconstruction from sparse views,
M. R. Barhdadi, H. Kurban, and H. Alnuweiri, “Physicsnerf: Physics-guided 3d reconstruction from sparse views,” 2025
2025
-
[1]
Neural point-based graphics,
K.-A. Aliev, A. Sevastopolsky, M. Kolos, D. Ulyanov, and V . Lem- pitsky, “Neural point-based graphics,” in European conference on computer vision. Springer, 2020, pp. 696–712
2020
-
[2]
Deep video portraits,
H. Kim, P . Garrido, A. Tewari, W. Xu, J. Thies, M. Niessner, P . P´erez, C. Richardt, M. Zollh ¨ofer, and C. Theobalt, “Deep video portraits,” ACM transactions on graphics (TOG) , vol. 37, no. 4, pp. 1–14, 2018
2018
-
[3]
Fov-nerf: Foveated neural radiance fields for virtual reality,
N. Deng, Z. He, J. Ye, B. Duinkharjav, P . Chakravarthula, X. Yang, and Q. Sun, “Fov-nerf: Foveated neural radiance fields for virtual reality,” IEEE Transactions on Visualization and Computer Graphics , vol. 28, no. 11, pp. 3854–3864, 2022
2022
-
[4]
Instant-3d: Instant neural radiance field training towards on-device ar/vr 3d reconstruction,
S. Li, C. Li, W. Zhu, B. Yu, Y. Zhao, C. Wan, H. You, H. Shi, and Y. Lin, “Instant-3d: Instant neural radiance field training towards on-device ar/vr 3d reconstruction,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , 2023, pp. 1–13
2023
Show all 300 references
-
[5]
Aligning cyber space with physical world: A comprehensive survey on embodied ai,
Y. Liu, W. Chen, Y. Bai, X. Liang, G. Li, W. Gao, and L. Lin, “Aligning cyber space with physical world: A comprehensive survey on embodied ai,” arXiv preprint arXiv:2407.06886, 2024
2024 arXiv
-
[6]
The essential role of causality in foundation world models for embodied ai,
T. Gupta, W. Gong, C. Ma, N. Pawlowski, A. Hilmkil, M. Scetbon, M. Rigter, A. Famoti, A. J. Llorens, J. Gaoet al., “The essential role of causality in foundation world models for embodied ai,” arXiv preprint arXiv:2402.06665, 2024
2024 arXiv
-
[7]
An embodied generalist agent in 3d world,
J. Huang, S. Yong, X. Ma, X. Linghu, P . Li, Y. Wang, Q. Li, S.-C. Zhu, B. Jia, and S. Huang, “An embodied generalist agent in 3d world,” arXiv preprint arXiv:2311.12871, 2023
2023 arXiv
-
[8]
3d-vla: A 3d vision-language-action generative world model,
H. Zhen, X. Qiu, P . Chen, J. Yang, X. Yan, Y. Du, Y. Hong, and C. Gan, “3d-vla: A 3d vision-language-action generative world model,” arXiv preprint arXiv:2403.09631, 2024
2024 arXiv
-
[9]
Recent advances in 3d gaussian splatting,
T. Wu, Y.-J. Yuan, L.-X. Zhang, J. Yang, Y.-P . Cao, L.-Q. Yan, and L. Gao, “Recent advances in 3d gaussian splatting,” Computa- tional Visual Media, vol. 10, no. 4, pp. 613–642, 2024
2024
-
[10]
3d gaussian splatting as new era: A survey,
B. Fei, J. Xu, R. Zhang, Q. Zhou, W. Yang, and Y. He, “3d gaussian splatting as new era: A survey,” IEEE Transactions on Visualization and Computer Graphics, 2024
2024
-
[11]
Stereo matching algorithm based on deep learning: A survey,
M. S. Hamid, N. Abd Manap, R. A. Hamzah, and A. F. Kadmin, “Stereo matching algorithm based on deep learning: A survey,” Journal of King Saud University-Computer and Information Sciences , vol. 34, no. 5, pp. 1663–1673, 2022
2022
-
[12]
Review of stereo matching algorithms based on deep learning,
K. Zhou, X. Meng, and B. Cheng, “Review of stereo matching algorithms based on deep learning,” Computational intelligence and neuroscience, vol. 2020, no. 1, p. 8562323, 2020
2020
-
[13]
A sur- vey on deep learning techniques for stereo-based depth estima- tion,
H. Laga, L. V . Jospin, F. Boussaid, and M. Bennamoun, “A sur- vey on deep learning techniques for stereo-based depth estima- tion,” IEEE transactions on pattern analysis and machine intelligence , vol. 44, no. 4, pp. 1738–1764, 2020
2020
-
[14]
Nerf: Neural radiance field in 3d vision, a comprehensive review,
K. Gao, Y. Gao, H. He, D. Lu, L. Xu, and J. Li, “Nerf: Neural radiance field in 3d vision, a comprehensive review,” arXiv preprint arXiv:2210.00379, 2022
2022 arXiv
-
[15]
3d gaussian splatting: Survey, technologies, challenges, and opportunities,
Y. Bao, T. Ding, J. Huo, Y. Liu, Y. Li, W. Li, Y. Gao, and J. Luo, “3d gaussian splatting: Survey, technologies, challenges, and opportunities,” IEEE Transactions on Circuits and Systems for Video Technology, 2025
2025
-
[16]
Nerf in robotics: A survey,
G. Wang, L. Pan, S. Peng, S. Liu, C. Xu, Y. Miao, W. Zhan, M. Tomizuka, M. Pollefeys, and H. Wang, “Nerf in robotics: A survey,” arXiv preprint arXiv:2405.01333, 2024
2024 arXiv
-
[17]
Nerf: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P . P . Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in European conference on computer vision. Springer, 2020, pp. 405–421
2020
-
[18]
Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape synthesis,
T. Shen, J. Gao, K. Yin, M.-Y. Liu, and S. Fidler, “Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape synthesis,” 2021
2021
-
[19]
3d gaussian splatting for real-time radiance field rendering,
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,” ACM Transactions on Graphics (ToG), vol. 42, no. 4, pp. 1–14, 2023
2023
-
[20]
Video diffusion models,
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” Advances in Neural Information Processing Systems, vol. 35, pp. 8633–8646, 2022
2022
-
[21]
Imagen video: High definition video generation with diffusion models,
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P . Kingma, B. Poole, M. Norouzi, D. J. Fleet et al., “Imagen video: High definition video generation with diffusion models,” arXiv preprint arXiv:2210.02303, 2022
-
[22]
Stable video diffusion: Scaling latent video diffusion models to large datasets,
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V . Voleti, A. Letts et al. , “Stable video diffusion: Scaling latent video diffusion models to large datasets,” arXiv preprint arXiv:2311.15127, 2023
2023 arXiv
-
[23]
Sift: Predicting amino acid changes that affect protein function,
P . C. Ng and S. Henikoff, “Sift: Predicting amino acid changes that affect protein function,” Nucleic acids research, vol. 31, no. 13, pp. 3812–3814, 2003
2003
-
[24]
R2d2: Reliable and repeatable detector and descriptor,
J. Revaud, C. De Souza, M. Humenberger, and P . Weinzaepfel, “R2d2: Reliable and repeatable detector and descriptor,”Advances in neural information processing systems, vol. 32, 2019
2019
-
[25]
Superpoint: Self- supervised interest point detection and description,
D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superpoint: Self- supervised interest point detection and description,” in Proceed- ings of the IEEE conference on computer vision and pattern recognition workshops, 2018, pp. 224–236
2018
-
[26]
Superglue: Learning feature matching with graph neural net- works,
P .-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superglue: Learning feature matching with graph neural net- works,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 4938–4947
2020
-
[27]
Loftr: Detector- free local feature matching with transformers,
J. Sun, Z. Shen, Y. Wang, H. Bao, and X. Zhou, “Loftr: Detector- free local feature matching with transformers,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 8922–8931
2021
-
[28]
Neural-guided ransac: Learn- ing where to sample model hypotheses,
E. Brachmann and C. Rother, “Neural-guided ransac: Learn- ing where to sample model hypotheses,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 4322–4331
2019
-
[29]
Lightglue: Local feature matching at light speed,
P . Lindenberger, P .-E. Sarlin, and M. Pollefeys, “Lightglue: Local feature matching at light speed,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 17 627–17 638
2023
-
[30]
Affineglue: Joint matching and robust estimation,
D. Barath, D. Mishkin, L. Cavalli, P .-E. Sarlin, P . Hruby, and M. Pollefeys, “Affineglue: Joint matching and robust estimation,” arXiv preprint arXiv:2307.15381, 2023
2023 arXiv
-
[31]
Structure-from-motion revis- ited,
J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revis- ited,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 4104–4113
2016
-
[32]
A survey of structure from motion*
O. ¨Ozyes ¸il, V . Voroninski, R. Basri, and A. Singer, “A survey of structure from motion*.” Acta Numerica, vol. 26, pp. 305–364, 2017
2017
-
[33]
Structure from motion photogrammetry in forestry: A review,
J. Iglhaut, C. Cabo, S. Puliti, L. Piermattei, J. O’Connor, and J. Rosette, “Structure from motion photogrammetry in forestry: A review,” Current Forestry Reports , vol. 5, no. 3, pp. 155–168, 2019
2019
-
[34]
Pixel- perfect structure-from-motion with featuremetric refinement,
P . Lindenberger, P .-E. Sarlin, V . Larsson, and M. Pollefeys, “Pixel- perfect structure-from-motion with featuremetric refinement,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5987–5997
2021
-
[35]
Bundle adjustment in the large,
S. Agarwal, N. Snavely, S. M. Seitz, and R. Szeliski, “Bundle adjustment in the large,” in European conference on computer vision. Springer, 2010, pp. 29–42
2010
-
[36]
Bundle adjustment rules,
C. Engels, H. Stew ´enius, and D. Nist ´er, “Bundle adjustment rules,” Photogrammetric computer vision, vol. 2, no. 32, 2006
2006
-
[37]
Robust bundle adjustment revisited,
C. Zach, “Robust bundle adjustment revisited,” in European Con- ference on Computer Vision. Springer, 2014, pp. 772–787
2014
-
[38]
Bundle adjustment—a modern synthesis,
B. Triggs, P . F. McLauchlan, R. I. Hartley, and A. W. Fitzgibbon, “Bundle adjustment—a modern synthesis,” in International work- shop on vision algorithms. Springer, 1999, pp. 298–372
1999
-
[39]
Pixelwise view selection for unstructured multi-view stereo,
J. L. Sch ¨onberger, E. Zheng, M. Pollefeys, and J.-M. Frahm, “Pixelwise view selection for unstructured multi-view stereo,” in European Conference on Computer Vision (ECCV), 2016. 16
2016
-
[40]
Visibility-aware multi-view stereo network,
J. Zhang, Y. Yao, S. Li, Z. Luo, and T. Fang, “Visibility-aware multi-view stereo network,” British Machine Vision Conference (BMVC), 2020
2020
-
[41]
Cost volume pyramid based depth inference for multi-view stereo,
J. Yang, W. Mao, J. M. Alvarez, and M. Liu, “Cost volume pyramid based depth inference for multi-view stereo,” in The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020
2020
-
[42]
Patchmatchnet: Learned multi-view patchmatch stereo,
F. Wang, S. Galliani, C. Vogel, P . Speciale, and M. Pollefeys, “Patchmatchnet: Learned multi-view patchmatch stereo,” 2021
2021
-
[43]
Cascade cost volume for high-resolution multi-view stereo and stereo matching,
X. Gu, Z. Fan, S. Zhu, Z. Dai, F. Tan, and P . Tan, “Cascade cost volume for high-resolution multi-view stereo and stereo matching,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2495–2504
2020
-
[44]
Dust3r: Geometric 3d vision made easy,
S. Wang, V . Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 20 697–20 709
2024
-
[45]
Mast3r-slam: Real- time dense slam with 3d reconstruction priors,
R. Murai, E. Dexheimer, and A. J. Davison, “Mast3r-slam: Real- time dense slam with 3d reconstruction priors,” arXiv preprint arXiv:2412.12392, 2024
2024 arXiv
-
[46]
Monst3r: A simple approach for estimating geometry in the presence of motion,
J. Zhang, C. Herrmann, J. Hur, V . Jampani, T. Darrell, F. Cole, D. Sun, and M.-H. Yang, “Monst3r: A simple approach for estimating geometry in the presence of motion,” arXiv preprint arXiv:2410.03825, 2024
2024 arXiv
-
[47]
Align3r: Aligned monocular depth estimation for dynamic videos,
J. Lu, T. Huang, P . Li, Z. Dou, C. Lin, Z. Cui, Z. Dong, S.-K. Yeung, W. Wang, and Y. Liu, “Align3r: Aligned monocular depth estimation for dynamic videos,” arXiv preprint arXiv:2412.03079 , 2024
2024 arXiv
-
[48]
Fast3r: Towards 3d recon- struction of 1000+ images in one forward pass,
J. Yang, A. Sax, K. J. Liang, M. Henaff, H. Tang, A. Cao, J. Chai, F. Meier, and M. Feiszli, “Fast3r: Towards 3d recon- struction of 1000+ images in one forward pass,” arXiv preprint arXiv:2501.13928, 2025
2025 arXiv
-
[49]
Transformer in transformer,
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang, “Transformer in transformer,” Advances in neural information processing systems , vol. 34, pp. 15 908–15 919, 2021
2021
-
[50]
A survey on vision transformer,
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xuet al., “A survey on vision transformer,”IEEE transactions on pattern analysis and machine intelligence , vol. 45, no. 1, pp. 87–110, 2022
2022
-
[51]
Reformer: The efficient transformer,
N. Kitaev, Ł. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” arXiv preprint arXiv:2001.04451, 2020
2001 arXiv
-
[52]
Point trans- former,
H. Zhao, L. Jiang, J. Jia, P . H. Torr, and V . Koltun, “Point trans- former,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 16 259–16 268
2021
-
[53]
Image transformer,
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran, “Image transformer,” in International conference on machine learning. PMLR, 2018, pp. 4055–4064
2018
-
[54]
Vggt: Visual geometry grounded transformer,
J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” in Proceedings of the Computer Vision and Pattern Recognition Confer- ence, 2025, pp. 5294–5306
2025
-
[55]
Nerf: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P . P . Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021
2021
-
[56]
3d gaus- sian splatting for real-time radiance field rendering,
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaus- sian splatting for real-time radiance field rendering,” ACM Trans- actions on Graphics , vol. 42, no. 4, July 2023. [Online]. Available: https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/
2023
-
[57]
Flexible isosurface extraction for gradient-based mesh optimization,
T. Shen, J. Munkberg, J. Hasselgren, K. Yin, Z. Wang, W. Chen, Z. Gojcic, S. Fidler, N. Sharp, and J. Gao, “Flexible isosurface extraction for gradient-based mesh optimization,” ACM Trans. Graph. , vol. 42, no. 4, jul 2023. [Online]. Available: https://doi.org/10.1145/3592430
2023 doi
-
[58]
Nerfies: Deformable neural radi- ance fields,
K. Park, U. Sinha, J. T. Barron, S. Bouaziz, D. B. Goldman, S. M. Seitz, and R. Martin-Brualla, “Nerfies: Deformable neural radi- ance fields,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 5865–5874
2021
-
[59]
Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monocular video,
E. Tretschk, A. Tewari, V . Golyanik, M. Zollh ¨ofer, C. Lassner, and C. Theobalt, “Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monocular video,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2...
2021
-
[60]
Hypernerf: A higher-dimensional representation for topologically varying neu- ral radiance fields,
K. Park, U. Sinha, P . Hedman, J. T. Barron, S. Bouaziz, D. B. Goldman, R. Martin-Brualla, and S. M. Seitz, “Hypernerf: A higher-dimensional representation for topologically varying neu- ral radiance fields,” arXiv preprint arXiv:2106.13228, 2021
2021 arXiv
-
[61]
Dylin: Making light field networks dynamic,
H. Yu, J. Julin, Z. A. Milacski, K. Niinuma, and L. A. Jeni, “Dylin: Making light field networks dynamic,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 12 397–12 406
2023
-
[62]
Spectromotion: Dynamic 3d reconstruction of specular scenes,
C.-D. Fan, C.-W. Chang, Y.-R. Liu, J.-Y. Lee, J.-L. Huang, Y.-C. Tseng, and Y.-L. Liu, “Spectromotion: Dynamic 3d reconstruction of specular scenes,” arXiv preprint arXiv:2410.17249, 2024
2024 arXiv
-
[63]
Neural scene flow fields for space-time view synthesis of dynamic scenes,
Z. Li, S. Niklaus, N. Snavely, and O. Wang, “Neural scene flow fields for space-time view synthesis of dynamic scenes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 6498–6508
2021
-
[64]
Dynamic view synthesis from dynamic monocular video,
C. Gao, A. Saraf, J. Kopf, and J.-B. Huang, “Dynamic view synthesis from dynamic monocular video,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 5712–5721
2021
-
[65]
Decoupling dynamic monocular videos for dynamic view synthesis,
M. You and J. Hou, “Decoupling dynamic monocular videos for dynamic view synthesis,” IEEE Transactions on Visualization and Computer Graphics, 2024
2024
-
[66]
Neural trajec- tory fields for dynamic novel view synthesis,
C. Wang, B. Eckart, S. Lucey, and O. Gallo, “Neural trajec- tory fields for dynamic novel view synthesis,” arXiv preprint arXiv:2105.05994, 2021
2021 arXiv
-
[67]
Neural scene chronology,
H. Lin, Q. Wang, R. Cai, S. Peng, H. Averbuch-Elor, X. Zhou, and N. Snavely, “Neural scene chronology,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 20 752–20 761
2023
-
[68]
Darenerf: Direction-aware representation for dynamic scenes,
A. Lou, B. Planche, Z. Gao, Y. Li, T. Luan, H. Ding, T. Chen, J. Noble, and Z. Wu, “Darenerf: Direction-aware representation for dynamic scenes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5031–5042
2024
-
[69]
Emernerf: Emer- gent spatial-temporal scene decomposition via self-supervision,
J. Yang, B. Ivanovic, O. Litany, X. Weng, S. W. Kim, B. Li, T. Che, D. Xu, S. Fidler, M. Pavone et al. , “Emernerf: Emer- gent spatial-temporal scene decomposition via self-supervision,” arXiv preprint arXiv:2311.02077, 2023
2023 arXiv
-
[70]
Gravity-aware monocular 3d human-object reconstruction,
R. Dabral, S. Shimada, A. Jain, C. Theobalt, and V . Golyanik, “Gravity-aware monocular 3d human-object reconstruction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 12 365–12 374
2021
-
[71]
D3d-hoi: Dynamic 3d human-object interactions from videos,
X. Xu, H. Joo, G. Mori, and M. Savva, “D3d-hoi: Dynamic 3d human-object interactions from videos,” arXiv preprint arXiv:2108.08420, 2021
2021 arXiv
-
[72]
Behave: Dataset and method for tracking human object interactions,
B. L. Bhatnagar, X. Xie, I. A. Petrov, C. Sminchisescu, C. Theobalt, and G. Pons-Moll, “Behave: Dataset and method for tracking human object interactions,” in Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , 2022, pp. 15 935– 15 946
2022
-
[73]
Intercap: Joint markerless 3d tracking of humans and objects in interaction,
Y. Huang, O. Taheri, M. J. Black, and D. Tzionas, “Intercap: Joint markerless 3d tracking of humans and objects in interaction,” in DAGM German Conference on Pattern Recognition. Springer, 2022, pp. 281–299
2022
-
[74]
Full-body articulated human-object inter- action,
N. Jiang, T. Liu, Z. Cao, J. Cui, Z. Zhang, Y. Chen, H. Wang, Y. Zhu, and S. Huang, “Full-body articulated human-object inter- action,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 9365–9376
2023
-
[75]
Stackflow: Monocular human-object reconstruction by stacked normalizing flow with offset,
C. Huo, Y. Shi, Y. Ma, L. Xu, J. Yu, and J. Wang, “Stackflow: Monocular human-object reconstruction by stacked normalizing flow with offset,” arXiv preprint arXiv:2407.20545, 2024
2024 arXiv
-
[76]
Monocular human-object recon- struction in the wild,
C. Huo, Y. Shi, and J. Wang, “Monocular human-object recon- struction in the wild,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 5547–5555
2024
-
[77]
I’m hoi: Inertia-aware monocular capture of 3d human- object interactions,
C. Zhao, J. Zhang, J. Du, Z. Shan, J. Wang, J. Yu, J. Wang, and L. Xu, “I’m hoi: Inertia-aware monocular capture of 3d human- object interactions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 729–741
2024
-
[78]
Sv4d: Dynamic 3d content generation with multi-frame and multi-view consistency,
Y. Xie, C.-H. Yao, V . Voleti, H. Jiang, and V . Jampani, “Sv4d: Dynamic 3d content generation with multi-frame and multi-view consistency,” arXiv preprint arXiv:2407.17470, 2024
2024 arXiv
-
[80]
Avatargo: Zero-shot 4d human-object interaction generation and anima- tion,
Y. Cao, L. Pan, K. Han, K.-Y. K. Wong, and Z. Liu, “Avatargo: Zero-shot 4d human-object interaction generation and anima- tion,” arXiv preprint arXiv:2410.07164, 2024
2024 arXiv
-
[81]
The one where they reconstructed 3d humans and environments in tv shows,
G. Pavlakos, E. Weber, M. Tancik, and A. Kanazawa, “The one where they reconstructed 3d humans and environments in tv shows,” in European Conference on Computer Vision . Springer, 2022, pp. 732–749. 17
2022
-
[82]
Joint optimization for 4d human-scene reconstruction in the wild,
Z. Liu, J. Lin, W. Wu, and B. Zhou, “Joint optimization for 4d human-scene reconstruction in the wild,” arXiv preprint arXiv:2501.02158, 2025
2025 arXiv
-
[83]
Odhsr: Online dense 3d reconstruction of humans and scenes from monocular videos,
Z. Zhang, M. Kaufmann, L. Xue, J. Song, and M. R. Oswald, “Odhsr: Online dense 3d reconstruction of humans and scenes from monocular videos,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 824–21 835
2025
-
[84]
Hosnerf: Dynamic human-object-scene neu- ral radiance fields from a single video,
J.-W. Liu, Y.-P . Cao, T. Yang, Z. Xu, J. Keppo, Y. Shan, X. Qie, and M. Z. Shou, “Hosnerf: Dynamic human-object-scene neu- ral radiance fields from a single video,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 18 483–18 494
2023
-
[85]
Neuman: Neural human radiance field from a single video,
W. Jiang, K. M. Yi, G. Samei, O. Tuzel, and A. Ranjan, “Neuman: Neural human radiance field from a single video,” in European Conference on Computer Vision. Springer, 2022, pp. 402–418
2022
-
[87]
Perpetual humanoid control for real-time simulated avatars,
Z. Luo, J. Cao, A. W. Winkler, K. Kitani, and W. Xu, “Perpetual humanoid control for real-time simulated avatars,” in Interna- tional Conference on Computer Vision (ICCV), 2023
2023
-
[89]
Isaac gym: High performance gpu-based physics simulation for robot learning,
V . Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa et al. , “Isaac gym: High performance gpu-based physics simulation for robot learning,” arXiv preprint arXiv:2108.10470, 2021
2021 arXiv
-
[90]
Reinforcement learning,
M. A. Wiering and M. Van Otterlo, “Reinforcement learning,” Adaptation, learning, and optimization, vol. 12, no. 3, p. 729, 2012
2012
-
[91]
Reinforcement learning: A survey,
L. P . Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” Journal of artificial intelligence research , vol. 4, pp. 237–285, 1996
1996
-
[92]
Reinforcement learning,
R. S. Sutton, A. G. Barto et al., “Reinforcement learning,” Journal of Cognitive Neuroscience, vol. 11, no. 1, pp. 126–134, 1999
1999
-
[94]
Pbr-nerf: Inverse rendering with physics-based neural fields,
S. Wu, S. Basu, T. Broedermann, L. V . Gool, and C. Sakaridis, “Pbr-nerf: Inverse rendering with physics-based neural fields,” 2025
2025
-
[95]
Cast: Component-aligned 3d scene reconstruction from an rgb image,
K. Yao, L. Zhang, X. Yan, Y. Zeng, Q. Zhang, W. Yang, L. Xu, J. Gu, and J. Yu, “Cast: Component-aligned 3d scene reconstruction from an rgb image,” 2025
2025
-
[96]
Sv3d: Novel multi- view synthesis and 3d generation from a single image using latent video diffusion,
V . Voleti, C.-H. Yao, M. Boss, A. Letts, D. Pankratz, D. Tochilkin, C. Laforte, R. Rombach, and V . Jampani, “Sv3d: Novel multi- view synthesis and 3d generation from a single image using latent video diffusion,” in European Conference on Computer Vision. Springer, 2025, pp. 439–457
2025
-
[97]
V3d: Video diffusion models are effective 3d generators,
Z. Chen, Y. Wang, F. Wang, Z. Wang, and H. Liu, “V3d: Video diffusion models are effective 3d generators,” arXiv preprint arXiv:2403.06738, 2024
2024 arXiv
-
[98]
Drea- mavatar: Text-and-shape guided 3d human avatar generation via diffusion models,
Y. Cao, Y.-P . Cao, K. Han, Y. Shan, and K.-Y. K. Wong, “Drea- mavatar: Text-and-shape guided 3d human avatar generation via diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 958–968
2024
-
[99]
4d-fy: Text-to-4d generation using hybrid score distillation sampling,
S. Bahmani, I. Skorokhodov, V . Rong, G. Wetzstein, L. Guibas, P . Wonka, S. Tulyakov, J. J. Park, A. Tagliasacchi, and D. B. Lin- dell, “4d-fy: Text-to-4d generation using hybrid score distillation sampling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...
2024
-
[100]
Tc4d: Trajectory- conditioned text-to-4d generation,
S. Bahmani, X. Liu, W. Yifan, I. Skorokhodov, V . Rong, Z. Liu, X. Liu, J. J. Park, S. Tulyakov, G. Wetzsteinet al., “Tc4d: Trajectory- conditioned text-to-4d generation,” in European Conference on Computer Vision. Springer, 2024, pp. 53–72
2024
-
[101]
4diffusion: Multi-view video diffusion model for 4d genera- tion,
H. Zhang, X. Chen, Y. Wang, X. Liu, Y. Wang, and Y. Qiao, “4diffusion: Multi-view video diffusion model for 4d genera- tion,” Advances in Neural Information Processing Systems , vol. 37, pp. 15 272–15 295, 2024
2024
-
[102]
Diffusion4d: Fast spatial-temporal consis- tent 4d generation via video diffusion models,
H. Liang, Y. Yin, D. Xu, H. Liang, Z. Wang, K. N. Plataniotis, Y. Zhao, and Y. Wei, “Diffusion4d: Fast spatial-temporal consis- tent 4d generation via video diffusion models,” arXiv preprint arXiv:2405.16645, 2024
2024 arXiv
-
[103]
Crowdmo- gen: Zero-shot text-driven collective motion generation,
Y. Cao, X. Guo, M. Zhang, H. Xie, C. Gu, and Z. Liu, “Crowdmo- gen: Zero-shot text-driven collective motion generation,” arXiv preprint arXiv:2407.06188, 2024
2024 arXiv
-
[104]
Guide3d: Create 3d avatars from text and image guidance,
Y. Cao, Y.-P . Cao, K. Han, Y. Shan, and K.-Y. K. Wong, “Guide3d: Create 3d avatars from text and image guidance,” arXiv preprint arXiv:2308.09705, 2023
2023 arXiv
-
[105]
Advances in 4d generation: A survey,
Q. Miao, K. Li, J. Quan, Z. Min, S. Ma, Y. Xu, Y. Yang, and Y. Luo, “Advances in 4d generation: A survey,” arXiv preprint arXiv:2503.14501, 2025
2025 arXiv
-
[106]
Generative ai meets 3d: A survey on text-to-3d in aigc era,
C. Li, C. Zhang, A. Waghwase, L.-H. Lee, F. Rameau, Y. Yang, S.-H. Bae, and C. S. Hong, “Generative ai meets 3d: A survey on text-to-3d in aigc era,” arXiv preprint arXiv:2305.06131, 2023
2023 arXiv
-
[107]
A comprehensive survey on 3d content generation,
J. Liu, X. Huang, T. Huang, L. Chen, Y. Hou, S. Tang, Z. Liu, W. Ouyang, W. Zuo, J. Jiang et al., “A comprehensive survey on 3d content generation,” arXiv preprint arXiv:2402.01166, 2024
2024 arXiv
-
[108]
Advances in 3d generation: A survey,
X. Li, Q. Zhang, D. Kang, W. Cheng, Y. Gao, J. Zhang, Z. Liang, J. Liao, Y.-P . Cao, and Y. Shan, “Advances in 3d generation: A survey,” arXiv preprint arXiv:2401.17807, 2024
2024 arXiv
-
[109]
Differentiable rendering: A survey,
H. Kato, D. Beker, M. Morariu, T. Ando, T. Matsuoka, W. Kehl, and A. Gaidon, “Differentiable rendering: A survey,” arXiv preprint arXiv:2006.12057, 2020
2006 arXiv
-
[110]
A survey of non- rigid 3d registration,
B. Deng, Y. Yao, R. M. Dyke, and J. Zhang, “A survey of non- rigid 3d registration,” in Computer Graphics Forum, vol. 41, no. 2. Wiley Online Library, 2022, pp. 559–589
2022
-
[111]
Inference time optimization using branchynet partitioning,
R. G. Pacheco and R. S. Couto, “Inference time optimization using branchynet partitioning,” in 2020 IEEE Symposium on Computers and Communications (ISCC). IEEE, 2020, pp. 1–6
2020
-
[112]
Geonet: Unsupervised learning of dense depth, optical flow and camera pose,
Z. Yin and J. Shi, “Geonet: Unsupervised learning of dense depth, optical flow and camera pose,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 1983–1992
2018
-
[113]
Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras,
A. Gordon, H. Li, R. Jonschkowski, and A. Angelova, “Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras,” in Proceedings of the IEEE/CVF interna- tional conference on computer vision, 2019, pp. 8977–8986
2019
-
[114]
Unsupervised scale-consistent depth and ego-motion learning from monocular video,
J. Bian, Z. Li, N. Wang, H. Zhan, C. Shen, M.-M. Cheng, and I. Reid, “Unsupervised scale-consistent depth and ego-motion learning from monocular video,” Advances in neural information processing systems, vol. 32, 2019
2019
-
[115]
Consis- tent video depth estimation,
X. Luo, J.-B. Huang, R. Szeliski, K. Matzen, and J. Kopf, “Consis- tent video depth estimation,” ACM Transactions on Graphics (ToG), vol. 39, no. 4, pp. 71–1, 2020
2020
-
[116]
Self-supervised learn- ing with geometric constraints in monocular video: Connecting flow, depth, and camera,
Y. Chen, C. Schmid, and C. Sminchisescu, “Self-supervised learn- ing with geometric constraints in monocular video: Connecting flow, depth, and camera,” in Proceedings of the IEEE/CVF interna- tional conference on computer vision, 2019, pp. 7063–7072
2019
-
[117]
Multi-view depth estimation using epipolar spatio-temporal networks,
X. Long, L. Liu, W. Li, C. Theobalt, and W. Wang, “Multi-view depth estimation using epipolar spatio-temporal networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 8258–8267
2021
-
[118]
Multi-frame self-supervised depth with transformers,
V . Guizilini, R. Ambrus , , D. Chen, S. Zakharov, and A. Gaidon, “Multi-frame self-supervised depth with transformers,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 160–170
2022
-
[119]
The temporal opportunist: Self-supervised multi-frame monocular depth,
J. Watson, O. Mac Aodha, V . Prisacariu, G. Brostow, and M. Fir- man, “The temporal opportunist: Self-supervised multi-frame monocular depth,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 1164–1174
2021
-
[120]
Simplerecon: 3d reconstruction without 3d convolu- tions,
M. Sayed, J. Gibson, J. Watson, V . Prisacariu, M. Firman, and C. Godard, “Simplerecon: 3d reconstruction without 3d convolu- tions,” in European Conference on Computer Vision. Springer, 2022, pp. 1–19
2022
-
[121]
Video depth es- timation by fusing flow-to-depth proposals,
J. Xie, C. Lei, Z. Li, L. E. Li, and Q. Chen, “Video depth es- timation by fusing flow-to-depth proposals,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2020, pp. 10 100–10 107
2020
-
[122]
Temporally consistent depth prediction with flow-guided memory units,
C. Eom, H. Park, and B. Ham, “Temporally consistent depth prediction with flow-guided memory units,” IEEE Transactions on Intelligent Transportation Systems, vol. 21, no. 11, pp. 4626–4636, 2019
2019
-
[123]
Don’t forget the past: Recurrent depth estimation from monocular video,
V . Patil, W. Van Gansbeke, D. Dai, and L. Van Gool, “Don’t forget the past: Recurrent depth estimation from monocular video,” IEEE Robotics and Automation Letters , vol. 5, no. 4, pp. 6813–6820, 2020
2020
-
[124]
Exploiting temporal consistency for real-time video depth estimation,
H. Zhang, C. Shen, Y. Li, Y. Cao, Y. Liu, and Y. Yan, “Exploiting temporal consistency for real-time video depth estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1725–1734
2019
-
[125]
Less is more: Consistent video depth estimation with masked frames 18 modeling,
Y. Wang, Z. Pan, X. Li, Z. Cao, K. Xian, and J. Zhang, “Less is more: Consistent video depth estimation with masked frames 18 modeling,” in Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 6347–6358
2022
-
[126]
Mamo: Leveraging memory and attention for monocular video depth estimation,
R. Yasarla, H. Cai, J. Jeong, Y. Shi, R. Garrepalli, and F. Porikli, “Mamo: Leveraging memory and attention for monocular video depth estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 8754–8764
2023
-
[127]
Monovit: Self-supervised monocular depth estimation with a vision transformer,
C. Zhao, Y. Zhang, M. Poggi, F. Tosi, X. Guo, Z. Zhu, G. Huang, Y. Tang, and S. Mattoccia, “Monovit: Self-supervised monocular depth estimation with a vision transformer,” in 2022 international conference on 3D vision (3DV). IEEE, 2022, pp. 668–678
2022
-
[128]
Temporally consistent online depth estima- tion in dynamic scenes,
Z. Li, W. Ye, D. Wang, F. X. Creighton, R. H. Taylor, G. Venkatesh, and M. Unberath, “Temporally consistent online depth estima- tion in dynamic scenes,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision , 2023, pp. 3018–3027
2023
-
[129]
Deepv2d: Video to depth with differ- entiable structure from motion,
Z. Teed and J. Deng, “Deepv2d: Video to depth with differ- entiable structure from motion,” in International Conference on Learning Representations, 2020
2020
-
[130]
Neural video depth stabilizer,
Y. Wang, M. Shi, J. Li, Z. Huang, Z. Cao, J. Zhang, K. Xian, and G. Lin, “Neural video depth stabilizer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 9466–9476
2023
-
[131]
Nvds +: Towards efficient and versatile neural stabilizer for video depth estimation,
Y. Wang, M. Shi, J. Li, C. Hong, Z. Huang, J. Peng, Z. Cao, J. Zhang, K. Xian, and G. Lin, “Nvds +: Towards efficient and versatile neural stabilizer for video depth estimation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , pp. 1–18, 2024
2024
-
[132]
Depthcrafter: Generating consistent long depth se- quences for open-world videos,
W. Hu, X. Gao, X. Li, S. Zhao, X. Cun, Y. Zhang, L. Quan, and Y. Shan, “Depthcrafter: Generating consistent long depth se- quences for open-world videos,” arXiv preprint arXiv:2409.02095, 2024
2024 arXiv
-
[133]
Learning temporally consistent video depth from video diffusion priors,
J. Shao, Y. Yang, H. Zhou, Y. Zhang, Y. Shen, M. Poggi, and Y. Liao, “Learning temporally consistent video depth from video diffusion priors,” arXiv preprint arXiv:2406.01493, 2024
2024 arXiv
-
[134]
Depth any video with scalable synthetic data,
H. Yang, D. Huang, W. Yin, C. Shen, H. Liu, X. He, B. Lin, W. Ouyang, and T. He, “Depth any video with scalable synthetic data,” arXiv preprint arXiv:2410.10815, 2024
2024 arXiv
-
[135]
Video depth anything: Consistent depth estimation for super- long videos,
S. Chen, H. Guo, S. Zhu, F. Zhang, Z. Huang, J. Feng, and B. Kang, “Video depth anything: Consistent depth estimation for super- long videos,” arXiv preprint arXiv:2501.12375, 2025
2025 arXiv
-
[136]
Depth anything v2,
L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” arXiv preprint arXiv:2406.09414, 2024
2024 arXiv
-
[137]
Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,
R. Mur-Artal and J. D. Tard ´os, “Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,” IEEE transactions on robotics, vol. 33, no. 5, pp. 1255–1262, 2017
2017
-
[138]
Line flow based simultaneous localization and mapping,
Q. Wang, Z. Yan, J. Wang, F. Xue, W. Ma, and H. Zha, “Line flow based simultaneous localization and mapping,” IEEE Transactions on Robotics, vol. 37, no. 5, pp. 1416–1432, 2021
2021
-
[139]
Stereo visual odometry with deep learning-based point and line feature matching using an attention graph neural network,
S. Kannapiran, N. Bendapudi, M.-Y. Yu, D. Parikh, S. Berman, A. Vora, and G. Pandey, “Stereo visual odometry with deep learning-based point and line feature matching using an attention graph neural network,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and S...
2023
-
[140]
Structure plp-slam: Efficient sparse mapping and localization using point, line and plane for monocular, rgb-d and stereo cameras,
F. Shu, J. Wang, A. Pagani, and D. Stricker, “Structure plp-slam: Efficient sparse mapping and localization using point, line and plane for monocular, rgb-d and stereo cameras,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 2105–2112
2023
-
[141]
Ul-slam: A universal monocular line-based slam via unifying structural and non- structural constraints,
H. Jiang, R. Qian, L. Du, J. Pu, and J. Feng, “Ul-slam: A universal monocular line-based slam via unifying structural and non- structural constraints,” IEEE Transactions on Automation Science and Engineering, 2024
2024
-
[142]
Lsd-slam: Large-scale direct monocular slam,
J. Engel, T. Sch ¨ops, and D. Cremers, “Lsd-slam: Large-scale direct monocular slam,” in European conference on computer vision. Springer, 2014, pp. 834–849
2014
-
[143]
Direct sparse odome- try,
J. Engel, V . Koltun, and D. Cremers, “Direct sparse odome- try,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 3, pp. 611–625, 2017
2017
-
[144]
Edplvo: Efficient direct point-line visual odometry,
L. Zhou, G. Huang, Y. Mao, S. Wang, and M. Kaess, “Edplvo: Efficient direct point-line visual odometry,” in 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2022, pp. 7559–7565
2022
-
[145]
D3vo: Deep depth, deep pose and deep uncertainty for monocular visual odometry,
N. Yang, L. v. Stumberg, R. Wang, and D. Cremers, “D3vo: Deep depth, deep pose and deep uncertainty for monocular visual odometry,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1281–1292
2020
-
[146]
Feature warping for robust speaker verification,
J. Pelecanos and S. Sridharan, “Feature warping for robust speaker verification,” in Proceedings of 2001 A speaker Odyssey: the speaker recognition workshop . European Speech Communication Association, 2001, pp. 213–218
2001
-
[147]
Tartanvo: A generalizable learning-based vo,
W. Wang, Y. Hu, and S. Scherer, “Tartanvo: A generalizable learning-based vo,” in Conference on Robot Learning. PMLR, 2021, pp. 1761–1772
2021
-
[148]
Deepvo: Towards end-to-end visual odometry with deep recurrent convolutional neural networks,
S. Wang, R. Clark, H. Wen, and N. Trigoni, “Deepvo: Towards end-to-end visual odometry with deep recurrent convolutional neural networks,” in 2017 IEEE international conference on robotics and automation (ICRA). IEEE, 2017, pp. 2043–2050
2017
-
[149]
Dytanvo: Joint refine- ment of visual odometry and motion segmentation in dynamic environments,
S. Shen, Y. Cai, W. Wang, and S. Scherer, “Dytanvo: Joint refine- ment of visual odometry and motion segmentation in dynamic environments,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 4048–4055
2023
-
[150]
Deep direct visual odometry,
C. Zhao, Y. Tang, Q. Sun, and A. V . Vasilakos, “Deep direct visual odometry,” IEEE transactions on intelligent transportation systems , vol. 23, no. 7, pp. 7733–7742, 2021
2021
-
[151]
Deep patch visual odometry,
Z. Teed, L. Lipson, and J. Deng, “Deep patch visual odometry,” Advances in Neural Information Processing Systems , vol. 36, pp. 39 033–39 051, 2023
2023
-
[152]
Airslam: An efficient and illumination-robust point-line visual slam system,
K. Xu, Y. Hao, S. Yuan, C. Wang, and L. Xie, “Airslam: An efficient and illumination-robust point-line visual slam system,” IEEE Transactions on Robotics, 2025
2025
-
[153]
Deep patch visual slam,
L. Lipson, Z. Teed, and J. Deng, “Deep patch visual slam,” in European Conference on Computer Vision. Springer, 2024, pp. 424– 440
2024
-
[154]
Xvo: General- ized visual odometry via cross-modal self-training,
L. Lai, Z. Shangguan, J. Zhang, and E. Ohn-Bar, “Xvo: General- ized visual odometry via cross-modal self-training,” in Proceed- ings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 10 094–10 105
2023
-
[155]
Anycam: Learning to recover camera poses and intrinsics from casual videos,
F. Wimbauer, W. Chen, D. Muhle, C. Rupprecht, and D. Cremers, “Anycam: Learning to recover camera poses and intrinsics from casual videos,” in Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2025
2025
-
[156]
Dynamic camera poses and where to find them,
C. Rockwell, J. Tung, T.-Y. Lin, M.-Y. Liu, D. F. Fouhey, and C.-H. Lin, “Dynamic camera poses and where to find them,” in Pro- ceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 12 444–12 455
2025
-
[157]
Efficient camera exposure control for visual odometry via deep reinforcement learning,
S. Zhang, J. He, Y. Zhu, J. Wu, and J. Yuan, “Efficient camera exposure control for visual odometry via deep reinforcement learning,” IEEE Robotics and Automation Letters, 2024
2024
-
[158]
Reinforcement learning meets visual odometry,
N. Messikommer, G. Cioffi, M. Gehrig, and D. Scaramuzza, “Reinforcement learning meets visual odometry,” in European Conference on Computer Vision. Springer, 2024, pp. 76–92
2024
-
[159]
Tracking everything everywhere all at once,
Q. Wang, Y.-Y. Chang, R. Cai, Z. Li, B. Hariharan, A. Holynski, and N. Snavely, “Tracking everything everywhere all at once,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 19 795–19 806
2023
-
[160]
Track every- thing everywhere fast and robustly,
Y. Song, J. Lei, Z. Wang, L. Liu, and K. Daniilidis, “Track every- thing everywhere fast and robustly,” in European Conference on Computer Vision. Springer, 2024, pp. 343–359
2024
-
[161]
Spatialtracker: Tracking any 2d pixels in 3d space,
Y. Xiao, Q. Wang, S. Zhang, N. Xue, S. Peng, Y. Shen, and X. Zhou, “Spatialtracker: Tracking any 2d pixels in 3d space,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 406–20 417
2024
-
[162]
Scene- tracker: Long-term scene flow estimation network,
B. Wang, J. Li, Y. Yu, L. Liu, Z. Sun, and D. Hu, “Scene- tracker: Long-term scene flow estimation network,”arXiv preprint arXiv:2403.19924, 2024
2024 arXiv
-
[163]
Delta: Dense efficient long-range 3d tracking for any video,
T. D. Ngo, P . Zhuang, C. Gan, E. Kalogerakis, S. Tulyakov, H.-Y. Lee, and C. Wang, “Delta: Dense efficient long-range 3d tracking for any video,” arXiv preprint arXiv:2410.24211, 2024
2024 arXiv
-
[164]
Seurat: From moving points to depth,
S. Cho, J. Huang, S. Kim, and J.-Y. Lee, “Seurat: From moving points to depth,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 7211–7221
2025
-
[165]
Tapip3d: Tracking any point in persistent 3d geometry,
B. Zhang, L. Ke, A. W. Harley, and K. Fragkiadaki, “Tapip3d: Tracking any point in persistent 3d geometry,” arXiv preprint arXiv:2504.14717, 2025
2025
-
[166]
Ego- points: Advancing point tracking for egocentric videos,
A. Darkhalil, R. Guerrier, A. W. Harley, and D. Damen, “Ego- points: Advancing point tracking for egocentric videos,” in 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2025, pp. 8556–8565
2025
-
[167]
Robust consistent video depth estimation,
J. Kopf, X. Rong, and J.-B. Huang, “Robust consistent video depth estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 1611–1621. 19
2021
-
[168]
Structure and motion from casual videos,
Z. Zhang, F. Cole, Z. Li, M. Rubinstein, N. Snavely, and W. T. Freeman, “Structure and motion from casual videos,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 20–37
2022
-
[169]
Megasam: Accurate, fast, and robust structure and motion from casual dynamic videos,
Z. Li, R. Tucker, F. Cole, Q. Wang, L. Jin, V . Ye, A. Kanazawa, A. Holynski, and N. Snavely, “Megasam: Accurate, fast, and robust structure and motion from casual dynamic videos,” arXiv preprint arXiv:2412.04463, 2024
2024 arXiv
-
[170]
Easi3r: Estimating disentangled motion from dust3r without training,
X. Chen, Y. Chen, Y. Xiu, A. Geiger, and A. Chen, “Easi3r: Estimating disentangled motion from dust3r without training,” arXiv preprint arXiv:2503.24391, 2025
2025
-
[171]
Geome- trycrafter: Consistent geometry estimation for open-world videos with diffusion priors,
T.-X. Xu, X. Gao, W. Hu, X. Li, S.-H. Zhang, and Y. Shan, “Geome- trycrafter: Consistent geometry estimation for open-world videos with diffusion priors,” arXiv preprint arXiv:2504.01016, 2025
2025 arXiv
-
[172]
3d reconstruction with spatial mem- ory,
H. Wang and L. Agapito, “3d reconstruction with spatial mem- ory,” arXiv preprint arXiv:2408.16061, 2024
2024 arXiv
-
[173]
Continuous 3d perception model with persistent state,
Q. Wang, Y. Zhang, A. Holynski, A. A. Efros, and A. Kanazawa, “Continuous 3d perception model with persistent state,” arXiv preprint arXiv:2501.12387, 2025
2025 arXiv
-
[174]
Point3r: Streaming 3d re- construction with explicit spatial pointer memory,
Y. Wu, W. Zheng, J. Zhou, and J. Lu, “Point3r: Streaming 3d re- construction with explicit spatial pointer memory,” arXiv preprint arXiv:2507.02863, 2025
2025
-
[175]
Streaming 4d visual geometry transformer,
D. Zhuo, W. Zheng, J. Guo, Y. Wu, J. Zhou, and J. Lu, “Streaming 4d visual geometry transformer,” arXiv preprint arXiv:2507.11539, 2025
2025 arXiv
-
[176]
π3: Scalable permutation- equivariant visual geometry learning,
Y. Wang, J. Zhou, H. Zhu, W. Chang, Y. Zhou, Z. Li, J. Chen, J. Pang, C. Shen, and T. He, “ π3: Scalable permutation- equivariant visual geometry learning,” 2025. [Online]. Available: https://arxiv.org/abs/2507.13347
2025 arXiv
-
[177]
Aether: Geometric-aware unified world modeling,
A. Team, H. Zhu, Y. Wang, J. Zhou, W. Chang, Y. Zhou, Z. Li, J. Chen, C. Shen, J. Pang et al., “Aether: Geometric-aware unified world modeling,” arXiv preprint arXiv:2503.18945, 2025
2025 arXiv
-
[178]
Geo4d: Leveraging video generators for geometric 4d scene reconstruc- tion,
Z. Jiang, C. Zheng, I. Laina, D. Larlus, and A. Vedaldi, “Geo4d: Leveraging video generators for geometric 4d scene reconstruc- tion,” arXiv preprint arXiv:2504.07961, 2025
2025 arXiv
-
[179]
Unigeo: Taming video diffusion for unified consistent geometry estimation,
Y.-T. Sun, X. Yu, Z. Huang, Y.-H. Huang, Y.-C. Guo, Z. Yang, Y.- P . Cao, and X. Qi, “Unigeo: Taming video diffusion for unified consistent geometry estimation,” arXiv preprint arXiv:2505.24521, 2025
2025 arXiv
-
[180]
Uni4d: Unifying visual foundation models for 4d modeling from a single video,
D. Y. Yao, A. J. Zhai, and S. Wang, “Uni4d: Unifying visual foundation models for 4d modeling from a single video,” in Pro- ceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 1116–1126
2025
-
[181]
Back on track: Bundle ad- justment for dynamic scene reconstruction,
W. Chen, G. Zhang, F. Wimbauer, R. Wang, N. Araslanov, A. Vedaldi, and D. Cremers, “Back on track: Bundle ad- justment for dynamic scene reconstruction,” arXiv preprint arXiv:2504.14516, 2025
2025
-
[182]
Stereo4d: Learning how things move in 3d from internet stereo videos,
L. Jin, R. Tucker, Z. Li, D. Fouhey, N. Snavely, and A. Holynski, “Stereo4d: Learning how things move in 3d from internet stereo videos,” arXiv preprint arXiv:2412.09621, 2024
2024 arXiv
-
[183]
Dynamic point maps: A versatile representation for dynamic 3d reconstruction,
E. Sucar, Z. Lai, E. Insafutdinov, and A. Vedaldi, “Dynamic point maps: A versatile representation for dynamic 3d reconstruction,” arXiv preprint arXiv:2503.16318, 2025
2025 arXiv
-
[184]
St4rtrack: Simultaneous 4d reconstruction and tracking in the world,
H. Feng, J. Zhang, Q. Wang, Y. Ye, P . Yu, M. J. Black, T. Darrell, and A. Kanazawa, “St4rtrack: Simultaneous 4d reconstruction and tracking in the world,” arXiv preprint arXiv:2504.13152, 2025
2025 arXiv
-
[185]
Pomato: Marrying pointmap matching with temporal motion for dynamic 3d reconstruction,
S. Zhang, Y. Ge, J. Tian, G. Xu, H. Chen, C. Lv, and C. Shen, “Pomato: Marrying pointmap matching with temporal motion for dynamic 3d reconstruction,” arXiv preprint arXiv:2504.05692 , 2025
2025 arXiv
-
[186]
Dˆ 2ust3r: Enhancing 3d reconstruction with 4d pointmaps for dynamic scenes,
J. Han, H. An, J. Jung, T. Narihira, J. Seo, K. Fukuda, C. Kim, S. Hong, Y. Mitsufuji, and S. Kim, “Dˆ 2ust3r: Enhancing 3d reconstruction with 4d pointmaps for dynamic scenes,” arXiv preprint arXiv:2504.06264, 2025
2025
-
[187]
Zero- shot monocular scene flow estimation in the wild,
Y. Liang, A. Badki, H. Su, J. Tompkin, and O. Gallo, “Zero- shot monocular scene flow estimation in the wild,” arXiv preprint arXiv:2501.10357, 2025
2025 arXiv
-
[188]
Fast encoder-based 3d from casual videos via point track processing,
Y. Kasten, W. Lu, and H. Maron, “Fast encoder-based 3d from casual videos via point track processing,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems
-
[189]
Transformers in vision: A survey,
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM computing surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022
2022
-
[190]
Spatialtrackerv2: 3d point tracking made easy,
Y. Xiao, J. Wang, N. Xue, N. Karaev, Y. Makarov, B. Kang, X. Zhu, H. Bao, Y. Shen, and X. Zhou, “Spatialtrackerv2: 3d point tracking made easy,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2025. [Online]. Available: https://arxiv.org/abs/2507.12462
2025 arXiv
-
[191]
Undeepvo: Monocular visual odometry through unsupervised deep learning,
R. Li, S. Wang, Z. Long, and D. Gu, “Undeepvo: Monocular visual odometry through unsupervised deep learning,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 7286–7291
2018
-
[192]
Deepv2d: Video to depth with differen- tiable structure from motion,
Z. Teed and J. Deng, “Deepv2d: Video to depth with differen- tiable structure from motion,” arXiv preprint arXiv:1812.04605 , 2018
2018 arXiv
-
[193]
Generalizing to the open world: Deep visual odometry with online adaptation,
S. Li, X. Wu, Y. Cao, and H. Zha, “Generalizing to the open world: Deep visual odometry with online adaptation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2021, pp. 13 184–13 193
2021
-
[194]
Improving monocular visual odometry using learned depth,
L. Sun, W. Yin, E. Xie, Z. Li, C. Sun, and C. Shen, “Improving monocular visual odometry using learned depth,” IEEE Transac- tions on Robotics, vol. 38, no. 5, pp. 3173–3186, 2022
2022
-
[195]
Towards better generaliza- tion: Joint depth-pose learning without posenet,
W. Zhao, S. Liu, Y. Shu, and Y.-J. Liu, “Towards better generaliza- tion: Joint depth-pose learning without posenet,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2020, pp. 9151–9161
2020
-
[196]
Deep unsupervised visual odometry via bundle adjusted pose graph optimization,
G. Lu, “Deep unsupervised visual odometry via bundle adjusted pose graph optimization,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 6131–6137
2023
-
[197]
Droid-slam: Deep visual slam for monoc- ular, stereo, and rgb-d cameras,
Z. Teed and J. Deng, “Droid-slam: Deep visual slam for monoc- ular, stereo, and rgb-d cameras,” Advances in neural information processing systems, vol. 34, pp. 16 558–16 569, 2021
2021
-
[198]
Nicer-slam: Neural implicit scene encoding for rgb slam,
Z. Zhu, S. Peng, V . Larsson, Z. Cui, M. R. Oswald, A. Geiger, and M. Pollefeys, “Nicer-slam: Neural implicit scene encoding for rgb slam,” in 2024 International Conference on 3D Vision (3DV). IEEE, 2024, pp. 42–52
2024
-
[199]
Dense rgb slam with neural implicit maps,
H. Li, X. Gu, W. Yuan, L. Yang, Z. Dong, and P . Tan, “Dense rgb slam with neural implicit maps,” arXiv preprint arXiv:2301.08930, 2023
2023 arXiv
-
[200]
Go-slam: Global optimization for consistent 3d instant reconstruction,
Y. Zhang, F. Tosi, S. Mattoccia, and M. Poggi, “Go-slam: Global optimization for consistent 3d instant reconstruction,” in Proceed- ings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3727–3737
2023
-
[201]
Glorie-slam: Globally optimized rgb-only im- plicit encoding point cloud slam,
G. Zhang, E. Sandstr ¨om, Y. Zhang, M. Patel, L. Van Gool, and M. R. Oswald, “Glorie-slam: Globally optimized rgb-only im- plicit encoding point cloud slam,”arXiv preprint arXiv:2403.19549, 2024
2024 arXiv
-
[202]
Splat-slam: Globally op- timized rgb-only slam with 3d gaussians,
E. Sandstr ¨om, K. Tateno, M. Oechsle, M. Niemeyer, L. Van Gool, M. R. Oswald, and F. Tombari, “Splat-slam: Globally op- timized rgb-only slam with 3d gaussians,” arXiv preprint arXiv:2405.16544, 2024
2024 arXiv
-
[203]
Gaussian splatting slam,
H. Matsuki, R. Murai, P . H. Kelly, and A. J. Davison, “Gaussian splatting slam,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 18 039–18 048
2024
-
[204]
How nerfs and 3d gaussian splat- ting are reshaping slam: A survey. arxiv 2024,
F. Tosi, Y. Zhang, Z. Gong, E. Sandstr ¨om, S. Mattoccia, M. Os- wald, and M. Poggi, “How nerfs and 3d gaussian splat- ting are reshaping slam: A survey. arxiv 2024,” arXiv preprint arXiv:2402.13255, 2024
2024 arXiv
-
[205]
Flowmap: High-quality camera poses, intrinsics, and depth via gradient descent,
C. Smith, D. Charatan, A. Tewari, and V . Sitzmann, “Flowmap: High-quality camera poses, intrinsics, and depth via gradient descent,” arXiv preprint arXiv:2404.15259, 2024
2024 arXiv
-
[206]
Grounding image matching in 3d with mast3r,
V . Leroy, Y. Cabon, and J. Revaud, “Grounding image matching in 3d with mast3r,” in European Conference on Computer Vision . Springer, 2024, pp. 71–91
2024
-
[207]
Mast3r-sfm: a fully-integrated solution for uncon- strained structure-from-motion,
B. Duisterhof, L. Zust, P . Weinzaepfel, V . Leroy, Y. Cabon, and J. Revaud, “Mast3r-sfm: a fully-integrated solution for uncon- strained structure-from-motion,” arXiv preprint arXiv:2409.19152, 2024
2024 arXiv
-
[208]
Light3r-sfm: Towards feed-forward structure-from-motion,
S. Elflein, Q. Zhou, S. Agostinho, and L. Leal-Taix ´e, “Light3r-sfm: Towards feed-forward structure-from-motion,” arXiv preprint arXiv:2501.14914, 2025
2025 arXiv
-
[209]
Must3r: Multi-view network for stereo 3d reconstruction,
Y. Cabon, L. Stoffl, L. Antsfeld, G. Csurka, B. Chidlovskii, J. Re- vaud, and V . Leroy, “Must3r: Multi-view network for stereo 3d reconstruction,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 1050–1060
2025
-
[210]
Pow3r: Empowering unconstrained 3d reconstruction with cam- era and scene priors,
W. Jang, P . Weinzaepfel, V . Leroy, L. Agapito, and J. Revaud, “Pow3r: Empowering unconstrained 3d reconstruction with cam- era and scene priors,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 1071–1081
2025
-
[211]
Regist3r: Incremen- tal registration with stereo foundation model,
S. Liu, W. Li, P . Qiao, and Y. Dou, “Regist3r: Incremen- tal registration with stereo foundation model,” arXiv preprint arXiv:2504.12356, 2025
2025
-
[212]
Surfels: Surface elements as rendering primitives,
H. Pfister, M. Zwicker, J. Van Baar, and M. Gross, “Surfels: Surface elements as rendering primitives,” inProceedings of the 27th annual 20 conference on Computer graphics and interactive techniques, 2000, pp. 335–342
2000
-
[213]
Differentiable surface splatting for point-based geometry pro- cessing,
W. Yifan, F. Serena, S. Wu, C. ¨Oztireli, and O. Sorkine-Hornung, “Differentiable surface splatting for point-based geometry pro- cessing,” ACM Transactions on Graphics (TOG) , vol. 38, no. 6, pp. 1–14, 2019
2019
-
[214]
Neural point cloud rendering via multi-plane projection,
P . Dai, Y. Zhang, Z. Li, S. Liu, and B. Zeng, “Neural point cloud rendering via multi-plane projection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 7830–7839
2020
-
[215]
Synsin: End- to-end view synthesis from a single image,
O. Wiles, G. Gkioxari, R. Szeliski, and J. Johnson, “Synsin: End- to-end view synthesis from a single image,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 7467–7477
2020
-
[216]
Pulsar: Efficient sphere-based neural rendering,
C. Lassner and M. Zollhofer, “Pulsar: Efficient sphere-based neural rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 1440–1449
2021
-
[217]
Point- based neural rendering with per-view optimization,
G. Kopanas, J. Philip, T. Leimk ¨uhler, and G. Drettakis, “Point- based neural rendering with per-view optimization,” inComputer Graphics Forum, vol. 40, no. 4. Wiley Online Library, 2021, pp. 29–43
2021
-
[218]
Adop: Approximate differentiable one-pixel point rendering,
D. R ¨uckert, L. Franke, and M. Stamminger, “Adop: Approximate differentiable one-pixel point rendering,” ACM Transactions on Graphics (ToG), vol. 41, no. 4, pp. 1–14, 2022
2022
-
[219]
Free view synthesis,
G. Riegler and V . Koltun, “Free view synthesis,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIX 16. Springer, 2020, pp. 623–640
2020
-
[220]
Stable view synthesis,
——, “Stable view synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 12 216–12 225
2021
-
[221]
Fwd: Real-time novel view synthesis with forward warping and depth,
A. Cao, C. Rockwell, and J. Johnson, “Fwd: Real-time novel view synthesis with forward warping and depth,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 713–15 724
2022
-
[222]
Botsch, L
M. Botsch, L. Kobbelt, M. Pauly, P . Alliez, and B. L ´evy, Polygon mesh processing. CRC press, 2010
2010
-
[223]
Local surface interpolation with b´ezier patches,
L. A. Shirman and C. H. Sequin, “Local surface interpolation with b´ezier patches,” Computer Aided Geometric Design, vol. 4, no. 4, pp. 279–295, 1987
1987
-
[224]
Dynamic surface func- tion networks for clothed human bodies,
A. Burov, M. Nießner, and J. Thies, “Dynamic surface func- tion networks for clothed human bodies,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 754–10 764
2021
-
[225]
Deferred neural render- ing: Image synthesis using neural textures,
J. Thies, M. Zollh ¨ofer, and M. Nießner, “Deferred neural render- ing: Image synthesis using neural textures,” Acm Transactions on Graphics (TOG), vol. 38, no. 4, pp. 1–12, 2019
2019
-
[226]
Opendr: An approximate differ- entiable renderer,
M. M. Loper and M. J. Black, “Opendr: An approximate differ- entiable renderer,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13. Springer, 2014, pp. 154–169
2014
-
[227]
Neural 3d mesh renderer,
H. Kato, Y. Ushiku, and T. Harada, “Neural 3d mesh renderer,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3907–3916
2018
-
[228]
Soft rasterizer: A differentiable renderer for image-based 3d reasoning,
S. Liu, T. Li, W. Chen, and H. Li, “Soft rasterizer: A differentiable renderer for image-based 3d reasoning,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 7708–7717
2019
-
[229]
Paparazzi: surface editing by way of multi-view image processing
H.-T. D. Liu, M. Tao, and A. Jacobson, “Paparazzi: surface editing by way of multi-view image processing.” ACM Trans. Graph. , vol. 37, no. 6, pp. 221–1, 2018
2018
-
[230]
Mitsuba 2: A retargetable forward and inverse renderer,
M. Nimier-David, D. Vicini, T. Zeltner, and W. Jakob, “Mitsuba 2: A retargetable forward and inverse renderer,” ACM Transactions on Graphics (TOG), vol. 38, no. 6, pp. 1–17, 2019
2019
-
[231]
Taichi: a language for high-performance computation on spa- tially sparse data structures,
Y. Hu, T.-M. Li, L. Anderson, J. Ragan-Kelley, and F. Durand, “Taichi: a language for high-performance computation on spa- tially sparse data structures,” ACM Transactions on Graphics (TOG), vol. 38, no. 6, p. 201, 2019
2019
-
[232]
Optical models for direct volume rendering,
N. Max, “Optical models for direct volume rendering,” IEEE Transactions on Visualization and Computer Graphics , vol. 1, no. 2, pp. 99–108, 1995
1995
-
[233]
Nerf in the wild: Neural radiance fields for unconstrained photo collections,
R. Martin-Brualla, N. Radwan, M. S. Sajjadi, J. T. Barron, A. Doso- vitskiy, and D. Duckworth, “Nerf in the wild: Neural radiance fields for unconstrained photo collections,” in CVPR, 2021, pp. 7210–7219
2021
-
[234]
Physg: Inverse rendering with spherical gaussians for physics-based material editing and relighting,
K. Zhang, F. Luan, Q. Wang, K. Bala, and N. Snavely, “Physg: Inverse rendering with spherical gaussians for physics-based material editing and relighting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 5453–5462
2021
-
[235]
Barf: Bundle- adjusting neural radiance fields,
C.-H. Lin, W.-C. Ma, A. Torralba, and S. Lucey, “Barf: Bundle- adjusting neural radiance fields,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 5741–5751
2021
-
[236]
Humannerf: Free-viewpoint ren- dering of moving people from monocular video,
C.-Y. Weng, B. Curless, P . P . Srinivasan, J. T. Barron, and I. Kemelmacher-Shlizerman, “Humannerf: Free-viewpoint ren- dering of moving people from monocular video,” inProceedings of the IEEE/CVF conference on computer vision and pattern Recognition, 2022, pp. 16 210–16 220
2022
-
[237]
Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,
P . Wang, L. Liu, Y. Liu, C. Theobalt, T. Komura, and W. Wang, “Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,” Advances in Neural Information Processing Systems, vol. 34, pp. 27 171–27 183, 2021
2021
-
[238]
Volume rendering of neural implicit surfaces,
L. Yariv, J. Gu, Y. Kasten, and Y. Lipman, “Volume rendering of neural implicit surfaces,” Advances in Neural Information Process- ing Systems, vol. 34, pp. 4805–4815, 2021
2021
-
[239]
GIRAFFE: Representing scenes as compositional generative neural feature fields,
M. Niemeyer and A. Geiger, “GIRAFFE: Representing scenes as compositional generative neural feature fields,” 2021
2021
-
[240]
Mip-nerf: A multiscale representa- tion for anti-aliasing neural radiance fields,
J. T. Barron, B. Mildenhall, M. Tancik, P . Hedman, R. Martin- Brualla, and P . P . Srinivasan, “Mip-nerf: A multiscale representa- tion for anti-aliasing neural radiance fields,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 5855–5864
2021
-
[241]
Mip-nerf 360: Unbounded anti-aliased neural radiance fields,
J. T. Barron, B. Mildenhall, D. Verbin, P . P . Srinivasan, and P . Hed- man, “Mip-nerf 360: Unbounded anti-aliased neural radiance fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5470–5479
2022
-
[242]
Ref-nerf: Structured view-dependent appear- ance for neural radiance fields,
D. Verbin, P . Hedman, B. Mildenhall, T. Zickler, J. T. Barron, and P . P . Srinivasan, “Ref-nerf: Structured view-dependent appear- ance for neural radiance fields,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022, pp. 5481–5490
2022
-
[243]
Dynibar: Neural dynamic image-based rendering,
Z. Li, Q. Wang, F. Cole, R. Tucker, and N. Snavely, “Dynibar: Neural dynamic image-based rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 4273–4284
2023
-
[244]
In-place scene labelling and understanding with implicit scene represen- tation,
S. Zhi, T. Laidlow, S. Leutenegger, and A. J. Davison, “In-place scene labelling and understanding with implicit scene represen- tation,” in ICCV, 2021, pp. 15 838–15 847
2021
-
[245]
Sparf: Neural radiance fields from sparse and noisy poses,
P . Truong, M.-J. Rakotosaona, F. Manhardt, and F. Tombari, “Sparf: Neural radiance fields from sparse and noisy poses,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 4190–4200
2023
-
[246]
Nerd: Neural reflectance decomposition from image collec- tions,
M. Boss, R. Braun, V . Jampani, J. T. Barron, C. Liu, and H. Lensch, “Nerd: Neural reflectance decomposition from image collec- tions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 12 684–12 694
2021
-
[247]
pixelnerf: Neural radiance fields from one or few images,
A. Yu, V . Ye, M. Tancik, and A. Kanazawa, “pixelnerf: Neural radiance fields from one or few images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 4578–4587
2021
-
[248]
IBRNet: Learning multi-view image-based rendering,
Q. Wang, Z. Wang, K. Genova, P . P . Srinivasan, H. Zhou, J. T. Bar- ron, R. Martin-Brualla, N. Snavely, and T. Funkhouser, “IBRNet: Learning multi-view image-based rendering,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 4690–4699
2021
-
[249]
Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps,
C. Reiser, S. Peng, Y. Liao, and A. Geiger, “Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 335–14 345
2021
-
[250]
Neural radiance flow for 4d view synthesis and video processing,
Y. Du, Y. Zhang, H.-X. Yu, J. B. Tenenbaum, and J. Wu, “Neural radiance flow for 4d view synthesis and video processing,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE Computer Society, 2021, pp. 14 304–14 314
2021
-
[251]
Neu- ral 3d video synthesis from multi-view video,
T. Li, M. Slavcheva, M. Zollhoefer, S. Green, C. Lassner, C. Kim, T. Schmidt, S. Lovegrove, M. Goesele, R. Newcombe et al., “Neu- ral 3d video synthesis from multi-view video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2022, pp. 5521–5531
2022
-
[252]
Animatable neural radiance fields for modeling dy- namic human bodies,
S. Peng, J. Dong, Q. Wang, S. Zhang, Q. Shuai, X. Zhou, and H. Bao, “Animatable neural radiance fields for modeling dy- namic human bodies,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 314–14 323
2021
-
[253]
Nerf in the palm of your hand: Corrective augmentation for robotics via novel-view synthesis,
A. Zhou, M. J. Kim, L. Wang, P . Florence, and C. Finn, “Nerf in the palm of your hand: Corrective augmentation for robotics via novel-view synthesis,” in Proceedings of the IEEE/CVF Conference 21 on Computer Vision and Pattern Recognition , 2023, pp. 17 907– 17 917
2023
-
[254]
Neat: Neural adaptive tomography,
D. R ¨uckert, Y. Wang, R. Li, R. Idoughi, and W. Heidrich, “Neat: Neural adaptive tomography,” ACM Transactions on Graphics (TOG), vol. 41, no. 4, pp. 1–13, 2022
2022
-
[255]
Gravitationally lensed black hole emission tomography,
A. Levis, P . P . Srinivasan, A. A. Chael, R. Ng, and K. L. Bouman, “Gravitationally lensed black hole emission tomography,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 19 841–19 850
2022
-
[256]
Mip-NeRF 360: Unbounded anti-aliased neural radiance fields,
J. T. Barron, B. Mildenhall, D. Verbin, P . P . Srinivasan, and P . Hed- man, “Mip-NeRF 360: Unbounded anti-aliased neural radiance fields,” CoRR, vol. abs/2111.12077, 2022
2022 arXiv
-
[257]
Screened poisson surface recon- struction,
M. Kazhdan and H. Hoppe, “Screened poisson surface recon- struction,” ACM Transactions on Graphics (ToG), vol. 32, no. 3, pp. 1–13, 2013
2013
-
[258]
Two algorithms for constructing a delaunay triangulation,
D.-T. Lee and B. J. Schachter, “Two algorithms for constructing a delaunay triangulation,” International Journal of Computer & Information Sciences, vol. 9, no. 3, pp. 219–242, 1980
1980
-
[259]
Structure-from-motion re- visited,
J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-motion re- visited,” in Conference on Computer Vision and Pattern Recognition (CVPR), 2016
2016
-
[260]
Mvsnet: Depth inference for unstructured multi-view stereo,
Y. Yao, Z. Luo, S. Li, T. Fang, and L. Quan, “Mvsnet: Depth inference for unstructured multi-view stereo,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 767– 783
2018
-
[261]
Neat: Learning neural implicit surfaces with arbitrary topologies from multi-view images,
X. Meng, W. Chen, and B. Yang, “Neat: Learning neural implicit surfaces with arbitrary topologies from multi-view images,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 248–258
2023
-
[262]
Neuralangelo: High-fidelity neural surface recon- struction,
Z. Li, T. M ¨uller, A. Evans, R. H. Taylor, M. Unberath, M.-Y. Liu, and C.-H. Lin, “Neuralangelo: High-fidelity neural surface recon- struction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 8456–8465
2023
-
[263]
Marching cubes: A high res- olution 3d surface construction algorithm,
W. E. Lorensen and H. E. Cline, “Marching cubes: A high res- olution 3d surface construction algorithm,” in Seminal graphics: pioneering efforts that shaped the field , 1998, pp. 347–353
1998
-
[264]
2d gaussian splatting for geometrically accurate radiance fields,
B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao, “2d gaussian splatting for geometrically accurate radiance fields,” in ACM SIGGRAPH 2024 conference papers, 2024, pp. 1–11
2024
-
[265]
Gaussian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes,
Z. Yu, T. Sattler, and A. Geiger, “Gaussian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes,” ACM Transactions on Graphics (TOG), vol. 43, no. 6, pp. 1–13, 2024
2024
-
[266]
Pgsr: Planar-based gaussian splatting for efficient and high-fidelity surface reconstruction,
D. Chen, H. Li, W. Ye, Y. Wang, W. Xie, S. Zhai, N. Wang, H. Liu, H. Bao, and G. Zhang, “Pgsr: Planar-based gaussian splatting for efficient and high-fidelity surface reconstruction,” IEEE Transactions on Visualization and Computer Graphics, 2024
2024
-
[267]
Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering,
A. Gu ´edon and V . Lepetit, “Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5354–5363
2024
-
[268]
3d gaus- sian splatting for fine-detailed surface reconstruction in large- scale scene,
S. Chen, Z. Li, Z. Chen, Q. Yan, G. Shen, and R. Duan, “3d gaus- sian splatting for fine-detailed surface reconstruction in large- scale scene,” arXiv preprint arXiv:2506.17636, 2025
2025 arXiv
-
[269]
Multiview geometric regularization of gaussian splatting for accurate radiance fields,
J. Kim, G. Park, and S. Lee, “Multiview geometric regularization of gaussian splatting for accurate radiance fields,” arXiv preprint arXiv:2506.13508, 2025
2025 arXiv
-
[270]
Tri 2 plane: Advancing neural im- plicit surface reconstruction for indoor scenes,
Y. Xie, H. Xiao, and W. Kang, “Tri 2 plane: Advancing neural im- plicit surface reconstruction for indoor scenes,” IEEE Transactions on Multimedia, 2025
2025
-
[271]
Esa-gs: Elongation splitting and assimilation in gaussian splatting for accurate sur- face reconstruction,
Y. Chen, W. Wu, Y. Peng, Y. Fei, and L. Zheng, “Esa-gs: Elongation splitting and assimilation in gaussian splatting for accurate sur- face reconstruction,” Computer Aided Geometric Design , p. 102434, 2025
2025
-
[272]
Gaussianudf: Inferring unsigned distance functions through 3d gaussian splatting,
S. Li, Y.-S. Liu, and Z. Han, “Gaussianudf: Inferring unsigned distance functions through 3d gaussian splatting,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 27 113–27 123
2025
-
[273]
Sof: Sorted opacity fields for fast unbounded surface reconstruction,
L. Radl, F. Windisch, T. Deixelberger, J. Hladky, M. Steiner, D. Schmalstieg, and M. Steinberger, “Sof: Sorted opacity fields for fast unbounded surface reconstruction,” arXiv preprint arXiv:2506.19139, 2025
2025
-
[274]
Multi-view surface reconstruction using normal and reflectance cues,
R. Bruneau, B. Brument, Y. Qu ´eau, J. M ´elou, F. B. Lauze, J.-D. Durou, and L. Calvet, “Multi-view surface reconstruction using normal and reflectance cues,” arXiv preprint arXiv:2506.04115 , 2025
2025
-
[275]
Quicksplat: Fast 3d surface reconstruction via learned gaussian initialization,
Y.-C. Liu, L. H ¨ollein, M. Nießner, and A. Dai, “Quicksplat: Fast 3d surface reconstruction via learned gaussian initialization,” arXiv preprint arXiv:2505.05591, 2025
2025 arXiv
-
[276]
Geometry field splatting with gaussian surfels,
K. Jiang, V . Sivaram, C. Peng, and R. Ramamoorthi, “Geometry field splatting with gaussian surfels,” in Proceedings of the Com- puter Vision and Pattern Recognition Conference , 2025, pp. 5752– 5762
2025
-
[277]
Solidgs: Consolidating gaussian surfel splatting for sparse-view surface reconstruction,
Z. Shen, Y. Liu, Z. Chen, Z. Li, J. Wang, Y. Liang, Z. Yu, J. Zhang, Y. Xu, S. Schaefer et al., “Solidgs: Consolidating gaussian surfel splatting for sparse-view surface reconstruction,” arXiv preprint arXiv:2412.15400, 2024
2024 arXiv
-
[278]
Sparseneus: Fast generalizable neural surface reconstruction from sparse views,
X. Long, C. Lin, P . Wang, T. Komura, and W. Wang, “Sparseneus: Fast generalizable neural surface reconstruction from sparse views,” in European Conference on Computer Vision . Springer, 2022, pp. 210–227
2022
-
[279]
Gens: Generalizable neural surface reconstruction from multi-view im- ages,
R. Peng, X. Gu, L. Tang, S. Shen, F. Yu, and R. Wang, “Gens: Generalizable neural surface reconstruction from multi-view im- ages,” Advances in Neural Information Processing Systems , vol. 36, pp. 56 932–56 945, 2023
2023
-
[280]
C2f2neus: Cascade cost frustum fusion for high fidelity and generalizable neural surface reconstruction,
L. Xu, T. Guan, Y. Wang, W. Liu, Z. Zeng, J. Wang, and W. Yang, “C2f2neus: Cascade cost frustum fusion for high fidelity and generalizable neural surface reconstruction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 18 291–18 301
2023
-
[281]
Uforecon: generalizable sparse-view surface reconstruction from arbitrary and unfavorable sets,
Y. Na, W. J. Kim, K. B. Han, S. Ha, and S.-E. Yoon, “Uforecon: generalizable sparse-view surface reconstruction from arbitrary and unfavorable sets,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5094–5104
2024
-
[282]
Surface-centric modeling for high-fidelity generalizable neural surface reconstruction,
R. Peng, S. Shen, K. Xiong, H. Gao, J. Jiao, X. Gu, and R. Wang, “Surface-centric modeling for high-fidelity generalizable neural surface reconstruction,” in European Conference on Computer Vi- sion. Springer, 2024, pp. 183–200
2024
-
[283]
Retr: Modeling rendering via transformer for generalizable neural surface reconstruction,
Y. Liang, H. He, and Y. Chen, “Retr: Modeling rendering via transformer for generalizable neural surface reconstruction,” Ad- vances in neural information processing systems , vol. 36, pp. 62 332– 62 351, 2023
2023
-
[284]
Lara: Efficient large-baseline radiance fields,
A. Chen, H. Xu, S. Esposito, S. Tang, and A. Geiger, “Lara: Efficient large-baseline radiance fields,” in European Conference on Computer Vision. Springer, 2024, pp. 338–355
2024
-
[285]
Special section on egocentric perception,
A. Furnari, D. Crandall, D. Damen, K. Grauman, and G. M. Farinella, “Special section on egocentric perception,” IEEE Trans- actions on Pattern Analysis and Machine Intelligence , vol. 45, no. 6, pp. 6602–6604, 2023
2023
-
[286]
An outlook into the future of egocentric vision,
C. Plizzari, G. Goletto, A. Furnari, S. Bansal, F. Ragusa, G. M. Farinella, D. Damen, and T. Tommasi, “An outlook into the future of egocentric vision,” International Journal of Computer Vision, vol. 132, no. 11, pp. 4880–4936, 2024
2024
-
[287]
Scene- script: Reconstructing scenes with an autoregressive structured language model,
A. Avetisyan, C. Xie, H. Howard-Jenkins, T.-Y. Yang, S. Aroudj, S. Patra, F. Zhang, D. Frost, L. Holland, C. Orme et al., “Scene- script: Reconstructing scenes with an autoregressive structured language model,” in European Conference on Computer Vision . Springer, 2024, pp. 247–263
2024
-
[288]
Ego- lifter: Open-world 3d segmentation for egocentric perception,
Q. Gu, Z. Lv, D. Frost, S. Green, J. Straub, and C. Sweeney, “Ego- lifter: Open-world 3d segmentation for egocentric perception,” in European Conference on Computer Vision . Springer, 2024, pp. 382–400
2024
-
[289]
Photoreal scene reconstruction from an egocentric device,
Z. Lv, M. Monge, K. Chen, Y. Zhu, M. Goesele, J. Engel, Z. Dong, and R. Newcombe, “Photoreal scene reconstruction from an egocentric device,” in ACM SIGGRAPH, 2025
2025
-
[290]
Spatial cognition from egocentric video: Out of sight, not out of mind,
C. Plizzari, S. Goel, T. Perrett, J. Chalk, A. Kanazawa, and D. Damen, “Spatial cognition from egocentric video: Out of sight, not out of mind,” arXiv preprint arXiv:2404.05072, 2024
2024 arXiv
-
[291]
Nerf++: Analyzing and improving neural radiance fields,
K. Zhang, G. Riegler, N. Snavely, and V . Koltun, “Nerf++: Analyzing and improving neural radiance fields,” CoRR, vol. abs/2010.07492, 2020
2010 arXiv
-
[292]
Zip-nerf: Anti-aliased grid-based neural radiance fields,
J. T. Barron, B. Mildenhall, D. Verbin, P . P . Srinivasan, and P . Hedman, “Zip-nerf: Anti-aliased grid-based neural radiance fields,” CoRR, vol. abs/2304.06706, 2023
2023 arXiv
-
[293]
City- gaussian: Real-time high-quality large-scale scene rendering with gaussians,
Y. Liu, C. Luo, L. Fan, N. Wang, J. Peng, and Z. Zhang, “City- gaussian: Real-time high-quality large-scale scene rendering with gaussians,” in European Conference on Computer Vision. Springer, 2024, pp. 265–282
2024
-
[294]
Citygaussianv2: Efficient and geometrically accurate reconstruction for large-scale scenes,
Y. Liu, C. Luo, Z. Mao, J. Peng, and Z. Zhang, “Citygaussianv2: Efficient and geometrically accurate reconstruction for large-scale scenes,” arXiv preprint arXiv:2411.00771, 2024
2024 arXiv
-
[295]
Octree- 22 gs: Towards consistent real-time rendering with lod-structured 3d gaussians,
K. Ren, L. Jiang, T. Lu, M. Yu, L. Xu, Z. Ni, and B. Dai, “Octree- 22 gs: Towards consistent real-time rendering with lod-structured 3d gaussians,” arXiv preprint arXiv:2403.17898, 2024
2024 arXiv
-
[296]
Citygs-x: A scalable architecture for efficient and geometrically accurate large-scale scene reconstruction,
Y. Gao, H. Li, J. Chen, Z. Zou, Z. Zhong, D. Zhang, X. Sun, and J. Han, “Citygs-x: A scalable architecture for efficient and geometrically accurate large-scale scene reconstruction,” arXiv preprint arXiv:2503.23044, 2025
2025 arXiv
-
[297]
Lodge: Level- of-detail large-scale gaussian splatting with efficient rendering,
J. Kulhanek, M.-J. Rakotosaona, F. Manhardt, C. Tsalicoglou, M. Niemeyer, T. Sattler, S. Peng, and F. Tombari, “Lodge: Level- of-detail large-scale gaussian splatting with efficient rendering,” arXiv preprint arXiv:2505.23158, 2025
2025
-
[298]
Block-nerf: Scal- able large-scene neural view synthesis,
M. Tancik, V . Casser, X. Yan, S. Pradhan, B. Mildenhall, P . P . Srinivasan, J. T. Barron, and H. Kretzschmar, “Block-nerf: Scal- able large-scene neural view synthesis,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2022
2022
-
[299]
Mega-nerf: Scal- able construction of large-scale nerfs for virtual fly-throughs,
H. Turki, D. Ramanan, and M. Satyanarayanan, “Mega-nerf: Scal- able construction of large-scale nerfs for virtual fly-throughs,” CoRR, vol. abs/2112.10703, 2022
2022 arXiv
-
[300]
Bungeenerf (city-nerf): Progressive neural radiance field for extreme multi-scale scene rendering,
Y. Xiangli, L. Xu, X. Pan, N. Zhao, A. Rao, C. Theobalt, B. Dai, and D. Lin, “Bungeenerf (city-nerf): Progressive neural radiance field for extreme multi-scale scene rendering,” in European Conf. Computer Vision (ECCV), 2022, pp. 106–122
2022
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.