Pith. sign in

REVIEW 3 major objections 5 minor 9 cited by

Reconstructing 4D Spatial Intelligence: A Survey

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Five levels map how machines build 4D scenes from video

desk verdict A current, well-organized survey whose five-level taxonomy is useful as a thematic map, but the 'progressive' hierarchy is asserted, not demonstrated. read the letter →

arxiv 2507.21045 v2 pith:IB3TXDV4 submitted 2025-07-28 cs.CV

classification cs.CV
keywords 4Dspatialintelligencescenereconstructionvideo-basedunderstandingtaxonomydynamicsceneshuman-objectinteractionphysics-basedneuralradiancefields
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that the scattered field of video-based 4D scene reconstruction is actually a progression with five levels of spatial intelligence: low-level 3D cues, 3D scene components, dynamic 4D scenes, interactions among components, and physical laws. The authors' claim is that organizing existing methods this way reveals the field's internal structure better than previous surveys, which covered stereo, 3D reconstruction, or dynamic scenes separately. A reader should care because the taxonomy turns hundreds of individual papers into a map with named rungs, each with its own open challenges, and it points toward what building the next level would require.

What carries the argument

The carrying object is the five-level taxonomy itself, defined in the introduction and used to organize every section: Level 1 low-level 3D cues, Level 2 3D scene components, Level 3 4D dynamic scenes, Level 4 interactions among components, and Level 5 physical laws and constraints. The taxonomy does the argument's work by assigning each surveyed method to a rung and by framing each section's open challenges as what must be solved before the next rung becomes reachable. Supporting machinery includes the 3D representations that appear across levels, such as neural radiance fields, 3D Gaussian splatting, signed distance functions, and parametric body models like SMPL, but these are tools rather than the survey's contribution.

What would settle it

A dependency analysis of the surveyed methods would falsify the progression claim if it found that a large share of methods at Levels 4 and 5 were developed without using outputs from Levels 1 and 2, or if the levels did not cluster in citation space.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that achieving full 4D spatial intelligence from video, capturing geometry, objects, motion, interaction, and physical behavior, is not a single problem but a layered one, and that the literature already reflects this layering. The five levels are: (1) low-level 3D cues such as depth, camera pose, point maps, and 3D tracking; (2) reconstruction of 3D scene components such as objects, humans, and structures; (3) reconstruction of dynamic 4D scenes, typically by canonical-space deformation or by adding time to the representation; (4) modeling of interactions among scene components, mostly human-centric; and (5) incorporation of physical laws and constraints so reconstructions behave plausibly under gravity, friction, and contact. The paper further claims that each level supports the next, and it ends each section by listing the challenges that block progress to the next level.

Load-bearing premise

The load-bearing premise is that the five levels really are a progression in which higher levels depend on lower ones, but the survey offers no formal dependency criterion, so if the ordering is only a convenient grouping the structural claim weakens.

Editorial extensions

If this is right

  • Researchers can use the five levels as a shared coordinate system for placing new methods and for spotting which rung a paper actually advances.
  • The survey identifies specific open challenges per level, including occlusions and dynamic motion at Level 1, fluids and topological change at Level 3, and physical contact at Level 4, so the gaps constitute a de facto research agenda.
  • Because each level is claimed to build on the previous one, progress at lower levels, such as unified feed-forward estimation of depth, pose, and tracking, should directly accelerate the higher levels.
  • The concluding discussion of a possible Level 6 implies that the hierarchy is intended to be extensible, with richer spatial intelligence beyond physics-grounded reconstruction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I would read the taxonomy as a claim about research dependencies rather than just a classification; if that is right, work at Level 1 has outsized downstream leverage on everything above it.
  • The taxonomy predicts that near-term breakthroughs will come at the boundaries between levels, for example feed-forward systems that jump from raw video to interaction modeling without explicitly reconstructing every intermediate representation.
  • A testable extension would be to derive the five levels automatically from citation or method-dependency data and compare the empirical clusters with the paper's assignments.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey organizes 4D scene reconstruction from video into five progressive levels: low-level 3D cues, 3D scene components, 4D dynamic scenes, interactions among components, and physical laws. For each level it reviews representative methods, describes paradigm architectures, and closes with challenges and future directions. The paper's central claim, stated in the abstract and Section 1, is that the five levels form a progressive hierarchy of 4D spatial intelligence, with each level building on the previous one.

Significance. The survey is timely and unusually broad: it covers roughly five hundred references, including many 2024–2025 preprints, and the level-by-level structure makes it a convenient entry point for newcomers. The paradigm figures (Figs. 2, 4, 5, 7, 8) are informative, and the maintained project page is a practical asset. The main conceptual contribution is the five-level taxonomy. Its value depends on whether 'progressive levels' is a genuine structural claim about the field; as submitted, that claim is not established, and some method assignments contradict it. With a clearly stated ordering criterion and a consistent assignment rule, the taxonomy could be a useful map of the field; without them, the contribution reduces to a topical grouping.

major comments (3)
  1. [Abstract and Section 1] The central claim that the five levels are 'progressive' is asserted but never defined, and the method assignments contradict a dependency reading. Level 5 entries such as DeepMimic [491], AMP [496], CALM [499], PULSE [88], ASAP [503], and UniPhys [504] in Section 6.1 are physics-based character animation and control methods trained on motion capture or reinforcement learning; they do not consume a Level 4 reconstructed interaction or a Level 3 reconstructed 4D scene. Likewise, HDM [452] and InterTrack [79] in Section 5.1 estimate human-object interaction directly from video frames without first solving Level 3 dynamic-scene reconstruction. If 'progressive' means that higher levels build on the outputs of lower levels, the taxonomy is internally inconsistent. Please either define an explicit ordering criterion (e.g., dependency, representational complexity, or task semantics) and justify each assignment against it, or reframe the contribution as a thematic organization into five topics of increasing complexity.
  2. [Section 6.1 and Scope] Section 6.1, 'Dynamic 4D human simulation with physics,' contains numerous methods that are not reconstructing 4D scenes from video, despite the Scope statement that the survey focuses on approaches for reconstructing 4D scenes from video inputs. DeepMimic, AMP, ASE, CALM, ControlVAE, PULSE, OmniGrasp, HOVER, ASAP, UniPhys, MaskedMimic, SuperPADL, PDP, and CLoSD learn policies from MoCap data, reinforcement learning, or text commands, not from video observations of a scene to be reconstructed. These methods may be relevant as downstream consumers of reconstructions, but as presented they address a different task (character control/synthesis). Please either narrow Level 5 to methods that operate on reconstructed 4D representations as input (e.g., PhysHOI [86], SkillMimic [481], PhysicsNeRF [93], PhyRecon [482]), or add a subsection that explicitly distinguishes reconstruction from control and justifies the inclusion of the control methods.
  3. [Section 1 and throughout] The paper does not provide a systematic, operational criterion for assigning a method to a level, and some assignments are hard to reconcile. For example, 3D tracking is presented in Level 1 as a low-level cue (Section 2.3), but the same capability reappears in Level 3 dynamic-reconstruction methods such as st4rtrack (Section 4.1); the boundary between Level 3 human-centric dynamic modeling (Section 4.2) and Level 4 interactions (Section 5.1) is not drawn in terms of what a method consumes or outputs. A survey taxonomy does not need a formal algorithm, but the central claim of progressivity requires a stated criterion. Please add a short 'taxonomy criteria' paragraph in Section 1 defining the level-assignment rule, and include a table that lists representative methods with their assigned levels.
minor comments (5)
  1. [Scope paragraph] Typo: '4D sptial intelligence' should be '4D spatial intelligence'.
  2. [Section 3.2] The sentence 'An overview of representative approaches in this category is shown in Fig. 2' should refer to Fig. 3; Fig. 2 is the low-level-cues paradigm figure from Section 2.
  3. [Section 5.1] The caption of Fig. 6 names InterDreamer, CIRCLE, and BUDDI, but InterDreamer is not discussed in the body text; please add a cross-reference or remove the name from the caption.
  4. [Section 2.3] The sentence about EgoPoints, 'It opens the door for future works,' is vague; please specify what future directions the new benchmark enables.
  5. [References] References [17] and [55] cite the same NeRF paper twice; please consolidate them into one canonical citation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the paper is a survey whose five-level taxonomy organizes existing methods and is not derived from fitted inputs or load-bearing self-citations.

full rationale

This is a survey paper; its central contribution is a taxonomy (five levels of 4D spatial intelligence) imposed on existing methods. There is no quantity derived from inputs, no fitted parameter that is later 'predicted,' and no uniqueness theorem or ansatz imported from the authors' prior work. The level definitions in the Introduction (Levels 1 through 5) are organizing criteria, not outputs of a calculation; whether the levels are 'progressive' is an interpretive claim justified (or not) by the survey's coverage, not by construction. The paper cites several works by its own authors (e.g., AvatrGo [80], DreamAvatar [98], CrowdMoGen [103], EgoLM [423]), but these are ordinary survey entries describing external methods; none supplies a premise on which the taxonomy depends, and removing them would not change the five-level structure. The reader's 'weakest assumption' and the skeptic note that some methods skip levels are correctness or conceptual-coherence concerns about the taxonomy's ordering claim; they are not circularity because the paper never reduces a conclusion to its own definition or to a self-citation chain. No equations are used to derive a result from an input; equations (1) through (5) describe existing representations (3DGS parameters, SMPL kinematics, HOSNeRF field) and do not feed into the taxonomy. Accordingly, no circular steps are identified.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

As a survey, the paper introduces no free parameters or new entities. The main assumptions are the progressiveness of the taxonomy, the representativeness of the reviewed methods, and the claimed gap in prior surveys.

assumptions (3)
  • domain assumption The five levels form a progressive hierarchy where each level builds on the previous one.
    The abstract and Section 1 label the levels as 'progressive' without defining the dependency relation or providing evidence of progression.
  • domain assumption The surveyed methods are representative and each method is correctly assigned to one level.
    The survey relies on its selection of papers from leading venues and the authors' judgment to assign methods to levels, without a formal classification criterion.
  • domain assumption Existing surveys do not provide a comprehensive hierarchical analysis of 4D scene reconstruction.
    The paper asserts this gap in the introduction but does not systematically compare prior survey taxonomies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Reconstructing 4D Spatial Intelligence: A Survey." pith.science (2026). https://pith.science/paper/IB3TXDV4

@misc{pith2026250721045,
  author       = {Pith},
  title        = {Pith review of: Reconstructing 4D Spatial Intelligence: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IB3TXDV4}},
  note         = {Machine review of arXiv:2507.21045}
}
read the original abstract

Reconstructing 4D spatial intelligence from visual observations has long been a central yet challenging task in computer vision, with broad real-world applications. These range from entertainment domains like movies, where the focus is often on reconstructing fundamental visual elements, to embodied AI, which emphasizes interaction modeling and physical realism. Fueled by rapid advances in 3D representations and deep learning architectures, the field has evolved quickly, outpacing the scope of previous surveys. Additionally, existing surveys rarely offer a comprehensive analysis of the hierarchical structure of 4D scene reconstruction. To address this gap, we present a new perspective that organizes existing methods into five progressive levels of 4D spatial intelligence: (1) Level 1 -- reconstruction of low-level 3D attributes (e.g., depth, pose, and point maps); (2) Level 2 -- reconstruction of 3D scene components (e.g., objects, humans, structures); (3) Level 3 -- reconstruction of 4D dynamic scenes; (4) Level 4 -- modeling of interactions among scene components; and (5) Level 5 -- incorporation of physical laws and constraints. We conclude the survey by discussing the key challenges at each level and highlighting promising directions for advancing toward even richer levels of 4D spatial intelligence. To track ongoing developments, we maintain an up-to-date project page: https://github.com/yukangcao/Awesome-4D-Spatial-Intelligence.

Figures

Figures reproduced from arXiv: 2507.21045 by the authors.

Figure 1
Figure 1. Classification of 4D spatial intelligence by level. Specifically, in this survey, we categorize the methods of reconstructing 3D spatial intelligence from video into five levels: (1) low-level 3D cues, (2) 3D scene components, (3) 4D dynamic scenes, (4) modeling of interactions among scene components, and (5) incorporation of physical laws and constraints. core structure of a 3D scene. Traditionally, this task has b… view at source ↗
Figure 2
Figure 2. The paradigms of methods for reconstructing low-level cues from video input. (I) Video-based depth reconstruction methods re￾cently leverage the diffusion model to obtain the depth maps; (II) Meth￾ods for reconstructing camera pose from video input typically employ the neural network to infer the camera pose based on the encoded image features; (III) 3D tracking methods uses point tracker and transformers to achieve… view at source ↗
Figure 3
Figure 3. The paradigms of methods for reconstructing 3D scene components from video input. 3D reconstruction methods for small￾scale and large-scale scenes often share similar architectures, differing primarily in the spatial extent they handle. As shown in the left panel (Image source: MipNeRF360 [256]), small-scale scenes correspond to the unaffected domain. large-scale scenes additionally incorporate a contracted domain. … view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The paradigms of methods for reconstructing dynamic scenes from video input. Methods in this domain typically adopt one of two strategies for temporal modeling: (I) explicitly incorporating time as an additional input to extend a static 3D representation, or (II) recon…
Figure 5
Figure 5. Figure 5: The illustrations of methods for reconstructing 4D dy￾namic humans from video input. Human-centric dynamic modeling approaches are generally categorized based on their representations: (I) methods that apply SMPL parametric model as their representation to derive the h…
Figure 6
Figure 6. Figure 6: Examples of methods for modeling SMPL-based human￾centric interaction. Image source: InterDreamer [444], CIRCLE [445], and BUDDI [446]. restricting their feasibility and applicability across diverse scenarios. To overcome this challenge, recent approaches like HDM [452…
Figure 7
Figure 7. Figure 7: The paradigms of methods for reconstructing appearance￾rich human-centric interaction. These methods generally build on SMPL-based linear blend skinning (LBS) deformation, extending the human body skeleton to include interacted objects. An example result is shown in th…
Figure 8
Figure 8. Figure 8: The paradigms of methods for inferring physically grounded 3D spatial understanding from videos. (I) Physical dynamic human modeling methods learn motion policies from real-world captures of human-object interactions, enabling deployment in simulators and trans￾fer to …

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. One Video, One World: Turning Monocular Video into Physical 4D Scenes

    cs.CV 2026-06 unverdicted novelty 8.0 of 10

    OVOW reconstructs instance-level, simulation-ready 4D mesh scenes from monocular video via a four-stage training-free pipeline and introduces a new benchmark for structured Video-to-4D evaluation.

  2. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

    cs.CV 2026-07 accept novelty 7.0 of 10

    ACE-Data-0 is a 150-hour home HOI dataset with millisecond-synced ego/exo video, mocap body/hands, object 6-DoF, audio, and tactile signals, plus a three-level benchmark exposing large SOTA gaps.

  3. CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

    cs.CV 2026-01 unverdicted novelty 7.0 of 10

    CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.

  4. Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    HA-HOI produces physically plausible 4D HOI animations from monocular videos by anchoring object reconstruction to human motion and refining the result in a physics-based humanoid-object simulator.

  5. Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Stitch4D reconstructs coherent 4D urban scenes from sparse non-overlapping camera placements by synthesizing bridge views and enforcing inter-location spatio-temporal consistency.

  6. PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception

    cs.CV 2025-10 unverdicted novelty 6.0 of 10

    PAGE-4D is a feedforward extension of VGGT that uses a dynamics-aware aggregator and mask to disentangle pose estimation from geometry reconstruction in videos with moving objects.

  7. PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A fine-tuned VGGT with a learned dynamics mask improves camera pose, depth, and point-cloud reconstruction on dynamic-scene benchmarks over the original static-scene model.

  8. Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation

    cs.CV 2026-04 conditional novelty 5.0 of 10

    Synthesizing intermediate bridge views between sparse, non-overlapping urban cameras and jointly optimizing them stabilizes 4D reconstruction where dense-view methods collapse.

  9. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

Reference graph

Works this paper leans on

300 extracted references · 10 canonical work pages · cited by 7 Pith papers

  1. [88]

    Universal humanoid motion representations for physics-based control,

    Z. Luo, J. Cao, J. Merel, A. Winkler, J. Huang, K. M. Kitani, and W. Xu, “Universal humanoid motion representations for physics-based control,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https: //openreview.net/forum?id=OrOd8PxOO2

  2. [79]

    Intertrack: Tracking hu- man object interaction without object templates,

    X. Xie, J. E. Lenssen, and G. Pons-Moll, “Intertrack: Tracking hu- man object interaction without object templates,” arXiv preprint arXiv:2408.13953, 2024

  3. [86]

    Physhoi: Physics-based imitation of dynamic human-object in- teraction,

    Y. Wang, J. Lin, A. Zeng, Z. Luo, J. Zhang, and L. Zhang, “Physhoi: Physics-based imitation of dynamic human-object in- teraction,” arXiv preprint arXiv:2312.04393, 2023

  4. [93]

    Physicsnerf: Physics-guided 3d reconstruction from sparse views,

    M. R. Barhdadi, H. Kurban, and H. Alnuweiri, “Physicsnerf: Physics-guided 3d reconstruction from sparse views,” 2025

  5. [1]

    Neural point-based graphics,

    K.-A. Aliev, A. Sevastopolsky, M. Kolos, D. Ulyanov, and V . Lem- pitsky, “Neural point-based graphics,” in European conference on computer vision. Springer, 2020, pp. 696–712

  6. [2]

    Deep video portraits,

    H. Kim, P . Garrido, A. Tewari, W. Xu, J. Thies, M. Niessner, P . P´erez, C. Richardt, M. Zollh ¨ofer, and C. Theobalt, “Deep video portraits,” ACM transactions on graphics (TOG) , vol. 37, no. 4, pp. 1–14, 2018

  7. [3]

    Fov-nerf: Foveated neural radiance fields for virtual reality,

    N. Deng, Z. He, J. Ye, B. Duinkharjav, P . Chakravarthula, X. Yang, and Q. Sun, “Fov-nerf: Foveated neural radiance fields for virtual reality,” IEEE Transactions on Visualization and Computer Graphics , vol. 28, no. 11, pp. 3854–3864, 2022

  8. [4]

    Instant-3d: Instant neural radiance field training towards on-device ar/vr 3d reconstruction,

    S. Li, C. Li, W. Zhu, B. Yu, Y. Zhao, C. Wan, H. You, H. Shi, and Y. Lin, “Instant-3d: Instant neural radiance field training towards on-device ar/vr 3d reconstruction,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , 2023, pp. 1–13

Show all 300 references
  1. [5]

    Aligning cyber space with physical world: A comprehensive survey on embodied ai,

    Y. Liu, W. Chen, Y. Bai, X. Liang, G. Li, W. Gao, and L. Lin, “Aligning cyber space with physical world: A comprehensive survey on embodied ai,” arXiv preprint arXiv:2407.06886, 2024

  2. [6]

    The essential role of causality in foundation world models for embodied ai,

    T. Gupta, W. Gong, C. Ma, N. Pawlowski, A. Hilmkil, M. Scetbon, M. Rigter, A. Famoti, A. J. Llorens, J. Gaoet al., “The essential role of causality in foundation world models for embodied ai,” arXiv preprint arXiv:2402.06665, 2024

  3. [7]

    An embodied generalist agent in 3d world,

    J. Huang, S. Yong, X. Ma, X. Linghu, P . Li, Y. Wang, Q. Li, S.-C. Zhu, B. Jia, and S. Huang, “An embodied generalist agent in 3d world,” arXiv preprint arXiv:2311.12871, 2023

  4. [8]

    3d-vla: A 3d vision-language-action generative world model,

    H. Zhen, X. Qiu, P . Chen, J. Yang, X. Yan, Y. Du, Y. Hong, and C. Gan, “3d-vla: A 3d vision-language-action generative world model,” arXiv preprint arXiv:2403.09631, 2024

  5. [9]

    Recent advances in 3d gaussian splatting,

    T. Wu, Y.-J. Yuan, L.-X. Zhang, J. Yang, Y.-P . Cao, L.-Q. Yan, and L. Gao, “Recent advances in 3d gaussian splatting,” Computa- tional Visual Media, vol. 10, no. 4, pp. 613–642, 2024

  6. [10]

    3d gaussian splatting as new era: A survey,

    B. Fei, J. Xu, R. Zhang, Q. Zhou, W. Yang, and Y. He, “3d gaussian splatting as new era: A survey,” IEEE Transactions on Visualization and Computer Graphics, 2024

  7. [11]

    Stereo matching algorithm based on deep learning: A survey,

    M. S. Hamid, N. Abd Manap, R. A. Hamzah, and A. F. Kadmin, “Stereo matching algorithm based on deep learning: A survey,” Journal of King Saud University-Computer and Information Sciences , vol. 34, no. 5, pp. 1663–1673, 2022

  8. [12]

    Review of stereo matching algorithms based on deep learning,

    K. Zhou, X. Meng, and B. Cheng, “Review of stereo matching algorithms based on deep learning,” Computational intelligence and neuroscience, vol. 2020, no. 1, p. 8562323, 2020

  9. [13]

    A sur- vey on deep learning techniques for stereo-based depth estima- tion,

    H. Laga, L. V . Jospin, F. Boussaid, and M. Bennamoun, “A sur- vey on deep learning techniques for stereo-based depth estima- tion,” IEEE transactions on pattern analysis and machine intelligence , vol. 44, no. 4, pp. 1738–1764, 2020

  10. [14]

    Nerf: Neural radiance field in 3d vision, a comprehensive review,

    K. Gao, Y. Gao, H. He, D. Lu, L. Xu, and J. Li, “Nerf: Neural radiance field in 3d vision, a comprehensive review,” arXiv preprint arXiv:2210.00379, 2022

  11. [15]

    3d gaussian splatting: Survey, technologies, challenges, and opportunities,

    Y. Bao, T. Ding, J. Huo, Y. Liu, Y. Li, W. Li, Y. Gao, and J. Luo, “3d gaussian splatting: Survey, technologies, challenges, and opportunities,” IEEE Transactions on Circuits and Systems for Video Technology, 2025

  12. [16]

    Nerf in robotics: A survey,

    G. Wang, L. Pan, S. Peng, S. Liu, C. Xu, Y. Miao, W. Zhan, M. Tomizuka, M. Pollefeys, and H. Wang, “Nerf in robotics: A survey,” arXiv preprint arXiv:2405.01333, 2024

  13. [17]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P . P . Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in European conference on computer vision. Springer, 2020, pp. 405–421

  14. [18]

    Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape synthesis,

    T. Shen, J. Gao, K. Yin, M.-Y. Liu, and S. Fidler, “Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape synthesis,” 2021

  15. [19]

    3d gaussian splatting for real-time radiance field rendering,

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,” ACM Transactions on Graphics (ToG), vol. 42, no. 4, pp. 1–14, 2023

  16. [20]

    Video diffusion models,

    J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” Advances in Neural Information Processing Systems, vol. 35, pp. 8633–8646, 2022

  17. [21]

    Imagen video: High definition video generation with diffusion models,

    J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P . Kingma, B. Poole, M. Norouzi, D. J. Fleet et al., “Imagen video: High definition video generation with diffusion models,” arXiv preprint arXiv:2210.02303, 2022

  18. [22]

    Stable video diffusion: Scaling latent video diffusion models to large datasets,

    A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V . Voleti, A. Letts et al. , “Stable video diffusion: Scaling latent video diffusion models to large datasets,” arXiv preprint arXiv:2311.15127, 2023

  19. [23]

    Sift: Predicting amino acid changes that affect protein function,

    P . C. Ng and S. Henikoff, “Sift: Predicting amino acid changes that affect protein function,” Nucleic acids research, vol. 31, no. 13, pp. 3812–3814, 2003

  20. [24]

    R2d2: Reliable and repeatable detector and descriptor,

    J. Revaud, C. De Souza, M. Humenberger, and P . Weinzaepfel, “R2d2: Reliable and repeatable detector and descriptor,”Advances in neural information processing systems, vol. 32, 2019

  21. [25]

    Superpoint: Self- supervised interest point detection and description,

    D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superpoint: Self- supervised interest point detection and description,” in Proceed- ings of the IEEE conference on computer vision and pattern recognition workshops, 2018, pp. 224–236

  22. [26]

    Superglue: Learning feature matching with graph neural net- works,

    P .-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superglue: Learning feature matching with graph neural net- works,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 4938–4947

  23. [27]

    Loftr: Detector- free local feature matching with transformers,

    J. Sun, Z. Shen, Y. Wang, H. Bao, and X. Zhou, “Loftr: Detector- free local feature matching with transformers,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 8922–8931

  24. [28]

    Neural-guided ransac: Learn- ing where to sample model hypotheses,

    E. Brachmann and C. Rother, “Neural-guided ransac: Learn- ing where to sample model hypotheses,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 4322–4331

  25. [29]

    Lightglue: Local feature matching at light speed,

    P . Lindenberger, P .-E. Sarlin, and M. Pollefeys, “Lightglue: Local feature matching at light speed,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 17 627–17 638

  26. [30]

    Affineglue: Joint matching and robust estimation,

    D. Barath, D. Mishkin, L. Cavalli, P .-E. Sarlin, P . Hruby, and M. Pollefeys, “Affineglue: Joint matching and robust estimation,” arXiv preprint arXiv:2307.15381, 2023

  27. [31]

    Structure-from-motion revis- ited,

    J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revis- ited,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 4104–4113

  28. [32]

    A survey of structure from motion*

    O. ¨Ozyes ¸il, V . Voroninski, R. Basri, and A. Singer, “A survey of structure from motion*.” Acta Numerica, vol. 26, pp. 305–364, 2017

  29. [33]

    Structure from motion photogrammetry in forestry: A review,

    J. Iglhaut, C. Cabo, S. Puliti, L. Piermattei, J. O’Connor, and J. Rosette, “Structure from motion photogrammetry in forestry: A review,” Current Forestry Reports , vol. 5, no. 3, pp. 155–168, 2019

  30. [34]

    Pixel- perfect structure-from-motion with featuremetric refinement,

    P . Lindenberger, P .-E. Sarlin, V . Larsson, and M. Pollefeys, “Pixel- perfect structure-from-motion with featuremetric refinement,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5987–5997

  31. [35]

    Bundle adjustment in the large,

    S. Agarwal, N. Snavely, S. M. Seitz, and R. Szeliski, “Bundle adjustment in the large,” in European conference on computer vision. Springer, 2010, pp. 29–42

  32. [36]

    Bundle adjustment rules,

    C. Engels, H. Stew ´enius, and D. Nist ´er, “Bundle adjustment rules,” Photogrammetric computer vision, vol. 2, no. 32, 2006

  33. [37]

    Robust bundle adjustment revisited,

    C. Zach, “Robust bundle adjustment revisited,” in European Con- ference on Computer Vision. Springer, 2014, pp. 772–787

  34. [38]

    Bundle adjustment—a modern synthesis,

    B. Triggs, P . F. McLauchlan, R. I. Hartley, and A. W. Fitzgibbon, “Bundle adjustment—a modern synthesis,” in International work- shop on vision algorithms. Springer, 1999, pp. 298–372

  35. [39]

    Pixelwise view selection for unstructured multi-view stereo,

    J. L. Sch ¨onberger, E. Zheng, M. Pollefeys, and J.-M. Frahm, “Pixelwise view selection for unstructured multi-view stereo,” in European Conference on Computer Vision (ECCV), 2016. 16

  36. [40]

    Visibility-aware multi-view stereo network,

    J. Zhang, Y. Yao, S. Li, Z. Luo, and T. Fang, “Visibility-aware multi-view stereo network,” British Machine Vision Conference (BMVC), 2020

  37. [41]

    Cost volume pyramid based depth inference for multi-view stereo,

    J. Yang, W. Mao, J. M. Alvarez, and M. Liu, “Cost volume pyramid based depth inference for multi-view stereo,” in The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020

  38. [42]

    Patchmatchnet: Learned multi-view patchmatch stereo,

    F. Wang, S. Galliani, C. Vogel, P . Speciale, and M. Pollefeys, “Patchmatchnet: Learned multi-view patchmatch stereo,” 2021

  39. [43]

    Cascade cost volume for high-resolution multi-view stereo and stereo matching,

    X. Gu, Z. Fan, S. Zhu, Z. Dai, F. Tan, and P . Tan, “Cascade cost volume for high-resolution multi-view stereo and stereo matching,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2495–2504

  40. [44]

    Dust3r: Geometric 3d vision made easy,

    S. Wang, V . Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 20 697–20 709

  41. [45]

    Mast3r-slam: Real- time dense slam with 3d reconstruction priors,

    R. Murai, E. Dexheimer, and A. J. Davison, “Mast3r-slam: Real- time dense slam with 3d reconstruction priors,” arXiv preprint arXiv:2412.12392, 2024

  42. [46]

    Monst3r: A simple approach for estimating geometry in the presence of motion,

    J. Zhang, C. Herrmann, J. Hur, V . Jampani, T. Darrell, F. Cole, D. Sun, and M.-H. Yang, “Monst3r: A simple approach for estimating geometry in the presence of motion,” arXiv preprint arXiv:2410.03825, 2024

  43. [47]

    Align3r: Aligned monocular depth estimation for dynamic videos,

    J. Lu, T. Huang, P . Li, Z. Dou, C. Lin, Z. Cui, Z. Dong, S.-K. Yeung, W. Wang, and Y. Liu, “Align3r: Aligned monocular depth estimation for dynamic videos,” arXiv preprint arXiv:2412.03079 , 2024

  44. [48]

    Fast3r: Towards 3d recon- struction of 1000+ images in one forward pass,

    J. Yang, A. Sax, K. J. Liang, M. Henaff, H. Tang, A. Cao, J. Chai, F. Meier, and M. Feiszli, “Fast3r: Towards 3d recon- struction of 1000+ images in one forward pass,” arXiv preprint arXiv:2501.13928, 2025

  45. [49]

    Transformer in transformer,

    K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang, “Transformer in transformer,” Advances in neural information processing systems , vol. 34, pp. 15 908–15 919, 2021

  46. [50]

    A survey on vision transformer,

    K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xuet al., “A survey on vision transformer,”IEEE transactions on pattern analysis and machine intelligence , vol. 45, no. 1, pp. 87–110, 2022

  47. [51]

    Reformer: The efficient transformer,

    N. Kitaev, Ł. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” arXiv preprint arXiv:2001.04451, 2020

  48. [52]

    Point trans- former,

    H. Zhao, L. Jiang, J. Jia, P . H. Torr, and V . Koltun, “Point trans- former,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 16 259–16 268

  49. [53]

    Image transformer,

    N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran, “Image transformer,” in International conference on machine learning. PMLR, 2018, pp. 4055–4064

  50. [54]

    Vggt: Visual geometry grounded transformer,

    J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” in Proceedings of the Computer Vision and Pattern Recognition Confer- ence, 2025, pp. 5294–5306

  51. [55]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P . P . Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021

  52. [56]

    3d gaus- sian splatting for real-time radiance field rendering,

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaus- sian splatting for real-time radiance field rendering,” ACM Trans- actions on Graphics , vol. 42, no. 4, July 2023. [Online]. Available: https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/

  53. [57]

    Flexible isosurface extraction for gradient-based mesh optimization,

    T. Shen, J. Munkberg, J. Hasselgren, K. Yin, Z. Wang, W. Chen, Z. Gojcic, S. Fidler, N. Sharp, and J. Gao, “Flexible isosurface extraction for gradient-based mesh optimization,” ACM Trans. Graph. , vol. 42, no. 4, jul 2023. [Online]. Available: https://doi.org/10.1145/3592430

  54. [58]

    Nerfies: Deformable neural radi- ance fields,

    K. Park, U. Sinha, J. T. Barron, S. Bouaziz, D. B. Goldman, S. M. Seitz, and R. Martin-Brualla, “Nerfies: Deformable neural radi- ance fields,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 5865–5874

  55. [59]

    Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monocular video,

    E. Tretschk, A. Tewari, V . Golyanik, M. Zollh ¨ofer, C. Lassner, and C. Theobalt, “Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monocular video,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2...

  56. [60]

    Hypernerf: A higher-dimensional representation for topologically varying neu- ral radiance fields,

    K. Park, U. Sinha, P . Hedman, J. T. Barron, S. Bouaziz, D. B. Goldman, R. Martin-Brualla, and S. M. Seitz, “Hypernerf: A higher-dimensional representation for topologically varying neu- ral radiance fields,” arXiv preprint arXiv:2106.13228, 2021

  57. [61]

    Dylin: Making light field networks dynamic,

    H. Yu, J. Julin, Z. A. Milacski, K. Niinuma, and L. A. Jeni, “Dylin: Making light field networks dynamic,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 12 397–12 406

  58. [62]

    Spectromotion: Dynamic 3d reconstruction of specular scenes,

    C.-D. Fan, C.-W. Chang, Y.-R. Liu, J.-Y. Lee, J.-L. Huang, Y.-C. Tseng, and Y.-L. Liu, “Spectromotion: Dynamic 3d reconstruction of specular scenes,” arXiv preprint arXiv:2410.17249, 2024

  59. [63]

    Neural scene flow fields for space-time view synthesis of dynamic scenes,

    Z. Li, S. Niklaus, N. Snavely, and O. Wang, “Neural scene flow fields for space-time view synthesis of dynamic scenes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 6498–6508

  60. [64]

    Dynamic view synthesis from dynamic monocular video,

    C. Gao, A. Saraf, J. Kopf, and J.-B. Huang, “Dynamic view synthesis from dynamic monocular video,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 5712–5721

  61. [65]

    Decoupling dynamic monocular videos for dynamic view synthesis,

    M. You and J. Hou, “Decoupling dynamic monocular videos for dynamic view synthesis,” IEEE Transactions on Visualization and Computer Graphics, 2024

  62. [66]

    Neural trajec- tory fields for dynamic novel view synthesis,

    C. Wang, B. Eckart, S. Lucey, and O. Gallo, “Neural trajec- tory fields for dynamic novel view synthesis,” arXiv preprint arXiv:2105.05994, 2021

  63. [67]

    Neural scene chronology,

    H. Lin, Q. Wang, R. Cai, S. Peng, H. Averbuch-Elor, X. Zhou, and N. Snavely, “Neural scene chronology,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 20 752–20 761

  64. [68]

    Darenerf: Direction-aware representation for dynamic scenes,

    A. Lou, B. Planche, Z. Gao, Y. Li, T. Luan, H. Ding, T. Chen, J. Noble, and Z. Wu, “Darenerf: Direction-aware representation for dynamic scenes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5031–5042

  65. [69]

    Emernerf: Emer- gent spatial-temporal scene decomposition via self-supervision,

    J. Yang, B. Ivanovic, O. Litany, X. Weng, S. W. Kim, B. Li, T. Che, D. Xu, S. Fidler, M. Pavone et al. , “Emernerf: Emer- gent spatial-temporal scene decomposition via self-supervision,” arXiv preprint arXiv:2311.02077, 2023

  66. [70]

    Gravity-aware monocular 3d human-object reconstruction,

    R. Dabral, S. Shimada, A. Jain, C. Theobalt, and V . Golyanik, “Gravity-aware monocular 3d human-object reconstruction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 12 365–12 374

  67. [71]

    D3d-hoi: Dynamic 3d human-object interactions from videos,

    X. Xu, H. Joo, G. Mori, and M. Savva, “D3d-hoi: Dynamic 3d human-object interactions from videos,” arXiv preprint arXiv:2108.08420, 2021

  68. [72]

    Behave: Dataset and method for tracking human object interactions,

    B. L. Bhatnagar, X. Xie, I. A. Petrov, C. Sminchisescu, C. Theobalt, and G. Pons-Moll, “Behave: Dataset and method for tracking human object interactions,” in Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , 2022, pp. 15 935– 15 946

  69. [73]

    Intercap: Joint markerless 3d tracking of humans and objects in interaction,

    Y. Huang, O. Taheri, M. J. Black, and D. Tzionas, “Intercap: Joint markerless 3d tracking of humans and objects in interaction,” in DAGM German Conference on Pattern Recognition. Springer, 2022, pp. 281–299

  70. [74]

    Full-body articulated human-object inter- action,

    N. Jiang, T. Liu, Z. Cao, J. Cui, Z. Zhang, Y. Chen, H. Wang, Y. Zhu, and S. Huang, “Full-body articulated human-object inter- action,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 9365–9376

  71. [75]

    Stackflow: Monocular human-object reconstruction by stacked normalizing flow with offset,

    C. Huo, Y. Shi, Y. Ma, L. Xu, J. Yu, and J. Wang, “Stackflow: Monocular human-object reconstruction by stacked normalizing flow with offset,” arXiv preprint arXiv:2407.20545, 2024

  72. [76]

    Monocular human-object recon- struction in the wild,

    C. Huo, Y. Shi, and J. Wang, “Monocular human-object recon- struction in the wild,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 5547–5555

  73. [77]

    I’m hoi: Inertia-aware monocular capture of 3d human- object interactions,

    C. Zhao, J. Zhang, J. Du, Z. Shan, J. Wang, J. Yu, J. Wang, and L. Xu, “I’m hoi: Inertia-aware monocular capture of 3d human- object interactions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 729–741

  74. [78]

    Sv4d: Dynamic 3d content generation with multi-frame and multi-view consistency,

    Y. Xie, C.-H. Yao, V . Voleti, H. Jiang, and V . Jampani, “Sv4d: Dynamic 3d content generation with multi-frame and multi-view consistency,” arXiv preprint arXiv:2407.17470, 2024

  75. [80]

    Avatargo: Zero-shot 4d human-object interaction generation and anima- tion,

    Y. Cao, L. Pan, K. Han, K.-Y. K. Wong, and Z. Liu, “Avatargo: Zero-shot 4d human-object interaction generation and anima- tion,” arXiv preprint arXiv:2410.07164, 2024

  76. [81]

    The one where they reconstructed 3d humans and environments in tv shows,

    G. Pavlakos, E. Weber, M. Tancik, and A. Kanazawa, “The one where they reconstructed 3d humans and environments in tv shows,” in European Conference on Computer Vision . Springer, 2022, pp. 732–749. 17

  77. [82]

    Joint optimization for 4d human-scene reconstruction in the wild,

    Z. Liu, J. Lin, W. Wu, and B. Zhou, “Joint optimization for 4d human-scene reconstruction in the wild,” arXiv preprint arXiv:2501.02158, 2025

  78. [83]

    Odhsr: Online dense 3d reconstruction of humans and scenes from monocular videos,

    Z. Zhang, M. Kaufmann, L. Xue, J. Song, and M. R. Oswald, “Odhsr: Online dense 3d reconstruction of humans and scenes from monocular videos,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 824–21 835

  79. [84]

    Hosnerf: Dynamic human-object-scene neu- ral radiance fields from a single video,

    J.-W. Liu, Y.-P . Cao, T. Yang, Z. Xu, J. Keppo, Y. Shan, X. Qie, and M. Z. Shou, “Hosnerf: Dynamic human-object-scene neu- ral radiance fields from a single video,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 18 483–18 494

  80. [85]

    Neuman: Neural human radiance field from a single video,

    W. Jiang, K. M. Yi, G. Samei, O. Tuzel, and A. Ranjan, “Neuman: Neural human radiance field from a single video,” in European Conference on Computer Vision. Springer, 2022, pp. 402–418

  81. [87]

    Perpetual humanoid control for real-time simulated avatars,

    Z. Luo, J. Cao, A. W. Winkler, K. Kitani, and W. Xu, “Perpetual humanoid control for real-time simulated avatars,” in Interna- tional Conference on Computer Vision (ICCV), 2023

  82. [89]

    Isaac gym: High performance gpu-based physics simulation for robot learning,

    V . Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa et al. , “Isaac gym: High performance gpu-based physics simulation for robot learning,” arXiv preprint arXiv:2108.10470, 2021

  83. [90]

    Reinforcement learning,

    M. A. Wiering and M. Van Otterlo, “Reinforcement learning,” Adaptation, learning, and optimization, vol. 12, no. 3, p. 729, 2012

  84. [91]

    Reinforcement learning: A survey,

    L. P . Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” Journal of artificial intelligence research , vol. 4, pp. 237–285, 1996

  85. [92]

    Reinforcement learning,

    R. S. Sutton, A. G. Barto et al., “Reinforcement learning,” Journal of Cognitive Neuroscience, vol. 11, no. 1, pp. 126–134, 1999

  86. [94]

    Pbr-nerf: Inverse rendering with physics-based neural fields,

    S. Wu, S. Basu, T. Broedermann, L. V . Gool, and C. Sakaridis, “Pbr-nerf: Inverse rendering with physics-based neural fields,” 2025

  87. [95]

    Cast: Component-aligned 3d scene reconstruction from an rgb image,

    K. Yao, L. Zhang, X. Yan, Y. Zeng, Q. Zhang, W. Yang, L. Xu, J. Gu, and J. Yu, “Cast: Component-aligned 3d scene reconstruction from an rgb image,” 2025

  88. [96]

    Sv3d: Novel multi- view synthesis and 3d generation from a single image using latent video diffusion,

    V . Voleti, C.-H. Yao, M. Boss, A. Letts, D. Pankratz, D. Tochilkin, C. Laforte, R. Rombach, and V . Jampani, “Sv3d: Novel multi- view synthesis and 3d generation from a single image using latent video diffusion,” in European Conference on Computer Vision. Springer, 2025, pp. 439–457

  89. [97]

    V3d: Video diffusion models are effective 3d generators,

    Z. Chen, Y. Wang, F. Wang, Z. Wang, and H. Liu, “V3d: Video diffusion models are effective 3d generators,” arXiv preprint arXiv:2403.06738, 2024

  90. [98]

    Drea- mavatar: Text-and-shape guided 3d human avatar generation via diffusion models,

    Y. Cao, Y.-P . Cao, K. Han, Y. Shan, and K.-Y. K. Wong, “Drea- mavatar: Text-and-shape guided 3d human avatar generation via diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 958–968

  91. [99]

    4d-fy: Text-to-4d generation using hybrid score distillation sampling,

    S. Bahmani, I. Skorokhodov, V . Rong, G. Wetzstein, L. Guibas, P . Wonka, S. Tulyakov, J. J. Park, A. Tagliasacchi, and D. B. Lin- dell, “4d-fy: Text-to-4d generation using hybrid score distillation sampling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...

  92. [100]

    Tc4d: Trajectory- conditioned text-to-4d generation,

    S. Bahmani, X. Liu, W. Yifan, I. Skorokhodov, V . Rong, Z. Liu, X. Liu, J. J. Park, S. Tulyakov, G. Wetzsteinet al., “Tc4d: Trajectory- conditioned text-to-4d generation,” in European Conference on Computer Vision. Springer, 2024, pp. 53–72

  93. [101]

    4diffusion: Multi-view video diffusion model for 4d genera- tion,

    H. Zhang, X. Chen, Y. Wang, X. Liu, Y. Wang, and Y. Qiao, “4diffusion: Multi-view video diffusion model for 4d genera- tion,” Advances in Neural Information Processing Systems , vol. 37, pp. 15 272–15 295, 2024

  94. [102]

    Diffusion4d: Fast spatial-temporal consis- tent 4d generation via video diffusion models,

    H. Liang, Y. Yin, D. Xu, H. Liang, Z. Wang, K. N. Plataniotis, Y. Zhao, and Y. Wei, “Diffusion4d: Fast spatial-temporal consis- tent 4d generation via video diffusion models,” arXiv preprint arXiv:2405.16645, 2024

  95. [103]

    Crowdmo- gen: Zero-shot text-driven collective motion generation,

    Y. Cao, X. Guo, M. Zhang, H. Xie, C. Gu, and Z. Liu, “Crowdmo- gen: Zero-shot text-driven collective motion generation,” arXiv preprint arXiv:2407.06188, 2024

  96. [104]

    Guide3d: Create 3d avatars from text and image guidance,

    Y. Cao, Y.-P . Cao, K. Han, Y. Shan, and K.-Y. K. Wong, “Guide3d: Create 3d avatars from text and image guidance,” arXiv preprint arXiv:2308.09705, 2023

  97. [105]

    Advances in 4d generation: A survey,

    Q. Miao, K. Li, J. Quan, Z. Min, S. Ma, Y. Xu, Y. Yang, and Y. Luo, “Advances in 4d generation: A survey,” arXiv preprint arXiv:2503.14501, 2025

  98. [106]

    Generative ai meets 3d: A survey on text-to-3d in aigc era,

    C. Li, C. Zhang, A. Waghwase, L.-H. Lee, F. Rameau, Y. Yang, S.-H. Bae, and C. S. Hong, “Generative ai meets 3d: A survey on text-to-3d in aigc era,” arXiv preprint arXiv:2305.06131, 2023

  99. [107]

    A comprehensive survey on 3d content generation,

    J. Liu, X. Huang, T. Huang, L. Chen, Y. Hou, S. Tang, Z. Liu, W. Ouyang, W. Zuo, J. Jiang et al., “A comprehensive survey on 3d content generation,” arXiv preprint arXiv:2402.01166, 2024

  100. [108]

    Advances in 3d generation: A survey,

    X. Li, Q. Zhang, D. Kang, W. Cheng, Y. Gao, J. Zhang, Z. Liang, J. Liao, Y.-P . Cao, and Y. Shan, “Advances in 3d generation: A survey,” arXiv preprint arXiv:2401.17807, 2024

  101. [109]

    Differentiable rendering: A survey,

    H. Kato, D. Beker, M. Morariu, T. Ando, T. Matsuoka, W. Kehl, and A. Gaidon, “Differentiable rendering: A survey,” arXiv preprint arXiv:2006.12057, 2020

  102. [110]

    A survey of non- rigid 3d registration,

    B. Deng, Y. Yao, R. M. Dyke, and J. Zhang, “A survey of non- rigid 3d registration,” in Computer Graphics Forum, vol. 41, no. 2. Wiley Online Library, 2022, pp. 559–589

  103. [111]

    Inference time optimization using branchynet partitioning,

    R. G. Pacheco and R. S. Couto, “Inference time optimization using branchynet partitioning,” in 2020 IEEE Symposium on Computers and Communications (ISCC). IEEE, 2020, pp. 1–6

  104. [112]

    Geonet: Unsupervised learning of dense depth, optical flow and camera pose,

    Z. Yin and J. Shi, “Geonet: Unsupervised learning of dense depth, optical flow and camera pose,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 1983–1992

  105. [113]

    Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras,

    A. Gordon, H. Li, R. Jonschkowski, and A. Angelova, “Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras,” in Proceedings of the IEEE/CVF interna- tional conference on computer vision, 2019, pp. 8977–8986

  106. [114]

    Unsupervised scale-consistent depth and ego-motion learning from monocular video,

    J. Bian, Z. Li, N. Wang, H. Zhan, C. Shen, M.-M. Cheng, and I. Reid, “Unsupervised scale-consistent depth and ego-motion learning from monocular video,” Advances in neural information processing systems, vol. 32, 2019

  107. [115]

    Consis- tent video depth estimation,

    X. Luo, J.-B. Huang, R. Szeliski, K. Matzen, and J. Kopf, “Consis- tent video depth estimation,” ACM Transactions on Graphics (ToG), vol. 39, no. 4, pp. 71–1, 2020

  108. [116]

    Self-supervised learn- ing with geometric constraints in monocular video: Connecting flow, depth, and camera,

    Y. Chen, C. Schmid, and C. Sminchisescu, “Self-supervised learn- ing with geometric constraints in monocular video: Connecting flow, depth, and camera,” in Proceedings of the IEEE/CVF interna- tional conference on computer vision, 2019, pp. 7063–7072

  109. [117]

    Multi-view depth estimation using epipolar spatio-temporal networks,

    X. Long, L. Liu, W. Li, C. Theobalt, and W. Wang, “Multi-view depth estimation using epipolar spatio-temporal networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 8258–8267

  110. [118]

    Multi-frame self-supervised depth with transformers,

    V . Guizilini, R. Ambrus , , D. Chen, S. Zakharov, and A. Gaidon, “Multi-frame self-supervised depth with transformers,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 160–170

  111. [119]

    The temporal opportunist: Self-supervised multi-frame monocular depth,

    J. Watson, O. Mac Aodha, V . Prisacariu, G. Brostow, and M. Fir- man, “The temporal opportunist: Self-supervised multi-frame monocular depth,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 1164–1174

  112. [120]

    Simplerecon: 3d reconstruction without 3d convolu- tions,

    M. Sayed, J. Gibson, J. Watson, V . Prisacariu, M. Firman, and C. Godard, “Simplerecon: 3d reconstruction without 3d convolu- tions,” in European Conference on Computer Vision. Springer, 2022, pp. 1–19

  113. [121]

    Video depth es- timation by fusing flow-to-depth proposals,

    J. Xie, C. Lei, Z. Li, L. E. Li, and Q. Chen, “Video depth es- timation by fusing flow-to-depth proposals,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2020, pp. 10 100–10 107

  114. [122]

    Temporally consistent depth prediction with flow-guided memory units,

    C. Eom, H. Park, and B. Ham, “Temporally consistent depth prediction with flow-guided memory units,” IEEE Transactions on Intelligent Transportation Systems, vol. 21, no. 11, pp. 4626–4636, 2019

  115. [123]

    Don’t forget the past: Recurrent depth estimation from monocular video,

    V . Patil, W. Van Gansbeke, D. Dai, and L. Van Gool, “Don’t forget the past: Recurrent depth estimation from monocular video,” IEEE Robotics and Automation Letters , vol. 5, no. 4, pp. 6813–6820, 2020

  116. [124]

    Exploiting temporal consistency for real-time video depth estimation,

    H. Zhang, C. Shen, Y. Li, Y. Cao, Y. Liu, and Y. Yan, “Exploiting temporal consistency for real-time video depth estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1725–1734

  117. [125]

    Less is more: Consistent video depth estimation with masked frames 18 modeling,

    Y. Wang, Z. Pan, X. Li, Z. Cao, K. Xian, and J. Zhang, “Less is more: Consistent video depth estimation with masked frames 18 modeling,” in Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 6347–6358

  118. [126]

    Mamo: Leveraging memory and attention for monocular video depth estimation,

    R. Yasarla, H. Cai, J. Jeong, Y. Shi, R. Garrepalli, and F. Porikli, “Mamo: Leveraging memory and attention for monocular video depth estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 8754–8764

  119. [127]

    Monovit: Self-supervised monocular depth estimation with a vision transformer,

    C. Zhao, Y. Zhang, M. Poggi, F. Tosi, X. Guo, Z. Zhu, G. Huang, Y. Tang, and S. Mattoccia, “Monovit: Self-supervised monocular depth estimation with a vision transformer,” in 2022 international conference on 3D vision (3DV). IEEE, 2022, pp. 668–678

  120. [128]

    Temporally consistent online depth estima- tion in dynamic scenes,

    Z. Li, W. Ye, D. Wang, F. X. Creighton, R. H. Taylor, G. Venkatesh, and M. Unberath, “Temporally consistent online depth estima- tion in dynamic scenes,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision , 2023, pp. 3018–3027

  121. [129]

    Deepv2d: Video to depth with differ- entiable structure from motion,

    Z. Teed and J. Deng, “Deepv2d: Video to depth with differ- entiable structure from motion,” in International Conference on Learning Representations, 2020

  122. [130]

    Neural video depth stabilizer,

    Y. Wang, M. Shi, J. Li, Z. Huang, Z. Cao, J. Zhang, K. Xian, and G. Lin, “Neural video depth stabilizer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 9466–9476

  123. [131]

    Nvds +: Towards efficient and versatile neural stabilizer for video depth estimation,

    Y. Wang, M. Shi, J. Li, C. Hong, Z. Huang, J. Peng, Z. Cao, J. Zhang, K. Xian, and G. Lin, “Nvds +: Towards efficient and versatile neural stabilizer for video depth estimation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , pp. 1–18, 2024

  124. [132]

    Depthcrafter: Generating consistent long depth se- quences for open-world videos,

    W. Hu, X. Gao, X. Li, S. Zhao, X. Cun, Y. Zhang, L. Quan, and Y. Shan, “Depthcrafter: Generating consistent long depth se- quences for open-world videos,” arXiv preprint arXiv:2409.02095, 2024

  125. [133]

    Learning temporally consistent video depth from video diffusion priors,

    J. Shao, Y. Yang, H. Zhou, Y. Zhang, Y. Shen, M. Poggi, and Y. Liao, “Learning temporally consistent video depth from video diffusion priors,” arXiv preprint arXiv:2406.01493, 2024

  126. [134]

    Depth any video with scalable synthetic data,

    H. Yang, D. Huang, W. Yin, C. Shen, H. Liu, X. He, B. Lin, W. Ouyang, and T. He, “Depth any video with scalable synthetic data,” arXiv preprint arXiv:2410.10815, 2024

  127. [135]

    Video depth anything: Consistent depth estimation for super- long videos,

    S. Chen, H. Guo, S. Zhu, F. Zhang, Z. Huang, J. Feng, and B. Kang, “Video depth anything: Consistent depth estimation for super- long videos,” arXiv preprint arXiv:2501.12375, 2025

  128. [136]

    Depth anything v2,

    L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” arXiv preprint arXiv:2406.09414, 2024

  129. [137]

    Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,

    R. Mur-Artal and J. D. Tard ´os, “Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,” IEEE transactions on robotics, vol. 33, no. 5, pp. 1255–1262, 2017

  130. [138]

    Line flow based simultaneous localization and mapping,

    Q. Wang, Z. Yan, J. Wang, F. Xue, W. Ma, and H. Zha, “Line flow based simultaneous localization and mapping,” IEEE Transactions on Robotics, vol. 37, no. 5, pp. 1416–1432, 2021

  131. [139]

    Stereo visual odometry with deep learning-based point and line feature matching using an attention graph neural network,

    S. Kannapiran, N. Bendapudi, M.-Y. Yu, D. Parikh, S. Berman, A. Vora, and G. Pandey, “Stereo visual odometry with deep learning-based point and line feature matching using an attention graph neural network,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and S...

  132. [140]

    Structure plp-slam: Efficient sparse mapping and localization using point, line and plane for monocular, rgb-d and stereo cameras,

    F. Shu, J. Wang, A. Pagani, and D. Stricker, “Structure plp-slam: Efficient sparse mapping and localization using point, line and plane for monocular, rgb-d and stereo cameras,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 2105–2112

  133. [141]

    Ul-slam: A universal monocular line-based slam via unifying structural and non- structural constraints,

    H. Jiang, R. Qian, L. Du, J. Pu, and J. Feng, “Ul-slam: A universal monocular line-based slam via unifying structural and non- structural constraints,” IEEE Transactions on Automation Science and Engineering, 2024

  134. [142]

    Lsd-slam: Large-scale direct monocular slam,

    J. Engel, T. Sch ¨ops, and D. Cremers, “Lsd-slam: Large-scale direct monocular slam,” in European conference on computer vision. Springer, 2014, pp. 834–849

  135. [143]

    Direct sparse odome- try,

    J. Engel, V . Koltun, and D. Cremers, “Direct sparse odome- try,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 3, pp. 611–625, 2017

  136. [144]

    Edplvo: Efficient direct point-line visual odometry,

    L. Zhou, G. Huang, Y. Mao, S. Wang, and M. Kaess, “Edplvo: Efficient direct point-line visual odometry,” in 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2022, pp. 7559–7565

  137. [145]

    D3vo: Deep depth, deep pose and deep uncertainty for monocular visual odometry,

    N. Yang, L. v. Stumberg, R. Wang, and D. Cremers, “D3vo: Deep depth, deep pose and deep uncertainty for monocular visual odometry,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1281–1292

  138. [146]

    Feature warping for robust speaker verification,

    J. Pelecanos and S. Sridharan, “Feature warping for robust speaker verification,” in Proceedings of 2001 A speaker Odyssey: the speaker recognition workshop . European Speech Communication Association, 2001, pp. 213–218

  139. [147]

    Tartanvo: A generalizable learning-based vo,

    W. Wang, Y. Hu, and S. Scherer, “Tartanvo: A generalizable learning-based vo,” in Conference on Robot Learning. PMLR, 2021, pp. 1761–1772

  140. [148]

    Deepvo: Towards end-to-end visual odometry with deep recurrent convolutional neural networks,

    S. Wang, R. Clark, H. Wen, and N. Trigoni, “Deepvo: Towards end-to-end visual odometry with deep recurrent convolutional neural networks,” in 2017 IEEE international conference on robotics and automation (ICRA). IEEE, 2017, pp. 2043–2050

  141. [149]

    Dytanvo: Joint refine- ment of visual odometry and motion segmentation in dynamic environments,

    S. Shen, Y. Cai, W. Wang, and S. Scherer, “Dytanvo: Joint refine- ment of visual odometry and motion segmentation in dynamic environments,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 4048–4055

  142. [150]

    Deep direct visual odometry,

    C. Zhao, Y. Tang, Q. Sun, and A. V . Vasilakos, “Deep direct visual odometry,” IEEE transactions on intelligent transportation systems , vol. 23, no. 7, pp. 7733–7742, 2021

  143. [151]

    Deep patch visual odometry,

    Z. Teed, L. Lipson, and J. Deng, “Deep patch visual odometry,” Advances in Neural Information Processing Systems , vol. 36, pp. 39 033–39 051, 2023

  144. [152]

    Airslam: An efficient and illumination-robust point-line visual slam system,

    K. Xu, Y. Hao, S. Yuan, C. Wang, and L. Xie, “Airslam: An efficient and illumination-robust point-line visual slam system,” IEEE Transactions on Robotics, 2025

  145. [153]

    Deep patch visual slam,

    L. Lipson, Z. Teed, and J. Deng, “Deep patch visual slam,” in European Conference on Computer Vision. Springer, 2024, pp. 424– 440

  146. [154]

    Xvo: General- ized visual odometry via cross-modal self-training,

    L. Lai, Z. Shangguan, J. Zhang, and E. Ohn-Bar, “Xvo: General- ized visual odometry via cross-modal self-training,” in Proceed- ings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 10 094–10 105

  147. [155]

    Anycam: Learning to recover camera poses and intrinsics from casual videos,

    F. Wimbauer, W. Chen, D. Muhle, C. Rupprecht, and D. Cremers, “Anycam: Learning to recover camera poses and intrinsics from casual videos,” in Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2025

  148. [156]

    Dynamic camera poses and where to find them,

    C. Rockwell, J. Tung, T.-Y. Lin, M.-Y. Liu, D. F. Fouhey, and C.-H. Lin, “Dynamic camera poses and where to find them,” in Pro- ceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 12 444–12 455

  149. [157]

    Efficient camera exposure control for visual odometry via deep reinforcement learning,

    S. Zhang, J. He, Y. Zhu, J. Wu, and J. Yuan, “Efficient camera exposure control for visual odometry via deep reinforcement learning,” IEEE Robotics and Automation Letters, 2024

  150. [158]

    Reinforcement learning meets visual odometry,

    N. Messikommer, G. Cioffi, M. Gehrig, and D. Scaramuzza, “Reinforcement learning meets visual odometry,” in European Conference on Computer Vision. Springer, 2024, pp. 76–92

  151. [159]

    Tracking everything everywhere all at once,

    Q. Wang, Y.-Y. Chang, R. Cai, Z. Li, B. Hariharan, A. Holynski, and N. Snavely, “Tracking everything everywhere all at once,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 19 795–19 806

  152. [160]

    Track every- thing everywhere fast and robustly,

    Y. Song, J. Lei, Z. Wang, L. Liu, and K. Daniilidis, “Track every- thing everywhere fast and robustly,” in European Conference on Computer Vision. Springer, 2024, pp. 343–359

  153. [161]

    Spatialtracker: Tracking any 2d pixels in 3d space,

    Y. Xiao, Q. Wang, S. Zhang, N. Xue, S. Peng, Y. Shen, and X. Zhou, “Spatialtracker: Tracking any 2d pixels in 3d space,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 406–20 417

  154. [162]

    Scene- tracker: Long-term scene flow estimation network,

    B. Wang, J. Li, Y. Yu, L. Liu, Z. Sun, and D. Hu, “Scene- tracker: Long-term scene flow estimation network,”arXiv preprint arXiv:2403.19924, 2024

  155. [163]

    Delta: Dense efficient long-range 3d tracking for any video,

    T. D. Ngo, P . Zhuang, C. Gan, E. Kalogerakis, S. Tulyakov, H.-Y. Lee, and C. Wang, “Delta: Dense efficient long-range 3d tracking for any video,” arXiv preprint arXiv:2410.24211, 2024

  156. [164]

    Seurat: From moving points to depth,

    S. Cho, J. Huang, S. Kim, and J.-Y. Lee, “Seurat: From moving points to depth,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 7211–7221

  157. [165]

    Tapip3d: Tracking any point in persistent 3d geometry,

    B. Zhang, L. Ke, A. W. Harley, and K. Fragkiadaki, “Tapip3d: Tracking any point in persistent 3d geometry,” arXiv preprint arXiv:2504.14717, 2025

  158. [166]

    Ego- points: Advancing point tracking for egocentric videos,

    A. Darkhalil, R. Guerrier, A. W. Harley, and D. Damen, “Ego- points: Advancing point tracking for egocentric videos,” in 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2025, pp. 8556–8565

  159. [167]

    Robust consistent video depth estimation,

    J. Kopf, X. Rong, and J.-B. Huang, “Robust consistent video depth estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 1611–1621. 19

  160. [168]

    Structure and motion from casual videos,

    Z. Zhang, F. Cole, Z. Li, M. Rubinstein, N. Snavely, and W. T. Freeman, “Structure and motion from casual videos,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 20–37

  161. [169]

    Megasam: Accurate, fast, and robust structure and motion from casual dynamic videos,

    Z. Li, R. Tucker, F. Cole, Q. Wang, L. Jin, V . Ye, A. Kanazawa, A. Holynski, and N. Snavely, “Megasam: Accurate, fast, and robust structure and motion from casual dynamic videos,” arXiv preprint arXiv:2412.04463, 2024

  162. [170]

    Easi3r: Estimating disentangled motion from dust3r without training,

    X. Chen, Y. Chen, Y. Xiu, A. Geiger, and A. Chen, “Easi3r: Estimating disentangled motion from dust3r without training,” arXiv preprint arXiv:2503.24391, 2025

  163. [171]

    Geome- trycrafter: Consistent geometry estimation for open-world videos with diffusion priors,

    T.-X. Xu, X. Gao, W. Hu, X. Li, S.-H. Zhang, and Y. Shan, “Geome- trycrafter: Consistent geometry estimation for open-world videos with diffusion priors,” arXiv preprint arXiv:2504.01016, 2025

  164. [172]

    3d reconstruction with spatial mem- ory,

    H. Wang and L. Agapito, “3d reconstruction with spatial mem- ory,” arXiv preprint arXiv:2408.16061, 2024

  165. [173]

    Continuous 3d perception model with persistent state,

    Q. Wang, Y. Zhang, A. Holynski, A. A. Efros, and A. Kanazawa, “Continuous 3d perception model with persistent state,” arXiv preprint arXiv:2501.12387, 2025

  166. [174]

    Point3r: Streaming 3d re- construction with explicit spatial pointer memory,

    Y. Wu, W. Zheng, J. Zhou, and J. Lu, “Point3r: Streaming 3d re- construction with explicit spatial pointer memory,” arXiv preprint arXiv:2507.02863, 2025

  167. [175]

    Streaming 4d visual geometry transformer,

    D. Zhuo, W. Zheng, J. Guo, Y. Wu, J. Zhou, and J. Lu, “Streaming 4d visual geometry transformer,” arXiv preprint arXiv:2507.11539, 2025

  168. [176]

    π3: Scalable permutation- equivariant visual geometry learning,

    Y. Wang, J. Zhou, H. Zhu, W. Chang, Y. Zhou, Z. Li, J. Chen, J. Pang, C. Shen, and T. He, “ π3: Scalable permutation- equivariant visual geometry learning,” 2025. [Online]. Available: https://arxiv.org/abs/2507.13347

  169. [177]

    Aether: Geometric-aware unified world modeling,

    A. Team, H. Zhu, Y. Wang, J. Zhou, W. Chang, Y. Zhou, Z. Li, J. Chen, C. Shen, J. Pang et al., “Aether: Geometric-aware unified world modeling,” arXiv preprint arXiv:2503.18945, 2025

  170. [178]

    Geo4d: Leveraging video generators for geometric 4d scene reconstruc- tion,

    Z. Jiang, C. Zheng, I. Laina, D. Larlus, and A. Vedaldi, “Geo4d: Leveraging video generators for geometric 4d scene reconstruc- tion,” arXiv preprint arXiv:2504.07961, 2025

  171. [179]

    Unigeo: Taming video diffusion for unified consistent geometry estimation,

    Y.-T. Sun, X. Yu, Z. Huang, Y.-H. Huang, Y.-C. Guo, Z. Yang, Y.- P . Cao, and X. Qi, “Unigeo: Taming video diffusion for unified consistent geometry estimation,” arXiv preprint arXiv:2505.24521, 2025

  172. [180]

    Uni4d: Unifying visual foundation models for 4d modeling from a single video,

    D. Y. Yao, A. J. Zhai, and S. Wang, “Uni4d: Unifying visual foundation models for 4d modeling from a single video,” in Pro- ceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 1116–1126

  173. [181]

    Back on track: Bundle ad- justment for dynamic scene reconstruction,

    W. Chen, G. Zhang, F. Wimbauer, R. Wang, N. Araslanov, A. Vedaldi, and D. Cremers, “Back on track: Bundle ad- justment for dynamic scene reconstruction,” arXiv preprint arXiv:2504.14516, 2025

  174. [182]

    Stereo4d: Learning how things move in 3d from internet stereo videos,

    L. Jin, R. Tucker, Z. Li, D. Fouhey, N. Snavely, and A. Holynski, “Stereo4d: Learning how things move in 3d from internet stereo videos,” arXiv preprint arXiv:2412.09621, 2024

  175. [183]

    Dynamic point maps: A versatile representation for dynamic 3d reconstruction,

    E. Sucar, Z. Lai, E. Insafutdinov, and A. Vedaldi, “Dynamic point maps: A versatile representation for dynamic 3d reconstruction,” arXiv preprint arXiv:2503.16318, 2025

  176. [184]

    St4rtrack: Simultaneous 4d reconstruction and tracking in the world,

    H. Feng, J. Zhang, Q. Wang, Y. Ye, P . Yu, M. J. Black, T. Darrell, and A. Kanazawa, “St4rtrack: Simultaneous 4d reconstruction and tracking in the world,” arXiv preprint arXiv:2504.13152, 2025

  177. [185]

    Pomato: Marrying pointmap matching with temporal motion for dynamic 3d reconstruction,

    S. Zhang, Y. Ge, J. Tian, G. Xu, H. Chen, C. Lv, and C. Shen, “Pomato: Marrying pointmap matching with temporal motion for dynamic 3d reconstruction,” arXiv preprint arXiv:2504.05692 , 2025

  178. [186]

    Dˆ 2ust3r: Enhancing 3d reconstruction with 4d pointmaps for dynamic scenes,

    J. Han, H. An, J. Jung, T. Narihira, J. Seo, K. Fukuda, C. Kim, S. Hong, Y. Mitsufuji, and S. Kim, “Dˆ 2ust3r: Enhancing 3d reconstruction with 4d pointmaps for dynamic scenes,” arXiv preprint arXiv:2504.06264, 2025

  179. [187]

    Zero- shot monocular scene flow estimation in the wild,

    Y. Liang, A. Badki, H. Su, J. Tompkin, and O. Gallo, “Zero- shot monocular scene flow estimation in the wild,” arXiv preprint arXiv:2501.10357, 2025

  180. [188]

    Fast encoder-based 3d from casual videos via point track processing,

    Y. Kasten, W. Lu, and H. Maron, “Fast encoder-based 3d from casual videos via point track processing,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems

  181. [189]

    Transformers in vision: A survey,

    S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM computing surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022

  182. [190]

    Spatialtrackerv2: 3d point tracking made easy,

    Y. Xiao, J. Wang, N. Xue, N. Karaev, Y. Makarov, B. Kang, X. Zhu, H. Bao, Y. Shen, and X. Zhou, “Spatialtrackerv2: 3d point tracking made easy,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2025. [Online]. Available: https://arxiv.org/abs/2507.12462

  183. [191]

    Undeepvo: Monocular visual odometry through unsupervised deep learning,

    R. Li, S. Wang, Z. Long, and D. Gu, “Undeepvo: Monocular visual odometry through unsupervised deep learning,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 7286–7291

  184. [192]

    Deepv2d: Video to depth with differen- tiable structure from motion,

    Z. Teed and J. Deng, “Deepv2d: Video to depth with differen- tiable structure from motion,” arXiv preprint arXiv:1812.04605 , 2018

  185. [193]

    Generalizing to the open world: Deep visual odometry with online adaptation,

    S. Li, X. Wu, Y. Cao, and H. Zha, “Generalizing to the open world: Deep visual odometry with online adaptation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2021, pp. 13 184–13 193

  186. [194]

    Improving monocular visual odometry using learned depth,

    L. Sun, W. Yin, E. Xie, Z. Li, C. Sun, and C. Shen, “Improving monocular visual odometry using learned depth,” IEEE Transac- tions on Robotics, vol. 38, no. 5, pp. 3173–3186, 2022

  187. [195]

    Towards better generaliza- tion: Joint depth-pose learning without posenet,

    W. Zhao, S. Liu, Y. Shu, and Y.-J. Liu, “Towards better generaliza- tion: Joint depth-pose learning without posenet,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2020, pp. 9151–9161

  188. [196]

    Deep unsupervised visual odometry via bundle adjusted pose graph optimization,

    G. Lu, “Deep unsupervised visual odometry via bundle adjusted pose graph optimization,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 6131–6137

  189. [197]

    Droid-slam: Deep visual slam for monoc- ular, stereo, and rgb-d cameras,

    Z. Teed and J. Deng, “Droid-slam: Deep visual slam for monoc- ular, stereo, and rgb-d cameras,” Advances in neural information processing systems, vol. 34, pp. 16 558–16 569, 2021

  190. [198]

    Nicer-slam: Neural implicit scene encoding for rgb slam,

    Z. Zhu, S. Peng, V . Larsson, Z. Cui, M. R. Oswald, A. Geiger, and M. Pollefeys, “Nicer-slam: Neural implicit scene encoding for rgb slam,” in 2024 International Conference on 3D Vision (3DV). IEEE, 2024, pp. 42–52

  191. [199]

    Dense rgb slam with neural implicit maps,

    H. Li, X. Gu, W. Yuan, L. Yang, Z. Dong, and P . Tan, “Dense rgb slam with neural implicit maps,” arXiv preprint arXiv:2301.08930, 2023

  192. [200]

    Go-slam: Global optimization for consistent 3d instant reconstruction,

    Y. Zhang, F. Tosi, S. Mattoccia, and M. Poggi, “Go-slam: Global optimization for consistent 3d instant reconstruction,” in Proceed- ings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3727–3737

  193. [201]

    Glorie-slam: Globally optimized rgb-only im- plicit encoding point cloud slam,

    G. Zhang, E. Sandstr ¨om, Y. Zhang, M. Patel, L. Van Gool, and M. R. Oswald, “Glorie-slam: Globally optimized rgb-only im- plicit encoding point cloud slam,”arXiv preprint arXiv:2403.19549, 2024

  194. [202]

    Splat-slam: Globally op- timized rgb-only slam with 3d gaussians,

    E. Sandstr ¨om, K. Tateno, M. Oechsle, M. Niemeyer, L. Van Gool, M. R. Oswald, and F. Tombari, “Splat-slam: Globally op- timized rgb-only slam with 3d gaussians,” arXiv preprint arXiv:2405.16544, 2024

  195. [203]

    Gaussian splatting slam,

    H. Matsuki, R. Murai, P . H. Kelly, and A. J. Davison, “Gaussian splatting slam,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 18 039–18 048

  196. [204]

    How nerfs and 3d gaussian splat- ting are reshaping slam: A survey. arxiv 2024,

    F. Tosi, Y. Zhang, Z. Gong, E. Sandstr ¨om, S. Mattoccia, M. Os- wald, and M. Poggi, “How nerfs and 3d gaussian splat- ting are reshaping slam: A survey. arxiv 2024,” arXiv preprint arXiv:2402.13255, 2024

  197. [205]

    Flowmap: High-quality camera poses, intrinsics, and depth via gradient descent,

    C. Smith, D. Charatan, A. Tewari, and V . Sitzmann, “Flowmap: High-quality camera poses, intrinsics, and depth via gradient descent,” arXiv preprint arXiv:2404.15259, 2024

  198. [206]

    Grounding image matching in 3d with mast3r,

    V . Leroy, Y. Cabon, and J. Revaud, “Grounding image matching in 3d with mast3r,” in European Conference on Computer Vision . Springer, 2024, pp. 71–91

  199. [207]

    Mast3r-sfm: a fully-integrated solution for uncon- strained structure-from-motion,

    B. Duisterhof, L. Zust, P . Weinzaepfel, V . Leroy, Y. Cabon, and J. Revaud, “Mast3r-sfm: a fully-integrated solution for uncon- strained structure-from-motion,” arXiv preprint arXiv:2409.19152, 2024

  200. [208]

    Light3r-sfm: Towards feed-forward structure-from-motion,

    S. Elflein, Q. Zhou, S. Agostinho, and L. Leal-Taix ´e, “Light3r-sfm: Towards feed-forward structure-from-motion,” arXiv preprint arXiv:2501.14914, 2025

  201. [209]

    Must3r: Multi-view network for stereo 3d reconstruction,

    Y. Cabon, L. Stoffl, L. Antsfeld, G. Csurka, B. Chidlovskii, J. Re- vaud, and V . Leroy, “Must3r: Multi-view network for stereo 3d reconstruction,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 1050–1060

  202. [210]

    Pow3r: Empowering unconstrained 3d reconstruction with cam- era and scene priors,

    W. Jang, P . Weinzaepfel, V . Leroy, L. Agapito, and J. Revaud, “Pow3r: Empowering unconstrained 3d reconstruction with cam- era and scene priors,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 1071–1081

  203. [211]

    Regist3r: Incremen- tal registration with stereo foundation model,

    S. Liu, W. Li, P . Qiao, and Y. Dou, “Regist3r: Incremen- tal registration with stereo foundation model,” arXiv preprint arXiv:2504.12356, 2025

  204. [212]

    Surfels: Surface elements as rendering primitives,

    H. Pfister, M. Zwicker, J. Van Baar, and M. Gross, “Surfels: Surface elements as rendering primitives,” inProceedings of the 27th annual 20 conference on Computer graphics and interactive techniques, 2000, pp. 335–342

  205. [213]

    Differentiable surface splatting for point-based geometry pro- cessing,

    W. Yifan, F. Serena, S. Wu, C. ¨Oztireli, and O. Sorkine-Hornung, “Differentiable surface splatting for point-based geometry pro- cessing,” ACM Transactions on Graphics (TOG) , vol. 38, no. 6, pp. 1–14, 2019

  206. [214]

    Neural point cloud rendering via multi-plane projection,

    P . Dai, Y. Zhang, Z. Li, S. Liu, and B. Zeng, “Neural point cloud rendering via multi-plane projection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 7830–7839

  207. [215]

    Synsin: End- to-end view synthesis from a single image,

    O. Wiles, G. Gkioxari, R. Szeliski, and J. Johnson, “Synsin: End- to-end view synthesis from a single image,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 7467–7477

  208. [216]

    Pulsar: Efficient sphere-based neural rendering,

    C. Lassner and M. Zollhofer, “Pulsar: Efficient sphere-based neural rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 1440–1449

  209. [217]

    Point- based neural rendering with per-view optimization,

    G. Kopanas, J. Philip, T. Leimk ¨uhler, and G. Drettakis, “Point- based neural rendering with per-view optimization,” inComputer Graphics Forum, vol. 40, no. 4. Wiley Online Library, 2021, pp. 29–43

  210. [218]

    Adop: Approximate differentiable one-pixel point rendering,

    D. R ¨uckert, L. Franke, and M. Stamminger, “Adop: Approximate differentiable one-pixel point rendering,” ACM Transactions on Graphics (ToG), vol. 41, no. 4, pp. 1–14, 2022

  211. [219]

    Free view synthesis,

    G. Riegler and V . Koltun, “Free view synthesis,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIX 16. Springer, 2020, pp. 623–640

  212. [220]

    Stable view synthesis,

    ——, “Stable view synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 12 216–12 225

  213. [221]

    Fwd: Real-time novel view synthesis with forward warping and depth,

    A. Cao, C. Rockwell, and J. Johnson, “Fwd: Real-time novel view synthesis with forward warping and depth,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 713–15 724

  214. [222]

    Botsch, L

    M. Botsch, L. Kobbelt, M. Pauly, P . Alliez, and B. L ´evy, Polygon mesh processing. CRC press, 2010

  215. [223]

    Local surface interpolation with b´ezier patches,

    L. A. Shirman and C. H. Sequin, “Local surface interpolation with b´ezier patches,” Computer Aided Geometric Design, vol. 4, no. 4, pp. 279–295, 1987

  216. [224]

    Dynamic surface func- tion networks for clothed human bodies,

    A. Burov, M. Nießner, and J. Thies, “Dynamic surface func- tion networks for clothed human bodies,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 754–10 764

  217. [225]

    Deferred neural render- ing: Image synthesis using neural textures,

    J. Thies, M. Zollh ¨ofer, and M. Nießner, “Deferred neural render- ing: Image synthesis using neural textures,” Acm Transactions on Graphics (TOG), vol. 38, no. 4, pp. 1–12, 2019

  218. [226]

    Opendr: An approximate differ- entiable renderer,

    M. M. Loper and M. J. Black, “Opendr: An approximate differ- entiable renderer,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13. Springer, 2014, pp. 154–169

  219. [227]

    Neural 3d mesh renderer,

    H. Kato, Y. Ushiku, and T. Harada, “Neural 3d mesh renderer,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3907–3916

  220. [228]

    Soft rasterizer: A differentiable renderer for image-based 3d reasoning,

    S. Liu, T. Li, W. Chen, and H. Li, “Soft rasterizer: A differentiable renderer for image-based 3d reasoning,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 7708–7717

  221. [229]

    Paparazzi: surface editing by way of multi-view image processing

    H.-T. D. Liu, M. Tao, and A. Jacobson, “Paparazzi: surface editing by way of multi-view image processing.” ACM Trans. Graph. , vol. 37, no. 6, pp. 221–1, 2018

  222. [230]

    Mitsuba 2: A retargetable forward and inverse renderer,

    M. Nimier-David, D. Vicini, T. Zeltner, and W. Jakob, “Mitsuba 2: A retargetable forward and inverse renderer,” ACM Transactions on Graphics (TOG), vol. 38, no. 6, pp. 1–17, 2019

  223. [231]

    Taichi: a language for high-performance computation on spa- tially sparse data structures,

    Y. Hu, T.-M. Li, L. Anderson, J. Ragan-Kelley, and F. Durand, “Taichi: a language for high-performance computation on spa- tially sparse data structures,” ACM Transactions on Graphics (TOG), vol. 38, no. 6, p. 201, 2019

  224. [232]

    Optical models for direct volume rendering,

    N. Max, “Optical models for direct volume rendering,” IEEE Transactions on Visualization and Computer Graphics , vol. 1, no. 2, pp. 99–108, 1995

  225. [233]

    Nerf in the wild: Neural radiance fields for unconstrained photo collections,

    R. Martin-Brualla, N. Radwan, M. S. Sajjadi, J. T. Barron, A. Doso- vitskiy, and D. Duckworth, “Nerf in the wild: Neural radiance fields for unconstrained photo collections,” in CVPR, 2021, pp. 7210–7219

  226. [234]

    Physg: Inverse rendering with spherical gaussians for physics-based material editing and relighting,

    K. Zhang, F. Luan, Q. Wang, K. Bala, and N. Snavely, “Physg: Inverse rendering with spherical gaussians for physics-based material editing and relighting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 5453–5462

  227. [235]

    Barf: Bundle- adjusting neural radiance fields,

    C.-H. Lin, W.-C. Ma, A. Torralba, and S. Lucey, “Barf: Bundle- adjusting neural radiance fields,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 5741–5751

  228. [236]

    Humannerf: Free-viewpoint ren- dering of moving people from monocular video,

    C.-Y. Weng, B. Curless, P . P . Srinivasan, J. T. Barron, and I. Kemelmacher-Shlizerman, “Humannerf: Free-viewpoint ren- dering of moving people from monocular video,” inProceedings of the IEEE/CVF conference on computer vision and pattern Recognition, 2022, pp. 16 210–16 220

  229. [237]

    Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,

    P . Wang, L. Liu, Y. Liu, C. Theobalt, T. Komura, and W. Wang, “Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,” Advances in Neural Information Processing Systems, vol. 34, pp. 27 171–27 183, 2021

  230. [238]

    Volume rendering of neural implicit surfaces,

    L. Yariv, J. Gu, Y. Kasten, and Y. Lipman, “Volume rendering of neural implicit surfaces,” Advances in Neural Information Process- ing Systems, vol. 34, pp. 4805–4815, 2021

  231. [239]

    GIRAFFE: Representing scenes as compositional generative neural feature fields,

    M. Niemeyer and A. Geiger, “GIRAFFE: Representing scenes as compositional generative neural feature fields,” 2021

  232. [240]

    Mip-nerf: A multiscale representa- tion for anti-aliasing neural radiance fields,

    J. T. Barron, B. Mildenhall, M. Tancik, P . Hedman, R. Martin- Brualla, and P . P . Srinivasan, “Mip-nerf: A multiscale representa- tion for anti-aliasing neural radiance fields,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 5855–5864

  233. [241]

    Mip-nerf 360: Unbounded anti-aliased neural radiance fields,

    J. T. Barron, B. Mildenhall, D. Verbin, P . P . Srinivasan, and P . Hed- man, “Mip-nerf 360: Unbounded anti-aliased neural radiance fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5470–5479

  234. [242]

    Ref-nerf: Structured view-dependent appear- ance for neural radiance fields,

    D. Verbin, P . Hedman, B. Mildenhall, T. Zickler, J. T. Barron, and P . P . Srinivasan, “Ref-nerf: Structured view-dependent appear- ance for neural radiance fields,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022, pp. 5481–5490

  235. [243]

    Dynibar: Neural dynamic image-based rendering,

    Z. Li, Q. Wang, F. Cole, R. Tucker, and N. Snavely, “Dynibar: Neural dynamic image-based rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 4273–4284

  236. [244]

    In-place scene labelling and understanding with implicit scene represen- tation,

    S. Zhi, T. Laidlow, S. Leutenegger, and A. J. Davison, “In-place scene labelling and understanding with implicit scene represen- tation,” in ICCV, 2021, pp. 15 838–15 847

  237. [245]

    Sparf: Neural radiance fields from sparse and noisy poses,

    P . Truong, M.-J. Rakotosaona, F. Manhardt, and F. Tombari, “Sparf: Neural radiance fields from sparse and noisy poses,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 4190–4200

  238. [246]

    Nerd: Neural reflectance decomposition from image collec- tions,

    M. Boss, R. Braun, V . Jampani, J. T. Barron, C. Liu, and H. Lensch, “Nerd: Neural reflectance decomposition from image collec- tions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 12 684–12 694

  239. [247]

    pixelnerf: Neural radiance fields from one or few images,

    A. Yu, V . Ye, M. Tancik, and A. Kanazawa, “pixelnerf: Neural radiance fields from one or few images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 4578–4587

  240. [248]

    IBRNet: Learning multi-view image-based rendering,

    Q. Wang, Z. Wang, K. Genova, P . P . Srinivasan, H. Zhou, J. T. Bar- ron, R. Martin-Brualla, N. Snavely, and T. Funkhouser, “IBRNet: Learning multi-view image-based rendering,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 4690–4699

  241. [249]

    Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps,

    C. Reiser, S. Peng, Y. Liao, and A. Geiger, “Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 335–14 345

  242. [250]

    Neural radiance flow for 4d view synthesis and video processing,

    Y. Du, Y. Zhang, H.-X. Yu, J. B. Tenenbaum, and J. Wu, “Neural radiance flow for 4d view synthesis and video processing,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE Computer Society, 2021, pp. 14 304–14 314

  243. [251]

    Neu- ral 3d video synthesis from multi-view video,

    T. Li, M. Slavcheva, M. Zollhoefer, S. Green, C. Lassner, C. Kim, T. Schmidt, S. Lovegrove, M. Goesele, R. Newcombe et al., “Neu- ral 3d video synthesis from multi-view video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2022, pp. 5521–5531

  244. [252]

    Animatable neural radiance fields for modeling dy- namic human bodies,

    S. Peng, J. Dong, Q. Wang, S. Zhang, Q. Shuai, X. Zhou, and H. Bao, “Animatable neural radiance fields for modeling dy- namic human bodies,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 314–14 323

  245. [253]

    Nerf in the palm of your hand: Corrective augmentation for robotics via novel-view synthesis,

    A. Zhou, M. J. Kim, L. Wang, P . Florence, and C. Finn, “Nerf in the palm of your hand: Corrective augmentation for robotics via novel-view synthesis,” in Proceedings of the IEEE/CVF Conference 21 on Computer Vision and Pattern Recognition , 2023, pp. 17 907– 17 917

  246. [254]

    Neat: Neural adaptive tomography,

    D. R ¨uckert, Y. Wang, R. Li, R. Idoughi, and W. Heidrich, “Neat: Neural adaptive tomography,” ACM Transactions on Graphics (TOG), vol. 41, no. 4, pp. 1–13, 2022

  247. [255]

    Gravitationally lensed black hole emission tomography,

    A. Levis, P . P . Srinivasan, A. A. Chael, R. Ng, and K. L. Bouman, “Gravitationally lensed black hole emission tomography,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 19 841–19 850

  248. [256]

    Mip-NeRF 360: Unbounded anti-aliased neural radiance fields,

    J. T. Barron, B. Mildenhall, D. Verbin, P . P . Srinivasan, and P . Hed- man, “Mip-NeRF 360: Unbounded anti-aliased neural radiance fields,” CoRR, vol. abs/2111.12077, 2022

  249. [257]

    Screened poisson surface recon- struction,

    M. Kazhdan and H. Hoppe, “Screened poisson surface recon- struction,” ACM Transactions on Graphics (ToG), vol. 32, no. 3, pp. 1–13, 2013

  250. [258]

    Two algorithms for constructing a delaunay triangulation,

    D.-T. Lee and B. J. Schachter, “Two algorithms for constructing a delaunay triangulation,” International Journal of Computer & Information Sciences, vol. 9, no. 3, pp. 219–242, 1980

  251. [259]

    Structure-from-motion re- visited,

    J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-motion re- visited,” in Conference on Computer Vision and Pattern Recognition (CVPR), 2016

  252. [260]

    Mvsnet: Depth inference for unstructured multi-view stereo,

    Y. Yao, Z. Luo, S. Li, T. Fang, and L. Quan, “Mvsnet: Depth inference for unstructured multi-view stereo,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 767– 783

  253. [261]

    Neat: Learning neural implicit surfaces with arbitrary topologies from multi-view images,

    X. Meng, W. Chen, and B. Yang, “Neat: Learning neural implicit surfaces with arbitrary topologies from multi-view images,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 248–258

  254. [262]

    Neuralangelo: High-fidelity neural surface recon- struction,

    Z. Li, T. M ¨uller, A. Evans, R. H. Taylor, M. Unberath, M.-Y. Liu, and C.-H. Lin, “Neuralangelo: High-fidelity neural surface recon- struction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 8456–8465

  255. [263]

    Marching cubes: A high res- olution 3d surface construction algorithm,

    W. E. Lorensen and H. E. Cline, “Marching cubes: A high res- olution 3d surface construction algorithm,” in Seminal graphics: pioneering efforts that shaped the field , 1998, pp. 347–353

  256. [264]

    2d gaussian splatting for geometrically accurate radiance fields,

    B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao, “2d gaussian splatting for geometrically accurate radiance fields,” in ACM SIGGRAPH 2024 conference papers, 2024, pp. 1–11

  257. [265]

    Gaussian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes,

    Z. Yu, T. Sattler, and A. Geiger, “Gaussian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes,” ACM Transactions on Graphics (TOG), vol. 43, no. 6, pp. 1–13, 2024

  258. [266]

    Pgsr: Planar-based gaussian splatting for efficient and high-fidelity surface reconstruction,

    D. Chen, H. Li, W. Ye, Y. Wang, W. Xie, S. Zhai, N. Wang, H. Liu, H. Bao, and G. Zhang, “Pgsr: Planar-based gaussian splatting for efficient and high-fidelity surface reconstruction,” IEEE Transactions on Visualization and Computer Graphics, 2024

  259. [267]

    Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering,

    A. Gu ´edon and V . Lepetit, “Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5354–5363

  260. [268]

    3d gaus- sian splatting for fine-detailed surface reconstruction in large- scale scene,

    S. Chen, Z. Li, Z. Chen, Q. Yan, G. Shen, and R. Duan, “3d gaus- sian splatting for fine-detailed surface reconstruction in large- scale scene,” arXiv preprint arXiv:2506.17636, 2025

  261. [269]

    Multiview geometric regularization of gaussian splatting for accurate radiance fields,

    J. Kim, G. Park, and S. Lee, “Multiview geometric regularization of gaussian splatting for accurate radiance fields,” arXiv preprint arXiv:2506.13508, 2025

  262. [270]

    Tri 2 plane: Advancing neural im- plicit surface reconstruction for indoor scenes,

    Y. Xie, H. Xiao, and W. Kang, “Tri 2 plane: Advancing neural im- plicit surface reconstruction for indoor scenes,” IEEE Transactions on Multimedia, 2025

  263. [271]

    Esa-gs: Elongation splitting and assimilation in gaussian splatting for accurate sur- face reconstruction,

    Y. Chen, W. Wu, Y. Peng, Y. Fei, and L. Zheng, “Esa-gs: Elongation splitting and assimilation in gaussian splatting for accurate sur- face reconstruction,” Computer Aided Geometric Design , p. 102434, 2025

  264. [272]

    Gaussianudf: Inferring unsigned distance functions through 3d gaussian splatting,

    S. Li, Y.-S. Liu, and Z. Han, “Gaussianudf: Inferring unsigned distance functions through 3d gaussian splatting,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 27 113–27 123

  265. [273]

    Sof: Sorted opacity fields for fast unbounded surface reconstruction,

    L. Radl, F. Windisch, T. Deixelberger, J. Hladky, M. Steiner, D. Schmalstieg, and M. Steinberger, “Sof: Sorted opacity fields for fast unbounded surface reconstruction,” arXiv preprint arXiv:2506.19139, 2025

  266. [274]

    Multi-view surface reconstruction using normal and reflectance cues,

    R. Bruneau, B. Brument, Y. Qu ´eau, J. M ´elou, F. B. Lauze, J.-D. Durou, and L. Calvet, “Multi-view surface reconstruction using normal and reflectance cues,” arXiv preprint arXiv:2506.04115 , 2025

  267. [275]

    Quicksplat: Fast 3d surface reconstruction via learned gaussian initialization,

    Y.-C. Liu, L. H ¨ollein, M. Nießner, and A. Dai, “Quicksplat: Fast 3d surface reconstruction via learned gaussian initialization,” arXiv preprint arXiv:2505.05591, 2025

  268. [276]

    Geometry field splatting with gaussian surfels,

    K. Jiang, V . Sivaram, C. Peng, and R. Ramamoorthi, “Geometry field splatting with gaussian surfels,” in Proceedings of the Com- puter Vision and Pattern Recognition Conference , 2025, pp. 5752– 5762

  269. [277]

    Solidgs: Consolidating gaussian surfel splatting for sparse-view surface reconstruction,

    Z. Shen, Y. Liu, Z. Chen, Z. Li, J. Wang, Y. Liang, Z. Yu, J. Zhang, Y. Xu, S. Schaefer et al., “Solidgs: Consolidating gaussian surfel splatting for sparse-view surface reconstruction,” arXiv preprint arXiv:2412.15400, 2024

  270. [278]

    Sparseneus: Fast generalizable neural surface reconstruction from sparse views,

    X. Long, C. Lin, P . Wang, T. Komura, and W. Wang, “Sparseneus: Fast generalizable neural surface reconstruction from sparse views,” in European Conference on Computer Vision . Springer, 2022, pp. 210–227

  271. [279]

    Gens: Generalizable neural surface reconstruction from multi-view im- ages,

    R. Peng, X. Gu, L. Tang, S. Shen, F. Yu, and R. Wang, “Gens: Generalizable neural surface reconstruction from multi-view im- ages,” Advances in Neural Information Processing Systems , vol. 36, pp. 56 932–56 945, 2023

  272. [280]

    C2f2neus: Cascade cost frustum fusion for high fidelity and generalizable neural surface reconstruction,

    L. Xu, T. Guan, Y. Wang, W. Liu, Z. Zeng, J. Wang, and W. Yang, “C2f2neus: Cascade cost frustum fusion for high fidelity and generalizable neural surface reconstruction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 18 291–18 301

  273. [281]

    Uforecon: generalizable sparse-view surface reconstruction from arbitrary and unfavorable sets,

    Y. Na, W. J. Kim, K. B. Han, S. Ha, and S.-E. Yoon, “Uforecon: generalizable sparse-view surface reconstruction from arbitrary and unfavorable sets,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5094–5104

  274. [282]

    Surface-centric modeling for high-fidelity generalizable neural surface reconstruction,

    R. Peng, S. Shen, K. Xiong, H. Gao, J. Jiao, X. Gu, and R. Wang, “Surface-centric modeling for high-fidelity generalizable neural surface reconstruction,” in European Conference on Computer Vi- sion. Springer, 2024, pp. 183–200

  275. [283]

    Retr: Modeling rendering via transformer for generalizable neural surface reconstruction,

    Y. Liang, H. He, and Y. Chen, “Retr: Modeling rendering via transformer for generalizable neural surface reconstruction,” Ad- vances in neural information processing systems , vol. 36, pp. 62 332– 62 351, 2023

  276. [284]

    Lara: Efficient large-baseline radiance fields,

    A. Chen, H. Xu, S. Esposito, S. Tang, and A. Geiger, “Lara: Efficient large-baseline radiance fields,” in European Conference on Computer Vision. Springer, 2024, pp. 338–355

  277. [285]

    Special section on egocentric perception,

    A. Furnari, D. Crandall, D. Damen, K. Grauman, and G. M. Farinella, “Special section on egocentric perception,” IEEE Trans- actions on Pattern Analysis and Machine Intelligence , vol. 45, no. 6, pp. 6602–6604, 2023

  278. [286]

    An outlook into the future of egocentric vision,

    C. Plizzari, G. Goletto, A. Furnari, S. Bansal, F. Ragusa, G. M. Farinella, D. Damen, and T. Tommasi, “An outlook into the future of egocentric vision,” International Journal of Computer Vision, vol. 132, no. 11, pp. 4880–4936, 2024

  279. [287]

    Scene- script: Reconstructing scenes with an autoregressive structured language model,

    A. Avetisyan, C. Xie, H. Howard-Jenkins, T.-Y. Yang, S. Aroudj, S. Patra, F. Zhang, D. Frost, L. Holland, C. Orme et al., “Scene- script: Reconstructing scenes with an autoregressive structured language model,” in European Conference on Computer Vision . Springer, 2024, pp. 247–263

  280. [288]

    Ego- lifter: Open-world 3d segmentation for egocentric perception,

    Q. Gu, Z. Lv, D. Frost, S. Green, J. Straub, and C. Sweeney, “Ego- lifter: Open-world 3d segmentation for egocentric perception,” in European Conference on Computer Vision . Springer, 2024, pp. 382–400

  281. [289]

    Photoreal scene reconstruction from an egocentric device,

    Z. Lv, M. Monge, K. Chen, Y. Zhu, M. Goesele, J. Engel, Z. Dong, and R. Newcombe, “Photoreal scene reconstruction from an egocentric device,” in ACM SIGGRAPH, 2025

  282. [290]

    Spatial cognition from egocentric video: Out of sight, not out of mind,

    C. Plizzari, S. Goel, T. Perrett, J. Chalk, A. Kanazawa, and D. Damen, “Spatial cognition from egocentric video: Out of sight, not out of mind,” arXiv preprint arXiv:2404.05072, 2024

  283. [291]

    Nerf++: Analyzing and improving neural radiance fields,

    K. Zhang, G. Riegler, N. Snavely, and V . Koltun, “Nerf++: Analyzing and improving neural radiance fields,” CoRR, vol. abs/2010.07492, 2020

  284. [292]

    Zip-nerf: Anti-aliased grid-based neural radiance fields,

    J. T. Barron, B. Mildenhall, D. Verbin, P . P . Srinivasan, and P . Hedman, “Zip-nerf: Anti-aliased grid-based neural radiance fields,” CoRR, vol. abs/2304.06706, 2023

  285. [293]

    City- gaussian: Real-time high-quality large-scale scene rendering with gaussians,

    Y. Liu, C. Luo, L. Fan, N. Wang, J. Peng, and Z. Zhang, “City- gaussian: Real-time high-quality large-scale scene rendering with gaussians,” in European Conference on Computer Vision. Springer, 2024, pp. 265–282

  286. [294]

    Citygaussianv2: Efficient and geometrically accurate reconstruction for large-scale scenes,

    Y. Liu, C. Luo, Z. Mao, J. Peng, and Z. Zhang, “Citygaussianv2: Efficient and geometrically accurate reconstruction for large-scale scenes,” arXiv preprint arXiv:2411.00771, 2024

  287. [295]

    Octree- 22 gs: Towards consistent real-time rendering with lod-structured 3d gaussians,

    K. Ren, L. Jiang, T. Lu, M. Yu, L. Xu, Z. Ni, and B. Dai, “Octree- 22 gs: Towards consistent real-time rendering with lod-structured 3d gaussians,” arXiv preprint arXiv:2403.17898, 2024

  288. [296]

    Citygs-x: A scalable architecture for efficient and geometrically accurate large-scale scene reconstruction,

    Y. Gao, H. Li, J. Chen, Z. Zou, Z. Zhong, D. Zhang, X. Sun, and J. Han, “Citygs-x: A scalable architecture for efficient and geometrically accurate large-scale scene reconstruction,” arXiv preprint arXiv:2503.23044, 2025

  289. [297]

    Lodge: Level- of-detail large-scale gaussian splatting with efficient rendering,

    J. Kulhanek, M.-J. Rakotosaona, F. Manhardt, C. Tsalicoglou, M. Niemeyer, T. Sattler, S. Peng, and F. Tombari, “Lodge: Level- of-detail large-scale gaussian splatting with efficient rendering,” arXiv preprint arXiv:2505.23158, 2025

  290. [298]

    Block-nerf: Scal- able large-scene neural view synthesis,

    M. Tancik, V . Casser, X. Yan, S. Pradhan, B. Mildenhall, P . P . Srinivasan, J. T. Barron, and H. Kretzschmar, “Block-nerf: Scal- able large-scene neural view synthesis,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2022

  291. [299]

    Mega-nerf: Scal- able construction of large-scale nerfs for virtual fly-throughs,

    H. Turki, D. Ramanan, and M. Satyanarayanan, “Mega-nerf: Scal- able construction of large-scale nerfs for virtual fly-throughs,” CoRR, vol. abs/2112.10703, 2022

  292. [300]

    Bungeenerf (city-nerf): Progressive neural radiance field for extreme multi-scale scene rendering,

    Y. Xiangli, L. Xu, X. Pan, N. Zhao, A. Rao, C. Theobalt, B. Dai, and D. Lin, “Bungeenerf (city-nerf): Progressive neural radiance field for extreme multi-scale scene rendering,” in European Conf. Computer Vision (ECCV), 2022, pp. 106–122

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.