Pith. sign in

REVIEW 3 major objections 4 minor 103 references

Globally Consistent RGB-D SLAM with 2D Gaussian Splatting

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Replacing 3D Gaussian ellipsoids with 2D Gaussian disks in an RGB-D SLAM map yields depth-consistent rendering, accurate on-manifold pose tracking, and online globally consistent map correction after loop closure.

desk verdict A solid 2DGS-SLAM system with a genuinely new on-manifold pose-Jacobian derivation and strong tracking results; the online global-consistency claim is real but slightly oversold by metrics that mostly come after offline refinement. read the letter →

arxiv 2506.00970 v1 pith:66A5BGZU submitted 2025-06-01 cs.RO

classification cs.RO
keywords RGB-DSLAM2DGaussiansplattingloopclosurecameraposeoptimizationglobalconsistencyradiancefieldmappingMASt3Rsurfacereconstruction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

2DGS-SLAM proposes that a scene map built from flat 2D Gaussian disks, rather than the 3D ellipsoids used by standard Gaussian splatting, removes a core inconsistency in radiance-field SLAM: because each disk's depth is computed by explicit ray-disk intersection, rendered depth images agree across viewpoints, which makes frame-to-map pose tracking with depth reliable. On this representation the paper builds a complete RGB-D SLAM system that tracks the camera purely by rendering and closes loops with a 3D foundation model that both finds revisited places and supplies a coarse relative pose. After pose graph optimization, the whole map is deformed by applying each keyframe's pose increment to the splats registered to that keyframe, and an active/inactive map state keeps tracking local and efficient while allowing old regions to be reactivated on revisits. The paper claims that this design beats rendering-based baselines in trajectory accuracy, matches or exceeds 3DGS methods in reconstruction quality with more globally consistent real-world maps, and runs faster with a more compact map than other radiance-field SLAM systems that support loop closure.

What carries the argument

The central object is the 2D Gaussian splat: a flat disk parameterized by a center, two tangent vectors with scales, a normal, color, and opacity, rendered by explicit ray-disk intersection rather than by projecting a 3D ellipsoid, which makes depth rendering multi-view consistent. The argument is carried by three mechanisms. First, analytic $\mathrm{SE}(3)$ Jacobians are derived for the 2DGS rendering process, propagating the rendering loss through the ray-splat homography to the splat's center and rotation in camera space, so pose tracking is an on-manifold gradient optimization that stays on the pose group. Second, loop closure is handled by MASt3R, which both retrieves candidate keyframes through aggregated image features and predicts dense point maps from which a coarse relative pose is extracted; that pose is scale-corrected with real depth and refined by tracking against the active map before being added as a pose graph constraint. Third, each splat records the ID of its closest observing keyframe, so after pose graph optimization the whole map is deformed by applying that keyframe's pose increment to the splat's position and orientation, with an active/inactive map state keeping the tracked subset local and reactivating splats when areas are revisited.

What would settle it

Render the online pre-refinement map after a loop closure in a scene where the same surface is observed on both sides of a long trajectory, and measure the depth misalignment of that overlapping region against a fresh depth capture; if surfaces observed by multiple keyframes retain errors comparable to the spread of their keyframe pose corrections while surfaces seen by a single keyframe align, the single-association update is insufficient. A cheaper check is to instrument the pipeline and compare per-splat residual misalignment against the number of observing keyframes and the variance of their pose increments.

Watch

Extended reading notes

Core claim

The central claim is that 2D Gaussian splatting is a better backbone than 3D Gaussian splatting for RGB-D SLAM because it supplies multi-view consistent depth rendering, and that this property can be exploited for both tracking and globally consistent mapping. The paper derives analytic $\mathrm{SE}(3)$ Jacobians for the 2DGS rasterizer, enabling rendering-based camera pose optimization directly on the manifold rather than through automatic differentiation of pose matrix elements. For global consistency, every splat stores the ID of the keyframe that observed it at closest range; after a detected loop, pose graph optimization updates keyframe poses and each splat is rigidly transformed by its associated keyframe's pose increment, while an active/inactive map state prevents stale regions from corrupting tracking and allows reactivation upon revisits. Experiments on the Replica, TUM-RGBD, and ScanNet datasets plus self-recorded robot data report sub-millimeter average trajectory error on Replica, reconstruction F1 scores matching or exceeding 3DGS baselines, higher-fidelity rendering in the no-refinement comparison, and roughly 6-7x faster runtime with a map far smaller than that of the loop-closure baselines.

Load-bearing premise

The global consistency claim rests on the assumption that updating each Gaussian splat by the pose correction of its single closest observing keyframe is enough to keep the map aligned, even when a splat is seen from many keyframes that drift differently; if that assumption fails, the map is fully consistent only after the offline 26,000-iteration refinement stage, not during online operation.

Editorial extensions

If this is right

  • Rendering-based RGB-D SLAM can dispense with external odometry: the same 2DGS map supports tracking, mapping, and loop closure in one representation.
  • Global map consistency is achievable online without submap management, because map correction is a direct transformation of splats keyed to keyframe pose increments.
  • Depth-consistent rendering carries over to downstream geometry: meshes extracted from the map are smoother and more globally consistent than those from 3DGS methods, particularly in real-world scenes with noisy depth.
  • Loop closure becomes computationally affordable for online robotics, with a 6-7x runtime speedup and a much smaller stored map than previous loop-closure-enabled radiance-field SLAM systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The single-association map correction suggests a testable extension: splats observed by many keyframes with divergent drift corrections could be updated by a weighted or averaged transform across their observing keyframes, potentially reducing residual misalignment before the offline refinement stage.
  • Because MASt3R supplies dense point maps, the same loop-closure pipeline could be adapted to monocular or RGB-only input by replacing the depth-based scale correction with predicted depth or scale-from-motion, making the design less dependent on an RGB-D sensor.
  • The active/inactive map state is a general strategy that could be transferred to other elastic map representations, such as surfels or neural points, to keep online tracking local while preserving global consistency after loop closure.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents 2DGS-SLAM, a coupled RGB-D SLAM system that uses 2D Gaussian splatting as its sole map representation. It derives analytic SE(3) pose Jacobians for the 2DGS rasterizer, introduces an active/inactive Gaussian map state, and adds MASt3R-based loop closure detection and relocalization followed by pose-graph optimization and elastic deformation of Gaussian splats. The system is evaluated on Replica, TUM-RGBD, ScanNet, and self-recorded robot sequences, reporting very low ATE on Replica (0.07 cm average), competitive tracking on TUM, a compact map (9.7 MB on ScanNet 0000), and strong rendering metrics. Reconstruction metrics in Table V are reported after an offline 26,000-iteration refinement stage described in Sec. IV-A.

Significance. If the claims are fully supported, the paper would make a useful contribution: it is the first 2DGS SLAM system with an on-manifold pose-gradient derivation, it avoids submap management by deforming individual Gaussian splats, and it provides open-source code. The analytic Jacobian derivation in Sec. III-B is a genuine technical step beyond prior 3DGS-based formulations, and the raw-map rendering comparison in Table VIII is valuable because it isolates the online map quality from post-processing. The main weakness is that the headline claim of online global consistency is not directly measured: the map correction in Sec. III-H uses a per-splat single-keyframe association, and the quantitative reconstruction results are obtained after offline refinement, so the evidence does not yet establish that the map is globally consistent at the moment a loop is closed.

major comments (3)
  1. [Sec. III-H, Eq. (32); Sec. IV-A] The online global-consistency claim is not supported by the reported evidence. Eq. (32) updates each Gaussian splat only by the pose increment of its single closest observing keyframe f_k^c. After pose-graph optimization, the other keyframes that observe the same splat generally receive different increments, so the deformed splat cannot satisfy all of them; splats on the same surface but associated with different keyframes will be displaced relative to each other. The paper reports no quantitative measure of this residual misalignment in overlap regions. Fig. 6 and Table VIII are qualitative or rendering-based and do not measure geometric consistency of the corrected map before refinement. This matters because the quantitative reconstruction results in Table V are obtained after the 26,000-iteration offline refinement described in Sec. IV-A, which can repair exactly this kind of inconsistency. Please report a direct metric of corrected-map consistency, e.g., the residual distance between corresponding splats in overlap regions or the ATE of the raw map immediately after pose-graph optimization, both before and after the offline refinement stage.
  2. [Sec. IV-A2 and Table I] The experimental evaluation lacks ablations and error bars, and two tracking hyperparameters are tuned per dataset. The text states that all settings are kept consistent across experiments, but then states that the depth loss weight lambda_d and the tracking success threshold epsilon_t are adjusted individually for each dataset. Because lambda_d directly controls the balance between color and depth in the tracking loss in Eq. (19), and epsilon_t gates loop closures in Sec. III-G, the reported gains could depend on this tuning. Please provide a sensitivity study for these two parameters and report multiple runs or uncertainty bounds for the main tables. An ablation isolating the contributions of the analytic SE(3) Jacobians, the normal/opacity masks in Eqs. (17)-(18), the active/inactive map update, and the loop closure module would also help support the central claims.
  3. [Sec. IV-C, Table V; Sec. IV-B, Table IV] Several claims in the abstract and conclusion are stated more strongly than the tables support. In Table V, 2DGS-SLAM has an average F1 of 89.7%, lower than LoopSplat's 90.4%, so the statement that the method outperforms 3DGS-based approaches in surface reconstruction is only true on a per-metric basis and not consistently. On ScanNet in Table IV, 2DGS-SLAM is second to GO-SLAM in average ATE, and in Table VII it is second to LoopSplat in PSNR and SSIM. The introduction's hedged claim of being 'on par' is acceptable, but the abstract and conclusion should be calibrated to these numbers, and the claim of 'more consistent global map reconstruction' should be backed by a quantitative consistency metric rather than only by the qualitative comparison in Fig. 7.
minor comments (4)
  1. [Sec. III-H, Eq. (32)] The notation is inconsistent: the text defines the updated center as mu'_k, but Eq. (32) writes x'_k. The symbol x is not introduced for the splat center in this section.
  2. [Sec. III-G] The scaled relative pose is introduced as T^r_lc in Eq. (29) but later written as T^t_lc after the scan-to-model refinement; please use distinct symbols consistently, since the two poses are different quantities.
  3. [Sec. IV-B] The text contains typos such as 'baselinws', 'theis', and 'caculate'; these should be corrected before publication.
  4. [Sec. IV and conclusion] The paper does not include a limitations section. Given that MASt3R-based relocalization is a central component, a short discussion of known failure cases, such as highly repetitive or textureless environments, would improve the manuscript.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation chain found; pose gradients and benchmark comparisons are externally grounded, and the disclosed refinement stage is an evidence gap, not a circular step.

full rationale

The paper's central derivations are self-contained against external inputs. The pose-optimization gradients in Sec. III-B (Eqs. 9-13) are obtained analytically from the 2DGS rendering process, with no fitted constant or target quantity entering the chain rule. Tracking and mapping losses (Eqs. 15, 19, 26) compare rendered outputs to sensor observations, and all headline results are measured on public benchmarks (Replica, TUM-RGBD, ScanNet) with standard metrics (ATE, Depth L1, F1, PSNR, SSIM, LPIPS); no benchmark outcome is fitted and then renamed as a prediction. Loop closure relies on the external MASt3R model for an initial relative pose, with the scale estimated from observed depth (Eq. 29) and refined by active-map tracking; this is a per-frame estimate, not a global fit. Self-citations such as PINGS, PIN-SLAM, and ActiveGS are used for context or inspiration, not as load-bearing evidence; no uniqueness theorem or unverified prior result is invoked to force the paper's choices. The map-correction step in Eqs. 31-32 applies only the closest keyframe's pose increment to each splat, which may leave residual inconsistency for splats observed by multiple keyframes with different increments; this is a real limitation for the online global-consistency claim, especially since some reconstruction tables use a disclosed 26,000-iteration offline refinement. However, this is an evidential or robustness gap, not circularity: Eq. 32 does not make any reported quantity equal to an input by construction, and Table VIII provides raw-map results without refinement. Overall, the paper's predictions are not forced by definition or by self-citation.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The system introduces no new physical entities. It relies on domain assumptions about 2DGS depth consistency and MASt3R's reliability, plus two per-dataset hyperparameter adjustments (λd and εt). All other hyperparameters in Table I are fixed across experiments. The core derivation is analytic and not fitted.

free parameters (2)
  • λd (tracking depth loss weight) = per-dataset
    Adjusted individually for each dataset (Sec. IV.A.2); value not stated in the paper, so the reader cannot reproduce without the code.
  • εt (tracking success threshold) = per-dataset
    Adjusted per dataset (Sec. IV.A.2); similarly not specified in the paper.
assumptions (4)
  • domain assumption 2DGS renders geometrically consistent depth and normal images for arbitrary viewpoints.
    This is a stated property of the representation (Sec. III-A) and is the foundation for the tracking improvements. It is inherited from the 2DGS paper [29].
  • domain assumption MASt3R provides metrically-scaled (up to scale) point clouds and relative poses for image pairs, and its features support image retrieval via ASMK.
    Used in Sec. III-G for loop detection and relocalization. The system relies on the pretrained model's generalization without additional training.
  • domain assumption The rendered depth and normal images from the active map are accurate enough for the proposed masks (Eqs. 17-18) to filter invalid pixels without biasing the tracking loss.
    The tracking loss in Eq. 19 uses masks based on rendered normal and opacity; if these are unreliable, the loss could be dominated by incorrect pixels.
  • standard math Lie algebra exp/log maps and the chain rule for SE(3) as used in the derivation (Sec. III-B).
    Standard mathematical background for the on-manifold pose optimization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Globally Consistent RGB-D SLAM with 2D Gaussian Splatting." pith.science (2026). https://pith.science/paper/66A5BGZU

@misc{pith2026250600970,
  author       = {Pith},
  title        = {Pith review of: Globally Consistent RGB-D SLAM with 2D Gaussian Splatting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/66A5BGZU}},
  note         = {Machine review of arXiv:2506.00970}
}
read the original abstract

Recently, 3D Gaussian splatting-based RGB-D SLAM displays remarkable performance of high-fidelity 3D reconstruction. However, the lack of depth rendering consistency and efficient loop closure limits the quality of its geometric reconstructions and its ability to perform globally consistent mapping online. In this paper, we present 2DGS-SLAM, an RGB-D SLAM system using 2D Gaussian splatting as the map representation. By leveraging the depth-consistent rendering property of the 2D variant, we propose an accurate camera pose optimization method and achieve geometrically accurate 3D reconstruction. In addition, we implement efficient loop detection and camera relocalization by leveraging MASt3R, a 3D foundation model, and achieve efficient map updates by maintaining a local active map. Experiments show that our 2DGS-SLAM approach achieves superior tracking accuracy, higher surface reconstruction quality, and more consistent global map reconstruction compared to existing rendering-based SLAM methods, while maintaining high-fidelity image rendering and improved computational efficiency.

Figures

Figures reproduced from arXiv: 2506.00970 by the authors.

Figure 1
Figure 1. Reconstruction results of 2DGS-SLAM on synthetic dataset Replica [72] and real-world dataset ScanNet [10]. We present the [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of 2D Gaussian splatting. The pose and shape of [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. System overview of 2DGS-SLAM. Our system consists of two parallel processes: a front-end and a back-end. Taking RGB-D [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Illustration of the state update process. Images (a)-(d) are [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: We input the current frame and the loop candidate into [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Result of loop closure and map correction. The left figures [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Qualitative comparison of reconstruction results on the ScanNet dataset. The first row presents the reconstructed meshes on sequence [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Qualitative comparison of Rendering results on the TUM dataset. For a fair comparison, we selected non-training views and used [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: (a) The wheeled robot platform used in our experiments [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

103 extracted references · 73 canonical work pages

  1. [1]

    Arandjelovic, P

    R. Arandjelovic, P. Gronat, A. Torii, T. Pajdla, and J. Sivic. NetVLAD: CNN Architecture for Weakly Supervised Place Recog- nition. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2016

  2. [2]

    Arandjelovic and A

    R. Arandjelovic and A. Zisserman. All About VLAD. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2013

  3. [3]

    Azinovi ´c, R

    D. Azinovi ´c, R. Martin-Brualla, D.B. Goldman, M. Nießner, and J. Thies. Neural RGB-D Surface Reconstruction. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2022

  4. [4]

    Bartolomei, L

    L. Bartolomei, L. Teixeira, and M. Chli. Fast Multi-UA V Decentral- ized Exploration of Forests. IEEE Robotics and Automation Letters (RA-L), 8(9):5576–5583, 2023

  5. [5]

    Behley and C

    J. Behley and C. Stachniss. Efficient Surfel-Based SLAM using 3D Laser Range Data in Urban Environments. In Proc. of Robotics: Science and Systems (RSS) , 2018

  6. [6]

    Blanco-Claraco

    J.L. Blanco-Claraco. A flexible framework for accurate lidar odom- etry, map manipulation, and localization. Intl. Journal of Robotics Research (IJRR), 0(0):02783649251316881, 2025

  7. [7]

    Campos, R

    C. Campos, R. Elvira, J.J.G. Rodr ´ıguez, J.M. Montiel, and J.D. Tard´os. Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam. IEEE Trans. on Robotics (TRO) , 37(6):1874–1890, 2021

  8. [8]

    X. Chen, T. L ¨abe, A. Milioto, T. R ¨ohling, O. Vysotska, A. Haag, J. Behley, and C. Stachniss. OverlapNet: Loop Closing for LiDAR- based SLAM. In Proc. of Robotics: Science and Systems (RSS) , 2020

Show all 103 references
  1. [9]

    Curless and M

    B. Curless and M. Levoy. A V olumetric Method for Building Complex Models from Range Images. In Proc. of the Intl. Conf. on Computer Graphics and Interactive Techniques (SIGGRAPH) , 1996

  2. [10]

    A. Dai, A. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner. ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2017

  3. [11]

    A. Dai, M. Nießner, M. Zollh ¨ofer, S. Izadi, and C. Theobalt. BundleFusion: Real-time Globally Consistent 3D Reconstruction using Online Surface Re-integration. ACM Trans. on Graphics (TOG), 36(3):1–18, 2017

  4. [12]

    P. Dai, J. Xu, W. Xie, X. Liu, H. Wang, and W. Xu. High- quality Surface Reconstruction using Gaussian Surfels. In Proc. of the Intl. Conf. on Computer Graphics and Interactive Techniques (SIGGRAPH), 2024

  5. [13]

    Della Corte, I

    B. Della Corte, I. Bogoslavskyi, C. Stachniss, and G. Grisetti. A General Framework for Flexible Multi-Cue Photometric Point Cloud Registration. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2018

  6. [14]

    Dellaert

    F. Dellaert. Factor graphs and GTSAM: A hands-on introduction. Georgia Institute of Technology, Tech. Rep , 2:4, 2012

  7. [15]

    T. Deng, G. Shen, T. Qin, J. Wang, W. Zhao, J. Wang, D. Wang, and W. Chen. PLGSLAM: Progressive Neural Scene Represenation with Local to Global Bundle Adjustment. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2024

  8. [16]

    Di Giammarino, L

    L. Di Giammarino, L. Brizi, T. Guadagnino, C. Stachniss, and G. Grisetti. Md-slam: Multi-cue direct slam. In Proc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS) , 2022

  9. [17]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proc. of the Intl. Conf. on Learning R...

  10. [18]

    Duisterhof, Z

    B. Duisterhof, Z. Lojze, W. Philippe, L. Vincent, C. Yohann, and R. Jerome. MASt3R-SfM: A Fully-Integrated Solution for Uncon- strained Structure-from-Motion. In Proc. of the Intl. Conf. on 3D Vision (3DV), 2025

  11. [19]

    Endres, J

    F. Endres, J. Hess, J. Sturm, D. Cremers, and W. Burgard. 3D Mapping with an RGB-D Camera. IEEE Trans. on Robotics (TRO) , 30(1):177–187, 2014

  12. [20]

    Forster, M

    C. Forster, M. Pizzoli, and D. Scaramuzza. SVO: Fast semi-direct monocular visual odometry. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2014

  13. [21]

    Galvez-L ´opez and J.D

    D. Galvez-L ´opez and J.D. Tard ´os. Bags of Binary Words for Fast Place Recognition in Image Sequences. IEEE Trans. on Robotics (TRO), 28(5):1188–1197, 2012

  14. [22]

    Giacomini, L

    E. Giacomini, L. Di Giammarino, L.D. Rebott, G. Grisetti, and M.R. Oswald. Splat-LOAM: Gaussian Splatting LiDAR Odometry and Mapping. arXiv preprint, arXiv:2503.17491, 2025

  15. [23]

    Glocker, J

    B. Glocker, J. Shotton, A. Criminisi, and S. Izadi. Real-time RGB-D camera relocalization via randomized ferns for keyframe encoding. IEEE Trans. on Visualization and Computer Graphics , 21(5):571– 583, 2014

  16. [24]

    Glover, W

    A. Glover, W. Maddern, M. Warren, S. Reid, M. Milford, and G. Wyeth. Openfabmap: An open source toolbox for appearance- 17 based loop closure detection. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2012

  17. [25]

    Guadagnino, B

    T. Guadagnino, B. Mersch, S. Gupta, I. Vizzo, G. Grisetti, and C. Stachniss. KISS-SLAM: A Simple, Robust, and Accurate 3D LiDAR SLAM System With Enhanced Generalization Capabilities. arXiv preprint, arXiv:2503.12660, 2025

  18. [26]

    S. Ha, J. Yeon, and H. Yu. RGBD GS-ICP SLAM. In Proc. of the Europ. Conf. on Computer Vision (ECCV) , 2024

  19. [27]

    Hornung, K

    A. Hornung, K. Wurm, M. Bennewitz, C. Stachniss, and W. Burgard. OctoMap: An Efficient Probabilistic 3D Mapping Framework Based on Octrees. Autonomous Robots , 34(3):189–206, 2013

  20. [28]

    J. Hu, M. Mao, H. Bao, G. Zhang, and Z. Cui. CP-SLAM: Collaborative Neural Point-based SLAM System. In Proc. of the Conf. on Neural Information Processing Systems (NeurIPS) , 2023

  21. [29]

    Huang, Z

    B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao. 2D Gaussian Splatting for Geometrically Accurate Radiance Fields. In Proc. of the Intl. Conf. on Computer Graphics and Interactive Techniques (SIGGRAPH), 2024

  22. [30]

    Huang, O

    C. Huang, O. Mees, A. Zeng, and W. Burgard. Audio visual language maps for robot navigation. In Proc. of the Intl. Symp. on Experimental Robotics (ISER) , 2023

  23. [31]

    Huang, L

    H. Huang, L. Li, C. Hui, and S.K. Yeung. Photo-SLAM: Real-time Simultaneous Localization and Photorealistic Mapping for Monocu- lar, Stereo, and RGB-D Cameras. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2024

  24. [32]

    Huang, B

    Y . Huang, B. Cui, L. Bai, Z. Chen, J. Wu, Z. Li, H. Liu, and H. Ren. Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-driven Surface Normal-aware Tracking and Mapping. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2025

  25. [33]

    Izquierdo and J

    S. Izquierdo and J. Civera. Close, But Not There: Boosting Geo- graphic Distance Sensitivity in Visual Place Recognition. In Proc. of the Europ. Conf. on Computer Vision (ECCV) , 2024

  26. [34]

    Izquierdo and J

    S. Izquierdo and J. Civera. Optimal Transport Aggregation for Visual Place Recognition. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2024

  27. [35]

    Jiaxin and S

    W. Jiaxin and S. Leutenegger. GSFusion: Online RGB-D Mapping Where Gaussian Splatting Meets TSDF Fusion. IEEE Robotics and Automation Letters (RA-L) , 9(12):11865–11872, 2024

  28. [36]

    L. Jin, X. Zhong, Y . Pan, J. Behley, C. Stachniss, and M. Popovic. ActiveGS: Active Scene Reconstruction using Gaussian Splatting. IEEE Robotics and Automation Letters (RA-L) , 10(5):4866–4873, 2025

  29. [37]

    Johari, C

    M.M. Johari, C. Carta, and F. Fleuret. ESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance Fields. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2023

  30. [38]

    Keetha, J

    N. Keetha, J. Karhade, K.M. Jatavallabhula, G. Yang, S. Scherer, D. Ramanan, and J. Luiten. SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2024

  31. [39]

    Keller, D

    M. Keller, D. Lefloch, M. Lambers, and S. Izadi. Real-time 3D Reconstruction in Dynamic Scenes using Point-based Fusion. In Proc. of the Intl. Conf. on 3D Vision (3DV) , 2013

  32. [40]

    Kerbl, G

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Trans. on Graphics (TOG), 42(4):1–14, 2023

  33. [41]

    C. Kerl, J. Sturm, and D. Cremers. Robust Odometry Estimation for RGB-D Cameras. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2013

  34. [42]

    C. Kerl, J. Sturm, and D. Cremers. Dense visual slam for rgb-d cameras. In Proc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS) , 2013

  35. [43]

    Labbe and F

    M. Labbe and F. Michaud. RTAB-Map: An open-source lidar and visual simultaneous localization and mapping library for large-scale and long-term online operation. Journal of Field Robotics (JFR) , 36(1):416–446, 2019

  36. [44]

    Leroy, Y

    V . Leroy, Y . Cabon, and J. Revaud. Grounding Image Matching in 3D with MASt3R. In Proc. of the Europ. Conf. on Computer Vision (ECCV), 2024

  37. [45]

    T.Y . Lim, B. Sun, M. Pollefeys, and H. Blum. Loop Closure from Two Views: Revisiting PGO for Scalable Trajectory Estimation through Monocular Priors. arXiv preprint, arXiv:2503.16275, 2025

  38. [46]

    L. Liso, E. Sandstr ¨om, V . Yugay, L. Van Gool, and M.R. Oswald. Loopy-slam: Dense neural slam with loop closures. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024

  39. [47]

    Loshchilov and F

    I. Loshchilov and F. Hutter. Decoupled weight decay regularization. In Proc. of the Intl. Conf. on Learning Representations (ICLR) , 2019

  40. [48]

    Maier, A

    D. Maier, A. Hornung, and M. Bennewitz. Real-time navigation in 3D environments based on depth camera data. In Proc. of the IEEE Intl. Conf. on Humanoid Robots , 2012

  41. [49]

    Y . Mao, X. Yu, K. Wang, Y . Wang, R. Xiong, and Y . Liao. NGEL-SLAM: Neural Implicit Representation-based Global Consis- tent Low-Latency SLAM System. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2024

  42. [50]

    Matsuki, R

    H. Matsuki, R. Murai, P.H. Kelly, and A.J. Davison. Gaussian splatting SLAM. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2024

  43. [51]

    Matthias, R

    Z. Matthias, R. Jussi, B. Mario, D. Carsten, and P. Mark. Perspective accurate splatting. In Proc. of Graphics Interface (GI) , 2004

  44. [52]

    Mildenhall, P

    B. Mildenhall, P. Srinivasan, M. Tancik, J. Barron, R. Ramamoorthi, and R. Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. In Proc. of the Europ. Conf. on Computer Vision (ECCV), 2020

  45. [53]

    Milford and G

    M. Milford and G. Wyeth. SeqSLAM: Visual route-based navigation for sunny summer days and stormy winter nights. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2012

  46. [54]

    Mur-Artal and J

    R. Mur-Artal and J. Tard ´os. ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo, and RGB-D Cameras. IEEE Trans. on Robotics (TRO), 33(5):1255–1262, 2017

  47. [55]

    Murai, E

    R. Murai, E. Dexheimer, and A.J. Davison. MASt3R-SLAM: Real- Time Dense SLAM with 3D Reconstruction Priors. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2025

  48. [56]

    Newcombe, S

    R.A. Newcombe, S. Izadi, O. Hilliges, D. Molyneaux, D. Kim, A.J. Davison, P. Kohli, J. Shotton, S. Hodges, and A. Fitzgibbon. KinectFusion: Real-Time Dense Surface Mapping and Tracking. In Proc. of the Intl. Symp. on Mixed and Augmented Reality (ISMAR) , 2011

  49. [57]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H.V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, R. Howes, P.Y . Huang, H. Xu, V . Sharma, S.W. Li, W. Galuba, M. Rabbat, M. Assran, N. Ballas, G. Synnaeve, I. Misra, H. Jegou, J. Mairal, P. Labatut, A....

  50. [58]

    Palazzolo, J

    E. Palazzolo, J. Behley, P. Lottes, P. Giguere, and C. Stachniss. ReFu- sion: 3D Reconstruction in Dynamic Environments for RGB-D Cam- eras Exploiting Residuals. In Proc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS) , 2019

  51. [59]

    Y . Pan, P. Xiao, Y . He, Z. Shao, and Z. Li. MULLS: Versatile LiDAR SLAM Via Multi-Metric Linear Least Square. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2021

  52. [60]

    Y . Pan, X. Zhong, L. Jin, L. Wiesmann, M. Popovi ´c, J. Behley, and C. Stachniss. PINGS: Gaussian Splatting Meets Distance Fields within a Point-Based Implicit Neural Map. In Proc. of Robotics: Science and Systems (RSS) , 2025

  53. [61]

    Y . Pan, X. Zhong, L. Wiesmann, T. Posewsky, J. Behley, and C. Stachniss. PIN-SLAM: LiDAR SLAM Using a Point-Based Im- plicit Neural Representation for Achieving Global Map Consistency. IEEE Trans. on Robotics (TRO) , 40:4045–4064, 2024

  54. [62]

    J.J. Park, P. Florence, J. Straub, R. Newcombe, and S. Lovegrove. DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2019

  55. [63]

    Z. Peng, T. Shao, L. Yong, J. Zhou, Y . Yang, J. Wang, and K. Zhou. RTG-SLAM: Real-time 3D Reconstruction at Scale using Gaussian Splatting. In Proc. of the Intl. Conf. on Computer Graphics and Interactive Techniques (SIGGRAPH) , 2024

  56. [64]

    Reijgwart, A

    V . Reijgwart, A. Millane, H. Oleynikova, R. Siegwart, C. Cadena, and J. Nieto. V oxgraph: Globally consistent, volumetric mapping using signed distance function submaps. IEEE Robotics and Automation Letters (RA-L) , 5(1):227–234, 2019

  57. [65]

    Rublee, V

    E. Rublee, V . Rabaud, K. Konolige, and G. Bradski. Orb: an efficient alternative to sift or surf. In Proc. of the IEEE Intl. Conf. on Computer Vision (ICCV), 2011

  58. [66]

    R. Rusu, N. Blodow, and M. Beetz. Fast point feature histograms (fpfh) for 3d registration. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2009

  59. [67]

    Sandstr ¨om, K

    E. Sandstr ¨om, K. Tateno, M. Oechsle, M. Niemeyer, L. Van Gool, M.R. Oswald, and F. Tombari. Splat-SLAM: Globally Opti- mized RGB-only SLAM with 3D Gaussians. arXiv preprint , arXiv:2405.16544, 2024. 18

  60. [68]

    Sandstr ¨om, Y

    E. Sandstr ¨om, Y . Li, L. Van Gool, and M. R. Oswald. Point-SLAM: Dense Neural Point Cloud-based SLAM. In Proc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV) , 2023

  61. [69]

    Schops, T

    T. Schops, T. Sattler, and M. Pollefeys. BAD SLAM: Bundle Adjusted Direct RGB-D SLAM. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2019

  62. [70]

    Segal, D

    A. Segal, D. Haehnel, and S. Thrun. Generalized-ICP. In Proc. of Robotics: Science and Systems (RSS) , 2009

  63. [71]

    Stachniss

    C. Stachniss. Springer Handbuch der Geod ¨asie, chapter Simultane- ous Localization and Mapping. Springer Verlag, 2016. In German, invited

  64. [72]

    Straub, T

    J. Straub, T. Whelan, L. Ma, Y . Chen, E. Wijmans, S. Green, J.J. Engel, R. Mur-Artal, C. Ren, S. Verma, et al. The Replica dataset: A digital replica of indoor spaces. arXiv preprint, arXiv:1906.05797, 2019

  65. [73]

    St ¨uckler and S

    J. St ¨uckler and S. Behnke. Multi-Resolution Surfel Maps for Efficient Dense 3D Modeling and Tracking. Journal of Visual Communication and Image Representation (JVCIR) , 25(1):137–147, 2014

  66. [74]

    Sturm, N

    J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers. A Benchmark for the Evaluation of RGB-D SLAM Systems. In Proc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS) , 2012

  67. [75]

    Sucar, S

    E. Sucar, S. Liu, J. Ortiz, and A.J. Davison. imap: Implicit mapping and positioning in real-time. In Proc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV) , 2021

  68. [76]

    S. Sun, M. Mielle, A.J. Lilienthal, and M. Magnusson. High-Fidelity SLAM Using Gaussian Splatting with Rendering-Guided Densifi- cation and Regularized Optimization. In Proc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS) , 2024

  69. [77]

    Y . Tang, J. Zhang, Z. Yu, H. Wang, and K. Xu. MIPS-Fusion: Multi- Implicit-Submaps for Scalable and Robust Online Neural RGB-D Reconstruction. ACM Trans. on Graphics (TOG) , 42(6):1–14, 2023

  70. [78]

    Teed and J

    Z. Teed and J. Deng. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. In Proc. of the Conf. on Neural Informa- tion Processing Systems (NeurIPS) , 2021

  71. [79]

    Tolias, Y

    G. Tolias, Y . Avrithis, and H. J´egou. To aggregate or not to aggregate: Selective match kernels for image search. In Proc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV) , 2013

  72. [80]

    S. Umeyama. Least-squares estimation of transformation parameters between two point patterns. IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI) , 13(4):376–380, 1991

  73. [81]

    Vincent, M.N

    L. Vincent, M.N. Francesc, and F. Pascal. EPnP: An Accurate O(n) Solution to the PnP Problem. Intl. Journal of Computer Vision (IJCV), 81:255–166, 2009

  74. [82]

    Vizzo, T

    I. Vizzo, T. Guadagnino, J. Behley, and C. Stachniss. VDBFusion: Flexible and Efficient TSDF Integration of Range Sensor Data. Sensors, 22(3):1296, 2022

  75. [83]

    Vizzo, T

    I. Vizzo, T. Guadagnino, B. Mersch, L. Wiesmann, J. Behley, and C. Stachniss. KISS-ICP: In Defense of Point-to-Point ICP – Simple, Accurate, and Robust Registration If Done the Right Way. IEEE Robotics and Automation Letters (RA-L) , 8(2):1029–1036, 2023

  76. [84]

    Vysotska and C

    O. Vysotska and C. Stachniss. Lazy Data Association For Image Sequences Matching Under Substantial Appearance Changes. IEEE Robotics and Automation Letters (RA-L) , 1(1):213–220, 2016

  77. [85]

    H. Wang, J. Wang, and L. Agapito. Co-SLAM: Joint Coordinate and Sparse Parametric Encodings for Neural Real-Time SLAM. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2023

  78. [86]

    S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud. DUst3R: Geometric 3D Vision Made Easy. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2024

  79. [87]

    Wang, A.C

    Z. Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Trans. on Image Processing , 13(4):600–612, 2004

  80. [88]

    Werby, C

    A. Werby, C. Huang, M. B ¨uchner, A. Valada, and W. Burgard. Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation. In Proc. of Robotics: Science and Systems (RSS) , 2024

  81. [89]

    Whelan, M

    T. Whelan, M. Kaess, M. Fallon, H. Johannsson, J. Leonard, and J. McDonald. Kintinuous: Spatially Extended KinectFusion. In Proc. of the RSS Workshop on RGB-D: Advanced Reasoning with Depth Cameras , 2012

  82. [90]

    Whelan, S

    T. Whelan, S. Leutenegger, R.S. Moreno, B. Glocker, and A. Davison. ElasticFusion: Dense SLAM Without A Pose Graph. In Proc. of Robotics: Science and Systems (RSS) , 2015

  83. [91]

    K. Wu, Z. Zhang, M. Tie, Z. Ai, Z. Gan, and W. Ding. VINGS- Mono: Visual-Inertial Gaussian Splatting Monocular SLAM in Large Scenes. arXiv preprint, arXiv:2501.08286, 2025

  84. [92]

    C. Yan, D. Qu, D. Xu, B. Zhao, Z. Wang, D. Wang, and X. Li. GS- SLAM: Dense Visual SLAM with 3D Gaussian Splatting. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024

  85. [93]

    X. Yang, H. Li, H. Zhai, Y . Ming, Y . Liu, and G. Zhang. V ox-Fusion: Dense Tracking and Mapping with V oxel-based Neural Implicit Representation. In Proc. of the Intl. Symp. on Mixed and Augmented Reality (ISMAR) , 2022

  86. [94]

    S. Yu, C. Cheng, Y . Zhou, X. Yang, and H. Wang. OpenGS- SLAM: RGB-Only Gaussian Splatting SLAM for Unbounded Out- door Scenes. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2025

  87. [95]

    Yugay, T

    V . Yugay, T. Gevers, and M.R. Oswald. MAGiC-SLAM: Multi-Agent Gaussian Globally Consistent SLAM. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2025

  88. [96]

    Yugay, Y

    V . Yugay, Y . Li, T. Gevers, and M.R. Oswald. Gaussian-SLAM: Photo-realistic Dense SLAM with Gaussian Splatting. arXiv preprint, arXiv:2312.10070, 2023

  89. [97]

    Zhang, E

    G. Zhang, E. Sandstr ¨om, Y . Zhang, M. Patel, L. Van Gool, and M.R. Oswald. GlORIE-SLAM: Globally Optimized Rgb-only Implicit Encoding Point Cloud SLAM. arXiv preprint , arXiv:2403.19549, 2024

  90. [98]

    Zhang, P

    R. Zhang, P. Isola, A.A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2018

  91. [99]

    Zhang, F

    Y . Zhang, F. Tosi, S. Mattoccia, and M. Poggi. GO-SLAM: Global Optimization for Consistent 3D Instant Reconstruction. In Proc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV) , 2023

  92. [100]

    Zhong, Y

    X. Zhong, Y . Pan, J. Behley, and C. Stachniss. SHINE-Mapping: Large-Scale 3D Mapping Using Sparse Hierarchical Implicit Neural Representations. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2023

  93. [101]

    X. Zhou, Z. Lin, X. Shan, Y . Wang, D. Sun, and M.H. Yang. DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2024

  94. [102]

    L. Zhu, Y . Li, E. Sandstr ¨om, S. Huang, K. Schindler, and I. Armeni. LoopSplat: Loop Closure by Registering 3D Gaussian Splats. In Proc. of the Intl. Conf. on 3D Vision (3DV) , 2025

  95. [103]

    Z. Zhu, S. Peng, V . Larsson, W. Xu, H. Bao, Z. Cui, M.R. Oswald, and M. Pollefeys. Nice-slam: Neural implicit scalable encoding for slam. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , 2022

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.