Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey

T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Semantic geometric SLAM is the only family that runs near real time on embedded hardware.

desk verdict A useful but uneven embedded semantic SLAM survey whose central claim is plausible yet rests on a benchmark with incomparable preprocessing conditions. read the letter →

arxiv 2505.12384 v1 pith:TJFSQDFM submitted 2025-05-18 cs.RO

classification cs.RO
keywords semanticSLAMembeddedsystemsNeRF3DGaussianSplattingresourceutilizationJetsonAGXOrindynamicenvironmentssegmentation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Robots that build maps with object labels—semantic SLAM—must fit accuracy, memory, and power budgets when deployed on embedded processors. The paper compares the three main architecture families—geometric, NeRF-based, and 3D Gaussian Splatting—on a single embedded platform, the Jetson AGX Orin, measuring localization error, segmentation quality, memory, speed, and energy. Its central finding is that semantic geometric SLAM is currently the most viable family for real-time embedded deployment, with Dynamic-VINS reaching 24.9 FPS at 12 W, while Gaussian-splatting systems run at 0.013–0.014 FPS and NeRF systems could not be run at all. NeRF and Gaussian Splatting still offer the richest semantic maps, so the paper frames the conclusion as a call for efficiency work and algorithm–hardware co-design rather than a permanent verdict.

What carries the argument

The argument is carried by a resource-utilization benchmark on a single embedded platform, the Jetson AGX Orin, that measures trajectory accuracy (ATE RMSE, the root-mean-square error between estimated and ground-truth trajectories), semantic quality (mIoU, the average overlap between predicted and true masks), memory, power, and speed across the three architecture families, using a dynamic RGB-D benchmark for geometric accuracy and the Replica dataset for semantic quality. The decisive mechanism is the semantic-integration mode: systems that apply segmentation only to keyframes or run it live, versus systems that consume precomputed semantic masks, determine whether semantics can stay within real-time budgets. These measurements are what separate the geometric family's near-real-time performance from the neural-representation families' high semantic detail.

What would settle it

Re-running the benchmark with all systems required to compute semantic segmentation live on the Jetson AGX Orin, and reporting per-sequence accuracy on the dynamic RGB-D benchmark including the hardest walking sequence (fr3/walking_xyz), would settle the ranking: if Gaussian-splatting systems then approach real time while geometric systems stall, or if the geometric advantage disappears on the hardest sequences, the conclusion fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that, on the Jetson AGX Orin, semantic geometric SLAM is currently the only architecture family that balances accuracy, memory, power, and speed well enough for real-time embedded deployment: Dynamic-VINS runs at 24.9 FPS with 8.29 GB RAM and 12 W, and VDO-SLAM at 20.9 FPS with 6.62 GB and 10.9 W, while the two Gaussian-splatting systems, GS3LAM and SGS-SLAM, run at 0.013–0.014 FPS with over 16 GB RAM and over 15 W even when consuming precomputed semantic masks. NeRF-based semantic SLAM systems could not be executed on the Orin under the paper's setup, and their reported accuracy on the Replica dataset, while sub-centimeter, does not close the embedded-deployment gap. The paper also observes that Gaussian-splatting methods are more efficient than NeRF methods while delivering comparable reconstruction quality and semantic consistency, but still far from real time on embedded hardware.

Load-bearing premise

The central comparison assumes that systems given precomputed semantic masks can be ranked for embedded deployment on the same footing as systems that must run segmentation live, and that the average over the chosen RGB-D dynamic sequences fairly represents dynamic-scene performance.

Editorial extensions

If this is right

  • Embedded semantic SLAM research should prioritize geometric pipelines with lightweight, keyframe-only semantic modules, since those delivered the only near-real-time results on the Orin.
  • Gaussian-splatting SLAM needs memory and compute reductions of more than an order of magnitude before real-time embedded deployment; at 0.013–0.014 FPS it is two orders of magnitude below real-time.
  • Panoptic semantics are not yet viable on embedded hardware, with Panoptic-SLAM reaching only 2.7 FPS, so the current embedded trade-off favors object-level or lightweight semantic segmentation.
  • NeRF-based semantic SLAM, despite sub-centimeter accuracy on Replica, could not be run on the Orin in this study, leaving its embedded feasibility unproven.
  • Algorithm–hardware co-design, including accelerators and lightweight model variants, is the paper's proposed path for making dense neural representations embeddable.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to rerun the comparison with hardware-accelerated neural-network inference, since the paper deliberately used unoptimized inference; such acceleration could change the ranking if geometric systems benefit more from faster segmentation.
  • A fairer embedded-deployment metric would count segmentation cost in the end-to-end loop for every system, since the current protocol gives precomputed-mask systems an advantage and the geometric-vs-Gaussian gap could narrow or widen under uniform treatment.
  • The dynamic-scene comparison is currently only geometric; a dynamic benchmark with semantic annotations could test whether Gaussian-splatting systems' high mIoU degrades when objects move and are occluded.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This manuscript surveys three families of semantic visual SLAM (geometric, NeRF, and Gaussian Splatting), then reports an experimental comparison of a selected subset (RDS-SLAM, VDO-SLAM, Dynamic-VINS, Panoptic-SLAM, GS3LAM, SGS-SLAM) on an NVIDIA Jetson AGX Orin using ATE, mIoU, RAM, power, and FPS metrics. The central conclusion is that Semantic-aware Geometric SLAM currently provides the most viable solution for real-time embedded deployment, while NeRF- and 3DGS-based methods are too resource-intensive for constrained platforms. The paper also argues that GS-based systems are more computationally efficient than NeRF-based ones, and it discusses future directions such as hardware-software co-design.

Significance. If accepted, the central conclusion would give the embedded SLAM community a clear research direction: focus on geometric pipelines with selective semantics rather than dense neural scene representations. The paper's strengths are its concrete resource measurements on the Jetson AGX Orin, the containerized experimental setup, the per-component timing breakdown in Figure 9, and the broad coverage of recent literature. The conclusion is plausible, but the support has load-bearing gaps: Table IV mixes end-to-end and off-device semantic processing, Table II averages over an unspecified TUM sequence subset, and the NeRF-versus-GS efficiency comparison is not backed by any Orin measurements of NeRF systems.

major comments (4)
  1. [Section IV-C, Table IV] Table IV is the main evidence for the paper's central claim, but it compares systems under different semantic-preprocessing conditions. As the text states, VDO-SLAM consumed precomputed Mask R-CNN masks in an offline preprocessing step, so its reported 20.9 FPS and 10.9 W exclude the semantic extraction cost; Section II-A further notes that VDO-SLAM is only applicable to pre-recorded benchmarks. The conclusion in Section V cites VDO-SLAM alongside Dynamic-VINS as evidence that selective semantic processing keeps costs manageable, yet VDO-SLAM's measurements are not end-to-end. Please provide end-to-end numbers with on-device segmentation, or explicitly add the segmentation time and energy to VDO-SLAM's totals, and flag in Table IV which systems perform semantics in real time.
  2. [Section IV-B, Table II] Table II reports a single average ATE over TUM RGB-D sequences, but the sequence subset is not stated. The ORB-SLAM2 baseline of 1.0 cm ATE RMSE is far below typical values on the dynamic walking sequences, which strongly suggests that the hardest sequences are excluded from the average. Since the paper claims that semantic geometric methods improve accuracy in dynamic environments, the table should either list per-sequence ATE for all TUM sequences used or clearly state the subset and justify it; otherwise the accuracy advantage is not demonstrated on the regime that motivates semantic SLAM.
  3. [Section V, with Section IV-C and Table III] The conclusion that "GS-enhanced SLAM systems typically offer better computational efficiency than NeRF-based methods" is not supported by the experiments in this paper, because no NeRF-based system was run on the Jetson AGX Orin: Section IV-C states that NIS-SLAM and vMAP could not be executed due to insufficient resources, and the NeRF entries in Tables II and III are literature values from [5] that were obtained on different platforms. Either run representative NeRF systems on the same Orin setup, or restrict the conclusion to what Table IV can support, namely geometric systems versus the two GS systems.
  4. [Section IV-A, Section IV-C] The power figures in Table IV are load-bearing for the embedded-deployment recommendation, but the manuscript does not describe how power was measured (e.g., wall meter, on-module sensors, averaging window, or load conditions). Without this methodology, even the end-to-end comparisons that remain after the VDO-SLAM issue are hard to interpret; please add a short measurement-protocol paragraph.
minor comments (4)
  1. [Section IV, first paragraph] RDS-SLAM is described as being built on the ORB-SLAM2 framework, but Table I and Section II-A state ORB-SLAM3; please reconcile this inconsistency.
  2. [Section IV-B, first sentence] The sentence "For these experiments, two widely used datasets TUM RGB-D and Replica" lacks a verb; rephrase to something like "For these experiments, we use two widely used datasets: TUM RGB-D and Replica."
  3. [Table I] The bullet symbol (•) is used in several columns without a legend; please state that a bullet means "yes" or otherwise explain the notation.
  4. [Figure 6 caption] The caption contains a typo: "an be incorporated" should read "can be incorporated."

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the survey's conclusions rest on measured or externally reported benchmark numbers, and the only author self-citation [16] is background context, not load-bearing.

full rationale

The paper contains no derived prediction or fitted-parameter claim whose output is equivalent to its input by construction. The central conclusion, that semantic-aware geometric SLAM is currently the most viable option for real-time embedded deployment, is supported by Tables II?IV, which report ATE, mIoU, RAM, power, and FPS values. Those values come either from the authors' own runs on the Jetson AGX Orin or from state-of-the-art reports marked with asterisks from the external survey [5]; they are not obtained from an equation that defines the conclusion into existence. The only self-citation is reference [16], by Salhi, Poreba, et al., used in Section I as background ('In the context of embedded devices, the survey of [16] presents the existing multimodal localization techniques'). That citation is descriptive context and does not justify the ranking, the accuracy comparison, or the architecture taxonomy, so it is not load-bearing. The paper itself flags the main validity limitation in Section IV-C: for VDO-SLAM, 'we followed the authors' recommendation and used Mask-RCNN pre-processing offline. This means that the second place achieved by this system is only possible because it only needs to load pre-computed segmentation masks.' This is an experimental-confounding issue that could affect the fairness of the FPS/power comparison, but it is not a circular derivation: no quantity is defined in terms of the target conclusion. Similarly, the ORB-SLAM2 baseline of 1.0 cm on TUM RGB-D may indicate selection of easier sequences, but that is a benchmark-representativeness concern, not a circularity concern. No self-definitional step, no fitted-input-renamed-as-prediction step, and no uniqueness theorem or ansatz imported from the authors' prior work appears in the manuscript. The survey is therefore self-contained in the sense relevant to circularity analysis; its weaknesses, if any, belong to experimental design and correctness risk, not to circular reasoning.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim relies on comparability assumptions about benchmark protocols rather than free parameters or invented entities. No numbers are fitted; the paper's conclusions rest on the representativeness of the selected systems, the TUM and Replica evaluation, and the direct comparability of FPS and power across different preprocessing setups.

assumptions (3)
  • domain assumption The TUM RGB-D results are averaged over sequences that include dynamic walking sequences; if the average excludes the hardest dynamic sequences, the ORB-SLAM2 baseline of 1.0 cm would be artificially low.
    Section IV-B states results will be reported as averages across sequences covering static to highly dynamic environments, but the ORB-SLAM2 baseline of 1.0 cm is inconsistent with typical failure on walking sequences.
  • ad hoc to paper Direct comparison of FPS and resource usage across systems with different preprocessing pipelines is meaningful for embedded deployment assessment.
    Section IV-C compares VDO-SLAM and GS systems that use offline or precomputed semantic masks to systems that run segmentation in real time; the comparison assumes this does not bias the efficiency ranking.
  • domain assumption Default settings and PyTorch inference without TensorRT are representative of a fair embedded deployment comparison.
    Section IV states tests were run with default settings and no TensorRT optimization; this may disadvantage some systems more than others but is disclosed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey." pith.science (2026). https://pith.science/paper/TJFSQDFM

@misc{pith2026250512384,
  author       = {Pith},
  title        = {Pith review of: Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TJFSQDFM}},
  note         = {Machine review of arXiv:2505.12384}
}
read the original abstract

In embedded systems, robots must perceive and interpret their environment efficiently to operate reliably in real-world conditions. Visual Semantic SLAM (Simultaneous Localization and Mapping) enhances standard SLAM by incorporating semantic information into the map, enabling more informed decision-making. However, implementing such systems on resource-limited hardware involves trade-offs between accuracy, computing efficiency, and power usage. This paper provides a comparative review of recent Semantic Visual SLAM methods with a focus on their applicability to embedded platforms. We analyze three main types of architectures - Geometric SLAM, Neural Radiance Fields (NeRF), and 3D Gaussian Splatting - and evaluate their performance on constrained hardware, specifically the NVIDIA Jetson AGX Orin. We compare their accuracy, segmentation quality, memory usage, and energy consumption. Our results show that methods based on NeRF and Gaussian Splatting achieve high semantic detail but demand substantial computing resources, limiting their use on embedded devices. In contrast, Semantic Geometric SLAM offers a more practical balance between computational cost and accuracy. The review highlights a need for SLAM algorithms that are better adapted to embedded environments, and it discusses key directions for improving their efficiency through algorithm-hardware co-design.

Figures

Figures reproduced from arXiv: 2505.12384 by the authors.

Figure 1
Figure 1. Timeline of key advancements in Semantic SLAM research from 2018 to 2024. The progression is categorized into three primary approaches: Semantic [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Typical pipeline of a geometric SLAM system, showing the funda [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Citation network and relationships in Semantic SLAM research. The visualization represents five distinct categories: Geometric SLAM approaches [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Architecture of RDS-SLAM [18] based on ORB-SLAM3. The [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Overview of VDO-SLAM architecture [19]. The system extends [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Overview of NeRF-enhanced SLAM architecture. The system in [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Overview of Gaussian Splatting-enhanced SLAM architecture. This [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Overview of Dynamic-VINS architecture [20]. The system operates [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: Performance comparison of Semantic Geometric SLAM approaches, showing processing time (in seconds, log scale) breakdown into mapping (red), tracking (blue) and segmentation (green) components. All systems, except for Dynamic-VINS, rely on keyframes to optimize processi…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robotic Contextual Awareness for Human-Robot Collaboration and Environmental Understanding

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Novel person re-identification with continual adaptation plus submap LiDAR SLAM, ground-aware filtering, Gaussian Scan Context, and multi-modal semantic mapping improve robotic contextual awareness for HRC and navigation.

Reference graph

Works this paper leans on

156 extracted references · 37 canonical work pages · cited by 1 Pith paper

  1. [17]

    A survey on real-time 3d scene reconstruction with slam methods in embedded systems,

    Q. Picard, S. Chevobbe, M. Darouich, and J.-Y . Didier, “A survey on real-time 3d scene reconstruction with slam methods in embedded systems,” ArXiv, vol. abs/2309.05349, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:261682162

  2. [5]

    How nerfs and 3d gaussian splatting are reshaping slam: a survey,

    F. Tosi, Y . Zhang, Z. Gong, E. Sandstr¨om, S. Mattoccia, M. R. Oswald, and M. Poggi, “How nerfs and 3d gaussian splatting are reshaping slam: a survey,” 2024. [Online]. Available: https://arxiv.org/abs/2402.13255

  3. [1]

    Semantic visual simultaneous localization and mapping: A survey,

    K. Chen, J. Zhang, J. Liu, Q. Tong, R. Liu, and S. Chen, “Semantic visual simultaneous localization and mapping: A survey,” 2022. [Online]. Available: https://arxiv.org/abs/2209.06428

  4. [2]

    A survey of visual slam in dynamic environment: The evolution from geometric to semantic approaches,

    Y . Wang, Y . Tian, J. Chen, K. Xu, and X. Ding, “A survey of visual slam in dynamic environment: The evolution from geometric to semantic approaches,” IEEE Transactions on Instrumentation and Measurement, vol. 73, pp. 1–21, 2024

  5. [3]

    A survey of image semantics-based visual simultaneous localization and mapping: Application-oriented solutions to autonomous navigation of mobile robots,

    L. Xia, J. Cui, R. Shen, X. Xu, Y . Gao, and X. Li, “A survey of image semantics-based visual simultaneous localization and mapping: Application-oriented solutions to autonomous navigation of mobile robots,” International Journal of Advanced Robotic Systems , vol. 17, p. 172988142091918, 05 2020

  6. [4]

    An overview on visual slam: From tradition to semantic,

    W. Chen, G. Shang, A. Ji, C. Zhou, X. Wang, C. Xu, Z. Li, and K. Hu, “An overview on visual slam: From tradition to semantic,” Remote Sensing, vol. 14, no. 13, 2022. [Online]. Available: https://www.mdpi.com/2072-4292/14/13/3010

  7. [6]

    Nerf in robotics: A survey,

    G. Wang, L. Pan, S. Peng, S. Liu, C. Xu, Y . Miao, W. Zhan, M. Tomizuka, M. Pollefeys, and H. Wang, “Nerf in robotics: A survey,” 2024. [Online]. Available: https://arxiv.org/abs/2405.01333

  8. [7]

    Slam meets nerf: A survey of implicit slam methods,

    K. Yang, Y . Cheng, Z. Chen, and J. Wang, “Slam meets nerf: A survey of implicit slam methods,” World Electric Vehicle Journal , vol. 15, no. 3, 2024. [Online]. Available: https://www.mdpi.com/2032-6653/ 15/3/85

Show all 156 references
  1. [8]

    Neural fields in robotics: A survey,

    M. Z. Irshad, M. Comi, Y .-C. Lin, N. Heppert, A. Valada, R. Ambrus, Z. Kira, and J. Tremblay, “Neural fields in robotics: A survey,” 2024. [Online]. Available: https://arxiv.org/abs/2410.20220

  2. [9]

    A survey on 3d gaussian splatting,

    G. Chen and W. Wang, “A survey on 3d gaussian splatting,” ArXiv, vol. abs/2401.03890, 2024. [Online]. Available: https://api. semanticscholar.org/CorpusID:266844057

  3. [10]

    3d gaussian splatting: Survey, technologies, challenges, and opportunities,

    Y . Bao, T. Ding, J. Huo, Y . Liu, Y . Li, W. Li, Y . Gao, and J. Luo, “3d gaussian splatting: Survey, technologies, challenges, and opportunities,” ArXiv, vol. abs/2407.17418, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:271404782

  4. [11]

    Customizable perturbation synthesis for robust slam benchmarking,

    X. Xu, T. Zhang, S. Wang, X. Li, Y . Chen, Y . Li, B. Raj, M. Johnson-Roberson, and X. Huang, “Customizable perturbation synthesis for robust slam benchmarking,” 2024. [Online]. Available: https://arxiv.org/abs/2402.08125

  5. [12]

    From perfect to noisy world simulation: Customizable embodied multi-modal perturbations for slam robustness benchmarking,

    ——, “From perfect to noisy world simulation: Customizable embodied multi-modal perturbations for slam robustness benchmarking,” 2024. [Online]. Available: https://arxiv.org/abs/2406.16850

  6. [13]

    Benchmarking implicit neural representation and geometric rendering in real-time rgb-d slam,

    T. Hua and L. Wang, “Benchmarking implicit neural representation and geometric rendering in real-time rgb-d slam,” 2024. [Online]. Available: https://arxiv.org/abs/2403.19473

  7. [14]

    Benchmarking neural radiance fields for autonomous robots: An overview,

    Y . Ming, X. Yang, W. Wang, Z. Chen, J. Feng, Y . Xing, and G. Zhang, “Benchmarking neural radiance fields for autonomous robots: An overview,” 2024. [Online]. Available: https://arxiv.org/abs/2405.05526

  8. [15]

    Evaluating modern approaches in 3d scene reconstruction: Nerf vs gaussian-based methods,

    Y . Zhou, Z. Zeng, A. Chen, X. Zhou, H. Ni, S. Zhang, P. Li, L. Liu, M. Zheng, and X. Chen, “Evaluating modern approaches in 3d scene reconstruction: Nerf vs gaussian-based methods,” in 2024 6th International Conference on Data-driven Optimization of Complex Systems (DOCS) . I...

  9. [16]

    Chapter 8 - multimodal localization for embedded systems: A survey,

    I. Salhi, M. Poreba, E. Piriou, V . Gouet-Brunet, and M. Ojail, “Chapter 8 - multimodal localization for embedded systems: A survey,” in Multimodal Scene Understanding , M. Y . Yang, B. Rosenhahn, and V . Murino, Eds. Academic Press, 2019, pp. 199–278. [Online]. Available: htt...

  10. [18]

    Rds-slam: Real-time dynamic slam using semantic segmentation methods,

    Y . Liu and J. Miura, “Rds-slam: Real-time dynamic slam using semantic segmentation methods,” IEEE Access , vol. 9, pp. 23 772– 23 785, 2021

  11. [19]

    VDO-SLAM: A Visual Dynamic Object-aware SLAM System,

    J. Zhang, M. Henein, R. Mahony, and V . Ila, “VDO-SLAM: A Visual Dynamic Object-aware SLAM System,” 2020

  12. [20]

    Rgb-d inertial odometry for a resource-restricted robot in dynamic environments,

    J. Liu, X. Li, Y . Liu, and H. Chen, “Rgb-d inertial odometry for a resource-restricted robot in dynamic environments,” IEEE Robotics and Automation Letters, vol. 7, no. 4, pp. 9573–9580, 2022

  13. [21]

    Panoptic-slam: Visual slam in dynamic environments using panoptic segmentation,

    G. F. Abati, J. C. V . Soares, V . S. Medeiros, M. A. Meggiolaro, and C. Semini, “Panoptic-slam: Visual slam in dynamic environments using panoptic segmentation,” 2024

  14. [22]

    GS$ˆ {3}$LAM: Gaussian semantic splatting SLAM,

    L. Li, L. Zhang, Z. Wang, and Y . Shen, “GS$ˆ {3}$LAM: Gaussian semantic splatting SLAM,” in ACM Multimedia 2024 , 2024. [Online]. Available: https://openreview.net/forum?id=juMYrkJlV3

  15. [23]

    Sni-slam: Semantic neural implicit slam,

    S. Zhu, G. Wang, H. Blum, J. Liu, L. Song, M. Pollefeys, and H. Wang, “Sni-slam: Semantic neural implicit slam,” 2024. [Online]. Available: https://arxiv.org/abs/2311.11016

  16. [24]

    Sgs-slam: Semantic gaussian splatting for neural dense slam,

    M. Li, S. Liu, H. Zhou, G. Zhu, N. Cheng, T. Deng, and H. Wang, “Sgs-slam: Semantic gaussian splatting for neural dense slam,” Feb 2024, european Conference on Computer Vision (ECCV) 2024. [Online]. Available: http://arxiv.org/abs/2402.03246v5

  17. [25]

    Octomap: an efficient probabilistic 3d mapping framework based on octrees,

    A. Hornung, K. M. Wurm, M. Bennewitz, C. Stachniss, and W. Burgard, “Octomap: an efficient probabilistic 3d mapping framework based on octrees,” Autonomous Robots , vol. 34, pp. 189 – 206, 2013. [Online]. Available: https://api.semanticscholar.org/ CorpusID:8655888

  18. [26]

    Faster r-cnn: Towards real-time object detection with region proposal networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” 2016. [Online]. Available: https://arxiv.org/abs/1506.01497

  19. [27]

    You only look once: Unified, real-time object detection,

    J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 779–788

  20. [28]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. ...

  21. [29]

    Dunet: A deformable network for retinal vessel segmentation,

    Q. Jin, Z. Meng, T. D. Pham, Q. Chen, L. Wei, and R. Su, “Dunet: A deformable network for retinal vessel segmentation,” Knowledge- Based Systems , vol. 178, pp. 149–162, 2019. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0950705119301984

  22. [30]

    Pyramid scene parsing network,

    H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in CVPR, 2017

  23. [31]

    Segnet: A deep convolutional encoder-decoder architecture for image segmentation,

    V . Badrinarayanan, A. Kendall, and R. Cipolla, “Segnet: A deep convolutional encoder-decoder architecture for image segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 39, pp. 2481–2495, 2015. [Online]. Available: https://api. semanticscholar....

  24. [32]

    Hardnet: A low memory traffic network,

    P. Chao, C.-Y . Kao, Y .-S. Ruan, C.-H. Huang, and Y .-L. Lin, “Hardnet: A low memory traffic network,” 2019. [Online]. Available: https://arxiv.org/abs/1909.00948

  25. [33]

    Bisenet v2: Bilateral network with guided aggregation for real- time semantic segmentation,

    C. Yu, C. Gao, J. Wang, G. Yu, C. Shen, and N. Sang, “Bisenet v2: Bilateral network with guided aggregation for real- time semantic segmentation,” 2020. [Online]. Available: https: //arxiv.org/abs/2004.02147 16

  26. [34]

    Mask r-cnn,

    K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in 2017 IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 2980–2988

  27. [35]

    Yolact++ better real-time instance segmentation,

    D. Bolya, C. Zhou, F. Xiao, and Y . J. Lee, “Yolact++ better real-time instance segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 2, pp. 1108–1121, 2022

  28. [37]

    Psmd-slam: Panoptic segmentation-aided multi-sensor fusion simultaneous localization and mapping in dynamic scenes,

    C. Song, B. Zeng, J. Cheng, F. Wu, and F. Hao, “Psmd-slam: Panoptic segmentation-aided multi-sensor fusion simultaneous localization and mapping in dynamic scenes,” Applied Sciences , vol. 14, no. 9, 2024. [Online]. Available: https://www.mdpi.com/2076-3417/14/9/3843

  29. [38]

    V olumetric semantically consistent 3d panoptic mapping,

    Y . Miao, I. Armeni, M. Pollefeys, and D. Barath, “V olumetric semantically consistent 3d panoptic mapping,” 2024. [Online]. Available: https://arxiv.org/abs/2309.14737

  30. [39]

    Panoptic Feature Pyramid Networks ,

    A. Kirillov, R. Girshick, K. He, and P. Dollar, “ Panoptic Feature Pyramid Networks ,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Los Alamitos, CA, USA: IEEE Computer Society, Jun. 2019, pp. 6392–

  31. [40]

    Segment everything everywhere all at once,

    X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, and Y . J. Lee, “Segment everything everywhere all at once,” inProceedings of the 37th International Conference on Neural Information Processing Systems, ser. NIPS ’23. Red Hook, NY , USA: Curran Associates Inc., 2023

  32. [41]

    Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,

    R. Mur-Artal and J. D. Tardos, “Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,” IEEE Transactions on Robotics , vol. 33, no. 5, p. 1255–1262, Oct. 2017. [Online]. Available: http://dx.doi.org/10.1109/TRO.2017.2705103

  33. [42]

    Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam,

    C. Campos, R. Elvira, J. J. G. Rodr ´ıguez, J. M. M. Montiel, and J. D. Tard ´os, “Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam,” IEEE Transactions on Robotics , vol. 37, no. 6, pp. 1874–1890, 2021

  34. [43]

    Towards real-time semantic rgb-d slam in dynamic environments,

    T. Ji, C. Wang, and L. Xie, “Towards real-time semantic rgb-d slam in dynamic environments,” 2021 IEEE International Conference on Robotics and Automation (ICRA) , pp. 11 175–11 181, 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:233025267

  35. [44]

    Rtsdm: A real-time semantic dense mapping system for uavs,

    Z. Li, J. Zhao, X. Zhou, S. Wei, P. Li, and F. Shuang, “Rtsdm: A real-time semantic dense mapping system for uavs,” Machines, vol. 10, no. 4, 2022. [Online]. Available: https://www.mdpi.com/ 2075-1702/10/4/285

  36. [45]

    Solo-slam: A parallel semantic slam algorithm for dynamic scenes,

    L. Sun, J. Wei, S. Su, and P. Wu, “Solo-slam: A parallel semantic slam algorithm for dynamic scenes,” Sensors, vol. 22, no. 18, 2022. [Online]. Available: https://www.mdpi.com/1424-8220/22/18/6977

  37. [46]

    Rdmo-slam: Real-time visual slam for dynamic environments using semantic label prediction with optical flow,

    Y . Liu and J. Miura, “Rdmo-slam: Real-time visual slam for dynamic environments using semantic label prediction with optical flow,” IEEE Access, vol. 9, pp. 106 981–106 997, 2021

  38. [47]

    D-vins: Dynamic adaptive visual–inertial slam with imu prior and semantic constraints in dynamic scenes,

    Y . Sun, Q. Wang, C. Yan, Y . Feng, R. Tan, X. Shi, and X. Wang, “D-vins: Dynamic adaptive visual–inertial slam with imu prior and semantic constraints in dynamic scenes,” Remote Sensing , vol. 15, no. 15, 2023. [Online]. Available: https://www.mdpi.com/2072-4292/ 15/15/3881

  39. [48]

    Semantic visual slam in dynamic environment,

    Wen, Li, Zhao et al., “Semantic visual slam in dynamic environment,” Auton Robot, vol. 45, p. 493–504, 2021

  40. [49]

    Fch-slam: A slam method for dynamic environments using semantic segmentation,

    Y . Wang, M. Mikawa, and M. Fujisawa, “Fch-slam: A slam method for dynamic environments using semantic segmentation,” in 2022 2nd In- ternational Conference on Image Processing and Robotics (ICIPRob) , 2022, pp. 1–6

  41. [50]

    Wf-slam: A robust vslam for dynamic scenarios via weighted features,

    Y . Zhong, S. Hu, G. Huang, L. Bai, and Q. Li, “Wf-slam: A robust vslam for dynamic scenarios via weighted features,” IEEE Sensors Journal, vol. 22, pp. 1–1, 06 2022

  42. [51]

    Slamantic - leveraging semantics to improve vslam in dynamic environments,

    M. Sch ¨orghuber, D. Steininger, Y . Cabon, M. Humenberger, and M. Gelautz, “Slamantic - leveraging semantics to improve vslam in dynamic environments,” in 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW) , 2019, pp. 3759–3768

  43. [52]

    Sad-slam: A visual slam based on semantic and depth information,

    X. Yuan and S. Chen, “Sad-slam: A visual slam based on semantic and depth information,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE Press, 2020, p. 4930–4935. [Online]. Available: https://doi.org/10.1109/IROS45743. 2020.9341180

  44. [53]

    Ds-slam: A semantic visual slam towards dynamic environments

    C. Yu, Z. Liu, X.-J. Liu, F. Xie, Y . Yang, Q. Wei, and Q. Fei, “Ds-slam: A semantic visual slam towards dynamic environments.” IEEE Press, 2018, p. 1168–1174. [Online]. Available: https://doi.org/10.1109/IROS.2018.8593691

  45. [54]

    Dynaslam: Tracking, mapping, and inpainting in dynamic scenes,

    B. Besc ´os, J. M. F ´acil, J. Civera, and J. Neira, “Dynaslam: Tracking, mapping, and inpainting in dynamic scenes,” IEEE Robotics and Automation Letters, vol. 3, pp. 4076–4083, 2018. [Online]. Available: https://api.semanticscholar.org/CorpusID:49207678

  46. [55]

    Sof-slam: A semantic visual slam for dynamic environments,

    L. Cui and C. Ma, “Sof-slam: A semantic visual slam for dynamic environments,” IEEE Access, vol. 7, pp. 166 528–166 539, 2019

  47. [56]

    Dynamic scene semantics slam based on semantic segmentation,

    S. Han and Z. Xi, “Dynamic scene semantics slam based on semantic segmentation,” IEEE Access, vol. PP, pp. 1–1, 03 2020

  48. [57]

    Ofm-slam: A visual semantic slam for dynamic indoor environments,

    X. Zhao, T. Zuo, and X. Hu, “Ofm-slam: A visual semantic slam for dynamic indoor environments,” Mathematical Problems in Engineering, vol. 2021, no. 1, p. 5538840, 2021. [Online]. Available: https://onlinelibrary.wiley.com/doi/abs/10.1155/2021/5538840

  49. [58]

    Ddl-slam: A robust rgb-d slam in dynamic environments combined with deep learning,

    Y . Ai, T. Rui, M. Lu, L. Fu, S. Liu, and S. Wang, “Ddl-slam: A robust rgb-d slam in dynamic environments combined with deep learning,” IEEE Access, vol. 8, pp. 162 335–162 342, 2020

  50. [59]

    D2slam: Semantic visual slam based on the influence of depth for dynamic environments,

    A. Beghdadi, M. Mallem, and L. Beji, “D2slam: Semantic visual slam based on the influence of depth for dynamic environments,” ArXiv, vol. abs/2210.08647, 2022. [Online]. Available: https://api. semanticscholar.org/CorpusID:252917876

  51. [60]

    A semantic SLAM system for dynamic environments,

    F. Hu, Q. Zong, X. Mou, Y . Chen, H. Wang, and M. He, “A semantic SLAM system for dynamic environments,” in International Conference on Automation and Intelligent Technology (ICAIT 2024) , R. Usubamatov, S. Feng, and X. Mei, Eds., vol. 13401, International Society for Optics a...

  52. [61]

    Learning from feedback: Semantic enhancement for object slam using foundation models,

    J. Hong, R. Choi, and J. J. Leonard, “Learning from feedback: Semantic enhancement for object slam using foundation models,”

  53. [62]

    V3d-slam: Robust rgb-d slam in dynamic environments with 3d semantic geometry voting,

    T. Dang and M. Huber, “V3d-slam: Robust rgb-d slam in dynamic environments with 3d semantic geometry voting,” 10 2024

  54. [63]

    3ds-slam: A 3d object detection based semantic slam towards dynamic indoor environments,

    G. S. Krishna, K. Supriya, and S. Baidya, “3ds-slam: A 3d object detection based semantic slam towards dynamic indoor environments,”

  55. [64]

    Blitz-slam: A semantic slam in dynamic environments,

    Y . Fan, Q. Zhang, Y . Tang, S. Liu, and H. Han, “Blitz-slam: A semantic slam in dynamic environments,” Pattern Recognition , vol. 121, p. 108225, 2022. [Online]. Available: https://www.sciencedirect. com/science/article/pii/S0031320321004064

  56. [65]

    By-slam: Dynamic visual slam system based on beblid and semantic information extraction,

    D. Zhu, P. Liu, Q. Qiu, J. Wei, and R. Gong, “By-slam: Dynamic visual slam system based on beblid and semantic information extraction,” Sensors, vol. 24, no. 14, 2024. [Online]. Available: https://www.mdpi.com/1424-8220/24/14/4693

  57. [66]

    Yolo-slam: A semantic slam system towards dynamic environment with geometric constraint,

    W. Wu, L. Guo, H. Gao, Z. You, Y . Liu, and Z. Chen, “Yolo-slam: A semantic slam system towards dynamic environment with geometric constraint,” Neural Computing and Applications , vol. 34, pp. 1–16, 04 2022

  58. [67]

    A dynamic object filtering approach based on object detection and geometric constraint between frames,

    J. Wei, S. Pan, W. Gao, and T. Zhao, “A dynamic object filtering approach based on object detection and geometric constraint between frames,” IET Image Processing, vol. 16, pp. 1636–1647, 2022. [Online]. Available: https://digital-library.theiet.org/doi/abs/10.1049/ipr2.12436

  59. [68]

    Orbslam-atlas: a robust and accurate multi-map system,

    R. Elvira, J. D. Tard ´os, and J. M. M. Montiel, “Orbslam-atlas: a robust and accurate multi-map system,” 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 6253–6259,

  60. [69]

    Solov2: Dynamic and fast instance segmentation,

    X. Wang, R. Zhang, T. Kong, L. Li, and C. Shen, “Solov2: Dynamic and fast instance segmentation,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 17 721– 17 73...

  61. [70]

    Mid-fusion: Octree-based object-level multi-instance dynamic slam,

    B. Xu, W. Li, D. Tzoumanikas, M. Bloesch, A. Davison, and S. Leutenegger, “Mid-fusion: Octree-based object-level multi-instance dynamic slam,” 12 2018

  62. [71]

    PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume,

    D. Sun, X. Yang, M.-Y . Liu, and J. Kautz, “PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume,” 2018

  63. [72]

    Suma++: Efficient lidar-based semantic slam,

    X. Chen, A. M. E. Palazzolo, P. Gigu `ere, J. Behley, and C. Stachniss, “Suma++: Efficient lidar-based semantic slam,” 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 4530–4537, 2019. [Online]. Available: https: //api.semanticscholar.org/C...

  64. [73]

    Efficient surfel-based slam using 3d laser range data in urban environments,

    J. Behley and C. Stachniss, “Efficient surfel-based slam using 3d laser range data in urban environments,” Robotics: Science and Systems XIV,

  65. [74]

    Dynamic-SLAM: Semantic Monocular Visual Localiza- tion and Mapping Based on Deep Learning in Dynamic Environment,

    L. X. et al., “Dynamic-SLAM: Semantic Monocular Visual Localiza- tion and Mapping Based on Deep Learning in Dynamic Environment,” Robot. Auton. Syst. , vol. 117, pp. 1–16, 2019. 17

  66. [75]

    SALSA: Semantic assisted life-long SLAM for indoor environments,

    A. Jhalani, H. Vhavle, and S. Mahajan, “SALSA: Semantic assisted life-long SLAM for indoor environments,” May 2020, 16-833 Robot Localization and Mapping (Spring 2020) Final Project at Carnegie Mellon University

  67. [76]

    Dynaslam ii: Tightly-coupled multi-object tracking and slam,

    B. Bescos, C. Campos, J. D. Tard ´os, and J. Neira, “Dynaslam ii: Tightly-coupled multi-object tracking and slam,” 2020. [Online]. Available: https://arxiv.org/abs/2010.07820

  68. [77]

    An fpga based energy efficient ds-slam accelerator for mobile robots in dynamic environment,

    Y . Wu, L. Luo, S. Yin, M. Yu, F. Qiao, H. Huang, X. Shi, Q. Wei, and X. Liu, “An fpga based energy efficient ds-slam accelerator for mobile robots in dynamic environment,” Applied Sciences, vol. 11, no. 4, 2021. [Online]. Available: https://www.mdpi.com/2076-3417/11/4/1828

  69. [78]

    Dp-slam: A visual slam with moving probability towards dynamic environments,

    A. Li, J. Wang, M. Xu, and Z. Chen, “Dp-slam: A visual slam with moving probability towards dynamic environments,” Information Sciences, vol. 556, pp. 128–142, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0020025520311841

  70. [79]

    Vins-mono: A robust and versatile monocular visual-inertial state estimator,

    T. Qin, P. Li, and S. Shen, “Vins-mono: A robust and versatile monocular visual-inertial state estimator,” IEEE Transactions on Robotics, vol. 34, pp. 1004–1020, 2017. [Online]. Available: https://api.semanticscholar.org/CorpusID:7334757

  71. [80]

    Rgbd-inertial trajectory estimation and mapping for ground robots,

    Z. Shan, R. Li, and S. Schwertfeger, “Rgbd-inertial trajectory estimation and mapping for ground robots,” Sensors, vol. 19, no. 10,

  72. [81]

    Semantic lidar odometry and mapping for mobile robots using rangenet++,

    X. Dong, G. He, P. Fan, F. Zhang, T. Li, J. Zhou, J. Xie, J. Zhang, J. Huang, and W. Shang, “Semantic lidar odometry and mapping for mobile robots using rangenet++,” in 2022 IEEE International Conference on Mechatronics and Automation (ICMA) , 2022, pp. 721– 726

  73. [82]

    Twistslam++: Fusing multiple modalities for accurate dynamic semantic slam,

    M. Gonzalez, E. Marchand, A. Kacete, and J. Royan, “Twistslam++: Fusing multiple modalities for accurate dynamic semantic slam,”

  74. [83]

    Twistslam: Constrained slam in dynamic environment,

    ——, “Twistslam: Constrained slam in dynamic environment,” IEEE Robotics and Automation Letters , vol. 7, pp. 1–1, 05 2022

  75. [84]

    S3lam: Structured scene slam,

    M. Gonzalez, ´E. Marchand, A. Kacete, and J. Royan, “S3lam: Structured scene slam,” 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 6389–6395, 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:237513510

  76. [85]

    3dssd: Point-based 3d single stage object detector,

    Z. Yang, Y . Sun, S. Liu, and J. Jia, “3dssd: Point-based 3d single stage object detector,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2020

  77. [86]

    Available: https://www.mdpi.com/1424-8220/19/10/ 2251

    [Online]. Available: https://www.mdpi.com/1424-8220/19/10/ 2251

  78. [87]

    An online semantic mapping system for extending and enhancing visual slam,

    T. Hempel and A. Al-Hamadi, “An online semantic mapping system for extending and enhancing visual slam,” Engineering Applications of Artificial Intelligence , vol. 111, p. 104830, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S095219762200094X

  79. [88]

    Factor graphs and gtsam: A hands-on introduction,

    F. Dellaert, “Factor graphs and gtsam: A hands-on introduction,”

  80. [89]

    Available: https://arxiv.org/abs/2209.07888

    [Online]. Available: https://arxiv.org/abs/2209.07888

  81. [90]

    So-slam: Semantic object slam with scale proportional and symmetrical texture constraints,

    Z. Liao, Y . Hu, J. Zhang, X. Qi, X. Zhang, and W. Wang, “So-slam: Semantic object slam with scale proportional and symmetrical texture constraints,” IEEE Robotics and Automation Letters , vol. 7, no. 2, pp. 4008–4015, 2022

  82. [91]

    Quadricslam: Dual quadrics from object detections as landmarks in object-oriented slam,

    L. Nicholson, M. Milford, and N. S ¨underhauf, “Quadricslam: Dual quadrics from object detections as landmarks in object-oriented slam,” IEEE Robotics and Automation Letters , vol. 4, no. 1, pp. 1–8, 2019

  83. [92]

    Visual localization and mapping in dynamic and changing environments,

    J. C. V . Soares, V . S. Medeiros, G. F. Abati, M. Becker, G. A. P. Caurin, M. Gattass, and M. A. Meggiolaro, “Visual localization and mapping in dynamic and changing environments,” Journal of Intelligent & Robotic Systems , vol. 109, pp. 1–20, 2022. [Online]. Available: https...

  84. [93]

    Detectron2,

    Y . Wu, A. Kirillov, F. Massa, W.-Y . Lo, and R. Girshick, “Detectron2,” https://github.com/facebookresearch/detectron2, 2019

  85. [94]

    A general optimization- based framework for global pose estimation with multiple sensors,

    T. Qin, S. Cao, J. Pan, and S. Shen, “A general optimization- based framework for global pose estimation with multiple sensors,” ArXiv, vol. abs/1901.03642, 2019. [Online]. Available: https://api. semanticscholar.org/CorpusID:57825739

  86. [95]

    An End-to-End Transformer Model for 3D Object Detection,

    I. Misra, R. Girdhar, and A. Joulin, “An End-to-End Transformer Model for 3D Object Detection,” in ICCV, 2021

  87. [96]

    Ydd-slam: Indoor dynamic visual slam fusing yolov5 with depth information,

    P. Cong, J. Liu, J. Li, Y . Xiao, X. Chen, X. Feng, and X. Zhang, “Ydd-slam: Indoor dynamic visual slam fusing yolov5 with depth information,” Sensors, vol. 23, no. 23, 2023. [Online]. Available: https://www.mdpi.com/1424-8220/23/23/9592

  88. [97]

    Dgs-slam: A fast and robust rgbd slam in dynamic environments combined by geometric and semantic information,

    L. Yan, X. Hu, L. Zhao, Y . Chen, P. Wei, and H. Xie, “Dgs-slam: A fast and robust rgbd slam in dynamic environments combined by geometric and semantic information,” Remote Sensing, vol. 14, no. 3,

  89. [98]

    PVO: Panoptic visual odometry,

    W. Ye, X. Lan, S. Chen, Y . Ming, X. Yu, H. Bao, Z. Cui, and G. Zhang, “PVO: Panoptic visual odometry,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 9579–9589

  90. [99]

    DROID-SLAM: Deep Visual SLAM for Monoc- ular, Stereo, and RGB-D Cameras,

    Z. Teed and J. Deng, “DROID-SLAM: Deep Visual SLAM for Monoc- ular, Stereo, and RGB-D Cameras,” Advances in neural information processing systems, 2021

  91. [100]

    Sd-slam: A semantic slam approach for dynamic scenes based on lidar point clouds,

    F. Li, C. Fu, D. Sun, J. Li, and J. Wang, “Sd-slam: A semantic slam approach for dynamic scenes based on lidar point clouds,” Big Data Research , vol. 36, p. 100463, 2024. [Online]. Available: https: //www.sciencedirect.com/science/article/pii/S221457962400039X

  92. [101]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection,

    S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu et al. , “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” arXiv preprint arXiv:2303.05499, 2023

  93. [102]

    Fusing panoptic segmentation and geometry information for robust visual slam in dynamic environments,

    H. Zhu, C. Yao, Z. Zhu, Z. Liu, and Z. Jia, “Fusing panoptic segmentation and geometry information for robust visual slam in dynamic environments,” in 2022 IEEE 18th International Conference on Automation Science and Engineering (CASE), 2022, pp. 1648–1653

  94. [103]

    Masked-attention mask transformer for universal image segmenta- tion,

    B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmenta- tion,” 2022. [Online]. Available: https://arxiv.org/abs/2112.01527

  95. [104]

    Be-slam: Bev-enhanced dynamic semantic slam with static object reconstruction,

    J. Luo, G. Wang, H. Liu, L. Wu, T. Huang, D. Xiao, H. Pu, and J. Luo, “Be-slam: Bev-enhanced dynamic semantic slam with static object reconstruction,” 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 10 105–10 112, 2024. [Online]. Available...

  96. [105]

    Dsp-slam: Object oriented slam with deep shape priors,

    J. Wang, M. R ¨unz, and L. Agapito, “Dsp-slam: Object oriented slam with deep shape priors,” in 2021 International Conference on 3D Vision (3DV). IEEE, 2021, pp. 1362–1371

  97. [106]

    Improving rgb-d slam accuracy in dynamic environments based on semantic and geometric constraints,

    X. Wang, S. Zheng, X. Lin, and F. Zhu, “Improving rgb-d slam accuracy in dynamic environments based on semantic and geometric constraints,” Measurement, vol. 217, p. 113084, 2023. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S0263224123006486

  98. [107]

    imap: Implicit mapping and positioning in real-time,

    E. Sucar, S. Liu, J. Ortiz, and A. J. Davison, “imap: Implicit mapping and positioning in real-time,” 2021. [Online]. Available: https://arxiv.org/abs/2103.12352

  99. [108]

    Feature-realistic neural fusion for real-time, open set scene understanding,

    K. Mazur, E. Sucar, and A. J. Davison, “Feature-realistic neural fusion for real-time, open set scene understanding,” 2022. [Online]. Available: https://arxiv.org/abs/2210.03043

  100. [109]

    Efficientnet: Rethinking model scaling for convolutional neural networks,

    M. Tan and Q. V . Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” 2020. [Online]. Available: https://arxiv.org/abs/1905.11946

  101. [110]

    Emerging properties in self-supervised vision transformers,

    M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,”

  102. [111]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Doll´ar, and R. Girshick, “Segment anything,” arXiv:2304.02643, 2023

  103. [113]

    vmap: Vectorised object mapping for neural field slam,

    X. Kong, S. Liu, M. Taher, and A. J. Davison, “vmap: Vectorised object mapping for neural field slam,” 2023. [Online]. Available: https://arxiv.org/abs/2302.01838

  104. [114]

    Ro-map: Real-time multi-object mapping with neural radiance fields,

    X. Han, H. Liu, Y . Ding, and L. Yang, “Ro-map: Real-time multi-object mapping with neural radiance fields,” IEEE Robotics and Automation Letters, vol. 8, no. 9, pp. 5950–5957, 2023

  105. [115]

    Ilabel: Interactive neural scene labelling,

    S. Zhi, E. Sucar, A. Mouton, I. Haughton, T. Laidlow, and A. J. Davison, “Ilabel: Interactive neural scene labelling,” 2021. [Online]. Available: https://arxiv.org/abs/2111.14637

  106. [116]

    Dn-slam: A visual slam with orb features and nerf mapping in dynamic environments,

    C. Ruan, Q. Zang, K. Zhang, and K. Huang, “Dn-slam: A visual slam with orb features and nerf mapping in dynamic environments,” IEEE Sensors Journal, vol. 24, no. 4, pp. 5279–5287, 2024

  107. [117]

    Dynamon: Motion- aware fast and robust camera localization for dynamic neural radiance fields,

    N. Schischka, H. Schieber, M. A. Karaoglu, M. Gorgulu, F. Gr ¨otzner, A. Ladikos, N. Navab, D. Roth, and B. Busam, “Dynamon: Motion- aware fast and robust camera localization for dynamic neural radiance fields,” IEEE Robotics and Automation Letters, vol. 10, no. 1, pp. 548– 55...

  108. [118]

    Ddn-slam: Real-time dense dynamic neural implicit slam,

    M. Li, Y . Zhou, G. Jiang, T. Deng, Y . Wang, and H. Wang, “Ddn-slam: Real-time dense dynamic neural implicit slam,” 2024. [Online]. Available: https://arxiv.org/abs/2401.01545

  109. [119]

    Nid-slam: Neural implicit representation-based rgb-d slam in dynamic environments,

    Z. Xu, J. Niu, Q. Li, T. Ren, and C. Chen, “Nid-slam: Neural implicit representation-based rgb-d slam in dynamic environments,”

  110. [120]

    Dvn-slam: Dynamic visual neural slam based on local-global encoding,

    W. Wu, G. Wang, T. Deng, S. Aegidius, S. Shanks, V . Modugno, D. Kanoulas, and H. Wang, “Dvn-slam: Dynamic visual neural slam based on local-global encoding,” 2024. [Online]. Available: https://arxiv.org/abs/2403.11776

  111. [121]

    Neural implicit dense semantic slam,

    Y . Haghighi, S. Kumar, J.-P. Thiran, and L. V . Gool, “Neural implicit dense semantic slam,” 2023. [Online]. Available: https: //arxiv.org/abs/2304.14560

  112. [122]

    Neds-slam: A neural explicit dense semantic slam framework using 3d gaussian splatting,

    Y . Ji, Y . Liu, G. Xie, B. Ma, Z. Xie, and H. Liu, “Neds-slam: A neural explicit dense semantic slam framework using 3d gaussian splatting,” IEEE Robotics and Automation Letters , vol. 9, pp. 8778–8785,

  113. [123]

    Depth anything: Unleashing the power of large-scale unlabeled data,

    L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” 2024. [Online]. Available: https://arxiv.org/abs/2401.10891

  114. [124]

    Hi-slam: Scaling- up semantics in slam with a hierarchically categorical gaussian splatting,

    B. Li, Z. Cai, Y .-F. Li, I. Reid, and H. Rezatofighi, “Hi-slam: Scaling- up semantics in slam with a hierarchically categorical gaussian splatting,” 2024. [Online]. Available: https://arxiv.org/abs/2409.12518

  115. [125]

    Nis-slam: Neural implicit semantic rgb-d slam for 3d consistent scene understanding,

    H. Zhai, G. Huang, Q. Hu, G. Li, H. Bao, and G. Zhang, “Nis-slam: Neural implicit semantic rgb-d slam for 3d consistent scene understanding,” 2024. [Online]. Available: https://arxiv.org/abs/ 2407.20853

  116. [126]

    Splatam: Splat, track & map 3d gaussians for dense rgb-d slam,

    N. Keetha, J. Karhade, K. M. Jatavallabhula, G. Yang, S. Scherer, D. Ramanan, and J. Luiten, “Splatam: Splat, track & map 3d gaussians for dense rgb-d slam,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024

  117. [127]

    Nerf: representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: representing scenes as neural radiance fields for view synthesis,” Commun. ACM , vol. 65, no. 1, p. 99–106, Dec. 2021. [Online]. Available: https://doi.org/10.1145/3503250

  118. [128]

    Dns slam: Dense neural semantic-informed slam,

    K. Li, M. Niemeyer, N. Navab, and F. Tombari, “Dns slam: Dense neural semantic-informed slam,” ArXiv, vol. abs/2312.00204,

  119. [129]

    Ro-map: Real-time multi-object mapping with neural radiance fields,

    X. Han, H. Liu, Y . Ding, and L. Yang, “Ro-map: Real-time multi-object mapping with neural radiance fields,” IEEE Robotics and Automation Letters, vol. 8, no. 9, p. 5950–5957, Sep. 2023. [Online]. Available: http://dx.doi.org/10.1109/LRA.2023.3302176

  120. [130]

    Available: https://arxiv.org/abs/2401.01189

    [Online]. Available: https://arxiv.org/abs/2401.01189

  121. [131]

    Stereo visual inertial odometry for robots with limited computational resources,

    S. Bahnam, S. Pfeiffer, and G. C. de Croon, “Stereo visual inertial odometry for robots with limited computational resources,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021, pp. 9154–9159

  122. [132]

    Semgauss-slam: Dense semantic gaussian splatting slam,

    S. Zhu, R. Qin, G. Wang, J. Liu, and H. Wang, “Semgauss-slam: Dense semantic gaussian splatting slam,” Mar 2024. [Online]. Available: http://arxiv.org/abs/2403.07494v3

  123. [133]

    GitHub - HiIAmTzeKean/Jetson-Nano-SLAM: A Nanyang Technological University Research Project. Multi-USB- Cam Implementation on Jetson Nano,

    HiIAmTzeKean, “GitHub - HiIAmTzeKean/Jetson-Nano-SLAM: A Nanyang Technological University Research Project. Multi-USB- Cam Implementation on Jetson Nano,” 2022, accessed Oct. 08, 2024. [Online]. Available: https://github.com/HiIAmTzeKean/ Jetson-Nano-SLAM

  124. [134]

    Available: https://api.semanticscholar.org/CorpusID: 268532088

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 268532088

  125. [135]

    Hw/sw codesign and fpga acceleration of visual odometry algorithms for rover navigation on mars,

    G. Lentaris, I. Stamoulias, D. Soudris, and M. Lourakis, “Hw/sw codesign and fpga acceleration of visual odometry algorithms for rover navigation on mars,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 26, no. 8, pp. 1563–1577, 2016

  126. [136]

    Fpga design of ekf block accelerator for 3d visual slam,

    D. T ¨ortei Tertei, J. Piat, and M. Devy, “Fpga design of ekf block accelerator for 3d visual slam,” Computers & Electrical Engineering, vol. 55, pp. 123–137, 2016. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0045790616301045

  127. [137]

    Panoslam: Panoptic 3d scene reconstruction via gaussian slam,

    C. Runnan, Z. Wang, J. Wang, B. Ma, M. Gong, W. Wang, and T. Liu, “Panoslam: Panoptic 3d scene reconstruction via gaussian slam,” 12 2024

  128. [138]

    Navion: A 2-mw fully integrated real-time visual-inertial odometry accelerator for autonomous navigation of nano drones,

    A. Suleiman, Z. Zhang, L. Carlone, S. Karaman, and V . Sze, “Navion: A 2-mw fully integrated real-time visual-inertial odometry accelerator for autonomous navigation of nano drones,” IEEE Journal of Solid- State Circuits, vol. 54, no. 4, pp. 1106–1119, 2019

  129. [139]

    An 879gops 243mw 80fps vga fully visual cnn-slam processor for wide-range autonomous exploration,

    Z. Li, Y . Chen, L. Gong, L. Liu, D. Sylvester, D. Blaauw, and H.-S. Kim, “An 879gops 243mw 80fps vga fully visual cnn-slam processor for wide-range autonomous exploration,” in 2019 IEEE International Solid-State Circuits Conference - (ISSCC) , 2019, pp. 134–136

  130. [140]

    Low latency visual inertial odom- etry with on-sensor accelerated optical flow for resource-constrained uavs,

    J. K ¨uhne, M. Magno, and L. Benini, “Low latency visual inertial odom- etry with on-sensor accelerated optical flow for resource-constrained uavs,” 06 2024

  131. [141]

    Available: https://api.semanticscholar.org/CorpusID: 265551655

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 265551655

  132. [142]

    An energy-efficient processor for real-time semantic lidar slam in mobile robots,

    J. Jung, S. Kim, B. Seo, W. Jang, S. Lee, J. Shin, D. Han, and K. Lee, “An energy-efficient processor for real-time semantic lidar slam in mobile robots,” IEEE Journal of Solid-State Circuits, vol. PP, pp. 1–13, 01 2024

  133. [143]

    3d gaussian splatting for real-time radiance field rendering,

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,” ACM Transactions on Graphics , vol. 42, no. 4, July 2023. [Online]. Available: https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/

  134. [144]

    The replica dataset: A digital replica of indoor spaces,

    J. Straub, T. Whelan, L. Ma, Y . Chen, E. Wijmans, S. Green, J. J. Engel, R. Mur-Artal, C. Ren, S. Verma, A. Clarkson, M. Yan, B. Budge, Y . Yan, X. Pan, J. Yon, Y . Zou, K. Leon, N. Carter, J. Briales, T. Gillingham, E. Mueggler, L. Pesqueira, M. Savva, D. Batra, H. M. Strasd...

  135. [145]

    Arm-vo: an efficient monocular visual odometry for ground vehicles on arm cpus,

    Z. Z. Nejad and A. H. Ahmadabadian, “Arm-vo: an efficient monocular visual odometry for ground vehicles on arm cpus,” Machine Vision and Applications, vol. 30, no. 6, pp. 1061–1070, 2019

  136. [147]

    SLAM Performance on Embedded Robots Undergraduate Student Research: Individual Project,

    N. Ghalehshahi, G. Tech, R. Hadidi, and H. Kim, “SLAM Performance on Embedded Robots Undergraduate Student Research: Individual Project,” 2024, accessed Oct. 08, 2024. [Online]. Available: https://ramyadhadidi.github.io/files/shoghi src esweek.pdf

  137. [150]

    Visual- inertial odometry on chip: An algorithm-and-hardware co-design ap- proach,

    Z. Zhang, A. Suleiman, L. Carlone, V . Sze, and S. Karaman, “Visual- inertial odometry on chip: An algorithm-and-hardware co-design ap- proach,” 07 2017

  138. [154]

    A low-power and real-time semantic lidar slam processor with point neural network segmentation and knn acceleration for mobile robots,

    J. Jung, S. Kim, B. Seo, W. Jang, S. Lee, J. Shin, D. Han, and K. J. Lee, “A low-power and real-time semantic lidar slam processor with point neural network segmentation and knn acceleration for mobile robots,” in 2024 IEEE Symposium in Low-Power and High-Speed Chips (COOL CHI...

  139. [156]

    A benchmark for the evaluation of rgb-d slam systems,

    J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers, “A benchmark for the evaluation of rgb-d slam systems,” in Proc. of the International Conference on Intelligent Robot Systems (IROS) , Oct. 2012

  140. [2009]

    Since 2022, he is a full Professor and deputy director of the U2IS (Robotics and AI) Lab at ENSTA Paris - Institut Polytechnique de Paris

    He worked for a few years as a research engineer at the French electrical company EDF and then as an Assistant, Associate, and Full Professor at Mines Paris - PSL University. Since 2022, he is a full Professor and deputy director of the U2IS (Robotics and AI) Lab at ENSTA Pari...

  141. [2012]

    Available: https://api.semanticscholar.org/CorpusID: 131215724

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 131215724

  142. [2018]

    Available: https://api.semanticscholar.org/CorpusID: 46954808

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 46954808

  143. [2019]

    Available: https://api.semanticscholar.org/CorpusID: 201698330

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 201698330

  144. [2021]

    Available: https://arxiv.org/abs/2104.14294

    [Online]. Available: https://arxiv.org/abs/2104.14294

  145. [2022]

    Available: https://www.mdpi.com/2072-4292/14/3/795

    [Online]. Available: https://www.mdpi.com/2072-4292/14/3/795

  146. [2023]

    Available: https://arxiv.org/abs/2310.06385

    [Online]. Available: https://arxiv.org/abs/2310.06385

  147. [2024]

    Available: https://arxiv.org/abs/2411.06752

    [Online]. Available: https://arxiv.org/abs/2411.06752

  148. [6401]

    Available: https://doi.ieeecomputersociety.org/10.1109/ CVPR.2019.00656

    [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/ CVPR.2019.00656

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.