Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Demonstrating CavePI: Autonomous Exploration of Underwater Caves by Semantic Guidance

T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper claims that a low-cost, one-person-portable AUV can navigate underwater caves by visually following the diver's guide line.

desk verdict CavePI is a credible, openly documented low-cost AUV platform for caveline following in clear water, but its cave-exploration claims rest on a perception system that failed exactly where it matters. read the letter →

arxiv 2502.05384 v4 pith:4LSETSJW submitted 2025-02-07 cs.RO

classification cs.RO
keywords autonomousunderwatervehiclecaveexplorationvisualservoingsemanticsegmentationcavelinedetectionpurepursuitcontroledgeAIdigitaltwin
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Underwater cave exploration can be delegated to a small, inexpensive AUV that follows the same guide line human divers use, rather than relying on expensive vehicles or full 3D mapping. The paper demonstrates CavePI, an 8.6 kg autonomous underwater vehicle that runs a lightweight real-time semantic segmentation model on an edge GPU to pick out the caveline and a PID pure-pursuit controller to steer along it. In controlled tank tests the vehicle keeps the line within roughly 13 cm after tuning, and in open spring-water trials it holds depth within ±10 cm despite currents. The paper also introduces a ROS/Gazebo digital twin for pre-mission validation. If these results hold, cave survey, hydrology, and archaeology teams gain a one-person-portable tool for missions too hazardous for divers.

What carries the argument

The central mechanism is the caveline itself: the guide line divers string through caves, which reduces a 3D cave passage to a 1D retraction. CavePI's down-facing camera feeds a MobileNetV3-DeepLabV3 segmentation network, a lightweight convolutional encoder with an atrous-convolution decoder, that labels each pixel as caveline or background at 18.2 FPS on a Jetson Nano. Post-processing extracts the caveline contours, and the controller steers toward the centroid of the farthest contour, using a PID-tuned Pure Pursuit law. This farthest-contour targeting, combined with the heading error signal, is what makes tracking-by-detection work without GPS or external localization.

What would settle it

A field trial in a turbid natural cave at night with a thin caveline, logging how long CavePI tracks without diver repositioning and measuring segmentation recall on held-out frames, would settle the claim: if the model misses or mislabels the line for most of the dive, or tracking ends within about a minute, the stated cave-exploration capability is not supported.

Watch

Extended reading notes

Core claim

The paper's claim is that underwater cave exploration can be driven by semantic guidance: instead of trying to map or localize in feature-deprived, GPS-denied water, CavePI detects the caveline, the guide line divers run from the cave entrance through the main passages, and treats it as the navigation reference. Onboard, a MobileNetV3-DeepLabV3 segmentation model classifies each pixel of the down-facing camera stream as caveline or background, running at 18.2 frames per second on a Jetson Nano; contours are then fed to a PID-tuned Pure Pursuit controller that steers the vehicle toward the farthest detected line segment. In a laboratory tank, after gain tuning, the mean tracking error was approximately 13 cm and depth error about 2 cm; in open spring-water trials the vehicle held depth within ±10 cm, with currents causing lateral drift and occasional loss of line. In natural caves at night the authors report that the lightweight model struggled to detect the caveline and occasionally confused tree roots for it, which they identify as a limitation to be addressed with a more powerful onboard computer. They conclude that these integrated design choices facilitate reliable AUV navigation under feature-deprived, GPS-denied, and low-visibility conditions with overhead obstacles.

Load-bearing premise

The whole demonstration depends on the onboard camera and segmentation model being able to see the caveline in the target cave's actual low-light, murky conditions; in the nighttime cave trials reported here, that assumption repeatedly failed.

Editorial extensions

If this is right

  • If the demonstrated tracking performance transfers to real missions, teams can survey karst aquifers, archaeological sites, and cave ecosystems with a robot that one person can carry, at a fraction of the cost of existing exploration AUVs.
  • The same semantic-guidance loop, segmentation of a 1D marker plus pure pursuit, can be retargeted to other indoor or overhead structures where guides are available, such as ship hulls, pipelines, and dams.
  • The digital twin gives a low-risk testbed for changing control gains or simulating sharp turns and dead ends before committing to a field dive.
  • Because the design, code, and data are released, other groups can reproduce the pipeline and extend it to new cave systems without rebuilding the vehicle.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's claims, the failure pattern suggests the bottleneck is perception, not control: in nighttime caves the controller lost the line only because the segmentation model could not see it, so improving low-light training data or adding temporal filtering may yield large gains without changing the robot.
  • The farthest-contour heuristic presumes the caveline is the most salient thin structure ahead of the robot; in root- and moss-filled caves, a temporal consistency check or geometric prior on line continuity would likely reject false positives before they reach the controller.
  • A direct test of the broader overhead-structure claim would be to run the same stack on a painted stripe inside a ship hull or pipeline, since the design already has the down-facing camera and depth control needed for such tasks.
  • The reported tracking-error standard deviation is comparable to the mean, so planning a cave mission on the mean alone would be risky; using the full error distribution, or adding sway control, is a natural next step.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents CavePI, a low-cost (~$3,500, 8.6 kg) AUV designed for semantic-guided navigation in underwater caves. The platform combines a downward camera, front camera, Ping2 sonar, Jetson Nano and Raspberry Pi-5 with a ROS2 backbone, and a 4-thruster layout. Navigation uses a MobileNetV3-DeepLabV3 semantic segmentation model fine-tuned on the authors' CL-ViT dataset to detect a diver-installed caveline, whose farthest contour centroid serves as a waypoint for a PID-controlled pure-pursuit heading controller; a depth PID and a Gazebo digital twin support the system. Evaluation includes FEA of the dome connector, controlled tank line-following with PID gain grid search (best mean tracking error 13.1 cm), a representative open-water spring trial (Fig. 14), nighttime cave trials (Sec. 6.2), and simulation. The paper claims reliable navigation under feature-deprived, GPS-denied, low-visibility conditions with overhead obstacles, and explicitly documents perception and control failure modes.

Significance. The paper's main strength is an open, reproducible system integration: hardware design, ROS code, and data are released, and the authors report both successes and failure cases (tank-edge false positives, root misidentification, current-induced drift, pitch-up, overshoot). If the claims are scoped appropriately, the platform is a useful low-cost testbed for semantic caveline following and for studying the perception-control trade-offs of edge-AI AUVs. The cave-deployment evidence, however, does not support the headline claim of reliable autonomous cave exploration: those trials used an earlier three-thruster vehicle, the segmentation model frequently failed, and tracking was maintained for at most one minute after manual repositioning. The paper should be judged as a system demonstration with partial field validation rather than as a validated autonomous cave-exploration result.

major comments (4)
  1. [Abstract and §6.2] The central claim of "reliable AUV navigation under feature-deprived, GPS-denied, and low-visibility conditions with overhead obstacles" is not supported by the cave-trial evidence. The nighttime cave trials used an earlier three-thruster configuration rather than the four-thruster CavePI described in §3, and the lightweight segmentation model "struggled to detect the caveline from camera images," misidentified submerged roots and moss as caveline (Fig. 16b), and maintained tracking only "up to one minute" after divers manually repositioned the vehicle. No detection-rate, false-positive-rate, mission-completion, or tracking-error statistics are reported for the cave environment. Although this limitation is openly acknowledged in §7.2, it is the load-bearing part of the headline claim, so the abstract and conclusion should be revised to state precisely what was demonstrated and on which platform configuration.
  2. [§5.2, Eqs. (3)–(8), Table 3] The reported tracking error δ is computed relative to the segmentation mask of the detected caveline, not an independent ground-truth line. The 13.1 cm mean tracking error therefore measures how well the controller follows the detector's output, not absolute line-following accuracy; if the detector systematically mislocates or fragments the caveline, δ can be small while the true offset is large. The paper should either provide ground-truth validation (e.g., manually annotated frames or an external localisation system) or explicitly qualify δ as a "segmentation-relative tracking error."
  3. [§4.1, Table 2] The segmentation model is fine-tuned on the authors' CL-ViT dataset (3,150 images) plus 150 lab images and evaluated on the authors' CL-Challenge benchmark; there is no independent or in-situ quantitative evaluation of the MobileNetV3-DeepLabV3 model in the low-light cave/grotto conditions where it is deployed. The reported mIoU of 48.95% is also below the 58.3% baseline from the same group's prior work, so the perception-generalisation claim rests largely on anecdotal failure observations. Please add per-environment segmentation metrics or clearly scope the claim to the environments where the detector was quantitatively assessed.
  4. [§6.1 and §6.2] The open-water evaluation is summarized with one representative 10-minute trial (Fig. 14), while the paper states that 15 open-water trials were conducted; without aggregate statistics (mean/median error, success rate, number of tracking losses across trials), the robustness claims for open-water operation are not quantitatively supported. This is less severe than the cave-trial issue but is still needed to support the claimed "long-term autonomous missions" in complex underwater environments.
minor comments (5)
  1. [§5.1] The text contains "V on Mises" and "V on-Mises" where "von Mises" is intended; please correct the spelling.
  2. [§4.2 and Algorithm 1] There is a typo "betweein" in the depth-control sentence, and Algorithm 1's line "Rotate 360°" is ambiguous because the surrounding text describes a circular search pattern rather than a single 360-degree rotation; please reword for clarity.
  3. [§5.2, Eqs. (3)–(7)] The notation I P is not formally defined, and the depth scale factor λ is first used in Eq. (3) but only defined later in Eq. (7); please introduce both symbols before first use.
  4. [Introduction] The phrase "provided in the the supplementary video" contains a duplicated article; please remove the duplicate.
  5. [§5.2] Figure 9(a) shows depth-control accuracy but the text does not report the corresponding mean depth error for that experiment; please add the value or refer to it explicitly when describing the depth controller's performance.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's central claims rest on direct experiments and openly acknowledged limitations, and the few self-citations are fixed benchmarks rather than load-bearing justifications.

full rationale

The paper's central contributions are hardware/software integration and empirical evaluation, not a formal derivation chain, so there is no prediction or first-principles result that could reduce to its inputs by construction. The perception model is fine-tuned on the authors' CL-ViT dataset and compared on the fixed CL-Challenge benchmark; although these resources come from the same research group, they are public, fixed datasets that any method can be evaluated against, and the benchmark comparison is peripheral to the main navigation claims. The PID gains are tuned by an openly reported grid search, and the resulting 13.1 cm tracking error is an in-sample optimum rather than a held-out prediction; this is a statistical-tuning concern, not circularity. Section 5.2's tracking-error computation uses the segmentation mask rather than an independent ground-truth caveline, so the reported error measures control consistency with the system's own perception; this is an experimental-validity limitation, not a circular derivation. The field results in Section 6.2 honestly report that the nighttime cave trials used an earlier three-thruster iteration and that the lightweight model 'struggled to detect the caveline from camera images,' and Section 7.2 concedes root misidentification and low-light failures; these statements weaken the abstract's 'reliable navigation under low-visibility conditions' claim, but acknowledging limitations is the opposite of circularity. No equation is shown to be equivalent to its own inputs, and no load-bearing argument depends on a self-citation chain. The correct finding is no significant circularity (score 0).

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim depends on the PID gains being reasonable for the platform, on the segmentation model transferring to new sites, and on the caveline being a reliable marker present in the cave. No new physical entities are introduced. The free parameters are the PID gains, tuned on the same test loop used for evaluation.

free parameters (2)
  • Heading PID gains (Kp, Kd) = Kp=3.4, Kd=0.9
    Selected by grid search on the 6-meter tank loop (Table 3) to minimize mean tracking error; used in all subsequent experiments.
  • Depth PID gains (Kp, Kd) = Kp=600, Kd=50
    Selected by grid search on the same tank loop (Table 4); the tracking and depth results are thus tuned on the evaluation setup.
assumptions (4)
  • domain assumption The caveline is a persistent, detectable marker that exists throughout the explored sections of underwater caves.
    The entire guidance paradigm depends on the presence of a physical line. This is the standard cave-diving convention [14], so the assumption is reasonable for the target domain.
  • domain assumption The fine-tuned MobileNetV3-DeepLabV3 segmentation model generalizes from the CL-ViT training set to unseen cave and spring environments.
    This is the load-bearing perception assumption; it is invoked in Section 4.1 and shown to fail in the low-light cave trials (Section 6.2).
  • domain assumption Pure pursuit with the farthest-detected-contour waypoint rule yields stable tracking for this 4-DOF AUV at the speeds used in the experiments.
    The controller in Algorithm 1 assumes the image center maps directly to the robot's heading and that the contour centroids provide a useful pursuit point; this is an engineering approximation that holds in the tested conditions but is not formally proven.
  • standard math Equations (3)-(8) rely on standard rigid-body transforms and pinhole camera projection.
    These are standard background results for visual odometry and coordinate frame transforms.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Demonstrating CavePI: Autonomous Exploration of Underwater Caves by Semantic Guidance." pith.science (2026). https://pith.science/paper/4LSETSJW

@misc{pith2026250205384,
  author       = {Pith},
  title        = {Pith review of: Demonstrating CavePI: Autonomous Exploration of Underwater Caves by Semantic Guidance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4LSETSJW}},
  note         = {Machine review of arXiv:2502.05384}
}
read the original abstract

Enabling autonomous robots to safely and efficiently navigate, explore, and map underwater caves is of significant importance to water resource management, hydrogeology, archaeology, and marine robotics. In this work, we demonstrate the system design and algorithmic integration of a visual servoing framework for semantically guided autonomous underwater cave exploration. We present the hardware and edge-AI design considerations to deploy this framework on a novel AUV (Autonomous Underwater Vehicle) named CavePI. The guided navigation is driven by a computationally light yet robust deep visual perception module, delivering a rich semantic understanding of the environment. Subsequently, a robust control mechanism enables CavePI to track the semantic guides and navigate within complex cave structures. We evaluate the system through field experiments in natural underwater caves and spring-water sites and further validate its ROS (Robot Operating System)-based digital twin in a simulation environment. Our results highlight how these integrated design choices facilitate reliable navigation under feature-deprived, GPS-denied, and low-visibility conditions.

Figures

Figures reproduced from arXiv: 2502.05384 by the authors.

Figure 1
Figure 1. The CavePI AUV navigates by leveraging the semantic [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The proposed CavePI system design is shown; (a) isometric 3D view of the robot; (b) side-view and top-view displaying [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 4
Figure 4. Data flow among major computational modules of [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (13 more)
Figure 5
Figure 5. Figure 5: Simplified model architecture for caveline segmenta [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: A few visual servoing test cases are shown; (a) AUV [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: (a) FEA mesh of the dome connector; (b) its total defor [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Subsequently, we add and vertical slopes in the tank [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 8
Figure 8. Figure 8: The laboratory setup used for tracking accuracy evaluation is shown. CavePI detects and follows the line laid on the tank floor, [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 10
Figure 10. Figure 10: Configurations of the coordinate frames are shown at: [PITH_FULL_IMAGE:figures/full_fig_p009_10.png]
Figure 9
Figure 9. Figure 9: Line-following and depth-holding accuracy are reported [PITH_FULL_IMAGE:figures/full_fig_p009_9.png]
Figure 11
Figure 11. Figure 11: The digital twin (DT) of CavePI, modeled in ROS, is [PITH_FULL_IMAGE:figures/full_fig_p010_11.png]
Figure 12
Figure 12. Figure 12: Interactive demonstration setup for CavePI’s virtual [PITH_FULL_IMAGE:figures/full_fig_p010_12.png]
Figure 14
Figure 14. Figure 14: Line-following and depth-holding accuracy are re [PITH_FULL_IMAGE:figures/full_fig_p011_14.png]
Figure 15
Figure 15. Figure 15: A few snapshots from our field trials for caveline tracking and following experiments with CavePI are shown. The setups [PITH_FULL_IMAGE:figures/full_fig_p012_15.png]
Figure 16
Figure 16. Figure 16: A few perception failure modes are shown. The down [PITH_FULL_IMAGE:figures/full_fig_p013_16.png]
Figure 17
Figure 17. Figure 17: A few tracking failure modes are shown: (a) CavePI [PITH_FULL_IMAGE:figures/full_fig_p013_17.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic Communication for the Internet of Underwater Things: Architectures, Applications, Challenges, and Future Directions

    eess.SP 2026-01 reject novelty 2.0 of 10

    A survey of semantic communication for underwater IoT that compiles architectures, applications, and future directions, but contains internally inconsistent performance claims and many non-archival citations.

Reference graph

Works this paper leans on

68 extracted references · 64 canonical work pages · cited by 1 Pith paper

  1. [1]

    CaveSeg: Deep Semantic Segmentation and Scene Parsing for Autonomous Underwater Cave Exploration

    Adnan Abdullah, Titon Barua, Reagan Tibbetts, Zijie Chen, Md Jahidul Islam, and Ioannis Rekleitis. CaveSeg: Deep Semantic Segmentation and Scene Parsing for Autonomous Underwater Cave Exploration. In IEEE International Con- ference on Robotics and Automation (ICRA). IEEE, 2024

  2. [2]

    Vi- sual Navigation Based on Deep Semantic Cues for Real- Time Autonomous Power Line Inspection

    Dimitrios Alexiou, Georgios Zampokas, Evangelos Skarta- dos, Kosmas Tsiakas, Ioannis Kostavelis, Dimitrios Giak- oumis, Antonios Gasteratos, and Dimitrios Tzovaras. Vi- sual Navigation Based on Deep Semantic Cues for Real- Time Autonomous Power Line Inspection. In International Conference on Unmanned Aircraft Systems (ICUAS) , pages 1262–1269, 2023

  3. [3]

    Finite Element Method-Based Kinematics and Closed-Loop Control of Soft, Continuum Manipulators

    Thor Morales Bieze, Frederick Largilliere, Alexandre Kruszewski, Zhongkai Zhang, Rochdi Merzouki, and Chris- tian Duriez. Finite Element Method-Based Kinematics and Closed-Loop Control of Soft, Continuum Manipulators. Soft robotics, 5(3):348–364, 2018

  4. [4]

    Image Classification System Based on Deep Learning Applied to The Recognition of Traffic Signs for Intelligent Robotic Ve- hicle Navigation Purposes

    Diego Renan Bruno and Fernando Santos Osorio. Image Classification System Based on Deep Learning Applied to The Recognition of Traffic Signs for Intelligent Robotic Ve- hicle Navigation Purposes. In Latin American Robotics Sym- posium (LARS) and Brazilian Symposium on Robotics (SBR), 2017

  5. [5]

    American Cave Diving Fatalities 1969-2007

    Peter L Buzzacott, Erin Zeigler, Petar Denoble, and Richard Vann. American Cave Diving Fatalities 1969-2007. Inter- national Journal of Aquatic Research and Education, 3(2):7, 2009

  6. [6]

    Sonar-Based Guidance of Unmanned Underwater Vehi- cles

    Massimo Caccia, Gabriele Bruzzone, and Gianmarco Verug- gio. Sonar-Based Guidance of Unmanned Underwater Vehi- cles. Advanced robotics, 15(5):551–573, 2001

  7. [7]

    On- board Visual-Based Navigation System for Power Line Fol- lowing with UA V.Int

    Alexander Cer ´on, Iv ´an Mondrag ´on, and Flavio Prieto. On- board Visual-Based Navigation System for Power Line Fol- lowing with UA V.Int. Journal of Advanced Robotic Systems, 15(2), 2018

  8. [8]

    Rethinking Atrous Convolution for Semantic Image Segmentation

    Liang-Chieh Chen. Rethinking Atrous Convolution for Semantic Image Segmentation. arXiv preprint arXiv:1706.05587, 2017

Show all 68 references
  1. [9]

    DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs

    Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. IEEE Transac- tions on Pattern Analysis and Machine Intelligence , 40(4): 83...

  2. [10]

    A Lane Detection Method Based on Semantic Segmentation

    Ling Ding, Huyin Zhang, Jinsheng Xiao, Cheng Shu, and Shejie Lu. A Lane Detection Method Based on Semantic Segmentation. Computer Modeling in Engineering & Sci- ences, 122(3):1039–1053, 2020

  3. [11]

    Introduction of INR18650-30Q

    Samsung Energy Business Division. Introduction of INR18650-30Q . https://bluerobotics.com/wp- content / uploads / 2018 / 10 / INR18650 - 30Q - Data- Sheet.pdf?x70095 , 2014. Accessed: 04-18- 2025

  4. [12]

    CLIP- Nav: Using CLIP for Zero-Shot Vision-and-Language Navi- gation

    Vishnu Sashank Dorbala, Gunnar Sigurdsson, Robinson Pi- ramuthu, Jesse Thomason, and Gaurav S Sukhatme. CLIP- Nav: Using CLIP for Zero-Shot Vision-and-Language Navi- gation. arXiv preprint arXiv:2211.16649, 2022

  5. [13]

    VTNet: Visual Transformer Network for Object Goal Navigation

    Heming Du, Xin Yu, and Liang Zheng. VTNet: Visual Transformer Network for Object Goal Navigation. arXiv preprint arXiv:2105.09447, 2021

  6. [14]

    Basic Cave Diving: A Blueprint for Survival

    Sheck Exley. Basic Cave Diving: A Blueprint for Survival . Cave Diving Section of the National Speleological Society, 1986

  7. [15]

    Introduction to Karst

    Derek Ford and Paul Williams. Introduction to Karst. John Wiley & Sons, Ltd, 2007

  8. [16]

    Horizon Line Detection in Marine Images: Which Method to Choose? International Journal on Advances in Intelligent Systems, 6(1), 2013

    Evgeny Gershikov, Tzvika Libe, and Samuel Kosolapov. Horizon Line Detection in Marine Images: Which Method to Choose? International Journal on Advances in Intelligent Systems, 6(1), 2013

  9. [17]

    Curee: A curious underwater robot for ecosystem exploration

    Yogesh Girdhar, Nathan McGuire, Levi Cai, Stew- art Jamieson, Seth McCammon, Brian Claus, John E San Soucie, Jessica E Todd, and T Aran Mooney. Curee: A curious underwater robot for ecosystem exploration. In2023 IEEE International Conference on Robotics and Automation (ICRA), ...

  10. [18]

    Deep Residual Learning for Image Recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016

  11. [19]

    Semantic SLAM for an AUV Using Object Recognition from Point Clouds

    Khadidja Himri, Pere Ridao, Nuno Gracias, Albert Palomer, Narc´ıs Palomeras, and Roger Pi. Semantic SLAM for an AUV Using Object Recognition from Point Clouds. IFAC- PapersOnLine, 51(29):360–365, 2018

  12. [20]

    Searching for Mo- bileNetV3

    Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al. Searching for Mo- bileNetV3. In IEEE/CVF International Conference on Com- puter Vision (ICCV), pages 1314–1324, 2019

  13. [21]

    Visual Language Maps for Robot Navigation

    Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard. Visual Language Maps for Robot Navigation. In IEEE International Conference on Robotics and Automation (ICRA), pages 10608–10615, 2023. 15

  14. [22]

    Fast Under- water Image Enhancement for Improved Visual Perception

    Md Jahidul Islam, Youya Xia, and Junaed Sattar. Fast Under- water Image Enhancement for Improved Visual Perception. IEEE Robotics and Automation Letters (RA-L) , 5(2):3227– 3234, 2020

  15. [23]

    Computer Vision Applications in Un- derwater Robotics and Oceanography

    Md Jahidul Islam, Alberto Quattrini Li, Yogesh A Girdhar, and Ioannis Rekleitis. Computer Vision Applications in Un- derwater Robotics and Oceanography. Computer Vision: Challenges, Trends, and Opportunities, 2024

  16. [24]

    Experimental Comparison of Open Source Visual-Inertial-Based State Estimation Al- gorithms in the Underwater Domain

    Bharat Joshi, Sharmin Rahman, Michail Kalaitzakis, Bren- nan Cain, James Johnson, Marios Xanthidis, Nare Kara- petyan, Alan Hernandez, Alberto Quattrini Li, Nikolaos Vitzilaios, and Ioannis Rekleitis. Experimental Comparison of Open Source Visual-Inertial-Based State Estimatio...

  17. [25]

    Reliability Based Analysis and Design of Anchor Retrofitted Concrete Gravity Dams

    Masoud R Kazemi. Reliability Based Analysis and Design of Anchor Retrofitted Concrete Gravity Dams. In World Con- ference on Earthquake Engineering, 2004

  18. [26]

    The Biological and Archaeological Significance of Coastal Caves and Karst Fea- tures

    Michael J Lace and John E Mylroie. The Biological and Archaeological Significance of Coastal Caves and Karst Fea- tures. In Coastal Karst Landforms, pages 111–126. Springer, 2013

  19. [27]

    Outdoor Place Recog- nition in Urban Environments Using Straight Lines

    Jin Han Lee, Sehyung Lee, Guoxuan Zhang, Jongwoo Lim, Wan Kyun Chung, and Il Hong Suh. Outdoor Place Recog- nition in Urban Environments Using Straight Lines. In IEEE International Conference on Robotics and Automation (ICRA), pages 5550–5557, 2014

  20. [28]

    Soft, Flexible Pressure Sensors for Pressure Monitoring Under Large Hydrostatic Pressure and Harsh Ocean Environments

    Yi Li, Andres Villada, Shao-Hao Lu, He Sun, Jianliang Xiao, and Xueju Wang. Soft, Flexible Pressure Sensors for Pressure Monitoring Under Large Hydrostatic Pressure and Harsh Ocean Environments. Soft Matter, 19(30):5772–5780, 2023

  21. [29]

    Numerical Investigation of Dimension- less Parameters in Carangiform Fish Swimming Hydrody- namics

    Marianela Machuca Mac ´ıas, Jos ´e Hermenegildo Garc ´ıa- Ortiz, Taygoara Felamingo Oliveira, and Antonio Cesar Pinho Brasil Junior. Numerical Investigation of Dimension- less Parameters in Carangiform Fish Swimming Hydrody- namics. Biomimetics, 9(1), 2024

  22. [30]

    Toward Autonomous Exploration in Confined Underwater Environments

    Angelos Mallios, Pere Ridao, David Ribas, Marc Carreras, and Richard Camilli. Toward Autonomous Exploration in Confined Underwater Environments. Journal of Field Robotics, 33(7):994–1012, 2016

  23. [31]

    Vision-Based Goal-Conditioned Policies for Underwater Navigation in the Presence of Obsta- cles, 2020

    Travis Manderson, Juan Camilo Gamboa Higuera, Stefan Wapnick, Jean-Franc ¸ois Tremblay, Florian Shkurti, David Meger, and Gregory Dudek. Vision-Based Goal-Conditioned Policies for Underwater Navigation in the Presence of Obsta- cles, 2020

  24. [32]

    Ux 1 system design-a robotic system for underwater mining ex- ploration

    Alfredo Martins, Jos ´e Almeida, Carlos Almeida, Andr ´e Dias, Nuno Dias, Jussi Aaltonen, Arttu Heininen, Kari T Koskinen, Claudio Rossi, Sergio Dominguez, et al. Ux 1 system design-a robotic system for underwater mining ex- ploration. In 2018 IEEE/RSJ International Conference...

  25. [33]

    Contour-based Approach for 3D Mapping of Un- derwater Galleries

    Quentin Massone, Sebastien Druon, Yohan Breux, and Jean Triboulet. Contour-based Approach for 3D Mapping of Un- derwater Galleries. In Global Oceans: Singapore–US Gulf Coast, 2020

  26. [34]

    Edge-Centric Real-Time Segmentation for Autonomous Un- derwater Cave Exploration

    Mohammadreza Mohammadi, Adnan Abdullah, Aishneet Juneja, Ioannis Rekleitis, Jahidul Islam, and Ramtin Zand. Edge-Centric Real-Time Segmentation for Autonomous Un- derwater Cave Exploration. In IEEE International Confer- ence on Machine Learning and Applications (ICMLA), 2024

  27. [35]

    Nvidia TensorRT

    Nvidia. Nvidia TensorRT. https://github.com/ NVIDIA/TensorRT, 2016. Accessed: 01-22-2025

  28. [36]

    Open Neural Network Exchange

    ONNX. Open Neural Network Exchange. https:// onnx.ai/, 2017. Accessed: 01-22-2025

  29. [37]

    Agronav: Autonomous Navigation Framework for Agricul- tural Robots and Vehicles using Semantic Segmentation and Semantic Line Detection

    Shivam K Panda, Yongkyu Lee, and M Khalid Jawed. Agronav: Autonomous Navigation Framework for Agricul- tural Robots and Vehicles using Semantic Segmentation and Semantic Line Detection. InIEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 6272–6281, 2023

  30. [38]

    Semantic Knowledge-Based Representa- tion for Improving Situation Awareness in Service Oriented Agents of Autonomous Underwater Vehicles

    Pedro Patr ´on, Emilio Miguelanez, Joel Cartwright, and Yvan R Petillot. Semantic Knowledge-Based Representa- tion for Improving Situation Awareness in Service Oriented Agents of Autonomous Underwater Vehicles. In OCEANS, 2008

  31. [39]

    Semantic-based Adaptive Mission Plan- ning for Unmanned Underwater Vehicles

    Pedro Patron et al. Semantic-based Adaptive Mission Plan- ning for Unmanned Underwater Vehicles. PhD thesis, Cite- seer, 2010

  32. [40]

    Thirty Years of American Cave Diving Fatalities

    Leah Potts, Peter Buzzacott, and Petar Denoble. Thirty Years of American Cave Diving Fatalities. Diving Hyperbaric Medicine, 46:150–154, 2016

  33. [41]

    Learning Transferable Visual Models from Natural Language Supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models from Natural Language Supervi- sion. In International conference on machine learning, ...

  34. [42]

    Sonar Visual Inertial SLAM of Underwater Structures

    Sharmin Rahman, Alberto Quattrini Li, and Ioannis Rek- leitis. Sonar Visual Inertial SLAM of Underwater Structures. In IEEE International Conference on Robotics and Automa- tion, pages 5190–5196, 2018

  35. [43]

    Contour based Reconstruction of Underwater Struc- tures Using Sonar, Visual, Inertial, and Depth Sensor

    Sharmin Rahman, Alberto Quattrini Li, and Ioannis Rek- leitis. Contour based Reconstruction of Underwater Struc- tures Using Sonar, Visual, Inertial, and Depth Sensor . In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 8048–8053, 2019

  36. [44]

    SVIn2: A Multi-sensor Fusion-based Underwater SLAM System

    Sharmin Rahman, Alberto Quattrini Li, and Ioannis Rek- leitis. SVIn2: A Multi-sensor Fusion-based Underwater SLAM System. International Journal of Robotics Research, 41(11-12):1022–1042, July 2022

  37. [45]

    Evaluating a pid, pure pursuit, and weighted steer- ing controller for an autonomous land vehicle

    Arturo L Rankin, Carl D Crane III, and David G Arm- strong II. Evaluating a pid, pure pursuit, and weighted steer- ing controller for an autonomous land vehicle. In Mobile Robots XII, volume 3210, pages 1–12. SPIE, 1998

  38. [46]

    Sunfish®: A human-portable ex- ploration auv for complex 3d environments

    Kristof Richmond, Chris Flesher, Laura Lindzey, Neal Tan- ner, and William C Stone. Sunfish®: A human-portable ex- ploration auv for complex 3d environments. In OCEANS 2018 MTS/IEEE Charleston, pages 1–9. IEEE, 2018

  39. [47]

    Autonomous Exploration and 3-D Mapping of Underwater Caves with the Human-portable SUNFISH® AUV

    Kristof Richmond, Chris Flesher, Neal Tanner, Vickie Siegel, and William C Stone. Autonomous Exploration and 3-D Mapping of Underwater Caves with the Human-portable SUNFISH® AUV. In Oceans: Singapore–US Gulf Coast , 2020. 16

  40. [48]

    Ad- vancements in The Field of Autonomous Underwater Vehi- cle

    Avilash Sahoo, Santosha K Dwivedy, and PS Robi. Ad- vancements in The Field of Autonomous Underwater Vehi- cle. Ocean Engineering, 181:145–160, 2019

  41. [49]

    A review of some pure-pursuit based path tracking techniques for control of autonomous vehicle

    Moveh Samuel, Mohamed Hussein, and Maziah Binti Mo- hamad. A review of some pure-pursuit based path tracking techniques for control of autonomous vehicle. International Journal of Computer Applications, 135(1):35–38, 2016

  42. [50]

    A comparison of expected flight times for intercept and pure pursuit missiles

    Louis L Scharf, William P Harthill, and Paul H Moose. A comparison of expected flight times for intercept and pure pursuit missiles. IEEE Transactions on Aerospace and Elec- tronic Systems, pages 672–673, 1969

  43. [51]

    Lm-Nav: Robotic Navigation with Large Pre-trained Models of Lan- guage, Vision, and Action

    Dhruv Shah, Bła ˙zej Osi´nski, Sergey Levine, et al. Lm-Nav: Robotic Navigation with Large Pre-trained Models of Lan- guage, Vision, and Action. In Conference on robot learning, pages 492–504, 2023

  44. [52]

    Underwater Multi-Robot Convoying using Visual Tracking by Detection

    Florian Shkurti, Wei-Di Chang, Peter Henderson, Md Jahidul Islam, Juan Camilo Gamboa Higuera, Jimmy Li, Travis Manderson, Anqi Xu, Gregory Dudek, and Junaed Sattar. Underwater Multi-Robot Convoying using Visual Tracking by Detection. In IEEE/RSJ International Conference on Int...

  45. [53]

    Experimental Evaluation of Under- water Semantic SLAM

    Thomas Jeongho Song. Experimental Evaluation of Under- water Semantic SLAM. Master’s thesis, Massachusetts In- stitute of Technology, 2024

  46. [54]

    Roboclip: One Demonstration is Enough to Learn Robot Policies

    Sumedh Sontakke, Jesse Zhang, S ´eb Arnold, Karl Pertsch, Erdem Bıyık, Dorsa Sadigh, Chelsea Finn, and Laurent Itti. Roboclip: One Demonstration is Enough to Learn Robot Policies. Adv. in Neural Information Processing Systems, 36, 2024

  47. [55]

    Deep Learning Waterline Detection for Low-Cost Autonomous Boats

    Lorenzo Steccanella, Domenico Bloisi, Jason Blum, and Alessandro Farinelli. Deep Learning Waterline Detection for Low-Cost Autonomous Boats. In Proceedings of the 15th International Intelligent Autonomous Systems Confer- ence (IAS-15), pages 613–625. Springer, 2019

  48. [56]

    Topological Structural Analysis of Digitized Binary Images by Border Following

    Satoshi Suzuki and KeiichiA be. Topological Structural Analysis of Digitized Binary Images by Border Following. Computer Vision, Graphics, and Image Processing , 30(1): 32–46, 1985

  49. [57]

    BlueME: Robust Underwater Robot-to- Robot Communication Using Compact Magnetoelectric An- tennas

    Mehron Talebi, Sultan Mahmud, Adam Khalifa, and Md Jahidul Islam. BlueME: Robust Underwater Robot-to- Robot Communication Using Compact Magnetoelectric An- tennas. arXiv preprint arXiv:2411.09241, 2024

  50. [58]

    Factor of Safety: Ratio for Safety in Design and Use

    SafetyCulture Content Team. Factor of Safety: Ratio for Safety in Design and Use. https://safetyculture. com/topics/factor-of-safety/ , 2024. Accessed: 01-29-2025

  51. [59]

    Dense, Sonar-based Reconstruction of Underwater Scenes

    Pedro V Teixeira, Dehann Fourie, Michael Kaess, and John J Leonard. Dense, Sonar-based Reconstruction of Underwater Scenes. In International Conference on Intelligent Robots and Systems (IROS), pages 8060–8066, 2019

  52. [60]

    Semantic Mapping for Autonomous Subsea Inter- vention

    Guillem Vallicrosa, Khadidja Himri, Pere Ridao, and Nuno Gracias. Semantic Mapping for Autonomous Subsea Inter- vention. Sensors, 21(20):6740, 2021

  53. [61]

    Real-Time Dense 3D Mapping of Un- derwater Environments

    Weihan Wang, Bharat Joshi, Nathaniel Burgdorfer, Kon- stantinos Batsos, Alberto Quattrini Li, Philippos Mordohai, and Ioannis Rekleitis. Real-Time Dense 3D Mapping of Un- derwater Environments. In IEEE International Conference on Robotics and Automation (ICRA), 2023

  54. [62]

    Underwater Cave Mapping and Recon- struction Using Stereo Vision

    Nicholas Weidner. Underwater Cave Mapping and Recon- struction Using Stereo Vision. Master’s thesis, Computer Science and Engineering Department, University of South Carolina, 2017

  55. [63]

    Underwater Cave Mapping us- ing Stereo Vision

    Nicholas Weidner, Sharmin Rahman, Alberto Quattrini Li, and Ioannis Rekleitis. Underwater Cave Mapping us- ing Stereo Vision. In IEEE International Conference on Robotics and Automation (ICRA), pages 5709 – 5715, 2017

  56. [64]

    A Natural Ges- ture Interface for Operating Robotic Systems

    Anqi Xu, Gregory Dudek, and Junaed Sattar. A Natural Ges- ture Interface for Operating Robotic Systems. In IEEE Int. Conf. on Robotics and Automation, pages 3557–3563, 2008

  57. [65]

    Image Based River Navigation System of Catamaran USV with Image Semantic Segmentation

    Ping Yang, Changhui Song, Linke Chen, and Weicheng Cui. Image Based River Navigation System of Catamaran USV with Image Semantic Segmentation. In WRC Symposium on Advanced Robotics and Automation, pages 147–151, 2022

  58. [66]

    Weakly Supervised Caveline Detection For AUV Navigation Inside Underwater Caves

    Boxiao Yu, Reagan Tibbetts, Titon Barna, Ailani Morales, Ioannis Rekleitis, and Md Jahidul Islam. Weakly Supervised Caveline Detection For AUV Navigation Inside Underwater Caves. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 9933–9940. IEEE, 2023

  59. [67]

    Segmentation of Side Scan Sonar Images on AUV

    Fei Yu, Yuemei Zhu, Qi Wang, Kaige Li, Meihan Wu, Guan- gliang Li, Tianhong Yan, and Bo He. Segmentation of Side Scan Sonar Images on AUV. In IEEE Underwater Technol- ogy (UT), 2019

  60. [68]

    Adaptive Semantic Segmentation for Unmanned Surface Vehicle Navigation

    Wenqiang Zhan, Changshi Xiao, Yuanqiao Wen, Chunhui Zhou, Haiwen Yuan, Supu Xiu, Xiong Zou, Cheng Xie, and Qiliang Li. Adaptive Semantic Segmentation for Unmanned Surface Vehicle Navigation. Electronics, 9(2):213, 2020. 17

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.