{"id":"fb1022b7-9908-42e1-8b45-85180e925c8e","arxiv_id":"2412.11840","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of sonar-based deep learning that identifies robustness, dataset scarcity, and sim-to-real gaps as the main obstacles to safe underwater autonomy.","lead":"This review maps how deep learning is used with sonar data on underwater robots, covering tasks, datasets, simulators, and robustness methods. It argues that the robustness of these AI models has been neglected and suggests steps toward safer autonomous underwater vehicles.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first comprehensive robustness survey' claim rests on Table I's classification of prior surveys, and that classification is not supported by a reproducible audit of their full texts.","rationale":"The reader's weakest assumption is exactly the right one: the novelty claim of being the first robustness-focused survey depends on the comparative verdict in Table I that no prior survey substantially covers OOD detection, adversarial attacks, or uncertainty quantification. My stress-test did not find a separate, more severe flaw; the survey is an honest organizational contribution with useful dataset and simulator tables, a constructive GitHub repository, and a credible workflow. The citation error in Table IV, where reference [156] is attributed to 'M. Shell et al.' in the table but to Q. Ma et al. in the text, is real but minor and correctness-limiting only in a bibliographic sense. The more load-bearing issue is evidentiary: the survey asserts a negative novelty result without documenting the search and screening process that would justify it. Absence claims in surveys require a reproducible protocol, especially when the paper's own text acknowledges that earlier surveys discuss denoising, data augmentation, and simulated data as ways to improve reliability under noise. Those topics are robustness-adjacent, so the paper's narrow definition of robustness (OOD, AA, UQ, verification) must be explicit and justified; otherwise Table I may understate prior work. Because this concern is about the strength of the novelty claim rather than the usefulness of the collected material, the appropriate action is to retain the conditional acceptance and require a full-text audit of Table I plus a documented search statement before publication. I do not recommend rejection: the dataset catalog, simulator comparison, and proposed workflow stand independently, and the paper is transparent about the scarcity it documents.","tokens_in":33159,"tokens_out":4852,"duration_ms":47495,"concrete_test":"Retrieve the full texts of the eight surveys listed in Table I and independently audit each for substantive coverage of robustness topics using keyword families: 'robustness', 'uncertainty quantification', 'UQ', 'out-of-distribution', 'OOD', 'adversarial attack', 'adversarial example', and 'neural network verification'. Any section-level discussion beyond a passing mention should flip the corresponding Table I cell. Separately, run a structured literature search on IEEE Xplore, Scopus, and Google Scholar with queries combining 'sonar' with 'neural network verification', 'adversarial attack', 'out-of-distribution', and 'uncertainty quantification' for deep learning; compare the resulting set against the 10 papers in Table IV. If the audit finds substantive prior coverage, the 'first' claim in the abstract and Section II should be softened and Table I corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central positioning is that it is the first sonar-based deep learning survey under the scope of robustness, and this is supported primarily by Table I, where all eight prior surveys are marked as lacking OOD, adversarial attack, and uncertainty quantification coverage (with only a 'brief mention' of UQ granted to Steiniger et al. [21]). This classification is load-bearing because the abstract, Section II, and Section V all repeat the 'first' claim. However, the paper itself notes that earlier surveys address robustness-adjacent content: [21] discusses data augmentation and GAN-based simulated data, [23] covers types of sonar noise and denoising, and [26] explicitly surveys denoising of sonar images. Whether those topics count as 'robustness' is a definitional choice, but the paper never states its inclusion/exclusion criteria for the robustness column, nor does it provide a documented search or full-text keyword audit for the eight prior surveys. The same gap applies to the stronger assertion in Section IV-G that zero papers exist on neural network verification for sonar: no database, query, date range, or screening protocol is described, so 'zero papers found' is only as strong as an unstated search. If any prior survey has substantive treatment of these topics, or if the unstated verification search misses existing work, the novelty claim is weakened and the survey's roadmap changes from 'first' to 'one of several'.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is an accepted-version survey of sonar-based deep learning for underwater robotics, explicitly organized around robustness. It reviews perception tasks (classification, detection, segmentation, SLAM), catalogs 19 open-source sonar datasets and 8 open-source simulators, summarizes synthetic-data generation, and surveys robustness methods under four headings: neural network verification, adversarial attacks, out-of-distribution detection, and uncertainty quantification. It closes with a proposed pre/post-deployment robustness workflow and a public GitHub repository of sonar datasets. The paper's central positioning is that it is the first comprehensive robustness-focused sonar DL survey, supported by a comparative table of prior surveys (Table I) and by counts of existing robustness work (Section IV-G, Table IV).","tokens_in":33414,"tokens_out":7144,"duration_ms":61937,"significance":"If the novelty claim survives scrutiny, the paper would be a genuinely useful consolidated reference for AUV practitioners: it gathers a dispersed body of datasets, simulators, and robustness methods, identifies research gaps, and contributes an open repository and a concrete workflow for robustness evaluation. The absence of derivations means the paper's soundness rests on factual accuracy and completeness of coverage rather than on proof; on breadth it is strong, but the central 'first comprehensive' and scarcity claims are currently not independently checkable. The paper also gives credit where due to recent open-data efforts and to robustness works by several groups, including the authors' own, and it is generally clearly written.","major_comments":[{"comment":"","section":"Section II, Table I and Section IV-G"},{"comment":"","section":"Section IV-D and Table IV"},{"comment":"","section":"Table IV and Section IV-D references"}],"minor_comments":[{"comment":"","section":"Section III-D"},{"comment":"","section":"Table II and Section IV-A"},{"comment":"","section":"Section IV-A"},{"comment":"","section":"Section I-A and Section III-A"},{"comment":"","section":"Fig. 7 caption"},{"comment":"","section":"Section IV-B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for the journal's scope and fills a real gap if its search and classification claims can be substantiated. The main risk is the unsupported 'first comprehensive' and 'zero papers' assertions; these are fixable with a documented methodology section. The authors' own datasets and methods (SWDD, SubPipe, ROSAR) are directly relevant and their inclusion is not inappropriate, but the editor may want the revision to make self-authored contributions clearly distinguishable from surveyed third-party work and to subject the authors' results to the same critical scrutiny as other methods."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a useful survey, not a breakthrough. Its real value is organizational: nineteen datasets compared side-by-side, a simulator table that actually lists which sensors each one renders, a robustness-method table (Table IV), and a GitHub repository that centralizes the datasets. I'll point new students here before sending them into the literature. The decision not to compare model performances across papers is also sensible, given how heterogeneous the datasets are.\n\nThe robustness framing is the right lens and does fill a gap. Prior sonar-DL surveys mostly compare models and lament the lack of open data; this one systematically walks through adversarial attacks, OOD detection, and uncertainty quantification as applied to sonar, and the proposed pre-deployment workflow is a reasonable synthesis.\n\nWhere it gets soft: the 'first comprehensive robustness survey' claim is load-bearing, and the support is not as solid as the abstract suggests. Table I marks every prior survey as lacking OOD, adversarial-attack, or UQ coverage, but the criteria for that classification are never stated, and no documented full-text audit is given. The same applies to the stronger claim of zero papers on neural network verification for sonar: no search database, query, date range, or screening protocol. That zero might be true, but as written it is an unverifiable assertion. The reasonable fix is to soften to 'to our knowledge' and describe the search. There is also a concrete citation error: Table IV lists 'M. Shell et al. [156]' for LASA, but reference [156] is Q. Ma, L. Jiang, and W. Yu. That should be corrected before publication.\n\nOn self-citations: SWDD, SubPipe, and ROSAR appear prominently, but they are genuinely relevant open datasets and robustness work, and the reader can check them. I do not see a fairness problem there.\n\nNone of this is fatal. The survey would still be useful if it were 'one of the first' rather than 'the first.' The missing search protocol is a minor revision issue, not a reason to reject.\n\nWho is this for? Practitioners building sonar-based DL systems, new researchers wanting a map of datasets and simulators, and referees who need a quick check on what robustness methods exist in this niche. It deserves a serious referee and, after the citation fix and a softened novelty claim, publication.","headline":"A genuinely useful consolidation of sonar-DL datasets, simulators, and robustness work; the 'first robustness survey' claim is weaker than advertised, but the practical value survives that.","tokens_in":33966,"tokens_out":1624,"would_cite":true,"duration_ms":27616,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims to be the first robustness-focused overview of sonar deep learning, reporting zero papers on neural network verification and only a handful each on adversarial attacks, out-of-distribution detection, and uncertainty…","keywords":["sonar-based deep learning","autonomous underwater vehicles","robustness","neural network verification","out-of-distribution detection","adversarial attacks","uncertainty quantification","sonar datasets"],"falsifier":"A systematic literature search for sonar combined with neural network verification before this paper's submission would falsify the zero-paper claim if it found any published verification of a sonar-trained network, and finding a prior survey with substantive out-of-distribution or adversarial coverage would falsify the first-survey claim.","tokens_in":32949,"feed_emoji":"🌊","tokens_out":7196,"duration_ms":63770,"temperature":0.7,"pith_summary":"The paper sets out to be the first survey that treats sonar-based deep learning through the lens of reliability rather than just model accuracy. It assembles the state of the art in sonar perception tasks, along with 19 open datasets and underwater simulators, and then tabulates how much robustness research exists. Its central finding is that the area is almost empty: no neural network verification studies for sonar, four adversarial-attack studies, four out-of-distribution studies, and two uncertainty-quantification studies. The authors argue this matters because autonomous underwater vehicles increasingly rely on real-time sonar deep learning for navigation, and noisy, variable sonar data can silently break a model. The paper closes with a pre-deployment workflow intended to make sonar deep learning safe enough to trust.","feed_headline":"Zero papers verify neural nets for sonar, survey finds","feed_subtitle":"First safety-focused survey of sonar deep learning maps 19 datasets and shows why AUV perception fails in the field.","key_machinery":"The organizing device is a four-part robustness lens, namely neural network verification, adversarial attacks, out-of-distribution detection, and uncertainty quantification, applied across four systematic comparisons: prior surveys, sonar datasets, underwater simulators, and existing robustness publications. The prescriptive output is the proposed pre-deployment workflow, which chains transfer learning, data augmentation, out-of-distribution and uncertainty checks, and either neural network verification or adversarial testing before a model is allowed to drive an autonomous vehicle.","core_discovery":"The central claim is that robustness of sonar-based deep learning is not merely under-explored but nearly untouched: the survey finds zero papers on neural network verification of sonar models, and only single-digit counts in adversarial attacks, out-of-distribution detection, and uncertainty quantification. The paper also documents a concrete failure mode: a model trained on same-location side-scan sonar data can drop from 98% to 15% average precision when deployed under a different date and vehicle altitude, showing that domain shift is severe. It positions itself as the first comprehensive overview to map this landscape, comparing prior surveys, datasets, simulators, and robustness methods, and it proposes a workflow for verifying and hardening models before deployment.","pith_inferences":["The 98%-to-15% precision collapse suggests that dataset shift is the dominant threat to sonar deep learning, so test-time adaptation methods developed for optical perception are a plausible next step that the paper does not explore.","The zero count for neural network verification likely reflects that available verifier tools target classification while most sonar tasks are detection and segmentation, making the adaptation of verifiers to one-stage detectors the highest-leverage test of the roadmap.","The proposed workflow could be turned into a benchmark: a standard sonar dataset plus prescribed perturbations such as black-line dropouts, altitude changes, and sonar brand changes would let the community quantify robustness improvements quantitatively.","The scarcity counts are a snapshot of the literature up to the paper's compilation, so the near-empty table is likely to fill quickly once the field's attention shifts toward reliability."],"forward_implications":["If the field adopts the proposed workflow, a sonar deep learning model would not be deployed until it passes out-of-distribution and uncertainty checks and either verification or adversarial testing, making safety a formal development step.","A shared community repository of open sonar datasets would allow different models to be compared on identical data, ending the current practice of comparing models trained on incompatible private datasets.","The documented sensitivity to sonar setup, including frequency, altitude, and colormap, implies that deployment documentation should record those parameters and that models should only be used inside a matching operating envelope.","Because neural network verification is completely empty in sonar, the first verifiable sonar model would open a new research direction rather than extend an existing one.","The paper's evidence that denoising alone cannot guarantee correct prediction redirects research attention from pre-processing toward intrinsic model robustness."],"supporting_citations":[{"why":"Prior survey of sonar automatic target recognition that the paper positions as not covering robustness, forming the baseline for the Table I novelty claim.","marker":"[20]"},{"why":"Prior survey of deep learning computer vision for sonar that only briefly mentions uncertainty, used to contrast the paper's robustness focus.","marker":"[21]"},{"why":"Pipeline inspection dataset that supplies the same-location, different-altitude training and validation images showing average precision dropping from 98% to 15%.","marker":"[57]"},{"why":"Adversarial re-training framework for side-scan sonar object detection, one of the four adversarial-attack works and the source of the natural black-line signal-loss example.","marker":"[100]"},{"why":"Perceptual Metric Prior method for out-of-distribution detection in synthetic aperture sonar, load-bearing for the OOD section's sonar-specific evidence.","marker":"[164]"},{"why":"Provides the PLUD loss and the first open-set long-tail sonar recognition benchmark, underpinning the out-of-distribution scarcity count.","marker":"[99]"},{"why":"Noise Adversarial Network for sonar object detection, one of the four adversarial-attack papers and evidence of an 8.9% mAP robustness gain.","marker":"[155]"},{"why":"CycleGAN-based uncertainty pipeline for forward-looking sonar detection and classification, one of only two uncertainty-quantification works in the table.","marker":"[177]"},{"why":"Self-supervised fish detection with uncertainty regularization on forward-looking sonar, the other uncertainty-quantification work supporting the scarcity count.","marker":"[178]"}],"fun_headline_variants":["Sonar deep learning: zero papers on neural-net verification","AUV sonar models unverified: survey finds zero checks","Domain shift reduces sonar model AP from 98% to 15%","Survey exposes missing robustness verification in sonar DL"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's claim to be the first robustness-focused survey depends on its Table I judgement that every earlier sonar deep learning survey lacks substantive treatment of out-of-distribution detection, adversarial attacks, and uncertainty quantification.","fun_headline_variants_meta":{"raw":{"variants":["Sonar deep learning: zero papers on neural-net verification","AUV sonar models unverified: survey finds zero checks","Domain shift reduces sonar model AP from 98% to 15%","Survey exposes missing robustness verification in sonar DL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000597,"raw_usage":{"total_tokens":2756,"prompt_tokens":872,"completion_tokens":1884,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":1814}},"tokens_in":488,"tokens_out":1884,"duration_ms":13514,"temperature":1.0,"reasoning_tokens":1814,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:30:15.144568+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search for sonar combined with neural network verification before this paper's submission would falsify the zero-paper claim if it found any published verification of a sonar-trained network, and finding a prior survey with substantive out-of-distribution or adversarial coverage would falsify the first-survey claim.","supporting_citations":[{"cited_title":"ROSAR: An Adversarial Re-Training Framework for Robust Side-Scan Sonar Object Detection","cited_arxiv_id":"2410.10554","evidence_quote":"Adversarial re-training framework for side-scan sonar object detection, one of the four adversarial-attack works and the source of the natural black-line signal-loss example."},{"cited_title":"A perceptual metric prior on deep latent space improves out-of-distribution synthetic aperture sonar image classification,","cited_arxiv_id":null,"evidence_quote":"Perceptual Metric Prior method for out-of-distribution detection in synthetic aperture sonar, load-bearing for the OOD section's sonar-specific evidence."},{"cited_title":"Deep learning with self-supervision and uncertainty regularization to count fish in underwater images,","cited_arxiv_id":null,"evidence_quote":"Self-supervised fish detection with uncertainty regularization on forward-looking sonar, the other uncertainty-quantification work supporting the scarcity count."}],"review_version":1}