{"id":"11888ad7-8ce2-492a-b81d-511248425013","arxiv_id":"2412.01461","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey of 38 USV vision datasets and the deep-learning methods applied to them, with a list of gaps that limit autonomous boat perception.","lead":"This paper reviews 38 public datasets and the deep-learning methods used for computer vision on unmanned surface vehicles, and it catalogs open problems such as missing 3D data, sparse annotations, and absent calibration. A generalist reader can use it as a map of who published which maritime datasets, which sensors are covered, and where the field is still thin.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal contradiction undermines the headline gap claim: the paper says public USV datasets lack calibration, yet its own Table 4 lists publicly accessible ROAM CRAS and Pohang Canal as providing calibration. The central data-infrastructure bottleneck needs re-verification before acceptance.","rationale":"I agree with the reader's conditional verdict but not exactly with the weakest assumption. The reader emphasizes missing search protocol; I find a more immediate, verifiable problem: the manuscript contradicts itself on the calibration-availability claim that supports the central bottleneck argument. This is an internal consistency issue rather than a speculation about omitted datasets. The no-public-3D-annotations claim is not directly contradicted by any entry I can see, which is why I recommend keeping the paper conditional rather than rejecting or accepting. A single verification of the two calibration-providing datasets would decide the matter. The synthetic FoggyShipInsseg entry is a second, smaller indicator that the inventory was not carefully checked, reinforcing the need for the verification step.","tokens_in":42884,"tokens_out":6967,"duration_ms":59773,"concrete_test":"Check the public availability and contents of ROAM CRAS and Pohang Canal by visiting the official project pages/repositories or contacting the authors. If either dataset is downloadable and includes calibration data (even without 3D perception annotations), Sections 2.1.2 and 4.1 must be revised to say public calibration exists but is rare or unannotated, and the central 'no public 3D data infrastructure' claim should be rephrased as 'no annotated 3D perception data.' If both are closed or lack calibration files, the contradiction is a wording error and the conditional acceptance can stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's contribution is a gap analysis: it concludes that USV perception is limited by data infrastructure, specifically 'no public 3D vision dataset' and 'almost no calibration data supplied' (Sections 4.1). These conclusions are load-bearing and depend on the accuracy of Tables 2-4. That dependency fails internally on the calibration claim. Section 2.1.2 states 'it lack of calibration data in the public dataset of USVs,' and Section 4.1 repeats that public multi-modal datasets have almost no calibration. But Section 2.1.4 states that ROAM CRAS (Campos et al., 2022) and Pohang Canal (Chung et al., 2023) 'also provided calibration data' and 'are publicly accessible.' Table 4 marks ROAM CRAS as Open/Limited and Pohang Canal as Open/Y and lists Calibration under sensors. The same internal reliability problem hits the 38 real-world dataset count: FoggyShipInsseg appears in Table 3 with sensor type 'Synthetic,' after Section 2 explicitly excludes synthetic datasets. A survey whose headline findings are absence claims must at least be consistent with its own inventory; currently it is not.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys vision datasets and deep learning techniques for Unmanned Surface Vehicles (USVs). It compiles 38 public datasets collected in real-world scenarios from 2015 to 2024, organizes them by task (object detection, segmentation, and other vision tasks), compares USV datasets with autonomous vehicle (AV) datasets, reviews single-sensor and multi-sensor deep learning methods, and concludes with a set of challenges and future directions, including the lack of public 3D perception datasets, sparse and inconsistent annotations, missing metadata, and privacy leaks.","tokens_in":43092,"tokens_out":3099,"duration_ms":27529,"significance":"If the inventory is accurate, this is a useful reference and gap analysis for the USV vision community. The paper's strengths include its broad coverage, a detailed chronological overview, a structured taxonomy of deep learning techniques, a quantitative comparison with AV datasets, and an explicit discussion of data privacy. However, the central contribution is the dataset inventory and the absence claims built on it, and those claims currently rest on an internally inconsistent set of tables and prose. The negative claims about calibration data and 3D perception benchmarks are plausible but need to be re-verified against the paper's own inventory before the survey can be relied upon.","major_comments":[{"comment":"The Section 2 introduction states that the authors 'disregard all synthetic datasets that are generated through simulation or deep learning generative models,' yet Table 3 lists FoggyShipInsseg (Sun et al., 2022b) with sensor type 'Synthetic' and includes it in the analysis. Since the paper's headline count of 38 real-world datasets depends on excluding synthetic data, this inconsistency directly affects the paper's central quantitative claim. Please either remove or reclassify FoggyShipInsseg, or revise the stated inclusion criterion to match the actual inventory.","section":"Section 2, Table 3"},{"comment":"The calibration gap is stated as a key finding: Section 2.1.2 says 'it lack of calibration data in the public dataset of USVs' and Section 4.1 says 'there was almost no calibration data supplied' for public multi-modal datasets. However, Section 2.1.4 explicitly states that ROAM CRAS (Campos et al., 2022) and Pohang Canal (Chung et al., 2023) 'also provided calibration data' and are publicly accessible, and Table 4 marks ROAM CRAS as Open/Limited and Pohang Canal as Open/Y while listing Calibration under their sensors. This internal contradiction undermines a load-bearing conclusion. The authors need to reconcile the prose with Table 4 and either strengthen or qualify the calibration-scarcity claim.","section":"Sections 2.1.2, 2.1.4, 4.1, Table 4"},{"comment":"The claimed count of 38 datasets is not verifiable from the manuscript itself. Section 2.1.4 discusses a 'Small ShipInsseg' dataset (Sun et al., 2023b) that does not appear in any table and is not counted in the 38, and the same paragraph cites 'Campos (Jeong et al., 2024)' in a way that conflates two distinct datasets (ROAM CRAS and Catabot). Since the survey's contribution is precisely the completeness and accuracy of its inventory, the authors should provide a transparent enumeration of all included datasets, reconcile prose with tables, and correct the mislabeled citation.","section":"Section 2.1.4, Tables 2-4"},{"comment":"The paper does not document its search protocol, inclusion/exclusion criteria, or the verification process for table attributes. This matters because the headline findings are absence claims: no public 3D perception datasets, almost no calibration data, and sparse metadata. Absence claims cannot be assessed without knowing the search scope, databases queried, keywords, date of search, and how each dataset entry was checked against the primary source. Additionally, there are unresolved date conflicts in the inventory: Kolomverse is cited as Nanda et al. (2024) but Table 2 lists year 2022, and SeaSAW is cited as Kaur et al. (2022) but Table 2 lists year 2023. Please add a methodology subsection and correct the year mismatches.","section":"Section 2, Tables 2-4"}],"minor_comments":[{"comment":"The table header contains a typo, 'detction', which should read 'detection'.","section":"Table 1"},{"comment":"The caption describes 'panotpic segmentation,' which should be 'panoptic segmentation.'","section":"Table 3"},{"comment":"Dataset names are inconsistently capitalized, for example 'MariShipInsSeg' in Section 2.1.2 versus 'MariShipInsseg' in Figure 3; please standardize all dataset names.","section":"Figure 3"},{"comment":"The year column for Kolomverse (2022) conflicts with the reference 'Nanda et al., 2024' cited in the same row, and SeaSAW shows 2023 while the reference is 'Kaur et al., 2022'. These should be checked against the original publications.","section":"Table 2"},{"comment":"The sentence 'Campos (Jeong et al., 2024) focuses on multi-domain inspection and maintenance of USVs' is confusing because Campos is the first author of the ROAM CRAS paper, while Jeong et al. (2024) is the Catabot reference. Please disambiguate the citation.","section":"Section 2.1.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a survey whose value depends entirely on the accuracy of its dataset inventory. The internal contradictions noted in the major comments, especially the synthetic-data inclusion and the calibration-data contradiction, mean the paper cannot be accepted in its current form. A careful revision that adds a methodology section, corrects the tables, and re-derives the gap claims from the corrected inventory would bring the paper within scope. The repeated self-citations are worth monitoring but are not by themselves disqualifying."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this survey is genuinely useful as a map of 38 USV vision datasets and the deep learning methods applied to them. It fills a real gap, and the qualitative observation that the field is limited by data infrastructure—sparse annotations, missing metadata, no public 3D benchmarks, privacy leaks—will ring true to anyone who works in the area. The AV comparison and the annotated-object inventory are nice additions. But the paper's own inventory contradicts its headline claims, and those contradictions need to be fixed before I'd trust it as a reference.\n\nThe load-bearing problem is calibration. Section 2.1.2 says public USV datasets lack calibration data, and Section 4.1 repeats that 'almost no calibration data' exists. But Section 2.1.4 and Table 4 list ROAM CRAS and Pohang Canal as publicly accessible with calibration data. That's not a minor typo; it's the central negative claim of the survey. If two public datasets provide calibration, the argument that multimodal fusion is blocked by missing calibration needs to be re-worded and re-verified.\n\nThe same issue hits the dataset count. Section 2 excludes synthetic datasets, but Table 3 includes FoggyShipInsseg with sensor type 'Synthetic'. The prose also mentions Small ShipInsSeg, which doesn't appear in any table or in the 38 count. Year mismatches for Kolomverse (2022 in Table 2, 2024 in the citation) and SeaSAW (2023 vs 2022) suggest the tables weren't carefully checked. There's also no search protocol or inclusion criteria, so we can't independently audit the completeness of the 38-dataset inventory.\n\nNone of this sinks the core contribution. The survey consolidates a scattered literature, the taxonomy of techniques is reasonable, and the qualitative gap analysis is credible. But a survey whose findings are absence claims—'no public 3D data', 'no calibration'—has to be internally consistent with its own tables. Right now it isn't.\n\nWho's this for? Anyone starting in USV perception who wants a quick map of datasets and methods. It deserves a serious referee, but it needs a careful revision pass before publication. I'd send it back for major revision, with instructions to fix the calibration contradiction, resolve the synthetic-dataset inclusion, document the search protocol, and reconcile all dataset counts.","headline":"A useful survey of USV vision datasets and methods whose central gap analysis is undercut by internal inconsistencies in the dataset inventory, especially the calibration claim.","tokens_in":43647,"tokens_out":3101,"would_cite":false,"duration_ms":26242,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A comprehensive review of 38 public USV vision datasets and the deep learning techniques trained on them finds that missing 3D data, calibration, annotations, metadata, and privacy protection—not model design—are what hold the field back.","keywords":["unmanned surface vehicles","vision datasets","deep learning","maritime perception","object detection","segmentation","multi-sensor fusion","survey"],"falsifier":"A reader could check release listings and preprint archives up to mid-2024: finding a public USV dataset that provides 3D object detection or depth ground truth with sensor calibration would directly refute the central claim, as would showing that any of the 38 listed 'not open' or 'not public' datasets is in fact openly downloadable and changes the counts on which the conclusions rest.","tokens_in":42654,"feed_emoji":"🚢","tokens_out":7829,"duration_ms":61033,"temperature":0.7,"pith_summary":"Unmanned surface vehicles (USVs) need vision to navigate, but this review argues that progress is currently blocked less by the algorithms than by the data available to train them. The authors compile and analyze 38 public datasets collected in real-world maritime conditions, covering object detection, segmentation, classification, tracking, and raw multi-sensor recordings. They report that no public USV dataset supports 3D perception tasks such as 3D detection, 3D segmentation, or depth estimation, and that calibration data, consistent object annotations, and privacy safeguards are largely absent. If the review's inventory is correct, then the field's next advances will come from building shared, well-annotated, privacy-aware maritime datasets rather than from novel network architectures.","feed_headline":"No public 3D dataset exists for USV vision, review finds","feed_subtitle":"Survey of 38 public datasets shows sparse annotations, missing calibration, and privacy leaks are the field's limits.","key_machinery":"The carrying mechanism is the dataset inventory itself: a structured catalog of 38 datasets (Tables 2–6) with uniform attributes including sensor type, resolution, FPS, tasks, number of annotated object classes, availability, field of view, and metadata. The work it does is to make gap analysis possible: by aligning all datasets on the same attribute axes, the review turns scattered release papers into a countable evidence base that supports negative claims such as 'no public 3D dataset' and comparative claims such as 'USV datasets trail AV datasets in every sensor and task category.' A second mechanism is the taxonomy of deep learning techniques, split into single-sensor and multi-modal branches, which shows that most USV-specific work reuses off-the-shelf architectures (YOLO, UNet, DeepLab, Mask R-CNN, transformers) and that innovation is concentrated in small modifications rather than fundamental building blocks.","core_discovery":"On its own terms, the paper's central discovery is a gap analysis: compared both with earlier maritime surveys and with the much larger autonomous-vehicle (AV) dataset ecosystem, USV vision is starved of exactly the data that modern perception systems depend on. The authors count 38 public datasets collected in real-world USV scenarios and categorize them by sensors, tasks, resolution, frame rate, number of annotated object classes, availability, field of view, area, location, and metadata. Their analysis finds zero public datasets supporting 3D object detection, zero supporting 3D segmentation, and zero supporting depth estimation, while 2D object detection has 15 public datasets and AV counterparts have 55. Among the larger multi-sensor collections (LiDAR, radar, stereo, sonar), the authors find that calibration data and annotations on those sensors are almost always missing, and that most multi-modal fusion methods reported in the literature are therefore validated on private, self-collected data. The authors further document sparse and inconsistent object annotations across datasets, minimal metadata describing weather, lighting, or water conditions, and a privacy deficit in which only one dataset claims to blur human identities and another visibly leaks a face. The conclusion the authors draw is that the central bottleneck for USV vision is data infrastructure, not model design.","pith_inferences":["Editorial inference: the review's exclusion of synthetic data makes its 'no 3D dataset' claim narrower than it sounds; the authors' own technique review shows depth-estimation work already falls back on synthetic depth because real ground truth is absent, so simulated maritime 3D data may be the fastest bridge to 3D perception for USVs.","Editorial inference: a concrete test of the paper's bottleneck thesis is the pace of fusion research: if a public multi-sensor dataset with calibration appeared, the number of camera-LiDAR-radar fusion papers validated on public data should rise sharply, since the paper's own survey shows such methods are currently limited mainly by data, not by available architectures.","Editorial inference: the privacy leak the paper points to suggests that future dataset design should integrate de-identification at capture time (e.g., camera placement and resolution choices) rather than as a post hoc blur, because blurring removes the very fine detail detection and segmentation models need."],"forward_implications":["Researchers and operators should assume that any USV perception system needing 3D information, such as depth or 3D bounding boxes, must generate its own annotations or adapt models from AV data, because no public maritime dataset currently provides them.","Multi-sensor fusion work (camera-LiDAR, camera-radar, camera-LiDAR-radar) will stay confined to private datasets until public releases include synchronized annotations and calibration matrices, which the paper shows are almost never supplied.","Comparable benchmarking across USV datasets is impossible until object-class definitions are standardized; the paper shows very little overlap in annotated objects, which blocks transfer learning and domain adaptation.","Privacy-aware data release becomes a prerequisite for future datasets: at least one widely used public dataset contains a recognizable human face, and no dataset explains its de-identification procedure, exposing the field to regulatory risk."],"supporting_citations":[{"why":"previous maritime vision dataset survey covering 15 datasets; the comparison baseline for coverage and the reason the authors can claim their review is broader.","marker":"Su et al. (2023)"},{"why":"previous survey of marine vision-based situational awareness for USVs using 12 datasets; establishes the prior state of USV-only reviews that the paper extends.","marker":"Qiao et al. (2021)"},{"why":"review of deep learning-based marine object detection across platforms; used to show that prior technique reviews were not USV-specific in this combined way.","marker":"Zhang et al. (2021a)"},{"why":"survey of small-object detection in maritime settings with 22 datasets; another prior dataset count the paper compares against.","marker":"Rekavandi et al. (2022)"},{"why":"review of ship detection with deep learning over 21 datasets; supplies the comparison of coverage and shows the gap in combined dataset-plus-technique surveys.","marker":"Er et al. (2023)"},{"why":"survey of deep learning for visual detection of marine organisms with 25 datasets; included in Table 1 comparison of prior maritime surveys.","marker":"Wang et al. (2023b)"},{"why":"LaRS dataset; the only public USV dataset the paper identifies that claims to blur human identities, supporting the privacy-gap finding.","marker":"Žust et al. (2023)"},{"why":"Catabot dataset; provides calibration data for camera-LiDAR fusion but is not publicly available, supporting the claim that calibration data is effectively absent from public releases.","marker":"Jeong et al. (2024)"},{"why":"MaSTr1478 dataset; used as the example of a public dataset with a leaked human face, grounding the privacy concern.","marker":"Žust and Kristan (2022)"}],"fun_headline_variants":["Zero public 3D datasets for USV vision, review finds","USV perception hit by data gap: no 3D, depth, or segmentation","Review: 38 USV datasets, but none support 3D or depth tasks","USV vision bottleneck is data infrastructure, not models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's gap analysis stands or falls on the completeness and accuracy of its list of 38 public USV vision datasets; if a public dataset that supports 3D perception or supplies calibration was missed, the headline claims about missing data would be wrong.","fun_headline_variants_meta":{"raw":{"variants":["Zero public 3D datasets for USV vision, review finds","USV perception hit by data gap: no 3D, depth, or segmentation","Review: 38 USV datasets, but none support 3D or depth tasks","USV vision bottleneck is data infrastructure, not models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000855,"raw_usage":{"total_tokens":3757,"prompt_tokens":1027,"completion_tokens":2730,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":2650}},"tokens_in":643,"tokens_out":2730,"duration_ms":19023,"temperature":1.0,"reasoning_tokens":2650,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:18:00.793612+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could check release listings and preprint archives up to mid-2024: finding a public USV dataset that provides 3D object detection or depth ground truth with sensor calibration would directly refute the central claim, as would showing that any of the 38 listed 'not open' or 'not public' datasets is in fact openly downloadable and changes the counts on which the conclusions rest.","supporting_citations":[{"cited_title":", author Liu, G","cited_arxiv_id":null,"evidence_quote":"previous survey of marine vision-based situational awareness for USVs using 12 datasets; establishes the prior state of USV-only reviews that the paper extends."},{"cited_title":"A Guide to Image and Video based Small Object Detection using Deep Learning : Case Study of Maritime Surveillance","cited_arxiv_id":"2207.12926","evidence_quote":"survey of small-object detection in maritime settings with 22 datasets; another prior dataset count the paper compares against."},{"cited_title":", author Chadda, A","cited_arxiv_id":null,"evidence_quote":"Catabot dataset; provides calibration data for camera-LiDAR fusion but is not publicly available, supporting the claim that calibration data is effectively absent from public releases."}],"review_version":1}