{"id":"cc2150ea-670e-43c0-b867-7b7fd4b1ae4d","arxiv_id":"2411.11287","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of weakly and self-supervised deep learning for underwater scene understanding, centered on datasets, techniques, and applications for reducing annotation effort.","lead":"This paper surveys machine learning methods that reduce the need for expert-labelled data in underwater image analysis, focusing on seagrass and coral monitoring. It organizes the field into a taxonomy and catalogs available datasets, gaps, and future directions.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'only one multi-species seagrass dataset' claim conflicts with the survey's own Table A1, which lists a 41-class (5 seagrass) image-level dataset; the data-scarcity conclusion needs a qualifying annotation-type caveat.","rationale":"The reader's verdict accepts the survey and identifies the dataset-search comprehensiveness as the weakest assumption. My stress-test finds a closely related but more specific problem: the survey's own Table A1 appears to contradict the text's 'only one multi-species seagrass dataset' claim. The Roelfsema et al. (2015) row lists 41 classes with 5 seagrass classes and 24,000 images under Image-Level Labels, which, on its face, is a multi-species seagrass dataset unless explicitly excluded by an annotation-type qualifier. No such qualifier is given in the sentence. This matters because the data-availability gap is the survey's principal contribution and the motivation for its call for weakly supervised methods. The qualitative SAM/CoralSCOP generalisation concern flagged by the reader is real but less load-bearing: the central argument does not strictly depend on those two models failing. The recommended verdict is CONDITIONAL because the survey is otherwise careful and useful, but the 'only one dataset' statement should be corrected or explicitly qualified before the data-scarcity conclusion is relied upon.","tokens_in":40710,"tokens_out":4758,"duration_ms":46484,"concrete_test":"Open the Roelfsema et al. (2015) PANGAEA record (doi 10.1594/PANGAEA.846147) and check whether each of the 24,000 images carries per-image labels for 5 seagrass species or only site/transect-level cover estimates. If per-image labels exist, reconcile Section 3.3 with Table A1 by adding a qualifier such as 'only patch-level multi-species seagrass dataset', and soften the Section 10.1.2 claim accordingly. If only aggregate cover data exist, add a sentence in Section 3.3 explaining why the row does not count as a multi-species seagrass image dataset. Either way, this check settles whether the 'only one dataset' claim is overstated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 3.3 and Section 10.1.2 make a load-bearing data-scarcity claim: only DeepSeagrass provides multiple seagrass taxonomic morphotypes. This supports the survey's central argument that weakly supervised methods are necessary because expert labels are scarce. Yet Table A1 lists Roelfsema et al. (2015), 'Benthic and substrate cover data derived from a time series of photo-transect surveys', as an Image-Level Labels dataset with 41 classes including 5 seagrass classes and 24,000 images. If that table row is accurate, the 'only one dataset' sentence is false as written; the authors must qualify it as 'only one dataset with patch-level annotations' or 'only one dataset designed for computer-vision training'. The survey provides no search protocol or inclusion criteria for Appendix A, so an independent reader cannot tell whether this is a deliberate scope decision or an omission. Because the headline conclusion about data availability is the survey's main novel contribution, this inconsistency is a real correctness risk, not a stylistic quibble.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys weakly supervised and self-supervised deep learning for underwater scene understanding, with a focus on two ecological tasks: multi-species seagrass meadow mapping and multi-species coral reef analysis. It reviews coastal ecosystem background, the visual challenges of underwater imagery, computer vision tasks and deep learning architectures, prior fully supervised and weakly supervised approaches in seagrass and coral domains, and relevant weakly supervised and self-supervised methods from the broader computer vision literature. The survey's principal novel contribution is its dataset availability analysis, which argues that expert label dependency is a central bottleneck and that suitable multi-species datasets are extremely scarce. Specifically, it claims that DeepSeagrass is the only dataset providing multiple seagrass morphotypes and that UCSD Mosaics is the only publicly available fine-grained multi-species coral dataset with dense ground truth masks. The paper also proposes a taxonomy, discusses transferability of generic weakly supervised methods to underwater imagery, and outlines future directions including temporal information, large language models, hierarchical labels, and human-robot collaboration.","tokens_in":41074,"tokens_out":5799,"duration_ms":56557,"significance":"If the data-scarcity claims are correctly scoped, this survey fills a real gap in the literature by synthesizing scattered work on reducing label dependency for underwater monitoring. The paper is timely, covers literature through late 2024, and provides useful background material in the taxonomy, the glossary, and the extensive dataset tables. It is also commendable for being explicit about scope choices, such as restricting attention to RGB imagery and to seagrass and coral ecosystems. The qualitative demonstrations in Figs. 19 and 20 are illustrative and appropriately framed as showing a failure to 'group pixels in a meaningful way' rather than as a quantitative benchmark. The main weakness is that the headline dataset-scarcity conclusions are stated more strongly than the survey's own appendix supports, and the absence of a documented search protocol makes these claims hard to verify. These issues are load-bearing because data scarcity is the motivating and organizing thesis of the paper.","major_comments":[{"comment":"The claim that only DeepSeagrass provides multiple seagrass taxa is contradicted by the survey's own Table A1, which lists Roelfsema et al. (2015) as an image-level-label dataset with 41 classes including 5 seagrass classes and 24,000 images. As written, the sentence 'with only one dataset providing multiple taxonomic morphotypes of seagrasses' is false without a qualifier such as 'with patch-level labels' or 'designed for computer vision training.' Section 10.1.2 similarly states that 'the only available multi-species seagrass datasets are labelled with patches,' which is also inconsistent with Table A1. Because the data-scarcity conclusion is the paper's main novel contribution, this inconsistency must be fixed by clarifying the annotation-type scope and explicitly discussing why Roelfsema et al. is excluded.","section":"Section 3.3 and Section 10.1.2"},{"comment":"The statement that 'the only publicly available dataset containing fine-grained multi-species coral images accompanied by dense ground truth masks is the UCSD Mosaics dataset' appears to conflict with Table A2, which lists #DeOlhoNosCorais under the 'Pixel-wise Labels' category with 21 classes and 1,411 images. If the authors regard #DeOlhoNosCorais as not fine-grained, not multi-species, or not comparable because of its social-media imagery, that judgment should be stated explicitly and defended. Otherwise the 'only publicly available' claim is not supported by the paper's own dataset survey and needs to be revised or qualified.","section":"Section 3.3 and Table A2"},{"comment":"The survey provides no search protocol, inclusion criteria, exclusion criteria, or date range for the literature and dataset search, so the reader cannot determine whether the 'only one dataset' claims reflect a deliberate scope decision or an omission. A reproducible methodology paragraph stating the databases searched, search terms, screening rules, and how annotation types (image-level, patch-level, point-level, pixel-wise) were assigned would substantially strengthen the central data-availability analysis and allow independent verification of the headline gaps.","section":"Appendix A and Section 3.3"}],"minor_comments":[{"comment":"The DeepSeagrass row lists '5 classes' while the text in Section 3.3 and Section 10.2.3 describes the four morphotypes Strappy, Ferny, Rounded, and Background; this numerical discrepancy should be reconciled.","section":"Table A1"},{"comment":"The phrase '66,946 single-species patch images' is confusing because DeepSeagrass labels are described elsewhere as morphotypes rather than species; clarifing whether each patch contains one taxon or one morphotype would improve precision.","section":"Section 3.3"},{"comment":"Several datasets under the heading 'Survey of Publicly Available Datasets of Coral Reef Imagery' are not coral-focused (e.g., DUO is an underwater object detection dataset for robot picking, and MAS3K is a marine animal segmentation dataset); a brief sentence on why these datasets are included would help avoid the appearance of scope inconsistency.","section":"Table A2"},{"comment":"The 'Desktop to Reality' subsection discusses an important problem but lacks citations to prior work on domain shift, sim-to-real transfer, or dataset bias in underwater robotics; adding a few references would strengthen this discussion.","section":"Section 10.1.3"},{"comment":"The qualitative failure demonstrations for Segment Anything and CoralSCOP are useful, but it would be clearer to state explicitly that these are illustrative examples rather than systematic evaluations, since the text currently implies more generality than the evidence supports.","section":"Figures 19 and 20"}],"recommendation":"major_revision","confidential_remarks":"The paper's dataset-scarcity claims are the key novel contribution, and the internal contradictions with Tables A1 and A2 are serious enough to require major revision. The authors should also consider whether the 'only one dataset' claims would benefit from an independent verification by the reviewers, given that the survey relies heavily on the authors' own prior publications (Raine et al., 2020-2024) as representative examples of weakly supervised underwater methods. That reliance is not improper, but the headline claims about uniqueness and scarcity should be scoped precisely and made reproducible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good news first: this is a real survey, not a padding exercise. The taxonomy is clearly laid out, the coverage of weakly and self-supervised methods for seagrass and coral imagery is well organized, and the dataset tables in Appendix A are the most useful part of the paper. The authors also state their scope and limitations honestly, which is rarer than it should be.\n\nThe soft spot is the data-scarcity claim that carries the paper's motivation. Section 3.3 says the only dataset providing multiple seagrass morphotypes is DeepSeagrass, and Section 10.1.2 says the only multi-species seagrass datasets are patch-labelled. But Table A1 lists Roelfsema et al. 2015 with image-level labels, 41 classes, 5 seagrass classes, and 24,000 images. That is a direct contradiction unless the authors add \"patch-level\" or \"designed for computer-vision training\" to the claim. Because the whole survey argues the field is label-limited, this needs a fix, not a hand-wave. Also, DeepSeagrass is described as four classes in Section 3.3 and five classes in Table A1. Minor, but embarrassing in a dataset survey.\n\nThe appendix has no stated inclusion protocol, so I cannot independently verify completeness. The citations seem to cover the known landscape, so I'd call that a minor weakness, but it does make the \"only one\" claims fragile. The qualitative figures showing Segment Anything and CoralSCOP failing on a handful of underwater images are illustrative rather than evidence; the text mostly treats them that way, so I won't push on that.\n\nThe self-citations (Raine et al. 2020–2024) are not a problem here. Those are external published works with their own data and results, and they are directly relevant to the survey's focus.\n\nBottom line: the central organization and gap analysis hold up. The \"only dataset\" sentence needs an annotation-type qualifier, and the DeepSeagrass class count must be reconciled. That is a normal survey revision, not a fundamental flaw. I would send it to peer review and expect accept after minor revision.","headline":"A genuinely useful survey of weakly supervised underwater scene understanding, but its data-scarcity claim is contradicted by its own appendix and needs a qualification before publication.","tokens_in":41375,"tokens_out":2180,"would_cite":true,"duration_ms":21519,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The label bottleneck, not the model zoo, is what limits automated underwater monitoring.","keywords":["underwater scene understanding","weakly supervised learning","self-supervised learning","semantic segmentation","coral reef monitoring","seagrass monitoring","point label propagation","dataset survey"],"falsifier":"Compile a complete registry of all underwater datasets released before November 2024 and check whether any contains multi-species seagrass labels beyond DeepSeagrass or dense multi-species coral masks beyond UCSD Mosaics; if such a dataset exists, the survey's central data-gap claim is overstated.","tokens_in":40566,"feed_emoji":"🌊","tokens_out":6662,"duration_ms":54214,"temperature":0.7,"pith_summary":"This survey argues that the bottleneck for automated analysis of underwater imagery is not model architecture but labelled data: the scenes are complex, species boundaries are fuzzy, and only marine ecologists can label pixels reliably, which crowdsourcing cannot fix. Its central claim is that weakly supervised and self-supervised deep learning can reduce this dependency, using coarser labels such as image-level classes, sparse points, or self-supervised features in place of dense expert masks. The survey positions these methods inside a taxonomy that joins underwater monitoring, computer vision, and deep learning, and reviews the available datasets to show where the gaps are. If the claim is right, marine surveys by robots and vehicles could be scaled to monitor reef health and blue carbon stocks without an unsustainable labelling burden.","feed_headline":"Only two datasets carry fine-grained species labels for underwater AI","feed_subtitle":"A survey of weak supervision shows expert labelling costs block reef and seagrass monitoring at scale.","key_machinery":"The organising device is a taxonomy that places weakly supervised underwater image analysis at the intersection of underwater environmental monitoring, computer vision, and deep learning, with 'reducing label dependency' as the explicit goal. The main algorithmic mechanisms reviewed are point label propagation via superpixels (turning sparse Coral Point Count labels into pseudo ground truth masks), image-level labels as supervision for coarse segmentation, self-supervised and vision-language foundation models used as feature extractors or supervisory signals, and human-in-the-loop selection of the most informative sparse points. These mechanisms are evaluated against a dataset survey that classifies annotations into image-level, patch-level, sparse point, and dense pixel-wise labels.","core_discovery":"The central discovery claimed is that the field of weakly supervised underwater scene understanding is small and data-limited: only DeepSeagrass provides multi-species seagrass labels (66,946 patches, four morphotypes), and only UCSD Mosaics provides fine-grained multi-species coral ground truth masks. The survey argues that existing weakly supervised methods—point label propagation, image-level label training, self-supervised foundation models, and human-in-the-loop sparse annotation—can productively reduce expert label dependency, and that the main barrier to progress is the absence of temporal, realistic, densely annotated multi-species image sequences rather than a lack of algorithmic ideas. It further contends that mainstream weakly supervised techniques designed for images with clear object boundaries do not transfer directly to underwater imagery, where foreground and background blur together.","pith_inferences":["If the survey's data-gap map is correct, the single highest-value release for the field would be a temporal, multi-species underwater video dataset with dense pixel-wise masks and GPS, repeated along the same transects over seasons; no current dataset provides this.","The extremely sparse label results for coral suggest an analogous human-in-the-loop protocol for seagrass—choosing the most informative 5–25 image-level or point labels—could make multi-species seagrass segmentation practical, since only patch-level labels exist today.","Mainstream weak supervision assumes central points and hard object boundaries; a direct transfer from remote-sensing or medical point-label methods to underwater imagery should be tested quantitatively on UCSD Mosaics, since the survey only demonstrates mismatch qualitatively."],"forward_implications":["Image-level labels can replace patch-level labels for training coarse multi-species seagrass segmentation, while still yielding a segmentation output at inference.","Point label propagation can turn the abundant Coral Point Count photo-quadrat data into pseudo ground truth for training coral segmentation models.","Foundation models trained on general imagery do not segment underwater scenes meaningfully without adaptation or a weak supervision signal, so domain-specific modelling remains necessary.","Human-in-the-loop selection of sparse points (5–25 per image) can reduce annotation effort for coral segmentation compared to the usual 100–300 points.","Progress in this area depends on new temporal, densely annotated multi-species datasets rather than on further algorithmic tweaks alone."],"supporting_citations":[{"why":"Supplies the DeepSeagrass dataset of 66,946 patch images in four seagrass morphotypes, the only multi-species seagrass label set the survey finds.","marker":"Raine et al, 2020"},{"why":"Source of the UCSD Mosaics dataset, the only publicly available fine-grained multi-species coral data with dense pixel-wise ground truth masks.","marker":"Edwards et al, 2017; Alonso et al, 2019"},{"why":"Introduces the multi-level superpixel point label propagation algorithm that the survey identifies as the main weakly supervised approach for coral segmentation.","marker":"Alonso and Murillo, 2018; Alonso et al, 2019"},{"why":"Improves superpixel-based point label propagation by clustering deep features and using point labels directly, making segments conform to coral boundaries.","marker":"Raine et al, 2022"},{"why":"Segment Anything, used in the survey's qualitative demonstration that general-purpose foundation models fail to group pixels meaningfully in underwater images.","marker":"Kirillov et al, 2023"},{"why":"CoralSCOP, the coral-specific foundation model that still degrades on deep, turbid, or dark imagery, framing the generalisation gap.","marker":"Zheng et al, 2024"},{"why":"Demonstrates human-in-the-loop selection of extremely sparse point labels (5–25 per image) for multi-species coral segmentation.","marker":"Raine et al, 2024b"},{"why":"Contributes the Looking for Seagrass dataset with pixel-wise binary seagrass masks, the common benchmark for binary seagrass segmentation.","marker":"Reus et al, 2018"}],"fun_headline_variants":["Underwater AI: only two datasets with species-level labels","Survey: Weak supervision can cut underwater labeling needs","Two datasets for species labels: underwater AI's bottleneck","Reducing expert labels for coral and seagrass surveys","Weakly supervised methods may ease underwater labeling burden"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's conclusions rest on the assumption that its search of the literature and datasets is complete and that the qualitative failures of Segment Anything and CoralSCOP shown on a few underwater images are representative of those models' general behaviour.","fun_headline_variants_meta":{"raw":{"variants":["Underwater AI: only two datasets with species-level labels","Survey: Weak supervision can cut underwater labeling needs","Two datasets for species labels: underwater AI's bottleneck","Reducing expert labels for coral and seagrass surveys","Weakly supervised methods may ease underwater labeling burden"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000306,"raw_usage":{"total_tokens":1771,"prompt_tokens":982,"completion_tokens":789,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":713}},"tokens_in":598,"tokens_out":789,"duration_ms":8245,"temperature":1.0,"reasoning_tokens":713,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:40:41.388177+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compile a complete registry of all underwater datasets released before November 2024 and check whether any contains multi-species seagrass labels beyond DeepSeagrass or dense multi-species coral masks beyond UCSD Mosaics; if such a dataset exists, the survey's central data-gap claim is overstated.","supporting_citations":[{"cited_title":"In: Proceedings of the Digital Image Computing: Techniques and Applications","cited_arxiv_id":null,"evidence_quote":"Supplies the DeepSeagrass dataset of 66,946 patch images in four seagrass morphotypes, the only multi-species seagrass label set the survey finds."},{"cited_title":"IEEE Robotics and Automation Letters 7(3):8291--8298","cited_arxiv_id":null,"evidence_quote":"Improves superpixel-based point label propagation by clustering deep features and using point labels directly, making segments conform to coral boundaries."},{"cited_title":"In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition","cited_arxiv_id":null,"evidence_quote":"CoralSCOP, the coral-specific foundation model that still degrades on deep, turbid, or dark imagery, framing the generalisation gap."},{"cited_title":"o ller T, J \\","cited_arxiv_id":null,"evidence_quote":"Contributes the Looking for Seagrass dataset with pixel-wise binary seagrass masks, the common benchmark for binary seagrass segmentation."}],"review_version":1}