Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

A Structured Review of Underwater Object Detection Challenges and Solutions: From Traditional to Large Vision Language Models

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A structured review argues that underwater object detection still lacks a complete solution, that DALL-E 3 synthetic images give only marginal YOLO11 gains, and that LVLMs localize well but hallucinate class names.

desk verdict Useful survey with two thin case studies; the quantitative claims about DALL-E 3 augmentation do not survive scrutiny, but the review half is honest and could stand after reframing. read the letter →

arxiv 2509.08490 v1 pith:YP4RKRBH submitted 2025-09-10 cs.CV cs.AI

classification cs.CVcs.AI
keywords underwaterobjectdetectionlargevision-languagemodelssyntheticdataaugmentationDALL-E3Florence-2LoRAfine-tuningimageenhancementdomainshift
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is a structured review of underwater object detection (UOD), organized around a five-part taxonomy of challenges: image quality degradation, target-related issues, data-related problems, computational constraints, and detection-methodology limits. Its central argument is that existing UOD methods address parts of these challenges but not the whole, and that large vision-language models (LVLMs) are a promising but unproven direction. To support that argument it runs two case studies: adding 1,200 DALL-E 3 synthetic images to 7,600 real images changes YOLO11 results only slightly (mAP@50 from 0.793 to 0.796, recall from 0.714 to 0.736, precision from 0.805 to 0.780), while a LoRA-fine-tuned Florence-2 draws good bounding boxes yet hallucinates class names so badly that mAP and recall could not be computed. The review's contribution is the taxonomy plus a candid statement of what current LVLMs can and cannot do underwater.

What carries the argument

The carrying mechanism is the five-category taxonomy of UOD challenges, which structures the whole review and lets the authors match each challenge family to a solution family (enhancement and restoration, image synthesis, detection architectures, domain adaptation, and efficient fine-tuning). Two case-study pipelines carry the empirical weight: DALL-E 3 text-to-image and image-to-image generation followed by OpenCV enhancement and manual annotation, then YOLO11 training on the original versus combined datasets; and Florence-2 with LoRA applied to attention projection, linear, and convolution layers, fine-tuning 1,929,928 parameters (0.7076% of the model).

What would settle it

Retrain YOLO11 on the original and combined datasets many times with different random seeds and compare the spread of mAP@50 and recall; if the 0.003 and 0.022 gaps overlap across runs, the claimed synthetic-data benefit is not established. Separately, map Florence-2's misspelled class names to ground-truth labels and recompute mAP and recall; if the corrected metrics do not beat a no-fine-tuning baseline, the 'strong localization' claim loses its support.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes a systematic map of UOD difficulties and shows where the field stands. It claims that the five challenge categories—image quality degradation, target-related issues, data-related challenges, computational and processing constraints, and detection methodology limitations—are best treated together, because solutions aimed at one category (such as enhancement before detection) can interact with another (such as artifacts that hurt small-object localization). The case studies carry the review's main empirical claims: LVLM-generated synthetic data can augment scarce underwater datasets, with Table 8 showing mAP@50 rising from 0.793 to 0.796 and recall from 0.714 to 0.736 at a cost in precision (0.805 to 0.780); and parameter-efficient LoRA fine-tuning of Florence-2 enables strong bounding-box localization for small underwater objects while producing misspelled, ungrounded class names that make standard metrics meaningless. The paper therefore presents LVLMs as a direction with demonstrated localization promise and equally demonstrated output-reliability problems, not as a finished solution.

Load-bearing premise

The case-study conclusions assume that the small Table 8 improvements from adding synthetic images are real effects of the data rather than run-to-run noise, and that Florence-2's qualitative bounding boxes demonstrate localization strength despite hallucinated class labels.

Editorial extensions

If this is right

  • If LVLM synthetic data is used to augment small underwater datasets, the expected effect is modest: Table 8 shows mAP@50 moving from 0.793 to 0.796 and recall from 0.714 to 0.736, with precision dropping from 0.805 to 0.780.
  • LoRA fine-tuning of Florence-2 on 8,800 images can produce strong bounding-box localization for small objects while updating only 0.7076% of the model parameters, making LVLM adaptation computationally feasible.
  • Class-name hallucination is the binding constraint for LVLM-based UOD: as long as outputs like 'echinullop' and 'starchin' replace correct labels, standard mAP and recall cannot be computed and practical deployment is blocked.
  • Real-time UOD with LVLMs remains unresolved; the paper points to efficient fine-tuning, prompt tuning, and lightweight architectures as the next steps rather than a demonstrated result.
  • The five-category taxonomy gives a shared vocabulary for matching a specific underwater failure (turbidity, small objects, class imbalance, latency, bounding-box overlap) to the solution families reviewed.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One testable reading of Table 8 is that the synthetic images teach YOLO11 to propose more candidate boxes (higher recall) at the cost of false positives (lower precision); per-class precision-recall curves for the four target classes would show whether the gain is concentrated in classes that are rare in the real data.
  • Because only single training runs are reported, the 0.003 mAP@50 gain could easily be seed noise; repeating both training conditions with several seeds and holding the real-image count fixed would settle whether the synthetic set itself, rather than simply more training images, drives the change.
  • Florence-2's misspelled labels look like a lexical or tokenization failure rather than a visual grounding failure, since localization is reportedly strong; a constrained decoding or label-mapping postprocessor could recover usable metrics and clarify whether the bottleneck is semantic hallucination or vocabulary coverage.
  • The taxonomy could double as an evaluation rubric for future UOD papers, requiring authors to state which challenge categories a method attacks and which it leaves untouched, which would make cross-paper comparison less fragmented.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript is a structured literature review of underwater object detection (UOD). It categorizes challenges into five areas—image quality degradation, target-related issues, data-related challenges, computational and processing constraints, and detection methodology limitations—and surveys solutions from traditional image enhancement and restoration through modern CNN-, transformer-, and hybrid-based detectors, culminating in a discussion of large vision-language models (LVLMs). The review contributes two original case studies: (1) augmenting a real underwater dataset (7,600 images) with 1,200 DALL-E 3 synthetic images and evaluating YOLO11 on the combined 8,800-image set, and (2) fine-tuning Florence-2 with LoRA on the same combined data. The paper claims that synthetic augmentation modestly improves YOLO11 detection metrics and that the LoRA-fine-tuned Florence-2 demonstrates strong localization, while also openly acknowledging limitations including the small performance gap and hallucinated class names.

Significance. If the empirical claims hold, the paper would provide a useful reference point for applying generative LVLMs and parameter-efficient fine-tuning to data-scarce UOD. The literature survey is broad and the taxonomy is sensible, and the authors are commendably explicit about the limitations of their own case studies. The main significance is therefore as a structured survey with two preliminary feasibility pilot studies; the quantitative conclusions of those pilots are not yet established at the level the paper asserts.

major comments (4)
  1. [Section 5.1.3, Table 8] The central claim that adding 1,200 DALL-E 3 synthetic images improves YOLO11 detection is not supported by the reported numbers. The differences in mAP@50 (0.796 vs. 0.793) and mAP@50-95 (0.505 vs. 0.501) are within the typical single-run stochastic variation of YOLO training, while recall (0.736 vs. 0.714) and precision (0.780 vs. 0.805) move in opposite directions. No training seeds, variance estimates, confidence intervals, or significance tests are reported. In addition, the comparison is confounded: the combined condition changes both the data source and the dataset size (7,600 to 8,800 images), so any observed improvement could result from simply having more training images or altered class composition rather than from synthetic data specifically. To support the stated conclusion, the authors should report multiple seeds with mean and standard deviation and include an ablation that isolates the synthetic contribution at a matched dataset size (e.g., adding 1,200 real images or subsampling the combined set to the original size).
  2. [Section 5.1.4] The discussion explicitly concedes that 'the limited number of synthetic images (1,200) might not have been sufficient to blend effectively with the larger real dataset (7,600).' This admission directly undercuts the stronger statement in Section 5.1.3 that the results 'highlight the strength of synthetic data augmentation.' As written, the paper is internally inconsistent about the strength of its own evidence. The conclusion should either be tempered to reflect that no reliable improvement was demonstrated, or the experimental evidence must be strengthened to justify the stronger claim.
  3. [Section 5.2.3] The claim that the LoRA-fine-tuned Florence-2 'demonstrated strong localization capabilities' and 'excelled' at drawing bounding boxes is not supported by any quantitative evaluation. The paper itself states that hallucinated class names made it impossible to compute meaningful mAP or recall values. Without either a controlled qualitative protocol (e.g., a predefined set of images, independent human rating, or localization-only IoU computed independently of class labels), this positive characterization remains anecdotal. The authors should either provide such an evaluation or present the case study strictly as a feasibility demonstration with no claims about the strength of localization.
  4. [Section 5.1.4] The paper claims that the quality of the enhanced synthetic images was 'rigorously evaluated using metrics such as PSNR and SSIM,' but no PSNR or SSIM values are reported anywhere in the manuscript. This makes the quality claim unverifiable. The authors should report the actual metric values and describe the reference images and protocol used, or remove the claim.
minor comments (5)
  1. [Section 5.1, Fig. 6] The pipeline defines three paths—P1 (synthetic only), P2 (real only), and P3 (mixed)—but Section 5.1.3 only evaluates P2 versus P3. P1 is never used in the experiments; either evaluate the synthetic-only condition or remove it from the pipeline description to avoid confusion.
  2. [Section 3.3.1 and Fig. 5] The abbreviation for Detection Transformer is written as 'DETR' in most places but as 'DeTR' in at least one instance and in the Fig. 5 caption. Please standardize the spelling.
  3. [Appendix A.3, Eq. (A.3)] Equation (A.3) reads 'V = 255 · G / ||G||', which is confusing because G appears on both sides with different meanings. Please use distinct symbols for the raw Gaussian kernel and the normalized vignette mask, and clarify how the Gaussian kernel is computed from the image dimensions.
  4. [Table 1] For the URPC entry, the dataset column lists five separate counts in parentheses, one per year/version, but the 'Image Number' column is empty; this is inconsistent with other rows and should be clarified either by giving a single aggregate count or by splitting the row by version.
  5. [Throughout] There are minor typographical issues, including 'Moreso' for 'More so,' 'Incase' for 'In case,' and inconsistent capitalization of 'scallop' in Figure 8 and Table 7. A careful proofreading pass would resolve these.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the case-study conclusions are empirical comparisons, not derivations that reduce to their inputs.

full rationale

This paper is a structured review with two empirical case studies, not a derivation chain, so the circularity burden is low and no circular step is exhibited. The main claim comparing YOLO11 on the original RF100-based dataset versus the combined dataset (Table 8) is a direct training comparison; the combined condition does change both data source and dataset size, and the reported gains are small and unreplicated, but that is a statistical and experimental-confounding weakness, not a case where a prediction is equivalent to an input by construction. No parameter is fitted to a subset and then renamed as a prediction, and no equation in the paper reduces to another equation by definition. The Florence-2 case study asserts 'strong localization capabilities' from qualitative bounding-box examples while the paper itself states that hallucinated class names made mAP and recall uncomputable; this is anecdotal evidence and an internal tension, not definitional circularity. The references to the authors' own prior works ([14], [56], [78], [102]) are background citations in a literature review and are not used to justify the paper's empirical conclusions or to forbid alternative explanations, so they are not load-bearing self-citations. The PSNR/SSIM质量控制 claim is mentioned without reporting values, but those metrics are not used to derive the mAP comparison, so the absence of values is a reporting gap rather than a circular reduction. Overall, the honest finding is no significant circularity; the concerns belong to evidence quality and statistical rigor, not to circular reasoning.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

No new physical or model entities are introduced; the paper's novelty claims concern the application of existing generative models and PEFT methods. The main ledger entries are hand-set constants in the synthetic-image pipeline and unstated experimental hyperparameters.

free parameters (6)
  • Enhancement weights alpha, beta, gamma, delta = 0.9 each
    Hand-set blending weights in Appendix A.1 (Eqs. A.1, A.2, A.4, A.5) that control tint, turbidity, scattering, and particle effects; no sensitivity analysis is given.
  • Gaussian blur kernel size = 15x15, sigma not reported
    Chosen in Appendix A.3 (Eq. A.19) to soften synthetic images; the sigma value is not stated.
  • Tint color = (20, 50, 100)
    Arbitrary BGR tint used in Appendix A.1 Eq. A.1 to simulate underwater color shift.
  • LoRA rank and layer targeting = not stated; 1,929,928 trainable params, 0.7076% of total
    Hyperparameters in Section 5.2.2 determine the capacity of adaptation, but the rank is not reported.
  • Learning rate and epochs = lr=5e-6, 20 epochs
    Training choices in Section 5.2.2 are stated without a search or justification.
  • DALL-E 3 prompts = not released
    Prompt engineering is the main hand-tuned control over the synthetic data distribution in Section 5.1, but no prompts or seeds are provided.
assumptions (5)
  • standard math Gaussian convolution and LAB color transfer equations in Appendix A are standard and internally consistent.
    The mathematical expressions for enhancement in Appendix A follow textbook image processing definitions.
  • domain assumption The Roboflow100 (RF100) underwater subset is a reliable reference benchmark with correct labels.
    The case study uses RF100 as the real-data baseline in Section 5.1.2 and Table 7, trusting the benchmark's annotations and composition.
  • domain assumption DALL-E 3 images plus OpenCV effects are sufficiently close to the real underwater distribution for evaluating detection augmentation.
    The entire synthetic augmentation experiment in Section 5.1 assumes these images transfer to real underwater scenes without a quantitative domain-gap analysis.
  • domain assumption Single-run mAP values are stable estimates of model performance.
    Table 8 reports one training run per dataset condition with no variance or significance testing, yet the conclusion of improvement relies on these single numbers.
  • domain assumption The Florence-2 pretrained checkpoint 'Florence-2-base-ft' revision 'refs/pr/6' is a suitable starting point for LoRA adaptation to UOD.
    Section 5.2.2 selects this checkpoint without comparing alternatives or analyzing its pretraining distribution.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Structured Review of Underwater Object Detection Challenges and Solutions: From Traditional to Large Vision Language Models." pith.science (2026). https://pith.science/paper/YP4RKRBH

@misc{pith2026250908490,
  author       = {Pith},
  title        = {Pith review of: A Structured Review of Underwater Object Detection Challenges and Solutions: From Traditional to Large Vision Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YP4RKRBH}},
  note         = {Machine review of arXiv:2509.08490}
}
read the original abstract

Underwater object detection (UOD) is vital to diverse marine applications, including oceanographic research, underwater robotics, and marine conservation. However, UOD faces numerous challenges that compromise its performance. Over the years, various methods have been proposed to address these issues, but they often fail to fully capture the complexities of underwater environments. This review systematically categorizes UOD challenges into five key areas: Image quality degradation, target-related issues, data-related challenges, computational and processing constraints, and limitations in detection methodologies. To address these challenges, we analyze the progression from traditional image processing and object detection techniques to modern approaches. Additionally, we explore the potential of large vision-language models (LVLMs) in UOD, leveraging their multi-modal capabilities demonstrated in other domains. We also present case studies, including synthetic dataset generation using DALL-E 3 and fine-tuning Florence-2 LVLM for UOD. This review identifies three key insights: (i) Current UOD methods are insufficient to fully address challenges like image degradation and small object detection in dynamic underwater environments. (ii) Synthetic data generation using LVLMs shows potential for augmenting datasets but requires further refinement to ensure realism and applicability. (iii) LVLMs hold significant promise for UOD, but their real-time application remains under-explored, requiring further research on optimization techniques.

Figures

Figures reproduced from arXiv: 2509.08490 by the authors.

Figure 1
Figure 1. Taxonomy of UOD Challenges and Solutions [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Some of the Challenges Faced by UOD from DUO Dataset [17] [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The relationship Between Image Enhancement, Restoration, Image Synthesis, and Robust Object [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Multi-Scale Feature Extraction and Fusion in Object Detection [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]
Figure 5
Figure 5. Figure 5: The Pipelines of Various Object Detection Model Architectures [PITH_FULL_IMAGE:figures/full_fig_p022_5.png]
Figure 6
Figure 6. Figure 6: The Workflow Showing the effect of Synthetic data augmentation and Enhancement in Improving [PITH_FULL_IMAGE:figures/full_fig_p027_6.png]
Figure 7
Figure 7. Figure 7: DALL-E 3’s text-to-image and image-to-image generation paths produced clear and high￾resolution underwater images, however, the generated images appeared overly pristine and lacked the typical noise and disturbances found in real underwater scenes. To address this, ima…
Figure 7
Figure 7. Figure 7: Synthetic Data Generation Pipeline 5.1.1. Image Enhancement on DALL-E 3’s Generated Images In this work, image enhancement plays a critical role in bridging the gap between syn￾thetic underwater images and real-world underwater scenes. OpenCV was utilized to apply subt…
Figure 8
Figure 8. Figure 8: showcases sample synthetic underwater images with improved visual quality and [PITH_FULL_IMAGE:figures/full_fig_p029_8.png]
Figure 8
Figure 8. Figure 8: Sample Synthetic Generated Underwater Images with desired Target Object Classes (echinus, [PITH_FULL_IMAGE:figures/full_fig_p030_8.png]
Figure 9
Figure 9. Figure 9: YOLO11 inference results on underwater images after training with a combined dataset, demon [PITH_FULL_IMAGE:figures/full_fig_p034_9.png]
Figure 10
Figure 10. Figure 10: The Florence-2 Architecture with Highlighted Layers Where LoRA is Applied [PITH_FULL_IMAGE:figures/full_fig_p037_10.png]
Figure 11
Figure 11. Figure 11: Fine-tuned Florence 2 Detection results Understanding Hallucination in Florence-2: Hallucination [139],[140] refers to a model’s tendency to generate outputs that are not grounded in the input data. In the case of Florence-2, hallucination manifested as misspelled cla…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Why Domain Matters: Domain-Aware Benchmarking of Underwater Object Detection and Annotation Quality

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Physically grounded domain labels for underwater images expose large, consistent gaps in both human annotation quality and detector mAP that aggregate metrics conceal.

Reference graph

Works this paper leans on

168 extracted references · 39 canonical work pages · cited by 1 Pith paper

  1. [1]

    Underwater video techniques for observing coastal marine biodiversity: A review of sixty years of publications (1952–2012).Fish- eries Research, 154:44–62, June 2014

    Delphine Mallet and Dominique Pelletier. Underwater video techniques for observing coastal marine biodiversity: A review of sixty years of publications (1952–2012).Fish- eries Research, 154:44–62, June 2014. ISSN 01657836. doi: 10.1016/j.fishres.2014.01

  2. [2]

    Adaptive low-level control of autonomous underwater vehicles using deep reinforce- ment learning.Robotics and Autonomous Systems, 107:71–86, September 2018

    IgnacioCarlucho, MarianoDePaula, SenWang, YvanPetillot, andGerardoG.Acosta. Adaptive low-level control of autonomous underwater vehicles using deep reinforce- ment learning.Robotics and Autonomous Systems, 107:71–86, September 2018. ISSN 09218890. doi: 10.1016/j.robot.2018.05.016. URLhttps://linkinghub.elsevier. com/retrieve/pii/S0921889018301519

  3. [3]

    Dwivedy, and P.S

    Avilash Sahoo, Santosha K. Dwivedy, and P.S. Robi. Advancements in the field of autonomous underwater vehicle.Ocean Engineering, 181:145–160, June 2019. ISSN 00298018. doi: 10.1016/j.oceaneng.2019.04.011. URLhttps://linkinghub. elsevier.com/retrieve/pii/S0029801819301623. 42

  4. [4]

    J. Y. Chiang and Ying-Ching Chen. Underwater Image Enhancement by Wavelength Compensation and Dehazing.IEEE Transactions on Image Processing, 21(4):1756– 1769, April 2012. ISSN 1057-7149, 1941-0042. doi: 10.1109/TIP.2011.2179666. URL http://ieeexplore.ieee.org/document/6104148/

  5. [5]

    Research Challenges, Recent Advances, and Popular Datasets in Deep Learning-Based Underwater Marine Object Detection: A Review.Sensors, 23(4):1990, February 2023

    Meng Joo Er, Jie Chen, Yani Zhang, and Wenxiao Gao. Research Challenges, Recent Advances, and Popular Datasets in Deep Learning-Based Underwater Marine Object Detection: A Review.Sensors, 23(4):1990, February 2023. ISSN 1424-8220. doi: 10.3390/s23041990. URLhttps://www.mdpi.com/1424-8220/23/4/1990

  6. [6]

    Vision-Language Models for Vision Tasks: A Survey, 2023

    Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. Vision-Language Models for Vision Tasks: A Survey, 2023. URLhttps://arxiv.org/abs/2304.00685. Version Number: 2

  7. [7]

    The Revolution of Multimodal LargeLanguageModels: ASurvey, 2024

    Davide Caffagni, Federico Cocchi, Luca Barsellotti, Nicholas Moratelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. The Revolution of Multimodal LargeLanguageModels: ASurvey, 2024. URLhttps://arxiv.org/abs/2402.12451. Version Number: 2

  8. [8]

    Moniruzzaman, Syed Mohammed Shamsul Islam, Mohammed Bennamoun, and Paul Lavery

    Md. Moniruzzaman, Syed Mohammed Shamsul Islam, Mohammed Bennamoun, and Paul Lavery. Deep Learning on Underwater Marine Object Detection: A Sur- vey. In Jacques Blanc-Talon, Rudi Penne, Wilfried Philips, Dan Popescu, and Paul Scheunders, editors,Advanced Concepts for Intelligent Vision Systems, vol- ume 10617, pages 150–160. Springer International Publishi...

Show all 168 references
  1. [9]

    Review on deep learning tech- niques for marine object recognition: Architectures and algorithms.Control En- gineering Practice, 118:104458, January 2022

    Ning Wang, Yuanyuan Wang, and Meng Joo Er. Review on deep learning tech- niques for marine object recognition: Architectures and algorithms.Control En- gineering Practice, 118:104458, January 2022. ISSN 09670661. doi: 10.1016/j. 43 conengprac.2020.104458. URLhttps://linkinghub...

  2. [10]

    Parah, and G

    Sheezan Fayaz, Shabir A. Parah, and G. J. Qureshi. Underwater object detec- tion: architectures and algorithms – a comprehensive review.Multimedia Tools and Applications, 81(15):20871–20916, June 2022. ISSN 1380-7501, 1573-7721. doi: 10.1007/s11042-022-12502-1. URLhttps://link...

  3. [11]

    Underwater object detection and datasets: a survey.Intelligent Marine Technology and Systems, 2(1):9, March 2024

    Muwei Jian, Nan Yang, Chen Tao, Huixiang Zhi, and Hanjiang Luo. Underwater object detection and datasets: a survey.Intelligent Marine Technology and Systems, 2(1):9, March 2024. ISSN 2948-1953. doi: 10.1007/s44295-024-00023-6. URLhttps: //link.springer.com/10.1007/s44295-024-00023-6

  4. [12]

    Underwater Object Detection in the Era of Artificial Intelligence: Current, Challenge, and Future, 2024

    Long Chen, Yuzhi Huang, Junyu Dong, Qi Xu, Sam Kwong, Huimin Lu, Huchuan Lu, and Chongyi Li. Underwater Object Detection in the Era of Artificial Intelligence: Current, Challenge, and Future, 2024. URLhttps://arxiv.org/abs/2410.05577. Version Number: 1

  5. [13]

    Fouda, Dinh-Thuan Do, Abdulaziz Almaleh, Abdullah M

    Anwar Khan, Mostafa M. Fouda, Dinh-Thuan Do, Abdulaziz Almaleh, Abdullah M. Alqahtani, and Atiq Ur Rahman. Underwater Target Detection Using Deep Learning: Methodologies, Challenges, Applications, and Future Evolution.IEEE Access, 12: 12618–12635, 2024. ISSN 2169-3536. doi: 10...

  6. [14]

    A systematic review and analysis of deep learning-based underwater object de- tection.Neurocomputing, 527:204–232, March 2023

    Shubo Xu, Minghua Zhang, Wei Song, Haibin Mei, Qi He, and Antonio Liotta. A systematic review and analysis of deep learning-based underwater object de- tection.Neurocomputing, 527:204–232, March 2023. ISSN 09252312. doi: 10. 1016/j.neucom.2023.01.056. URLhttps://linkinghub.els...

  7. [15]

    Rethinking general underwater object detection: Datasets, challenges, and solutions.Neurocomputing, 517:243–256, January 2023

    Chenping Fu, Risheng Liu, Xin Fan, Puyang Chen, Hao Fu, Wanqi Yuan, Ming Zhu, 44 and Zhongxuan Luo. Rethinking general underwater object detection: Datasets, challenges, and solutions.Neurocomputing, 517:243–256, January 2023. ISSN 09252312. doi: 10.1016/j.neucom.2022.10.039. ...

  8. [16]

    Dipta Gomes, A. F. M. Saifuddin Saif, and Dip Nandi. Robust Underwater Object Detection with Autonomous Underwater Vehicle: A Comprehensive Study. InProceed- ings of the International Conference on Computing Advancements, pages 1–10, Dhaka Bangladesh, January 2020. ACM. ISBN 9...

  9. [17]

    A Dataset and Benchmark of Underwater Object Detection for Robot Picking

    Chongwei Liu, Haojie Li, Shuchang Wang, Ming Zhu, Dong Wang, Xin Fan, and Zhihui Wang. A Dataset and Benchmark of Underwater Object Detection for Robot Picking. In2021 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), pages 1–6, Shenzhen, China, July 2021. ...

  10. [18]

    Cherian, Eswaran Poovammal, Ninan Sajeeth Philip, Kadiyala Ramana, Saurabh Singh, and In-Ho Ra

    Aswathy K. Cherian, Eswaran Poovammal, Ninan Sajeeth Philip, Kadiyala Ramana, Saurabh Singh, and In-Ho Ra. Deep Learning Based Filtering Algorithm for Noise Removal in Underwater Images.Water, 13(19):2742, October 2021. ISSN 2073-4441. doi: 10.3390/w13192742. URLhttps://www.md...

  11. [19]

    URLhttps://linkinghub.elsevier.com/retrieve/pii/S0165783614000356

  12. [20]

    URLhttps://ieeexplore.ieee.org/ document/9455997/

    doi: 10.1109/ICMEW53276.2021.9455997. URLhttps://ieeexplore.ieee.org/ document/9455997/

  13. [21]

    Underwater object detection and temporal signal detection in turbid water using 3D-integral imaging and deep learning.Optics Express, 32(2):1789, January

    Rakesh Joshi, Kashif Usmani, Gokul Krishnan, Fletcher Blackmon, and Bahram Ja- vidi. Underwater object detection and temporal signal detection in turbid water using 3D-integral imaging and deep learning.Optics Express, 32(2):1789, January

  14. [22]

    Underwater Image Enhancement via Triple-Branch Dense Block and Generative Adversarial Network

    Peng Yang, Chunhua He, Shaojuan Luo, Tao Wang, and Heng Wu. Underwater Image Enhancement via Triple-Branch Dense Block and Generative Adversarial Network. Journal of Marine Science and Engineering, 11(6):1124, May 2023. ISSN 2077-1312. doi: 10.3390/jmse11061124. URLhttps://www...

  15. [23]

    A Survey of Target Detection and Recognition Methods in Underwater Turbid Areas.Applied 45 Sciences, 12(10):4898, May 2022

    Xin Yuan, Linxu Guo, Citong Luo, Xiaoteng Zhou, and Changli Yu. A Survey of Target Detection and Recognition Methods in Underwater Turbid Areas.Applied 45 Sciences, 12(10):4898, May 2022. ISSN 2076-3417. doi: 10.3390/app12104898. URL https://www.mdpi.com/2076-3417/12/10/4898

  16. [24]

    Adeoluwa, Carson D

    Oladipupo O. Adeoluwa, Carson D. Moseley, Seongsin M. Kim, Patrick Kung, and Sevgi Z. Gurbuz. Evaluation of Laser Image Enhancement and Restoration for Un- derwater Object Recognition.IEEE Sensors Journal, 23(21):26136–26153, November

  17. [25]

    MAT: Motion-aware multi-object tracking.Neurocomputing, 476:75– 86, March 2022

    Shoudong Han, Piao Huang, Hongwei Wang, En Yu, Donghaisheng Liu, and Xi- aofeng Pan. MAT: Motion-aware multi-object tracking.Neurocomputing, 476:75– 86, March 2022. ISSN 09252312. doi: 10.1016/j.neucom.2021.12.104. URLhttps: //linkinghub.elsevier.com/retrieve/pii/S0925231221019627

  18. [26]

    Artificial Intelligence Based Object Detection and Tracking for a Small Underwater Robot.Processes, 11(2):312, January 2023

    Min-Fan Ricky Lee and Ying-Chu Chen. Artificial Intelligence Based Object Detection and Tracking for a Small Underwater Robot.Processes, 11(2):312, January 2023. ISSN 46 2227-9717. doi: 10.3390/pr11020312. URLhttps://www.mdpi.com/2227-9717/11/ 2/312

  19. [27]

    Underwater Small Target Detection Based on Deformable Convolutional Pyra- mid

    Shuhan Qi, Jianjun Du, Mingyan Wu, Hong Yi, Linlin Tang, Tao Qian, and Xuan Wang. Underwater Small Target Detection Based on Deformable Convolutional Pyra- mid. InICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2784–2...

  20. [28]

    Underwater Object Detection Method Based on Im- proved Faster RCNN.Applied Sciences, 13(4):2746, February 2023

    Hao Wang and Nanfeng Xiao. Underwater Object Detection Method Based on Im- proved Faster RCNN.Applied Sciences, 13(4):2746, February 2023. ISSN 2076-3417. doi: 10.3390/app13042746. URLhttps://www.mdpi.com/2076-3417/13/4/2746

  21. [29]

    Syn2Real Domain Generalization for Underwater Mine-like Object Detection Using Side-Scan Sonar, 2024

    Aayush Agrawal, Aniruddh Sikdar, Rajini Makam, Suresh Sundaram, Suresh Kumar Besai, and Mahesh Gopi. Syn2Real Domain Generalization for Underwater Mine-like Object Detection Using Side-Scan Sonar, 2024. URLhttps://arxiv.org/abs/2410. 12953. Version Number: 1

  22. [30]

    JosephL.Walker, ZhengZeng, ChengchenL.Wu, JulesS.Jaffe, KaitlinE.Frasier, and Stuart S. Sandin. Underwater Object Detection Under Domain Shift.IEEE Journal of Oceanic Engineering, 49(4):1209–1219, October 2024. ISSN 0364-9059, 1558-1691, 2373-7786. doi: 10.1109/JOE.2024.342545...

  23. [31]

    Yongcan Yu, Jianhu Zhao, Chao Huang, and Xi Zhao. Treat Noise as Domain Shift: Noise Feature Disentanglement for Underwater Perception and Maritime Surveys in Side-Scan Sonar Images.IEEE Transactions on Geoscience and Remote Sensing, 61: 1–15, 2023. ISSN 0196-2892, 1558-0644. ...

  24. [32]

    ADOD: Adaptive Domain-Aware Object Detection with Residual Attention for Underwater Environments

    Lyes Saad Saoud, Zhenwei Niu, Atif Sultan, Lakmal Seneviratne, and Irfan Hus- sain. ADOD: Adaptive Domain-Aware Object Detection with Residual Attention for Underwater Environments. In2023 21st International Conference on Advanced Robotics (ICAR), pages 633–638, Abu Dhabi, Uni...

  25. [33]

    ISBN 9798350342291

    IEEE. ISBN 9798350342291. doi: 10.1109/ICAR58858.2023.10436502. URL https://ieeexplore.ieee.org/document/10436502/

  26. [34]

    MMDetection: Open MMLab Detection Toolbox and Benchmark, 2019

    Kai Chen, Jiaqi Wang, and et al Pang. MMDetection: Open MMLab Detection Toolbox and Benchmark, 2019. URLhttps://arxiv.org/abs/1906.07155. Version Number: 1

  27. [35]

    Moeslund

    Malte Pedersen, Joakim Bruslund Haurum, Rikke Gade, and Thomas B. Moeslund. Detection of marine animals in a new underwater dataset with varying visibility. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 18–26, Long ...

  28. [36]

    Roboflow 100: A Rich, Multi-Domain Object Detection Benchmark,

    Floriana Ciaglia, Francesco Saverio Zuppichini, Paul Guerrie, Mark McQuade, and Jacob Solawetz. Roboflow 100: A Rich, Multi-Domain Object Detection Benchmark,

  29. [37]

    48 2012, Tsukuba International Congress Center, Tsukuba Science City, Japan

    International Association for Pattern Recognition and IEEE Computer Society, edi- tors.21st International Conference on Pattern Recognition (ICPR), 2012: 11 - 15 Nov. 48 2012, Tsukuba International Congress Center, Tsukuba Science City, Japan. IEEE, Piscataway, NJ, 2012. ISBN ...

  30. [38]

    A New Dataset, Poisson GAN and AquaNet for Underwater Object Grabbing.IEEE Transactions on Circuits and Systems for Video Technology, 32(5):2831–2844, May 2022

    Chongwei Liu, Zhihui Wang, Shijie Wang, Tao Tang, Yulong Tao, Caifei Yang, Haojie 47 Li, Xing Liu, and Xin Fan. A New Dataset, Poisson GAN and AquaNet for Underwater Object Grabbing.IEEE Transactions on Circuits and Systems for Video Technology, 32(5):2831–2844, May 2022. ISSN...

  31. [39]

    Underwater Species Detection using Channel Sharpening Attention

    Lihao Jiang, Yi Wang, Qi Jia, Shengwei Xu, Yu Liu, Xin Fan, Haojie Li, Risheng Liu, Xinwei Xue, and Ruili Wang. Underwater Species Detection using Channel Sharpening Attention. InProceedings of the 29th ACM International Conference on Multimedia, pages 4259–4267, Virtual Event...

  32. [40]

    Improving GAN-based Domain Adaptation for Object Detection

    Maximilian Menke, Thomas Wenzel, and Andreas Schwung. Improving GAN-based Domain Adaptation for Object Detection. In2022 IEEE 25th International Confer- ence on Intelligent Transportation Systems (ITSC), pages 3880–3885, Macau, China, October 2022. IEEE. ISBN 978-1-66546-880-0...

  33. [41]

    Research on Underwater Object Detection Based on Improved YOLOv4

    Wang Hao and Nangfeng Xiao. Research on Underwater Object Detection Based on Improved YOLOv4. In2021 8th International Conference on Information, Cybernet- ics, and Computational Social Systems (ICCSS), pages 166–171, Beijing, China, De- cember2021.IEEE. ISBN978-1-66540-245-3....

  34. [42]

    Underwater Object Detection Algo- rithm Based on Improved YOLOv7-tiny

    Jing Ling, Can Zhang, and Dengrong Du. Underwater Object Detection Algo- rithm Based on Improved YOLOv7-tiny. In2023 IEEE 11th International Confer- ence on Computer Science and Network Technology (ICCSNT), pages 28–31, Dalian, China, October 2023. IEEE. ISBN 9798350311594. do...

  35. [43]

    WildFish: A Large Benchmark for Fish Recognition in the Wild

    Peiqin Zhuang, Yali Wang, and Yu Qiao. WildFish: A Large Benchmark for Fish Recognition in the Wild. InProceedings of the 26th ACM international conference on Multimedia, pages 1301–1309, Seoul Republic of Korea, October 2018. ACM. ISBN 978-1-4503-5665-7. doi: 10.1145/3240508....

  36. [44]

    Robust Bounding Box Regression for Small Object Detection

    Ziqi Guo, Chu He, Lian Zhou, Qingyi Zhang, and Shilei Sun. Robust Bounding Box Regression for Small Object Detection. In2023 IEEE International Conference on Image Processing (ICIP), pages 2290–2294, Kuala Lumpur, Malaysia, October

  37. [45]

    An Improved YOLO Al- gorithm for Fast and Accurate Underwater Object Detection.Symmetry, 14(8): 1669, August 2022

    Shijia Zhao, Jiachun Zheng, Shidan Sun, and Lei Zhang. An Improved YOLO Al- gorithm for Fast and Accurate Underwater Object Detection.Symmetry, 14(8): 1669, August 2022. ISSN 2073-8994. doi: 10.3390/sym14081669. URLhttps: //www.mdpi.com/2073-8994/14/8/1669

  38. [47]

    IDA-UIE: An Iterative Framework for Deep Network-based Degradation Aware Underwater Image Enhancement, 2024

    Pranjali Singh and Prithwijit Guha. IDA-UIE: An Iterative Framework for Deep Network-based Degradation Aware Underwater Image Enhancement, 2024. URL https://arxiv.org/abs/2406.18628. Version Number: 1

  39. [48]

    Ancuti, Cosmin Ancuti, Christophe De Vleeschouwer, and Philippe Bekaert

    Codruta O. Ancuti, Cosmin Ancuti, Christophe De Vleeschouwer, and Philippe Bekaert. Color Balance and Fusion for Underwater Image Enhancement.IEEE Transactions on Image Processing, 27(1):379–393, January 2018. ISSN 1057-7149, 1941-0042. doi: 10.1109/TIP.2017.2759252. URLhttps:...

  40. [49]

    Color correction and adaptive contrast enhancement for underwater image enhance- ment.Computers & Electrical Engineering, 91:106981, May 2021

    Weidong Zhang, Xipeng Pan, Xiwang Xie, Lingqiao Li, Zimin Wang, and Chu Han. Color correction and adaptive contrast enhancement for underwater image enhance- ment.Computers & Electrical Engineering, 91:106981, May 2021. ISSN 00457906. doi: 10.1016/j.compeleceng.2021.106981. UR...

  41. [50]

    Scale-aware feature pyramid architecture for marine object detection.Neural Computing and Applications, 33(8):3637–3653, April 2021

    Fengqiang Xu, Huibing Wang, Jinjia Peng, and Xianping Fu. Scale-aware feature pyramid architecture for marine object detection.Neural Computing and Applications, 33(8):3637–3653, April 2021. ISSN 0941-0643, 1433-3058. doi: 10.1007/s00521-020-05217-7. URLhttps://link.springer.c...

  42. [51]

    Underwater image enhancement based on colour correction and fusion.IET Image Processing, 15(11):2591–2603, Septem- ber 2021

    Daqi Zhu, Zhiqiang Liu, and Youmin Zhang. Underwater image enhancement based on colour correction and fusion.IET Image Processing, 15(11):2591–2603, Septem- ber 2021. ISSN 1751-9659, 1751-9667. doi: 10.1049/ipr2.12247. URLhttps: //onlinelibrary.wiley.com/doi/10.1049/ipr2.12247

  43. [52]

    ISBN 978-1-72819-835-4

    IEEE. ISBN 978-1-72819-835-4. doi: 10.1109/ICIP49359.2023.10222753. URL https://ieeexplore.ieee.org/document/10222753/

  44. [53]

    Algorithms for improving the qual- ity of underwater optical images: A comprehensive review.Signal Processing, 219: 109408, June 2024

    Xuecheng Shuang, Jin Zhang, and Yu Tian. Algorithms for improving the qual- ity of underwater optical images: A comprehensive review.Signal Processing, 219: 109408, June 2024. ISSN 01651684. doi: 10.1016/j.sigpro.2024.109408. URL https://linkinghub.elsevier.com/retrieve/pii/S0...

  45. [54]

    DAE-GAN: UnderwaterImage Super-Resolution Based on Symmetric Degradation Attention Enhanced Generative Adversarial Network.Symmetry, 16(5):588, May 2024

    MiaoweiGao, ZhongguoLi, QiWang, andWenbinFan. DAE-GAN: UnderwaterImage Super-Resolution Based on Symmetric Degradation Attention Enhanced Generative Adversarial Network.Symmetry, 16(5):588, May 2024. ISSN 2073-8994. doi: 10.3390/ sym16050588. URLhttps://www.mdpi.com/2073-8994/16/5/588

  46. [55]

    A Pixel Distribution Remapping and Multi-Prior Retinex Variational Model for Underwater Image Enhancement.IEEE Transactions on Multimedia, 26:7838–7849, 2024

    Jingchun Zhou, Shiyin Wang, Zifan Lin, Qiuping Jiang, and Ferdous Sohel. A Pixel Distribution Remapping and Multi-Prior Retinex Variational Model for Underwater Image Enhancement.IEEE Transactions on Multimedia, 26:7838–7849, 2024. ISSN 1520-9210, 1941-0077. doi: 10.1109/TMM.2...

  47. [56]

    From shallow sea to deep sea: research progress in underwater image restora- tion.Frontiers in Marine Science, 10:1163831, May 2023

    Wei Song, Yaling Liu, Dongmei Huang, Bing Zhang, Zhihao Shen, and Huifang Xu. From shallow sea to deep sea: research progress in underwater image restora- tion.Frontiers in Marine Science, 10:1163831, May 2023. ISSN 2296-7745. doi: 10.3389/fmars.2023.1163831. URLhttps://www.fr...

  48. [57]

    Yan-Tsung Peng, Xiangyun Zhao, and Pamela C. Cosman. Single underwater image enhancement using depth estimation based on blurriness. In2015 IEEE International Conference on Image Processing (ICIP), pages 4952–4956, Quebec City, QC, Septem- ber 2015. IEEE. ISBN 978-1-4799-8339-...

  49. [58]

    Yi-Ning Fan, Geng-Kun Wu, Jia-Zheng Han, Bei-Ping Zhang, and Jie Xu. Inno- vative underwater image enhancement algorithm: Combined application of adaptive white balance color compensation and pyramid image fusion to submarine algal mi- croscopy.Image and Vision Computing, 156:...

  50. [59]

    Polarimetric underwater image recovery for color image with crosstalk compensation.Optics and Lasers in Engineering, 124: 105833, January 2020

    Tiegen Liu, Zijian Guan, Xiaobo Li, Zhenzhou Cheng, Yingdong Han, Jingyu Yang, Kun Li, Junying Zhao, and Haofeng Hu. Polarimetric underwater image recovery for color image with crosstalk compensation.Optics and Lasers in Engineering, 124: 105833, January 2020. ISSN 01438166. d...

  51. [60]

    Un- derwater image restoration based on progressive guidance.Signal Processing, 223: 109569, October 2024

    Jianghe Zhang, Weiling Chen, Zuxin Lin, Hongan Wei, and Tiesong Zhao. Un- derwater image restoration based on progressive guidance.Signal Processing, 223: 109569, October 2024. ISSN 01651684. doi: 10.1016/j.sigpro.2024.109569. URL https://linkinghub.elsevier.com/retrieve/pii/S...

  52. [61]

    Guojia Hou, Nan Li, Peixian Zhuang, Kunqian Li, Haihan Sun, and Chongyi Li. Non-Uniform Illumination Underwater Image Restoration via Illumination Channel Sparsity Prior.IEEE Transactions on Circuits and Systems for Video Technology, 34 (2):799–814, February 2024. ISSN 1051-82...

  53. [62]

    UnitModule: A lightweight jointimageenhancementmoduleforunderwaterobjectdetection.Pattern Recognition, 51 151:110435, July 2024

    Zhuoyan Liu, Bo Wang, Ye Li, Jiaxian He, and Yunfeng Li. UnitModule: A lightweight jointimageenhancementmoduleforunderwaterobjectdetection.Pattern Recognition, 51 151:110435, July 2024. ISSN 00313203. doi: 10.1016/j.patcog.2024.110435. URL https://linkinghub.elsevier.com/retri...

  54. [63]

    Image descat- tering and absorption compensation in underwater polarimetric imaging.Optics and Lasers in Engineering, 132:106115, September 2020

    Xianping Fu, Zheng Liang, Xueyan Ding, Xinyue Yu, and Yafei Wang. Image descat- tering and absorption compensation in underwater polarimetric imaging.Optics and Lasers in Engineering, 132:106115, September 2020. ISSN 01438166. doi: 10.1016/ j.optlaseng.2020.106115. URLhttps://...

  55. [64]

    Range-intensity-profile prior dehazing method for underwater range-gated imaging.Optics Express, 29(5):7630, March 2021

    Minmin Wang, Xinwei Wang, Yue Zhang, Liang Sun, Pingshun Lei, Yuqing Yang, Jianan Chen, Jun He, and Yan Zhou. Range-intensity-profile prior dehazing method for underwater range-gated imaging.Optics Express, 29(5):7630, March 2021. ISSN 1094-4087. doi: 10.1364/OE.417131. URLhtt...

  56. [65]

    GPLM: Enhancing underwater images with Global Pyramid Linear Modulation.Image and Vision Computing, 154:105361, 53 February 2025

    Jinxin Shao, Haosu Zhang, and Jianming Miao. GPLM: Enhancing underwater images with Global Pyramid Linear Modulation.Image and Vision Computing, 154:105361, 53 February 2025. ISSN 0262-8856. doi: 10.1016/j.imavis.2024.105361. URLhttps:// linkinghub.elsevier.com/retrieve/pii/S0...

  57. [66]

    Rajasekar, A

    M. Rajasekar, A. Celine Kavida, and M. Anto Bennet. A pattern analysis based un- derwater video segmentation system for target object detection.Multidimensional Systems and Signal Processing, 31(4):1579–1602, October 2020. ISSN 0923-6082, 1573-0824. doi: 10.1007/s11045-020-007...

  58. [67]

    Adaptive histogram fusion-based colour restoration and enhancement for underwater images.International Journal of Security and Networks, 16(1):49, 2021

    Jingchun Zhou, Dehuan Zhang, and Weishi Zhang. Adaptive histogram fusion-based colour restoration and enhancement for underwater images.International Journal of Security and Networks, 16(1):49, 2021. ISSN 1747-8405, 1747-8413. doi: 10.1504/ IJSN.2021.112848. URLhttp://www.inde...

  59. [68]

    Color image simulation for underwater optics.Applied Optics, 51(23):5633, August 2012

    Matthieu Boffety, Frédéric Galland, and Anne-Gaëlle Allais. Color image simulation for underwater optics.Applied Optics, 51(23):5633, August 2012. ISSN 1559-128X, 2155-3165. doi: 10.1364/AO.51.005633. URLhttps://opg.optica.org/abstract. cfm?URI=ao-51-23-5633

  60. [69]

    Single under- water image restoration by blue-green channels dehazing and red channel correction

    Chongyi Li, Jichang Quo, Yanwei Pang, Shanji Chen, and Jian Wang. Single under- water image restoration by blue-green channels dehazing and red channel correction. In2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1731–1735, Shangh...

  61. [70]

    UnderwaterImageRestoration using Deep Networks to Estimate Background Light and Scene Depth

    KemingCao, Yan-TsungPeng, andPamelaC.Cosman. UnderwaterImageRestoration using Deep Networks to Estimate Background Light and Scene Depth. In2018 IEEE Southwest Symposium on Image Analysis and Interpretation (SSIAI), pages 1–4, Las Vegas, NV, April 2018. IEEE. ISBN 978-1-5386-6...

  62. [71]

    Kingma and Max Welling

    Diederik P. Kingma and Max Welling. An Introduction to Variational Autoencoders. Foundations and Trends®in Machine Learning, 12(4):307–392, 2019. ISSN 1935- 8237, 1935-8245. doi: 10.1561/2200000056. URLhttp://www.nowpublishers.com/ article/Details/MAL-056

  63. [72]

    Generative Adversarial Networks, 2022

    Gilad Cohen and Raja Giryes. Generative Adversarial Networks, 2022. URLhttps: //arxiv.org/abs/2203.00667. Version Number: 1

  64. [73]

    Diffusion Models Beat GANs on Image Synthesis,

    Prafulla Dhariwal and Alex Nichol. Diffusion Models Beat GANs on Image Synthesis,

  65. [74]

    Underwater Image Restoration and Enhancement Based on a Fusion Algorithm With Color Balance, Contrast Optimiza- tion, and Histogram Stretching.IEEE Access, 9:31792–31804, 2021

    Weilin Luo, Shunqiang Duan, and Jiwen Zheng. Underwater Image Restoration and Enhancement Based on a Fusion Algorithm With Color Balance, Contrast Optimiza- tion, and Histogram Stretching.IEEE Access, 9:31792–31804, 2021. ISSN 2169-

  66. [75]

    Underwater scene prior inspired deep underwater image and video enhancement.Pattern Recognition, 98:107038, February

    Chongyi Li, Saeed Anwar, and Fatih Porikli. Underwater scene prior inspired deep underwater image and video enhancement.Pattern Recognition, 98:107038, February

  67. [76]

    A Novel Underwa- ter Image Synthesis Method Based on a Pixel-Level Self-Supervised Training Strat- egy

    Zhiheng Wu, Zhengxing Wu, Yue Lu, Jian Wang, and Junzhi Yu. A Novel Underwa- ter Image Synthesis Method Based on a Pixel-Level Self-Supervised Training Strat- egy. In2021 IEEE International Conference on Real-time Computing and Robotics (RCAR), pages 1254–1259, Xining, China, ...

  68. [77]

    A Perception-Aware Decomposition and Fusion Framework for Underwater Image En- hancement.IEEE Transactions on Circuits and Systems for Video Technology, 33 (3):988–1002, March 2023

    Yaozu Kang, Qiuping Jiang, Chongyi Li, Wenqi Ren, Hantao Liu, and Pengjun Wang. A Perception-Aware Decomposition and Fusion Framework for Underwater Image En- hancement.IEEE Transactions on Circuits and Systems for Video Technology, 33 (3):988–1002, March 2023. ISSN 1051-8215,...

  69. [78]

    Jingchun Zhou, Jiaming Sun, Chongyi Li, Qiuping Jiang, Man Zhou, Kin-Man Lam, Weishi Zhang, and Xianping Fu. HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image En- hancement.International Journal of Computer Vision, 1...

  70. [79]

    An underwater image enhancement method based on multi-scale layer decomposition and fusion.Signal Processing, 227:109690, February

    Jie Yang and Jun Wang. An underwater image enhancement method based on multi-scale layer decomposition and fusion.Signal Processing, 227:109690, February

  71. [80]

    An improved attention mechanism based YOLOv4 for small target detection at sea

    Ju He, Jianfeng Chen, Jintao Xu, and Muhammad Saad Ayub. An improved attention mechanism based YOLOv4 for small target detection at sea. InProceedings of the 15th International Conference on Digital Image Processing, pages 1–8, Nanjing China, May 2023. ACM. ISBN 9798400708237....

  72. [81]

    YOLOv7-CHS: An Emerging Model for Underwater Object Detection.Journal of Marine Science and Engineering, 11(10):1949, October 2023

    Liang Zhao, Qing Yun, Fucai Yuan, Xu Ren, Junwei Jin, and Xianchao Zhu. YOLOv7-CHS: An Emerging Model for Underwater Object Detection.Journal of Marine Science and Engineering, 11(10):1949, October 2023. ISSN 2077-1312. doi: 10.3390/jmse11101949. URLhttps://www.mdpi.com/2077-1...

  73. [82]

    Dynamic YOLO for small underwater object de- tection.Artificial Intelligence Review, 57(7):165, June 2024

    Jie Chen and Meng Joo Er. Dynamic YOLO for small underwater object de- tection.Artificial Intelligence Review, 57(7):165, June 2024. ISSN 1573-7462. doi: 10.1007/s10462-024-10788-1. URLhttps://link.springer.com/10.1007/ s10462-024-10788-1

  74. [83]

    SFDet: spatial to frequency atten- tion for small-object detection in underwater images.Journal of Elec- tronic Imaging, 33(02):023057–023057, April 2024

    Dazhi Chen and Gang Gou. SFDet: spatial to frequency atten- tion for small-object detection in underwater images.Journal of Elec- tronic Imaging, 33(02):023057–023057, April 2024. ISSN 1017-9909. doi: 10.1117/1.JEI.33.2.023057. URLhttps://www.spiedigitallibrary.org/ 56 journal...

  75. [84]

    PE- Transformer: Path enhanced transformer for improving underwater object detection

    Jinxiong Gao, Yonghui Zhang, Xu Geng, Hao Tang, and Uzair Aslam Bhatti. PE- Transformer: Path enhanced transformer for improving underwater object detection. Expert Systems with Applications, 246:123253, July 2024. ISSN 09574174. doi: 10. 1016/j.eswa.2024.123253. URLhttps://li...

  76. [85]

    POSEIDON: A Data Augmentation Tool for Small Object Detection Datasets in Maritime Environments.Sensors, 23(7):3691, April 2023

    Pablo Ruiz-Ponce, David Ortiz-Perez, Jose Garcia-Rodriguez, and Benjamin Kiefer. POSEIDON: A Data Augmentation Tool for Small Object Detection Datasets in Maritime Environments.Sensors, 23(7):3691, April 2023. ISSN 1424-8220. doi: 10.3390/s23073691. URLhttps://www.mdpi.com/142...

  77. [86]

    Mul- tiple information perception-based attention in YOLO for underwater object detec- tion.The Visual Computer, 40(3):1415–1438, March 2024

    Xin Shen, Huibing Wang, Tianxiang Cui, Zhicheng Guo, and Xianping Fu. Mul- tiple information perception-based attention in YOLO for underwater object detec- tion.The Visual Computer, 40(3):1415–1438, March 2024. ISSN 0178-2789, 1432-

  78. [87]

    Boosting R-CNN: Reweighting R-CNN Samples by RPN’s Error for Underwater Object Detection, 2022

    Pinhao Song, Pengteng Li, Linhui Dai, Tao Wang, and Zhan Chen. Boosting R-CNN: Reweighting R-CNN Samples by RPN’s Error for Underwater Object Detection, 2022. URLhttps://arxiv.org/abs/2206.13728. Version Number: 3

  79. [88]

    DeepSeaNet: Improving Underwater Object Detection using Effi- cientDet

    Sanyam Jain. DeepSeaNet: Improving Underwater Object Detection using Effi- cientDet. In2024 4th International Conference on Applied Artificial Intelligence (ICAPAI), pages 1–11, Halden, Norway, April 2024. IEEE. ISBN 9798350349764. doi: 10.1109/ICAPAI61893.2024.10541265. URLht...

  80. [89]

    URLhttps://ieeexplore.ieee.org/ document/9517333/

    doi: 10.1109/RCAR52367.2021.9517333. URLhttps://ieeexplore.ieee.org/ document/9517333/

  81. [90]

    Physics-Inspired Synthesized Underwater Image Dataset, 2024

    Reina Kaneko, Takumi Ueda, Hiroshi Higashi, and Yuichi Tanaka. Physics-Inspired Synthesized Underwater Image Dataset, 2024. URLhttps://arxiv.org/abs/2404. 03998. Version Number: 2. 55

  82. [91]

    Lightweight Underwater Object Detection Based on YOLO v4 and Multi-Scale Attentional Fea- ture Fusion.Remote Sensing, 13(22):4706, November 2021

    Minghua Zhang, Shubo Xu, Wei Song, Qi He, and Quanmiao Wei. Lightweight Underwater Object Detection Based on YOLO v4 and Multi-Scale Attentional Fea- ture Fusion.Remote Sensing, 13(22):4706, November 2021. ISSN 2072-4292. doi: 10.3390/rs13224706. URLhttps://www.mdpi.com/2072-4...

  83. [92]

    Novel Dynamic Feature Fusion Stragegy for Detection of Small Underwater Marine Ob- ject

    Jie Chen, Meng Joo Er, Yani Zhang, Wenxiao Gao, and Jianguo Wu. Novel Dynamic Feature Fusion Stragegy for Detection of Small Underwater Marine Ob- ject. In2022 5th International Conference on Intelligent Autonomous Systems (ICoIAS), pages 24–30, Dalian, China, September 2022. ...

  84. [93]

    Manimurugan, C

    S. Manimurugan, C. Narmatha, Majed M. Aborokbah, Naveen Chilamkurti, Sub- ramaniam Ganesan, Rajendran Thavasimuthu, P. Karthikeyan, and M Ammad Ud- din. HLASwin-T-ACoat-Net Based Underwater Object Detection.IEEE Access, 12: 32200–32217, 2024. ISSN 2169-3536. doi: 10.1109/ACCES...

  85. [94]

    Bounding Box Repairing Algorithm for Underwater Object Detection Based on IoU Optimization

    Bingchuan Chen, Lei Ma, and Jinmeng Wu. Bounding Box Repairing Algorithm for Underwater Object Detection Based on IoU Optimization. In2020 7th Interna- tional Conference on Information Science and Control Engineering (ICISCE), pages 369–373, Changsha, China, December 2020. IEE...

  86. [95]

    Underwater Biological Detection Algorithm Based on Improved Faster-RCNN.Water, 13(17):2420, September 2021

    Pengfei Shi, Xiwang Xu, Jianjun Ni, Yuanxue Xin, Weisheng Huang, and Song Han. Underwater Biological Detection Algorithm Based on Improved Faster-RCNN.Water, 13(17):2420, September 2021. ISSN 2073-4441. doi: 10.3390/w13172420. URLhttps: //www.mdpi.com/2073-4441/13/17/2420

  87. [96]

    Improved YOLOv8 Algorithm for Water Surface Object De- tection.Sensors, 24(15):5059, August 2024

    Jie Wang and Hong Zhao. Improved YOLOv8 Algorithm for Water Surface Object De- tection.Sensors, 24(15):5059, August 2024. ISSN 1424-8220. doi: 10.3390/s24155059. URLhttps://www.mdpi.com/1424-8220/24/15/5059

  88. [97]

    FBDPN: CNN- Transformer hybrid feature boosting and differential pyramid network for underwater object detection.Expert Systems with Applications, 256:124978, December 2024

    Xun Ji, Shijie Chen, Li-Ying Hao, Jingchun Zhou, and Long Chen. FBDPN: CNN- Transformer hybrid feature boosting and differential pyramid network for underwater object detection.Expert Systems with Applications, 256:124978, December 2024. ISSN 09574174. doi: 10.1016/j.eswa.2024...

  89. [98]

    Edge- guided representation learning for underwater object detection.CAAI Transactions on Intelligence Technology, 9(5):1078–1091, October 2024

    Linhui Dai, Hong Liu, Pinhao Song, Hao Tang, Runwei Ding, and Shengquan Li. Edge- guided representation learning for underwater object detection.CAAI Transactions on Intelligence Technology, 9(5):1078–1091, October 2024. ISSN 2468-2322, 2468-2322. doi: 10.1049/cit2.12325. URLh...

  90. [99]

    Zhuo Wang, Haojie Chen, Hongde Qin, and Qin Chen. Self-Supervised Pre-Training Joint Framework: Assisting Lightweight Detection Network for Underwater Object Detection.Journal of Marine Science and Engineering, 11(3):604, March 2023. ISSN 2077-1312. doi: 10.3390/jmse11030604. ...

  91. [100]

    Underwater Object Detection Using TC-YOLO with Attention Mechanisms.Sensors, 23(5):2567, February 2023

    Kun Liu, Lei Peng, and Shanran Tang. Underwater Object Detection Using TC-YOLO with Attention Mechanisms.Sensors, 23(5):2567, February 2023. ISSN 1424-8220. doi: 10.3390/s23052567. URLhttps://www.mdpi.com/1424-8220/23/5/2567. 59

  92. [101]

    Lightweight enhanced YOLOv8n underwater object detection network for low light envi- ronments.Scientific Reports, 14(1):27922, November 2024

    Jifeng Ding, Junquan Hu, Jiayuan Lin, and Xiaotong Zhang. Lightweight enhanced YOLOv8n underwater object detection network for low light envi- ronments.Scientific Reports, 14(1):27922, November 2024. ISSN 2045-2322. doi: 10.1038/s41598-024-79211-7. URLhttps://www.nature.com/ar...

  93. [102]

    Effi- cient Small-Object Detection in Underwater Images Using the Enhanced YOLOv8 Network.Applied Sciences, 14(3):1095, January 2024

    Minghua Zhang, Zhihua Wang, Wei Song, Danfeng Zhao, and Huijuan Zhao. Effi- cient Small-Object Detection in Underwater Images Using the Enhanced YOLOv8 Network.Applied Sciences, 14(3):1095, January 2024. ISSN 2076-3417. doi: 10.3390/app14031095. URLhttps://www.mdpi.com/2076-34...

  94. [103]

    Underwater Object Detection Based on Image Enhancement and Multi-Branch Structure

    Zheng Cui, Xian Wang, Hao Duan, Sen Wang, Chunxi Yang, and Jing Na. Underwater Object Detection Based on Image Enhancement and Multi-Branch Structure. In2024 43rd Chinese Control Conference (CCC), pages 8393–8398, Kunming, China, July

  95. [104]

    ISBN 978-988-758-158-1

    IEEE. ISBN 978-988-758-158-1. doi: 10.23919/CCC63176.2024.10661592. URL https://ieeexplore.ieee.org/document/10661592/

  96. [105]

    A gated cross-domain collaborative network for underwater object detection.Pattern Recognition, 149: 110222, May 2024

    Linhui Dai, Hong Liu, Pinhao Song, and Mengyuan Liu. A gated cross-domain collaborative network for underwater object detection.Pattern Recognition, 149: 110222, May 2024. ISSN 00313203. doi: 10.1016/j.patcog.2023.110222. URL https://linkinghub.elsevier.com/retrieve/pii/S00313...

  97. [106]

    CEH-YOLO: A composite enhanced YOLO-based model for underwater object detection.Ecological Informatics, 82:102758, September 2024

    Jiangfan Feng and Tao Jin. CEH-YOLO: A composite enhanced YOLO-based model for underwater object detection.Ecological Informatics, 82:102758, September 2024. ISSN 15749541. doi: 10.1016/j.ecoinf.2024.102758. URLhttps://linkinghub. elsevier.com/retrieve/pii/S1574954124003005

  98. [107]

    Enhancing Underwater Object Detection: Leveraging YOLOv8m for Im- proved Subaquatic Monitoring.SN Computer Science, 5(6):793, August 2024

    Abhishek Bajpai, Naveen Tiwari, Aditya Yadav, Divyansh Chaurasia, and Mohit Kumar. Enhancing Underwater Object Detection: Leveraging YOLOv8m for Im- proved Subaquatic Monitoring.SN Computer Science, 5(6):793, August 2024. ISSN 2661-8907. doi: 10.1007/s42979-024-03170-z. URLhtt...

  99. [108]

    Vi- sualBERT: A Simple and Performant Baseline for Vision and Language, 2019

    Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Vi- sualBERT: A Simple and Performant Baseline for Vision and Language, 2019. URL https://arxiv.org/abs/1908.03557. Version Number: 1

  100. [109]

    Learning Transferable Visual Models From Natural Lan- guage Supervision, 2021

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Lan- guage Supervision, 2021. URLhttps:/...

  101. [110]

    Flamingo: a Visual Language Model for Few-Shot Learning, 2022

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Mari- anne Monteiro, Jacob Menick, Sebasti...

  102. [111]

    The Dawn of LMMs: Preliminary Explorations with GPT- 4V(ision), 2023

    Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The Dawn of LMMs: Preliminary Explorations with GPT- 4V(ision), 2023. URLhttps://arxiv.org/abs/2309.17421. Version Number: 2

  103. [112]

    BLIP-2: Bootstrapping 61 Language-Image Pre-training with Frozen Image Encoders and Large Language Mod- els, 2023

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping 61 Language-Image Pre-training with Frozen Image Encoders and Large Language Mod- els, 2023. URLhttps://arxiv.org/abs/2301.12597. Version Number: 3

  104. [113]

    Underwater object detection method based on learnable query recall mechanism and lightweight adapter.PLOS ONE, 19(2): e0298739, February 2024

    Xi Lin, Xixia Huang, and Le Wang. Underwater object detection method based on learnable query recall mechanism and lightweight adapter.PLOS ONE, 19(2): e0298739, February 2024. ISSN 1932-6203. doi: 10.1371/journal.pone.0298739. URL https://dx.plos.org/10.1371/journal.pone.0298739

  105. [114]

    Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint...

  106. [115]

    MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning, 2023

    Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elho- seiny. MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning, 2023. URLhttps://arxiv....

  107. [116]

    InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks, 2023

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks, 2023. URLhtt...

  108. [117]

    Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, 2023

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, 2023. URLhttps:// arxiv.org/abs/2308.12966. Version Number: 3. 62

  109. [118]

    Optimization and Application of Improved YOLOv9s-UI for Underwater Object Detection.Applied Sciences, 14 (16):7162, August 2024

    Wei Pan, Jiabao Chen, Bangjun Lv, and Likun Peng. Optimization and Application of Improved YOLOv9s-UI for Underwater Object Detection.Applied Sciences, 14 (16):7162, August 2024. ISSN 2076-3417. doi: 10.3390/app14167162. URLhttps: //www.mdpi.com/2076-3417/14/16/7162

  110. [119]

    Junjie Wen, Jinqiang Cui, Benyun Zhao, Bingxin Han, Xuchen Liu, Zhi Gao, and Ben M. Chen. EnYOLO: A Real-Time Framework for Domain-Adaptive Underwater Object Detection with Image Enhancement. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 12613–1...

  111. [120]

    ISBN 9798350384574

    IEEE. ISBN 9798350384574. doi: 10.1109/ICRA57147.2024.10610639. URL https://ieeexplore.ieee.org/document/10610639/

  112. [121]

    Cluster-based fusion detection of soft and hard decisions for underwater non- cooperative targets.Signal Processing, 217:109327, April 2024

    Xiaoli Du, Yuyan Zhang, Yintang Wen, Zhixia Yang, Xiaoyuan Luo, and Jing Yan. Cluster-based fusion detection of soft and hard decisions for underwater non- cooperative targets.Signal Processing, 217:109327, April 2024. ISSN 01651684. doi: 10.1016/j.sigpro.2023.109327. URLhttps...

  113. [122]

    Towards Domain Generalization In Underwater Object Detection

    Hong Liu, Pinhao Song, and Runwei Ding. Towards Domain Generalization In Underwater Object Detection. In2020 IEEE International Conference on Image Processing (ICIP), pages 1971–1975, Abu Dhabi, United Arab Emirates, October 60

  114. [123]

    ISBN 978-1-72816-395-6

    IEEE. ISBN 978-1-72816-395-6. doi: 10.1109/ICIP40778.2020.9191364. URL https://ieeexplore.ieee.org/document/9191364/

  115. [124]

    Achieving domain generalization for underwater object detection by domain mixup and contrastive learning.Neurocomputing, 528:20–34, April 2023

    Yang Chen, Pinhao Song, Hong Liu, Linhui Dai, Xiaochuan Zhang, Runwei Ding, and Shengquan Li. Achieving domain generalization for underwater object detection by domain mixup and contrastive learning.Neurocomputing, 528:20–34, April 2023. ISSN 09252312. doi: 10.1016/j.neucom.20...

  116. [125]

    Dai, and et al Firat

    Rohan Anil, Andrew M. Dai, and et al Firat. PaLM 2 Technical Report, 2023. URL https://arxiv.org/abs/2305.10403. Version Number: 3

  117. [126]

    Wei Zhang, Miaoxin Cai, Tong Zhang, Yin Zhuang, and Xuerui Mao. EarthGPT: A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing Domain.IEEE Transactions on Geoscience and Remote Sensing, 63 62:1–20, 2024. ISSN 0196-2892, 1558-0644. d...

  118. [127]

    LLMRA: Multi-modal Large Language Model based Restoration Assistant, 2024

    Xiaoyu Jin, Yuan Shi, Bin Xia, and Wenming Yang. LLMRA: Multi-modal Large Language Model based Restoration Assistant, 2024. URLhttps://arxiv.org/abs/ 2401.11401. Version Number: 1

  119. [128]

    ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine- Grained Reward Modeling, 2024

    Siming Yan, Min Bai, Weifeng Chen, Xiong Zhou, Qixing Huang, and Li Erran Li. ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine- Grained Reward Modeling, 2024. URLhttps://arxiv.org/abs/2402.06118. Ver- sion Number: 3

  120. [129]

    Rela- tionVLM: Making Large Vision-Language Models Understand Visual Relations, 2024

    Zhipeng Huang, Zhizheng Zhang, Zheng-Jun Zha, Yan Lu, and Baining Guo. Rela- tionVLM: Making Large Vision-Language Models Understand Visual Relations, 2024. URLhttps://arxiv.org/abs/2403.12801. Version Number: 1

  121. [130]

    Gemini: A Family of Highly Capable Multimodal Models, 2023

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, and et al Yu. Gemini: A Family of Highly Capable Multimodal Models, 2023. URLhttps: //arxiv.org/abs/2312.11805. Version Number: 4

  122. [131]

    What Does DALL-E 2 Know About Radiology?Journal of Medical Internet Research, 25:e43110, March 2023

    LisaCAdams, FelixBusch, DanielTruhn, MarcusRMakowski, HugoJWLAerts, and Keno K Bressem. What Does DALL-E 2 Know About Radiology?Journal of Medical Internet Research, 25:e43110, March 2023. ISSN 1438-8871. doi: 10.2196/43110. URL https://www.jmir.org/2023/1/e43110

  123. [132]

    Creating Image Datasets in Agricultural Envi- ronments using DALL.E: Generative AI-Powered Large Language Model, 2023

    Ranjan Sapkota and Manoj Karkee. Creating Image Datasets in Agricultural Envi- ronments using DALL.E: Generative AI-Powered Large Language Model, 2023. URL https://arxiv.org/abs/2307.08789. Version Number: 4

  124. [133]

    Nascimento

    Chihcheng Hsieh, Catarina Moreira, Isabel Blanco Nobre, Sandra Costa Sousa, Chun Ouyang, Margot Brereton, Joaquim Jorge, and Jacinto C. Nascimento. DALL-M: Context-Aware Clinical Data Augmentation with LLMs, October 2024. URLhttp: //arxiv.org/abs/2407.08227. arXiv:2407.08227 [cs]. 64

  125. [134]

    MoE-LLaVA: Mixture of Experts for Large Vision-Language Models, 2024

    Bin Lin, Zhenyu Tang, Yang Ye, Jinfa Huang, Junwu Zhang, Yatian Pang, Peng Jin, Munan Ning, Jiebo Luo, and Li Yuan. MoE-LLaVA: Mixture of Experts for Large Vision-Language Models, 2024. URLhttps://arxiv.org/abs/2401.15947. Version Number: 5

  126. [135]

    The Llama 3 Herd of Models,

    Aaron Grattafiori, Abhimanyu Dubey, and et al Jauhri. The Llama 3 Herd of Models,

  127. [136]

    Version Number: 3

    URLhttps://arxiv.org/abs/2407.21783. Version Number: 3

  128. [137]

    Baichuan-Omni Technical Report,

    Yadong Li, Haoze Sun, Mingan Lin, and et al Li. Baichuan-Omni Technical Report,

  129. [138]

    Version Number: 4

    URLhttps://arxiv.org/abs/2410.08565. Version Number: 4

  130. [139]

    Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model, 2024

    Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model, 2024. URL https://arxiv.org/abs/2408.11039. Version...

  131. [140]

    Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models, 2024

    Matt Deitke, Christopher Clark, and et al Lee. Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models, 2024. URLhttps: //arxiv.org/abs/2409.17146. Version Number: 2

  132. [141]

    Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks, 2023

    Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan. Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks, 2023. URLhttps://arxiv.org/abs/2311.06242. Version Number: 1

  133. [142]

    DeepSeek-VL2: Mixture- of-Experts Vision-Language Models for Advanced Multimodal Understanding, 2024

    Zhiyu Wu, Xiaokang Chen, Zizheng Pan, and et al Liu. DeepSeek-VL2: Mixture- of-Experts Vision-Language Models for Advanced Multimodal Understanding, 2024. URLhttps://arxiv.org/abs/2412.10302. Version Number: 1

  134. [143]

    GPT-4 Technical Report,

    OpenAI, Josh Achiam, Steven Adler, and et al Agarwal. GPT-4 Technical Report,

  135. [144]

    Version Number: 6

    URLhttps://arxiv.org/abs/2303.08774. Version Number: 6

  136. [145]

    LAPT: Label-driven Auto- mated Prompt Tuning for OOD Detection with Vision-Language Models, 2024

    Yabin Zhang, Wenjie Zhu, Chenhang He, and Lei Zhang. LAPT: Label-driven Auto- mated Prompt Tuning for OOD Detection with Vision-Language Models, 2024. URL https://arxiv.org/abs/2407.08966. Version Number: 1

  137. [146]

    Segment Anything Model for automated image data annotation: empirical studies using text prompts from Grounding DINO,

    Fuseini Mumuni and Alhassan Mumuni. Segment Anything Model for automated image data annotation: empirical studies using text prompts from Grounding DINO,

  138. [147]

    A systematic re- view: object detection.AI & SOCIETY, April 2025

    Aishvi Guleria, Kamya Varshney, Garima ., and Shweta Jindal. A systematic re- view: object detection.AI & SOCIETY, April 2025. ISSN 0951-5666, 1435-

  139. [150]

    CoLLaVO: Crayon Large Language and Vision mOdel, 2024

    Byung-Kwan Lee, Beomchan Park, Chae Won Kim, and Yong Man Ro. CoLLaVO: Crayon Large Language and Vision mOdel, 2024. URLhttps://arxiv.org/abs/ 2402.11248. Version Number: 4

  140. [155]

    InternLM- XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolu- tions from 336 Pixels to 4K HD, 2024

    Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, a...

  141. [156]

    Michael Dorkenwald, Nimrod Barazani, Cees G. M. Snoek, and Yuki M. Asano. PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),pages13548–13558, Seattle, WA, USA, June 2024. IEEE. ISB...

  142. [157]

    Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models, 2024

    AlexHavrilla, AndrewDai, LauraO’Mahony, KoenOostermeijer, VeraZisler, AlonAl- balak, FabrizioMilo, SharathChandraRaparthy, KanishkGandhi, BaberAbbasi, Duy Phung, Maia Iyer, Dakota Mahan, Chase Blagden, Srishti Gureja, Mohammed Hamdy, Wen-Ding Li, Giovanni Paolini, Pawan Sasank...

  143. [158]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models, 2021. URLhttps://arxiv.org/abs/2106.09685. Version Number: 2

  144. [159]

    A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.ACM Transactions on Information Systems, 43(2):1–55, March 2025

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, 65 Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.ACM Transactions on I...

  145. [160]

    Hallucination is Inevitable: An Innate Limitation of Large Language Models, 2024

    Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. Hallucination is Inevitable: An Innate Limitation of Large Language Models, 2024. URLhttps://arxiv.org/abs/ 2401.11817. Version Number: 1

  146. [161]

    An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine- tuning, 2023

    Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang. An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine- tuning, 2023. URLhttps://arxiv.org/abs/2308.08747. Version Number: 5

  147. [162]

    A survey of efficient fine-tuning methods for Vision-Language Models — Prompt and Adapter.Computers & Graphics, 119:103885, April 2024

    Jialu Xing, Jianping Liu, Jian Wang, Lulu Sun, Xi Chen, Xunxun Gu, and Yingfei Wang. A survey of efficient fine-tuning methods for Vision-Language Models — Prompt and Adapter.Computers & Graphics, 119:103885, April 2024. ISSN 00978493. doi: 10.1016/j.cag.2024.01.012. URLhttps:...

  148. [163]

    Multi-scale fea- ture fusion with task-specific data synthesis for pneumonia pathogen classification

    Yinzhe Cui, Jing Liu, Ze Teng, Shuangfeng Yang, Hongfeng Li, Pingkang Li, Ji- abin Lu, Yajuan Gao, Yun Peng, Hongbin Han, and Wanyi Fu. Multi-scale fea- ture fusion with task-specific data synthesis for pneumonia pathogen classification. Image and Vision Computing, page 105662...

  149. [164]

    Diffusion Models: A Comprehensive Survey of Methods and Applications.ACM Computing Surveys, 56(4):1–39, April 2024

    LingYang, ZhilongZhang, YangSong, ShendaHong, RunshengXu, YueZhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. Diffusion Models: A Comprehensive Survey of Methods and Applications.ACM Computing Surveys, 56(4):1–39, April 2024. ISSN 0360-0300, 1557-7341. doi: 10.1145/3626235. U...

  150. [167]

    Version Number: 2

    URLhttps://arxiv.org/abs/2406.19057. Version Number: 2

  151. [2019]

    URLhttps://openaccess.thecvf.com/content_CVPRW_2019/ html/AAMVEM/Pedersen_Detection_of_Marine_Animals_in_a_New_Underwater_ Dataset_with_CVPRW_2019_paper.html

    IEEE Xplore. URLhttps://openaccess.thecvf.com/content_CVPRW_2019/ html/AAMVEM/Pedersen_Detection_of_Marine_Animals_in_a_New_Underwater_ Dataset_with_CVPRW_2019_paper.html

  152. [2020]

    doi: 10.1016/j.patcog.2019.107038

    ISSN00313203. doi: 10.1016/j.patcog.2019.107038. URLhttps://linkinghub. elsevier.com/retrieve/pii/S0031320319303401

  153. [2021]

    Version Number: 4

    URLhttps://arxiv.org/abs/2105.05233. Version Number: 4

  154. [2022]

    Version Number: 3

    URLhttps://arxiv.org/abs/2211.13523. Version Number: 3

  155. [2023]

    doi: 10.1109/JSEN.2023.3313108

    ISSN 1530-437X, 1558-1748, 2379-9153. doi: 10.1109/JSEN.2023.3313108. URL https://ieeexplore.ieee.org/document/10250195/

  156. [2024]

    doi: 10.1364/OE.510681

    ISSN 1094-4087. doi: 10.1364/OE.510681. URLhttps://opg.optica.org/ abstract.cfm?URI=oe-32-2-1789

  157. [2025]

    doi: 10.1016/j.sigpro.2024.109690

    ISSN 01651684. doi: 10.1016/j.sigpro.2024.109690. URLhttps://linkinghub. elsevier.com/retrieve/pii/S0165168424003104. 54

  158. [2315]

    URLhttps://link.springer.com/10

    doi: 10.1007/s00371-023-02858-2. URLhttps://link.springer.com/10. 1007/s00371-023-02858-2

  159. [3536]

    URLhttps://ieeexplore.ieee.org/ document/9359796/

    doi: 10.1109/ACCESS.2021.3060947. URLhttps://ieeexplore.ieee.org/ document/9359796/

  160. [5655]

    URLhttps://link.springer.com/10

    doi: 10.1007/s00146-025-02372-0. URLhttps://link.springer.com/10. 1007/s00146-025-02372-0. Publisher: Springer Science and Business Media LLC. Appendix A. The Process of Image Enhancement Appendix A.1. Aesthetic Enhancement In our case, the process of Underwater Aesthetic Enha...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.