REVIEW 3 major objections 4 minor 74 references
Unveiling the Underwater World: CLIP Perception Model-Guided Underwater Image Enhancement
T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Adding a CLIP perception model trained on human opinion scores as a loss and curriculum regularizer improves underwater enhancement quality and generalization over state-of-the-art methods.
desk verdict Solid but modest plug-and-play CLIP losses improve UIE networks on full-reference metrics; the paper's generalization claim is contradicted by its own no-reference numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the CLIP perception model: a frozen CLIP image encoder and text encoder with a learnable antonymic prompt pair (tokens T_p and T_n, both in $R^{{Nx512}}$), whose softmax of cosine similarities S_out = $e^{{S_p}}$/($e^{{S_p}}$+$e^{{S_n}}$) yields a predicted quality score in [0,1]. It is trained by L2 regression to human MOS labels on UEQAB (PLCC 0.83, SROCC 0.80 on the test set). Once trained, the prompt embeddings are frozen, and the model serves two roles: (a) it defines L_clip, a margin-style loss that drives the enhanced image's score above $\alpha$ times the reference's deficiency; and (b) it evaluates each of the six non-easy negatives N_q at each epoch, classifying each as 'hard' (if S_out^Y' > S_out^Nq, weight 1+gamma) or 'very-hard' (otherwise, weight 1-gamma), with gamma=0.25. The L_CR term then computes a weighted ratio of VGG-19 feature L1 distances between positive/negative/anchor, pulling the anchor to the reference and pushing it from a weighted combination of degraded images.
What would settle it
Train the same NU2Net-based UIE network with L1 + L_clip + L_CR but replace the CLIP score used to classify hard/very-hard negatives with random labels (or with the negative index order), keeping everything else identical. If the PSNR/SSIM/LPIPS results on U90 do not degrade materially compared to the proposed weighting, then the CLIP-based curriculum, not the mere presence of extra negatives, is not the source of the gains. Additionally, run a human preference study on 50 randomly sampled enhanced images from U45 and SQUID comparing the proposed method against NU2Net; if human raters do not choose the proposed method more often, the central perceptual-quality claim lacks support.
Extended reading notes
Core claim
The central discovery is that a CLIP perception model trained on human opinions can act as a differentiable quality oracle for underwater images, and that using it simultaneously as a loss and as a curriculum weighting of negatives yields better enhancement than the same network trained with L1 alone or with earlier quality-assessment losses. Concretely, the paper's L_clip = max(0, (1-S_out(Y')) - alpha(1-S_out(Y))) with alpha=0.975 encourages the enhanced image Y' to achieve a CLIP score exceeding the reference, while L_CR contrasts VGG-19 features of the anchor Y' against the reference Y, the 6 non-easy negatives from fixed UIE methods, and the input; the weight of each non-easy negative is 1+gamma or 1-gamma depending on whether its CLIP score is below or above the current anchor's, making difficult negatives gradually easier over training. The ablation tables show each module helps, and that the CLIP-based version of L_CR outperforms the PSNR-based version, which is the load-bearing evidence for using CLIP as the quality criterion.
Load-bearing premise
The CLIP perception model trained on UEQAB mean opinion scores accurately judges perceptual quality on all the other test sets (U45, SQUID, C60), so that both the L_clip loss and the weighting of negatives in L_CR are guided by a trustworthy human-aligned quality score.
Editorial extensions
If this is right
- Any existing UIE network (WaterNet, FUnIE, Shallow-UWnet, PUGAN, NU2Net) gains PSNR and SSIM when L_clip and L_CR are added, so the two loss modules are architecture-agnostic additions.
- The CLIP perception model can serve as a no-reference evaluation score for underwater images, complementing UCIQE and UIQM, since it matches human MOS more closely on UEQAB.
- Training with content-matched negatives at varying difficulty levels constrains the solution space and prevents both under- and over-enhancement, as shown by the ablation with single-negative and nine-negative configurations.
- The method generalizes to no-reference test sets with diverse distortion types (color casts, low contrast, blur) without retraining on those sets.
- Because L_clip is differentiable, the same perception model could be reused for other restoration tasks that optimize human-perceived quality.
Reading between the lines
- Since the paper validates the CLIP perception model's correlation with human scores only on UEQAB, a natural test is to collect human MOS on the U45, SQUID, and C60 enhancement outputs; if CLIP scores disagree with human rankings there, the reported no-reference CLIP-Score improvements should be interpreted as model preference rather than human-perceived quality.
- The binary hard/very-hard split with fixed gamma could be replaced by a continuous weight based on the CLIP score difference, which might yield smoother curriculum dynamics and faster convergence.
- The six negative-generation methods are fixed and heuristic; using the CLIP perception model to also select or generate negatives (e.g., through learned degradation models) could strengthen the regularization and push the gain further.
- The plug-and-play claim suggests a drop-in use for other underwater tasks like depth estimation or saliency detection, a direction the paper gestures at only through the SVAM-Net saliency experiment.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a plug-and-play perceptual training scheme for underwater image enhancement (UIE). The authors first train a CLIP-based perception model on the UEQAB MOS dataset using learnable antonymic prompt pairs (Section III.B, Eqs. 1-4). This model is then used in two ways inside the enhancement network: as a CLIP perception loss Lclip (Eq. 5) and as a difficulty-weighting criterion for curriculum contrastive regularization LCR (Eqs. 6-7), where negatives come from fixed UIE methods (UDCP, IBLA, DCP, HE, FUnIE, USUIR). The enhancement backbone is NU2Net, and the method is also applied to WaterNet, FUnIE, Shallow-UWnet, and PUGAN. Full-reference evaluation on U90 reports PSNR 23.115, SSIM 0.929, LPIPS 0.133 versus NU2Net's 22.669, 0.924, 0.154. No-reference evaluation is reported on U45, SQUID, and C60 using UCIQE, UIQM, and a CLIP-Score, plus qualitative comparisons and a saliency-detection application test.
Significance. If the central claim is read narrowly, the paper provides a consistent ablation story: adding Lclip and LCR to NU2Net improves full-reference metrics on U90 (Table IV), and adding both losses to four other UIE backbones improves them as well (Table VIII). The CLIP perception model itself achieves competitive PLCC/SROCC on the UEQAB test set (Table I), and the application test with a saliency detector is a useful auxiliary demonstration. The main weaknesses are that the headline full-reference gains are small and are reported without any variance or significance analysis, and the no-reference evaluation is partly circular because the reported CLIP-Score is computed by the same model that is used as the optimization target and as the curriculum-weighting criterion. The abstract's generalization claim is not supported by Table III, which shows the proposed method below NU2Net on several no-reference metrics. These issues are fixable by narrowing claims, adding multi-seed statistics, and validating the CLIP model on the no-reference datasets, so the work is potentially publishable but needs substantive revision.
major comments (3)
- [Section IV.A.3 and Table III] Reporting CLIP-Score as a no-reference evaluation metric is circular for this method. The CLIP perception model that computes CLIP-Score is the same model optimized by Lclip (Eq. 5) and used to set the curriculum weights in Eq. (6), and it is validated only on UEQAB in Table I; Section IV.J itself concedes that the model "may benefit from a more accurate selection of initialization prompts." Moreover, Table III does not show consistent superiority on the no-reference datasets: the proposed method is below NU2Net in CLIP-Score on U45 (56.02 vs 56.89) and C60 (50.41 vs 50.81), and below NU2Net in UIQM on SQUID (2.360 vs 2.480) and C60 (2.810 vs 2.900). Therefore the abstract's claim that the method "outperforms state-of-the-art methods in terms of visual quality and generalization ability" is not supported. The authors should either add human MOS studies on U45/SQUID/C60, report an external no-reference metric that is not part of the training objective, or restrict the generalization claim to the full-reference U90 result.
- [Tables II, IV, and VIII] The full-reference gains on U90 are small (PSNR +0.446 dB, SSIM +0.005, LPIPS -0.021 relative to NU2Net), and every number in the paper appears to come from a single training run. No standard deviation, confidence interval, paired per-image test, or training-seed information is reported anywhere, even though the network structure is identical to NU2Net and gains of a few tenths of a dB could plausibly arise from run-to-run variation. Please report mean and standard deviation over multiple seeds and, where possible, paired per-image differences on U90 to establish that the proposed loss terms are the cause of the observed improvement.
- [Section IV.A.1] The SQUID test protocol is underspecified: the text says "we select 16 representative examples as the test set, which is the same as [16]," but gives no criterion for selecting these 16 examples. Because the generalization claim depends on the no-reference datasets U45, SQUID, and C60, the selection procedure must be described in advance or the full SQUID set should be used; otherwise the result cannot be distinguished from favorable subset selection.
minor comments (4)
- [Section IV.E] The text is internally inconsistent: it states "the scheme L1 + Lclip actually means NU2Net method [20]," but Table V shows that NU2Net's reported performance corresponds to L1 + LUranker (22.669), while L1 + Lclip gives 22.677 in the same table. The paragraph also names the best scheme as "L1 + LUranker" twice, with one instance presumably meant to be L1 + LBFEN. Please rewrite this paragraph and reconcile it with Table V.
- [Section IV.C] The text refers to "The SUIQD dataset" in the discussion of the no-reference datasets; this appears to be a typo for SQUID and should be corrected.
- [Algorithm 2 and Section IV.F] There is a discrepancy in the number of negatives: Algorithm 2 sets z=6 non-easy negatives and Section IV.A.2 lists six fixed UIE methods, but Section IV.F says "the number of negatives is set as 7" and labels the scheme "+CR(1:7)." Please clarify whether z excludes or includes the easy negative and align the notation.
- [Table III caption] The caption states that the top three results are marked with red, blue, and green, but the text does not explain or discuss these color markers, and the printed table does not make the ranking visually clear. Either explain the color coding in the text or remove it.
Circularity Check
Full-reference improvements are independently supported, but the non-reference CLIP-Score evaluation reuses the model that the training loss explicitly optimizes, giving partial circularity.
-
fitted input called prediction
[Section IV.A.3 (Evaluation metrics); Section III.C Eq.(5); Section III.D Eq.(6); Table III]
"For the test datasets C60, U45, and SQUID, which do not contain reference images, ..., the proposed CLIP perception model is also employed for its stronger linear correlation with human visual perception. ... Lclip = max(0, ((1−SY′out)−α(1−SYout))). ... CLIP-Score↑"
The CLIP-Score column in Table III is the output Sout of Eq.(3), produced by prompts fit to UEQAB MOS labels via Eq.(4). During enhancement training, Eq.(5) explicitly drives the enhanced image's Sout upward (α=0.975), and Eq.(6) sets the curriculum-negative weights using that same Sout. Therefore, reporting CLIP-Score on U45/SQUID/C60 as evidence of human perceptual quality and generalization is using the training objective as the test metric: the reported number is the quantity the method was optimized to increase, not an independent assessment of those datasets. Section IV.J concedes that the model 'may benefit from a more accurate selection of initialization prompts,' and no human validation is given for the three no-reference datasets.
full rationale
The central enhancement claim is not circular. The proposed network is NU2Net augmented with Lclip and LCR, and the full-reference results on U90 (PSNR 23.115 vs 22.669, SSIM 0.929 vs 0.924, LPIPS 0.133 vs 0.154) are measured with metrics that are not part of the loss or the CLIP perception model. The ablations in Table IV and the plug-in experiments on WaterNet, FUnIE, Shallow-UWnet, and PUGAN in Table VIII show consistent improvements from the added modules, so the main derivation is self-contained and externally checkable. The circularity is limited to the non-reference evaluation path: the CLIP perception model is trained on UEQAB MOS labels (Eq.4), used as the training loss (Eq.5) and as the negative-weighting criterion (Eq.6), and then the same model's score is reported as CLIP-Score in Table III to claim generalization on U45/SQUID/C60. That is a fitted quantity used as an evaluation measure without independent human validation on those datasets, and the paper's own limitation section concedes incomplete prompt optimization. No load-bearing self-citation chain or imported uniqueness theorem is involved, and the full-reference claim does not reduce to the fitted CLIP model. The appropriate score is therefore 4: partial circularity in one evaluation metric, while the central full-reference contribution retains independent content.
Assumptions & free parameters
free parameters (8)
- CLIP prompt embeddings T_p and T_n =
Not reported (learned on UEQAB)
- alpha in Lclip =
0.975
- gamma in curriculum weights =
0.25
- lambda1 (Lclip weight) =
0.025
- lambda2 (LCR weight) =
0.1
- Number of non-easy negatives z =
6 (7 total including easy negative)
- VGG feature layer weights xi_i =
1/32, 1/16, 1/8, 1/4, 1
- Initial prompt texts =
"Clear Underwater photo.", "Turbid Underwater photo."
assumptions (5)
- domain assumption CLIP's frozen image and text encoders provide a transferable representation for underwater perceptual quality.
- domain assumption UEQAB MOS scores are reliable ground-truth labels for perceptual quality.
- domain assumption UIEB pseudo ground-truth images are valid training targets.
- domain assumption VGG-19 features are valid perceptual distances for underwater enhancement.
- domain assumption All baseline methods were trained fairly on the same datasets and devices.
Cite this review
Pith. "Pith review of Unveiling the Underwater World: CLIP Perception Model-Guided Underwater Image Enhancement." pith.science (2026). https://pith.science/paper/5PEDV66X
@misc{pith2026250706234,
author = {Pith},
title = {Pith review of: Unveiling the Underwater World: CLIP Perception Model-Guided Underwater Image Enhancement},
year = {2026},
howpublished = {\url{https://pith.science/paper/5PEDV66X}},
note = {Machine review of arXiv:2507.06234}
}
read the original abstract
High-quality underwater images are essential for both machine vision tasks and viewers with their aesthetic appeal.However, the quality of underwater images is severely affected by light absorption and scattering. Deep learning-based methods for Underwater Image Enhancement (UIE) have achieved good performance. However, these methods often overlook considering human perception and lack sufficient constraints within the solution space. Consequently, the enhanced images often suffer from diminished perceptual quality or poor content restoration.To address these issues, we propose a UIE method with a Contrastive Language-Image Pre-Training (CLIP) perception loss module and curriculum contrastive regularization. Above all, to develop a perception model for underwater images that more aligns with human visual perception, the visual semantic feature extraction capability of the CLIP model is leveraged to learn an appropriate prompt pair to map and evaluate the quality of underwater images. This CLIP perception model is then incorporated as a perception loss module into the enhancement network to improve the perceptual quality of enhanced images. Furthermore, the CLIP perception model is integrated with the curriculum contrastive regularization to enhance the constraints imposed on the enhanced images within the CLIP perceptual space, mitigating the risk of both under-enhancement and over-enhancement. Specifically, the CLIP perception model is employed to assess and categorize the learning difficulty level of negatives in the regularization process, ensuring comprehensive and nuanced utilization of distorted images and negatives with varied quality levels. Extensive experiments demonstrate that our method outperforms state-of-the-art methods in terms of visual quality and generalization ability.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[21]
Exploring clip for assessing the look and feel of images,
J. Wang, K. C. Chan, and C. C. Loy, “Exploring clip for assessing the look and feel of images,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 2, 2023, pp. 2555–2563
2023
-
[22]
Curricular contrastive regularization for physics-aware single image dehazing,
Y . Zheng, J. Zhan, S. He, J. Dong, and Y . Du, “Curricular contrastive regularization for physics-aware single image dehazing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2023, pp. 5785–5794
work page 2023
-
[16]
Underwater image enhancement via medium transmission-guided multi-color space embedding,
C. Li, S. Anwar, J. Hou, R. Cong, C. Guo, and W. Ren, “Underwater image enhancement via medium transmission-guided multi-color space embedding,” IEEE Transactions on Image Processing, vol. 30, pp. 4985– 5000, 2021
2021
-
[1]
Uncertainty inspired underwater image enhancement,
Z. Fu, W. Wang, Y . Huang, X. Ding, and K.-K. Ma, “Uncertainty inspired underwater image enhancement,” in European conference on computer vision, 2022, pp. 465–482
work page 2022
-
[2]
An in-depth survey of underwater image enhancement and restoration,
M. Yang, J. Hu, C. Li, G. Rohde, Y . Du, and K. Hu, “An in-depth survey of underwater image enhancement and restoration,” IEEE Access, vol. 7, pp. 123 638–123 657, 2019
work page 2019
-
[3]
A survey on underwater image en- hancement techniques,
P. Sahu, N. Gupta, and N. Sharma, “A survey on underwater image en- hancement techniques,” International Journal of Computer Applications, vol. 87, no. 13, 2014
work page 2014
-
[4]
M. J. Kaiser et al. , Marine ecology: processes, systems, and impacts . Oxford University Press, USA, 2011
work page 2011
-
[5]
R. Long, “The marine strategy framework directive: a new european approach to the regulation of the marine environment, marine natural resources and marine ecological services,” Journal of Energy & Natural Resources Law, vol. 29, no. 1, pp. 1–44, 2011
work page 2011
Show all 74 references
-
[6]
Towards an urban marine ecology: characterizing the drivers, patterns and processes of marine ecosystems in coastal cities,
P. A. Todd, E. C. Heery, L. H. Loke, R. H. Thurstan, D. J. Kotze, and C. Swan, “Towards an urban marine ecology: characterizing the drivers, patterns and processes of marine ecosystems in coastal cities,” Oikos, vol. 128, no. 9, pp. 1215–1242, 2019
2019
-
[7]
Underwater image processing: state of the art of restoration and image enhancement methods,
R. Schettini and S. Corchs, “Underwater image processing: state of the art of restoration and image enhancement methods,” EURASIP journal on advances in signal processing , vol. 2010, pp. 1–14, 2010
2010
-
[8]
Enhancing underwa- ter images and videos by fusion,
C. Ancuti, C. O. Ancuti, T. Haber, and P. Bekaert, “Enhancing underwa- ter images and videos by fusion,” in 2012 IEEE conference on computer vision and pattern recognition . IEEE, 2012, pp. 81–88
2012
-
[9]
Enhancing the low quality images using unsupervised colour correction method,
K. Iqbal, M. Odetayo, A. James, R. A. Salam, and A. Z. H. Talib, “Enhancing the low quality images using unsupervised colour correction method,” in 2010 IEEE International Conference on Systems, Man and Cybernetics. IEEE, 2010, pp. 1703–1709
2010
-
[10]
Single image haze removal using dark channel prior,
K. He, J. Sun, and X. Tang, “Single image haze removal using dark channel prior,” IEEE transactions on pattern analysis and machine intelligence, vol. 33, no. 12, pp. 2341–2353, 2010
2010
-
[11]
Underwater depth estimation and image restoration based on single images,
P. L. Drews, E. R. Nascimento, S. S. Botelho, and M. F. M. Campos, “Underwater depth estimation and image restoration based on single images,” IEEE computer graphics and applications , vol. 36, no. 2, pp. 24–35, 2016
2016
-
[12]
Underwater image enhancement by dehazing with minimum information loss and histogram distribution prior,
C.-Y . Li, J.-C. Guo, R.-M. Cong, Y .-W. Pang, and B. Wang, “Underwater image enhancement by dehazing with minimum information loss and histogram distribution prior,” IEEE Transactions on Image Processing , vol. 25, no. 12, pp. 5664–5677, 2016
2016
-
[13]
Object detection in 20 years: A survey,
Z. Zou, K. Chen, Z. Shi, Y . Guo, and J. Ye, “Object detection in 20 years: A survey,” Proceedings of the IEEE , 2023
2023
-
[14]
Underwater image enhancement using a multiscale dense generative adversarial network,
Y . Guo, H. Li, and P. Zhuang, “Underwater image enhancement using a multiscale dense generative adversarial network,” IEEE Journal of Oceanic Engineering, vol. 45, no. 3, pp. 862–870, 2019
2019
-
[15]
Water- GAN: Unsupervised generative network to enable real-time color correc- tion of monocular underwater images,
J. Li, K. A. Skinner, R. M. Eustice, and M. Johnson-Roberson, “Water- GAN: Unsupervised generative network to enable real-time color correc- tion of monocular underwater images,” IEEE Robotics and Automation letters, vol. 3, no. 1, pp. 387–394, 2017
2017
-
[17]
Sea-thru: A method for removing water from underwater images,
D. Akkaynak and T. Treibitz, “Sea-thru: A method for removing water from underwater images,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 1682–1691
2019
-
[18]
An underwater image enhancement benchmark dataset and beyond,
C. Li, C. Guo, W. Ren, R. Cong, J. Hou, S. Kwong, and D. Tao, “An underwater image enhancement benchmark dataset and beyond,” IEEE Transactions on Image Processing , vol. 29, pp. 4376–4389, 2019
2019
-
[19]
U-shape transformer for underwater image enhancement,
L. Peng, C. Zhu, and L. Bian, “U-shape transformer for underwater image enhancement,” IEEE Transactions on Image Processing , 2023. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 12
2023
-
[20]
Underwater ranker: Learn which is better and how to be better,
C. Guo, R. Wu, X. Jin, L. Han, W. Zhang, Z. Chai, and C. Li, “Underwater ranker: Learn which is better and how to be better,” in Proceedings of the AAAI conference on artificial intelligence , vol. 37, no. 1, 2023, pp. 702–709
2023
-
[23]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
-
[24]
Contrastive learning for compact single image dehazing,
H. Wu, Y . Qu, S. Lin, J. Zhou, R. Qiao, Z. Zhang, Y . Xie, and L. Ma, “Contrastive learning for compact single image dehazing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 10 551–10 560
2021
-
[25]
Generalization of the dark channel prior for single image restoration,
Y .-T. Peng, K. Cao, and P. C. Cosman, “Generalization of the dark channel prior for single image restoration,” IEEE Transactions on Image Processing, vol. 27, no. 6, pp. 2856–2868, 2018
2018
-
[26]
Automatic red- channel underwater image restoration,
A. Galdran, D. Pardo, A. Pic ´on, and A. Alvarez-Gila, “Automatic red- channel underwater image restoration,” Journal of Visual Communica- tion and Image Representation , vol. 26, pp. 132–145, 2015
2015
-
[27]
Underwater image restoration based on image blurriness and light absorption,
Y .-T. Peng and P. C. Cosman, “Underwater image restoration based on image blurriness and light absorption,” IEEE transactions on image processing, vol. 26, no. 4, pp. 1579–1594, 2017
2017
-
[28]
Shallow-uwnet: Compressed model for underwater image enhancement (student abstract),
A. Naik, A. Swarnakar, and K. Mittal, “Shallow-uwnet: Compressed model for underwater image enhancement (student abstract),” in Pro- ceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 18, 2021, pp. 15 853–15 854
2021
-
[29]
UIECˆ 2-Net: Cnn-based under- water image enhancement using two color space,
Y . Wang, J. Guo, H. Gao, and H. Yue, “UIECˆ 2-Net: Cnn-based under- water image enhancement using two color space,” Signal Processing: Image Communication, vol. 96, p. 116250, 2021
2021
-
[30]
Uncertainty inspired underwater image enhancement,
Z. Fu, W. Wang, Y . Huang, X. Ding, and K.-K. Ma, “Uncertainty inspired underwater image enhancement,” in European Conference on Computer Vision. Springer, 2022, pp. 465–482
2022
-
[31]
Enhancing underwater imagery using generative adversarial networks,
C. Fabbri, M. J. Islam, and J. Sattar, “Enhancing underwater imagery using generative adversarial networks,” in 2018 IEEE international conference on robotics and automation (ICRA). IEEE, 2018, pp. 7159– 7165
2018
-
[32]
Open-vocabulary detr with conditional matching,
Y . Zang, W. Li, K. Zhou, C. Huang, and C. C. Loy, “Open-vocabulary detr with conditional matching,” in European Conference on Computer Vision. Springer, 2022, pp. 106–122
2022
-
[33]
F-vlm: Open-vocabulary object detection upon frozen vision and language models,
W. Kuo, Y . Cui, X. Gu, A. Piergiovanni, and A. Angelova, “F-vlm: Open-vocabulary object detection upon frozen vision and language models,” arXiv preprint arXiv:2209.15639 , 2022
2022 arXiv
-
[34]
Extract free dense labels from clip,
C. Zhou, C. C. Loy, and B. Dai, “Extract free dense labels from clip,” in European Conference on Computer Vision. Springer, 2022, pp. 696– 712
2022
-
[35]
Blind image quality assessment via vision-language correspondence: A multitask learning perspective,
W. Zhang, G. Zhai, Y . Wei, X. Yang, and K. Ma, “Blind image quality assessment via vision-language correspondence: A multitask learning perspective,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 14 071–14 081
2023
-
[36]
VILA: Learning image aesthetics from user comments with vision-language pretraining,
J. Ke, K. Ye, J. Yu, Y . Wu, P. Milanfar, and F. Yang, “VILA: Learning image aesthetics from user comments with vision-language pretraining,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 10 041–10 051
2023
-
[37]
A reference-free underwater image quality assessment metric in frequency domain,
N. Yang, Q. Zhong, K. Li, R. Cong, Y . Zhao, and S. Kwong, “A reference-free underwater image quality assessment metric in frequency domain,” Signal Processing: Image Communication, vol. 94, p. 116218, 2021
2021
-
[39]
Learning to prompt for vision- language models,
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision- language models,” International Journal of Computer Vision , vol. 130, no. 9, pp. 2337–2348, 2022
2022
-
[40]
Conditional prompt learning for vision-language models,
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Conditional prompt learning for vision-language models,” in Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition , 2022, pp. 16 816– 16 825
2022
-
[41]
A simple framework for contrastive learning of visual representations,
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning . PMLR, 2020, pp. 1597–1607
2020
-
[42]
Bootstrap your own latent-a new approach to self-supervised learning,
J.-B. Grill, F. Strub, F. Altch ´e, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar et al. , “Bootstrap your own latent-a new approach to self-supervised learning,” Advances in neural information processing systems, vol. 33, pp. ...
2020
-
[43]
Momentum contrast for unsupervised visual representation learning,
K. He, H. Fan, Y . Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 9729–9738
2020
-
[44]
Curriculum learning,
Y . Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proceedings of the 26th annual international conference on machine learning , 2009, pp. 41–48
2009
-
[45]
Contrastive semi- supervised learning for underwater image restoration via reliable bank,
S. Huang, K. Wang, H. Liu, J. Chen, and Y . Li, “Contrastive semi- supervised learning for underwater image restoration via reliable bank,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 18 145–18 155
2023
-
[46]
Underwater image restoration via contrastive learning and a real-world dataset,
J. Han, M. Shoeiby, T. Malthus, E. Botha, J. Anstee, S. Anwar, R. Wei, M. A. Armin, H. Li, and L. Petersson, “Underwater image restoration via contrastive learning and a real-world dataset,” Remote Sensing, vol. 14, no. 17, p. 4297, 2022
2022
-
[47]
Denseclip: Language-guided dense prediction with context-aware prompting,
Y . Rao, W. Zhao, G. Chen, Y . Tang, Z. Zhu, G. Huang, J. Zhou, and J. Lu, “Denseclip: Language-guided dense prediction with context-aware prompting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 082–18 091
2022
-
[48]
Natural language processing: State of the art, current trends and challenges,
D. Khurana, A. Koli, K. Khatter, and S. Singh, “Natural language processing: State of the art, current trends and challenges,” Multimedia tools and applications , vol. 82, no. 3, pp. 3713–3744, 2023
2023
-
[49]
Human perceptual quality driven underwater image enhancement framework,
M. Li, Y . Lin, L. Shen, Z. Wang, K. Wang, and Z. Wang, “Human perceptual quality driven underwater image enhancement framework,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1– 15, 2022
2022
-
[50]
A fusion adversarial underwater image enhancement network with a public test dataset. arxiv 2019,
H. Li, J. Li, and W. Wang, “A fusion adversarial underwater image enhancement network with a public test dataset. arxiv 2019,” arXiv preprint arXiv:1906.06819
2019 arXiv
-
[51]
The unreasonable effectiveness of deep features as a perceptual metric,
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 586–595
2018
-
[52]
An underwater color image quality evaluation metric,
M. Yang and A. Sowmya, “An underwater color image quality evaluation metric,” IEEE Transactions on Image Processing , vol. 24, no. 12, pp. 6062–6071, 2015
2015
-
[53]
Human-visual-system-inspired underwater image quality measures,
K. Panetta, C. Gao, and S. Agaian, “Human-visual-system-inspired underwater image quality measures,” IEEE Journal of Oceanic Engi- neering, vol. 41, no. 3, pp. 541–551, 2015
2015
-
[54]
Transmission estimation in underwater single images,
P. Drews, E. Nascimento, F. Moraes, S. Botelho, and M. Campos, “Transmission estimation in underwater single images,” in Proceedings of the IEEE international conference on computer vision workshops , 2013, pp. 825–830
2013
-
[55]
Under- water image enhancement via minimal color loss and locally adaptive contrast enhancement,
W. Zhang, P. Zhuang, H.-H. Sun, G. Li, S. Kwong, and C. Li, “Under- water image enhancement via minimal color loss and locally adaptive contrast enhancement,” IEEE Transactions on Image Processing, vol. 31, pp. 3997–4010, 2022
2022
-
[56]
Fast underwater image enhancement for improved visual perception,
M. J. Islam, Y . Xia, and J. Sattar, “Fast underwater image enhancement for improved visual perception,” IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 3227–3234, 2020
2020
-
[57]
Subjective and objective de-raining quality assessment towards authentic rain image,
Q. Wu, L. Wang, K. N. Ngan, H. Li, F. Meng, and L. Xu, “Subjective and objective de-raining quality assessment towards authentic rain image,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 30, no. 11, pp. 3883–3897, 2020
2020
-
[58]
PUGAN: Physical model-guided underwater image en- hancement using gan with dual-discriminators,
R. Cong, W. Yang, W. Zhang, C. Li, C.-L. Guo, Q. Huang, and S. Kwong, “PUGAN: Physical model-guided underwater image en- hancement using gan with dual-discriminators,” IEEE Transactions on Image Processing, vol. 32, pp. 4472–4485, 2023
2023
-
[59]
SV AM: Saliency-guided visual attention modeling by autonomous underwater robots,
M. J. Islam, R. Wang, and J. Sattar, “SV AM: Saliency-guided visual attention modeling by autonomous underwater robots,” arXiv preprint arXiv:2011.06252, 2020
2011 arXiv
-
[60]
Re- inforced swin-convs transformer for simultaneous underwater sensing scene image enhancement and super-resolution,
T. Ren, H. Xu, G. Jiang, M. Yu, X. Zhang, B. Wang, and T. Luo, “Re- inforced swin-convs transformer for simultaneous underwater sensing scene image enhancement and super-resolution,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–16, 2022
2022
-
[61]
UW- CycleGAN: Model-driven cyclegan for underwater image restoration,
H. Yan, Z. Zhang, J. Xu, T. Wang, P. An, A. Wang, and Y . Duan, “UW- CycleGAN: Model-driven cyclegan for underwater image restoration,” IEEE Transactions on Geoscience and Remote Sensing , vol. 61, pp. 1– 17, 2023
2023
-
[62]
Iterative prompt learn- ing for unsupervised backlit image enhancement,
Z. Liang, C. Li, S. Zhou, R. Feng, and C. C. Loy, “Iterative prompt learn- ing for unsupervised backlit image enhancement,” in 2023 IEEE/CVF JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 International Conference on Computer Vision (ICCV) , 2023, pp. 8060– 8069
2023
-
[63]
Unsu- pervised underwater image restoration: From a homology perspective,
Z. Fu, H. Lin, Y . Yang, S. Chai, L. Sun, Y . Huang, and X. Ding, “Unsu- pervised underwater image restoration: From a homology perspective,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 1, 2022, pp. 643–651
2022
-
[64]
Image enhancement by histogram transformation,
R. Hummel, “Image enhancement by histogram transformation,” Un- known, 1975
1975
-
[65]
Underwater image enhancement quality evaluation: Benchmark dataset and objective met- ric,
Q. Jiang, Y . Gu, C. Li, R. Cong, and F. Shao, “Underwater image enhancement quality evaluation: Benchmark dataset and objective met- ric,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 32, no. 9, pp. 5959–5974, 2022
2022
-
[66]
MirrorDiffusion: Stabilizing diffusion process in zero-shot image translation by prompts redescription and beyond,
Y . Lin, X. Xian, Y . Shi, and L. Lin, “MirrorDiffusion: Stabilizing diffusion process in zero-shot image translation by prompts redescription and beyond,” IEEE Signal Processing Letters , 2024
2024
-
[67]
Exploring negatives in contrastive learning for unpaired image-to-image translation,
Y . Lin, S. Zhang, T. Chen, Y . Lu, G. Li, and Y . Shi, “Exploring negatives in contrastive learning for unpaired image-to-image translation,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 1186–1194
2022
-
[68]
Perceptual losses for real-time style transfer and super-resolution,
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11- 14, 2016, Proceedings, Part II 14 . Springer, 2016, pp. 694–711
2016
-
[69]
AUV navigation and localiza- tion: A review,
L. Paull, S. Saeedi, M. Seto, and H. Li, “AUV navigation and localiza- tion: A review,” IEEE Journal of oceanic engineering , vol. 39, no. 1, pp. 131–149, 2013
2013
-
[70]
RRNet: Re- lational reasoning network with parallel multiscale attention for salient object detection in optical remote sensing images,
R. Cong, Y . Zhang, L. Fang, J. Li, Y . Zhao, and S. Kwong, “RRNet: Re- lational reasoning network with parallel multiscale attention for salient object detection in optical remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–11, 2021
2021
-
[71]
A reinforcement learning paradigm of configuring visual enhancement for object detection in underwater scenes,
H. Wang, S. Sun, X. Bai, J. Wang, and P. Ren, “A reinforcement learning paradigm of configuring visual enhancement for object detection in underwater scenes,” IEEE Journal of Oceanic Engineering , vol. 48, no. 2, pp. 443–461, 2023
2023
-
[72]
Nested network with two-stream pyramid for salient object detection in optical remote sensing images,
C. Li, R. Cong, J. Hou, S. Zhang, Y . Qian, and S. Kwong, “Nested network with two-stream pyramid for salient object detection in optical remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 57, no. 11, pp. 9156–9166, 2019
2019
-
[73]
Diving deeper into underwater image enhance- ment: A survey,
S. Anwar and C. Li, “Diving deeper into underwater image enhance- ment: A survey,” Signal Processing: Image Communication , vol. 89, p. 115978, 2020
2020
-
[74]
Underwater image enhancement via adaptive group attention-based multiscale cascade transformer,
Z. Huang, J. Li, Z. Hua, and L. Fan, “Underwater image enhancement via adaptive group attention-based multiscale cascade transformer,” IEEE Transactions on Instrumentation and Measurement , vol. 71, pp. 1–18, 2022
2022
-
[75]
Multiscale underwater image enhancement in rgb and hsv color spaces,
C. Liu, X. Shu, L. Pan, J. Shi, and B. Han, “Multiscale underwater image enhancement in rgb and hsv color spaces,” IEEE Transactions on Instrumentation and Measurement , vol. 72, pp. 1–14, 2023
2023
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.