REVIEW 5 major objections 6 minor 75 references
DGIQA: Depth-guided Feature Attention and Refinement for Generalizable Image Quality Assessment
T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Adding depth-guided cross-attention to a no-reference image quality model makes its quality scores generalize to unseen distortions such as low light, haze, and lens flares.
desk verdict Depth-guided IQA with a promising architecture, yet the key ablation is confounded by parameter count and the abstract oversells the results. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing component is the Depth-CAR block, a depth-guided cross-attention and refinement mechanism: features from the depth stream provide the queries, while RGB features supply keys and values, so the attention weights emphasize close, salient objects and structural boundaries that human viewers tend to prioritize. The TCB block is the second mechanism; it applies squeeze-and-excitation channel recalibration to transformer patch embeddings followed by 3x3 convolutions, distilling global context into local hierarchical features and shrinking parameter count by 23.3%. A final dilated-convolution stack with rates 2 and 4 widens the receptive field before global pooling and a fully connected layer predict the quality score. The training objective combines MSE with a consistency loss that penalizes disagreement between an image and its horizontal flip.
What would settle it
Retrain DGIQA with the estimated depth maps replaced by random noise or by depth from a deliberately corrupted estimator; if cross-dataset SROCC and PLCC on the unseen natural distortion sets stay the same or improve, then depth content is not what drives the reported gains, and the depth-guidance claim collapses.
Extended reading notes
Core claim
The paper's central claim is that depth-guided attention improves NR-IQA generalization, and DGIQA is the empirical demonstration. Two parallel Swin-Transformer backbones extract features from the RGB image and its estimated depth map; four TCB blocks recalibrate channels and add local convolutions; then the Depth-CAR block uses depth-derived features as queries in scaled dot-product attention over RGB keys and values, followed by self-attention refinement and a dilated-convolution stack before score prediction. On the benchmark evaluations the model ranks in the top two across all seven datasets and takes first place on LIVE, CSIQ, Kadid10k, LIVE-C, and Koniq10k. In cross-dataset transfers it reports gains of roughly 0.8–5.0% in SROCC and up to 6.4% in PLCC over state-of-the-art baselines, and on unseen natural distortions its predicted-score distributions for low-light, haze, lens-flare, and motion-blur images overlap 3–95% less with high-quality distributions than the baselines do, with the strongest separation on hazy scenes (1.73% overlap). The authors also introduce 'density separation' as a new criterion for measuring NR-IQA generalization.
Load-bearing premise
The load-bearing premise is that depth maps computed by a pretrained monocular estimator on distorted images preserve enough structural information for depth-guided attention to help quality prediction; the paper itself notes that on flat or depth-uniform scenes, or when estimated depth is noisy, the depth guidance does not help and can even cause the model to diverge.
Editorial extensions
If this is right
- If the depth-guidance claim holds, NR-IQA models can be made substantially more robust to unseen real-world distortions by adding a depth branch, without requiring new subjective datasets.
- The TCB design shows that fusing transformer global features with CNN local features can improve quality prediction while cutting parameters, pointing toward cheaper IQA models.
- The reported cross-dataset gains imply that models trained on one large authentic dataset such as Koniq10k transfer better to other authentic and synthetic domains when depth-guided.
- Density separation on external datasets such as LOL, IHAZE, Flare7k, and GoPro offers a practical protocol for testing generalization to natural distortions beyond standard benchmarks.
Reading between the lines
- Beyond the paper: if depth maps remain stable under corruptions that destroy RGB texture (haze, blur, low light), depth-guided attention could extend naturally to video quality assessment, where temporal depth coherence would give a stronger prior than per-frame RGB statistics.
- Beyond the paper: the density-separation measure could be adopted as a standard generalization test for NR-IQA, but its overlap percentages depend on the choice of high/low quality thresholds and should be calibrated against multiple baselines before being used as a headline metric.
- Beyond the paper: because the depth estimator is trained mostly on clean scenes, heavily noisy or synthetic-distortion images may yield unreliable depth; testing DGIQA with depth maps from a distortion-robust estimator would clarify whether the improvement comes from depth quality or from the attention mechanism itself.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DGIQA, a no-reference image quality assessment model that combines two Swin Transformer backbones (one for RGB, one for depth), Transformer-CNN Bridge (TCB) blocks, a depth-guided cross-attention and refinement (Depth-CAR) module, and a dilated convolution stack. The model is trained with MSE plus a consistency loss, and is evaluated on seven benchmark datasets, cross-dataset transfer, and five natural distortion datasets (low light, haze, flare, motion blur). The authors claim state-of-the-art performance on synthetic and authentic benchmarks, superior cross-dataset generalization, and introduce a 'density separation' criterion to quantify generalization on unseen distortions. The paper includes ablations, Grad-CAM visualizations, t-SNE plots, and an FR-IQA adaptation.
Significance. If the central claims were properly supported, depth-guided attention would be a useful and reasonably novel contribution to NR-IQA, and the TCB block is a sensible parameter-reduction idea. The paper reports a 10-split evaluation protocol and makes model and code publicly available, which are strengths. However, the SOTA claim is weakened by Table 2, where DGIQA is not the best method on several datasets, and the key ablation in Tables 4-5 is confounded because removing Depth-CAR also removes the second Swin-T backbone (65M vs 103M parameters). The proposed density-separation criterion is not formally defined and is not validated against human opinion or established generalization metrics. The contribution is therefore plausible but not yet established.
major comments (5)
- [Abstract; Table 2] The SOTA claim in the Abstract ('achieves state-of-the-art (SOTA) performance on both synthetic and authentic benchmark datasets') is not supported by Table 2. On TID2013, MANIQA reports SROCC/PLCC of 0.937/0.943 while DGIQA reports 0.934/0.940; on LIVE-FB, TOPIQ reports 0.652/0.745 and Re-IQA 0.645/0.733 versus DGIQA 0.591/0.685; and on LIVE-C, LIQE reports 0.904/0.910 versus DGIQA 0.891/0.910. The claim should be revised to 'competitive' or 'top-two', and the comparisons need variance information.
- [Sec. 4.4, Tables 4-5] The ablation that attributes the cross-dataset gains to Depth-CAR is confounded by model capacity. In Table 5, configuration #1 (DGIQA w/o D-CAR) has 65M parameters and configuration #4 (full DGIQA) has 103M parameters; removing Depth-CAR also removes the entire second Swin-T backbone that consumes the depth map. The reported improvements of +3.2% SROCC on LIVE-C and +8.2% on LIVE-FB could therefore be due to increased trainable capacity rather than to depth-guided attention. A controlled comparison, for example a second RGB stream with the same capacity or a frozen depth backbone, is necessary to support the central novelty claim.
- [Sec. 5.2] The proposed 'density separation' criterion is not defined formally, and it is used as the main evidence for generalization on natural distortions. The paper does not specify how the Gaussian overlap is computed, how the 41-50% / 20-21% / 5-13% / 65-95% / 3-60% figures are derived, or whether these differences are statistically significant across the 10 training splits. Since this criterion is introduced by the authors and is not validated against human opinion scores or an established metric, it should be presented as an auxiliary diagnostic rather than as standalone evidence of SOTA generalization.
- [Sec. 4.2; Appendix B] The depth maps are generated by DepthAnything on distorted inputs, and Appendix B shows that severe white noise degrades the structural information in these maps. The paper acknowledges this limitation in Sec. 5.2, but there is no quantitative analysis of how often DepthAnything fails or how much such failures affect DGIQA's predictions. To support the claim that depth guidance is robust, the authors should report performance under corrupted or noisy depth inputs, or compare against a model trained on clean depth maps.
- [Sec. 4.3, Table 3] The cross-dataset claim is stated as outperforming SOTA 'in most dataset pairs', but Table 3(a) shows that on LIVE-C, LoDa achieves SROCC 0.811 versus DGIQA 0.808, and DGIQA does not consistently beat all baselines across pairs. Furthermore, no standard deviations or significance tests are reported for any of the 10-split means, so small differences such as 0.808 vs 0.811 are not interpretable. Please provide per-split statistics or error bars, and qualify the statement accordingly.
minor comments (6)
- [Sec. 2; Sec. 4.2] There are typos: 'seimesi architecture' should be 'Siamese architecture' in Sec. 2, and 'albumentation library' should be 'Albumentations library' in Sec. 4.2.
- [Appendix A.2] Equation (7) writes 'LSME' instead of 'LMSE' for the mean-squared-error loss term.
- [Sec. 4.4, Table 4] The checkmarks in Table 4 are not legible in the typeset version; the four configurations should be explicitly named in the caption or in the table.
- [Sec. 3.1; Eq. (4)] The notation for dilated convolutions is inconsistent: the text says 'dilation rates of 2 and 4', but Eq. (4) writes 'dltn=2,4' without a clear definition of the symbol 'dltn'.
- [Sec. 4.4] The claim that TCB blocks 'reduce the model parameters by 23.3%' is not directly verifiable from Tables 4-5; the parameter count drops from 127M (w/o TCB) to 103M (full), which is about 18.9%, so please clarify the reference configuration for this percentage.
- [Sec. 5.2] The phrase 'a new evaluation criteria' should be 'a new evaluation criterion'.
Circularity Check
No significant circularity: DGIQA's claims are empirical and benchmarked against external datasets; the only self-citation is peripheral and no prediction reduces to its inputs.
full rationale
The paper's central claims are empirical architecture and comparison results, not a derivation chain whose conclusions are equivalent to its premises. The quality score is trained against external MOS/DMOS labels, and the proposed Depth-CAR and TCB blocks are defined by ordinary attention and squeeze-excitation equations that do not encode the benchmark outcomes. Depth maps come from the external, pretrained DepthAnything model, not from the IQA labels or from fitted IQA parameters. The new 'density separation' criterion in Sec. 5.2 is an evaluation metric, not a training target, and it is not used to fit the model. The only self-citation is reference [20], a prior saliency-attention paper by one of the co-authors; it supports a general statement about human visual attention but is not load-bearing for any architectural choice or performance claim. The ablation confound noted by a skeptical reader—removing Depth-CAR also removes the second Swin-T backbone, changing parameter count from 65M to 103M—is a legitimate experimental-validity concern about attributing gains to depth semantics versus capacity, but it is not circularity: it does not make the claimed improvement equivalent by construction to the input. Hyperparameter tuning of crops and λ is model selection, not fitted-input-called-prediction. No circular step can be quoted from the paper.
Assumptions & free parameters
free parameters (6)
- Consistency loss weight lambda =
0.3
- Number of crops at evaluation =
25
- TCB channel reduction rate r =
16
- Dilation rates in dilation block =
2 and 4
- Learning rates =
1e-4 synthetic, 1e-5 authentic
- Batch size and max epochs =
16, 200
assumptions (5)
- domain assumption Two Swin Transformer backbones pretrained on ImageNet-21k provide transferable features for both RGB and depth.
- domain assumption DepthAnything produces depth maps accurate enough on distorted images for the cross-attention to help.
- domain assumption An 80:20 random split with ten different seeds estimates generalization without leakage.
- domain assumption Horizontal-flip consistency loss improves generalization beyond the MSE loss alone.
- ad hoc to paper Density separation of predicted-score distributions is a valid measure of IQA generalization.
invented entities (1)
-
Density separation criterion
Cite this review
Pith. "Pith review of DGIQA: Depth-guided Feature Attention and Refinement for Generalizable Image Quality Assessment." pith.science (2026). https://pith.science/paper/U4NATM4F
@misc{pith2026250524002,
author = {Pith},
title = {Pith review of: DGIQA: Depth-guided Feature Attention and Refinement for Generalizable Image Quality Assessment},
year = {2026},
howpublished = {\url{https://pith.science/paper/U4NATM4F}},
note = {Machine review of arXiv:2505.24002}
}
read the original abstract
A long-held challenge in no-reference image quality assessment (NR-IQA) learning from human subjective perception is the lack of objective generalization to unseen natural distortions. To address this, we integrate a novel Depth-Guided cross-attention and refinement (Depth-CAR) mechanism, which distills scene depth and spatial features into a structure-aware representation for improved NR-IQA. This brings in the knowledge of object saliency and relative contrast of the scene for more discriminative feature learning. Additionally, we introduce the idea of TCB (Transformer-CNN Bridge) to fuse high-level global contextual dependencies from a transformer backbone with local spatial features captured by a set of hierarchical CNN (convolutional neural network) layers. We implement TCB and Depth-CAR as multimodal attention-based projection functions to select the most informative features, which also improve training time and inference efficiency. Experimental results demonstrate that our proposed DGIQA model achieves state-of-the-art (SOTA) performance on both synthetic and authentic benchmark datasets. More importantly, DGIQA outperforms SOTA models on cross-dataset evaluations as well as in assessing natural image distortions such as low-light effects, hazy conditions, and lens flares.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
Arniqa: Learning Distortion Mani- fold for Image Quality Assessment
Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini, and Alberto Del Bimbo. Arniqa: Learning Distortion Mani- fold for Image Quality Assessment. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 189–198, 2024. 2, 6, 8, 15
work page 2024
-
[2]
Ancuti, Cosmin Ancuti, Radu Timofte, and Christophe De Vleeschouwer
Codruta O. Ancuti, Cosmin Ancuti, Radu Timofte, and Christophe De Vleeschouwer. I-HAZE: A Dehazing Bench- mark With Real Hazy and Haze-Free Indoor Images. In arXiv:1804.05091v1, 2018. 2, 7, 15
arXiv 2018
-
[3]
On the Use of Deep Learning for Blind Image Quality Assessment
Simone Bianco, Luigi Celona, Paolo Napoletano, and Rai- mondo Schettini. On the Use of Deep Learning for Blind Image Quality Assessment. Signal, Image and Video Pro- cessing, 12:355–362, 2018. 2
work page 2018
-
[4]
Deep Neural Net- works for No-Reference and Full-Reference Image Qual- ity Assessment
Sebastian Bosse, Dominique Maniry, Klaus-Robert M ¨uller, Thomas Wiegand, and Wojciech Samek. Deep Neural Net- works for No-Reference and Full-Reference Image Qual- ity Assessment. IEEE Transactions on image processing , 27(1):206–219, 2017. 2
work page 2017
-
[5]
Iglovikov, Eugene Khved- chenya, Alex Parinov, Mikhail Druzhinin, and Alexandr A
Alexander Buslaev, Vladimir I. Iglovikov, Eugene Khved- chenya, Alex Parinov, Mikhail Druzhinin, and Alexandr A. Kalinin. Albumentations: Fast and Flexible Image Augmen- tations. Information, 11(2), 2020. 5
work page 2020
-
[6]
TOPIQ: A Top-down Approach from Semantics to Distortions for Im- age Quality Assessment, 2023
Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu, Liang Liao, Wenxiu Sun, Qiong Yan, and Weisi Lin. TOPIQ: A Top-down Approach from Semantics to Distortions for Im- age Quality Assessment, 2023. 6
work page 2023
-
[7]
Progressively Complementarity- Aware Fusion Network for RGB-D Salient Object Detection
Hao Chen and Youfu Li. Progressively Complementarity- Aware Fusion Network for RGB-D Salient Object Detection. In 2018 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 3051–3060, 2018. 1
work page 2018
-
[8]
Hao Chen, Youfu Li, and Dan Su. Multi-modal fusion net- work with multi-scale multi-path and cross-modal interac- tions for RGB-D salient object detection. Pattern Recogni- tion, 86:376–385, 2019. 1
work page 2019
Show all 75 references
-
[9]
PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts
Zewen Chen, Haina Qin, Juan Wang, Chunfeng Yuan, Bing Li, Weiming Hu, and Liang Wang. PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts. In European Conference on Computer Vision, pages 247–264. Springer, 2025. 3
2025
-
[10]
Deep Retinex Decomposition for Low-Light Enhancement
Chen Wei, Wenjing Wang, Wenhan Yang, Jiaying Liu. Deep Retinex Decomposition for Low-Light Enhancement. In British Machine Vision Conference, 2018. 2, 7
2018
-
[11]
Flare7K: A Phenomenological Night- time Flare Removal Dataset
Yuekun Dai, Chongyi Li, Shangchen Zhou, Ruicheng Feng, and Chen Change Loy. Flare7K: A Phenomenological Night- time Flare Removal Dataset. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Bench- marks Track, 2022. 2, 7, 15
2022
-
[12]
Uni- versal Blind Image Quality Assessment Metrics Via Natural Scene Statistics and Multiple Kernel Learning
Xinbo Gao, Fei Gao, Dacheng Tao, and Xuelong Li. Uni- versal Blind Image Quality Assessment Metrics Via Natural Scene Statistics and Multiple Kernel Learning. IEEE Trans- actions on neural networks and learning systems , 24(12),
-
[13]
Massive On- line Crowdsourced Study of Subjective and Objective Pic- ture Quality
Deepti Ghadiyaram and Alan C Bovik. Massive On- line Crowdsourced Study of Subjective and Objective Pic- ture Quality. IEEE Transactions on Image Processing , 25(1):372–387, 2015. 2, 5, 12, 14
2015
-
[14]
Perceptual Quality Prediction on Authentically Distorted Images Using a Bag of Features Approach
Deepti Ghadiyaram and Alan C Bovik. Perceptual Quality Prediction on Authentically Distorted Images Using a Bag of Features Approach. Journal of vision, 17(1):32–32, 2017. 2
2017
-
[15]
No-Reference Image Quality Assessment Via Transformers, Relative Ranking, and Self-Consistency
S Alireza Golestaneh, Saba Dadsetan, and Kris M Kitani. No-Reference Image Quality Assessment Via Transformers, Relative Ranking, and Self-Consistency. In Proceedings of the IEEE/CVF winter conference on applications of com- puter vision, pages 1220–1230, 2022. 3, 4, 6
2022
-
[16]
”Magnitude, Precision, and Realism of Depth Perception in Stereoscopic Vision”
Paul B Hibbard, Alice E Haines, and Rebecca L Hornsey. ”Magnitude, Precision, and Realism of Depth Perception in Stereoscopic Vision”. Cogn. Res. Princ. Implic. , 2(1):25, May 2017. 1
2017
-
[17]
Image Quality Metrics: PSNR vs
Alain Hor ´e and Djemel Ziou. Image Quality Metrics: PSNR vs. SSIM. In 2010 20th International Conference on Pattern Recognition, pages 2366–2369, 2010. 15
2010
-
[18]
KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment
Vlad Hosu, Hanhe Lin, Tamas Sziranyi, and Dietmar Saupe. KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment. IEEE Trans- actions on Image Processing, 29:4041–4056, 2020. 2, 5
2020
-
[19]
Squeeze-and-Excitation Net- works
Jie Hu, Li Shen, and Gang Sun. Squeeze-and-Excitation Net- works. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7132–7141, 2018. 4
2018
-
[20]
SV AM: Saliency-guided Visual Attention Modeling by Autonomous Underwater Robots
Md Jahidul Islam, Ruobing Wang, and Junaed Sattar. SV AM: Saliency-guided Visual Attention Modeling by Autonomous Underwater Robots. In Robotics: Science and Systems (RSS), NY , USA, 2022. 6, 12
2022
-
[21]
GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer, 2024
Ding Jia, Jianyuan Guo, Kai Han, Han Wu, Chao Zhang, Chang Xu, and Xinghao Chen. GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer, 2024. 2
2024
-
[22]
Tongue Image Quality Assessment Based on a Deep Convolutional Neural Network
Tao Jiang, Xiao-juan Hu, Xing-hua Yao, Li-ping Tu, Jing- bin Huang, Xu-xiang Ma, Ji Cui, Qing-feng Wu, and Jia-tuo Xu. Tongue Image Quality Assessment Based on a Deep Convolutional Neural Network. BMC Medical Informatics and Decision Making, 21(1):147, 2021. 2
2021
-
[23]
Convo- lutional Neural Networks for No-Reference Image Quality Assessment
Le Kang, Peng Ye, Yi Li, and David Doermann. Convo- lutional Neural Networks for No-Reference Image Quality Assessment. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1733–1740,
-
[24]
Musiq: Multi-Scale Image Quality Transformer
Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-Scale Image Quality Transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5148–5157, 2021. 3, 6
2021
-
[25]
Divide and Conquer: Ill-Light Image Enhancement Via Hybrid Deep Network
Rizwan Khan, You Yang, Qiong Liu, and Zahid Hussain Qaisar. Divide and Conquer: Ill-Light Image Enhancement Via Hybrid Deep Network. Expert Systems with Applica- tions, 182:115034, 2021. 2, 7, 15
2021
-
[26]
Most Appar- ent Distortion: Full-Reference Image Quality Assessment and The Role of Strategy
Eric C Larson and Damon M Chandler. Most Appar- ent Distortion: Full-Reference Image Quality Assessment and The Role of Strategy. Journal of electronic imaging , 19(1):011006–011006, 2010. 2, 5, 14, 15
2010
-
[27]
Statistical Eval- uation of No-Reference Image Quality Assessment Metrics For Remote Sensing Images
Shuang Li, Zewei Yang, and Hongsheng Li. Statistical Eval- uation of No-Reference Image Quality Assessment Metrics For Remote Sensing Images. ISPRS International Journal of Geo-Information, 6(5):133, 2017. 1
2017
-
[28]
Swinir: Image Restoration Using Swin Transformer
Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. Swinir: Image Restoration Using Swin Transformer. In Proceedings of the IEEE/CVF international conference on computer vision , pages 1833– 1844, 2021. 1
2021
-
[29]
KADID-10k: A Large-Scale Artificially Distorted IQA Database
Hanhe Lin, Vlad Hosu, and Dietmar Saupe. KADID-10k: A Large-Scale Artificially Distorted IQA Database. InEleventh International Conference on Quality of Multimedia Experi- ence (QoMEX), pages 1–3. IEEE, 2019. 2, 5, 14, 15
2019
-
[30]
Hallucinated-IQA: No-Reference Image Quality Assessment Via Adversarial Learning
Kwan-Yee Lin and Guanxiang Wang. Hallucinated-IQA: No-Reference Image Quality Assessment Via Adversarial Learning. In Proceedings of the IEEE conference on com- puter vision and pattern recognition , pages 732–741, 2018. 2
2018
-
[31]
Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 5
2021
-
[32]
End-to-End Blind Im- age Quality Assessment Using Deep Neural Networks.IEEE Transactions on Image Processing, 27(3):1202–1213, 2017
Kede Ma, Wentao Liu, Kai Zhang, Zhengfang Duanmu, Zhou Wang, and Wangmeng Zuo. End-to-End Blind Im- age Quality Assessment Using Deep Neural Networks.IEEE Transactions on Image Processing, 27(3):1202–1213, 2017. 2
2017
-
[33]
No-Reference Image Qality Assessment in the Spa- tial Domain
Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-Reference Image Qality Assessment in the Spa- tial Domain. IEEE Transactions on image processing , 21(12):4695–4708, 2012. 2, 6
2012
-
[34]
Blind Im- age Quality Assessment: From Natural Scene Statistics to Perceptual Quality
Anush Krishna Moorthy and Alan Conrad Bovik. Blind Im- age Quality Assessment: From Natural Scene Statistics to Perceptual Quality. IEEE transactions on Image Processing, 20(12):3350–3364, 2011. 2
2011
-
[35]
Deep Multi-Scale Convolutional Neural Network for Dynamic Scene Deblurring
Seungjun Nah, Tae Hyun Kim, and Kyoung Mu Lee. Deep Multi-Scale Convolutional Neural Network for Dynamic Scene Deblurring. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3883–3891,
-
[36]
Image Database TID2013: Peculiarities, Results and Perspectives
Nikolay Ponomarenko, Lina Jin, Oleg Ieremeiev, Vladimir Lukin, Karen Egiazarian, Jaakko Astola, Benoit V ozel, Kacem Chehdi, Marco Carli, Federica Battisti, et al. Image Database TID2013: Peculiarities, Results and Perspectives. Signal processing: Image communication , 30:57–7...
2015
-
[37]
PieAPP: Perceptual Image-Error Assessment Through Pairwise Preference
Ekta Prashnani, Hong Cai, Yasamin Mostofi, and Pradeep Sen. PieAPP: Perceptual Image-Error Assessment Through Pairwise Preference. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 1808– 1817, 2018. 15
2018
-
[38]
Blind Image Quality Assessment: A Natural Scene Statis- tics Approach in the DCT Domain
Michele A Saad, Alan C Bovik, and Christophe Charrier. Blind Image Quality Assessment: A Natural Scene Statis- tics Approach in the DCT Domain. IEEE transactions on Image Processing, 21(8):3339–3352, 2012. 1, 2, 6
2012
-
[39]
Re- iqa: Unsupervised Learning for Image Quality Assessment In the Wild
Avinab Saha, Sandeep Mishra, and Alan C Bovik. Re- iqa: Unsupervised Learning for Image Quality Assessment In the Wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5846–5855,
-
[40]
Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In 2017 IEEE Interna- tional Conference on Computer Vision (ICCV) , pages 618– 626,...
2017
-
[41]
A Statistical Evaluation of Recent Full Reference Image Qual- ity Assessment Algorithms
Hamid R Sheikh, Muhammad F Sabir, and Alan C Bovik. A Statistical Evaluation of Recent Full Reference Image Qual- ity Assessment Algorithms. IEEE Transactions on image processing, 15(11):3440–3451, 2006. 2, 5, 12, 15
2006
-
[42]
Wood, Thomas San- ford, Ge Wang, and Pingkun Yan
Xinrui Song, Hanqing Chao, Xuanang Xu, Hengtao Guo, Sheng Xu, Baris Turkbey, Bradford J. Wood, Thomas San- ford, Ge Wang, and Pingkun Yan. Cross-Modal Attention for Multi-Modal Image Registration. Medical Image Analy- sis, 82:102612, 2022. 4
2022
-
[43]
Blindly Assess Image Qual- ity in the Wild Guided by a Self-Adaptive Hyper Network
Shaolin Su, Qingsen Yan, Yu Zhu, Cheng Zhang, Xin Ge, Jinqiu Sun, and Yanning Zhang. Blindly Assess Image Qual- ity in the Wild Guided by a Self-Adaptive Hyper Network. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 3667–3676, 2020. 2, 6
2020
-
[44]
NIMA: Neural Im- age Assessment
Hossein Talebi and Peyman Milanfar. NIMA: Neural Im- age Assessment. IEEE transactions on image processing , 27(8):3998–4011, 2018. 2
2018
-
[45]
Visualizing Data Using t-SNE
Laurens Van der Maaten and Geoffrey Hinton. Visualizing Data Using t-SNE. Journal of machine learning research , 9(11), 2008. 7
2008
-
[46]
Gomez, Łukasz Kaiser, and Il- lia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Il- lia Polosukhin. Attention is All You Need. In Advances in Neural Information Processing Systems , volume 30, pages 5998–6008. Curran Associates, Inc., 2017. 4
2017
-
[47]
Ex- ploring Clip for Assessing the Look and Feel of Images
Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. Ex- ploring Clip for Assessing the Look and Feel of Images. In Proceedings of the AAAI Conference on Artificial Intel- ligence, volume 37, pages 2555–2563, 2023. 3, 6
2023
- [48]
-
[49]
Real-esrgan: Training Real-world Blind Super-resolution With Pure Synthetic Data
Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-esrgan: Training Real-world Blind Super-resolution With Pure Synthetic Data. In Proceedings of the IEEE/CVF international conference on computer vision , pages 1905– 1914, 2021. 1
1905
-
[50]
Modern Image Quality Assessment
Zhou Wang and Alan Conrad Bovik. Modern Image Quality Assessment. PhD thesis, Springer, 2006. 1
2006
-
[51]
Why Is Im- age Quality Assessment So Difficult? In IEEE International conference on acoustics, speech, and signal processing, vol- ume 4, pages IV–3313
Zhou Wang, Alan C Bovik, and Ligang Lu. Why Is Im- age Quality Assessment So Difficult? In IEEE International conference on acoustics, speech, and signal processing, vol- ume 4, pages IV–3313. IEEE, 2002. 1
2002
-
[52]
Image Quality Assessment: From Error Visi- bility to Structural Similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image Quality Assessment: From Error Visi- bility to Structural Similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 5, 12, 15
2004
-
[53]
A Comprehensive Study of Multimodal Large Lan- guage Models for Image Quality Assessment
Tianhe Wu, Kede Ma, Jie Liang, Yujiu Yang, and Lei Zhang. A Comprehensive Study of Multimodal Large Lan- guage Models for Image Quality Assessment. arXiv preprint arXiv:2403.10854, 2024. 3
2024 arXiv
-
[54]
Do- main Fingerprints for No-Reference Image Quality Assess- ment
Weihao Xia, Yujiu Yang, Jing-Hao Xue, and Jing Xiao. Do- main Fingerprints for No-Reference Image Quality Assess- ment. IEEE Transactions on Circuits and Systems for Video Technology, 31(4):1332–1341, 2020. 2, 3
2020
-
[55]
Blind Image Quality Assessment Based on High Order Statistics Aggregation
Jingtao Xu, Peng Ye, Qiaohong Li, Haiqing Du, Yong Liu, and David Doermann. Blind Image Quality Assessment Based on High Order Statistics Aggregation. IEEE Trans- actions on Image Processing, 25(9):4444–4457, 2016. 2
2016
-
[56]
Boosting Image Quality Assessment through Efficient Transformer Adapta- tion with Local Feature Enhancement
Kangmin Xu, Liang Liao, Jing Xiao, Chaofeng Chen, Haon- ing Wu, Qiong Yan, and Weisi Lin. Boosting Image Quality Assessment through Efficient Transformer Adapta- tion with Local Feature Enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...
2024
-
[57]
Align-IQA: Aligning Im- age Quality Assessment Models with Diverse Human Pref- erences via Customizable Guidance
Junfeng Yang, Jing Fu, Zhen Zhang, Limei Liu, Qin Li, Wei Zhang, and Wenzhi Cao. Align-IQA: Aligning Im- age Quality Assessment Models with Diverse Human Pref- erences via Customizable Guidance. In Proceedings of the 32nd ACM International Conference on Multimedia , pages 1000...
2024
-
[58]
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10371–10381, 2024. 5, 12, 17
2024
-
[59]
Maniqa: Multi-Dimension Attention Network for No- Reference Image Quality Assessment
Sidi Yang, Tianhe Wu, Shuwei Shi, Shanshan Lao, Yuan Gong, Mingdeng Cao, Jiahao Wang, and Yujiu Yang. Maniqa: Multi-Dimension Attention Network for No- Reference Image Quality Assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pag...
2022
-
[60]
No-Reference Image Quality Assessment Using Visual Codebooks
Peng Ye and David Doermann. No-Reference Image Quality Assessment Using Visual Codebooks. IEEE Transactions on Image Processing, 21(7):3129–3138, 2012. 2
2012
-
[61]
Un- supervised Feature Learning Framework for No-Reference Image Quality Assessment
Peng Ye, Jayant Kumar, Le Kang, and David Doermann. Un- supervised Feature Learning Framework for No-Reference Image Quality Assessment. In IEEE conference on computer vision and pattern recognition , pages 1098–1105. IEEE,
-
[62]
From Patches to Pictures (PaQ-2-PiQ): Mapping the Perceptual Space of Pic- ture Quality
Zhenqiang Ying, Haoran Niu, Praful Gupta, Dhruv Maha- jan, Deepti Ghadiyaram, and Alan Bovik. From Patches to Pictures (PaQ-2-PiQ): Mapping the Perceptual Space of Pic- ture Quality. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3575–3585,
-
[63]
Transformer for Image Quality Assessment
Junyong You and Jari Korhonen. Transformer for Image Quality Assessment. In IEEE international conference on image processing (ICIP), pages 1389–1393. IEEE, 2021. 3
2021
-
[64]
Multi-Scale Context Aggregation By Dilated Convo- lutions
F Yu. Multi-Scale Context Aggregation By Dilated Convo- lutions. arXiv preprint arXiv:1511.07122, 2015. 3, 4
2015 arXiv
-
[65]
Perceptual Image Qual- ity Assessment: A Survey
Guangtao Zhai and Xiongkuo Min. Perceptual Image Qual- ity Assessment: A Survey. Science China Information Sci- ences, 63:1–52, 2020. 1, 5, 12
2020
-
[66]
A Feature- Enriched Completely Blind Image Quality Evaluator
Lin Zhang, Lei Zhang, and Alan C Bovik. A Feature- Enriched Completely Blind Image Quality Evaluator. IEEE Transactions on Image Processing, 24(8):2579–2591, 2015. 2
2015
-
[67]
FSIM: A Feature Similarity Index for Image Quality Assess- ment
Lin Zhang, Lei Zhang, Xuanqin Mou, and David Zhang. FSIM: A Feature Similarity Index for Image Quality Assess- ment. IEEE transactions on Image Processing, 20(8):2378– 2386, 2011. 15
2011
-
[68]
Feature Calibrating and Fusing Network for RGB-D Salient Object Detection
Qiang Zhang, Qi Qin, Yang Yang, Qiang Jiao, and Jungong Han. Feature Calibrating and Fusing Network for RGB-D Salient Object Detection. IEEE Transactions on Circuits and Systems for Video Technology, 34(3):1493–1507, 2024. 1
2024
-
[69]
The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR, 2018. 15
2018
-
[70]
Blind Image Quality Assessment Using a Deep Bi- linear Convolutional Neural Network
Weixia Zhang, Kede Ma, Jia Yan, Dexiang Deng, and Zhou Wang. Blind Image Quality Assessment Using a Deep Bi- linear Convolutional Neural Network. IEEE Transactions on Circuits and Systems for Video Technology, 30(1):36–47,
-
[71]
Blind Image Quality Assessment Via Vision- Language Correspondence: A Multitask Learning Perspec- tive
Weixia Zhang, Guangtao Zhai, Ying Wei, Xiaokang Yang, and Kede Ma. Blind Image Quality Assessment Via Vision- Language Correspondence: A Multitask Learning Perspec- tive. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , pages 14071–14081,
-
[72]
Blind Image Quality Assessment via Vision- Language Correspondence: A Multitask Learning Perspec- tive
Weixia Zhang, Guangtao Zhai, Ying Wei, Xiaokang Yang, and Kede Ma. Blind Image Quality Assessment via Vision- Language Correspondence: A Multitask Learning Perspec- tive. In IEEE Conference on Computer Vision and Pattern Recognition, pages 14071–14081, 2023. 6
2023
-
[73]
CMPFFNet: Cross-Modal and Progressive Feature Fusion Network for RGB-D Indoor Scene Semantic Segmentation
Wujie Zhou, Yuxiang Xiao, Weiqing Yan, and Lu Yu. CMPFFNet: Cross-Modal and Progressive Feature Fusion Network for RGB-D Indoor Scene Semantic Segmentation. IEEE Transactions on Automation Science and Engineering, 21(4):5523–5533, 2024. 2, 4
2024
-
[74]
MetaIQA: Deep Meta-Learning for No- Reference Image Quality Assessment
Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, and Guangming Shi. MetaIQA: Deep Meta-Learning for No- Reference Image Quality Assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14143–14152, 2020. 2, 5 Appendix A. Hyperpa...
2020
-
[75]
horseshoe
clearly demonstrate degraded structural information, re- flecting the impact of heavy noise on feature extraction. Figure 13 provides examples from the LIVEFB [62] dataset, which contains two inherent categories of distor- tions: Motion Blur and Voc emotic AVA. The depth maps ...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.