REVIEW 3 major objections 5 minor 71 references
Enhancing Content Representation for AR Image Quality Assessment Using Knowledge Distillation
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A knowledge-distilled transformer metric outperforms all prior AR image quality models on ARIQA.
desk verdict A plausible KD-based AR-IQA model with a new SOTA claim, but the evaluation is undercut by test-fold hyperparameter tuning and missing error bars. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a three-encoder transformer pipeline whose pivot is the class token. Three tiny vision transformers, pretrained with self-supervised self-distillation, encode the AR foreground, the background, and the superimposed image; for the first two, an MLP projection head attached to the class token predicts the content category (web, natural, or graphics for the foreground; indoor or outdoor for the background), and the projected tokens become teacher targets. A negative cosine similarity loss with a stop-gradient aligns the superimposed image's class token to both teacher targets, teaching the student encoder to keep foreground and background content separated even when the two are visually confusable. Around that pivot, distortions are quantified as $\ell^1$ shift representations, the patch-wise absolute differences between superimposed and reference embeddings, and two cross-attention decoders use the reference embeddings as queries and the shift representations as keys and values to generate quality-aware features that two regressors turn into a predicted MOS. The design is kept lightweight by using only the first two of the twelve transformer layers.
What would settle it
Fix all hyperparameters before evaluating, for instance by setting the blending weight to 0.5, using the first two ViT layers, and selecting regularization strengths on a development split, then retrain on the same five ARIQA folds and report SRCC; if the gap over ARIQA+'s 0.8124 disappears, the reported advantage depends on test-set-informed tuning rather than on the architecture.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that in full-reference AR image quality assessment, the bottleneck is not the distortion measure but the distorted image's representation: a frozen generic encoder fails on AR content because visual confusion pushes the superimposed image far from the training distribution. The solution is to fine-tune the three content-aware encoders and distill knowledge from the reference encoders into the superimposed encoder by maximizing the cosine similarity between the class tokens of the superimposed image and the projected class tokens of the foreground and background, using a stop-gradient to avoid collapsing the references. The distilled encoder is combined with $\ell^1$-based shift representations that capture patch-level residuals between the superimposed and reference embeddings, cross-attention decoders that query the references with the shifts, and lightweight regressors to produce the final MOS. In the paper's experiments, TransformAR-KD and TransformAR-KD+ outperform all compared classical and learned metrics across SRCC, KRCC, and PLCC on the five ARIQA folds, and the 50-fold evaluation reports even higher average correlations.
Load-bearing premise
The reported state-of-the-art numbers assume that the model's tuning choices, especially the 0.51 weight that blends the foreground and background quality scores and the decision to use only the first two transformer layers, were made without looking at the ARIQA test folds, because the paper says those settings came from iterative experimentation and performance evaluation.
Editorial extensions
If this is right
- On the standard five ARIQA folds, TransformAR-KD+ raises SRCC from 0.8124 to 0.8411, KRCC from 0.6184 to 0.6538, and PLCC from 0.8136 to 0.8416 relative to the previous best model, ARIQA+.
- On 50 random scene-disjoint folds, the best variant averages SRCC 0.8566, KRCC 0.6712, and PLCC 0.8580, so the advantage is not an artifact of a single split.
- The ablation study attributes the largest gains to the l1 shift representations, whose removal roughly halves accuracy, and to the cross-attention decoder, whose removal also drops all variants; label smoothing and the Huber loss add smaller, consistent gains.
- Because the architecture uses only the first two transformer layers, the metric remains lightweight, which matters for real-time or on-device AR quality monitoring.
- The class-token distillation also sharpens attention maps: the fine-tuned encoders focus on the AR object and spread attention across its parts instead of being pulled toward bright background regions like the sky.
Reading between the lines
- Because the classification heads and distillation share the same class token, the paper does not isolate whether the gain comes from semantic category knowledge or from any auxiliary supervision; a variant with shuffled labels would settle that.
- All reported numbers come from 560 stimuli built from 20 backgrounds and 20 foregrounds, so the ranking on novel AR content or on binocular AR displays remains untested; the paper's own future-work note singles out binocular visual confusion.
- A natural stress test is class-holdout evaluation, training on two of the three foreground categories and testing on the third, which would show whether the distilled class knowledge transfers or simply memorizes the training categories.
- If the hyperparameters were tuned on ARIQA's test folds, the true generalization would be optimistic; fixing the blending weight and the layer choice on a separate development set before touching the test folds would quantify that bias.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a full-reference deep image quality assessment method for augmented-reality images, called TransformAR, with two enhanced variants TransformAR-KD and TransformAR-KD+. The architecture uses a DINO-pretrained ViT-S/16 encoder (restricted to its first two layers) to embed the AR foreground, the real background, and the superimposed image; computes L1 shift representations between reference and distorted embeddings; passes these through cross-attention decoders; and regresses quality scores that are combined as pMOS = ζ·S_as + (1−ζ)·S_bs. Knowledge distillation aligns the superimposed image's class token with projections of the foreground and background class tokens, and the training loss combines Huber loss, negative cosine similarity, cross-entropy, and elastic-net regularization, with label smoothing added to the MOS targets. Experiments on the ARIQA dataset are reported on the five folds of Duan et al. and on an additional 50-fold split, with claims of state-of-the-art performance: TransformAR-KD+ reaches SRCC 0.8411, KRCC 0.6538, PLCC 0.8416 on the five folds, exceeding ARIQA+ (0.8124, 0.6184, 0.8136).
Significance. If the reported results are valid, the paper makes a useful contribution to an under-studied problem: AR image quality assessment with visual confusion. The use of self-supervised ViT features, a lightweight two-layer encoder, knowledge distillation from foreground/background content, and explicit handling of overfitting are all sensible design choices, and the ablation study indicates that each component contributes. The paper also provides qualitative attention-map and UMAP analyses that support the claimed mechanism of improved representation. However, the central state-of-the-art claim currently rests on an evaluation protocol in which several hyperparameters were selected after inspecting performance on the same test folds, and the 50-fold results lack baselines and error bars. These issues are load-bearing: without a cleaner protocol, the reported gains over ARIQA+ cannot be distinguished from selection bias. The contribution is nonetheless plausible and the methodological problems are addressable within the scope of a revision.
major comments (3)
- [Section II-D, Eq. (13)] The headline five-fold comparison in Table I is not reliable because hyperparameters appear to have been selected using test-fold performance. The paper states that ζ was 'empirically set to 0.51 after iterative experimentation and performance evaluation' (Section II-D), and similar empirical choices are reported for the Huber loss δ (Section III-A), the elastic-net α (Section III-B), the loss weights λ_i (Section III-B), and the number of ViT layers used (Section II-A2). Since the five ARIQA folds collectively contain all 560 stimuli, selecting these values by looking at five-fold results can directly inflate the reported SRCC gain over ARIQA+ (0.8411 vs. 0.8124). Please provide a validation protocol in which hyperparameters are chosen on training/validation folds only, and report per-fold results with means and standard deviations.
- [Section IV-A, Table I] The 50-fold rows report only TransformAR, TransformAR-KD, and TransformAR-KD+ (SRCC 0.8267, 0.8563, 0.8566), with no baseline methods and no variance information. The text claims that in the 50-fold setting 'TransformAR shows good performance, surpassing all previous methods', but no previous-method 50-fold numbers are given. Without mean±std for all methods and a paired significance test, the 50-fold results cannot support a superiority claim. Please add 50-fold results for ARIQA+, LPIPS, CFIQA, and other baselines, together with standard deviations or confidence intervals.
- [Section IV-B, Table II] The ablation study is not matched to the headline evaluation. Section IV-B states that the ablations use 'five folds from X', which are not necessarily the same five folds used by Duan et al. and reported in Table I. Consequently, the 'all combined' rows (e.g., TransformAR-KD SRCC 0.8637) are not the same experiment as the Table I result for TransformAR-KD (SRCC 0.8390), so the ablation does not validate the state-of-the-art claim. Please run the ablations on the same five folds as Table I, or provide a matched per-fold comparison so that removing a component can be compared directly with the full model on identical test data.
minor comments (5)
- [Section IV, opening paragraph] The text says 'Four evaluation metrics are used' but then lists only SRCC, KRCC, and PLCC; please either add the fourth metric or correct the sentence.
- [Section III-C] The label-smoothing mechanism is described as applied 'when the model begins to overfit', but no operational criterion is given for detecting that point. Please specify the epoch or validation-based rule so that the procedure is reproducible.
- [Section II-A2] The justification for using only the first two ViT layers is one sentence ('we have noticed that our framework is prone to overfitting when using all twelve transformer layers'). Since this choice affects the central results, please document the comparison that led to this decision and state whether it was made on training/validation data rather than on the ARIQA test folds.
- [Section II-B, Eq. (5)] The notation in Eq. (5) is unclear: the left-hand side pairs f^i_cls with f^i_j, but the symbols x^i_cls and x^i_j are not defined in the text. Please define these terms explicitly.
- [References] TransformAR is cited as reference [69], which is prior work by the same authors. Please clarify in the text which components are new in this manuscript relative to [69], since the current presentation blurs the boundary between the new contributions and the earlier conference paper.
Circularity Check
The reported SOTA is partly a selection artifact: the mixing weight ζ in Eq. (13) and other hyperparameters are tuned after performance evaluation on the same five ARIQA folds used for the headline numbers.
-
fitted input called prediction
[Section II-D, Eq. (13)]
"The next step involves aggregating these scores using a hyperparameter ζ, which yields the predicted Mean Opinion Score (pMOS) according to the following formula pMOS = ζS as + (1−ζ)Sbs. (13) Here, ζ is empirically set to 0.51 after iterative experimentation and performance evaluation."
The headline result (TransformAR-KD+ SRCC 0.8411 on the five Duan et al. folds, Table I) is reported on the same folds used to select ζ. Because ζ directly weights the two quality branches in pMOS, tuning it after 'performance evaluation' on those folds means the reported correlation is partly a selection outcome rather than an independent prediction. No held-out validation split is described; Section III-E states that the same five folds of [24] were used for comparison. Thus the improvement over ARIQA+ is not a free prediction of the model but includes the effect of fitting the mixing hyperparameter to the test folds.
-
fitted input called prediction
[Section II-A2]
"However, we have noticed that our framework is prone to overfitting when using all twelve transformer layers in the ViT-S/16. Therefore, we empirically chose only the first two layers, which makes our framework extremely lightweight."
The encoder depth is an architectural choice made after observing overfitting behavior on the ARIQA data, and other hyperparameters (δ, α, λ_i) are likewise reported as empirically set. Since the final configuration is selected using the same dataset on which the SOTA claim is made, the reported margin includes model-selection optimism. This is not a definitional equivalence, but it is a fitted input presented as part of a predictive architecture rather than as a validated configuration.
full rationale
The paper's central derivation chain is a standard encoder–decoder architecture with knowledge distillation, and most components (ViT features, l1 shift representations, cross-attention decoders, distillation losses) are independently motivated and not circular by construction. However, the final predicted MOS in Eq. (13) contains ζ, which the paper says is 'empirically set to 0.51 after iterative experimentation and performance evaluation.' Because the same five ARIQA folds are used both for this tuning and for the reported state-of-the-art comparison, the claimed margin over ARIQA+ is partially a selection artifact rather than a clean out-of-sample prediction. The choice of only the first two ViT layers and other empirically set hyperparameters reinforce this concern. The paper does not provide a validation split, error bars for the five-fold results, or a 50-fold comparison table for baseline methods, so the reader cannot separate genuine generalization from test-fold fitting. This is not full circularity—the quality features and distillation mechanism are not defined in terms of the target MOS—but it is enough to reduce the strength of the central 'prediction' claim.
Assumptions & free parameters
free parameters (6)
- zeta (mixing weight) =
0.51
- Huber loss delta =
1
- Elastic net alpha =
0.7
- Loss weights lambda =
lambda0=lambda1=lambda2=1, lambda3=0.05
- Number of ViT layers used =
2 (of 12)
- Label smoothing noise scale =
uniform(-1,1) * N(0,1)
assumptions (4)
- domain assumption DINO ViT-S/16 pretrained on ImageNet provides useful feature representations for AR images
- domain assumption The categorical labels (web, natural, graphics for foreground; indoor, outdoor for background) are useful auxiliary supervision for quality prediction
- domain assumption The l1 distance between patch embeddings from early ViT layers is a sufficient shift representation for perceptual distortion
- domain assumption Human MOS values in ARIQA are reliable ground truth
Cite this review
Pith. "Pith review of Enhancing Content Representation for AR Image Quality Assessment Using Knowledge Distillation." pith.science (2026). https://pith.science/paper/KUI7FLUB
@misc{pith2026241206003,
author = {Pith},
title = {Pith review of: Enhancing Content Representation for AR Image Quality Assessment Using Knowledge Distillation},
year = {2026},
howpublished = {\url{https://pith.science/paper/KUI7FLUB}},
note = {Machine review of arXiv:2412.06003}
}
read the original abstract
Augmented Reality (AR) is a major immersive media technology that enriches our perception of reality by overlaying digital content (the foreground) onto physical environments (the background). It has far-reaching applications, from entertainment and gaming to education, healthcare, and industrial training. Nevertheless, challenges such as visual confusion and classical distortions can result in user discomfort when using the technology. Evaluating AR quality of experience becomes essential to measure user satisfaction and engagement, facilitating the refinement necessary for creating immersive and robust experiences. Though, the scarcity of data and the distinctive characteristics of AR technology render the development of effective quality assessment metrics challenging. This paper presents a deep learning-based objective metric designed specifically for assessing image quality for AR scenarios. The approach entails four key steps, (1) fine-tuning a self-supervised pre-trained vision transformer to extract prominent features from reference images and distilling this knowledge to improve representations of distorted images, (2) quantifying distortions by computing shift representations, (3) employing cross-attention-based decoders to capture perceptual quality features, and (4) integrating regularization techniques and label smoothing to address the overfitting problem. To validate the proposed approach, we conduct extensive experiments on the ARIQA dataset. The results showcase the superior performance of our proposed approach across all model variants, namely TransformAR, TransformAR-KD, and TransformAR-KD+ in comparison to existing state-of-the-art methods.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
O. Cakmakci and J. Rolland, “Head-worn displays: a review,” Journal of display technology , vol. 2, no. 3, pp. 199–216, 2006
work page 2006
-
[2]
Augmented reality and virtual reality displays: perspectives and challenges,
T. Zhan, K. Yin, J. Xiong, Z. He, and S.-T. Wu, “Augmented reality and virtual reality displays: perspectives and challenges,” Iscience, vol. 23, no. 8, 2020
work page 2020
-
[3]
R. Vertucci, S. D’Onofrio, S. Ricciardi, and M. De Nino, “History of augmented reality,” in Springer Handbook of Augmented Reality . Springer, 2023, pp. 35–50
work page 2023
-
[4]
Deep-based quality assessment of medical images through domain adaptation,
M. Tliba, A. Sekhri, M. A. Kerkouri, and A. Chetouani, “Deep-based quality assessment of medical images through domain adaptation,” in 2022 IEEE International Conference on Image Processing (ICIP), 2022, pp. 3692–3696
work page 2022
-
[5]
International Telecommu- nication Union
(2021) ITU-T Recommendation G.1035. International Telecommu- nication Union. [Online]. Available: https://www.itu.int/rec/T-REC-G. 1035-202111-I/en
work page 2021
-
[6]
Qualinet white paper on definitions of immersive media experience (imex),
A. Perkis, C. Timmerer, S. Barakovi ´c, J. B. Husi ´c, S. Bech, S. Bosse, J. Botev, K. Brunnstr ¨om, L. Cruz, K. De Moor et al., “Qualinet white paper on definitions of immersive media experience (imex),” arXiv preprint arXiv:2007.07032, 2020
arXiv 2007
-
[7]
M. S. Van Gisbergen, Contextual connected media: How rearranging a media puzzle, brings virtual reality into being . NHTV , 2016
work page 2016
-
[8]
V . Mazin, M. J. Cree, L. Streeter, K. Nezhivleva, and A. Mozhaeva, “Research and application of the adaptive model of the human visual sys- tem for improving the effectiveness of objective video quality metrics,” in 2023 33rd Conference of Open Innovations Association (FRUCT) . IEEE, 2023, pp. 192–197
work page 2023
Show all 71 references
-
[9]
Dashrestreamer: Framework for creation of impaired video clips under realistic network conditions,
K. Hod ˇzi´c, M. Cosovic, S. Mrdovic, J. J. Quinlan, and D. Raca, “Dashrestreamer: Framework for creation of impaired video clips under realistic network conditions,” ACM Transactions on Multimedia Com- puting, Communications and Applications , 2018
2018
-
[10]
Iv- psnr—the objective quality metric for immersive video applications,
A. Dziembowski, D. Mieloch, J. Stankowski, and A. Grzelka, “Iv- psnr—the objective quality metric for immersive video applications,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 32, no. 11, pp. 7575–7591, 2022
2022
-
[11]
A bayesian quality-of-experience model for adaptive streaming videos,
Z. Duanmu, W. Liu, D. Chen, Z. Li, Z. Wang, Y . Wang, and W. Gao, “A bayesian quality-of-experience model for adaptive streaming videos,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 18, no. 3s, pp. 1–24, 2023
2023
-
[12]
Itu-t recommendation g.1036,
International Telecommunication Union, “Itu-t recommendation g.1036,” https://www.itu.int/rec/T-REC-G.1036-202207-I, 2022, accessed on Oc- tober 27, 2023
2022
-
[13]
Objective and subjective quality assessment of 360- degree images,
A. Sendjasni, “Objective and subjective quality assessment of 360- degree images,” Ph.D. dissertation, Universit´e de Poitiers and Norwegian University of Science and Technology, 2023
2023
-
[14]
Attentive deep image quality assessment for omnidirectional stitching,
H. Duan, X. Min, W. Sun, Y . Zhu, X.-P. Zhang, and G. Zhai, “Attentive deep image quality assessment for omnidirectional stitching,” IEEE Journal of Selected Topics in Signal Processing , 2023
2023
-
[15]
Perceptual quality assessment of omnidirectional audio-visual signals,
X. Zhu, H. Duan, Y . Cao, Y . Zhu, Y . Zhu, J. Liu, L. Chen, X. Min, and G. Zhai, “Perceptual quality assessment of omnidirectional audio-visual signals,” arXiv preprint arXiv:2307.10813 , 2023
2023 arXiv
-
[16]
Acr360: A dataset on subjective 360° video quality assessment using acr methods,
M. Elwardy, H.-J. Zepernick, Y . Hu, and T. M. C. Chu, “Acr360: A dataset on subjective 360° video quality assessment using acr methods,” in 16th International Conference on Signal Processing and Communi- cation System (ICSPCS) . IEEE, 2023, pp. 1–8
2023
-
[17]
Subjective quality assessment of stereoscopic omnidirectional image,
J. Xu, C. Lin, W. Zhou, and Z. Chen, “Subjective quality assessment of stereoscopic omnidirectional image,” in Advances in Multimedia Information Processing–PCM 2018: 19th Pacific-Rim Conference on Multimedia, Hefei, China, 2018, Proceedings, Part I 19 . Springer, 2018, pp. 589–599
2018
-
[18]
Study of 3D virtual reality picture quality,
M. Chen, Y . Jin, T. Goodall, X. Yu, and A. C. Bovik, “Study of 3D virtual reality picture quality,” IEEE Journal of Selected Topics in Signal Processing, vol. 14, no. 1, pp. 89–102, 2019
2019
-
[19]
3d22 mx: Performance subjective evalu- ation of 3D/stereoscopic image processing and analysis,
J. J. Moreno Escobar, E. Y . Aguilar del Villar, O. Morales Matamoros, and L. Chanona Hern ´andez, “3d22 mx: Performance subjective evalu- ation of 3D/stereoscopic image processing and analysis,” Mathematics, vol. 11, no. 1, p. 171, 2022
2022
-
[20]
Subjective quality database and objective study of compressed point clouds with 6dof head- mounted display,
X. Wu, Y . Zhang, C. Fan, J. Hou, and S. Kwong, “Subjective quality database and objective study of compressed point clouds with 6dof head- mounted display,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 12, pp. 4630–4644, 2021
2021
-
[21]
Subjective and objective visual quality assessment of textured 3D meshes,
J. Guo, V . Vidal, I. Cheng, A. Basu, A. Baskurt, and G. Lavoue, “Subjective and objective visual quality assessment of textured 3D meshes,” ACM Transactions on Applied Perception (TAP), vol. 14, no. 2, pp. 1–20, 2016
2016
-
[22]
Quality evaluation of 3D objects in mixed reality for different lighting conditions,
J. Guti ´errez, T. Vigier, and P. Le Callet, “Quality evaluation of 3D objects in mixed reality for different lighting conditions,” Electronic Imaging, vol. 32, pp. 1–7, 2020
2020
-
[23]
Towards subjective quality assessment of point cloud imaging in augmented reality,
E. Alexiou, E. Upenik, and T. Ebrahimi, “Towards subjective quality assessment of point cloud imaging in augmented reality,” in IEEE 19th Int. Workshop on Multimedia Signal Processing (MMSP), 2017, pp. 1–6
2017
-
[24]
Confusing image quality assessment: Toward better augmented reality experience,
H. Duan, X. Min, Y . Zhu et al., “Confusing image quality assessment: Toward better augmented reality experience,” IEEE Transactions on Image Processing, vol. 31, pp. 7206–7221, 2022
2022
-
[25]
Image quality assess- ment: from error visibility to structural similarity,
Z. Wang, A. Bovik, H. Sheikh, and E. Simoncelli, “Image quality assess- ment: from error visibility to structural similarity,” IEEE Transactions on Image Processing (TIP) , vol. 13, no. 4, pp. 600–612, 2004
2004
-
[26]
Multiscale structural similarity for image quality assessment,
Z. Wang, E. Simoncelli, and A. Bovik, “Multiscale structural similarity for image quality assessment,” in Proceedings of the Asilomar Confer- ence on Signals, Systems & Computers , vol. 2, 2003, pp. 1398–1402
2003
-
[27]
Information content weighting for perceptual image quality assessment,
Z. Wang and Q. Li, “Information content weighting for perceptual image quality assessment,” IEEE Transactions on Image Processing (TIP) , vol. 20, no. 5, pp. 1185–1198, 2010
2010
-
[28]
Fsim: A feature similarity index for image quality assessment,
L. Zhang, L. Zhang, X. Mou, and D. Zhang, “Fsim: A feature similarity index for image quality assessment,” IEEE Transactions on Image Processing (TIP), vol. 20, no. 8, pp. 2378–2386, 2011
2011
-
[29]
Simultaneous estimation of image quality and distortion via multi-task convolutional neural networks,
L. Kang, P. Ye, Y . Li, and D. Doermann, “Simultaneous estimation of image quality and distortion via multi-task convolutional neural networks,” in Proceedings of the IEEE international conference on image processing (ICIP) , 2015, pp. 2791–2795
2015
-
[30]
Image quality assessment: a sparse learning way,
Y . Yuan, Q. Guo, and X. Lu, “Image quality assessment: a sparse learning way,” Neurocomputing, vol. 159, pp. 227–241, 2015
2015
-
[31]
Deep neural networks for no- reference and full-reference image quality assessment,
S. Bosse, D. Maniry, K. M ¨uller et al. , “Deep neural networks for no- reference and full-reference image quality assessment,” IEEE Transac- tions on Image Processing (TIP) , vol. 27, no. 1, pp. 206–219, 2017
2017
-
[32]
The unreasonable effectiveness of deep features as a perceptual metric,
R. Zhang, P. Isola, A. Efros et al. , “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , Salt Lake City, UT, USA, 2018, pp. 586–595
2018
-
[33]
Squeezenet: Alexnet-level accuracy with 50x fewer param- eters and ¡ 0.5 mb model size,
F. N. Iandola, S. Han, M. Moskewicz, K. Ashraf, W. Dally, and K. Keutzer, “Squeezenet: Alexnet-level accuracy with 50x fewer param- eters and ¡ 0.5 mb model size,” in Proceedings of the IEEE Conference IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY 14 on Comput...
2017
-
[34]
Imagenet classification with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. Hinton, “Imagenet classification with deep convolutional neural networks,” in Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , Lake Tahoe, NV , USA, 2012, pp. 1097–1105
2012
-
[35]
Very deep convolutional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556 , 2014
2014 arXiv
-
[36]
Perceptual image quality assessment with transformers,
M. Cheon, S. Yoon, B. Kang, and J. Lee, “Perceptual image quality assessment with transformers,” in Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR) , Virtual Event, 2021, pp. 433–442
2021
-
[37]
No-reference image qual- ity assessment via transformers, relative ranking, and self-consistency,
S. A. Golestaneh, S. Dadsetan, and K. Kitani, “No-reference image qual- ity assessment via transformers, relative ranking, and self-consistency,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa Village, HI, USA, 2022, pp. 1220– 1230
2022
-
[38]
Image quality assessment with transformers and multi-metric fusion modules,
W. Jiang, L. Li, Y . Ma, Y . Zhai et al., “Image quality assessment with transformers and multi-metric fusion modules,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 2022, pp. 1805–1809
2022
-
[39]
Perceptual video coding based on ssim-inspired divisive normalization,
S. Wang, A. Rehman, Z. Wang, S. Ma, and W. Gao, “Perceptual video coding based on ssim-inspired divisive normalization,” IEEE Transactions on Image Processing, vol. 22, no. 4, pp. 1418–1429, 2012
2012
-
[40]
Perceptual quality assessment for multi-exposure image fusion,
K. Ma, K. Zeng, and Z. Wang, “Perceptual quality assessment for multi-exposure image fusion,” IEEE Transactions on Image Processing, vol. 24, no. 11, pp. 3345–3356, 2015
2015
-
[41]
Automatic contrast enhancement technology with saliency preservation,
K. Gu, G. Zhai, X. Yang, W. Zhang, and C. W. Chen, “Automatic contrast enhancement technology with saliency preservation,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 25, no. 9, pp. 1480–1494, 2014
2014
-
[42]
Image quality assessment based on a degradation model,
N. Damera-Venkata, T. Kite, W. Geisler et al., “Image quality assessment based on a degradation model,” IEEE Transactions on Image Processing (TIP), vol. 9, no. 4, pp. 636–650, 2000
2000
-
[43]
Image information and visual quality,
H. Sheikh and A. Bovik, “Image information and visual quality,” IEEE Transactions on Image Processing (TIP) , vol. 15, no. 2, pp. 430–444, 2006
2006
-
[44]
Image quality assessment based on gradient similarity,
A. Liu, W. Lin, and M. Narwaria, “Image quality assessment based on gradient similarity,” IEEE Transactions on Image Processing (TIP) , vol. 21, no. 4, pp. 1500–1512, 2011
2011
-
[45]
Gradient magnitude similarity deviation: A highly efficient perceptual image quality index,
W. Xue, L. Zhang, X. Mou, and A. Bovik, “Gradient magnitude similarity deviation: A highly efficient perceptual image quality index,” IEEE Transactions on Image Processing (TIP) , vol. 23, no. 2, pp. 684– 695, 2013
2013
-
[46]
Perceptual fidelity aware mean squared error,
W. Xue, X. Mou, L. Zhang, and X. Feng, “Perceptual fidelity aware mean squared error,” in Proceedings of the IEEE International Confer- ence on Computer Vision (ICCV), Sydney, Australia, 2013, pp. 705–712
2013
-
[47]
An efficient color image quality metric with local-tuned-global model,
K. Gu, G. Zhai, X. Yang, and W. Zhang, “An efficient color image quality metric with local-tuned-global model,” in Proceedings of the IEEE International Conference on Image Processing (ICIP) , Paris, France, 2014, pp. 506–510
2014
-
[48]
Vsi: A visual saliency-induced index for perceptual image quality assessment,
L. Zhang, Y . Shen, and H. Li, “Vsi: A visual saliency-induced index for perceptual image quality assessment,” IEEE Transactions on Image Processing, vol. 23, no. 10, pp. 4270–4281, 2014
2014
-
[49]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV , USA, 2016, pp. 770– 778
2016
-
[50]
Influence of affective image content on subjective quality assessment,
I. van der Linde and R. M. Doe, “Influence of affective image content on subjective quality assessment,” JOSA A, vol. 29, no. 9, pp. 1948–1955, 2012
1948
-
[51]
Do vision transformers see like convolutional neural networks?
M. Raghu, T. Unterthiner, S. Kornblith, C. Zhang, and A. Dosovitskiy, “Do vision transformers see like convolutional neural networks?” Ad- vances in Neural Information Processing Systems , vol. 34, pp. 12 116– 12 128, 2021
2021
-
[52]
Umap: Uniform manifold approximation and projection for dimension reduction,
L. McInnes, J. Healy, and J. Melville, “Umap: Uniform manifold approximation and projection for dimension reduction,” arXiv preprint arXiv:1802.03426, 2018
2018 arXiv
-
[53]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
-
[54]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020
2010 arXiv
-
[55]
To compress or not to compress–self- supervised learning and information theory: A review,
R. Shwartz-Ziv and Y . LeCun, “To compress or not to compress–self- supervised learning and information theory: A review,” arXiv preprint arXiv:2304.09355, 2023
2023 arXiv
-
[56]
A cookbook of self-supervised learning,
R. Balestriero, M. Ibrahim, V . Sobal, A. Morcos, S. Shekhar, T. Gold- stein, F. Bordes, A. Bardes, G. Mialon, Y . Tian et al., “A cookbook of self-supervised learning,” arXiv preprint arXiv:2304.12210 , 2023
2023 arXiv
-
[57]
A simple framework for contrastive learning of visual representations,
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in Proceedings of the International Conference on Machine Learning (ICML), Vienna, Austria, 2020, pp. 1597–1607
2020
-
[58]
Emerging properties in self-supervised vision transformers,
M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Virtual Event, 2021, pp. 9650–9660
2021
-
[59]
Barlow twins: Self- supervised learning via redundancy reduction,
J. Zbontar, L. Jing, I. Misra, Y . LeCun, and S. Deny, “Barlow twins: Self- supervised learning via redundancy reduction,” in Proceedings of the International Conference on Machine Learning (ICML) , Virtual Event, 2021, pp. 12 310–12 320
2021
-
[60]
Vicreg: Variance-invariance- covariance regularization for self-supervised learning,
A. Bardes, J. Ponce, and Y . LeCun, “Vicreg: Variance-invariance- covariance regularization for self-supervised learning,” arXiv preprint arXiv:2105.04906, 2021
2021 arXiv
-
[61]
Exploring simple siamese representation learning,
X. Chen and K. He, “Exploring simple siamese representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Nashville, TN, USA, 2021, pp. 15 750–15 758
2021
-
[62]
Regularization and variable selection via the elastic net,
H. Zou and T. Hastie, “Regularization and variable selection via the elastic net,” Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 67, no. 2, pp. 301–320, 2005
2005
-
[63]
Robust estimation of a location parameter,
P. Huber, “Robust estimation of a location parameter,” in Breakthroughs in statistics: Methodology and distribution . Springer, 1992, pp. 492– 518
1992
-
[64]
Decoupled weight decay regularization,
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101 , 2017
2017 arXiv
-
[65]
Mean deviation similarity index: Efficient and reliable full-reference image quality evaluator,
H. Z. Nafchi, A. Shahkolaei, R. Hedjam, and M. Cheriet, “Mean deviation similarity index: Efficient and reliable full-reference image quality evaluator,” Ieee Access, vol. 4, pp. 5579–5590, 2016
2016
-
[66]
Image quality assessment based on dct subband similarity,
A. Balanov, A. Schwartz, Y . Moshe, and N. Peleg, “Image quality assessment based on dct subband similarity,” in 2015 IEEE international conference on image processing (ICIP) . IEEE, 2015, pp. 2105–2109
2015
-
[67]
A Haar wavelet-based perceptual similarity index for image quality assessment,
R. Reisenhofer, S. Bosse, G. Kutyniok, and T. Wiegand, “A Haar wavelet-based perceptual similarity index for image quality assessment,” Signal Processing: Image Communication , vol. 61, pp. 33–43, 2018
2018
-
[68]
Sr-sim: A fast and high performance iqa index based on spectral residual,
L. Zhang and H. Li, “Sr-sim: A fast and high performance iqa index based on spectral residual,” in 19th IEEE international conference on image processing. IEEE, 2012, pp. 1473–1476
2012
-
[69]
Towards light-weight transformer-based quality assessment metric for augmented reality,
A. Sekhri, S. A. Amirshahi, and M.-C. Larabi, “Towards light-weight transformer-based quality assessment metric for augmented reality,” in 2024 IEEE 26th International Workshop on Multimedia Signal Process- ing (MMSP), 2024, pp. 1–6
2024
-
[70]
Libsvm: a library for support vector machines,
C. Chang and C. Lin, “Libsvm: a library for support vector machines,” ACM Transactions on Intelligent Systems and Technology , vol. 2, no. 3, pp. 1–27, 2011
2011
-
[71]
Pytorch image quality: Metrics for image quality assessment,
S. Kastryulin, J. Zakirov, D. Prokopenko, and D. V . Dylov, “Pytorch image quality: Metrics for image quality assessment,” 2022. [Online]. Available: https://arxiv.org/abs/2208.14818
2022 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.