REVIEW 5 major objections 4 minor 105 references
Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality
T0 review · 5 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper argues that Mean Opinion Score can no longer be the sole supervisory signal for multimedia quality assessment, and that models must integrate context-awareness, reasoning, and multimodality.
desk verdict A clear, well-written position paper that usefully synthesizes MOS's known limitations and proposes a concrete annotation roadmap, but overstates the 'must' without testing the feasibility of its own proposal. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Mean Opinion Score itself, treated as a reductive representation of human judgment, and the proposed replacement is a structured prediction paradigm built on three pillars: context-awareness (explicit metadata conditioning such as task type and device), reasoning (chain-of-thought justifications grounded in visual evidence), and multimodality (joint embedding of visual features, text, saliency cues, and artifact maps). The pillars are meant to do the work that MOS cannot: expose semantic failures, preserve rater disagreement, support personalization, enable human-in-the-loop verification in medical and safety-critical settings, and provide actionable feedback. The argument is carried by empirical citations showing MOS-trained models' brittleness plus a roadmap for benchmarks that annotate contexts, personas, rationales, semantic flags, and rating distributions.
What would settle it
Train a no-reference quality model on existing MOS labels and a second model on the proposed richer annotations, then score both on a held-out set of AI-generated images with known semantic errors. If the richer model does not detect meaningfully more semantic failures or align better with expert rationales, the central claim is refuted.
Extended reading notes
Core claim
The central claim is that MOS, while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. More precisely, the paper contends that a model's output should be a structured judgment, a contextualized score conditioned on task and user intent, a natural-language rationale, a visual attention map, and an artifact mask, rather than a single number. The paper grounds this in documented failure modes: MOS-trained models misclassify semantically broken AI-generated images as high quality, lose substantial Spearman correlation in cross-dataset tests, cannot distinguish acceptable from unacceptable degradations across use cases, and discard rater disagreement that is itself informative. It then argues that context-awareness, reasoning, and multimodality are not optional enhancements but necessary foundations for robust, human-aligned quality systems.
Load-bearing premise
The load-bearing premise is that richer annotations, such as context metadata, persona-conditioned scores, rationales, semantic flags, and rating distributions, can be collected at scale without introducing new biases and that models trained on them will actually beat MOS-supervised models on robustness and human alignment.
Editorial extensions
If this is right
- Benchmarks would need to add task metadata, persona-conditioned scores, structured rationales, semantic-error flags, and full rating distributions instead of a single mean.
- Evaluation would move beyond PLCC and SROCC to include semantic alignment, reasoning fidelity, contextual sensitivity, and uncertainty calibration.
- Quality models would output a contextualized score plus a rationale, attention map, and artifact mask, enabling debugging and human-in-the-loop workflows.
- In constrained settings a pure MOS fast path remains acceptable, while high-stakes use cases would route through the full reasoning path.
- New datasets and shared challenges organized along these lines would let the field test semantic robustness and explainability directly.
Reading between the lines
- Not tested in the paper: the proposal could be checked cheaply by re-annotating a small slice of an existing quality database with personas and rationales; if annotators cannot agree on rationales, the reasoning pillar becomes harder to sustain.
- Editorial extension: if richer annotations do scale, quality assessment could converge with general vision-language reasoning benchmarks, where quality becomes one more capability of instruction-following models rather than a specialized regression head.
- Editorial extension: preserving rating distributions could turn quality models into disagreement detectors, flagging divisive content for platform moderation or medical second opinions, a use case the paper mentions but does not develop.
- The paper positions the shift as urgent, but a decisive empirical test, whether a reasoning, context-conditioned model beats a MOS-trained model on out-of-distribution semantic failures, is left as future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that Mean Opinion Score (MOS), the scalar average used to supervise quality assessment models, is no longer sufficient as the sole supervisory signal. The authors critique MOS along several axes—generalization and semantic blindness, the limitations of correlation-based evaluation, the collapse of inter-subject disagreement, and the absence of interpretability—and then propose a modeling paradigm built on three pillars: context-awareness, reasoning, and multimodality. They outline structured outputs (contextualized score, rationale, attention map, artifact mask), benchmark reforms (multi-condition and persona-conditioned annotations, rationales, distributional labels, scalable collection infrastructure), and new evaluation metrics. No new experiments or datasets are presented; the argument is supported by citations to existing studies.
Significance. The paper offers a useful synthesis of a growing literature criticizing MOS and a coherent roadmap for reform. Its strengths are the catalog of documented failure modes and the concrete, structured proposal for richer annotations and outputs. However, the positive proposal is normative: the 'must' claim depends on an untested empirical premise that the proposed labels can be collected at scale and that models trained on them will improve robustness and human alignment. If that premise is validated by future work, the roadmap could influence benchmark design; as it stands, the paper is better read as a position statement than as an established finding. The explicit enumeration of failure-mode categories and structured output modalities is valuable for framing community discussion.
major comments (5)
- [§4–§5 (esp. §4.4, §5.1, §5.4)] The central normative claim that quality assessment models 'must integrate' context, reasoning, and multimodality rests on an unverified premise: that the annotation protocols of §5.1 (multi-condition labels, persona-conditioned ratings, structured rationales, semantic-error flags, distributional labels) can be collected at scale without introducing new biases, and that models trained with the §4.4 structured objectives will actually improve robustness and human alignment. No pilot, proof-of-concept, or even a synthetic demonstration is provided. The statement in §5.4 that these reforms 'are not speculative enhancements' is therefore an assertion rather than a supported conclusion, and the 'must' in the Abstract and §1 is stronger than the evidence warrants.
- [§5.3] The 'Reasoning Metrics' paragraph names coherence, validity, groundness, factuality, and utility but does not define any of them for the quality-assessment setting, does not specify how they would be computed for QA rationales, and provides no evidence that they correlate with human quality judgments or with perceptual quality. Citing QAFactEval [95] and a survey [96] establishes that general reasoning metrics exist, but not that they transfer to the proposed structured outputs of §4.4.
- [§5.2] The scalable data-collection claims are supported by references from outside multimedia quality assessment: the 'up to 40%' label-efficiency improvement cites a general active-learning survey [91], gamification evidence comes from image-description tasks [92], and the LEAD inter-rater reliability figure is from a psychiatric diagnostic panel [93]. None of these establishes QA-specific feasibility, so the scalability premise of the roadmap is currently unsubstantiated.
- [§6] The rebuttal concedes that 'in constrained scenarios, pure MOS suffices.' This directly qualifies the universal 'must' formulated in the Abstract and §1. The paper should either restrict its central claim to high-stakes and context-sensitive applications, where the argument is strongest, or provide clear criteria for when the additional capabilities are necessary; otherwise the thesis oscillates between a strong and a weak reading.
- [§3.1–3.2] Several quantitative claims supporting the critique are imported from preprints or secondary reports without error bars, confidence intervals, or significance information: the 25% cross-dataset SROCC drop (§3.1), the 0.15 SROCC gain of a CLIP-guided NR model (§3.1), the 30% drop for Re-IQA (§3.2), and the PLCC≈0.92-to-SROCC<0.70 contrast (§3.2). For a position paper these numbers are illustrative, but presenting them as precise figures gives a false sense of measurement certainty; please present them as approximate ranges or clearly note the original conditions and variability.
minor comments (4)
- [§3.3, §3.5, §2] There are several typos and missing spaces: 'somerecent works' (§3.3), 'adress' (§3.5), and 'LLaV A' (reference [44] in §2).
- [Figure 1 caption] The caption reads 'a reduced pipeline from our paradigm'; 'reduced' appears to mean 'simplified,' which would be clearer.
- [§5.1] The phrase 'and explanations to subject before experience' is grammatically unclear; likely 'provide explanations to the subject before the experiment/rating task' is intended.
- [References] The ISO 9241 reference appears as a footnote URL rather than a formal reference entry, which is inconsistent with the reference style used elsewhere; consider adding it to the bibliography.
Circularity Check
No significant circularity: the paper is a literature-based position statement, and the paper's self-citations are background support rather than load-bearing premises.
full rationale
This is a position paper rather than a derivation: it contains no equations, no fitted parameters, no benchmark results of its own, and no prediction that reduces to an input. The central claim that modern quality assessment models must integrate context-awareness, reasoning, and multimodality is an argued normative position supported by a broad review of external literature; it is not derived from the paper's own definitions, from any fitted quantity, or from a self-citation chain. The self-citations that appear (refs. 11, 85, 86, 87, 102) are used only as supporting examples for background claims: medical imaging is high-stakes, saliency and scanpath cues are relevant modalities, and medical workflows use explainability. None of these supplies the load-bearing premise that MOS is insufficient or that the three-pillar framework is necessary. The roadmap in Sections 4 and 5 is explicitly a proposal with no proof-of-concept; that is an evidence gap or correctness risk, not circularity. Section 6 even concedes that 'in constrained scenarios, pure MOS suffices,' which qualifies the universal 'must' but does not make the argument circular. Thus the paper is self-contained in its argumentative structure, and no circular step can be identified.
Assumptions & free parameters
assumptions (3)
- domain assumption MOS is the dominant supervisory signal in multimedia quality assessment and reduces human judgments to a single scalar.
- domain assumption The empirical findings cited, such as a 25% cross-dataset SROCC drop and a 0.15 SROCC improvement with CLIP guidance, are accurate and representative.
- ad hoc to paper Models trained with context, reasoning, and multimodal supervision will be more robust, interpretable, and human-aligned than MOS-only models.
Cite this review
Pith. "Pith review of Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality." pith.science (2026). https://pith.science/paper/X255IZDA
@misc{pith2026250519696,
author = {Pith},
title = {Pith review of: Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality},
year = {2026},
howpublished = {\url{https://pith.science/paper/X255IZDA}},
note = {Machine review of arXiv:2505.19696}
}
read the original abstract
This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensitive human judgments to a single scalar, obscuring semantic failures, user intent, and the rationale behind quality decisions. We contend that modern quality assessment models must integrate three interdependent capabilities: (1) context-awareness, to adapt evaluations to task-specific goals and viewing conditions; (2) reasoning, to produce interpretable, evidence-grounded justifications for quality judgments; and (3) multimodality, to align perceptual and semantic cues using vision-language models. We critique the limitations of current MOS-centric benchmarks and propose a roadmap for reform: richer datasets with contextual metadata and expert rationales, and new evaluation metrics that assess semantic alignment, reasoning fidelity, and contextual sensitivity. By reframing quality assessment as a contextual, explainable, and multimodal modeling task, we aim to catalyze a shift toward more robust, human-aligned, and trustworthy evaluation systems.
Figures
Reference graph
Works this paper leans on
-
[95]
QAFactEval: Improved QA-based factual consistency evaluation for summarization,
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong, “QAFactEval: Improved QA-based factual consistency evaluation for summarization,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir ...
work page 2022
-
[96]
Evaluating step-by-step reasoning traces: A survey,
Jinu Lee and J. Hockenmaier, “Evaluating step-by-step reasoning traces: A survey,” ArXiv, vol. abs/2502.12289, 2025
arXiv 2025
-
[91]
A survey on deep active learning: Recent advances and new frontiers,
Dongyuan Li, Zhen Wang, Yankai Chen, Renhe Jiang, Weiping Ding, and Manabu Okumura, “A survey on deep active learning: Recent advances and new frontiers,” IEEE Transactions on Neural Networks and Learning Systems, 2024
work page 2024
-
[92]
Tomislav Ivanjko, “Crowdsourcing image descriptions using gamification: a comparison between game-generated labels and professional descriptors,” in 2019 42nd International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO). IEEE, 2019, pp. 537–541
work page 2019
-
[93]
Veerle C. Eijsbroek, Katarina Kjell, H. Andrew Schwartz, Jan R. Boehnke, Eiko I. Fried, Daniel N. Klein, Peik Gustafsson, Isabelle Augenstein, Patrick M.M. Bossuyt, and Oscar N.E. Kjell, “The leading guideline: Reporting standards for expert panel, best-estimate diagnosis, and longitudinal expert all data (lead) methods,” Comprehensive Psychiatry, p. 152603, 2025
work page 2025
-
[1]
Audio-visual multimedia quality assessment: A comprehensive survey,
Zahid Akhtar and Tiago H Falk, “Audio-visual multimedia quality assessment: A comprehensive survey,” IEEE access, vol. 5, pp. 21090–21117, 2017
2017
-
[2]
Perceptual image quality assessment: a survey,
Guangtao Zhai and Xiongkuo Min, “Perceptual image quality assessment: a survey,” Science China Information Sciences, vol. 63, pp. 1–52, 2020
2020
-
[3]
Perceptual video quality assessment: A survey,
Xiongkuo Min, Huiyu Duan, Wei Sun, Yucheng Zhu, and Guangtao Zhai, “Perceptual video quality assessment: A survey,” Science China Information Sciences, vol. 67, no. 11, pp. 211301, 2024
2024
Show all 105 references
-
[4]
Image aesthetic assessment: An experimental survey,
Yubin Deng, Chen Change Loy, and Xiaoou Tang, “Image aesthetic assessment: An experimental survey,”IEEE Signal Processing Magazine, vol. 34, no. 4, pp. 80–106, 2017
2017
-
[5]
A brief survey on adaptive video streaming quality assessment,
Wei Zhou, Xiongkuo Min, Hong Li, and Qiuping Jiang, “A brief survey on adaptive video streaming quality assessment,” Journal of Visual Communication and Image Representation, vol. 86, pp. 103526, 2022
2022
-
[6]
Real-time quality-and energy-aware bitrate ladder construction for live video streaming,
Mohammad Ghasempour, Hadi Amirpour, and Christian Timmerer, “Real-time quality-and energy-aware bitrate ladder construction for live video streaming,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems, 2025
2025
-
[7]
Telepresence video quality assessment,
Zhenqiang Ying, Deepti Ghadiyaram, and Alan Bovik, “Telepresence video quality assessment,” in European Conference on Computer Vision. Springer, 2022, pp. 327–347
2022
-
[8]
A systematic literature review: Real-time 3d reconstruction method for telepresence system,
Fazliaty Edora Fadzli, Ajune Wanis Ismail, and Shafina Abd Karim Ishigaki, “A systematic literature review: Real-time 3d reconstruction method for telepresence system,” Plos one, vol. 18, no. 11, pp. e0287155, 2023
2023
-
[9]
Quality assessment of videos on social media platforms related to gestational diabetes mellitus in china: A cross-section study,
Qin-Yu Cai, Jing Tang, Si-Zhe Meng, Yi Sun, Xia Lan, and Tai-Hang Liu, “Quality assessment of videos on social media platforms related to gestational diabetes mellitus in china: A cross-section study,” Heliyon, vol. 10, no. 7, pp. e29020, 2024
2024
-
[10]
Robust dual-modal image quality assessment aware deep learning network for traffic targets detection of autonomous vehicles,
Keke Geng, Ge Dong, and Wenhan Huang, “Robust dual-modal image quality assessment aware deep learning network for traffic targets detection of autonomous vehicles,” Multimedia Tools and Applications, vol. 81, no. 5, pp. 6801–6826, 2022
2022
-
[11]
Deep-based quality assessment of medical images through domain adaptation,
Marouane Tliba, Aymen Sekhri, Mohamed Amine Kerkouri, and Aladine Chetouani, “Deep-based quality assessment of medical images through domain adaptation,” in 2022 IEEE International Conference on Image Processing (ICIP), 2022, pp. 3692–3696
2022
-
[12]
Representation learning optimization for 3d point cloud quality assessment without reference,
Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, and Frederic Dufaux, “Representation learning optimization for 3d point cloud quality assessment without reference,” 10 2022, pp. 3702–3706
2022
-
[13]
Recommendation p.910: Subjective video quality assessment methods for multimedia applications,
International Telecommunication Union Telecommunication Standardization Sector (ITU-T), “Recommendation p.910: Subjective video quality assessment methods for multimedia applications,” https://www.itu.int/rec/T-REC-P. 910-202310-I/en, 2023, ITU-T Recommendation P.910 (10/23)
2023
-
[14]
Mean opinion score (mos) revisited: methods and applications, limitations and alternatives,
Robert C Streijl, Stefan Winkler, and David S Hands, “Mean opinion score (mos) revisited: methods and applications, limitations and alternatives,” Multimedia Systems, vol. 22, no. 2, pp. 213–227, 2016
2016
-
[15]
William james, gustav fechner, and early psychophysics,
Stephanie L Hawkins, “William james, gustav fechner, and early psychophysics,” Frontiers in physiology, vol. 2, pp. 68, 2011
2011
-
[16]
A brief review of the history and application of psychometrics and scaling to image quality assessment,
Norman Burningham, “A brief review of the history and application of psychometrics and scaling to image quality assessment,” in PICS, 1999, pp. 169–172
1999
-
[17]
Just noticeable difference,
Melissa K Stern and James H Johnson, “Just noticeable difference,” The corsini encyclopedia of psychology, pp. 1–2, 2010
2010
-
[18]
From pairwise comparisons and rating to a unified quality scale,
Maria Perez-Ortiz, Aliaksei Mikhailiuk, Emin Zerman, Vedad Hulusic, Giuseppe Valenzise, and Rafał K Mantiuk, “From pairwise comparisons and rating to a unified quality scale,” IEEE Transactions on Image Processing, vol. 29, pp. 1139–1151, 2019
2019
-
[19]
A law of comparative judgment,
Louis L Thurstone, “A law of comparative judgment,” in Scaling, pp. 81–92. Routledge, 2017
2017
-
[20]
Recommendation bt.500: Methodology for the subjective assessment of the quality of television pictures,
International Telecommunication Union Radiocommunication Sector (ITU-R), “Recommendation bt.500: Methodology for the subjective assessment of the quality of television pictures,” https://www.itu.int/rec/R-REC-BT.500 , 2000, ITU-R Recommendation BT.500-10
2000
-
[21]
Subjective and objective quality assessment of image: A survey,
Pedram Mohammadi, Abbas Ebrahimi-Moghadam, and Shahram Shirani, “Subjective and objective quality assessment of image: A survey,” arXiv preprint arXiv:1406.7799, 2014
2014 arXiv
-
[22]
Experimental comparison of psnr and ssim metrics for video quality estimation,
Z. Kotevski and P. Mitrevski, “Experimental comparison of psnr and ssim metrics for video quality estimation,” in ICT Innovations 2009, Danco Davcev and Juan M. Gómez, Eds., pp. 205–213. Springer, Berlin, Heidelberg, 2010
2009
-
[23]
Image qualityassessment: From errorvisibilitytostructural similarity,
BovikAC WangZhou, HR Sheikh, et al., “Image qualityassessment: From errorvisibilitytostructural similarity,” IEEE Transon ImageProcessing, vol. 13, no. 4, pp. 600, 2004
2004
-
[24]
Toward a practical perceptual video quality metric, 2016,
Zhi Li, Anne Aaron, Ioannis Katsavounidis, Anush Moorthy, and Megha Manohara, “Toward a practical perceptual video quality metric, 2016,” Dostupno na: http://techblog. netflix. com/2016/06/toward-practical-perceptual-video. html [16.8. 2022.], 2016. 10
2016
-
[25]
Reduced-reference image quality assessment based on perceptual image hashing,
Xudong Lv and Z. Jane Wang, “Reduced-reference image quality assessment based on perceptual image hashing,” 2009 16th IEEE International Conference on Image Processing (ICIP), pp. 4361–4364, 2009
2009
-
[26]
Reduced reference image quality assessment via sub-image similarity based redundancy measurement,
Xuanqin Mou, Wufeng Xue, and Lei Zhang, “Reduced reference image quality assessment via sub-image similarity based redundancy measurement,” in Electronic imaging, 2012
2012
-
[27]
No-reference image quality assessment in the spatial domain,
Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik, “No-reference image quality assessment in the spatial domain,” IEEE Transactions on Image Processing, vol. 21, no. 12, pp. 4695–4708, 2012
2012
-
[28]
Making a “completely blind
Anish Mittal, Rajiv Soundararajan, and Alan C Bovik, “Making a “completely blind” image quality analyzer,” IEEE Signal processing letters, vol. 20, no. 3, pp. 209–212, 2012
2012
-
[29]
Arniqa: Learning distortion manifold for image quality assessment,
Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini, and Alberto Del Bimbo, “Arniqa: Learning distortion manifold for image quality assessment,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 189–198
2024
-
[30]
Blind image quality assessment using a deep bilinear convolutional neural network,
Weixia Zhang, Kede Ma, Jia Yan, Dexiang Deng, and Zhou Wang, “Blind image quality assessment using a deep bilinear convolutional neural network,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, pp. 36–47, 2019
2019
-
[31]
Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,
Vlad Hosu, Hanhe Lin, Tamas Sziranyi, and Dietmar Saupe, “Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,” IEEE Transactions on Image Processing, vol. 29, pp. 4041–4056, 2020
2020
-
[32]
The unreasonable effectiveness of deep features as a perceptual metric,
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 586–595
2018
-
[33]
Topiq: A top-down approach from semantics to distortions for image quality assessment,
Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu, Liang Liao, Wenxiu Sun, Qiong Yan, and Weisi Lin, “Topiq: A top-down approach from semantics to distortions for image quality assessment,” IEEE Transactions on Image Processing, 2024
2024
-
[34]
Musiq: Multi-scale image quality transformer,
Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang, “Musiq: Multi-scale image quality transformer,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5148–5157
2021
-
[35]
Perceptual image quality assessment with transform- ers,
Manri Cheon, Sung-Jun Yoon, Byungyeon Kang, and Junwoo Lee, “Perceptual image quality assessment with transform- ers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2021, pp. 433–442
2021
-
[36]
Re-iqa: Unsupervised learning for image quality assessment in the wild,
Avinab Saha, Sandeep Mishra, and Alan C Bovik, “Re-iqa: Unsupervised learning for image quality assessment in the wild,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 5846–5855
2023
-
[37]
Learning transferable visual models from natural language supervision,
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning....
2021
-
[38]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in International conference on machine learning . PMLR, 2022, pp. 12888–12900
2022
-
[39]
Benchmark evaluations, applications, and challenges of large vision language models: A survey,
Zongxia Li, Xiyang Wu, Hongyang Du, Huy Nghiem, and Guangyao Shi, “Benchmark evaluations, applications, and challenges of large vision language models: A survey,” arXiv preprint arXiv:2501.02189, vol. 1, 2025
2025 arXiv
-
[40]
Exploring clip for assessing the look and feel of images,
Jianyi Wang, Kelvin CK Chan, and Chen Change Loy, “Exploring clip for assessing the look and feel of images,” in AAAI, 2023
2023
-
[41]
Quality-aware image-text alignment for real-world image quality assessment,
Lorenzo Agnolucci, Leonardo Galteri, and Marco Bertini, “Quality-aware image-text alignment for real-world image quality assessment,” arXiv preprint arXiv:2403.11176, vol. 5, no. 6, 2024
2024 arXiv
-
[42]
Q-align: Teaching lmms for visual scoring via discrete text-defined levels,
Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, et al., “Q-align: Teaching lmms for visual scoring via discrete text-defined levels,” arXiv preprint arXiv:2312.17090, 2023
2023 arXiv
-
[43]
Flamingo: a visual language model for few-shot learning,
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al., “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems, vol. ...
2022
-
[44]
Visual instruction tuning,
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee, “Visual instruction tuning,” Advances in neural information processing systems, vol. 36, pp. 34892–34916, 2023
2023
-
[45]
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond,
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou, “Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond,” arXiv preprint arXiv:2308.12966, 2023
2023 arXiv
-
[46]
A comprehensive study of multimodal large language models for image quality assessment,
Tianhe Wu, Kede Ma, Jie Liang, Yujiu Yang, and Lei Zhang, “A comprehensive study of multimodal large language models for image quality assessment,” in European Conference on Computer Vision. Springer, 2024, pp. 143–160
2024
-
[47]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,
DeepSeek-AI, “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” 2025
2025
-
[48]
Qwq-32b: Embracing the power of reinforcement learning,
Qwen Team, “Qwq-32b: Embracing the power of reinforcement learning,” March 2025. 11
2025
-
[49]
Vlm-r1: A stable and generalizable r1-style large vision-language model,
Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, et al., “Vlm-r1: A stable and generalizable r1-style large vision-language model,” arXiv preprint arXiv:2504.07615, 2025
2025 arXiv
-
[50]
Deepseekmath: Pushing the limits of mathematical reasoning in open language models,
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al., “Deepseekmath: Pushing the limits of mathematical reasoning in open language models,” arXiv preprint arXiv:2402.03300, 2024
2024 arXiv
-
[51]
Proximal policy optimization algorithms,
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[52]
Blind image quality assessment: From natural scene statistics to perceptual quality,
Anush Krishna Moorthy and Alan Conrad Bovik, “Blind image quality assessment: From natural scene statistics to perceptual quality,” IEEE Transactions on Image Processing, vol. 20, no. 12, pp. 3350–3364, 2011
2011
-
[53]
Cnn-based cross-dataset no-reference image quality assessment,
Dan Yang, Veli-Tapani Peltoketo, and Joni-Kristian Kämäräinen, “Cnn-based cross-dataset no-reference image quality assessment,” in 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), 2019, pp. 3913–3921
2019
-
[54]
Semantically-aware game image quality assessment,
Kai Zhu, Vignesh Edithal, Le Zhang, Ilia Blank, and Imran Junejo, “Semantically-aware game image quality assessment,” arXiv preprint arXiv:2505.11724, 2025
2025 arXiv
-
[55]
Pytorch image quality: Metrics for image quality assessment,
Sergey Kastryulin, Jamil Zakirov, Denis Prokopenko, and Dmitry V Dylov, “Pytorch image quality: Metrics for image quality assessment,” arXiv preprint arXiv:2208.14818, 2022
2022 arXiv
-
[56]
Do image and video quality metrics model low-level human vision?,
Dounia Hammou, Yancheng Cai, Pavan Madhusudanarao, Christos G Bampis, and Rafał K Mantiuk, “Do image and video quality metrics model low-level human vision?,” arXiv preprint arXiv:2503.16264, 2025
2025
-
[57]
A survey on image quality assessment: Insights, analysis, and future outlook,
Chengqian Ma, Zhengyi Shi, Zhiqiang Lu, Shenghao Xie, Fei Chao, and Yao Sui, “A survey on image quality assessment: Insights, analysis, and future outlook,” arXiv preprint arXiv:2502.08540, 2025
2025 arXiv
-
[58]
Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,
V . Hosu, H. Lin, T. Sziranyi, and D. Saupe, “Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,” IEEE Transactions on Image Processing, vol. 29, pp. 4041–4056, 2020
2020
-
[59]
Quality assessment of in-the-wild videos,
Dingquan Li, Tingting Jiang, and Ming Jiang, “Quality assessment of in-the-wild videos,” in Proceedings of the 27th ACM international conference on multimedia, 2019, pp. 2351–2359
2019
-
[60]
Perceptual quality prediction on authentically distorted images using a bag of features approach,
Deepti Ghadiyaram and Alan C Bovik, “Perceptual quality prediction on authentically distorted images using a bag of features approach,” Journal of vision, vol. 17, no. 1, pp. 32–32, 2017
2017
-
[61]
Exploring semantic feature discrimination for perceptual image super-resolution and opinion-unaware no-reference image quality assessment,
Guanglu Dong, Xiangyu Liao, Mingyang Li, Guihuan Guo, and Chao Ren, “Exploring semantic feature discrimination for perceptual image super-resolution and opinion-unaware no-reference image quality assessment,” arXiv preprint arXiv:2503.19295, 2025
2025 arXiv
-
[62]
No-reference image quality assessment combining swin-transformer and natural scene statistics,
Yuxuan Yang, Zhichun Lei, and Changlu Li, “No-reference image quality assessment combining swin-transformer and natural scene statistics,” Sensors, vol. 24, no. 16, pp. 5221, 2024
2024
-
[63]
Ie-iqa: Intelligibility enriched generalizable no-reference image quality assessment,
Tianshu Song, Leida Li, Hancheng Zhu, and Jiansheng Qian, “Ie-iqa: Intelligibility enriched generalizable no-reference image quality assessment,” Frontiers in Neuroscience, vol. 15, pp. 739138, 2021
2021
-
[64]
Massive online crowdsourced study of subjective and objective picture quality,
Deepti Ghadiyaram and Alan C Bovik, “Massive online crowdsourced study of subjective and objective picture quality,” IEEE Transactions on Image Processing, vol. 25, no. 1, pp. 372–387, 2015
2015
-
[65]
Biq2021: a large-scale blind image quality assessment database,
Nisar Ahmed and Shahzad Asif, “Biq2021: a large-scale blind image quality assessment database,” Journal of Electronic Imaging, vol. 31, no. 5, pp. 053010–053010, 2022
2022
-
[66]
No reference opinion unaware quality assessment of authentically distorted images,
Nithin C Babu, Vignesh Kannan, and Rajiv Soundararajan, “No reference opinion unaware quality assessment of authentically distorted images,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2459–2468
2023
-
[67]
Patch-vq:’patching up’the video quality problem,
Zhenqiang Ying, Maniratnam Mandal, Deepti Ghadiyaram, and Alan Bovik, “Patch-vq:’patching up’the video quality problem,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 14019– 14029
2021
-
[68]
The konstanz natural video database (konvid-1k),
Vlad Hosu, Franz Hahn, Mohsen Jenadeleh, Hanhe Lin, Hui Men, Tamás Szirányi, Shujun Li, and Dietmar Saupe, “The konstanz natural video database (konvid-1k),” in 2017 Ninth international conference on quality of multimedia experience (QoMEX). IEEE, 2017, pp. 1–6
2017
-
[69]
Agiqa-3k: An open database for ai-generated image quality assessment,
Chunyi Li, Zicheng Zhang, Haoning Wu, Wei Sun, Xiongkuo Min, Xiaohong Liu, Guangtao Zhai, and Weisi Lin, “Agiqa-3k: An open database for ai-generated image quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 8, pp. 6833–6846, 2023
2023
-
[70]
Kvq: Kwai video quality assessment for short-form videos,
Yiting Lu, Xin Li, Yajing Pei, Kun Yuan, Qizhi Xie, Yunpeng Qu, Ming Sun, Chao Zhou, and Zhibo Chen, “Kvq: Kwai video quality assessment for short-form videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 25963–25973
2024
-
[71]
A multi-annotated and multi-modal dataset for wide-angle video quality assessment,
Bo Hu, Wei Wang, Chunyi Li, Lihuo He, Leida Li, and Xinbo Gao, “A multi-annotated and multi-modal dataset for wide-angle video quality assessment,” arXiv preprint arXiv:2501.12082, 2025
2025 arXiv
-
[72]
Nits-iqa database: a new image quality assessment database,
Jayesh Ruikar and Saurabh Chaudhury, “Nits-iqa database: a new image quality assessment database,” Sensors, vol. 23, no. 4, pp. 2279, 2023. 12
2023
-
[73]
Exiqa: Explainable image quality assessment using distortion attributes,
Sepehr Kazemi Ranjbar and Emad Fatemizadeh, “Exiqa: Explainable image quality assessment using distortion attributes,” arXiv e-prints, pp. arXiv–2409, 2024
2024
-
[74]
Explainable and generalizable blind image quality assessment via semantic attribute reasoning,
Yipo Huang, Leida Li, Yuzhe Yang, Yaqian Li, and Yandong Guo, “Explainable and generalizable blind image quality assessment via semantic attribute reasoning,” IEEE Transactions on Multimedia, vol. 25, pp. 7672–7685, 2022
2022
-
[75]
Explainability for deep learning in mammography image quality assessment,
Narbota Amanova, Jörg Martin, and Clemens Elster, “Explainability for deep learning in mammography image quality assessment,” Machine Learning: Science and Technology, vol. 3, no. 2, pp. 025015, 2022
2022
-
[76]
Pipal: a large-scale image quality assessment dataset for perceptual image restoration,
Gu Jinjin, Cai Haoming, Chen Haoyu, Ye Xiaoxing, Jimmy S Ren, and Dong Chao, “Pipal: a large-scale image quality assessment dataset for perceptual image restoration,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI ...
2020
-
[77]
Maniqa: Multi-dimension attention network for no-reference image quality assessment,
Sidi Yang, Tianhe Wu, Shuwei Shi, Shanshan Lao, Yuan Gong, Mingdeng Cao, Jiahao Wang, and Yujiu Yang, “Maniqa: Multi-dimension attention network for no-reference image quality assessment,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 20...
2022
-
[78]
Large multi-modality model assisted ai-generated image quality assessment,
Puyi Wang, Wei Sun, Zicheng Zhang, Jun Jia, Yanwei Jiang, Zhichao Zhang, Xiongkuo Min, and Guangtao Zhai, “Large multi-modality model assisted ai-generated image quality assessment,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 7803–7812
2024
-
[79]
Pea265: Perceptual assessment of video compression artifacts,
Liqun Lin, Shiqi Yu, Liping Zhou, Weiling Chen, Tiesong Zhao, and Zhou Wang, “Pea265: Perceptual assessment of video compression artifacts,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, no. 11, pp. 3898–3910, 2020
2020
-
[80]
Saliency-aware spatio-temporal artifact detection for compressed video quality assessment,
Liqun Lin, Yang Zheng, Weiling Chen, Chengdong Lan, and Tiesong Zhao, “Saliency-aware spatio-temporal artifact detection for compressed video quality assessment,” IEEE Signal Processing Letters, vol. 30, pp. 693–697, 2023
2023
-
[81]
Clip-agiqa: Boosting the performance of ai-generated image quality assessment with clip,
Zhenchen Tang, Zichuan Wang, Bo Peng, and Jing Dong, “Clip-agiqa: Boosting the performance of ai-generated image quality assessment with clip,” in International Conference on Pattern Recognition. Springer, 2025, pp. 48–61
2025
-
[82]
Bringing textual prompt to ai-generated image quality assessment,
Bowen Qu, Haohui Li, and Wei Gao, “Bringing textual prompt to ai-generated image quality assessment,” in 2024 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 2024, pp. 1–6
2024
-
[83]
Metaiqa: Deep meta-learning for no- reference image quality assessment,
Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, and Guangming Shi, “Metaiqa: Deep meta-learning for no- reference image quality assessment,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 14143–14152
2020
-
[84]
Sgiqa: semantic-guided no-reference image quality assessment,
Linpeng Pan, Xiaozhe Zhang, Fengying Xie, Haopeng Zhang, and Yushan Zheng, “Sgiqa: semantic-guided no-reference image quality assessment,” IEEE Transactions on Broadcasting, 2024
2024
-
[85]
A domain adaptive deep learning solution for scanpath prediction of paintings,
Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, and Alessandro Bruno, “A domain adaptive deep learning solution for scanpath prediction of paintings,” in Proceedings of the 19th International Conference on Content-Based Multimedia Indexing, 2022, pp. 57–63
2022
-
[86]
Self supervised scanpath prediction framework for painting images,
Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, and Alessandro Bruno, “Self supervised scanpath prediction framework for painting images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1539–1548
2022
-
[87]
On the use of a scanpath predictor and convolutional neural network for blind image quality assessment,
Aladine Chetouani and Leida Li, “On the use of a scanpath predictor and convolutional neural network for blind image quality assessment,” Signal Processing: Image Communication, vol. 89, pp. 115963, 2020
2020
-
[88]
Emotion detection through facial expressions: A survey of ai-based methods,
Simranjit Singh, Amrik Singh, and Baljinder Kaur, “Emotion detection through facial expressions: A survey of ai-based methods,” IJSAT-International Journal on Science and Technology, vol. 16, no. 1
-
[89]
Human emotion detection and face recognition system,
Bikash Kumar Jha, Bharat Paudel, Adarsh Mishra, Aabik Maharjan, and Pralhad Chapagain, “Human emotion detection and face recognition system,” International Journal on Engineering Technology, vol. 2, no. 2, pp. 90–97, 2025
2025
-
[90]
A subjective study of image quality assessment metrics using crowdsourcing,
Jun Xu, Weisi Lin, and Leida Zhang, “A subjective study of image quality assessment metrics using crowdsourcing,” IEEE Transactions on Image Processing, vol. 30, pp. 1788–1800, 2021
2021
-
[94]
Scaling synthetic data creation with 1,000,000,000 personas,
Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu, “Scaling synthetic data creation with 1,000,000,000 personas,” 2025. 13
2025
-
[97]
Color image database tid2013: Peculiarities and preliminary results,
Nikolay Ponomarenko, Oleg Ieremeiev, Vladimir Lukin, Karen Egiazarian, Lina Jin, Jaakko Astola, Benoit V ozel, Kacem Chehdi, Marco Carli, Federica Battisti, and C.-C. Jay Kuo, “Color image database tid2013: Peculiarities and preliminary results,” in European Workshop on Visual...
2013
-
[98]
Image-guided outdoor lidar perception quality assessment for autonomous driving,
Ce Zhang and Azim Eskandarian, “Image-guided outdoor lidar perception quality assessment for autonomous driving,” IEEE Transactions on Intelligent Vehicles, pp. 1–12, 2024
2024
-
[99]
Explainable artificial intelligence: Importance, use domains, stages, output shapes, and challenges,
Naeem Ullah, Javed Ali Khan, Ivanoe De Falco, and Giovanna Sannino, “Explainable artificial intelligence: Importance, use domains, stages, output shapes, and challenges,” ACM Comput. Surv., vol. 57, no. 4, Dec. 2024
2024
-
[100]
Can surgeons trust ai? perspectives on machine learning in surgery and the importance of explainable artificial intelligence (xai),
J. M. Brandenburg, B. P. Müller-Stich, M. Wagner, et al., “Can surgeons trust ai? perspectives on machine learning in surgery and the importance of explainable artificial intelligence (xai),” Langenbeck’s Archives of Surgery, vol. 410, pp. 53, 2025
2025
-
[101]
Transparency and explainability of ai systems: Ethical guidelines in practice,
N. Balasubramaniam, M. Kauppinen, K. Hiekkanen, and S. Kujala, “Transparency and explainability of ai systems: Ethical guidelines in practice,” in Requirements Engineering: Foundation for Software Quality, Vincenzo Gervasi and Andreas V ogelsang, Eds., vol. 13216 ofLecture Not...
2022
-
[102]
Automatic diagnosis of knee osteoarthritis severity using swin transformer,
Aymen Sekhri, Mohamed A Kerkouri, Aladine Chetouani, Marouane Tliba, Yassine Nasser, Rachid Jennane, and Alessandro Bruno, “Automatic diagnosis of knee osteoarthritis severity using swin transformer,” in Proceedings of the 20th International Conference on Content-Based Multime...
2023
-
[103]
Appealing, but misleading: a warning against a naive ai realism,
P. Engel-Hermann and A. Skulmowski, “Appealing, but misleading: a warning against a naive ai realism,” AI Ethics, 2024
2024
-
[104]
Mixture of experts: a literature survey,
Saeed Masoudnia and Reza Ebrahimpour, “Mixture of experts: a literature survey,” Artificial Intelligence Review, vol. 42, pp. 275–293, 2014
2014
-
[105]
Lora: Low- rank adaptation of large language models,
Edward J. Hu, Yelong Shen, Phil Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shawn Wang, and Weizhu Chen, “Lora: Low- rank adaptation of large language models,” in Proceedings of the International Conference on Learning Representations (ICLR), 2022, vol. 1, p. 3. 14
2022
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.