Pith. sign in

REVIEW 5 major objections 4 minor 105 references

Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality

T0 review · 5 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper argues that Mean Opinion Score can no longer be the sole supervisory signal for multimedia quality assessment, and that models must integrate context-awareness, reasoning, and multimodality.

desk verdict A clear, well-written position paper that usefully synthesizes MOS's known limitations and proposes a concrete annotation roadmap, but overstates the 'must' without testing the feasibility of its own proposal. read the letter →

arxiv 2505.19696 v1 pith:X255IZDA submitted 2025-05-26 cs.CV cs.MMeess.IV

classification cs.CVcs.MMeess.IV
keywords imagequalityassessmentmeanopinionscorecontext-awarenessreasoningmultimodalityvision-languagemodelsbenchmarkreformexplainability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a position statement, not a new algorithm: it argues that the Mean Opinion Score (MOS), the single scalar average that has anchored image, video, and audio quality assessment for decades, is no longer sufficient as the only supervisory signal for quality models. The authors claim that a scalar label hides semantic failures such as hallucinated text, erases viewing context and user intent, discards disagreement among raters, and gives no rationale for a judgment. They propose that next-generation quality assessment must combine three capabilities: context-awareness (conditioning on task, device, and environment), reasoning (producing evidence-grounded explanations), and multimodality (jointly using visual, textual, attention, and artifact-map cues). The paper's contribution, if accepted, is to reframe quality assessment from correlation-matching to structured, explainable, and context-sensitive prediction. The authors outline benchmark reforms with richer annotations, persona-conditioned labels, and new reasoning metrics, but do not yet demonstrate that such annotations are feasible or that models trained on them outperform MOS-supervised models.

What carries the argument

The central object is the Mean Opinion Score itself, treated as a reductive representation of human judgment, and the proposed replacement is a structured prediction paradigm built on three pillars: context-awareness (explicit metadata conditioning such as task type and device), reasoning (chain-of-thought justifications grounded in visual evidence), and multimodality (joint embedding of visual features, text, saliency cues, and artifact maps). The pillars are meant to do the work that MOS cannot: expose semantic failures, preserve rater disagreement, support personalization, enable human-in-the-loop verification in medical and safety-critical settings, and provide actionable feedback. The argument is carried by empirical citations showing MOS-trained models' brittleness plus a roadmap for benchmarks that annotate contexts, personas, rationales, semantic flags, and rating distributions.

What would settle it

Train a no-reference quality model on existing MOS labels and a second model on the proposed richer annotations, then score both on a held-out set of AI-generated images with known semantic errors. If the richer model does not detect meaningfully more semantic failures or align better with expert rationales, the central claim is refuted.

Watch

Extended reading notes

Core claim

The central claim is that MOS, while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. More precisely, the paper contends that a model's output should be a structured judgment, a contextualized score conditioned on task and user intent, a natural-language rationale, a visual attention map, and an artifact mask, rather than a single number. The paper grounds this in documented failure modes: MOS-trained models misclassify semantically broken AI-generated images as high quality, lose substantial Spearman correlation in cross-dataset tests, cannot distinguish acceptable from unacceptable degradations across use cases, and discard rater disagreement that is itself informative. It then argues that context-awareness, reasoning, and multimodality are not optional enhancements but necessary foundations for robust, human-aligned quality systems.

Load-bearing premise

The load-bearing premise is that richer annotations, such as context metadata, persona-conditioned scores, rationales, semantic flags, and rating distributions, can be collected at scale without introducing new biases and that models trained on them will actually beat MOS-supervised models on robustness and human alignment.

Editorial extensions

If this is right

  • Benchmarks would need to add task metadata, persona-conditioned scores, structured rationales, semantic-error flags, and full rating distributions instead of a single mean.
  • Evaluation would move beyond PLCC and SROCC to include semantic alignment, reasoning fidelity, contextual sensitivity, and uncertainty calibration.
  • Quality models would output a contextualized score plus a rationale, attention map, and artifact mask, enabling debugging and human-in-the-loop workflows.
  • In constrained settings a pure MOS fast path remains acceptable, while high-stakes use cases would route through the full reasoning path.
  • New datasets and shared challenges organized along these lines would let the field test semantic robustness and explainability directly.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Not tested in the paper: the proposal could be checked cheaply by re-annotating a small slice of an existing quality database with personas and rationales; if annotators cannot agree on rationales, the reasoning pillar becomes harder to sustain.
  • Editorial extension: if richer annotations do scale, quality assessment could converge with general vision-language reasoning benchmarks, where quality becomes one more capability of instruction-following models rather than a specialized regression head.
  • Editorial extension: preserving rating distributions could turn quality models into disagreement detectors, flagging divisive content for platform moderation or medical second opinions, a use case the paper mentions but does not develop.
  • The paper positions the shift as urgent, but a decisive empirical test, whether a reasoning, context-conditioned model beats a MOS-trained model on out-of-distribution semantic failures, is left as future work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. This position paper argues that Mean Opinion Score (MOS), the scalar average used to supervise quality assessment models, is no longer sufficient as the sole supervisory signal. The authors critique MOS along several axes—generalization and semantic blindness, the limitations of correlation-based evaluation, the collapse of inter-subject disagreement, and the absence of interpretability—and then propose a modeling paradigm built on three pillars: context-awareness, reasoning, and multimodality. They outline structured outputs (contextualized score, rationale, attention map, artifact mask), benchmark reforms (multi-condition and persona-conditioned annotations, rationales, distributional labels, scalable collection infrastructure), and new evaluation metrics. No new experiments or datasets are presented; the argument is supported by citations to existing studies.

Significance. The paper offers a useful synthesis of a growing literature criticizing MOS and a coherent roadmap for reform. Its strengths are the catalog of documented failure modes and the concrete, structured proposal for richer annotations and outputs. However, the positive proposal is normative: the 'must' claim depends on an untested empirical premise that the proposed labels can be collected at scale and that models trained on them will improve robustness and human alignment. If that premise is validated by future work, the roadmap could influence benchmark design; as it stands, the paper is better read as a position statement than as an established finding. The explicit enumeration of failure-mode categories and structured output modalities is valuable for framing community discussion.

major comments (5)
  1. [§4–§5 (esp. §4.4, §5.1, §5.4)] The central normative claim that quality assessment models 'must integrate' context, reasoning, and multimodality rests on an unverified premise: that the annotation protocols of §5.1 (multi-condition labels, persona-conditioned ratings, structured rationales, semantic-error flags, distributional labels) can be collected at scale without introducing new biases, and that models trained with the §4.4 structured objectives will actually improve robustness and human alignment. No pilot, proof-of-concept, or even a synthetic demonstration is provided. The statement in §5.4 that these reforms 'are not speculative enhancements' is therefore an assertion rather than a supported conclusion, and the 'must' in the Abstract and §1 is stronger than the evidence warrants.
  2. [§5.3] The 'Reasoning Metrics' paragraph names coherence, validity, groundness, factuality, and utility but does not define any of them for the quality-assessment setting, does not specify how they would be computed for QA rationales, and provides no evidence that they correlate with human quality judgments or with perceptual quality. Citing QAFactEval [95] and a survey [96] establishes that general reasoning metrics exist, but not that they transfer to the proposed structured outputs of §4.4.
  3. [§5.2] The scalable data-collection claims are supported by references from outside multimedia quality assessment: the 'up to 40%' label-efficiency improvement cites a general active-learning survey [91], gamification evidence comes from image-description tasks [92], and the LEAD inter-rater reliability figure is from a psychiatric diagnostic panel [93]. None of these establishes QA-specific feasibility, so the scalability premise of the roadmap is currently unsubstantiated.
  4. [§6] The rebuttal concedes that 'in constrained scenarios, pure MOS suffices.' This directly qualifies the universal 'must' formulated in the Abstract and §1. The paper should either restrict its central claim to high-stakes and context-sensitive applications, where the argument is strongest, or provide clear criteria for when the additional capabilities are necessary; otherwise the thesis oscillates between a strong and a weak reading.
  5. [§3.1–3.2] Several quantitative claims supporting the critique are imported from preprints or secondary reports without error bars, confidence intervals, or significance information: the 25% cross-dataset SROCC drop (§3.1), the 0.15 SROCC gain of a CLIP-guided NR model (§3.1), the 30% drop for Re-IQA (§3.2), and the PLCC≈0.92-to-SROCC<0.70 contrast (§3.2). For a position paper these numbers are illustrative, but presenting them as precise figures gives a false sense of measurement certainty; please present them as approximate ranges or clearly note the original conditions and variability.
minor comments (4)
  1. [§3.3, §3.5, §2] There are several typos and missing spaces: 'somerecent works' (§3.3), 'adress' (§3.5), and 'LLaV A' (reference [44] in §2).
  2. [Figure 1 caption] The caption reads 'a reduced pipeline from our paradigm'; 'reduced' appears to mean 'simplified,' which would be clearer.
  3. [§5.1] The phrase 'and explanations to subject before experience' is grammatically unclear; likely 'provide explanations to the subject before the experiment/rating task' is intended.
  4. [References] The ISO 9241 reference appears as a footnote URL rather than a formal reference entry, which is inconsistent with the reference style used elsewhere; consider adding it to the bibliography.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the paper is a literature-based position statement, and the paper's self-citations are background support rather than load-bearing premises.

full rationale

This is a position paper rather than a derivation: it contains no equations, no fitted parameters, no benchmark results of its own, and no prediction that reduces to an input. The central claim that modern quality assessment models must integrate context-awareness, reasoning, and multimodality is an argued normative position supported by a broad review of external literature; it is not derived from the paper's own definitions, from any fitted quantity, or from a self-citation chain. The self-citations that appear (refs. 11, 85, 86, 87, 102) are used only as supporting examples for background claims: medical imaging is high-stakes, saliency and scanpath cues are relevant modalities, and medical workflows use explainability. None of these supplies the load-bearing premise that MOS is insufficient or that the three-pillar framework is necessary. The roadmap in Sections 4 and 5 is explicitly a proposal with no proof-of-concept; that is an evidence gap or correctness risk, not circularity. Section 6 even concedes that 'in constrained scenarios, pure MOS suffices,' which qualifies the universal 'must' but does not make the argument circular. Thus the paper is self-contained in its argumentative structure, and no circular step can be identified.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new fitted parameters, mathematical axioms, or invented entities. Its free-parameter and entity ledgers are empty; the axioms listed capture the domain assumptions it imports from prior literature and the untested efficacy assumption of its own roadmap.

assumptions (3)
  • domain assumption MOS is the dominant supervisory signal in multimedia quality assessment and reduces human judgments to a single scalar.
    This is the foundational premise of the paper, supported by citations [13,14] but not independently verified here.
  • domain assumption The empirical findings cited, such as a 25% cross-dataset SROCC drop and a 0.15 SROCC improvement with CLIP guidance, are accurate and representative.
    The paper relies on external studies, including preprints [54], without replication or meta-analysis.
  • ad hoc to paper Models trained with context, reasoning, and multimodal supervision will be more robust, interpretable, and human-aligned than MOS-only models.
    This is the core efficacy claim of the proposed paradigm; the paper provides no experimental evidence for it.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality." pith.science (2026). https://pith.science/paper/X255IZDA

@misc{pith2026250519696,
  author       = {Pith},
  title        = {Pith review of: Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/X255IZDA}},
  note         = {Machine review of arXiv:2505.19696}
}
read the original abstract

This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensitive human judgments to a single scalar, obscuring semantic failures, user intent, and the rationale behind quality decisions. We contend that modern quality assessment models must integrate three interdependent capabilities: (1) context-awareness, to adapt evaluations to task-specific goals and viewing conditions; (2) reasoning, to produce interpretable, evidence-grounded justifications for quality judgments; and (3) multimodality, to align perceptual and semantic cues using vision-language models. We critique the limitations of current MOS-centric benchmarks and propose a roadmap for reform: richer datasets with contextual metadata and expert rationales, and new evaluation metrics that assess semantic alignment, reasoning fidelity, and contextual sensitivity. By reframing quality assessment as a contextual, explainable, and multimodal modeling task, we aim to catalyze a shift toward more robust, human-aligned, and trustworthy evaluation systems.

Figures

Figures reproduced from arXiv: 2505.19696 by the authors.

Figure 1
Figure 1. An illustrative example comparing the a reduced pipeline from our paradigm (b) with a traditional MOS [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

105 extracted references · 63 canonical work pages

  1. [95]

    QAFactEval: Improved QA-based factual consistency evaluation for summarization,

    Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong, “QAFactEval: Improved QA-based factual consistency evaluation for summarization,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir ...

  2. [96]

    Evaluating step-by-step reasoning traces: A survey,

    Jinu Lee and J. Hockenmaier, “Evaluating step-by-step reasoning traces: A survey,” ArXiv, vol. abs/2502.12289, 2025

  3. [91]

    A survey on deep active learning: Recent advances and new frontiers,

    Dongyuan Li, Zhen Wang, Yankai Chen, Renhe Jiang, Weiping Ding, and Manabu Okumura, “A survey on deep active learning: Recent advances and new frontiers,” IEEE Transactions on Neural Networks and Learning Systems, 2024

  4. [92]

    Crowdsourcing image descriptions using gamification: a comparison between game-generated labels and professional descriptors,

    Tomislav Ivanjko, “Crowdsourcing image descriptions using gamification: a comparison between game-generated labels and professional descriptors,” in 2019 42nd International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO). IEEE, 2019, pp. 537–541

  5. [93]

    The leading guideline: Reporting standards for expert panel, best-estimate diagnosis, and longitudinal expert all data (lead) methods,

    Veerle C. Eijsbroek, Katarina Kjell, H. Andrew Schwartz, Jan R. Boehnke, Eiko I. Fried, Daniel N. Klein, Peik Gustafsson, Isabelle Augenstein, Patrick M.M. Bossuyt, and Oscar N.E. Kjell, “The leading guideline: Reporting standards for expert panel, best-estimate diagnosis, and longitudinal expert all data (lead) methods,” Comprehensive Psychiatry, p. 152603, 2025

  6. [1]

    Audio-visual multimedia quality assessment: A comprehensive survey,

    Zahid Akhtar and Tiago H Falk, “Audio-visual multimedia quality assessment: A comprehensive survey,” IEEE access, vol. 5, pp. 21090–21117, 2017

  7. [2]

    Perceptual image quality assessment: a survey,

    Guangtao Zhai and Xiongkuo Min, “Perceptual image quality assessment: a survey,” Science China Information Sciences, vol. 63, pp. 1–52, 2020

  8. [3]

    Perceptual video quality assessment: A survey,

    Xiongkuo Min, Huiyu Duan, Wei Sun, Yucheng Zhu, and Guangtao Zhai, “Perceptual video quality assessment: A survey,” Science China Information Sciences, vol. 67, no. 11, pp. 211301, 2024

Show all 105 references
  1. [4]

    Image aesthetic assessment: An experimental survey,

    Yubin Deng, Chen Change Loy, and Xiaoou Tang, “Image aesthetic assessment: An experimental survey,”IEEE Signal Processing Magazine, vol. 34, no. 4, pp. 80–106, 2017

  2. [5]

    A brief survey on adaptive video streaming quality assessment,

    Wei Zhou, Xiongkuo Min, Hong Li, and Qiuping Jiang, “A brief survey on adaptive video streaming quality assessment,” Journal of Visual Communication and Image Representation, vol. 86, pp. 103526, 2022

  3. [6]

    Real-time quality-and energy-aware bitrate ladder construction for live video streaming,

    Mohammad Ghasempour, Hadi Amirpour, and Christian Timmerer, “Real-time quality-and energy-aware bitrate ladder construction for live video streaming,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems, 2025

  4. [7]

    Telepresence video quality assessment,

    Zhenqiang Ying, Deepti Ghadiyaram, and Alan Bovik, “Telepresence video quality assessment,” in European Conference on Computer Vision. Springer, 2022, pp. 327–347

  5. [8]

    A systematic literature review: Real-time 3d reconstruction method for telepresence system,

    Fazliaty Edora Fadzli, Ajune Wanis Ismail, and Shafina Abd Karim Ishigaki, “A systematic literature review: Real-time 3d reconstruction method for telepresence system,” Plos one, vol. 18, no. 11, pp. e0287155, 2023

  6. [9]

    Quality assessment of videos on social media platforms related to gestational diabetes mellitus in china: A cross-section study,

    Qin-Yu Cai, Jing Tang, Si-Zhe Meng, Yi Sun, Xia Lan, and Tai-Hang Liu, “Quality assessment of videos on social media platforms related to gestational diabetes mellitus in china: A cross-section study,” Heliyon, vol. 10, no. 7, pp. e29020, 2024

  7. [10]

    Robust dual-modal image quality assessment aware deep learning network for traffic targets detection of autonomous vehicles,

    Keke Geng, Ge Dong, and Wenhan Huang, “Robust dual-modal image quality assessment aware deep learning network for traffic targets detection of autonomous vehicles,” Multimedia Tools and Applications, vol. 81, no. 5, pp. 6801–6826, 2022

  8. [11]

    Deep-based quality assessment of medical images through domain adaptation,

    Marouane Tliba, Aymen Sekhri, Mohamed Amine Kerkouri, and Aladine Chetouani, “Deep-based quality assessment of medical images through domain adaptation,” in 2022 IEEE International Conference on Image Processing (ICIP), 2022, pp. 3692–3696

  9. [12]

    Representation learning optimization for 3d point cloud quality assessment without reference,

    Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, and Frederic Dufaux, “Representation learning optimization for 3d point cloud quality assessment without reference,” 10 2022, pp. 3702–3706

  10. [13]

    Recommendation p.910: Subjective video quality assessment methods for multimedia applications,

    International Telecommunication Union Telecommunication Standardization Sector (ITU-T), “Recommendation p.910: Subjective video quality assessment methods for multimedia applications,” https://www.itu.int/rec/T-REC-P. 910-202310-I/en, 2023, ITU-T Recommendation P.910 (10/23)

  11. [14]

    Mean opinion score (mos) revisited: methods and applications, limitations and alternatives,

    Robert C Streijl, Stefan Winkler, and David S Hands, “Mean opinion score (mos) revisited: methods and applications, limitations and alternatives,” Multimedia Systems, vol. 22, no. 2, pp. 213–227, 2016

  12. [15]

    William james, gustav fechner, and early psychophysics,

    Stephanie L Hawkins, “William james, gustav fechner, and early psychophysics,” Frontiers in physiology, vol. 2, pp. 68, 2011

  13. [16]

    A brief review of the history and application of psychometrics and scaling to image quality assessment,

    Norman Burningham, “A brief review of the history and application of psychometrics and scaling to image quality assessment,” in PICS, 1999, pp. 169–172

  14. [17]

    Just noticeable difference,

    Melissa K Stern and James H Johnson, “Just noticeable difference,” The corsini encyclopedia of psychology, pp. 1–2, 2010

  15. [18]

    From pairwise comparisons and rating to a unified quality scale,

    Maria Perez-Ortiz, Aliaksei Mikhailiuk, Emin Zerman, Vedad Hulusic, Giuseppe Valenzise, and Rafał K Mantiuk, “From pairwise comparisons and rating to a unified quality scale,” IEEE Transactions on Image Processing, vol. 29, pp. 1139–1151, 2019

  16. [19]

    A law of comparative judgment,

    Louis L Thurstone, “A law of comparative judgment,” in Scaling, pp. 81–92. Routledge, 2017

  17. [20]

    Recommendation bt.500: Methodology for the subjective assessment of the quality of television pictures,

    International Telecommunication Union Radiocommunication Sector (ITU-R), “Recommendation bt.500: Methodology for the subjective assessment of the quality of television pictures,” https://www.itu.int/rec/R-REC-BT.500 , 2000, ITU-R Recommendation BT.500-10

  18. [21]

    Subjective and objective quality assessment of image: A survey,

    Pedram Mohammadi, Abbas Ebrahimi-Moghadam, and Shahram Shirani, “Subjective and objective quality assessment of image: A survey,” arXiv preprint arXiv:1406.7799, 2014

  19. [22]

    Experimental comparison of psnr and ssim metrics for video quality estimation,

    Z. Kotevski and P. Mitrevski, “Experimental comparison of psnr and ssim metrics for video quality estimation,” in ICT Innovations 2009, Danco Davcev and Juan M. Gómez, Eds., pp. 205–213. Springer, Berlin, Heidelberg, 2010

  20. [23]

    Image qualityassessment: From errorvisibilitytostructural similarity,

    BovikAC WangZhou, HR Sheikh, et al., “Image qualityassessment: From errorvisibilitytostructural similarity,” IEEE Transon ImageProcessing, vol. 13, no. 4, pp. 600, 2004

  21. [24]

    Toward a practical perceptual video quality metric, 2016,

    Zhi Li, Anne Aaron, Ioannis Katsavounidis, Anush Moorthy, and Megha Manohara, “Toward a practical perceptual video quality metric, 2016,” Dostupno na: http://techblog. netflix. com/2016/06/toward-practical-perceptual-video. html [16.8. 2022.], 2016. 10

  22. [25]

    Reduced-reference image quality assessment based on perceptual image hashing,

    Xudong Lv and Z. Jane Wang, “Reduced-reference image quality assessment based on perceptual image hashing,” 2009 16th IEEE International Conference on Image Processing (ICIP), pp. 4361–4364, 2009

  23. [26]

    Reduced reference image quality assessment via sub-image similarity based redundancy measurement,

    Xuanqin Mou, Wufeng Xue, and Lei Zhang, “Reduced reference image quality assessment via sub-image similarity based redundancy measurement,” in Electronic imaging, 2012

  24. [27]

    No-reference image quality assessment in the spatial domain,

    Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik, “No-reference image quality assessment in the spatial domain,” IEEE Transactions on Image Processing, vol. 21, no. 12, pp. 4695–4708, 2012

  25. [28]

    Making a “completely blind

    Anish Mittal, Rajiv Soundararajan, and Alan C Bovik, “Making a “completely blind” image quality analyzer,” IEEE Signal processing letters, vol. 20, no. 3, pp. 209–212, 2012

  26. [29]

    Arniqa: Learning distortion manifold for image quality assessment,

    Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini, and Alberto Del Bimbo, “Arniqa: Learning distortion manifold for image quality assessment,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 189–198

  27. [30]

    Blind image quality assessment using a deep bilinear convolutional neural network,

    Weixia Zhang, Kede Ma, Jia Yan, Dexiang Deng, and Zhou Wang, “Blind image quality assessment using a deep bilinear convolutional neural network,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, pp. 36–47, 2019

  28. [31]

    Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,

    Vlad Hosu, Hanhe Lin, Tamas Sziranyi, and Dietmar Saupe, “Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,” IEEE Transactions on Image Processing, vol. 29, pp. 4041–4056, 2020

  29. [32]

    The unreasonable effectiveness of deep features as a perceptual metric,

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 586–595

  30. [33]

    Topiq: A top-down approach from semantics to distortions for image quality assessment,

    Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu, Liang Liao, Wenxiu Sun, Qiong Yan, and Weisi Lin, “Topiq: A top-down approach from semantics to distortions for image quality assessment,” IEEE Transactions on Image Processing, 2024

  31. [34]

    Musiq: Multi-scale image quality transformer,

    Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang, “Musiq: Multi-scale image quality transformer,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5148–5157

  32. [35]

    Perceptual image quality assessment with transform- ers,

    Manri Cheon, Sung-Jun Yoon, Byungyeon Kang, and Junwoo Lee, “Perceptual image quality assessment with transform- ers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2021, pp. 433–442

  33. [36]

    Re-iqa: Unsupervised learning for image quality assessment in the wild,

    Avinab Saha, Sandeep Mishra, and Alan C Bovik, “Re-iqa: Unsupervised learning for image quality assessment in the wild,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 5846–5855

  34. [37]

    Learning transferable visual models from natural language supervision,

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning....

  35. [38]

    Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in International conference on machine learning . PMLR, 2022, pp. 12888–12900

  36. [39]

    Benchmark evaluations, applications, and challenges of large vision language models: A survey,

    Zongxia Li, Xiyang Wu, Hongyang Du, Huy Nghiem, and Guangyao Shi, “Benchmark evaluations, applications, and challenges of large vision language models: A survey,” arXiv preprint arXiv:2501.02189, vol. 1, 2025

  37. [40]

    Exploring clip for assessing the look and feel of images,

    Jianyi Wang, Kelvin CK Chan, and Chen Change Loy, “Exploring clip for assessing the look and feel of images,” in AAAI, 2023

  38. [41]

    Quality-aware image-text alignment for real-world image quality assessment,

    Lorenzo Agnolucci, Leonardo Galteri, and Marco Bertini, “Quality-aware image-text alignment for real-world image quality assessment,” arXiv preprint arXiv:2403.11176, vol. 5, no. 6, 2024

  39. [42]

    Q-align: Teaching lmms for visual scoring via discrete text-defined levels,

    Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, et al., “Q-align: Teaching lmms for visual scoring via discrete text-defined levels,” arXiv preprint arXiv:2312.17090, 2023

  40. [43]

    Flamingo: a visual language model for few-shot learning,

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al., “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems, vol. ...

  41. [44]

    Visual instruction tuning,

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee, “Visual instruction tuning,” Advances in neural information processing systems, vol. 36, pp. 34892–34916, 2023

  42. [45]

    Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond,

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou, “Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond,” arXiv preprint arXiv:2308.12966, 2023

  43. [46]

    A comprehensive study of multimodal large language models for image quality assessment,

    Tianhe Wu, Kede Ma, Jie Liang, Yujiu Yang, and Lei Zhang, “A comprehensive study of multimodal large language models for image quality assessment,” in European Conference on Computer Vision. Springer, 2024, pp. 143–160

  44. [47]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

    DeepSeek-AI, “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” 2025

  45. [48]

    Qwq-32b: Embracing the power of reinforcement learning,

    Qwen Team, “Qwq-32b: Embracing the power of reinforcement learning,” March 2025. 11

  46. [49]

    Vlm-r1: A stable and generalizable r1-style large vision-language model,

    Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, et al., “Vlm-r1: A stable and generalizable r1-style large vision-language model,” arXiv preprint arXiv:2504.07615, 2025

  47. [50]

    Deepseekmath: Pushing the limits of mathematical reasoning in open language models,

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al., “Deepseekmath: Pushing the limits of mathematical reasoning in open language models,” arXiv preprint arXiv:2402.03300, 2024

  48. [51]

    Proximal policy optimization algorithms,

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017

  49. [52]

    Blind image quality assessment: From natural scene statistics to perceptual quality,

    Anush Krishna Moorthy and Alan Conrad Bovik, “Blind image quality assessment: From natural scene statistics to perceptual quality,” IEEE Transactions on Image Processing, vol. 20, no. 12, pp. 3350–3364, 2011

  50. [53]

    Cnn-based cross-dataset no-reference image quality assessment,

    Dan Yang, Veli-Tapani Peltoketo, and Joni-Kristian Kämäräinen, “Cnn-based cross-dataset no-reference image quality assessment,” in 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), 2019, pp. 3913–3921

  51. [54]

    Semantically-aware game image quality assessment,

    Kai Zhu, Vignesh Edithal, Le Zhang, Ilia Blank, and Imran Junejo, “Semantically-aware game image quality assessment,” arXiv preprint arXiv:2505.11724, 2025

  52. [55]

    Pytorch image quality: Metrics for image quality assessment,

    Sergey Kastryulin, Jamil Zakirov, Denis Prokopenko, and Dmitry V Dylov, “Pytorch image quality: Metrics for image quality assessment,” arXiv preprint arXiv:2208.14818, 2022

  53. [56]

    Do image and video quality metrics model low-level human vision?,

    Dounia Hammou, Yancheng Cai, Pavan Madhusudanarao, Christos G Bampis, and Rafał K Mantiuk, “Do image and video quality metrics model low-level human vision?,” arXiv preprint arXiv:2503.16264, 2025

  54. [57]

    A survey on image quality assessment: Insights, analysis, and future outlook,

    Chengqian Ma, Zhengyi Shi, Zhiqiang Lu, Shenghao Xie, Fei Chao, and Yao Sui, “A survey on image quality assessment: Insights, analysis, and future outlook,” arXiv preprint arXiv:2502.08540, 2025

  55. [58]

    Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,

    V . Hosu, H. Lin, T. Sziranyi, and D. Saupe, “Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,” IEEE Transactions on Image Processing, vol. 29, pp. 4041–4056, 2020

  56. [59]

    Quality assessment of in-the-wild videos,

    Dingquan Li, Tingting Jiang, and Ming Jiang, “Quality assessment of in-the-wild videos,” in Proceedings of the 27th ACM international conference on multimedia, 2019, pp. 2351–2359

  57. [60]

    Perceptual quality prediction on authentically distorted images using a bag of features approach,

    Deepti Ghadiyaram and Alan C Bovik, “Perceptual quality prediction on authentically distorted images using a bag of features approach,” Journal of vision, vol. 17, no. 1, pp. 32–32, 2017

  58. [61]

    Exploring semantic feature discrimination for perceptual image super-resolution and opinion-unaware no-reference image quality assessment,

    Guanglu Dong, Xiangyu Liao, Mingyang Li, Guihuan Guo, and Chao Ren, “Exploring semantic feature discrimination for perceptual image super-resolution and opinion-unaware no-reference image quality assessment,” arXiv preprint arXiv:2503.19295, 2025

  59. [62]

    No-reference image quality assessment combining swin-transformer and natural scene statistics,

    Yuxuan Yang, Zhichun Lei, and Changlu Li, “No-reference image quality assessment combining swin-transformer and natural scene statistics,” Sensors, vol. 24, no. 16, pp. 5221, 2024

  60. [63]

    Ie-iqa: Intelligibility enriched generalizable no-reference image quality assessment,

    Tianshu Song, Leida Li, Hancheng Zhu, and Jiansheng Qian, “Ie-iqa: Intelligibility enriched generalizable no-reference image quality assessment,” Frontiers in Neuroscience, vol. 15, pp. 739138, 2021

  61. [64]

    Massive online crowdsourced study of subjective and objective picture quality,

    Deepti Ghadiyaram and Alan C Bovik, “Massive online crowdsourced study of subjective and objective picture quality,” IEEE Transactions on Image Processing, vol. 25, no. 1, pp. 372–387, 2015

  62. [65]

    Biq2021: a large-scale blind image quality assessment database,

    Nisar Ahmed and Shahzad Asif, “Biq2021: a large-scale blind image quality assessment database,” Journal of Electronic Imaging, vol. 31, no. 5, pp. 053010–053010, 2022

  63. [66]

    No reference opinion unaware quality assessment of authentically distorted images,

    Nithin C Babu, Vignesh Kannan, and Rajiv Soundararajan, “No reference opinion unaware quality assessment of authentically distorted images,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2459–2468

  64. [67]

    Patch-vq:’patching up’the video quality problem,

    Zhenqiang Ying, Maniratnam Mandal, Deepti Ghadiyaram, and Alan Bovik, “Patch-vq:’patching up’the video quality problem,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 14019– 14029

  65. [68]

    The konstanz natural video database (konvid-1k),

    Vlad Hosu, Franz Hahn, Mohsen Jenadeleh, Hanhe Lin, Hui Men, Tamás Szirányi, Shujun Li, and Dietmar Saupe, “The konstanz natural video database (konvid-1k),” in 2017 Ninth international conference on quality of multimedia experience (QoMEX). IEEE, 2017, pp. 1–6

  66. [69]

    Agiqa-3k: An open database for ai-generated image quality assessment,

    Chunyi Li, Zicheng Zhang, Haoning Wu, Wei Sun, Xiongkuo Min, Xiaohong Liu, Guangtao Zhai, and Weisi Lin, “Agiqa-3k: An open database for ai-generated image quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 8, pp. 6833–6846, 2023

  67. [70]

    Kvq: Kwai video quality assessment for short-form videos,

    Yiting Lu, Xin Li, Yajing Pei, Kun Yuan, Qizhi Xie, Yunpeng Qu, Ming Sun, Chao Zhou, and Zhibo Chen, “Kvq: Kwai video quality assessment for short-form videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 25963–25973

  68. [71]

    A multi-annotated and multi-modal dataset for wide-angle video quality assessment,

    Bo Hu, Wei Wang, Chunyi Li, Lihuo He, Leida Li, and Xinbo Gao, “A multi-annotated and multi-modal dataset for wide-angle video quality assessment,” arXiv preprint arXiv:2501.12082, 2025

  69. [72]

    Nits-iqa database: a new image quality assessment database,

    Jayesh Ruikar and Saurabh Chaudhury, “Nits-iqa database: a new image quality assessment database,” Sensors, vol. 23, no. 4, pp. 2279, 2023. 12

  70. [73]

    Exiqa: Explainable image quality assessment using distortion attributes,

    Sepehr Kazemi Ranjbar and Emad Fatemizadeh, “Exiqa: Explainable image quality assessment using distortion attributes,” arXiv e-prints, pp. arXiv–2409, 2024

  71. [74]

    Explainable and generalizable blind image quality assessment via semantic attribute reasoning,

    Yipo Huang, Leida Li, Yuzhe Yang, Yaqian Li, and Yandong Guo, “Explainable and generalizable blind image quality assessment via semantic attribute reasoning,” IEEE Transactions on Multimedia, vol. 25, pp. 7672–7685, 2022

  72. [75]

    Explainability for deep learning in mammography image quality assessment,

    Narbota Amanova, Jörg Martin, and Clemens Elster, “Explainability for deep learning in mammography image quality assessment,” Machine Learning: Science and Technology, vol. 3, no. 2, pp. 025015, 2022

  73. [76]

    Pipal: a large-scale image quality assessment dataset for perceptual image restoration,

    Gu Jinjin, Cai Haoming, Chen Haoyu, Ye Xiaoxing, Jimmy S Ren, and Dong Chao, “Pipal: a large-scale image quality assessment dataset for perceptual image restoration,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI ...

  74. [77]

    Maniqa: Multi-dimension attention network for no-reference image quality assessment,

    Sidi Yang, Tianhe Wu, Shuwei Shi, Shanshan Lao, Yuan Gong, Mingdeng Cao, Jiahao Wang, and Yujiu Yang, “Maniqa: Multi-dimension attention network for no-reference image quality assessment,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 20...

  75. [78]

    Large multi-modality model assisted ai-generated image quality assessment,

    Puyi Wang, Wei Sun, Zicheng Zhang, Jun Jia, Yanwei Jiang, Zhichao Zhang, Xiongkuo Min, and Guangtao Zhai, “Large multi-modality model assisted ai-generated image quality assessment,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 7803–7812

  76. [79]

    Pea265: Perceptual assessment of video compression artifacts,

    Liqun Lin, Shiqi Yu, Liping Zhou, Weiling Chen, Tiesong Zhao, and Zhou Wang, “Pea265: Perceptual assessment of video compression artifacts,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, no. 11, pp. 3898–3910, 2020

  77. [80]

    Saliency-aware spatio-temporal artifact detection for compressed video quality assessment,

    Liqun Lin, Yang Zheng, Weiling Chen, Chengdong Lan, and Tiesong Zhao, “Saliency-aware spatio-temporal artifact detection for compressed video quality assessment,” IEEE Signal Processing Letters, vol. 30, pp. 693–697, 2023

  78. [81]

    Clip-agiqa: Boosting the performance of ai-generated image quality assessment with clip,

    Zhenchen Tang, Zichuan Wang, Bo Peng, and Jing Dong, “Clip-agiqa: Boosting the performance of ai-generated image quality assessment with clip,” in International Conference on Pattern Recognition. Springer, 2025, pp. 48–61

  79. [82]

    Bringing textual prompt to ai-generated image quality assessment,

    Bowen Qu, Haohui Li, and Wei Gao, “Bringing textual prompt to ai-generated image quality assessment,” in 2024 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 2024, pp. 1–6

  80. [83]

    Metaiqa: Deep meta-learning for no- reference image quality assessment,

    Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, and Guangming Shi, “Metaiqa: Deep meta-learning for no- reference image quality assessment,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 14143–14152

  81. [84]

    Sgiqa: semantic-guided no-reference image quality assessment,

    Linpeng Pan, Xiaozhe Zhang, Fengying Xie, Haopeng Zhang, and Yushan Zheng, “Sgiqa: semantic-guided no-reference image quality assessment,” IEEE Transactions on Broadcasting, 2024

  82. [85]

    A domain adaptive deep learning solution for scanpath prediction of paintings,

    Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, and Alessandro Bruno, “A domain adaptive deep learning solution for scanpath prediction of paintings,” in Proceedings of the 19th International Conference on Content-Based Multimedia Indexing, 2022, pp. 57–63

  83. [86]

    Self supervised scanpath prediction framework for painting images,

    Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, and Alessandro Bruno, “Self supervised scanpath prediction framework for painting images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1539–1548

  84. [87]

    On the use of a scanpath predictor and convolutional neural network for blind image quality assessment,

    Aladine Chetouani and Leida Li, “On the use of a scanpath predictor and convolutional neural network for blind image quality assessment,” Signal Processing: Image Communication, vol. 89, pp. 115963, 2020

  85. [88]

    Emotion detection through facial expressions: A survey of ai-based methods,

    Simranjit Singh, Amrik Singh, and Baljinder Kaur, “Emotion detection through facial expressions: A survey of ai-based methods,” IJSAT-International Journal on Science and Technology, vol. 16, no. 1

  86. [89]

    Human emotion detection and face recognition system,

    Bikash Kumar Jha, Bharat Paudel, Adarsh Mishra, Aabik Maharjan, and Pralhad Chapagain, “Human emotion detection and face recognition system,” International Journal on Engineering Technology, vol. 2, no. 2, pp. 90–97, 2025

  87. [90]

    A subjective study of image quality assessment metrics using crowdsourcing,

    Jun Xu, Weisi Lin, and Leida Zhang, “A subjective study of image quality assessment metrics using crowdsourcing,” IEEE Transactions on Image Processing, vol. 30, pp. 1788–1800, 2021

  88. [94]

    Scaling synthetic data creation with 1,000,000,000 personas,

    Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu, “Scaling synthetic data creation with 1,000,000,000 personas,” 2025. 13

  89. [97]

    Color image database tid2013: Peculiarities and preliminary results,

    Nikolay Ponomarenko, Oleg Ieremeiev, Vladimir Lukin, Karen Egiazarian, Lina Jin, Jaakko Astola, Benoit V ozel, Kacem Chehdi, Marco Carli, Federica Battisti, and C.-C. Jay Kuo, “Color image database tid2013: Peculiarities and preliminary results,” in European Workshop on Visual...

  90. [98]

    Image-guided outdoor lidar perception quality assessment for autonomous driving,

    Ce Zhang and Azim Eskandarian, “Image-guided outdoor lidar perception quality assessment for autonomous driving,” IEEE Transactions on Intelligent Vehicles, pp. 1–12, 2024

  91. [99]

    Explainable artificial intelligence: Importance, use domains, stages, output shapes, and challenges,

    Naeem Ullah, Javed Ali Khan, Ivanoe De Falco, and Giovanna Sannino, “Explainable artificial intelligence: Importance, use domains, stages, output shapes, and challenges,” ACM Comput. Surv., vol. 57, no. 4, Dec. 2024

  92. [100]

    Can surgeons trust ai? perspectives on machine learning in surgery and the importance of explainable artificial intelligence (xai),

    J. M. Brandenburg, B. P. Müller-Stich, M. Wagner, et al., “Can surgeons trust ai? perspectives on machine learning in surgery and the importance of explainable artificial intelligence (xai),” Langenbeck’s Archives of Surgery, vol. 410, pp. 53, 2025

  93. [101]

    Transparency and explainability of ai systems: Ethical guidelines in practice,

    N. Balasubramaniam, M. Kauppinen, K. Hiekkanen, and S. Kujala, “Transparency and explainability of ai systems: Ethical guidelines in practice,” in Requirements Engineering: Foundation for Software Quality, Vincenzo Gervasi and Andreas V ogelsang, Eds., vol. 13216 ofLecture Not...

  94. [102]

    Automatic diagnosis of knee osteoarthritis severity using swin transformer,

    Aymen Sekhri, Mohamed A Kerkouri, Aladine Chetouani, Marouane Tliba, Yassine Nasser, Rachid Jennane, and Alessandro Bruno, “Automatic diagnosis of knee osteoarthritis severity using swin transformer,” in Proceedings of the 20th International Conference on Content-Based Multime...

  95. [103]

    Appealing, but misleading: a warning against a naive ai realism,

    P. Engel-Hermann and A. Skulmowski, “Appealing, but misleading: a warning against a naive ai realism,” AI Ethics, 2024

  96. [104]

    Mixture of experts: a literature survey,

    Saeed Masoudnia and Reza Ebrahimpour, “Mixture of experts: a literature survey,” Artificial Intelligence Review, vol. 42, pp. 275–293, 2014

  97. [105]

    Lora: Low- rank adaptation of large language models,

    Edward J. Hu, Yelong Shen, Phil Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shawn Wang, and Weizhu Chen, “Lora: Low- rank adaptation of large language models,” in Proceedings of the International Conference on Learning Representations (ICLR), 2022, vol. 1, p. 3. 14

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.