Pith. sign in

REVIEW 2 cited by

A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.10854 v3 pith:4K3UZX5H submitted 2024-03-16 cs.CV

A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment

classification cs.CV
keywords qualitymllmspromptingimagelanguagemodelsthreevisual
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality Assessment (IQA) remains largely unexplored. In this paper, we conduct a comprehensive and systematic study of prompting MLLMs for IQA. We first investigate nine prompting systems for MLLMs as the combinations of three standardized testing procedures in psychophysics (i.e., the single-stimulus, double-stimulus, and multiple-stimulus methods) and three popular prompting strategies in natural language processing (i.e., the standard, in-context, and chain-of-thought prompting). We then present a difficult sample selection procedure, taking into account sample diversity and uncertainty, to further challenge MLLMs equipped with the respective optimal prompting systems. We assess three open-source and one closed-source MLLMs on several visual attributes of image quality (e.g., structural and textural distortions, geometric transformations, and color differences) in both full-reference and no-reference scenarios. Experimental results show that only the closed-source GPT-4V provides a reasonable account for human perception of image quality, but is weak at discriminating fine-grained quality variations (e.g., color differences) and at comparing visual quality of multiple images, tasks humans can perform effortlessly.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA

    cs.CV 2025-09 conditional novelty 5.0

    With a learned 30-pixel border prompt added to input images, a frozen mPLUG-Owl2-7B reaches 0.932 SRCC on KADID-10k using about 156K trainable parameters.

  2. HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment

    cs.CV 2025-08 unverdicted novelty 5.0

    HiRQA is a self-supervised NR-IQA framework trained on synthetic distortions, using a higher-order ranking loss, embedding distance loss, and text-guided contrastive alignment, claimed to generalize to authentic distortions.