Pith. sign in

REVIEW 2 cited by

Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06141 v1 pith:O2P2WCOS submitted 2025-03-08 cs.CV

Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model

classification cs.CV
keywords imagequalityassessmentimagesaestheticattributeslevelmllms
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The rapid expansion of mobile internet has resulted in a substantial increase in user-generated content (UGC) images, thereby making the thorough assessment of UGC images both urgent and essential. Recently, multimodal large language models (MLLMs) have shown great potential in image quality assessment (IQA) and image aesthetic assessment (IAA). Despite this progress, effectively scoring the quality and aesthetics of UGC images still faces two main challenges: 1) A single score is inadequate to capture the hierarchical human perception. 2) How to use MLLMs to output numerical scores, such as mean opinion scores (MOS), remains an open question. To address these challenges, we introduce a novel dataset, named Realistic image Quality and Aesthetic (RealQA), including 14,715 UGC images, each of which is annoted with 10 fine-grained attributes. These attributes span three levels: low level (e.g., image clarity), middle level (e.g., subject integrity) and high level (e.g., composition). Besides, we conduct a series of in-depth and comprehensive investigations into how to effectively predict numerical scores using MLLMs. Surprisingly, by predicting just two extra significant digits, the next token paradigm can achieve SOTA performance. Furthermore, with the help of chain of thought (CoT) combined with the learnt fine-grained attributes, the proposed method can outperform SOTA methods on five public datasets for IQA and IAA with superior interpretability and show strong zero-shot generalization for video quality assessment (VQA). The code and dataset will be released.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

    cs.CV 2026-07 conditional novelty 6.0

    Peak-End-Net uses image-aesthetic priors and peak-end-rule frame weighting to set a new state of the art on VADB and zero-shot DIVIDE-3K video aesthetic assessment.

  2. PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models

    cs.CV 2026-02 conditional novelty 6.0

    PhotoAgent uses a vision-language model, Monte-Carlo tree search, and a learned UGC aesthetic reward to autonomously choose and sequence photo edits.