Pith. sign in

REVIEW 4 cited by

UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09619 v1 pith:H3WQZRBD submitted 2024-04-15 cs.CV cs.AI

classification cs.CVcs.AI
keywords aestheticassessmentmllmsuniaa-benchuniaa-llavaimagemulti-modalperception
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As an alternative to expensive expert evaluation, Image Aesthetic Assessment (IAA) stands out as a crucial task in computer vision. However, traditional IAA methods are typically constrained to a single data source or task, restricting the universality and broader application. In this work, to better align with human aesthetics, we propose a Unified Multi-modal Image Aesthetic Assessment (UNIAA) framework, including a Multi-modal Large Language Model (MLLM) named UNIAA-LLaVA and a comprehensive benchmark named UNIAA-Bench. We choose MLLMs with both visual perception and language ability for IAA and establish a low-cost paradigm for transforming the existing datasets into unified and high-quality visual instruction tuning data, from which the UNIAA-LLaVA is trained. To further evaluate the IAA capability of MLLMs, we construct the UNIAA-Bench, which consists of three aesthetic levels: Perception, Description, and Assessment. Extensive experiments validate the effectiveness and rationality of UNIAA. UNIAA-LLaVA achieves competitive performance on all levels of UNIAA-Bench, compared with existing MLLMs. Specifically, our model performs better than GPT-4V in aesthetic perception and even approaches the junior-level human. We find MLLMs have great potential in IAA, yet there remains plenty of room for further improvement. The UNIAA-LLaVA and UNIAA-Bench will be released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ArtiMuse is an MLLM that jointly scores image aesthetics and writes expert-style 8-attribute critiques, trained on a new 10,000-image expert-annotated dataset with a token-based continuous scoring method.

  2. Scaling-up Perceptual Video Quality Assessment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new pipeline plus datasets (OmniVQA-Chat-400K, OmniVQA-MOS-20K, OmniVQA-FG-Benchmark) yield LMMs with state-of-the-art video quality understanding and rating.

  3. Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A 2B multimodal model trained with HCM-GRPO, a GRPO variant with partial-credit rewards and hard-case oversampling, outperforms larger models on the authors' private physical-plausibility test set.

  4. Aesthetic Image Captioning with Saliency Enhanced MLLMs

    cs.CV 2025-09 conditional novelty 5.0 of 10

    ASE-MLLM fuses EAT-derived aesthetic saliency into MLLMs via cross-attention, improving aesthetic captioning over fine-tuning alone, but the SOTA claim rests on comparing to non-fine-tuned baselines.

Pith tools