Pith. sign in

REVIEW 6 cited by

UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09619 v1 pith:H3WQZRBD submitted 2024-04-15 cs.CV cs.AI

classification cs.CVcs.AI
keywords aestheticassessmentmllmsuniaa-benchuniaa-llavaimagemulti-modalperception
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As an alternative to expensive expert evaluation, Image Aesthetic Assessment (IAA) stands out as a crucial task in computer vision. However, traditional IAA methods are typically constrained to a single data source or task, restricting the universality and broader application. In this work, to better align with human aesthetics, we propose a Unified Multi-modal Image Aesthetic Assessment (UNIAA) framework, including a Multi-modal Large Language Model (MLLM) named UNIAA-LLaVA and a comprehensive benchmark named UNIAA-Bench. We choose MLLMs with both visual perception and language ability for IAA and establish a low-cost paradigm for transforming the existing datasets into unified and high-quality visual instruction tuning data, from which the UNIAA-LLaVA is trained. To further evaluate the IAA capability of MLLMs, we construct the UNIAA-Bench, which consists of three aesthetic levels: Perception, Description, and Assessment. Extensive experiments validate the effectiveness and rationality of UNIAA. UNIAA-LLaVA achieves competitive performance on all levels of UNIAA-Bench, compared with existing MLLMs. Specifically, our model performs better than GPT-4V in aesthetic perception and even approaches the junior-level human. We find MLLMs have great potential in IAA, yet there remains plenty of room for further improvement. The UNIAA-LLaVA and UNIAA-Bench will be released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation

    cs.CV 2025-09 conditional novelty 7.0 of 10

    A new 340K-image human-annotated dataset, a trained vision-language assessor, and an automated benchmark reveal that even state-of-the-art T2I models produce artifacts in roughly a third of output images.

  2. ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ArtiMuse is an MLLM that jointly scores image aesthetics and writes expert-style 8-attribute critiques, trained on a new 10,000-image expert-annotated dataset with a token-based continuous scoring method.

  3. Scaling-up Perceptual Video Quality Assessment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new pipeline plus datasets (OmniVQA-Chat-400K, OmniVQA-MOS-20K, OmniVQA-FG-Benchmark) yield LMMs with state-of-the-art video quality understanding and rating.

  4. Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new MLLM with multi-scale text-guided self-supervised learning reports state-of-the-art results across aesthetic scoring, commenting, and personalized image aesthetic assessment benchmarks.

  5. Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A 2B multimodal model trained with HCM-GRPO, a GRPO variant with partial-credit rewards and hard-case oversampling, outperforms larger models on the authors' private physical-plausibility test set.

  6. Aesthetic Image Captioning with Saliency Enhanced MLLMs

    cs.CV 2025-09 conditional novelty 5.0 of 10

    ASE-MLLM fuses EAT-derived aesthetic saliency into MLLMs via cross-attention, improving aesthetic captioning over fine-tuning alone, but the SOTA claim rests on comparing to non-fine-tuned baselines.

Pith tools