Pith. sign in

REVIEW 13 cited by

AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.08276 v1 pith:KXRSE3YZ submitted 2024-01-16 cs.CV cs.CL

AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception

classification cs.CV cs.CL
keywords perceptionmllmsaestheticaesbenchaestheticsbenchmarkimagedevelopment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

With collective endeavors, multimodal large language models (MLLMs) are undergoing a flourishing development. However, their performances on image aesthetics perception remain indeterminate, which is highly desired in real-world applications. An obvious obstacle lies in the absence of a specific benchmark to evaluate the effectiveness of MLLMs on aesthetic perception. This blind groping may impede the further development of more advanced MLLMs with aesthetic perception capacity. To address this dilemma, we propose AesBench, an expert benchmark aiming to comprehensively evaluate the aesthetic perception capacities of MLLMs through elaborate design across dual facets. (1) We construct an Expert-labeled Aesthetics Perception Database (EAPD), which features diversified image contents and high-quality annotations provided by professional aesthetic experts. (2) We propose a set of integrative criteria to measure the aesthetic perception abilities of MLLMs from four perspectives, including Perception (AesP), Empathy (AesE), Assessment (AesA) and Interpretation (AesI). Extensive experimental results underscore that the current MLLMs only possess rudimentary aesthetic perception ability, and there is still a significant gap between MLLMs and humans. We hope this work can inspire the community to engage in deeper explorations on the aesthetic potentials of MLLMs. Source data will be available at https://github.com/yipoh/AesBench.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

    cs.CL 2026-06 unverdicted novelty 7.0

    PlanBench-V is a new benchmark and dataset for evaluating VLMs on spatial planning map interpretation via a four-stage framework of Perception, Reasoning, Association, and Implementation.

  2. MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models

    cs.CV 2026-05 conditional novelty 7.0

    MultiEmo-Bench supplies 10,344 images with aggregated multi-label emotion votes from 20 annotators each to evaluate MLLMs on dominant emotion and full distribution prediction.

  3. Tunable Polariton Canalization in Natural van der Waals Oxide

    physics.optics 2026-04 unverdicted novelty 7.0

    Natural alpha-V2O5 exhibits continuously tunable polariton canalization with unidirectional Poynting vector propagation controlled by incident infrared frequency.

  4. Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging

    cs.CV 2026-07 conditional novelty 6.0

    PRAC mines preference-rich images and merges LoRA adapters from aesthetically similar users to achieve state-of-the-art personalized aesthetic rating prediction.

  5. COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

    cs.AI 2026-06 unverdicted novelty 6.0

    COMPASS is a unified multimodal framework using a shared expert token τ_c to ground composition-intent for both perception and generation, backed by the new Comp-11 dataset.

  6. What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

    cs.CL 2026-06 conditional novelty 6.0

    VLM accuracy can be predicted from a scalar capability score derived from LLM text benchmarks plus multimodal data volume via a fitted transfer-absorption scaling law.

  7. Redefining Quality Criteria and Distance-Aware Score Modeling for Image Editing Assessment

    cs.CV 2026-04 unverdicted novelty 6.0

    DS-IEQA jointly learns evaluation criteria via feedback-driven prompt optimization and continuous score modeling via token-decoupled distance regression, ranking 4th in the 2026 NTIRE X-AIGC Quality Assessment Track 2...

  8. The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor

    cs.HC 2026-01 conditional novelty 6.0

    LAION-Aesthetics Predictor reinforces Western and male biases by preferentially selecting images associated with women and realistic Western/Japanese art while excluding men, LGBTQ+ references, and other styles.

  9. Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

    cs.LG 2025-10 conditional novelty 6.0

    A dynamic replay and reweighting scheduler (RECAP) preserves general capabilities during RLVR while keeping reasoning performance at least as good as reasoning-only finetuning.

  10. Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

    cs.CL 2026-06 unverdicted novelty 5.0

    MLLMs generate verbose, comprehensive, and repetitive aesthetic critiques unlike selective human ones, and reference-based metrics fail to detect this because they capture model house style instead of image-specific content.

  11. Tunable Polariton Canalization in Natural van der Waals Oxide

    physics.optics 2026-04 unverdicted novelty 5.0

    Untreated alpha-V2O5 exhibits frequency-tunable in-plane polariton canalization with unidirectional Poynting-vector flow, mapped by infrared nano-imaging and a permittivity phase diagram.

  12. What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?

    cs.CV 2026-04 unverdicted novelty 5.0

    Vision-language models encode diverse aesthetic attributes that propagate through decoder layers and enable effective personalized image aesthetics assessment using lightweight linear models without fine-tuning.

  13. Stable diffusion models reveal a persisting human and AI gap in visual creativity

    cs.AI 2025-11 unverdicted novelty 5.0

    Human visual artists outperform non-artists and AI image generators on creativity ratings, with more human prompting improving AI output but not closing the gap, while human and AI raters disagree on what counts as creative.