Pith. sign in

REVIEW 10 cited by

A-Bench: Are LMMs Masters at Evaluating AI-generated Images?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.03070 v2 pith:J3XSTYHN submitted 2024-06-05 cs.CV cs.AI

classification cs.CVcs.AI
keywords aigislmmsa-benchmodelsai-generatedaigibenchmarkevaluating
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

How to accurately and efficiently assess AI-generated images (AIGIs) remains a critical challenge for generative models. Given the high costs and extensive time commitments required for user studies, many researchers have turned towards employing large multi-modal models (LMMs) as AIGI evaluators, the precision and validity of which are still questionable. Furthermore, traditional benchmarks often utilize mostly natural-captured content rather than AIGIs to test the abilities of LMMs, leading to a noticeable gap for AIGIs. Therefore, we introduce A-Bench in this paper, a benchmark designed to diagnose whether LMMs are masters at evaluating AIGIs. Specifically, A-Bench is organized under two key principles: 1) Emphasizing both high-level semantic understanding and low-level visual quality perception to address the intricate demands of AIGIs. 2) Various generative models are utilized for AIGI creation, and various LMMs are employed for evaluation, which ensures a comprehensive validation scope. Ultimately, 2,864 AIGIs from 16 text-to-image models are sampled, each paired with question-answers annotated by human experts, and tested across 18 leading LMMs. We hope that A-Bench will significantly enhance the evaluation process and promote the generation quality for AIGIs. The benchmark is available at https://github.com/Q-Future/A-Bench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new large dataset and the FSCD model improve automated quality scoring of AI-generated talking-head videos, beating 15 baselines in correlation with human ratings.

  2. Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Fine-tuning Qwen-2.5-VL on the new FakeXplained dataset of 8,772 AI-generated images with box-and-caption artifact annotations yields an explainable detector with 98.1% accuracy and 37.8% IoU.

  3. Affordance Benchmark for MLLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 2,000-question benchmark finds multimodal AI models recognize object affordances far worse than humans, with top model Gemini-2.0-Pro at 18.05% versus 85.34% human best.

  4. Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    An interaction-augmented scene graph pipeline with chain-of-thought graph construction and reward-based tuning improves VLM reasoning on several benchmarks.

  5. Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new benchmark and fine-tuned models show that multi-task training improves average perceptual-similarity accuracy on known tasks but does not generalize to out-of-distribution perceptual tasks.

  6. OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?

    cs.CV 2024-12 conditional novelty 6.0 of 10

    OBI-Bench evaluates 23 large multimodal models on five oracle bone inscription tasks and finds they lag on fine-grained perception but approach untrained-human level in deciphering.

  7. HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator

    cs.CV 2024-11 conditional novelty 6.0 of 10

    HEIE is an MLLM-based evaluator that predicts defect heatmaps, plausibility scores, and natural-language explanations for AI-generated images, along with a new explainability dataset.

  8. MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis

    cs.CL 2024-11 conditional novelty 6.0 of 10

    MEMO-Bench scores 12 text-to-image models and 16 multimodal LLMs on emotion generation and recognition, finding stronger performance on positive emotions and weak fine-grained intensity estimation.

  9. OBIFormer: A Fast Attentive Denoising Framework for Oracle Bone Inscriptions

    cs.CV 2025-04 conditional novelty 5.0 of 10

    OBIFormer reports state-of-the-art PSNR and SSIM on oracle bone inscription denoising benchmarks while using fewer parameters than prior transformer-based methods.

  10. Embodied Image Quality Assessment for Robotic Intelligence

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Images are labeled by downstream robot task reward, yielding a 12,500-image benchmark on which human-oriented quality metrics fail and robot quality preferences diverge sharply from human ones.

Pith tools