REVIEW 10 cited by
A-Bench: Are LMMs Masters at Evaluating AI-generated Images?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
How to accurately and efficiently assess AI-generated images (AIGIs) remains a critical challenge for generative models. Given the high costs and extensive time commitments required for user studies, many researchers have turned towards employing large multi-modal models (LMMs) as AIGI evaluators, the precision and validity of which are still questionable. Furthermore, traditional benchmarks often utilize mostly natural-captured content rather than AIGIs to test the abilities of LMMs, leading to a noticeable gap for AIGIs. Therefore, we introduce A-Bench in this paper, a benchmark designed to diagnose whether LMMs are masters at evaluating AIGIs. Specifically, A-Bench is organized under two key principles: 1) Emphasizing both high-level semantic understanding and low-level visual quality perception to address the intricate demands of AIGIs. 2) Various generative models are utilized for AIGI creation, and various LMMs are employed for evaluation, which ensures a comprehensive validation scope. Ultimately, 2,864 AIGIs from 16 text-to-image models are sampled, each paired with question-answers annotated by human experts, and tested across 18 leading LMMs. We hope that A-Bench will significantly enhance the evaluation process and promote the generation quality for AIGIs. The benchmark is available at https://github.com/Q-Future/A-Bench.
Forward citations
Cited by 10 Pith papers
-
Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
A new large dataset and the FSCD model improve automated quality scoring of AI-generated talking-head videos, beating 15 baselines in correlation with human ratings.
-
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
Fine-tuning Qwen-2.5-VL on the new FakeXplained dataset of 8,772 AI-generated images with box-and-caption artifact annotations yields an explainable detector with 98.1% accuracy and 37.8% IoU.
-
Affordance Benchmark for MLLMs
A new 2,000-question benchmark finds multimodal AI models recognize object affordances far worse than humans, with top model Gemini-2.0-Pro at 18.05% versus 85.34% human best.
-
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
An interaction-augmented scene graph pipeline with chain-of-thought graph construction and reward-based tuning improves VLM reasoning on several benchmarks.
-
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
A new benchmark and fine-tuned models show that multi-task training improves average perceptual-similarity accuracy on known tasks but does not generalize to out-of-distribution perceptual tasks.
-
OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?
OBI-Bench evaluates 23 large multimodal models on five oracle bone inscription tasks and finds they lag on fine-grained perception but approach untrained-human level in deciphering.
-
HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
HEIE is an MLLM-based evaluator that predicts defect heatmaps, plausibility scores, and natural-language explanations for AI-generated images, along with a new explainability dataset.
-
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
MEMO-Bench scores 12 text-to-image models and 16 multimodal LLMs on emotion generation and recognition, finding stronger performance on positive emotions and weak fine-grained intensity estimation.
-
OBIFormer: A Fast Attentive Denoising Framework for Oracle Bone Inscriptions
OBIFormer reports state-of-the-art PSNR and SSIM on oracle bone inscription denoising benchmarks while using fewer parameters than prior transformer-based methods.
-
Embodied Image Quality Assessment for Robotic Intelligence
Images are labeled by downstream robot task reward, yielding a 12,500-image benchmark on which human-oriented quality metrics fail and robot quality preferences diverge sharply from human ones.
Discussion (0). Continue with ORCID to comment.