Pith. sign in

REVIEW 11 cited by

GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02782 v3 pith:L6JO7INT submitted 2025-04-03 cs.CV

classification cs.CV
keywords gpt-4oimagegenerationbenchmarkeditinggpt-imgevalacrossarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent breakthroughs in OpenAI's GPT4o model have demonstrated surprisingly good capabilities in image generation and editing, resulting in significant excitement in the community. This technical report presents the first-look evaluation benchmark (named GPT-ImgEval), quantitatively and qualitatively diagnosing GPT-4o's performance across three critical dimensions: (1) generation quality, (2) editing proficiency, and (3) world knowledge-informed semantic synthesis. Across all three tasks, GPT-4o demonstrates strong performance, significantly surpassing existing methods in both image generation control and output quality, while also showcasing exceptional knowledge reasoning capabilities. Furthermore, based on the GPT-4o's generated data, we propose a classification-model-based approach to investigate the underlying architecture of GPT-4o, where our empirical results suggest the model consists of an auto-regressive (AR) combined with a diffusion-based head for image decoding, rather than the VAR-like architectures. We also provide a complete speculation on GPT-4o's overall architecture. In addition, we conduct a series of analyses to identify and visualize GPT-4o's specific limitations and the synthetic artifacts commonly observed in its image generation. We also present a comparative study of multi-round image editing between GPT-4o and Gemini 2.0 Flash, and discuss the safety implications of GPT-4o's outputs, particularly their detectability by existing image forensic models. We hope that our work can offer valuable insight and provide a reliable benchmark to guide future research, foster reproducibility, and accelerate innovation in the field of image generation and beyond. The codes and datasets used for evaluating GPT-4o can be found at https://github.com/PicoTrex/GPT-ImgEval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    UDM-GRPO is the first RL integration for uniform discrete diffusion models, using final clean samples as actions and forward-process trajectory reconstruction to raise GenEval accuracy from 69% to 96% and OCR accuracy...

  2. AI-generated Images Challenge Visual Trust in High-risk Scenarios

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On SafeIMG, a new safety-focused benchmark of 1,131 GPT Image 2 images, the best VLM detects 49.5% of generated images and the best specialized detector 33.1%, versus 81.7% for humans.

  3. Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

    physics.chem-ph 2026-07 conditional novelty 6.0 of 10

    A two-stage AI pipeline — spectral hypothesis generation followed by mass-constrained molecular refinement — reconstructs organic structures from multimodal spectra, with 93.8% top-1 accuracy on simulated QM9 data and...

  4. Reconstruction Alignment Improves Unified Multimodal Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    RECA, a self-supervised post-training objective that conditions unified multimodal models on their own visual understanding embeddings to reconstruct input images, improves text-to-image and editing benchmarks across ...

  5. UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A reinforcement-learning reward based on bipartite face matching improves multi-identity consistency and reduces identity confusion in image customization models.

  6. Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A curated GPT-4o synthetic image dataset improves open-source generation models on instruction-following, surreal scenes, and multi-reference synthesis, plus two new benchmarks to measure those skills.

  7. X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

    cs.CV 2025-07 conditional novelty 6.0 of 10

    GRPO reinforcement learning applied to a discrete autoregressive image generator with a diffusion decoder improves instruction following, image quality, and long-text rendering in a unified multimodal model.

  8. ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 91K GPT-4o-generated image and editing dataset, and a fine-tuned open model Janus-4o, report improved text-to-image scores and new editing ability.

  9. ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A ComfyUI-based multi-agent system with semantic workflow modules and tree-based local-feedback planning reports near-perfect pass rates on ComfyBench and competitive scores on GenEval and Reason-Edit.

  10. $I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion

    cs.CL 2025-05 reject novelty 5.0 of 10

    A pairwise-conditioned diffusion model generates instructional illustrations from procedural text and is finetuned with a text-image alignment reward.

  11. CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CAD combines cross-modal lip-speech alignment with per-modality artifact distillation and reports 99.96% AUC on IDForge-v2, with strong cross-dataset results.

Pith tools