{"work":{"id":"80b0a285-2676-4088-8aec-195dac7e8ceb","openalex_id":null,"doi":null,"arxiv_id":"2504.02782","raw_key":null,"title":"GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation","authors":null,"authors_text":"Z","year":2025,"venue":"cs.CV","abstract":"The recent breakthroughs in OpenAI's GPT4o model have demonstrated surprisingly good capabilities in image generation and editing, resulting in significant excitement in the community. This technical report presents the first-look evaluation benchmark (named GPT-ImgEval), quantitatively and qualitatively diagnosing GPT-4o's performance across three critical dimensions: (1) generation quality, (2) editing proficiency, and (3) world knowledge-informed semantic synthesis. Across all three tasks, GPT-4o demonstrates strong performance, significantly surpassing existing methods in both image generation control and output quality, while also showcasing exceptional knowledge reasoning capabilities. Furthermore, based on the GPT-4o's generated data, we propose a classification-model-based approach to investigate the underlying architecture of GPT-4o, where our empirical results suggest the model consists of an auto-regressive (AR) combined with a diffusion-based head for image decoding, rather than the VAR-like architectures. We also provide a complete speculation on GPT-4o's overall architecture. In addition, we conduct a series of analyses to identify and visualize GPT-4o's specific limitations and the synthetic artifacts commonly observed in its image generation. We also present a comparative study of multi-round image editing between GPT-4o and Gemini 2.0 Flash, and discuss the safety implications of GPT-4o's outputs, particularly their detectability by existing image forensic models. We hope that our work can offer valuable insight and provide a reliable benchmark to guide future research, foster reproducibility, and accelerate innovation in the field of image generation and beyond. The codes and datasets used for evaluating GPT-4o can be found at https://github.com/PicoTrex/GPT-ImgEval.","external_url":"https://arxiv.org/abs/2504.02782","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-05T11:41:02.715765+00:00","pith_arxiv_id":"2504.02782","created_at":"2026-05-10T05:36:01.895461+00:00","updated_at":"2026-07-05T11:41:02.715765+00:00","title_quality_ok":true,"display_title":"Gpt-imgeval: A comprehen- sive benchmark for diagnosing gpt4o in image generation","render_title":"Gpt-imgeval: A comprehen- sive benchmark for diagnosing gpt4o in image generation"},"hub":{"state":{"work_id":"80b0a285-2676-4088-8aec-195dac7e8ceb","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":10,"external_cited_by_count":null,"distinct_field_count":1,"first_pith_cited_at":"2024-11-23T19:10:32+00:00","last_pith_cited_at":"2026-05-21T08:00:49+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-06-26T10:24:54.916602+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":4},{"context_role":"dataset","n":1},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":5},{"context_polarity":"use_method","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}