Pith. sign in

REVIEW 1 major objections 1 minor 1 cited by

AEGIS benchmark shows current tools detect AI-generated academic images at only 48.8 percent accuracy.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-05-22 10:38 UTC pith:G4ATTNW5

load-bearing objection AEGIS introduces a specialized benchmark for forensic analysis of AI-generated academic images, highlighting performance shortfalls but relying on unverified assumptions about its synthetic forgeries representing real cases. the 1 major comments →

arxiv 2604.28177 v2 pith:G4ATTNW5 submitted 2026-04-30 cs.CV cs.CY

AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images

classification cs.CV cs.CY
keywords AI-generated imagesimage forensicsacademic publishingforgery detectionbenchmarklocalizationmultimodal modelsdetection accuracy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces AEGIS as a benchmark to test how well AI systems can detect and analyze images created by generative models for use in academic papers. It organizes content across seven scholarly fields with 39 detailed subtypes and applies four common forgery approaches drawn from 25 different generators. A sympathetic reader would care because unreliable detection could allow fabricated visuals to enter research publications and undermine the trustworthiness of scientific records. The work evaluates models on detection, explanation of decisions, and precise identification of altered areas rather than a single yes-or-no judgment.

Core claim

AEGIS serves as a diagnostic testbed for academic image forensics by covering seven categories with 39 fine-grained subtypes, modeling four prevalent forgery strategies across 25 generative models, and jointly measuring detection, reasoning, and localization, where GPT-5.1 reaches 48.80 percent overall performance, expert models reach 30.09 percent IoU on localization, multimodal large language models reach 84.74 percent on textual artifact recognition, and expert detectors reach 79.54 percent on binary authenticity detection.

What carries the argument

The AEGIS benchmark that combines domain-specific academic image categories, simulated forgery strategies, and a three-part evaluation of detection, reasoning, and localization.

Load-bearing premise

The selected academic categories, 39 subtypes, and four forgery strategies reflect the main real-world difficulties in spotting AI-generated images in scholarly publishing.

What would settle it

A forensic method that reaches above 80 percent accuracy on both overall detection and localization across all seven categories and four forgery strategies in the AEGIS set would challenge the claim of fundamental limitations.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Forensic accuracy stays below 50 percent for images from 11 of the 25 generative models tested.
  • Multimodal language models and specialized expert detectors show complementary performance on different parts of the forensic task.
  • Localization of altered regions in academic images remains especially difficult for all tested systems.
  • Advances in image generation have outpaced current forensic capabilities in the academic domain.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Future work could track progress by repeatedly testing new models on the same AEGIS set over time.
  • A combined system that routes different subtasks to the strongest model type for each might improve overall results.
  • The same diagnostic approach could be applied to other image-heavy fields such as medical or legal documents.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. The paper introduces AEGIS, a holistic benchmark for evaluating forensic analysis of AI-generated academic images. It advances prior work through three contributions: (1) domain-specific complexity via seven academic categories and 39 fine-grained subtypes, where GPT-5.1 achieves 48.80% overall performance and expert models reach only IoU 30.09% for localization; (2) diverse forgery simulations of four prevalent academic strategies across 25 generative models, with 11 yielding below 50% forensic accuracy; and (3) multi-dimensional evaluation of detection, reasoning, and localization across 25 MLLMs, nine expert models, and one unified model. AEGIS is positioned as a diagnostic testbed exposing fundamental limitations in academic image forensics.

Significance. If the benchmark's categories, subtypes, and forgery strategies accurately sample real-world academic image forgery distributions, the work would provide a valuable standardized testbed for the forensics community. Its multi-dimensional evaluation framework (detection + reasoning + localization) and broad model coverage reveal complementary model-family strengths, such as MLLMs at 84.74% in textual artifact recognition versus expert detectors at 79.54% in binary detection. This could usefully guide future tool development for academic publishing integrity.

major comments (1)
  1. [Abstract] Abstract: The central claim that AEGIS 'exposes fundamental limitations' and 'intrinsic forensic difficulty' (with GPT-5.1 at 48.80% and IoU 30.09%) rests on the unvalidated assertion that the seven categories, 39 subtypes, and four forgery strategies 'prevalent' in academic publishing faithfully represent real forensic challenges. No corpus analysis of retracted papers, publisher reports, or expert-labeled real cases is described to quantify coverage, artifact fidelity, or alignment with actual detectable cues such as metadata or semantic inconsistencies. Without this grounding, the reported performance ceilings may reflect benchmark construction choices rather than general limitations.
minor comments (1)
  1. [Abstract] Abstract: Clarify the exact model versions and release dates for references such as 'GPT-5.1' and the 25 generative models to ensure reproducibility of the reported metrics.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for their constructive and detailed review of our manuscript on AEGIS. We address the major comment below and outline revisions that will strengthen the presentation of the benchmark's scope and construction.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The central claim that AEGIS 'exposes fundamental limitations' and 'intrinsic forensic difficulty' (with GPT-5.1 at 48.80% and IoU 30.09%) rests on the unvalidated assertion that the seven categories, 39 subtypes, and four forgery strategies 'prevalent' in academic publishing faithfully represent real forensic challenges. No corpus analysis of retracted papers, publisher reports, or expert-labeled real cases is described to quantify coverage, artifact fidelity, or alignment with actual detectable cues such as metadata or semantic inconsistencies. Without this grounding, the reported performance ceilings may reflect benchmark construction choices rather than general limitations.

    Authors: We thank the referee for this important observation. The seven categories, 39 subtypes, and four forgery strategies were derived from a synthesis of academic publishing guidelines (e.g., COPE and journal integrity policies), documented cases of image-related misconduct in the literature, and input from experts in scientific visualization and research integrity. These choices target representative challenges such as figure duplication, synthetic data insertion, and composite manipulation that appear across disciplines. Nevertheless, the current manuscript does not present a quantitative corpus analysis of retracted papers or publisher databases to measure exact coverage or alignment with cues like metadata. In the revised version, we will expand the Benchmark Construction section with additional references to prior studies on academic image integrity violations and include a dedicated Limitations subsection. This subsection will explicitly state that AEGIS provides a curated diagnostic testbed for prevalent strategies rather than a statistically exhaustive sample of all real-world instances, thereby clarifying that the reported performance figures (e.g., GPT-5.1 at 48.80%) reflect difficulty within the defined scope. revision: yes

Circularity Check

0 steps flagged

Empirical benchmark evaluation with no derivations or self-referential reductions

full rationale

The paper presents AEGIS as an empirical benchmark for AI-generated academic image forensics, describing coverage of seven categories with 39 subtypes, four forgery strategies across 25 models, and evaluation of detection/reasoning/localization on 25 MLLMs plus expert models. No equations, mathematical derivations, fitted parameters, or predictive claims appear in the provided text. Performance figures such as 48.80% overall accuracy and 30.09% IoU are reported as direct evaluation outcomes rather than outputs derived from the benchmark construction itself. The selection of subtypes and strategies is framed as a design choice to expose difficulty, without any reduction to self-definition, self-citation chains, or renaming of prior results. The work is self-contained as a diagnostic testbed whose claims rest on the empirical measurements obtained from the constructed dataset.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 1 invented entities

The central claim depends on the representativeness of the chosen academic categories, subtypes, and forgery strategies for exposing real forensic difficulties, with no independent evidence provided in the abstract for this coverage.

axioms (1)
  • domain assumption The selected seven academic categories with 39 subtypes and four forgery strategies across 25 generative models capture intrinsic forensic difficulty in academic images.
    Invoked to support claims that even advanced models show limited performance and that forensics lag behind generative advances.
invented entities (1)
  • AEGIS benchmark no independent evidence
    purpose: Holistic evaluation framework for forensic analysis of AI-generated academic images
    Newly introduced testbed whose diagnostic value rests on the domain assumptions above.

pith-pipeline@v0.9.0 · 5813 in / 1269 out tokens · 38022 ms · 2026-05-22T10:38:13.075457+00:00 · methodology

0 comments
read the original abstract

We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS features three key advances: (1) Domain-Specific Complexity: covering seven academic categories with 39 fine-grained subtypes, exposing intrinsic forensic difficulty, where even GPT-5.1 reaches 48.80% overall performance and expert models achieve only limited localization accuracy (IoU 30.09%); (2) Diverse Forgery Simulations: modeling four prevalent academic forgery strategies across 25 generative models, with 11 yielding average forensic accuracy below 50%, showing that forensics lag behind generative advances; and (3) Multi-Dimensional Forensic Evaluation: jointly assessing detection, reasoning, and localization, revealing complementary strengths between model families, with multimodal large language models (MLLMs) at 84.74% accuracy in textual artifact recognition and expert detectors peaking at 79.54% accuracy in binary authenticity detection. By evaluating 25 leading MLLMs, nine expert models, and one unified multimodal understanding and generation model, AEGIS serves as a diagnostic testbed exposing fundamental limitations in academic image forensics.

Figures

Figures reproduced from arXiv: 2604.28177 by Bo Zhang, Haihong E, Haiyang Sun, Haocheng Gao, Jiacheng Liu, Junpeng Ding, Liangjia Wang, Peilin Gao, Ronghui Xi, Tzu-Yen Ma, Yiling Huang, Yizhuo Zhao, Yuan Liu, Yuanze Li, Yujie Wang, Yuyue Zhang, Zhongjun Yang, Zichen Tang, Zijie Xi, Zirui Wang, Zixin Ding.

Figure 1
Figure 1. Figure 1: AEGIS investigates whether current models can effectively audit AI-generated academic images in academic papers by performing holistic forensic analy￾sis across four complementary tasks. while, leveraging visual understanding and reason￾ing, multimodal large language models (MLLMs) have been increasingly applied to image forgery analysis, either directly (Wen et al., 2025) or in conjunction with expert mod… view at source ↗
Figure 2
Figure 2. Figure 2: Hierarchical taxonomy of AEGIS. We organize academic images into seven categories and 39 fine￾grained subtypes based on their structural and semantic characteristics. Real (left) and fake (right) examples are shown for comparison, where “fake” refers to AI-generated forgeries. Artifact Recognition, (3) Manipulation Clas￾sification, and (4) Tampering Pinpointing. These tasks comprehensively assess models’ c… view at source ↗
Figure 3
Figure 3. Figure 3: Construction pipeline of AEGIS. Stage 1: Paper Parsing extracts figures, captions, and panels from papers. Stage 2: Expert Curation retains qualified academic panels while excluding non-academic ones. Stage 3: Forgery Strategy Simulation synthesizes AI-generated academic image forgeries via four representative strategies. 2.2 Dataset Construction Data Curation. Papers were parsed into figures and panels, t… view at source ↗
Figure 4
Figure 4. Figure 4: Evaluation design of AEGIS. Four evaluation question types are designed to support staged forensic analysis, from global authenticity assessment to fine-grained region-level localization, including Forgery Scope Discrimination, Textual Artifact Recognition, Manipulation Classification, and Tampering Pinpointing. structing missing local content. By applying masks to specific regions of authentic images, we … view at source ↗
Figure 5
Figure 5. Figure 5: Fine-grained experimental analysis of AEGIS. TCF: Text Constraint Fabrication; IIF: Image Inference Forgery; TRR: Targeted Region Restoration; TRE: Targeted Region Editing. SMG: Stained Micrograph; MG: Micrograph. Seedream: Doubao-Seedream; SD: Stable Diffusion; SN: SenseNova. 3.2 Experimental Analysis 3.2.1 Domain-Specific Complexity Challenges Current Models We first investigate how the intrinsic complex… view at source ↗
Figure 6
Figure 6. Figure 6: Impact of post-processing perturbations on AEGIS. FSD: Forgery Scope Discrimination; TAR: Textual Artifact Recognition; MC: Manipulation Clas￾sification; TP: Tampering Pinpointing. 25 55 85 FSD TAR MC A c c u r a c y ( % ) Doubao-Seed-1.6-thinking (Zero-Shot) GPT-5.1 (Zero-Shot) Qwen3-VL-Plus (Few-Shot) Doubao-Seed-1.6-thinking (Few-Shot) GPT-5.1 (Few-Shot) Qwen3-VL-Plus (Zero-Shot) view at source ↗
Figure 7
Figure 7. Figure 7: Impact of Few-Shot prompting on AEGIS. degrade sharply under post-processing perturba￾tions (i.e., Gaussian blurring, JPEG compression and image scaling; view at source ↗
Figure 8
Figure 8. Figure 8: Example of a retracted paper published in view at source ↗
Figure 9
Figure 9. Figure 9: Example of a retracted paper from the Lippin view at source ↗
Figure 10
Figure 10. Figure 10: Distribution of real images and images generated by four forgery strategie in AEGIS view at source ↗
Figure 11
Figure 11. Figure 11: Distribution of seven academic image cat￾egories in AEGIS. A.4 Data Curation Paper Parsing. AEGIS collected 4,362 high￾quality academic papers from the open-access PMC repository and performed document-level parsing to construct the initial visual corpus. The paper selection criteria were as follows: • Structural Completeness. Each paper must con￾tain at least four independent figures to ensure sufficient… view at source ↗
Figure 13
Figure 13. Figure 13: Success case by GPT-5.1 on Forgery Scope Discrimination. The panel belongs to the Micrograph category view at source ↗
Figure 14
Figure 14. Figure 14: Failure case by GPT-5.1 on Forgery Scope Discrimination. The panel belongs to the Stained Micrograph category view at source ↗
Figure 15
Figure 15. Figure 15: Failure case by GPT-5.1 on Forgery Scope Discrimination. The panel belongs to the Diagram category view at source ↗
Figure 16
Figure 16. Figure 16: Failure case by GPT-5.1 on Forgery Scope Discrimination. The panel belongs to the Chart category view at source ↗
Figure 17
Figure 17. Figure 17: Failure case by GPT-5.1 on Forgery Scope Discrimination. The panel belongs to the Medical Imaging category view at source ↗
Figure 18
Figure 18. Figure 18: Success case by GPT-5.1 on Textual Artifact Recognition. The panel belongs to the Diagram category view at source ↗
Figure 19
Figure 19. Figure 19: Failure case by GPT-5.1 on Textual Artifact Recognition. The panel belongs to the Stained Micrograph category and contains a localized forgery generated via Targeted Region Editing view at source ↗
Figure 20
Figure 20. Figure 20: Success case by GPT-5.1 on Manipulation Classification. The panel belongs to the Micrograph category and contains a localized forgery generated via Targeted Region Editing view at source ↗
Figure 21
Figure 21. Figure 21: Failure case by GPT-5.1 on Manipulation Classification. The panel belongs to the Medical Imaging category and contains a localized forgery generated via Targeted Region Restoration view at source ↗
Figure 22
Figure 22. Figure 22: Failure case by GPT-5.1 on Manipulation Classification. The panel belongs to the Medical Imaging category and contains a localized forgery generated via Targeted Region Restoration view at source ↗
Figure 23
Figure 23. Figure 23: Success case by GPT-5.1 on Tampering Pinpointing. The panel belongs to the Physical Object category and contains a localized forgery generated via Targeted Region Restoration view at source ↗
Figure 24
Figure 24. Figure 24: Failure case by GPT-5.1 on Tampering Pinpointing. The panel belongs to the Stained Micrograph category and contains a localized forgery generated via Targeted Region Restoration view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

    cs.CV 2026-07 conditional novelty 6.0

    Strengthening fine-grained, semantic-anomaly, and pixel-level perception with verifiable rewards, then value-aware on-policy self-distillation, improves generalizable MLLM AI-image detection and adaptation.

Reference graph

Works this paper leans on

3 extracted references · 3 canonical work pages · cited by 1 Pith paper · 1 internal anchor

  1. [1]

    dots.ocr: Multilingual document layout parsing in a single vision-language model, 2025

    On the detection of synthetic images generated by diffusion models. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE. Google DeepMind. 2025a. Gemini 2.5 flash and pro are now generally available, and we’re introducing 2.5 flash-lite, our most cost-efficient and fastest 2.5 model yet. htt...

  2. [2]

    DINOv3

    Llava-next: Improved reasoning, ocr, and world knowledge. https://llava-vl.github.io/ blog/2024-01-30-llava-next/ . Accessed: 2025- 04-05. Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022. Repaint: Inpainting using denoising diffusion prob- abilistic models. InProceedings of the IEEE/CVF Conference on Compu...

  3. [3]

    Lower FID scores indi- cate higher visual fidelity and closer alignment with real-image statistics

    evaluates the distributional similarity be- tween generated images and real academic im- ages by computing the Fréchet distance between their feature embeddings extracted from a pre- trained Inception model. Lower FID scores indi- cate higher visual fidelity and closer alignment with real-image statistics. • CLIP Score.CLIP Score (Hessel et al., 2021) mea...