Pith. sign in

REVIEW 5 cited by

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.22126 v1 pith:RVJEKDA2 submitted 2025-05-28 cs.CV cs.AI

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model

classification cs.CV cs.AI
keywords generationscientificbenchmarkimagemodelsgpt-4o-imagehumanillustration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent years have seen rapid advances in AI-driven image generation. Early diffusion models emphasized perceptual quality, while newer multimodal models like GPT-4o-image integrate high-level reasoning, improving semantic understanding and structural composition. Scientific illustration generation exemplifies this evolution: unlike general image synthesis, it demands accurate interpretation of technical content and transformation of abstract ideas into clear, standardized visuals. This task is significantly more knowledge-intensive and laborious, often requiring hours of manual work and specialized tools. Automating it in a controllable, intelligent manner would provide substantial practical value. Yet, no benchmark currently exists to evaluate AI on this front. To fill this gap, we introduce SridBench, the first benchmark for scientific figure generation. It comprises 1,120 instances curated from leading scientific papers across 13 natural and computer science disciplines, collected via human experts and MLLMs. Each sample is evaluated along six dimensions, including semantic fidelity and structural accuracy. Experimental results reveal that even top-tier models like GPT-4o-image lag behind human performance, with common issues in text/visual clarity and scientific correctness. These findings highlight the need for more advanced reasoning-driven visual generation capabilities.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models

    cs.CV 2026-04 unverdicted novelty 7.0

    GENFIG1 is a new benchmark that tests whether vision-language models can create effective Figure 1 visuals capturing the central scientific idea from paper text.

  2. SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

    cs.CL 2026-07 conditional novelty 6.0

    Scientific-figure editing can be learned from arXiv revision pairs: a skill-evolving SVG agent follows edit instructions and transfers learned skills across LLM backbones.

  3. SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces SciIR-82k dataset and SciIR-Bench for scientific image reasoning generation organized by Peirce's semiotic triad, with fine-tuning raising model score from 35% to 43%.

  4. Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces FEPBench benchmark to evaluate T2I models on instruction faithfulness, reasoning enrichment, and semantic precision for natural-science illustrations using atom set annotations.

  5. SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

    cs.CV 2026-02 conditional novelty 6.0

    A structure-first benchmark finds that text-to-image models preserve little recoverable graph structure in scientific diagrams, with the best model scoring 0.116 graph-level versus 0.01-0.09 for others.