REVIEW 4 major objections 8 minor 3 cited by
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
T0 review · 4 major / 8 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read SridBench, the first benchmark for scientific figure generation, finds that even the top model, GPT-4o-image, remains at 'fair' level.
desk verdict Useful dataset and a plausible central finding, but the 'first benchmark' claim is undercut by the paper's own citation of ScImage, and the automatic judge is too thinly validated to support the reported numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the triplet structure (reference image, caption, related section text) and the six-dimension scoring protocol. Each instance is a generation task: a model is given only the caption and section text and must produce a diagram. A multimodal judge (GPT-4o) then scores the output against the original figure on six 1–5 scales, a protocol the authors validate on a subset of 100 instances against human experts. The filter pipeline uses MLLMs to keep only concept, framework, flow, and structure diagrams, excluding photos, experimental-result graphs, and statistical plots.
What would settle it
A fresh, independent human re-scoring of a random sample of, say, 200 instances from the full SridBench set that shows systematic discrepancies between GPT-4o's scores and human judgments—particularly on GPT-4o-image outputs—would show that the reported gap between models and humans is not reliably measured.
Extended reading notes
Core claim
SridBench is positioned as the first evaluation resource specifically for generating scientific research illustrations from textual descriptions. The benchmark comprises 1,120 triplets of image, caption, and related section text, collected from authoritative peer-reviewed publications and filtered by human experts with MLLM assistance. The paper's main empirical claim is that no tested model is close to human-level: GPT-4o-image achieves roughly 'fair' scores across the six dimensions, whereas Gemini-2.0-Flash scores below 2 on every dimension, and Emu-3 fails to produce relevant content. The bottlenecks the authors identify are missing or inaccurate textual elements and scientific errors in the generated diagrams.
Load-bearing premise
The entire evaluation assumes that GPT-4o's automatic scoring—validated against human experts on only 100 of the 1,120 instances—remains accurate for all 1,120 instances, including images produced by GPT-4o-image, a model from the same family.
Editorial extensions
If this is right
- Any future image-generation model can be benchmarked on SridBench and compared on a common six-dimension rubric.
- The observed gap between text accuracy and text completeness indicates that models tend to omit textual details rather than garble them, pointing to a specific technical target.
- Open-source and non-specialist models scoring near 1 show that scientific illustration remains effectively unsolved for most systems.
- The six dimensions could serve as a template for evaluating other forms of technical or instructional diagram generation.
Reading between the lines
- The validity of the headline result depends on the 100-instance human validation of GPT-4o as judge; extending that validation to a larger random sample would strengthen or weaken the reported rankings.
- Using GPT-4o to grade images produced by GPT-4o-image, a closely related model, leaves open a same-family bias that the small human check may not fully detect.
- The 'first benchmark' claim is contingent on the definition of scientific illustration; if one includes chart or graph generation benchmarks, the novelty is narrower, though this paper specifically targets schematic figures.
- A concrete next step suggested by the paper's data is feeding models the full LaTeX source of the section instead of plain text, which might improve text completeness.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SridBench, a benchmark of 1,120 instances for scientific research illustration generation, spanning 13 disciplines in natural and computer science. Each instance is a triple of (figure, caption, related section), collected from arXiv and Nature and screened by human experts and MLLMs. The authors propose a six-dimension evaluation protocol (completeness and accuracy of textual information, diagrammatic structural integrity, diagrammatic logic, cognitive readability, aesthetic feeling) and report results for GPT-4o-image, Gemini-2.0-Flash, and Emu-3, with GPT-4o serving as an automated judge. The central findings are that GPT-4o-image scores around 3/5 ('fair') and falls far short of human experts, while Gemini-2.0-Flash scores below 2 and Emu-3 is qualitatively unusable. A 100-instance human comparison is used to justify the use of GPT-4o as the automated judge.
Significance. If the benchmark is properly constructed and the evaluation method is reliable, SridBench would be a useful resource for tracking progress in a practically important task, and the six-dimension rubric is a reasonable starting point. The human comparison on 100 instances is a positive step toward grounding the automated judge. However, the significance is conditional on resolving the novelty claim (the paper's own reference [14] ScImage appears to be a scientific text-to-image generation benchmark) and on providing stronger evidence that the GPT-4o judge generalizes across the full dataset. The empirical finding that GPT-4o-image is only at a 'fair' level is plausible but currently lacks error bars and significance tests.
major comments (4)
- [Abstract, Sections 1 and 2 (Research Gaps), and Conclusion] The paper repeatedly claims that 'no benchmark currently exists' and that SridBench is 'the first benchmark' for scientific figure generation. This is contradicted by the manuscript's own reference [14], titled 'ScImage: How good are multimodal large language models at scientific text-to-image generation?' which, by its title, is exactly a generation benchmark for scientific figures. The text in Section 2, however, describes ScImage as focused on 'understanding capabilities.' This is an internal inconsistency. The authors must position SridBench relative to ScImage, compare dataset construction, metrics, and scope, and clearly state what is genuinely novel. As written, the headline novelty claim is unsupported.
- [Section 4.2, Figure 3(b)] The automated judge (GPT-4o) is validated on only 100 of the 1,120 instances (50 natural science, 50 computer science). The paper states that 'GPT-4o scores are broadly in line with those of human experts' without reporting any quantitative agreement metric (e.g., correlation, mean absolute error, or per-dimension breakdown). Furthermore, GPT-4o is used to judge images generated by GPT-4o-image, a model from the same family. Since the central quantitative results—GPT-4o-image scoring around 3/5 and the human-model gap—are derived entirely from this judge, the paper must either provide a more extensive validation (e.g., on a larger and more diverse sample) or report confidence intervals and agreement statistics. Without this, the size of the human-model gap is not established.
- [Sections 4.2–4.5, Figures 3–6] The reported scores are averages without error bars, variance, or significance tests. For example, Figure 3(a) shows average scores per dimension per model, but there is no indication of the variability across the 1,120 instances, the number of generation runs, or whether the differences between models and human experts are statistically significant. The headline claim that 'GPT-4o-image falls far short of human-level performance' rests on these averages. The authors should report standard deviations, confidence intervals, or perform appropriate statistical tests to support the strength of the conclusions.
- [Section 4.1] Only two models are fully evaluated quantitatively (GPT-4o-image and Gemini-2.0-Flash); Emu-3 is excluded from the quantitative analysis due to generation time. This is a narrow model set for a 'comprehensive empirical study' as claimed in the contributions. At minimum, the paper should state this limitation explicitly and discuss how the findings might generalize to other model families.
minor comments (8)
- [Section 1] 'Scientific illustration are essential tools' should be 'Scientific illustrations are essential tools.'
- [Section 2] 'Stable Diffusion series have demonstrated' should be 'Stable Diffusion series has demonstrated' to agree in number.
- [Section 3.1] The sentence 'The text should also be able to support and cover the elements that generate the illustration' is unclear and should be rephrased.
- [Section 4.4] The phrase 'GPT-4o-image alternative shows absolutely no difference' is confusing; it appears to mean that performance shows no significant difference across subjects, but the wording should be revised.
- [Section 4.2] The sentence 'Therefore, we use GPT-4o for automated scoring' appears in the middle of a paragraph about the human comparison; consider moving it to the evaluation methodology section for clarity.
- [Appendix A] In the Image Judgement prompt, 'the first one by an anthropologist' should likely be 'the first one by a human expert.'
- [General] The paper does not provide a link to the dataset or code, which limits reproducibility; a public release plan should be included.
- [References] Reference [1] appears unrelated to the diffusion model introduction; please verify and correct.
Circularity Check
The paper's central 'first benchmark / blank field' claim is manufactured by reclassifying its own cited ScImage [14]—a scientific text-to-image generation benchmark—as an 'understanding' benchmark; the empirical scoring chain is not circular.
-
renaming known result
[Section 1 (Introduction), Section 2 (Related Work, Research Gaps), and References [14]; echoed in Abstract and Section 5]
"most of the current research efforts are focused on the understanding of scientific images and the generation of image captions (such as SciFIBench [13], FigCaps-HF [28], etc. [29]). The field of evaluating the generation of scientific research drawings is almost blank. ... mainly focused on benchmarking the understanding capabilities of multimodal models (e.g., SciFIBench [13], ScImage [14]). ... [14] L. Zhang, ... "ScImage: How good are multimodal large language models at scientific text-to-image generation?""
The paper's load-bearing novelty claim is that the field of evaluating scientific figure generation is 'almost blank' and that SridBench is 'the first benchmark.' That claim is obtained by assigning ScImage [14] to 'understanding capabilities' in the Introduction, while the paper's own bibliography titles [14] as a benchmark for 'scientific text-to-image generation.' If ScImage is a generation benchmark, the field is not blank and SridBench is not first; the central novelty step reduces to a reclassification of the cited prior work. The empirical evaluation chain—human validation on 100 samples followed by GPT-4o scoring—is not itself circular, so the overall circularity is partial rather than total.
full rationale
Most of the paper's derivation chain is self-contained: the 1,120 triples are collected from papers using an MLLM filter plus human-expert screening, the generation prompts are specified, and the automated evaluation uses a GPT-4o scorer whose agreement with human experts is checked on a 100-instance subset. That is a validation/generalization step rather than a construction-level circularity. The same-family overlap (GPT-4o judging GPT-4o-image) is a legitimate validity risk, but it does not make the reported scores equal to the inputs by construction, and the human subset provides independent grounding for the broad conclusion that models lag humans. The circular/renaming content is concentrated in the novelty argument: ScImage [14], a scientific text-to-image generation benchmark, is categorized as an 'understanding' benchmark so that the paper can declare the field blank and itself 'the first benchmark.' This reclassification drives the score of 4 rather than a 0-2 non-finding.
Assumptions & free parameters
assumptions (4)
- domain assumption Filtering to four schematic types (concept, model frame, process flow, structure) defines the scope of 'scientific illustration'.
- domain assumption GPT-4o's scores are a valid proxy for human expert evaluation of the six dimensions.
- domain assumption Reference figures from high-citation papers in top venues are suitable ground truth for generation quality.
- domain assumption Generating a figure from caption plus surrounding section text is a faithful operationalization of scientific illustration drawing.
Cite this review
Pith. "Pith review of SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model." pith.science (2026). https://pith.science/paper/RVJEKDA2
@misc{pith2026250522126,
author = {Pith},
title = {Pith review of: SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/RVJEKDA2}},
note = {Machine review of arXiv:2505.22126}
}
read the original abstract
Recent years have seen rapid advances in AI-driven image generation. Early diffusion models emphasized perceptual quality, while newer multimodal models like GPT-4o-image integrate high-level reasoning, improving semantic understanding and structural composition. Scientific illustration generation exemplifies this evolution: unlike general image synthesis, it demands accurate interpretation of technical content and transformation of abstract ideas into clear, standardized visuals. This task is significantly more knowledge-intensive and laborious, often requiring hours of manual work and specialized tools. Automating it in a controllable, intelligent manner would provide substantial practical value. Yet, no benchmark currently exists to evaluate AI on this front. To fill this gap, we introduce SridBench, the first benchmark for scientific figure generation. It comprises 1,120 instances curated from leading scientific papers across 13 natural and computer science disciplines, collected via human experts and MLLMs. Each sample is evaluated along six dimensions, including semantic fidelity and structural accuracy. Experimental results reveal that even top-tier models like GPT-4o-image lag behind human performance, with common issues in text/visual clarity and scientific correctness. These findings highlight the need for more advanced reasoning-driven visual generation capabilities.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 3 Pith papers
-
SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions
Scientific-figure editing can be learned from arXiv revision pairs: a skill-evolving SVG agent follows edit instructions and transfers learned skills across LLM backbones.
-
SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing
A structure-first benchmark finds that text-to-image models preserve little recoverable graph structure in scientific diagrams, with the best model scoring 0.116 graph-level versus 0.01-0.09 for others.
-
AI4Research: A Survey of Artificial Intelligence for Scientific Research
A survey that organizes AI-for-research work into five tasks, comprehension, survey, discovery, writing, and peer review, and compiles associated tools and benchmarks.
Reference graph
Works this paper leans on
-
[14]
Scimage: How good are multimodal large language models at scientific text-to-image generation?
L. Zhang, S. Eger, Y . Cheng, W. Zhai, J. Belouadi, C. Leiter, S. P. Ponzetto, F. Moafian, and Z. Zhao, “Scimage: How good are multimodal large language models at scientific text-to-image generation?” 2024. [Online]. Available: https://arxiv.org/abs/2412.02368
arXiv 2024
-
[1]
DSM Refinement with Deep Encoder-Decoder Networks
N. Metzger, “Dsm refinement with deep encoder-decoder networks,” 2020. [Online]. Available: https://arxiv.org/abs/2012.07427
work page Pith review arXiv 2020
-
[2]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” 2020. [Online]. Available: https://arxiv.org/abs/2006.11239
arXiv 2020
-
[3]
High-resolution image synthesis with latent diffusion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” inCVPR, 2022, pp. 10 684–10 695
work page 2022
-
[4]
Scaling rectified flow transformers for high-resolution image synthesis,
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y . Levi, D. Lorenz, A. Sauer, F. Boeselet al., “Scaling rectified flow transformers for high-resolution image synthesis,” in Forty-first International Conference on Machine Learning, 2024
work page 2024
-
[5]
Hierarchical text-conditional image generation with clip latents,
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, “Hierarchical text-conditional image generation with clip latents,”arXiv preprint arXiv:2204.06125, vol. 1, no. 2, p. 3, 2022
arXiv 2022
-
[6]
Improving image generation with better captions,
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y . Guo et al., “Improving image generation with better captions,”Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, vol. 2, no. 3, p. 8, 2023
work page 2023
-
[7]
Black Forest Labs, “Flux,” https://github.com/black-forest-labs/flux, 2024, accessed: 2024-11- 05
work page 2024
Show all 30 references
-
[8]
Emu3: Next-token prediction is all you need,
X. Wang, X. Zhang, Z. Luo, Q. Sun, Y . Cui, J. Wang, F. Zhang, Y . Wang, Z. Li, Q. Yuet al., “Emu3: Next-token prediction is all you need,”arXiv preprint arXiv:2409.18869, 2024
2024 arXiv
-
[9]
Neighboring autoregressive modeling for efficient visual generation,
Y . He, Y . He, S. He, F. Chen, H. Zhou, K. Zhang, and B. Zhuang, “Neighboring autoregressive modeling for efficient visual generation,” 2025. [Online]. Available: https://arxiv.org/abs/2503.10696
2025 arXiv
-
[10]
Visual autoregressive modeling: Scalable image generation via next-scale prediction,
K. Tian, Y . Jiang, Z. Yuan, B. Peng, and L. Wang, “Visual autoregressive modeling: Scalable image generation via next-scale prediction,”Advances in neural information processing systems, vol. 37, pp. 84 839–84 865, 2024
2024
-
[11]
Janus-pro: Uni- fied multimodal understanding and generation with data and model scaling,
X. Chen, Z. Wu, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, and C. Ruan, “Janus-pro: Uni- fied multimodal understanding and generation with data and model scaling,”arXiv preprint arXiv:2501.17811, 2025
2025 arXiv
-
[12]
Addendum to gpt-4o system card: 4o image generation,
OpenAI, “Addendum to gpt-4o system card: 4o image generation,” 2025, accessed: 2025-04-02. [Online]. Available: https://openai.com/index/ gpt-4o-image-generation-system-card-addendum/
2025
-
[13]
Scifibench: Benchmarking large multimodal models for scientific figure interpretation,
J. Roberts, K. Han, N. Houlsby, and S. Albanie, “Scifibench: Benchmarking large multimodal models for scientific figure interpretation,” 2024. [Online]. Available: https://arxiv.org/abs/2405.08807
2024 arXiv
-
[15]
Sdxl: Improving latent diffusion models for high-resolution image synthesis,
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rom- bach, “Sdxl: Improving latent diffusion models for high-resolution image synthesis,”arXiv preprint arXiv:2307.01952, 2023
2023 arXiv
-
[16]
Mlperf inference benchmark,
V . J. Reddi, C. Cheng, D. Kanter, P. Mattson, G. Schmuelling, C.-J. Wu, B. Anderson, M. Breughe, M. Charlebois, W. Chou, R. Chukka, C. Coleman, S. Davis, P. Deng, G. Diamos, J. Duke, D. Fick, J. S. Gardner, I. Hubara, S. Idgunji, T. B. Jablin, J. Jiao, T. S. John, P. Kanwar, ...
2020 arXiv
-
[17]
Designbench: Exploring and benchmarking dall- e 3 for imagining visual design,
K. Lin, Z. Yang, L. Li, J. Wang, and L. Wang, “Designbench: Exploring and benchmarking dall- e 3 for imagining visual design,” 2023. [Online]. Available: https://arxiv.org/abs/2310.15144
2023 arXiv
-
[18]
Vila-u: a unified foundation model integrating visual understanding and generation,
Y . Wu, Z. Zhang, J. Chen, H. Tang, D. Li, Y . Fang, L. Zhu, E. Xie, H. Yin, L. Yiet al., “Vila-u: a unified foundation model integrating visual understanding and generation,”arXiv preprint arXiv:2409.04429, 2024
2024 arXiv
-
[19]
Generative multimodal models are in-context learners,
Q. Sun, Y . Cui, X. Zhang, F. Zhang, Q. Yu, Z. Luo, Y . Wang, Y . Rao, J. Liu, T. Huang, and X. Wang, “Generative multimodal models are in-context learners,” 2023
2023
-
[20]
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,
K. Huang, K. Sun, E. Xie, Z. Li, and X. Liu, “T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,”Advances in Neural Information Processing Systems, vol. 36, pp. 78 723–78 747, 2023
2023
-
[21]
Qwen2-vl: En- hancing vision-language model’s perception of the world at any resolution,
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, Y . Fan, K. Dang, M. Du, X. Ren, R. Men, D. Liu, C. Zhou, J. Zhou, and J. Lin, “Qwen2-vl: En- hancing vision-language model’s perception of the world at any resolution,”arXiv preprint arXiv:2409...
2024 arXiv
-
[22]
Qwen2.5-vl technical report,
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y . Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y . Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin, “Qwen2.5-vl technical report,”arXiv prepri...
2025 arXiv
-
[23]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Luet al., “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 202...
2024
-
[24]
Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling,
Z. Chen, W. Wang, Y . Cao, Y . Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liuet al., “Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling,”arXiv preprint arXiv:2412.05271, 2024
2024 arXiv
-
[25]
Chain-of- thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhouet al., “Chain-of- thought prompting elicits reasoning in large language models,”NeurIPS, 2022
2022
-
[26]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Biet al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,”arXiv preprint arXiv:2501.12948, 2025
2025 arXiv
-
[27]
QwQ: Reflect Deeply on the Boundaries of the Unknown,
Q Team, “QwQ: Reflect Deeply on the Boundaries of the Unknown,” Nov. 2024, accessed: 2025-01-01. [Online]. Available: https://qwenlm.github.io/blog/qwq-32b-preview/
2024
-
[28]
Figcaps-hf: A figure-to-caption generative framework and benchmark with human feedback,
A. Singh, P. Agarwal, Z. Huang, A. Singh, T. Yu, S. Kim, V . Bursztyn, N. Vlassis, and R. A. Rossi, “Figcaps-hf: A figure-to-caption generative framework and benchmark with human feedback,” 2023. [Online]. Available: https://arxiv.org/abs/2307.10867
2023 arXiv
-
[29]
Chartbench: A benchmark for complex visual reasoning in charts,
Z. Xu, S. Du, Y . Qi, C. Xu, C. Yuan, and J. Guo, “Chartbench: A benchmark for complex visual reasoning in charts,” 2024. [Online]. Available: https://arxiv.org/abs/2312.15915
2024 arXiv
-
[30]
Gemini: a family of highly capable multimodal models,
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millicanet al., “Gemini: a family of highly capable multimodal models,”arXiv preprint arXiv:2312.11805, 2023. 12 A Prompts used in data process and judgement Image FilterPlea...
2023 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.