{"id":"2a7933ff-5fa7-461b-bb65-fb472ff81905","arxiv_id":"2505.22126","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SridBench provides a large multi-discipline benchmark for scientific illustration generation and shows current image generation models, especially GPT-4o-image, remain far below human expert quality.","lead":"This paper introduces SridBench, a benchmark of 1,120 scientific figure generation tasks drawn from 13 disciplines, with six scoring dimensions. It evaluates GPT-4o-image and Gemini-2.0-Flash and finds that even the best model scores only around fair, far below human-created figures.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first benchmark' claim is contradicted by the paper's own citation of ScImage [14], which is a scientific text-to-image generation benchmark; if ScImage already evaluates generation, the central novelty claim is false and the manuscript is internally inconsistent.","rationale":"I read the paper in good faith. The construction effort is substantial: 1,120 triples across 13 disciplines, human expert screening, and a six-dimension evaluation protocol are real contributions, and the qualitative failure-mode analysis of GPT-4o-image is informative. The reader's conditional verdict already flags the lack of dataset/code release, the limited 100-sample judge validation, and possible novelty overstatement relative to ScImage. My stress-test sharpens the novelty point: it is not merely an overstatement but an internal inconsistency, because the manuscript itself cites ScImage as 'understanding' work when its title and scope are scientific text-to-image generation. This is the single most load-bearing concern because it attacks the central claim directly and can be settled by reading one cited paper. The judge-proxy issue remains real but secondary: even if the human validation on 100 instances is accepted, the headline novelty claim still fails if ScImage already exists. I do not recommend changing the reader's verdict because the conditional already requires resolving novelty and release issues; however, the revision must specifically address ScImage before the benchmark can be considered a first or definitive contribution.","tokens_in":9925,"tokens_out":6088,"duration_ms":67399,"concrete_test":"Read the full text of arXiv:2412.02368 (ScImage). If it contains a benchmark and evaluation protocol for generating scientific figures from text (rather than only understanding or captioning them), then the statements 'no benchmark currently exists' and 'first benchmark' in the abstract, Section 1, and Section 5 are false; the paper should be revised to compare with ScImage and to limit the novelty claim to specific properties of SridBench (e.g., 1,120 triples, 13 disciplines, six-dimension protocol). If ScImage only benchmarks understanding, the novelty claim survives; run this check before accepting the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 2, the authors write that existing work is 'mainly focused on benchmarking the understanding capabilities of multimodal models (e.g., SciFIBench [13], ScImage [14])' and conclude in the Research Gaps paragraph that 'the field of evaluating the generation of scientific research drawings is almost blank.' Yet reference [14] is titled 'ScImage: How good are multimodal large language models at scientific text-to-image generation?' and, by its title, is exactly a benchmark for scientific figure generation, not understanding. The abstract, introduction, and conclusion repeat that 'no benchmark currently exists' and call SridBench 'the first benchmark.' This is not a mere marketing overstatement; it is an internal contradiction with the paper's own cited literature. If ScImage already provides a scientific text-to-image generation benchmark, the central claim that SridBench fills an empty field is false, and the paper must at minimum position SridBench relative to ScImage, compare dataset construction and metrics, and demonstrate what is genuinely new. This concern is more basic than the judge-proxy worry: it invalidates the headline novelty claim on the face of the manuscript, independent of any statistical generalization argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SridBench, a benchmark of 1,120 instances for scientific research illustration generation, spanning 13 disciplines in natural and computer science. Each instance is a triple of (figure, caption, related section), collected from arXiv and Nature and screened by human experts and MLLMs. The authors propose a six-dimension evaluation protocol (completeness and accuracy of textual information, diagrammatic structural integrity, diagrammatic logic, cognitive readability, aesthetic feeling) and report results for GPT-4o-image, Gemini-2.0-Flash, and Emu-3, with GPT-4o serving as an automated judge. The central findings are that GPT-4o-image scores around 3/5 ('fair') and falls far short of human experts, while Gemini-2.0-Flash scores below 2 and Emu-3 is qualitatively unusable. A 100-instance human comparison is used to justify the use of GPT-4o as the automated judge.","tokens_in":10162,"tokens_out":3753,"duration_ms":35547,"significance":"If the benchmark is properly constructed and the evaluation method is reliable, SridBench would be a useful resource for tracking progress in a practically important task, and the six-dimension rubric is a reasonable starting point. The human comparison on 100 instances is a positive step toward grounding the automated judge. However, the significance is conditional on resolving the novelty claim (the paper's own reference [14] ScImage appears to be a scientific text-to-image generation benchmark) and on providing stronger evidence that the GPT-4o judge generalizes across the full dataset. The empirical finding that GPT-4o-image is only at a 'fair' level is plausible but currently lacks error bars and significance tests.","major_comments":[{"comment":"The paper repeatedly claims that 'no benchmark currently exists' and that SridBench is 'the first benchmark' for scientific figure generation. This is contradicted by the manuscript's own reference [14], titled 'ScImage: How good are multimodal large language models at scientific text-to-image generation?' which, by its title, is exactly a generation benchmark for scientific figures. The text in Section 2, however, describes ScImage as focused on 'understanding capabilities.' This is an internal inconsistency. The authors must position SridBench relative to ScImage, compare dataset construction, metrics, and scope, and clearly state what is genuinely novel. As written, the headline novelty claim is unsupported.","section":"Abstract, Sections 1 and 2 (Research Gaps), and Conclusion"},{"comment":"The automated judge (GPT-4o) is validated on only 100 of the 1,120 instances (50 natural science, 50 computer science). The paper states that 'GPT-4o scores are broadly in line with those of human experts' without reporting any quantitative agreement metric (e.g., correlation, mean absolute error, or per-dimension breakdown). Furthermore, GPT-4o is used to judge images generated by GPT-4o-image, a model from the same family. Since the central quantitative results—GPT-4o-image scoring around 3/5 and the human-model gap—are derived entirely from this judge, the paper must either provide a more extensive validation (e.g., on a larger and more diverse sample) or report confidence intervals and agreement statistics. Without this, the size of the human-model gap is not established.","section":"Section 4.2, Figure 3(b)"},{"comment":"The reported scores are averages without error bars, variance, or significance tests. For example, Figure 3(a) shows average scores per dimension per model, but there is no indication of the variability across the 1,120 instances, the number of generation runs, or whether the differences between models and human experts are statistically significant. The headline claim that 'GPT-4o-image falls far short of human-level performance' rests on these averages. The authors should report standard deviations, confidence intervals, or perform appropriate statistical tests to support the strength of the conclusions.","section":"Sections 4.2–4.5, Figures 3–6"},{"comment":"Only two models are fully evaluated quantitatively (GPT-4o-image and Gemini-2.0-Flash); Emu-3 is excluded from the quantitative analysis due to generation time. This is a narrow model set for a 'comprehensive empirical study' as claimed in the contributions. At minimum, the paper should state this limitation explicitly and discuss how the findings might generalize to other model families.","section":"Section 4.1"}],"minor_comments":[{"comment":"'Scientific illustration are essential tools' should be 'Scientific illustrations are essential tools.'","section":"Section 1"},{"comment":"'Stable Diffusion series have demonstrated' should be 'Stable Diffusion series has demonstrated' to agree in number.","section":"Section 2"},{"comment":"The sentence 'The text should also be able to support and cover the elements that generate the illustration' is unclear and should be rephrased.","section":"Section 3.1"},{"comment":"The phrase 'GPT-4o-image alternative shows absolutely no difference' is confusing; it appears to mean that performance shows no significant difference across subjects, but the wording should be revised.","section":"Section 4.4"},{"comment":"The sentence 'Therefore, we use GPT-4o for automated scoring' appears in the middle of a paragraph about the human comparison; consider moving it to the evaluation methodology section for clarity.","section":"Section 4.2"},{"comment":"In the Image Judgement prompt, 'the first one by an anthropologist' should likely be 'the first one by a human expert.'","section":"Appendix A"},{"comment":"The paper does not provide a link to the dataset or code, which limits reproducibility; a public release plan should be included.","section":"General"},{"comment":"Reference [1] appears unrelated to the diffusion model introduction; please verify and correct.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's claim of being the 'first' scientific figure generation benchmark is factually undermined by its own reference [14] (ScImage). This is a serious novelty issue that must be resolved in revision, either by positioning SridBench as complementary or by demonstrating concrete differences. The evaluator validation is also thin; a more convincing agreement analysis is needed. The paper may be suitable for a benchmark-focused venue if these issues are addressed. The lack of a dataset/code release is an additional concern for reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: SridBench is a decent dataset-building effort with a plausible central finding, but the title claim to be the first scientific-figure-generation benchmark is contradicted by the paper's own reference list, and the automatic judge is not validated enough to support the reported numbers.\n\nWhat is genuinely new: the 1,120 caption-section-image triples across 13 disciplines, with six scoring dimensions, is a concrete resource the community doesn't have. The curation process, with human experts filtering for schematic rather than photographic or empirical figures, is sensible. The small human comparison on 100 instances gives some independent grounding: the paper claims GPT-4o scores are 'broadly in line' with human experts, and the qualitative observation that generated figures fail on text completeness and domain correctness (sun orbiting Earth, wrong reaction conditions) is believable and useful.\n\nWhere it gets soft. First, the novelty claim. Section 2 lumps ScImage [14] with understanding benchmarks, but the reference title is 'ScImage: How good are multimodal large language models at scientific text-to-image generation?' That is a generation benchmark. The abstract and conclusion repeat 'no benchmark currently exists' and 'the first benchmark.' That is an internal contradiction with the paper's own citation. Even if SridBench differs in scope, the authors need to say how, not claim the field is blank. Second, the judge: GPT-4o scores all 1,120 instances but was validated on only 100, and the paper gives no correlation or agreement statistic, just 'broadly in line.' The fact that GPT-4o also judges images from GPT-4o-image, the best-performing model, is a same-family bias worth addressing, even if the human subset suggests the direction is right. Third, the paper does not release the dataset or code, and reports no error bars or significance tests. Fourth, Emu-3 is listed as evaluated but was dropped from quantitative analysis; that leaves only two models.\n\nThe core finding, that modern models lag human experts on this task, is plausible and likely correct. The flaws are fixable but real: the novelty claim needs a factual correction, the judge validation needs numbers and ideally an independent judge, and the data should be released.\n\nWho this is for: people working on evaluation of multimodal generation, especially scientific figures. Worth a serious referee, but I would not accept it as is. My recommendation: send to peer review with a request for major revision, and make the authors position SridBench relative to ScImage. I would not cite it until the dataset and code are out.","headline":"Useful dataset and a plausible central finding, but the 'first benchmark' claim is undercut by the paper's own citation of ScImage, and the automatic judge is too thinly validated to support the reported numbers.","tokens_in":10685,"tokens_out":3372,"would_cite":false,"duration_ms":34153,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SridBench, the first benchmark for scientific figure generation, finds that even the top model, GPT-4o-image, remains at 'fair' level.","keywords":["scientific illustration generation","benchmark dataset","multimodal large language models","image generation evaluation","text-to-image","six-dimension scoring","GPT-4o-image","computer science and natural science"],"falsifier":"A fresh, independent human re-scoring of a random sample of, say, 200 instances from the full SridBench set that shows systematic discrepancies between GPT-4o's scores and human judgments—particularly on GPT-4o-image outputs—would show that the reported gap between models and humans is not reliably measured.","tokens_in":9749,"feed_emoji":"🎨","tokens_out":9097,"duration_ms":87966,"temperature":0.7,"pith_summary":"The paper introduces SridBench, a benchmark of 1,120 scientific-illustration drawing tasks drawn from published research across 13 natural-science and computer-science disciplines. Each task gives a model a diagram's caption and the surrounding section text from the original paper and asks it to redraw the figure. Six dimensions — completeness and accuracy of textual information, structural integrity, diagrammatic logic, cognitive readability, and aesthetic feeling — are scored from 1 to 5. The central finding is that even the strongest current model, GPT-4o-image, averages around 3 ('fair'), while other proprietary and open-source models score near 1 or lower. The authors argue this is the first benchmark of its kind and that the results show current models fall far short of human-level scientific drawing.","feed_headline":"Best AI model scores 'fair' on scientific illustration benchmark","feed_subtitle":"SridBench tests 1,120 drawing tasks across 13 disciplines; even top models fall far short of human experts.","key_machinery":"The carrying object is the triplet structure (reference image, caption, related section text) and the six-dimension scoring protocol. Each instance is a generation task: a model is given only the caption and section text and must produce a diagram. A multimodal judge (GPT-4o) then scores the output against the original figure on six 1–5 scales, a protocol the authors validate on a subset of 100 instances against human experts. The filter pipeline uses MLLMs to keep only concept, framework, flow, and structure diagrams, excluding photos, experimental-result graphs, and statistical plots.","core_discovery":"SridBench is positioned as the first evaluation resource specifically for generating scientific research illustrations from textual descriptions. The benchmark comprises 1,120 triplets of image, caption, and related section text, collected from authoritative peer-reviewed publications and filtered by human experts with MLLM assistance. The paper's main empirical claim is that no tested model is close to human-level: GPT-4o-image achieves roughly 'fair' scores across the six dimensions, whereas Gemini-2.0-Flash scores below 2 on every dimension, and Emu-3 fails to produce relevant content. The bottlenecks the authors identify are missing or inaccurate textual elements and scientific errors in the generated diagrams.","pith_inferences":["The validity of the headline result depends on the 100-instance human validation of GPT-4o as judge; extending that validation to a larger random sample would strengthen or weaken the reported rankings.","Using GPT-4o to grade images produced by GPT-4o-image, a closely related model, leaves open a same-family bias that the small human check may not fully detect.","The 'first benchmark' claim is contingent on the definition of scientific illustration; if one includes chart or graph generation benchmarks, the novelty is narrower, though this paper specifically targets schematic figures.","A concrete next step suggested by the paper's data is feeding models the full LaTeX source of the section instead of plain text, which might improve text completeness."],"forward_implications":["Any future image-generation model can be benchmarked on SridBench and compared on a common six-dimension rubric.","The observed gap between text accuracy and text completeness indicates that models tend to omit textual details rather than garble them, pointing to a specific technical target.","Open-source and non-specialist models scoring near 1 show that scientific illustration remains effectively unsolved for most systems.","The six dimensions could serve as a template for evaluating other forms of technical or instructional diagram generation."],"supporting_citations":[{"why":"The state-of-the-art image-generation model evaluated in the paper; its 'fair'-level scores are the main empirical result.","marker":"[12]"},{"why":"Prior benchmark for scientific figure interpretation, used to justify the claim that no generation-focused benchmark exists.","marker":"[13]"},{"why":"Prior work on scientific text-to-image generation, positioned as not being a systematic benchmark and therefore motivating SridBench.","marker":"[14]"},{"why":"Emu3, an autoregressive model evaluated as a baseline; its near-zero performance supports the finding that most models cannot do this task.","marker":"[8]"},{"why":"Gemini-based model evaluated in the study, scoring below 2 on all dimensions and demonstrating the gap.","marker":"[30]"},{"why":"Chain-of-thought prompting, cited as the reasoning capability that new models like GPT-4o-image integrate and that the paper argues is essential for scientific illustration.","marker":"[25]"}],"fun_headline_variants":["First benchmark for scientific figure generation: AI lags","SridBench: AI still can't draw science figures well","Even top AI models fail science illustration benchmark","New benchmark shows AI far from human-level science drawing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire evaluation assumes that GPT-4o's automatic scoring—validated against human experts on only 100 of the 1,120 instances—remains accurate for all 1,120 instances, including images produced by GPT-4o-image, a model from the same family.","fun_headline_variants_meta":{"raw":{"variants":["First benchmark for scientific figure generation: AI lags","SridBench: AI still can't draw science figures well","Even top AI models fail science illustration benchmark","New benchmark shows AI far from human-level science drawing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1304,"prompt_tokens":878,"completion_tokens":426,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":363}},"tokens_in":494,"tokens_out":426,"duration_ms":5239,"temperature":1.0,"reasoning_tokens":363,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:13:51.930182+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A fresh, independent human re-scoring of a random sample of, say, 200 instances from the full SridBench set that shows systematic discrepancies between GPT-4o's scores and human judgments—particularly on GPT-4o-image outputs—would show that the reported gap between models and humans is not reliably measured.","supporting_citations":[{"cited_title":"Addendum to gpt-4o system card: 4o image generation,","cited_arxiv_id":null,"evidence_quote":"The state-of-the-art image-generation model evaluated in the paper; its 'fair'-level scores are the main empirical result."},{"cited_title":"Chain-of- thought prompting elicits reasoning in large language models,","cited_arxiv_id":null,"evidence_quote":"Chain-of-thought prompting, cited as the reasoning capability that new models like GPT-4o-image integrate and that the paper argues is essential for scientific illustration."}],"review_version":1}