Pith. sign in

REVIEW 4 major objections 8 minor 3 cited by

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model

T0 review · 4 major / 8 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read SridBench, the first benchmark for scientific figure generation, finds that even the top model, GPT-4o-image, remains at 'fair' level.

desk verdict Useful dataset and a plausible central finding, but the 'first benchmark' claim is undercut by the paper's own citation of ScImage, and the automatic judge is too thinly validated to support the reported numbers. read the letter →

arxiv 2505.22126 v1 pith:RVJEKDA2 submitted 2025-05-28 cs.CV cs.AI

classification cs.CVcs.AI
keywords scientificillustrationgenerationbenchmarkdatasetmultimodallargelanguagemodelsimageevaluationtext-to-imagesix-dimensionscoringGPT-4o-imagecomputerscienceandnatural
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces SridBench, a benchmark of 1,120 scientific-illustration drawing tasks drawn from published research across 13 natural-science and computer-science disciplines. Each task gives a model a diagram's caption and the surrounding section text from the original paper and asks it to redraw the figure. Six dimensions — completeness and accuracy of textual information, structural integrity, diagrammatic logic, cognitive readability, and aesthetic feeling — are scored from 1 to 5. The central finding is that even the strongest current model, GPT-4o-image, averages around 3 ('fair'), while other proprietary and open-source models score near 1 or lower. The authors argue this is the first benchmark of its kind and that the results show current models fall far short of human-level scientific drawing.

What carries the argument

The carrying object is the triplet structure (reference image, caption, related section text) and the six-dimension scoring protocol. Each instance is a generation task: a model is given only the caption and section text and must produce a diagram. A multimodal judge (GPT-4o) then scores the output against the original figure on six 1–5 scales, a protocol the authors validate on a subset of 100 instances against human experts. The filter pipeline uses MLLMs to keep only concept, framework, flow, and structure diagrams, excluding photos, experimental-result graphs, and statistical plots.

What would settle it

A fresh, independent human re-scoring of a random sample of, say, 200 instances from the full SridBench set that shows systematic discrepancies between GPT-4o's scores and human judgments—particularly on GPT-4o-image outputs—would show that the reported gap between models and humans is not reliably measured.

Watch

Extended reading notes

Core claim

SridBench is positioned as the first evaluation resource specifically for generating scientific research illustrations from textual descriptions. The benchmark comprises 1,120 triplets of image, caption, and related section text, collected from authoritative peer-reviewed publications and filtered by human experts with MLLM assistance. The paper's main empirical claim is that no tested model is close to human-level: GPT-4o-image achieves roughly 'fair' scores across the six dimensions, whereas Gemini-2.0-Flash scores below 2 on every dimension, and Emu-3 fails to produce relevant content. The bottlenecks the authors identify are missing or inaccurate textual elements and scientific errors in the generated diagrams.

Load-bearing premise

The entire evaluation assumes that GPT-4o's automatic scoring—validated against human experts on only 100 of the 1,120 instances—remains accurate for all 1,120 instances, including images produced by GPT-4o-image, a model from the same family.

Editorial extensions

If this is right

  • Any future image-generation model can be benchmarked on SridBench and compared on a common six-dimension rubric.
  • The observed gap between text accuracy and text completeness indicates that models tend to omit textual details rather than garble them, pointing to a specific technical target.
  • Open-source and non-specialist models scoring near 1 show that scientific illustration remains effectively unsolved for most systems.
  • The six dimensions could serve as a template for evaluating other forms of technical or instructional diagram generation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The validity of the headline result depends on the 100-instance human validation of GPT-4o as judge; extending that validation to a larger random sample would strengthen or weaken the reported rankings.
  • Using GPT-4o to grade images produced by GPT-4o-image, a closely related model, leaves open a same-family bias that the small human check may not fully detect.
  • The 'first benchmark' claim is contingent on the definition of scientific illustration; if one includes chart or graph generation benchmarks, the novelty is narrower, though this paper specifically targets schematic figures.
  • A concrete next step suggested by the paper's data is feeding models the full LaTeX source of the section instead of plain text, which might improve text completeness.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 8 minor

Summary. The paper introduces SridBench, a benchmark of 1,120 instances for scientific research illustration generation, spanning 13 disciplines in natural and computer science. Each instance is a triple of (figure, caption, related section), collected from arXiv and Nature and screened by human experts and MLLMs. The authors propose a six-dimension evaluation protocol (completeness and accuracy of textual information, diagrammatic structural integrity, diagrammatic logic, cognitive readability, aesthetic feeling) and report results for GPT-4o-image, Gemini-2.0-Flash, and Emu-3, with GPT-4o serving as an automated judge. The central findings are that GPT-4o-image scores around 3/5 ('fair') and falls far short of human experts, while Gemini-2.0-Flash scores below 2 and Emu-3 is qualitatively unusable. A 100-instance human comparison is used to justify the use of GPT-4o as the automated judge.

Significance. If the benchmark is properly constructed and the evaluation method is reliable, SridBench would be a useful resource for tracking progress in a practically important task, and the six-dimension rubric is a reasonable starting point. The human comparison on 100 instances is a positive step toward grounding the automated judge. However, the significance is conditional on resolving the novelty claim (the paper's own reference [14] ScImage appears to be a scientific text-to-image generation benchmark) and on providing stronger evidence that the GPT-4o judge generalizes across the full dataset. The empirical finding that GPT-4o-image is only at a 'fair' level is plausible but currently lacks error bars and significance tests.

major comments (4)
  1. [Abstract, Sections 1 and 2 (Research Gaps), and Conclusion] The paper repeatedly claims that 'no benchmark currently exists' and that SridBench is 'the first benchmark' for scientific figure generation. This is contradicted by the manuscript's own reference [14], titled 'ScImage: How good are multimodal large language models at scientific text-to-image generation?' which, by its title, is exactly a generation benchmark for scientific figures. The text in Section 2, however, describes ScImage as focused on 'understanding capabilities.' This is an internal inconsistency. The authors must position SridBench relative to ScImage, compare dataset construction, metrics, and scope, and clearly state what is genuinely novel. As written, the headline novelty claim is unsupported.
  2. [Section 4.2, Figure 3(b)] The automated judge (GPT-4o) is validated on only 100 of the 1,120 instances (50 natural science, 50 computer science). The paper states that 'GPT-4o scores are broadly in line with those of human experts' without reporting any quantitative agreement metric (e.g., correlation, mean absolute error, or per-dimension breakdown). Furthermore, GPT-4o is used to judge images generated by GPT-4o-image, a model from the same family. Since the central quantitative results—GPT-4o-image scoring around 3/5 and the human-model gap—are derived entirely from this judge, the paper must either provide a more extensive validation (e.g., on a larger and more diverse sample) or report confidence intervals and agreement statistics. Without this, the size of the human-model gap is not established.
  3. [Sections 4.2–4.5, Figures 3–6] The reported scores are averages without error bars, variance, or significance tests. For example, Figure 3(a) shows average scores per dimension per model, but there is no indication of the variability across the 1,120 instances, the number of generation runs, or whether the differences between models and human experts are statistically significant. The headline claim that 'GPT-4o-image falls far short of human-level performance' rests on these averages. The authors should report standard deviations, confidence intervals, or perform appropriate statistical tests to support the strength of the conclusions.
  4. [Section 4.1] Only two models are fully evaluated quantitatively (GPT-4o-image and Gemini-2.0-Flash); Emu-3 is excluded from the quantitative analysis due to generation time. This is a narrow model set for a 'comprehensive empirical study' as claimed in the contributions. At minimum, the paper should state this limitation explicitly and discuss how the findings might generalize to other model families.
minor comments (8)
  1. [Section 1] 'Scientific illustration are essential tools' should be 'Scientific illustrations are essential tools.'
  2. [Section 2] 'Stable Diffusion series have demonstrated' should be 'Stable Diffusion series has demonstrated' to agree in number.
  3. [Section 3.1] The sentence 'The text should also be able to support and cover the elements that generate the illustration' is unclear and should be rephrased.
  4. [Section 4.4] The phrase 'GPT-4o-image alternative shows absolutely no difference' is confusing; it appears to mean that performance shows no significant difference across subjects, but the wording should be revised.
  5. [Section 4.2] The sentence 'Therefore, we use GPT-4o for automated scoring' appears in the middle of a paragraph about the human comparison; consider moving it to the evaluation methodology section for clarity.
  6. [Appendix A] In the Image Judgement prompt, 'the first one by an anthropologist' should likely be 'the first one by a human expert.'
  7. [General] The paper does not provide a link to the dataset or code, which limits reproducibility; a public release plan should be included.
  8. [References] Reference [1] appears unrelated to the diffusion model introduction; please verify and correct.

Circularity Check

1 steps flagged · score 4.0 of 10

The paper's central 'first benchmark / blank field' claim is manufactured by reclassifying its own cited ScImage [14]—a scientific text-to-image generation benchmark—as an 'understanding' benchmark; the empirical scoring chain is not circular.

  1. renaming known result [Section 1 (Introduction), Section 2 (Related Work, Research Gaps), and References [14]; echoed in Abstract and Section 5]
    "most of the current research efforts are focused on the understanding of scientific images and the generation of image captions (such as SciFIBench [13], FigCaps-HF [28], etc. [29]). The field of evaluating the generation of scientific research drawings is almost blank. ... mainly focused on benchmarking the understanding capabilities of multimodal models (e.g., SciFIBench [13], ScImage [14]). ... [14] L. Zhang, ... "ScImage: How good are multimodal large language models at scientific text-to-image generation?""

    The paper's load-bearing novelty claim is that the field of evaluating scientific figure generation is 'almost blank' and that SridBench is 'the first benchmark.' That claim is obtained by assigning ScImage [14] to 'understanding capabilities' in the Introduction, while the paper's own bibliography titles [14] as a benchmark for 'scientific text-to-image generation.' If ScImage is a generation benchmark, the field is not blank and SridBench is not first; the central novelty step reduces to a reclassification of the cited prior work. The empirical evaluation chain—human validation on 100 samples followed by GPT-4o scoring—is not itself circular, so the overall circularity is partial rather than total.

full rationale

Most of the paper's derivation chain is self-contained: the 1,120 triples are collected from papers using an MLLM filter plus human-expert screening, the generation prompts are specified, and the automated evaluation uses a GPT-4o scorer whose agreement with human experts is checked on a 100-instance subset. That is a validation/generalization step rather than a construction-level circularity. The same-family overlap (GPT-4o judging GPT-4o-image) is a legitimate validity risk, but it does not make the reported scores equal to the inputs by construction, and the human subset provides independent grounding for the broad conclusion that models lag humans. The circular/renaming content is concentrated in the novelty argument: ScImage [14], a scientific text-to-image generation benchmark, is categorized as an 'understanding' benchmark so that the paper can declare the field blank and itself 'the first benchmark.' This reclassification drives the score of 4 rather than a 0-2 non-finding.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The benchmark introduces no fitted parameters or new physical entities. Its claims rest on domain assumptions about what counts as a scientific illustration, how generation should be prompted, and whether an LLM judge can stand in for human experts; these are stated in the prompts and Section 4.2 rather than independently established.

assumptions (4)
  • domain assumption Filtering to four schematic types (concept, model frame, process flow, structure) defines the scope of 'scientific illustration'.
    The image filter prompt in Appendix A restricts accepted figures to these types and excludes photos, statistical graphs, tables, formulas, and pseudocode; this shapes the benchmark's claims about scientific figure generation.
  • domain assumption GPT-4o's scores are a valid proxy for human expert evaluation of the six dimensions.
    Section 4.2 validates GPT-4o against human experts on 100 instances and reports broad alignment, then applies GPT-4o to all 1,120 instances; the paper also notes a slight overrating in completeness and accuracy.
  • domain assumption Reference figures from high-citation papers in top venues are suitable ground truth for generation quality.
    Section 3.1 filters by citation count, venue, and human expert selection; this assumes expert-published figures represent the standard for scientific illustration.
  • domain assumption Generating a figure from caption plus surrounding section text is a faithful operationalization of scientific illustration drawing.
    The generation prompt in Appendix A provides only section and caption, treating this reconstruction task as representative of how researchers create figures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model." pith.science (2026). https://pith.science/paper/RVJEKDA2

@misc{pith2026250522126,
  author       = {Pith},
  title        = {Pith review of: SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RVJEKDA2}},
  note         = {Machine review of arXiv:2505.22126}
}
read the original abstract

Recent years have seen rapid advances in AI-driven image generation. Early diffusion models emphasized perceptual quality, while newer multimodal models like GPT-4o-image integrate high-level reasoning, improving semantic understanding and structural composition. Scientific illustration generation exemplifies this evolution: unlike general image synthesis, it demands accurate interpretation of technical content and transformation of abstract ideas into clear, standardized visuals. This task is significantly more knowledge-intensive and laborious, often requiring hours of manual work and specialized tools. Automating it in a controllable, intelligent manner would provide substantial practical value. Yet, no benchmark currently exists to evaluate AI on this front. To fill this gap, we introduce SridBench, the first benchmark for scientific figure generation. It comprises 1,120 instances curated from leading scientific papers across 13 natural and computer science disciplines, collected via human experts and MLLMs. Each sample is evaluated along six dimensions, including semantic fidelity and structural accuracy. Experimental results reveal that even top-tier models like GPT-4o-image lag behind human performance, with common issues in text/visual clarity and scientific correctness. These findings highlight the need for more advanced reasoning-driven visual generation capabilities.

Figures

Figures reproduced from arXiv: 2505.22126 by the authors.

Figure 1
Figure 1. General description of SridBench. We collected triple data from 13 directions in natural [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The framework of our Benchmark of Scientific Research Illustration Drawing of Image [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. (a). On the computer science and natural science data, the average score of GPT-4o-image [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: On different subjects of natural science data, the average score of GPT-4o-image and [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: On different subjects of computer science data, the average score of GPT-4o-image and [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: On different types of computer science data, the average score of GPT-4o-image and [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Computer science paper illustrations generated by different image generation models under [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Natural science paper illustrations generated by different image generation models under [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Illustrations generated by GPT-4o-image (left) and their reference from original paper [PITH_FULL_IMAGE:figures/full_fig_p009_9.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Scientific-figure editing can be learned from arXiv revision pairs: a skill-evolving SVG agent follows edit instructions and transfers learned skills across LLM backbones.

  2. SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A structure-first benchmark finds that text-to-image models preserve little recoverable graph structure in scientific diagrams, with the best model scoring 0.116 graph-level versus 0.01-0.09 for others.

  3. AI4Research: A Survey of Artificial Intelligence for Scientific Research

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey that organizes AI-for-research work into five tasks, comprehension, survey, discovery, writing, and peer review, and compiles associated tools and benchmarks.

Reference graph

Works this paper leans on

30 extracted references · 11 canonical work pages · cited by 3 Pith papers

  1. [14]

    Scimage: How good are multimodal large language models at scientific text-to-image generation?

    L. Zhang, S. Eger, Y . Cheng, W. Zhai, J. Belouadi, C. Leiter, S. P. Ponzetto, F. Moafian, and Z. Zhao, “Scimage: How good are multimodal large language models at scientific text-to-image generation?” 2024. [Online]. Available: https://arxiv.org/abs/2412.02368

  2. [1]

    DSM Refinement with Deep Encoder-Decoder Networks

    N. Metzger, “Dsm refinement with deep encoder-decoder networks,” 2020. [Online]. Available: https://arxiv.org/abs/2012.07427

  3. [2]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” 2020. [Online]. Available: https://arxiv.org/abs/2006.11239

  4. [3]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” inCVPR, 2022, pp. 10 684–10 695

  5. [4]

    Scaling rectified flow transformers for high-resolution image synthesis,

    P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y . Levi, D. Lorenz, A. Sauer, F. Boeselet al., “Scaling rectified flow transformers for high-resolution image synthesis,” in Forty-first International Conference on Machine Learning, 2024

  6. [5]

    Hierarchical text-conditional image generation with clip latents,

    A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, “Hierarchical text-conditional image generation with clip latents,”arXiv preprint arXiv:2204.06125, vol. 1, no. 2, p. 3, 2022

  7. [6]

    Improving image generation with better captions,

    J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y . Guo et al., “Improving image generation with better captions,”Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, vol. 2, no. 3, p. 8, 2023

  8. [7]

    Black Forest Labs, “Flux,” https://github.com/black-forest-labs/flux, 2024, accessed: 2024-11- 05

Show all 30 references
  1. [8]

    Emu3: Next-token prediction is all you need,

    X. Wang, X. Zhang, Z. Luo, Q. Sun, Y . Cui, J. Wang, F. Zhang, Y . Wang, Z. Li, Q. Yuet al., “Emu3: Next-token prediction is all you need,”arXiv preprint arXiv:2409.18869, 2024

  2. [9]

    Neighboring autoregressive modeling for efficient visual generation,

    Y . He, Y . He, S. He, F. Chen, H. Zhou, K. Zhang, and B. Zhuang, “Neighboring autoregressive modeling for efficient visual generation,” 2025. [Online]. Available: https://arxiv.org/abs/2503.10696

  3. [10]

    Visual autoregressive modeling: Scalable image generation via next-scale prediction,

    K. Tian, Y . Jiang, Z. Yuan, B. Peng, and L. Wang, “Visual autoregressive modeling: Scalable image generation via next-scale prediction,”Advances in neural information processing systems, vol. 37, pp. 84 839–84 865, 2024

  4. [11]

    Janus-pro: Uni- fied multimodal understanding and generation with data and model scaling,

    X. Chen, Z. Wu, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, and C. Ruan, “Janus-pro: Uni- fied multimodal understanding and generation with data and model scaling,”arXiv preprint arXiv:2501.17811, 2025

  5. [12]

    Addendum to gpt-4o system card: 4o image generation,

    OpenAI, “Addendum to gpt-4o system card: 4o image generation,” 2025, accessed: 2025-04-02. [Online]. Available: https://openai.com/index/ gpt-4o-image-generation-system-card-addendum/

  6. [13]

    Scifibench: Benchmarking large multimodal models for scientific figure interpretation,

    J. Roberts, K. Han, N. Houlsby, and S. Albanie, “Scifibench: Benchmarking large multimodal models for scientific figure interpretation,” 2024. [Online]. Available: https://arxiv.org/abs/2405.08807

  7. [15]

    Sdxl: Improving latent diffusion models for high-resolution image synthesis,

    D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rom- bach, “Sdxl: Improving latent diffusion models for high-resolution image synthesis,”arXiv preprint arXiv:2307.01952, 2023

  8. [16]

    Mlperf inference benchmark,

    V . J. Reddi, C. Cheng, D. Kanter, P. Mattson, G. Schmuelling, C.-J. Wu, B. Anderson, M. Breughe, M. Charlebois, W. Chou, R. Chukka, C. Coleman, S. Davis, P. Deng, G. Diamos, J. Duke, D. Fick, J. S. Gardner, I. Hubara, S. Idgunji, T. B. Jablin, J. Jiao, T. S. John, P. Kanwar, ...

  9. [17]

    Designbench: Exploring and benchmarking dall- e 3 for imagining visual design,

    K. Lin, Z. Yang, L. Li, J. Wang, and L. Wang, “Designbench: Exploring and benchmarking dall- e 3 for imagining visual design,” 2023. [Online]. Available: https://arxiv.org/abs/2310.15144

  10. [18]

    Vila-u: a unified foundation model integrating visual understanding and generation,

    Y . Wu, Z. Zhang, J. Chen, H. Tang, D. Li, Y . Fang, L. Zhu, E. Xie, H. Yin, L. Yiet al., “Vila-u: a unified foundation model integrating visual understanding and generation,”arXiv preprint arXiv:2409.04429, 2024

  11. [19]

    Generative multimodal models are in-context learners,

    Q. Sun, Y . Cui, X. Zhang, F. Zhang, Q. Yu, Z. Luo, Y . Wang, Y . Rao, J. Liu, T. Huang, and X. Wang, “Generative multimodal models are in-context learners,” 2023

  12. [20]

    T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,

    K. Huang, K. Sun, E. Xie, Z. Li, and X. Liu, “T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,”Advances in Neural Information Processing Systems, vol. 36, pp. 78 723–78 747, 2023

  13. [21]

    Qwen2-vl: En- hancing vision-language model’s perception of the world at any resolution,

    P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, Y . Fan, K. Dang, M. Du, X. Ren, R. Men, D. Liu, C. Zhou, J. Zhou, and J. Lin, “Qwen2-vl: En- hancing vision-language model’s perception of the world at any resolution,”arXiv preprint arXiv:2409...

  14. [22]

    Qwen2.5-vl technical report,

    S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y . Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y . Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin, “Qwen2.5-vl technical report,”arXiv prepri...

  15. [23]

    Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

    Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Luet al., “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 202...

  16. [24]

    Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling,

    Z. Chen, W. Wang, Y . Cao, Y . Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liuet al., “Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling,”arXiv preprint arXiv:2412.05271, 2024

  17. [25]

    Chain-of- thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhouet al., “Chain-of- thought prompting elicits reasoning in large language models,”NeurIPS, 2022

  18. [26]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

    D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Biet al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,”arXiv preprint arXiv:2501.12948, 2025

  19. [27]

    QwQ: Reflect Deeply on the Boundaries of the Unknown,

    Q Team, “QwQ: Reflect Deeply on the Boundaries of the Unknown,” Nov. 2024, accessed: 2025-01-01. [Online]. Available: https://qwenlm.github.io/blog/qwq-32b-preview/

  20. [28]

    Figcaps-hf: A figure-to-caption generative framework and benchmark with human feedback,

    A. Singh, P. Agarwal, Z. Huang, A. Singh, T. Yu, S. Kim, V . Bursztyn, N. Vlassis, and R. A. Rossi, “Figcaps-hf: A figure-to-caption generative framework and benchmark with human feedback,” 2023. [Online]. Available: https://arxiv.org/abs/2307.10867

  21. [29]

    Chartbench: A benchmark for complex visual reasoning in charts,

    Z. Xu, S. Du, Y . Qi, C. Xu, C. Yuan, and J. Guo, “Chartbench: A benchmark for complex visual reasoning in charts,” 2024. [Online]. Available: https://arxiv.org/abs/2312.15915

  22. [30]

    Gemini: a family of highly capable multimodal models,

    G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millicanet al., “Gemini: a family of highly capable multimodal models,”arXiv preprint arXiv:2312.11805, 2023. 12 A Prompts used in data process and judgement Image FilterPlea...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.