Pith. sign in

Paper Citation Record · LEDGER

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

As of 20 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2412.02368.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02368 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:36:24.852757Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:19:03.335574Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:54:20.682844Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact4
  • verified fuzzy16
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c5372279-12b2-4b0a-95e3-63ce7951c627 · outbound

This paper cites Automatikz: Text-guided synthesis of scientific vector graphics with tikz.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Automatikz: Text-guided synthesis of scientific vector graphics with tikz

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.602380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.582070Z digest=sha256:d2871914a8d7a3fa411e58e354c611d918691153b125c52aae3e07dc5f07c7b9

Observation 661d87b5-b107-46a6-8f3e-344ef429b8df · outbound

This paper cites DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.586355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.586355Z digest=sha256:e5532a84c8c81e5543535702a8743bc04ef5a3934b97de4a856ac93cbb8d4eae

Observation c3b372f0-0d0a-4809-bdd9-20ef8b921b47 · outbound

This paper cites Scene text visual question answering.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Scene text visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.594425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.589521Z digest=sha256:cf1f5df4a706120aa52f966b52e982e147d1a439464d5d20619f56298e058042

Observation b8e1369e-e701-4757-b420-93dc7a76d9da · outbound

This paper cites Elicit: Language models as research tools.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Elicit: Language models as research tools

Reference 4

Resolution
verified exact
doi, observed 2026-08-11T23:36:24.958657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.591998Z digest=sha256:8eb420f5db712423d4fd3ca5f214b8eafb945239d0eab74606fcdc9415bfe4f6

Observation c122e92b-403c-4fd4-9363-a2db1914d52b · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.594724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.594724Z digest=sha256:1174cfe8fd4eeff3b30b5334f38149816f0b164e183158c31f59b051828847c8

Observation ef81c52e-34e9-40e8-87fb-e3d08db2bbe1 · outbound

This paper cites Transformers go for the LOL s: Generating (humourous) titles from scientific abstracts end-to-end.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Transformers go for the LOL s: Generating (humourous) titles from scientific abstracts end-to-end

Reference 6

Resolution
verified exact
doi, observed 2026-08-11T23:36:24.951208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.597894Z digest=sha256:ca6c0c69a9ef485eb8aed186dae272f763a98e31f300b0cc9708330fe940586f

Observation 70b5358e-5c68-488a-acfa-67c1e082bdc4 · outbound

This paper cites Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.586914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.600980Z digest=sha256:b2f6e4fa87a0f95388b558394698dc6d72ed5d49bf4b3b2959831bddb7341c07

Observation 0a8ae3c3-3b0d-4f63-a50a-4f3a9c678043 · outbound

This paper cites Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.603625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.603625Z digest=sha256:30b389272b522de431032af3171552d7460f3ecada427fa6cef4d5e543f8cb35

Observation bb9af1da-ee0c-4a7f-90fe-2960baab0669 · outbound

This paper cites Exams-v: A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Exams-v: A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.574864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.608636Z digest=sha256:ff5098151358a1a2aed7b45cf2f57931823505f3cc5202a7af81bb0c7f2dfbb6

Observation 07437ee4-269f-4035-b987-8e24aa871502 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Imagenet: A large-scale hierarchical image database

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.611608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.611608Z digest=sha256:98caeff247084844b70395f59b11bb742044d7e5729976a44fdc5df0b20d70f6

Observation 0ed76c6a-7ea1-43c3-8341-31fa89ae0b0e · outbound

This paper cites The mnist database of handwritten digit images for machine learning research [best of the web].

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? The mnist database of handwritten digit images for machine learning research [best of the web]

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.614395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.614395Z digest=sha256:15f42f45ed1922598f3a1a1b44d8aafa20f913cad36a104a40fe0867ea1bd870

Observation 7945bba2-7c8f-43b0-b9c0-8be88271c234 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.617536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.617536Z digest=sha256:37da4802205cea1f2e7535f7ec322b5f7b94455feb0c02e9a822a18f700758d5

Observation f787e6ca-3b28-449a-9e3d-edbaf58ea912 · outbound

This paper cites CLIPS core: A reference-free evaluation metric for image captioning.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? CLIPS core: A reference-free evaluation metric for image captioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.620687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.620687Z digest=sha256:79b4ae2e2409b31ccef650b3e24c8dde99f852ea8684523c7055be8b2a5bfa5e

Observation a3b40108-ae91-4a10-8f24-37c94459b2f7 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.623681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.623681Z digest=sha256:4733245ad9061848955c0eb5ec535b8def4ef57ef1403d98cf1cf54609b37da4

Observation afe80c1e-7d9e-4037-bd1b-ae7277da6592 · outbound

This paper cites Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.626213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.626213Z digest=sha256:c786e354a7c755dee13cb4f164140a6301c0df1c199a0f76ac9437f5576a7dd8

Observation 261e4018-77df-49e3-ac78-748fb089be3a · outbound

This paper cites Learning multiple layers of features from tiny images.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Learning multiple layers of features from tiny images

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.629335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.629335Z digest=sha256:99b408ad44569be6700716695c5d7651ae5565692807f66c4a37297be333ac37

Observation 56486e10-e646-4248-8dfb-7633f3915b8d · outbound

This paper cites West, and Bill Howe.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? West, and Bill Howe

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.557982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.706202Z digest=sha256:7816368e35afc7180e97dfee57616267a5de9f9e8819a59fb87ba244879c2d60

Observation c742c38a-ca43-4ae6-b5cd-3cef0c8e6df2 · outbound

This paper cites Holistic evaluation of text-to-image models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Holistic evaluation of text-to-image models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.549898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.709142Z digest=sha256:9ba51275274ee0e8246c67a1b84743e2e5d33f0c11925391995a84e7c4c507a2

Observation 4065fcbb-6140-4768-80a3-d1444fff5df0 · outbound

This paper cites PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.712170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.712170Z digest=sha256:bfbeb4f68a1c2cf094d989de94f6e42cefc20106e7a5c2a38f8db5fff519a6a0

Observation f5c09fb1-54bf-423b-9565-843611a298aa · outbound

This paper cites The E val4 NLP 2023 shared task on prompting large language models as explainable metrics.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? The E val4 NLP 2023 shared task on prompting large language models as explainable metrics

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.715531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.715531Z digest=sha256:d1ec0b80dbfad4ccb9488f1b523dd41a0e1843fa35cac6522ce6783b82b66bb3

Observation 601ddb87-e901-45fb-aa69-054dc81ebdb7 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.718590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.718590Z digest=sha256:5fd09b5f35582129c3292a7e0bfe702f14b2932601c53cb2bd7ba8d977103bf9

Observation 4eda8e9c-c886-4254-a71d-24491027ff13 · outbound

This paper cites MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.722836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.722836Z digest=sha256:77eb03bb799e98eadc0a81d311988f4a416a70cdc2da06b3af0fb1fb3e2bc438

Observation d7f6339b-00bf-495b-8543-e0446f83a03a · outbound

This paper cites an unresolved cited work.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:36:25.541809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.725965Z digest=sha256:85ee32d534b3db976fe531e87982d3b20c41297342d59b6254678336dfa03bb4

Observation 6c95bd47-95a3-4194-87cb-380e7dec5d08 · outbound

This paper cites Microsoft coco: Common objects in context.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Microsoft coco: Common objects in context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.728607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.728607Z digest=sha256:eda062fde56767d2e0ef59543bae68bde141ab6149e5d9c7b09a676ee9cd3a61

Observation 62b7b0d7-f757-4fcd-90b3-606beac84412 · outbound

This paper cites Figurefirst: A layout-first approach for scientific figures.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Figurefirst: A layout-first approach for scientific figures

Reference 25

Resolution
verified exact
doi, observed 2026-08-11T23:36:24.932781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.731346Z digest=sha256:8fa7225f4ffdccbcce9a31b95d3fa94d8841eb30a7f255afe401c280d2fa3dcf

Observation 69122b81-1fe2-49af-85bd-5b64ed6f81de · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.734072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.734072Z digest=sha256:b6234691326c4bade7703e248e63dc6ba79ec581ee421c47ad8ae301ce642fde

Observation 8faf7e59-5c32-495b-abf5-dd27bb19d193 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.527485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.736948Z digest=sha256:0101fe031694eed4dfd9afbd2ae22166851396c68d9be21b00e64fd71a54b180

Observation 7e9585a5-dfb1-4cf9-82fa-09a236012bc5 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.518354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.739695Z digest=sha256:67359d6b73905cdb5bbbf50fea62a3d90cae712bad20465ff094ff8e7e8e62dc

Observation eb18100c-eed6-430b-a3fd-c97b21c86cb0 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.742162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.742162Z digest=sha256:5469ab56b454a15ac12d2044668f31ead2d69b09186c474815500e45703512d5

Observation e1125bea-1dd7-4c59-a8c5-d20b99b7276f · outbound

This paper cites State of What Art? A Call for Multi-Prompt LLM Evaluation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.744889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.744889Z digest=sha256:b572be8e23e06cad0c2339a11ca75c736c548b03818494ccec049a0caeef3297

Observation c32e0fe4-2bdf-4e5c-b6cf-6c2597bb6c0f · outbound

This paper cites Evaluating large language models for structured science summarization in the open research knowledge graph.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Evaluating large language models for structured science summarization in the open research knowledge graph

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.747047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.747047Z digest=sha256:f9aa42526991ef3f237e877c565e54bcd0a36ea7cf79d273f64f9b2f7326233f

Observation f8fd8537-aa7e-47e4-8735-70dad3f503c5 · outbound

This paper cites Introducing openai o1 (preview).

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Introducing openai o1 (preview)

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.509261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.749337Z digest=sha256:ad5605f85ab5fd029b2b203e06487b991ea4298d5298acdafb1dba96a02f124a

Observation 93aa862c-c758-4532-8ab0-5fb96977256b · outbound

This paper cites Zero-Shot Text-to-Image Generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Zero-Shot Text-to-Image Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.751632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.751632Z digest=sha256:f5cc8602d7c9fd74d407b44a5c9a92ed712c9d224aa42b87c344d4c694b7442f

Observation 3ce14d11-9b2e-4f13-96b3-afe3728de606 · outbound

This paper cites SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.754369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.754369Z digest=sha256:bda19fbef10f817b93e712ff61fb46e1565ad56f947a0edf46643965474c7967

Observation f4c579dd-bdd0-4f06-a76b-dcbcce905dd6 · outbound

This paper cites Ocr-vqgan: Taming text-within-image generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Ocr-vqgan: Taming text-within-image generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.500047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.757290Z digest=sha256:cfa34701cd93348c42b73e09062a0f93871e4009495fdc7d06cf741bb783ca20

Observation a53ccde3-da8e-44e1-bcad-a2c239059785 · outbound

This paper cites Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2).

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.759801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.759801Z digest=sha256:25733452da3840c80a8a95ff61783311c82674f7b5d6e61c44ddc31dba115044

Observation a1c67c02-191b-4eec-ac8a-825d45a73b44 · outbound

This paper cites Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.762652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.762652Z digest=sha256:841680e4a041a7b933e85e23ea578697e7da6964652980ca98f855588b216903

Observation 8339022e-882e-42c5-8e43-30b00e4c53fd · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.765471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.765471Z digest=sha256:0b48492ac21b5066ed7fbb267318ed03cb9f5e29360121a039e617390df54ecc

Observation 8c211054-9ab3-4b9d-a2da-813c14d3d249 · outbound

This paper cites ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.768284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.768284Z digest=sha256:be1e8489b920a9d8504cdeba8eefb2167f8638739810a59b7bc199a429456e7c

Observation 92c77b8d-fda4-4811-ae56-944a4fcbd4ee · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.771267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.771267Z digest=sha256:73ddffed04a9b38d2b3c2a89b7a499a4e5e56f3c7abe56d6351f13285b749cd7

Observation 6a9f03a6-efa5-4f09-8288-b1a2a5634fc3 · outbound

This paper cites Towards vqa models that can read.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Towards vqa models that can read

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.488831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.774247Z digest=sha256:5f4f2b05cfbe7466ec362247284d6129e1322ea47c5b6e02620b3539c816a653

Observation 652b5bc1-f02f-4ac5-8b8b-675c166aaf60 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Gemma: Open Models Based on Gemini Research and Technology

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.776853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.776853Z digest=sha256:8869c496f7eb6a55dbce4191392e44dae02a1faa1ad7857abf107ab0c2f94317

Observation e10d8a00-7504-4c61-ae91-60687ef018cc · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.780407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.780407Z digest=sha256:7266593c17c1636e66b94b32f95c728ea631f077511909d01f60360b69cc587f

Observation 1e37139c-7570-4237-af61-b7157016c499 · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.783054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.783054Z digest=sha256:eea70e870a79103104422f4c4ab248dff38b903397f1823e7d1a03cab5afc5d8

Observation 47dad0a8-ded6-4426-b2ae-bfdabd94f879 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.786161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.786161Z digest=sha256:6af7309a5d89979b249d1c3f79fb6481434c0dc14ecccd5c1d56c7d3b28bd62d

Observation a39e7aa2-91d3-4b09-a154-17c5a68cdf70 · outbound

This paper cites Plots made quickly: An efficient approach for generating visualizations from natural language queries.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Plots made quickly: An efficient approach for generating visualizations from natural language queries

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.474081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.789043Z digest=sha256:dd52fda14702b278129b115bb22f974821c5f787c4edf0c8f8affbc6b250b161

Observation 1859436c-9ba3-4901-92fc-0cacfaa5973f · outbound

This paper cites Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.465937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.791619Z digest=sha256:015c3ac2c758f0b7a239302bed3836b5b1ada08c0b9700ba84c8d5c21604aca5

Observation 34ec3669-99c7-447b-9a7c-3238a9de9f3f · outbound

This paper cites Scibench: Evaluating college-level scientific problem-solving abilities of large language models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Scibench: Evaluating college-level scientific problem-solving abilities of large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.457618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.794196Z digest=sha256:c99250f63377472e688bcc90509bec4b3c338a44ba07784ae221ace8c5c19f1b

Observation f031aa26-6506-42b8-babb-7d68b49cb617 · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.796788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.796788Z digest=sha256:f352defd50f87599ae891370015880a103a2b9e56fb6d423919f57594b24b078

Observation 280d9344-caab-4daf-958c-6f8f7b06dd35 · outbound

This paper cites seaborn: statistical data visualization.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? seaborn: statistical data visualization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.799390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.799390Z digest=sha256:3bedc6119d9f29107526a296b287dfe17d45850d28e630a58e8b906a0ea628bf

Observation 96bdf0c0-9074-46f8-87a8-f9e5f2768fa0 · outbound

This paper cites Evaluating and analyzing relationship hallucinations in large vision-language models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Evaluating and analyzing relationship hallucinations in large vision-language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.449624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.802121Z digest=sha256:c018412203101caf4547ca6ba1f279d0c8d7c884992324427ba0e284d27bdae2

Observation 804bfb6d-e949-4db8-9e41-b897d6bfa0fc · outbound

This paper cites An automatic graph generation method for scholarly papers based on table structure analysis.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? An automatic graph generation method for scholarly papers based on table structure analysis

Reference 52

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T23:36:25.142576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.804221Z digest=sha256:13aeb40160bc7163e0a4bd1b8bb44fe44f3b1d1a2aa3ca474b24254f1ba7908a

Observation 56badd52-6031-4f40-be3d-4bd11ee4d2ff · outbound

This paper cites Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.441426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.806384Z digest=sha256:5243d1c55645d4f514afadcfeb853c9d7e36319c16dfccfe9115608faca26d9c

Observation 330f4ae8-8e4a-49b0-af2b-331f6281e39f · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.808524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.808524Z digest=sha256:fe4af2a14957ec43c139c9abf6c974cc7f7849f991e21ba7ed51155e8528f68a

Observation c5c4a680-b8d6-41f9-af77-8246d708752a · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.810662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.810662Z digest=sha256:3fcbd432e29d38475e3f9e573e0b68b78faff4f5079daa1cb86d4669c3f6a64e

Observation 6efe08f8-6c48-428d-860e-93166fb6f430 · outbound

This paper cites VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.813015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.813015Z digest=sha256:5479f75d2ec1a229c01fabdd5256c84a4f6d896d86439fea585f727a84aa4a13

Observation f83bdc48-8871-4b57-83c1-8fd35ba4ae31 · outbound

This paper cites write newline.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.815914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.815914Z digest=sha256:1086cd3b6f98b3996db5cdfda97c843b38b9c26350ea9d08e0113157484432bd

Observation 9f72e5de-83b8-4647-8246-9b592ddc57b3 · outbound

This paper cites @esa (Ref.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? @esa (Ref

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.819576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.819576Z digest=sha256:10d5fdceda16b7fcc74854b7c907936b7bf4518a7dd39785fb1f6323ddf9b0b3

Observation 4b8b12fa-ce42-423d-815a-2323c93d2a44 · outbound

This paper cites an unresolved cited work.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.822587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.822587Z digest=sha256:0a30bf196f4d4b2d49b2c852c5f43620daf402c48bc138291657471a138380ce

Observation b38d3007-f889-430b-b41f-c538f42b203b · outbound

This paper cites !1A Qa.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? !1A Qa

Reference 61

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:36:25.047164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.852757Z digest=sha256:16e2947792ca2d07be50c5e235f3e3d7e82c30a58ad23996ec8b417058ae99fc

Pith citing papers

Observation 7f637f93-847b-483a-a77e-e272dc0cf8d2 · inbound

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model cites this paper.

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:03.335574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:03.335574Z digest=sha256:ce10e9c258d17512145159cf9d547e22637373973368d73d471a0d2945853d2e

Observation 0b3e74ec-ec27-4832-9a1c-3f71c099d409 · inbound

Modeling Human Perspectives with Socio-Demographic Representations cites this paper.

Modeling Human Perspectives with Socio-Demographic Representations ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:11:20.843879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T05:42:04.716416Z digest=sha256:f60bfec5393a4fb4a780b615d50d9a91a21a7eea9b9721c889be49138533a75e

Observation 79eff9a3-0a08-4a2d-9908-10ce9e71b0e3 · inbound

Quantifying and Predicting Disagreement in Graded Human Ratings cites this paper.

Quantifying and Predicting Disagreement in Graded Human Ratings ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.840823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:f46f0ea8349ce8a2cc7a4cb5c91527e7c5310c8ba67e75eb9179f58fb9643c51

Observation ac12c993-210a-4983-a4e0-6ff6f93ff381 · inbound

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation cites this paper.

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.684407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T06:48:22.594002Z digest=sha256:0e5acd7931d2d6ac838170d23329ac6ddd651bcdfddf8f09690c96066eb7c018