Pith. sign in

Paper Citation Record · LEDGER

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

As of 12 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2412.02368.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02368 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:36:24.852757Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:19:03.335574Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:54:20.682844Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact4
  • verified fuzzy16
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c5372279-12b2-4b0a-95e3-63ce7951c627 · outbound

This paper cites Automatikz: Text-guided synthesis of scientific vector graphics with tikz.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Automatikz: Text-guided synthesis of scientific vector graphics with tikz

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.602380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.582070Z digest=sha256:8722dde622f2aaf33d4c17de148105e258fe7d446821e99056c5186e119f33ad

Observation 661d87b5-b107-46a6-8f3e-344ef429b8df · outbound

This paper cites DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.586355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.586355Z digest=sha256:2e40890d9bca2595ab450bbb3ea5e73887c1773ecd75c152cb32fe3fa6ef9eae

Observation c3b372f0-0d0a-4809-bdd9-20ef8b921b47 · outbound

This paper cites Scene text visual question answering.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Scene text visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.594425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.589521Z digest=sha256:89e96553689c43b8c8e39bc432e3ce558f82bf936bbdbebc9136ce22f81be407

Observation b8e1369e-e701-4757-b420-93dc7a76d9da · outbound

This paper cites Elicit: Language models as research tools.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Elicit: Language models as research tools

Reference 4

Resolution
verified exact
doi, observed 2026-08-11T23:36:24.958657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.591998Z digest=sha256:a30888c0bebe2f745b0b43b5a4ac8804286806926e0cda50c262a01f104f27b2

Observation c122e92b-403c-4fd4-9363-a2db1914d52b · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.594724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.594724Z digest=sha256:53b67a3d801d6a4d52823549229d02abacfd976ed677c28078ad02b208319eb6

Observation ef81c52e-34e9-40e8-87fb-e3d08db2bbe1 · outbound

This paper cites Transformers go for the LOL s: Generating (humourous) titles from scientific abstracts end-to-end.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Transformers go for the LOL s: Generating (humourous) titles from scientific abstracts end-to-end

Reference 6

Resolution
verified exact
doi, observed 2026-08-11T23:36:24.951208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.597894Z digest=sha256:cde0aed03d67f8a74db245062c48acf7ca48fecdbde56354e1e0b63187850cfb

Observation 70b5358e-5c68-488a-acfa-67c1e082bdc4 · outbound

This paper cites Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.586914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.600980Z digest=sha256:12e825c7eddf3fcd97ce96f5fd0cd7219533500165b5e7ae42bab582531ec1b9

Observation 0a8ae3c3-3b0d-4f63-a50a-4f3a9c678043 · outbound

This paper cites Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.603625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.603625Z digest=sha256:e4e32cac36f10d1695c6585fbffa2c23c361dc7aa66f63a9593cd57cdc549753

Observation bb9af1da-ee0c-4a7f-90fe-2960baab0669 · outbound

This paper cites Exams-v: A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Exams-v: A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.574864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.608636Z digest=sha256:8ba411c59f52beef05320cc1bf083641b126349ac6e83ecce0a0674970e29a78

Observation 07437ee4-269f-4035-b987-8e24aa871502 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Imagenet: A large-scale hierarchical image database

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.611608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.611608Z digest=sha256:6f8dd529ba6f112ecc4bd35a54c16da5b33fe4f3890379d6a908e963d45b6429

Observation 0ed76c6a-7ea1-43c3-8341-31fa89ae0b0e · outbound

This paper cites The mnist database of handwritten digit images for machine learning research [best of the web].

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? The mnist database of handwritten digit images for machine learning research [best of the web]

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.614395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.614395Z digest=sha256:28a398bec755ccc96194859a49d5f5999800b09742a4d911a320a5a635aca4cc

Observation 7945bba2-7c8f-43b0-b9c0-8be88271c234 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.617536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.617536Z digest=sha256:65965585751c1b5e351f691aa5346d8faab6d0eb755a35c4c63b365cc1189f41

Observation f787e6ca-3b28-449a-9e3d-edbaf58ea912 · outbound

This paper cites CLIPS core: A reference-free evaluation metric for image captioning.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? CLIPS core: A reference-free evaluation metric for image captioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.620687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.620687Z digest=sha256:aca98c1c4f7ea57e3c692d30f4ee4e05b4c881d1a15d5cf4c88d5297d0a2c8ff

Observation a3b40108-ae91-4a10-8f24-37c94459b2f7 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.623681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.623681Z digest=sha256:2ff0f76faa13fcad6b81d3f359c6a2aa5981621ed2159c19152dc60e08bb7be6

Observation afe80c1e-7d9e-4037-bd1b-ae7277da6592 · outbound

This paper cites Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.626213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.626213Z digest=sha256:98727d0f1867ff1c677a351a1085a39fd7d20018ccb47427e7982676e012353f

Observation 261e4018-77df-49e3-ac78-748fb089be3a · outbound

This paper cites Learning multiple layers of features from tiny images.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Learning multiple layers of features from tiny images

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.629335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.629335Z digest=sha256:6facf2b93526d133da396689134d32c60911442d846840cdc840d1a5760fcbfd

Observation 56486e10-e646-4248-8dfb-7633f3915b8d · outbound

This paper cites West, and Bill Howe.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? West, and Bill Howe

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.557982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.706202Z digest=sha256:b3e313b4f883ab208503d6b1973b691de0c3c897b623472b017ec246e1839940

Observation c742c38a-ca43-4ae6-b5cd-3cef0c8e6df2 · outbound

This paper cites Holistic evaluation of text-to-image models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Holistic evaluation of text-to-image models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.549898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.709142Z digest=sha256:c5f6dccf3080250cdde98255fb9b9cc3a7da875be8e067f438b6b31e94de7999

Observation 4065fcbb-6140-4768-80a3-d1444fff5df0 · outbound

This paper cites PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.712170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.712170Z digest=sha256:7bbe5aa666a229aafc1c64b6b71158ba64ac78e5437860ca6fe9b7efdbd581d3

Observation f5c09fb1-54bf-423b-9565-843611a298aa · outbound

This paper cites The E val4 NLP 2023 shared task on prompting large language models as explainable metrics.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? The E val4 NLP 2023 shared task on prompting large language models as explainable metrics

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.715531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.715531Z digest=sha256:2649afd04bfcaf0728684cf6566669b2fa04c88676d65d419f7b45a33ae428b9

Observation 601ddb87-e901-45fb-aa69-054dc81ebdb7 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.718590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.718590Z digest=sha256:361d889627f26f8f5fa1913bbd02e5e31ed2b1552b7943a01209f88c8430017c

Observation 4eda8e9c-c886-4254-a71d-24491027ff13 · outbound

This paper cites MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.722836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.722836Z digest=sha256:f4da8be5a68d1eb78df66f840b7e708d4066d59788bfb72fb1de5ea43e56dc5f

Observation d7f6339b-00bf-495b-8543-e0446f83a03a · outbound

This paper cites an unresolved cited work.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:36:25.541809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.725965Z digest=sha256:80a6ea6fe6b9d76d605097021db1cc22a5c16b71f14535b9f9be03f31e08cc2d

Observation 6c95bd47-95a3-4194-87cb-380e7dec5d08 · outbound

This paper cites Microsoft coco: Common objects in context.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Microsoft coco: Common objects in context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.728607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.728607Z digest=sha256:04f2bf3ffeb12c464d9eb5a96e55026735073ca71aa60d2fdd092e33d2d54dca

Observation 62b7b0d7-f757-4fcd-90b3-606beac84412 · outbound

This paper cites Figurefirst: A layout-first approach for scientific figures.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Figurefirst: A layout-first approach for scientific figures

Reference 25

Resolution
verified exact
doi, observed 2026-08-11T23:36:24.932781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.731346Z digest=sha256:1a4b64812dbbd9d0733bd3a06f39004c08176e6067703c619b0f25c0c9b0b9c3

Observation 69122b81-1fe2-49af-85bd-5b64ed6f81de · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.734072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.734072Z digest=sha256:a93e51771b62ede45e30b7a1cf78249183a46a9d64dde2cfad3290a8dc5a26f0

Observation 8faf7e59-5c32-495b-abf5-dd27bb19d193 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.527485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.736948Z digest=sha256:d75c7194daa6874ac41777fb828664c7dc8e694cabcca19a2d43f652c47b5727

Observation 7e9585a5-dfb1-4cf9-82fa-09a236012bc5 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.518354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.739695Z digest=sha256:6527cc6a150519b918bac7e479f5aee473c903eb8e943034547237dfeb80a542

Observation eb18100c-eed6-430b-a3fd-c97b21c86cb0 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.742162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.742162Z digest=sha256:a4bf46db6a002052a15daf9977d6bcc5a09f17344768c03d838f7fe543bb8e3c

Observation e1125bea-1dd7-4c59-a8c5-d20b99b7276f · outbound

This paper cites State of What Art? A Call for Multi-Prompt LLM Evaluation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.744889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.744889Z digest=sha256:f503376a11182077b8e9b03234154ab16c8426253ae6db85cfafd341835075ef

Observation c32e0fe4-2bdf-4e5c-b6cf-6c2597bb6c0f · outbound

This paper cites Evaluating large language models for structured science summarization in the open research knowledge graph.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Evaluating large language models for structured science summarization in the open research knowledge graph

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.747047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.747047Z digest=sha256:b31677faded4f2ba6252e7d18e23abf12570eb7267a4a5c1423903165174cb27

Observation f8fd8537-aa7e-47e4-8735-70dad3f503c5 · outbound

This paper cites Introducing openai o1 (preview).

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Introducing openai o1 (preview)

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.509261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.749337Z digest=sha256:37defc6662bb688c65f53da67baec17005c275506d37591934326ce4a42635a2

Observation 93aa862c-c758-4532-8ab0-5fb96977256b · outbound

This paper cites Zero-Shot Text-to-Image Generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Zero-Shot Text-to-Image Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.751632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.751632Z digest=sha256:077a3a7bf2dead46ecdbd8ead2365ffcf9c8289145169ba22146c9882a48a4b5

Observation 3ce14d11-9b2e-4f13-96b3-afe3728de606 · outbound

This paper cites SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.754369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.754369Z digest=sha256:ce930ba34879b7b9d3f238fbd95ab3936e7734c874d046e45d91fb64d604a08f

Observation f4c579dd-bdd0-4f06-a76b-dcbcce905dd6 · outbound

This paper cites Ocr-vqgan: Taming text-within-image generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Ocr-vqgan: Taming text-within-image generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.500047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.757290Z digest=sha256:32418087be7691ccef44879cde58e28c9c947f204a999501cb5289db2126d805

Observation a53ccde3-da8e-44e1-bcad-a2c239059785 · outbound

This paper cites Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2).

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.759801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.759801Z digest=sha256:34687ff37f52b2c5cf85d2c6d14f0b431cdc1b9b837433d4ce5a44b324847a6a

Observation a1c67c02-191b-4eec-ac8a-825d45a73b44 · outbound

This paper cites Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.762652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.762652Z digest=sha256:604b353124c0de99e7386339c9173712b2248a4d75058a78eafff61dcec3fdda

Observation 8339022e-882e-42c5-8e43-30b00e4c53fd · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.765471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.765471Z digest=sha256:bd537a1b9dd1fae11ab801c7462afcf28409dce4c6763106c83626069de79897

Observation 8c211054-9ab3-4b9d-a2da-813c14d3d249 · outbound

This paper cites ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.768284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.768284Z digest=sha256:c87d4e4b2a326fe1b8fb2497219ff47863352710bee4aa7e8cdc6d8c0fb95bac

Observation 92c77b8d-fda4-4811-ae56-944a4fcbd4ee · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.771267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.771267Z digest=sha256:b9dd8c7a30cb4dbe5b31110f6cc12fc36fad48515c64ca0f4b41a88b3d9cb2d1

Observation 6a9f03a6-efa5-4f09-8288-b1a2a5634fc3 · outbound

This paper cites Towards vqa models that can read.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Towards vqa models that can read

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.488831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.774247Z digest=sha256:6fcde7436e71bf4d9b5ad9b30dae1f8685dc9356f57705ba8c693f808e84b77c

Observation 652b5bc1-f02f-4ac5-8b8b-675c166aaf60 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Gemma: Open Models Based on Gemini Research and Technology

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.776853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.776853Z digest=sha256:5b7975a66b8f65776a1e8726183d8aa34b4b1fcc0e2fa9ff5a58c790ae7b1c1f

Observation e10d8a00-7504-4c61-ae91-60687ef018cc · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.780407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.780407Z digest=sha256:57fb600402e8c42d745df5a6e9990012612803114ed59f8b47370e10a7a7e11d

Observation 1e37139c-7570-4237-af61-b7157016c499 · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.783054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.783054Z digest=sha256:3e3c0db79b877847829b17b885dd0c4c7eb2028a14834712159b701bf93d4be6

Observation 47dad0a8-ded6-4426-b2ae-bfdabd94f879 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.786161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.786161Z digest=sha256:cea2472c57f3e1dabf50034bd1bbcce59cd62b22098474bf55c8ab946a8fc3eb

Observation a39e7aa2-91d3-4b09-a154-17c5a68cdf70 · outbound

This paper cites Plots made quickly: An efficient approach for generating visualizations from natural language queries.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Plots made quickly: An efficient approach for generating visualizations from natural language queries

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.474081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.789043Z digest=sha256:4b7760f091362859547e88c6cfd25a9f533a0cb21e8804aa446be303e5c31ed9

Observation 1859436c-9ba3-4901-92fc-0cacfaa5973f · outbound

This paper cites Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.465937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.791619Z digest=sha256:81a58b1ad4056259d3b6b2c2c5bf0edd7f10654e67aaadb01fa62b5eaac8f3a5

Observation 34ec3669-99c7-447b-9a7c-3238a9de9f3f · outbound

This paper cites Scibench: Evaluating college-level scientific problem-solving abilities of large language models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Scibench: Evaluating college-level scientific problem-solving abilities of large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.457618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.794196Z digest=sha256:8619d69d2068329fe4ed90a3355bef447695b373d78eefa4280b080f35d77e42

Observation f031aa26-6506-42b8-babb-7d68b49cb617 · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.796788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.796788Z digest=sha256:512f789cdd0926d09320c4529f46126246f4907af2dfd14a058f5a7a5efc8e33

Observation 280d9344-caab-4daf-958c-6f8f7b06dd35 · outbound

This paper cites seaborn: statistical data visualization.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? seaborn: statistical data visualization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.799390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.799390Z digest=sha256:4232d21cf857ce6dd40de4aa07afe546a071cc325911966bab8331e2914a9b10

Observation 96bdf0c0-9074-46f8-87a8-f9e5f2768fa0 · outbound

This paper cites Evaluating and analyzing relationship hallucinations in large vision-language models.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Evaluating and analyzing relationship hallucinations in large vision-language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.449624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.802121Z digest=sha256:014b3eacc874fb33098e4f1c22dc43c625117285b63b0f42372cd0fca6e019d5

Observation 804bfb6d-e949-4db8-9e41-b897d6bfa0fc · outbound

This paper cites An automatic graph generation method for scholarly papers based on table structure analysis.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? An automatic graph generation method for scholarly papers based on table structure analysis

Reference 52

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T23:36:25.142576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.804221Z digest=sha256:b7573f33d605ae019daf09c793d5e1e49ad39c3c7d33c678eeb77bcc33f21b4e

Observation 56badd52-6031-4f40-be3d-4bd11ee4d2ff · outbound

This paper cites Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:36:25.441426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.806384Z digest=sha256:c0a94f959b39009f00cd3cc2714937cbae77a494d8dc7391190da7cf7748475a

Observation 330f4ae8-8e4a-49b0-af2b-331f6281e39f · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.808524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.808524Z digest=sha256:62de5c37bf8cadb2903fabb6491ce0914800dc038c58e4de35f7c2f1f08f21b2

Observation c5c4a680-b8d6-41f9-af77-8246d708752a · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.810662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.810662Z digest=sha256:bde829d1c394fd1b0a4c68772aff812b5c5e2a4c60239c8b64c4c53f05d61f32

Observation 6efe08f8-6c48-428d-860e-93166fb6f430 · outbound

This paper cites VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.813015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.813015Z digest=sha256:a707518a8a0439b2c03ac810a553a64f40466ad2cd5ed78b8113933ae99e7f64

Observation f83bdc48-8871-4b57-83c1-8fd35ba4ae31 · outbound

This paper cites write newline.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.815914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.815914Z digest=sha256:566b27c04aef16597026d216ad2151e5a334e8614e12d1f531dfed612d34458c

Observation 9f72e5de-83b8-4647-8246-9b592ddc57b3 · outbound

This paper cites @esa (Ref.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? @esa (Ref

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.819576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.819576Z digest=sha256:8310a4a544be73aafea1915f3325d329b0704348b555ae040acb7a26510a30bc

Observation 4b8b12fa-ce42-423d-815a-2323c93d2a44 · outbound

This paper cites an unresolved cited work.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.822587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.822587Z digest=sha256:68406bc47e372bf5497feaa545183564cfc8baf7192b112b7176efba9e05d357

Observation b38d3007-f889-430b-b41f-c538f42b203b · outbound

This paper cites !1A Qa.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? !1A Qa

Reference 61

Resolution
verified exact
raw_fallback, observed 2026-08-11T23:36:25.047164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:36:24.852757Z digest=sha256:a080a7e04fe9d0ddfefcb8181c74282bf701b6ef3133be5c25248dd538e9618f

Pith citing papers

Observation 7f637f93-847b-483a-a77e-e272dc0cf8d2 · inbound

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model cites this paper.

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:03.335574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:03.335574Z digest=sha256:64b4ad0db100d27d7b536f102ccaa7f28546f8ed7452b99402e5255bd8610558

Observation 0b3e74ec-ec27-4832-9a1c-3f71c099d409 · inbound

Modeling Human Perspectives with Socio-Demographic Representations cites this paper.

Modeling Human Perspectives with Socio-Demographic Representations ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:11:20.843879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T05:42:04.716416Z digest=sha256:7c592786bba749c58f1d8e6572ef0756156f28f9041fa775ad310211fe22df46

Observation 79eff9a3-0a08-4a2d-9908-10ce9e71b0e3 · inbound

Quantifying and Predicting Disagreement in Graded Human Ratings cites this paper.

Quantifying and Predicting Disagreement in Graded Human Ratings ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.840823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:5528533f700189b0f12689342104d425b4c14f2b70d385665600245f2de37267

Observation ac12c993-210a-4983-a4e0-6ff6f93ff381 · inbound

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation cites this paper.

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.684407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T06:48:22.594002Z digest=sha256:d9e8bf70b014066dbde0b4a6472b7c8925de5ef1cbc204ff927f8a680027b127