Pith. sign in

Paper Citation Record · LEDGER

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education

As of 7 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2605.31212.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.31212 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:12:32.101737Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2bd62212-ac62-4e01-a251-6a01f61377ff · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education Emerging Properties in Unified Multimodal Pretraining

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:12:46.510598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:a7b6b98cb69552d98dc84f084bb71792da31ea3ab2b4fe3682433d498517e13b

Observation b4123e56-afa1-4e98-bdbf-2ea4d7300008 · outbound

This paper cites InForty-first Interna- tional Conference on Machine Learning.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education InForty-first Interna- tional Conference on Machine Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:30e4bbf156cf444dcb4b58e1768dc7e23464c27a3986a4b6a85b69f6b5dc83e9

Observation d672b9c9-6341-43e7-a694-db9f63f4200e · outbound

This paper cites International journal of Stem education, 2:1–13.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education International journal of Stem education, 2:1–13

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:99ad91ecceacfd9b24992e31c73c07f0feeeea2f9245e0b565383d690ce6d3a8

Observation ef16ec1b-2857-4d2c-918e-ebe381ed82d6 · outbound

This paper cites John Hoven and Barry Garelick.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education John Hoven and Barry Garelick

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:d35542282fc50da627e9257e87a60927554f1e4a312dcc5e7540c24fb90d70ef

Observation bece03ae-f791-4e47-a05d-eb6c929e49fb · outbound

This paper cites Featured Certification.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education Featured Certification

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:e1915d5aa9cf84376888466aa554a45a59da849318037bdbcb6a0cf2bb862d01

Observation 2e536e8c-260d-4890-9277-919708796976 · outbound

This paper cites Ministry of General Education and Instruction, Re- public of South Sudan.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education Ministry of General Education and Instruction, Re- public of South Sudan

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:5a485e717a0f0b0f4ca143d2f1f5cbb838e074c0667e5732ea1e67330eccdba9

Observation 15d0897b-ccfe-4f79-a439-0d17abd1a35f · outbound

This paper cites How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:12:46.508029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:0b2b490eacb26cb967407a2c42c94199d4117b24c99df33b2276459ebc968b71

Observation 985f120f-c920-4c7c-ae3a-11c5d6b448a3 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:12:46.513148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:e3f0f589a0ce4916866a0bac5ae1f7e623c3abe17904969a4563e1c11caa8fab

Observation 33863086-c1c3-466c-b1c4-c3a011572f8b · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education Show-o2: Improved Native Unified Multimodal Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:12:46.505173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:27436fd2fd90b783ffd68473b1e60af43e0a4c76cadb4aa03b95fd8b4c48b1a0

Observation 4e5cfc5c-ec8f-46b6-8332-ce4a111e244b · outbound

This paper cites Create a cartoon style image to visualize this equation:3+4=7.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education Create a cartoon style image to visualize this equation:3+4=7

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:82126ebe2678f39e935027366668afc9ad3e524c656be602c631850ce0e7ca59

Observation 33211178-dcb5-474c-9bed-44af3499055d · outbound

This paper cites an unresolved cited work.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:ca891f7d0f0d06c5983c6e31b25814fc1c6f0621e231c7cc4596d6f95790524b

Observation f2b12a3d-7efd-4760-aabf-ffb44f7a3351 · outbound

This paper cites an unresolved cited work.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:9a5e97f487112512800530aa51e8ffa3200cb36ee62e1914e6137b99898487aa

Observation b534a410-cf84-4250-a111-220d13a8ed0f · outbound

This paper cites If the prompt involve same type of objects in different color, group objects of the same color together.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education If the prompt involve same type of objects in different color, group objects of the same color together

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:dd4ebd8f0c41f7be8ca0f5c90be671c002952d8cdd1aebdc5848a2ee67b38bf7

Observation f5274e41-e819-4d57-ac05-8d65a97227aa · outbound

This paper cites If there are too many objects, you can use a top-down view as indicated in the Background prompt.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education If there are too many objects, you can use a top-down view as indicated in the Background prompt

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:9051c0e5eb1b41a6d7513ec4c853390af6c8b1331426c19446adcab416cf0fbd

Observation 87fa670b-5a2e-44fc-85ba-682bfaa1239b · outbound

This paper cites Example:.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education Example:

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:234298c680a75ec5ca2202e49f49e1c3f95390e17b4fcfdadd35894f8eb35470

Observation 55115d66-1702-45fb-8801-93a5a323dde3 · outbound

This paper cites A short distance away, there are eight green balloons also floating.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education A short distance away, there are eight green balloons also floating

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-06-28T23:12:32.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:582a67680e3e7b645c70ad1e2594523946b09bc48355d344c990fd940e3e40f5

Pith citing papers

No inbound Pith citation observations are available.