Pith. sign in

Paper Citation Record · LEDGER

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2506.21596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21596 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:05.141147Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fae8a539-5ada-4a15-8b65-ee2a7d11eaa7 · outbound

This paper cites Visual instruction tuning,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Visual instruction tuning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:03.563420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:03.563420Z digest=sha256:29dd9afef86987573b8424e40da36669c619d4d8bc70359e8354cae2fa75a7d3

Observation 00f0d1c9-a049-4ace-ba4c-bde3bd43e63d · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Improved Baselines with Visual Instruction Tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:03.623288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:03.623288Z digest=sha256:56e2df5b043df94e9f646dcb6d48507b2abcad044495df29b2705e912de17bc5

Observation 4d946e54-d937-4ff2-b146-3688b3a99d89 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:03.679170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:03.679170Z digest=sha256:937ab167ec9a99e767909c0a155dc93961fbe953eb42216f3cf371c1745d9acc

Observation 7bd79d0f-16e8-494e-a535-8c84e774f444 · outbound

This paper cites Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:09.040655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:03.747155Z digest=sha256:07593e44f8e9fbe9ab1ef0a213cbb93f162f9cc2e572a421520626d04f3ef440

Observation 7ce0486e-a7fe-4841-9a56-18d74747f6c9 · outbound

This paper cites On opportunities and challenges of large multimodal foundation models in education,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering On opportunities and challenges of large multimodal foundation models in education,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.905573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:03.838513Z digest=sha256:5f8a2ace4d339df97463239389de31e6b6e57485e85c915bc179b7f58a7510ce

Observation 84789112-6207-439f-af96-0dcb1a5516bb · outbound

This paper cites Are you smarter than a sixth grader? textbook question an- swering for multimodal machine comprehension,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Are you smarter than a sixth grader? textbook question an- swering for multimodal machine comprehension,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.772329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:03.910260Z digest=sha256:f48bae5df0bfb6b32e5a088deecd43817847be07c2e38bfc9a55512391c1b38e

Observation 151a1e35-aeab-43f6-92c5-1e26822f2f62 · outbound

This paper cites A review on vision-language-based approaches: Challenges and applications,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering A review on vision-language-based approaches: Challenges and applications,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.624559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:03.979107Z digest=sha256:b91744aa38647fbef2dd758c7caf2d98e2e00cca37807eacdf0576d3209780f2

Observation c3a08743-84a5-47c9-9a33-7446f9f640aa · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.480404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.026641Z digest=sha256:fe87a64fc3508989cf5b8ba1f2749fdfb67fe0856ba831b9bbe28e5d4d504668

Observation f5eeea5c-cbeb-44a8-ae56-fd4e724b9905 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Minigpt-4: Enhancing vision-language understanding with advanced large language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.330726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.086129Z digest=sha256:9c99bbe863b5158c273d69a0c1a39e63300e86f9cea0b543e802ff85f48bb5bf

Observation 27591440-ac1b-49a3-ab2e-1d280f14405a · outbound

This paper cites Making the v in VQA matter: Elevating the role of image understanding in visual question answering,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Making the v in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.189944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.143669Z digest=sha256:765ba76d790c5b52827b9ff0637208d4485be6b547b957f0341ee0e639f54659

Observation 8849f29d-7b38-406b-bb3d-46f2ce5e6da2 · outbound

This paper cites Ok-VQA: A visual question answering benchmark requiring external knowledge,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Ok-VQA: A visual question answering benchmark requiring external knowledge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.013450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.184296Z digest=sha256:a5215367383111431557214b25e6c3172d968165c188efe6ff6e2d226b1d1d04

Observation d620872c-86af-4ed9-9521-eeafac3d6cac · outbound

This paper cites Microsoft COCO: Common objects in context,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Microsoft COCO: Common objects in context,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:07.691252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.276954Z digest=sha256:543904d3dc8f47a638b50e4d51e885803bb0d74eac83820831a81df78ad22d4f

Observation f6143b45-6167-4c8f-ac0a-421e7d3d17e2 · outbound

This paper cites SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:04.357027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:04.357027Z digest=sha256:9d6e96ca991381090149e72a96945fe3d43f34c01b8acc42da624f1355bc8046

Observation 0ea08724-8e05-4920-9e19-52d5ca82ee55 · outbound

This paper cites Eduvqa: A multimodal visual question answering framework for smart education,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Eduvqa: A multimodal visual question answering framework for smart education,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:07.417384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.425031Z digest=sha256:2d6bfc0edf0409df148c6ce8d225d5982ff2edd548fbb51bc04288aba6bf0aed

Observation d54e57bf-b35c-4e3e-a1f2-94a3233472c4 · outbound

This paper cites Enhancing textual textbook question answering with large language models and retrieval augmented generation,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Enhancing textual textbook question answering with large language models and retrieval augmented generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:07.151392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.508308Z digest=sha256:9083d6dac45ef656bb9284b637cc2e424e9f9edbde5b7e88ccf45d4e7ff9cf3a

Observation 57cf384d-eed3-4364-b65d-c3f0844c66a5 · outbound

This paper cites Enhancing textbook question answering with knowledge graph-augmented large language models,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Enhancing textbook question answering with knowledge graph-augmented large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.928385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.614008Z digest=sha256:dc3d0f1ac40479f17cc33fe995a7cd88beba1f7b21c3e93e7787339c3ad90d2f

Observation f75ebdc2-61c3-4db1-8337-07a1e4745380 · outbound

This paper cites ISAAQ - mastering textbook questions with pre-trained transformers and bottom-up and top-down attention,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering ISAAQ - mastering textbook questions with pre-trained transformers and bottom-up and top-down attention,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.642236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.676803Z digest=sha256:3f6b81627d62f0fde095a9f503efcdba47e75957b4b013ff82ce8d87a6da1a45

Observation 921db7cb-37bc-4d69-abda-1a6b2cb389f5 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Imagebind: One embedding space to bind them all,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.355230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.754929Z digest=sha256:33221fd9a96a0e8dc01a74cdb58cc19e29f310adb726577afb6b29efbe853a1d

Observation 44b78eb2-0f23-4782-ab3a-a9a17102ef91 · outbound

This paper cites KDB.AI: The scalable vector database for ai,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering KDB.AI: The scalable vector database for ai,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.116786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.846424Z digest=sha256:27b9d733233ea36729514d80c42baf4536aae0f9601af090c3755549190de49c

Observation 0c0e506f-a6b0-4bfb-a07f-4847e7c6b754 · outbound

This paper cites GPT-4v(ision) system card,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering GPT-4v(ision) system card,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:05.828269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.897266Z digest=sha256:d40df5d71d9be3c23674e91e71d6d72af67397ecb27176dc2c07674fd9fe7c59

Observation 8b00df41-b0d4-4458-a67e-8354fd5278a6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Gemini: A Family of Highly Capable Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:04.903476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:04.903476Z digest=sha256:3fbeb0df9df143ca2e26d8ab341de8b9c269d970f505fff9a3453906bd4df60d

Observation af04fca8-030e-4bf4-95f4-b09c94cea7f5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Learning transferable visual models from natural language supervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:05.582385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:04.997136Z digest=sha256:111f81a77980e1039201dc2e99fdbd7fedb023716e4308ca55089984c36d1694

Observation 8e02cfc0-9382-4039-afa1-b29ccdf59431 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:05.393712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:53:05.141147Z digest=sha256:fbc994f40d75ac05f7478a7c433522e095454017c85c02ae24ef03f21507b9e0

Pith citing papers

No inbound Pith citation observations are available.