Pith. sign in

Paper Citation Record · LEDGER

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering

As of 13 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2506.21596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21596 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:05.141147Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fae8a539-5ada-4a15-8b65-ee2a7d11eaa7 · outbound

This paper cites Visual instruction tuning,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Visual instruction tuning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:03.563420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:03.563420Z digest=sha256:0bad04aa5324ae57fcb0f08ec0fdb7700482548f0030423c4edcbc9d62a10d95

Observation 00f0d1c9-a049-4ace-ba4c-bde3bd43e63d · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Improved Baselines with Visual Instruction Tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:03.623288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:03.623288Z digest=sha256:14e233599ff5fc970db5f91a2647d2d9b081e13d78c1bb3ae79dd3b845cffbcf

Observation 4d946e54-d937-4ff2-b146-3688b3a99d89 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:03.679170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:03.679170Z digest=sha256:4b1c3bac913d5c322615fb3a4f9c8b0ace81922e5cb2d7564f88ef046bb7c613

Observation 7bd79d0f-16e8-494e-a535-8c84e774f444 · outbound

This paper cites Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:09.040655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:03.747155Z digest=sha256:47203e8e6ae0c4711f121017ca534f991f67354fd49a5873e9005c51cd1b7534

Observation 7ce0486e-a7fe-4841-9a56-18d74747f6c9 · outbound

This paper cites On opportunities and challenges of large multimodal foundation models in education,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering On opportunities and challenges of large multimodal foundation models in education,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.905573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:03.838513Z digest=sha256:b250d97d880f44d0061a42a89c8d64e6af935cbb7952ee11c4712698a36eab71

Observation 84789112-6207-439f-af96-0dcb1a5516bb · outbound

This paper cites Are you smarter than a sixth grader? textbook question an- swering for multimodal machine comprehension,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Are you smarter than a sixth grader? textbook question an- swering for multimodal machine comprehension,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.772329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:03.910260Z digest=sha256:243e9daff755c11429f6786c3dd50b52903532f0df1f98946edf4a998c66548a

Observation 151a1e35-aeab-43f6-92c5-1e26822f2f62 · outbound

This paper cites A review on vision-language-based approaches: Challenges and applications,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering A review on vision-language-based approaches: Challenges and applications,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.624559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:03.979107Z digest=sha256:ecd0d97d0aa3f99a0bbfa60ab7e2422e02ac8ae526c3cd1039b333bb1246f8a1

Observation c3a08743-84a5-47c9-9a33-7446f9f640aa · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.480404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.026641Z digest=sha256:471a3e8f9a9595c7956f05bd83247560a079a8cabd3d1f759bfdee873a32cdeb

Observation f5eeea5c-cbeb-44a8-ae56-fd4e724b9905 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Minigpt-4: Enhancing vision-language understanding with advanced large language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.330726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.086129Z digest=sha256:0029adef0070a1568195c4d94a9411404797ee576f6089bdca45e913d2ad1da6

Observation 27591440-ac1b-49a3-ab2e-1d280f14405a · outbound

This paper cites Making the v in VQA matter: Elevating the role of image understanding in visual question answering,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Making the v in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.189944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.143669Z digest=sha256:d061568c43fa74e9bbd061716eef4d6c5b4f8db913b5d795ab74180d622e8f50

Observation 8849f29d-7b38-406b-bb3d-46f2ce5e6da2 · outbound

This paper cites Ok-VQA: A visual question answering benchmark requiring external knowledge,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Ok-VQA: A visual question answering benchmark requiring external knowledge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.013450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.184296Z digest=sha256:4fe899e233794b86950fa2af3958ea5c4f8e7d830e441db898f9c800768eb622

Observation d620872c-86af-4ed9-9521-eeafac3d6cac · outbound

This paper cites Microsoft COCO: Common objects in context,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Microsoft COCO: Common objects in context,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:07.691252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.276954Z digest=sha256:9430d932d16cc49b5906384b3c09809df4372b83c05066e1fe731c78d4e4aee3

Observation f6143b45-6167-4c8f-ac0a-421e7d3d17e2 · outbound

This paper cites SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:04.357027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:04.357027Z digest=sha256:6571686879d1f1580a1442fb12f90c81c13ea706aae9965d4af395a72387ac90

Observation 0ea08724-8e05-4920-9e19-52d5ca82ee55 · outbound

This paper cites Eduvqa: A multimodal visual question answering framework for smart education,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Eduvqa: A multimodal visual question answering framework for smart education,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:07.417384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.425031Z digest=sha256:4258939a63d74be0afc58f0b620ff34f1b94e2de3aff45eb7b6a13a637d2fab7

Observation d54e57bf-b35c-4e3e-a1f2-94a3233472c4 · outbound

This paper cites Enhancing textual textbook question answering with large language models and retrieval augmented generation,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Enhancing textual textbook question answering with large language models and retrieval augmented generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:07.151392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.508308Z digest=sha256:c84dbbc37881c2d3035789b445dd97a07869540e344a867ba691a23804396000

Observation 57cf384d-eed3-4364-b65d-c3f0844c66a5 · outbound

This paper cites Enhancing textbook question answering with knowledge graph-augmented large language models,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Enhancing textbook question answering with knowledge graph-augmented large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.928385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.614008Z digest=sha256:d06c5641c8fbc069f5a9f7a68a55401a357fe4284383f73ae0c0f8068ecf6792

Observation f75ebdc2-61c3-4db1-8337-07a1e4745380 · outbound

This paper cites ISAAQ - mastering textbook questions with pre-trained transformers and bottom-up and top-down attention,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering ISAAQ - mastering textbook questions with pre-trained transformers and bottom-up and top-down attention,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.642236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.676803Z digest=sha256:d5975ca989de6b8f4b3471b56ae1a04c260a757df966fc46379392f9f91e1c6f

Observation 921db7cb-37bc-4d69-abda-1a6b2cb389f5 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Imagebind: One embedding space to bind them all,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.355230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.754929Z digest=sha256:f5dca81e870e311e0f1d7a35d0e7557f2b8d795975714854d970b4970509b7d0

Observation 44b78eb2-0f23-4782-ab3a-a9a17102ef91 · outbound

This paper cites KDB.AI: The scalable vector database for ai,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering KDB.AI: The scalable vector database for ai,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.116786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.846424Z digest=sha256:764cd64fdf3bcd29fa9a28feab0c8027fb859b643b065b36a0814cb1e22d4e8d

Observation 0c0e506f-a6b0-4bfb-a07f-4847e7c6b754 · outbound

This paper cites GPT-4v(ision) system card,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering GPT-4v(ision) system card,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:05.828269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.897266Z digest=sha256:bff8c3ed59bd6d521d16aac0d396e8de57db420e4b8da7c9658efda1f7d00220

Observation 8b00df41-b0d4-4458-a67e-8354fd5278a6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Gemini: A Family of Highly Capable Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:04.903476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:04.903476Z digest=sha256:bb291ecc5dfbdaad4c21193a1316c632391d83725553d7bfd8e72f3f2a731e18

Observation af04fca8-030e-4bf4-95f4-b09c94cea7f5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Learning transferable visual models from natural language supervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:05.582385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:04.997136Z digest=sha256:2016596570b8c879ea4f38fcd312307ebf041471e039d518c9c60836e6408fba

Observation 8e02cfc0-9382-4039-afa1-b29ccdf59431 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:05.393712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:05.141147Z digest=sha256:3131fdf69bb073f124abdd837fa29086b1a5b67ce9f4d2b03a82bd9f72ec05ab

Pith citing papers

No inbound Pith citation observations are available.