Pith. sign in

Paper Citation Record · LEDGER

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

As of 22 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2501.00958.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00958 v4

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:41:33.683491Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T15:27:04.228144Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:27:04.347986Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19a81898-c503-4dc9-b66c-74b2259c93a4 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.415370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.415370Z digest=sha256:e8dd97e6f6e8facfa96126d3c5642faaac3cf9ab25a2ac5592284c2c71bde3e1

Observation 5d468dde-5869-4eb3-bf8b-6d5243185b2a · outbound

This paper cites Phi-4 Technical Report.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Phi-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.420477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.420477Z digest=sha256:6fbfe642c07b887c0ced751d25d170cfef781b5cb21b93a4ec9d0812e81fa20c

Observation d1b21053-f891-4028-9c21-1c167550c89f · outbound

This paper cites GPT-4 Technical Report.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.424920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.424920Z digest=sha256:88d8eb75a21f21b6c4f74f79bae58a099c9cd7cdb227b21f25792f9f0b7be098

Observation 70e7d395-7f3f-44d5-b60a-900eb5e1c84b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.429340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.429340Z digest=sha256:00c12fe6c9f38c8c9d9ee71d6b7742387b493d2625a15fbeb795df141c5458c2

Observation b4ba0d46-c611-4779-94ff-6deaecfd7e9f · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.433950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.433950Z digest=sha256:5ae2e14bd411fee766065680fc3abf6fdd4bcbcda8c36e8b95a1e4b8b08e970c

Observation 6da73a83-1a9c-4f35-8ccd-c6d10f6daf40 · outbound

This paper cites MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:41:34.063728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.438886Z digest=sha256:cde5960ab6de78aa9aa0a987f6c9e69348b513d678a24280d847399c5859e37c

Observation d5cfa623-c584-4114-977b-0878e5c3958d · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.444839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.444839Z digest=sha256:570b9cc42d0789cd3d375e02182a2a31101555717f9363ec626e1c5e08fa95b5

Observation 63a009ad-b5fc-4f62-b699-44e810313bc0 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Coyo-700m: Image-text pair dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.424654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.449374Z digest=sha256:c2ccca2a34ce7989353954b99e22b6c5f9e230cb7e2151328bf5d7cce8e9d721

Observation 27405bc8-1e8b-4909-8ce9-6709ea4a9fa5 · outbound

This paper cites CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.453802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.453802Z digest=sha256:672646d6078f72ac6ba9fbe5c6d0e9ebac66f9232d0e9a7bf852a0f54b797059

Observation abc1fe38-9012-4cf2-a9ab-826fe649d5d7 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.458104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.458104Z digest=sha256:c276f4a83c8a033d375c2dbee8664aa35ae4bf2a715721f6763e62d7869b787c

Observation 305a66cc-66bb-4c0f-98c6-b63c226491ae · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.462287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.462287Z digest=sha256:ed13550f0b3d603d4a2c6e2d3ac619818c4edf5bf56f813b800b5aee6d64b37d

Observation 65f0a421-4009-4948-907f-b4dabf5d67c3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.466554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.466554Z digest=sha256:e6e46e999b78f87ffc103bbad8cc31d3b3aaffeef04622602177a71f6c78033c

Observation 7417e222-25cc-4ed3-a0a5-430725dcbf4c · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.470574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.470574Z digest=sha256:b50435762e63273f3a3a873ff73bbe3f8c3a67f40b4bd425a647047dd71196e2

Observation cb38d3f1-76b2-4abc-9760-f168676a69de · outbound

This paper cites Textbooks Are All You Need.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Textbooks Are All You Need

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.474634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.474634Z digest=sha256:e6008b5399e2e8de38d41e8f3dace54157e65a9e56fa76c8c051ea6ef789d35b

Observation ea6ee0ee-fb3f-409d-b746-f4e57e05b61f · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.478758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.478758Z digest=sha256:68849b4b0b704e90aa6a9b251dd1b8c53f8d435f0bff335764a28c0fa8272da9

Observation e84cd31b-bb6e-491f-acc8-75453cb690aa · outbound

This paper cites Language is not all you need: Aligning perception with language mod- els.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Language is not all you need: Aligning perception with language mod- els

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.412410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.482505Z digest=sha256:d638957fe26b7c2a1e916898b1b3eff5ca13d52f5f3e00d2c4cf8a3ae452aee2

Observation 3c082ae1-28ad-4531-9aaf-9217b1f59e70 · outbound

This paper cites Phi-2: The surprising power of small language models.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Phi-2: The surprising power of small language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.400963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.486102Z digest=sha256:b2c0f119a3aeb0eb4dc5606b65c58e7b5c7873d71fd4995dd825ad6bfcac1add

Observation 796e83b5-d867-4a60-93dd-1c144f0fab90 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.490244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.490244Z digest=sha256:27083111f0ece25fe927457e3ab09fbeb1900db85d56594af437dbc6fb8b4d6c

Observation 1ef98b67-3213-4d26-8d38-47b74fb39781 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.494187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.494187Z digest=sha256:742690bfcc54077f2ef9ecf7152227650c90b572fb73469340caf684cd89455c

Observation 4662635b-f70f-440c-bc44-1957f06b8306 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Building and better understanding vision-language models: insights and future directions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.498202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.498202Z digest=sha256:7487e6dabf94c44da555370a543dfc199fc0aa27aba3243f2401501b3bacdfb6

Observation 4f70afbb-80fb-401c-b613-9d37cd13c1db · outbound

This paper cites What matters when building vision-language models?.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining What matters when building vision-language models?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.502091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.502091Z digest=sha256:26eca51c4ca5e211002631f0e869ced56fef485d2123bb3850ba4b2d11973d37

Observation 0fe4f996-0c70-4d58-b9c0-8393bc5eaab2 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.505809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.505809Z digest=sha256:6fe61f5036eafa1b9e75b0b1f2d11ca20f5701d79967fccc7e7fa9be069fed22

Observation 78360fe2-975a-4e0a-a6f7-72197a3a3d47 · outbound

This paper cites Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.376772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.510065Z digest=sha256:5ed0ab95087caa62513368602a370040514f89b2b906f748c2390e9d68170fe8

Observation e6e3b362-b6b0-4b4d-99be-20510be8ac04 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.514036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.514036Z digest=sha256:0717f9ab343ec2895a3f4ff30902a52c690f543a77d44b926ba0cd4649de87c3

Observation ac0ba78b-5f72-4496-89f4-a62cafe8728f · outbound

This paper cites OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.517914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.517914Z digest=sha256:5e84002a648a0ddd43a14db0bebe03adbcb4c94dbce66c8463948b94b5f5d598

Observation 63661e11-afca-46a1-be82-c457fa8f58d7 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Textbooks Are All You Need II: phi-1.5 technical report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.522200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.522200Z digest=sha256:0ef6db0da947a5eb124f1917d4b3cf8f63b70e4928025bf42b4cfa5f807e043b

Observation 1a1f334b-6289-441d-b46a-ea7f2305610f · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.526669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.526669Z digest=sha256:de2968acf0d6f8ef71c6bd28ca57abae6950a458f8e15dbe7b5a7a837be70336

Observation 166456da-11b0-4c38-8319-91d3bab432f0 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Vila: On pre-training for visual language models, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.357362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.531559Z digest=sha256:6559cc1d020d5c2624365e7eb61c72f391afb88bb0910ac14edffe5a4f2bf49d

Observation 53517bd6-ca20-41c7-a2d1-d9d3fada8173 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Improved baselines with visual instruction tuning, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.535728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.535728Z digest=sha256:e5db43dfa5c1d2c477629ac0432b046d6d0efff3725544fb009422f0236582ea

Observation 68b767fd-cd5b-4228-a018-fd99a2f01690 · outbound

This paper cites Visual instruction tuning, 2023.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Visual instruction tuning, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.540060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.540060Z digest=sha256:4a3361cbcb27177c90c438698186fc63ca77182c79d835207a2433bb6df2f684

Observation da4c6b75-3f40-4f4a-8471-3b20b200386d · outbound

This paper cites Improved baselines with visual instruction tuning.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Improved baselines with visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.332208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.543877Z digest=sha256:277bbd0931d3eaa1abf0dea22b5d5a23e9db316e646d77a78031664062885c4e

Observation ae97e300-461c-405d-a3c3-14944bc3b69a · outbound

This paper cites Visual instruction tuning.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.547555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.547555Z digest=sha256:fbafca88fdbf3aec75fbdf0ca76c944275adad6a33a29d6cadff861d25c85443

Observation d3b828b9-517c-4837-9239-0b8ee1b52986 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.551296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.551296Z digest=sha256:8e53a308d682a385cce42dc2b99158331ad96f949c5ac830fb7edadf7c952644

Observation 2c5fd610-f5d1-4690-bab9-f151740d2b5d · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.555750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.555750Z digest=sha256:dd0661f80fe4844ebe83a9b747d5151f1375ad48fcf47a65beb62ed385099486

Observation 446e088e-95ca-4ec4-852f-adb4f460886f · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.559914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.559914Z digest=sha256:853282d4c8d3a2d953d798f0ddd4cb56d2c0597d8cbbfdf241d233e696a5bae9

Observation 7ebac317-42ee-47c1-a01c-9d8aa13303b0 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.563686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.563686Z digest=sha256:950e5cdefdc2940d9b381ed60e62bf4b8aa4a790c0c75336c571d5098c2fd524

Observation 6f540502-10a1-44fb-802d-6ea3756ed07e · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.297900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.567799Z digest=sha256:b54eb3e0e9eb441064e366f264a09b4b2a4c1c283e5f84c8d7aad5f829325865

Observation b0f5f245-7ac8-4402-9e10-dad28df63f41 · outbound

This paper cites True few- shot learning with language models.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining True few- shot learning with language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.285874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.571772Z digest=sha256:7c0ce0af8d87895f301dcdd94defa80e114d2bdae55e85ca74b02d5457a15908

Observation 136864f4-bcbf-4e70-8305-93e908a04fe7 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Learning transferable visual models from natural language supervi- sion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.577332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.577332Z digest=sha256:2447a325d7a959994a4c7b989843e4ad9898fb177d3c7505441cef38c78a33e2

Observation 7b8dd9a1-41b7-46e7-ac8d-34b2652dea38 · outbound

This paper cites How2: A Large-scale Dataset for Multimodal Language Understanding.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining How2: A Large-scale Dataset for Multimodal Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.581898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.581898Z digest=sha256:14bcb27705e3381fda99d5be427ebdb54ffcb73b27e477ffd2252e3b6b0059dc

Observation e4f81f18-01e3-481b-8907-b962753dbe7b · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.586566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.586566Z digest=sha256:526a6f752071c8855b79fef3cc155500914d529b7eb09c68246c84f2d0826efb

Observation 03c399bb-ea41-46a5-8535-4ea5308be216 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.266562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.591393Z digest=sha256:eca7d2b48f02be05e594692f42525791e3c293b58b18e1f958f2319c39f87d01

Observation cfb94201-92a2-4bc9-b51b-4bb9643bf3c9 · outbound

This paper cites Towards vqa models that can read.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Towards vqa models that can read

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.594974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.594974Z digest=sha256:6278f7b3b8117d3ee481e96f2da7a3feb3162e0cfbfa54806d32d8faac100b33

Observation 7d7a3e7c-3e33-47ab-9117-8c6093582396 · outbound

This paper cites Generative multimodal mod- els are in-context learners.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Generative multimodal mod- els are in-context learners

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.246865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.598821Z digest=sha256:a6eee1dd23093ef92544bfa748ba2ae733bb6391f0c46e806080692a6fee8313

Observation a4bf3cd4-a804-47e7-8021-08b2e75e73b2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.603647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.603647Z digest=sha256:3448a3c6ecc693b202febf835a342e7b5d9d7f374a81cf653a3d23181a74a5e3

Observation 2862ef4c-f3d6-410f-9b37-35db36abbf42 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining LLaMA: Open and Efficient Foundation Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.608533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.608533Z digest=sha256:ad7d3b4682632815c4ae74951d965759ab7a902e4c8f319ddbe7c8eeb73a0150

Observation 535b67e5-2d7b-49b8-ad45-1e3e1a625524 · outbound

This paper cites Mobile- clip: Fast image-text models through multi-modal reinforced training.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Mobile- clip: Fast image-text models through multi-modal reinforced training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.236099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.612968Z digest=sha256:2cc8952fdcf655b58df73d89c778437dcf936916bdca7b60b990c032062982a0

Observation 01e51f8b-1e81-4081-893f-cc2d29d6c965 · outbound

This paper cites PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:41:33.815663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.616945Z digest=sha256:be2752d17dff0eaf00fbbd5acc4a7e84c1674cbec5f018186d6425b6f940bd2d

Observation 81a89059-0ce5-4ce6-9ef3-025dadc1d6b9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.621736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.621736Z digest=sha256:fbf061c7ddc7f72ff47056e8242ec58abfd3b086647eef0adde066f52c2264f0

Observation 368c1881-c794-4ecb-a8f0-b373eb63b056 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Image quality assessment: from error visibility to structural similarity

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.625319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.625319Z digest=sha256:a499a985615714a18aa9611d500fe310f4192368df9f9103ca44423060877631

Observation 7ca87b8d-1d5c-48a3-93f7-d08752ffe494 · outbound

This paper cites Qwen2 Technical Report.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Qwen2 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.628692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.628692Z digest=sha256:c1cf5329e8db08b71e00f19827ef561a39291ceb20d5aa67b9c235b7bfc493a8

Observation 5cf22d33-fa1a-4350-b76e-3cc94dd0751f · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining An empirical study of gpt-3 for few-shot knowledge-based vqa

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.216387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.632037Z digest=sha256:a862d0d74faa70122ad8743e18c4c9cb5ced6649c952816d85ff71ab7d8e0952

Observation 0df0f12a-c6bb-45d4-98c9-88d37be09f77 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.635688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.635688Z digest=sha256:40bde6776d7be68fdd13ea3f884ee97653bbc391dd5ebdb3a3c3a75cad5537c8

Observation 6081b138-d8ae-424a-b59a-646d4feaedbb · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.639213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.639213Z digest=sha256:5df4969d8a22b7e1d2df8bfd99263f0918eb641bf05453f77c378264d29fa72c

Observation 1004c520-fb61-44b1-afe2-bb02954dee39 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.205725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.643096Z digest=sha256:d644e441bc5c6a06796ffc96cabe9d5653fb3061ec4d22acea905dcede2e8bd5

Observation 4e868f3e-21e9-4443-8232-f5746ad3b439 · outbound

This paper cites Merlot reserve: Neu- ral script knowledge through vision and language and sound.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Merlot reserve: Neu- ral script knowledge through vision and language and sound

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.195186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.646884Z digest=sha256:6d887d7512164df4733aac6e70e7ac727691d4c2f47594c29737c3c26cfd0750

Observation c8521ea9-ad14-473a-8b8a-b88338adf286 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.650394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.650394Z digest=sha256:bb84cc7317876e1c560115cc3b8b283f2c87f42e2c4b1cb37ee111e2ab7481e0

Observation 456ffcc3-2a1e-40b8-8d45-15eeca180e0c · outbound

This paper cites Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.654630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.654630Z digest=sha256:c24e3565ba7fc8451a81db8f119cf86e14be8a42718e4ffaa433fe8459b4322a

Observation f56c5425-d8ee-434c-a30c-f86198eea148 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.658362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.658362Z digest=sha256:a723b6ffeb07f4ebc4e69b30c35bc74701a48bed58e7f4876c9f4e54402d2904

Observation 6aafec04-ee17-439a-be23-2f4777742297 · outbound

This paper cites Multimodal c4: An open, billion-scale corpus of images interleaved with text.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Multimodal c4: An open, billion-scale corpus of images interleaved with text

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.183742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.661842Z digest=sha256:c99082cf64f0c2e4386c077215200de67595e7b66960eaa1f3406a7abf9d0d59

Observation 34633937-b0ca-4b6d-98f8-076e42a0e327 · outbound

This paper cites Implementation Details When synthesizing the Knowledge Taxonomy, we utilize GPT-4o to construct the taxonomy.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Implementation Details When synthesizing the Knowledge Taxonomy, we utilize GPT-4o to construct the taxonomy

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.170531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.665444Z digest=sha256:f76eb3a20574ff0d8b0b36a7be64d0f4a542e76f82bf28e73c6f970cddba7f1f

Observation c2c8b565-e288-43bc-b52a-676f0f40eb11 · outbound

This paper cites an unresolved cited work.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:41:34.157996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.670801Z digest=sha256:e1418d3057a726b4ccadf00c335f2d924e988e0ac222bb0b2f7925a5f3759fa9

Observation 65ef5647-1766-4148-ae45-097d828e9668 · outbound

This paper cites We will continue to improve the quality and knowledge density of our textbook.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining We will continue to improve the quality and knowledge density of our textbook

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:41:34.146027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.674940Z digest=sha256:5a16ba914d6d846105e4dbc537fa1d93808571fe8405a1e85074ccdb9cab7bb9

Observation bf2a67c5-9339-456f-881b-0d87115b81c5 · outbound

This paper cites an unresolved cited work.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:41:34.134239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.679491Z digest=sha256:5de20c6a0176537a4a299d67efa9c96fa64757a935df0e8ae6b272d3d4c3b90c

Observation 735ffa26-4f41-48f0-845d-2f5b7dd4669c · outbound

This paper cites an unresolved cited work.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:41:34.122765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:41:33.683491Z digest=sha256:c7e46d1df1f0885f91da6615cb42eda74a748d86e3a3ccc8e5520b29b0a3b423

Pith citing papers

Observation 07cc1ea5-4d39-47ed-aa23-194287dc94d8 · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.350557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:526cce264e359210024d87c1a857df404f729c50dbca628222ffb0438b98fef8

Observation 94d03705-9203-4199-b635-f0431c84279a · inbound

Logics-Parsing-Omni Technical Report cites this paper.

Logics-Parsing-Omni Technical Report 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:40:01.727666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T13:37:44.189839Z digest=sha256:f03f1143a533fab9c70d9ba1bdd421b9d5f25db5983f8c78846697a8a58452b9

Observation 3509a9af-5078-47ba-9cd8-a82bdce4b7d4 · inbound

Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding cites this paper.

Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:27.436177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T05:01:38.118237Z digest=sha256:4cd57fcb243e5804aa461ef3115b89d2ca3401cf8132d744ac35b9ca3912d179