Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:41:33.683491Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2501.00958.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:41:33.683491Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-16T15:27:04.228144Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T15:27:04.347986Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 19a81898-c503-4dc9-b66c-74b2259c93a4 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d468dde-5869-4eb3-bf8b-6d5243185b2a · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Phi-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b21053-f891-4028-9c21-1c167550c89f · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e7d395-7f3f-44d5-b60a-900eb5e1c84b · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Flamingo: a visual language model for few-shot learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4ba0d46-c611-4779-94ff-6deaecfd7e9f · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da73a83-1a9c-4f35-8ccd-c6d10f6daf40 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d5cfa623-c584-4114-977b-0878e5c3958d · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a009ad-b5fc-4f62-b699-44e810313bc0 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Coyo-700m: Image-text pair dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 27405bc8-1e8b-4909-8ce9-6709ea4a9fa5 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc1fe38-9012-4cf2-a9ab-826fe649d5d7 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 305a66cc-66bb-4c0f-98c6-b63c226491ae · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65f0a421-4009-4948-907f-b4dabf5d67c3 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7417e222-25cc-4ed3-a0a5-430725dcbf4c · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb38d3f1-76b2-4abc-9760-f168676a69de · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Textbooks Are All You Need
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea6ee0ee-fb3f-409d-b746-f4e57e05b61f · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e84cd31b-bb6e-491f-acc8-75453cb690aa · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Language is not all you need: Aligning perception with language mod- els
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3c082ae1-28ad-4531-9aaf-9217b1f59e70 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Phi-2: The surprising power of small language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 796e83b5-d867-4a60-93dd-1c144f0fab90 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Scaling up visual and vision-language representa- tion learning with noisy text supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef98b67-3213-4d26-8d38-47b74fb39781 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4662635b-f70f-440c-bc44-1957f06b8306 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Building and better understanding vision-language models: insights and future directions
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f70afbb-80fb-401c-b613-9d37cd13c1db · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining What matters when building vision-language models?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fe4f996-0c70-4d58-b9c0-8393bc5eaab2 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78360fe2-975a-4e0a-a6f7-72197a3a3d47 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e6e3b362-b6b0-4b4d-99be-20510be8ac04 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac0ba78b-5f72-4496-89f4-a62cafe8728f · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63661e11-afca-46a1-be82-c457fa8f58d7 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Textbooks Are All You Need II: phi-1.5 technical report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a1f334b-6289-441d-b46a-ea7f2305610f · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Towards General Text Embeddings with Multi-stage Contrastive Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 166456da-11b0-4c38-8319-91d3bab432f0 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Vila: On pre-training for visual language models, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 53517bd6-ca20-41c7-a2d1-d9d3fada8173 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Improved baselines with visual instruction tuning, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68b767fd-cd5b-4228-a018-fd99a2f01690 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Visual instruction tuning, 2023
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4c6b75-3f40-4f4a-8471-3b20b200386d · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Improved baselines with visual instruction tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ae97e300-461c-405d-a3c3-14944bc3b69a · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Visual instruction tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b828b9-517c-4837-9239-0b8ee1b52986 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5fd610-f5d1-4690-bab9-f151740d2b5d · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 446e088e-95ca-4ec4-852f-adb4f460886f · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ebac317-42ee-47c1-a01c-9d8aa13303b0 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f540502-10a1-44fb-802d-6ea3756ed07e · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b0f5f245-7ac8-4402-9e10-dad28df63f41 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining True few- shot learning with language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 136864f4-bcbf-4e70-8305-93e908a04fe7 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Learning transferable visual models from natural language supervi- sion
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8dd9a1-41b7-46e7-ac8d-34b2652dea38 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining How2: A Large-scale Dataset for Multimodal Language Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f81f18-01e3-481b-8907-b962753dbe7b · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c399bb-ea41-46a5-8535-4ea5308be216 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cfb94201-92a2-4bc9-b51b-4bb9643bf3c9 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Towards vqa models that can read
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d7a3e7c-3e33-47ab-9117-8c6093582396 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Generative multimodal mod- els are in-context learners
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a4bf3cd4-a804-47e7-8021-08b2e75e73b2 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2862ef4c-f3d6-410f-9b37-35db36abbf42 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining LLaMA: Open and Efficient Foundation Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 535b67e5-2d7b-49b8-ad45-1e3e1a625524 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Mobile- clip: Fast image-text models through multi-modal reinforced training
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 01e51f8b-1e81-4081-893f-cc2d29d6c965 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 81a89059-0ce5-4ce6-9ef3-025dadc1d6b9 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 368c1881-c794-4ecb-a8f0-b373eb63b056 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Image quality assessment: from error visibility to structural similarity
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ca87b8d-1d5c-48a3-93f7-d08752ffe494 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Qwen2 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf22d33-fa1a-4350-b76e-3cc94dd0751f · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining An empirical study of gpt-3 for few-shot knowledge-based vqa
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0df0f12a-c6bb-45d4-98c9-88d37be09f77 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6081b138-d8ae-424a-b59a-646d4feaedbb · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1004c520-fb61-44b1-afe2-bb02954dee39 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4e868f3e-21e9-4443-8232-f5746ad3b439 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Merlot reserve: Neu- ral script knowledge through vision and language and sound
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c8521ea9-ad14-473a-8b8a-b88338adf286 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456ffcc3-2a1e-40b8-8d45-15eeca180e0c · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56c5425-d8ee-434c-a30c-f86198eea148 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aafec04-ee17-439a-be23-2f4777742297 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Multimodal c4: An open, billion-scale corpus of images interleaved with text
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 34633937-b0ca-4b6d-98f8-076e42a0e327 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Implementation Details When synthesizing the Knowledge Taxonomy, we utilize GPT-4o to construct the taxonomy
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c2c8b565-e288-43bc-b52a-676f0f40eb11 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 65ef5647-1766-4148-ae45-097d828e9668 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining We will continue to improve the quality and knowledge density of our textbook
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bf2a67c5-9339-456f-881b-0d87115b81c5 · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 735ffa26-4f41-48f0-845d-2f5b7dd4669c · outbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 07cc1ea5-4d39-47ed-aa23-194287dc94d8 · inbound
MMSearch-R1: Incentivizing LMMs to Search 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 94d03705-9203-4199-b635-f0431c84279a · inbound
Logics-Parsing-Omni Technical Report 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3509a9af-5078-47ba-9cd8-a82bdce4b7d4 · inbound
Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.