Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:22:35.866665Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 4 inbound Pith citation observations for arXiv:2509.01644.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:22:35.866665Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T07:35:07.825460Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T07:35:28.745644Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 04be0f62-5e78-4fb7-a8dd-7e5e6a24ee74 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aac679f-1a31-42c2-92f5-b743f973cc60 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Lan- guage models are few-shot learners
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5add7096-c187-495c-85b7-1c4837441c51 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Open-llava-next: An open- source implementation of llava-next series for facilitating the large multi-modal model community
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abb71275-78ed-4e94-a184-7080bc75fb7f · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Generative pre- training from pixels
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 313fd4f8-c736-42c0-8bfc-c223068f7d4e · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a38f1a5-bc7e-40b7-adb5-408051470dd2 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2af4d339-31f2-45cc-9cb4-a95d41a1886c · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Virtex: Learning visual representations from textual annotations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 487eeaa0-cee4-4056-9ee8-d9d6f8a06f5b · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learning Musical Representations for Music Performance Question Answering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b107f77a-1c8a-42a2-9d10-8475833fc3d6 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 505638ea-f1f5-4a45-97f1-1463a38f024b · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Scalable Pre-training of Large Autoregressive Image Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f53512-35cf-4345-918e-a3e7c47c7c9c · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Improving clip training with language rewrites
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2644c2e3-a988-4830-a232-90903decb198 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Data Filtering Networks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0318e99b-ce9f-42c5-bb61-293b850ed45c · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 941b4be7-1e48-4b85-8b2e-757a4e93f2f2 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98faa756-30e0-4b35-a4e4-807594854e0b · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning DataComp: In search of the next generation of multimodal datasets
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a14ad9-84c8-4574-90ca-4e2cff0b5ee6 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395ade9c-bbc6-4a3a-95c3-67e5ce3beaec · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9bd0af9-2c21-43c1-9fe0-b36231603e0d · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Open- clip
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd03223c-6af5-4d78-8340-399ab782ead3 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Scaling up visual and vision-language representation learning with noisy text supervision
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0aff7eb-aadf-4f01-9269-4ea65727829c · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learning visual features from large weakly supervised data
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 931daa9a-01c9-4e79-b00e-68952f45e7d7 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Deep visual-semantic align- ments for generating image descriptions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19d4ccc2-4a75-4acd-8614-b093f8c1c5e7 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Vilt: Vision- and-language transformer without convolution or region su- pervision
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12c60015-cc86-464e-9daf-49dda71f9333 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Veclip: Improving clip training via visual-enriched captions
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc92feb3-0176-4b9b-bf84-89e1c99efe61 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learning visual n-grams from web data
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14433b13-cfb1-4179-bd76-b63cca80af55 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32a9aefd-6dd5-494d-9197-e9b69b23443e · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be819195-7e08-4fa5-8a42-58e0aca29630 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b6daa7-0aa3-4079-bc33-fc9a67b7ded3 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Align before fuse: Vision and language representation learn- ing with momentum distillation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67f44edf-b97d-46a6-bf67-be7c43d53446 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2124472e-1761-4fa3-88f0-08854232d528 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning What If We Recaption Billions of Web Images with LLaMA-3?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159157b4-97f0-45f4-a65a-0f524f7c7493 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning An inverse scal- ing law for clip training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ca4bc67-469d-426d-8095-ec448cac8178 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Evaluating object hallucination in large vision-language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccc04d01-7477-4b74-a313-2d1d61bf5473 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Improved baselines with visual instruction tuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be78f29c-2021-4df5-adb2-06593a868d3c · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Llava-next: Im- proved reasoning, ocr, and world knowledge
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1eb0e4c0-8a68-481b-b4f0-9fd423270969 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Visual Instruction Tuning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5172db-a689-4111-ae86-087c771bf7ab · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e47fd057-702d-4d69-8b35-91d7acde64ef · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d66bc4c3-8c74-4396-8636-f1dcaca02fe3 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning MLLMs-Augmented Visual-Language Representation Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4e4d833-649b-466e-a76b-920ad2148b11 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01b1815e-e04a-4ccc-92a6-5e6a9a142855 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4278d2c-71e8-4bb0-af6d-0ff39448a37f · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05f11dd9-6210-4a1f-afb4-27d7e2e95ff0 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c9f56c5-0976-4beb-8888-4e92d67123fa · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learn- ing transferable visual models from natural language super- vision
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5987f3ac-96e9-4bf3-b8c5-8a67bc3d6004 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Improving language understanding by gen- erative pre-training
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a84847-cf8c-4886-b920-c439921bc361 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Language models are unsu- pervised multitask learners
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1861d3a1-de43-44ad-b365-45865d87b0b4 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learning visual representations with caption annotations
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6ccd962-bd75-4097-a000-c4602d3e3c7b · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b35f7dc4-eb34-448f-9acf-bc9f496caaf1 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Towards vqa models that can read
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d429898-6337-47ec-82bf-e4572a281f2c · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b994c06d-2bf8-42f8-a806-75fb8d5c903c · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Emu: Generative Pretraining in Multimodality
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f30213f-1104-46d1-b48d-583098ef147b · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bba2b83-8108-49f8-8411-f070ad9b491f · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Image captioners are scalable vision learners too
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e361bb20-d1f6-4f69-89b4-bea689bb6f11 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Show and tell: A neural image caption gen- erator
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5cd19f60-9f85-4e89-884c-0cb945a16a3c · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d0e87c1-b4d3-4fdc-85cd-1b468b421cd9 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Simvlm: Simple visual language model pretraining with weak supervision
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53b3c1f6-142c-4f3f-af8c-5d3e7362bde7 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb02969-488b-488d-8a65-8b1db26bcfe2 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6618a9d2-9863-41de-a2ce-e5512b08acca · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Demystifying clip data
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ea1aedb-6682-474a-a017-ff205f8672d1 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Show, attend and tell: Neural image caption gen- eration with visual attention
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f74dd6b-fb83-4269-9b00-d7cd63c16282 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15110c82-2031-4011-877f-ca696559a30a · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Sigmoid loss for language image pre-training
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32d4ff18-a9f5-40b9-baf0-6c7177c5d672 · outbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Dreamlip: Language- image pre-training with long captions
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d54b4bd1-d104-473b-9408-1017bb394642 · inbound
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ceff4272-2fc6-4feb-aacd-0a7e9e4d51ac · inbound
EXAONE 4.5 Technical Report OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2aebc1d-71bf-4d4c-9d10-796e6daf5d92 · inbound
Let ViT Speak: Generative Language-Image Pre-training OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 872b850a-b17a-41ea-a8c2-cce065bdfcc7 · inbound
Let ViT Speak: Generative Language-Image Pre-training OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.