Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T04:09:36.019146Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 41 inbound Pith citation observations for arXiv:2403.09611.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T04:09:36.019146Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:03:25.101268Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 137 outbound references displayed
External citation measurements
11
pith, observed 2026-08-05T02:28:24.338817Z
Observation 52984cd2-2a08-4832-8309-08850565750d · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff5ae10a-e89a-4ee7-a990-5b5971cbec47 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICCV (2019)
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f98930c-b76d-4ffd-af2b-2078e34f6aab · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31f07fd6-8f3e-49ba-9eed-7e8a91a09192 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e91d606a-c567-4ecb-9f5e-7ae270836808 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52e6c210-4f2d-4713-84c1-1bcd47a01b8e · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: EMNLP (2013)
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f1acd2c-a893-421b-9128-92049224e440 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training AAAI (2020)
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e079e4cb-7c98-495b-8c24-8982cc7f0483 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Training Diffusion Models with Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 970824c6-f955-4db4-9983-7c6969419393 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training On the Opportunities and Risks of Foundation Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37cee271-5c8f-466c-95a1-258b6d9700d2 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training NeurIPS (2020)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 671852e6-e722-4d51-bdf2-25cfe45285fd · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training https://github.com/kakaobrain/coyo-dataset (2022)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7bdbda1d-41da-4d34-8bf9-82923ec64273 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Honeybee: Locality-enhanced Projector for Multimodal LLM
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42425d51-6414-454d-84a8-ec8c278ea8ee · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2021)
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3fc9b4e-5f33-4700-8d05-4bb087fed4f7 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e2c4f3b-0dc4-4d48-9875-ca59ea3711f1 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5e08aa1-5855-456e-9fed-a694872c8e5b · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICCV (2023)
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28579ef0-fbde-4206-a8f7-2108b1517492 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 634ff84b-25a1-48e8-8fe9-f4856ffe6db7 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5addcd36-ca0f-4d0c-913a-87b665f0e296 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training JMLR (2023)
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6f2d595-4bbe-4f80-8f1a-d7393d254b97 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08f49448-5684-42c0-b86b-f5d11a8854df · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scaling Instruction-Finetuned Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ee52c6a-7a32-48d4-8a35-5a7120a86125 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99ab4eca-51ac-4c8a-9a8f-62e3b3f9b773 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ead157dd-857d-4ad3-a04f-660a503e8b3c · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a60623f-18ef-4d72-9ccb-63e5341c61d3 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 942439c7-c3f0-438d-a349-3cdbb21074e3 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e9e65cb-c432-462a-80a8-2739a352a78c · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02ce5c85-46d4-4d83-ab83-ad289e7fc8c3 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training PaLM-E: An Embodied Multimodal Language Model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b6f2cf3-f31f-4742-aaab-f310aba4c423 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICML (2022)
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc51c86e-8447-4528-87ac-473afe76094a · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scalable Pre-training of Large Autoregressive Image Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f77a8e15-076d-4275-ad1e-8aeae032a19a · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Data Filtering Networks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7623c334-3a45-4175-ac6c-ea1bce611f53 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80c1ef7a-c135-48fa-9e2b-aa9d77b07421 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81b7d2a8-df2a-4a4d-838e-2ac8de554557 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Guiding Instruction-based Image Editing via Multimodal Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13e48b14-591c-44f2-82bc-19a8764a8f82 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training https://doi.org/10.5281/zenodo.10256836,https://zenodo
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af8ae0c2-4493-4ead-8c91-db74540d08e0 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c10f305-4288-4827-a656-462842bf09e0 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49dccf1b-219c-4baa-9507-872f9a74bfe4 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2017)
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39177b3e-2dd3-453d-be19-866453919463 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2018)
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09e96b13-8f7b-4534-955e-ea6b92dfd300 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2022)
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 948384f8-c340-4e11-819f-f329b793ebde · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2016)
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10fc5fcb-2c2c-449c-bf9d-7a35a5b434e0 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Efficient Multimodal Learning from Data-centric Perspective
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a7dad1e-d0d9-456e-91c1-11bf8ec428f5 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scaling Laws for Autoregressive Generative Modeling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0854ebc-7f93-4679-b5ac-fb21eb035aac · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aaf443ae-937c-4c18-b6b5-7042acc417bc · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8cb4cf0-c05b-4ebd-b965-5fd44d7a30fc · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2019)
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3af3a71-f461-4258-b78c-40d68e878ec3 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training https://huggingface.co/blog/idefics (2023)
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4187dedb-bbe1-475d-9bc6-3132a3ed0cbc · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6cd8749b-43f6-4a2f-9ad3-2b64a24af156 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb5610d7-05a4-4afe-91d7-69ad82224a29 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0595f8f-a6e0-475f-bc56-8139c0607170 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2018)
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51721e93-c87c-456e-9d5b-a5110834c7b1 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ECCV (2016)
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55fcfc5b-c42b-4508-ab2d-47d917163cbf · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ECCV (2022)
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 574486fd-fdde-468e-b95a-9d2b89837cd2 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Generating Images with Multimodal Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d4ca5d2-cea4-450d-ae1e-0f88769819bf · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICLR (2023) 20 B
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b836275-5181-4428-b431-2b6e90226b10 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training LISA: Reasoning Segmentation via Large Language Model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7fd15e9-a78d-4c20-b9bd-1793625f42e7 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VeCLIP: Improving CLIP Training via Visual-enriched Captions
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c8ec99a-7469-475b-b139-f4cd71a4b265 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e5fe84a-a7cc-4f80-a240-335a5e24ed15 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICLR (2021)
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88cc9aa3-66f0-45ff-8787-07577844ec3b · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38f5fa54-ffdb-4752-abd5-3518a7d1718d · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9dbfeace-473f-4087-bcf4-30b28aa6b0bd · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 125869bc-ad0a-4f6c-ae49-1192b81d123b · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Multimodal Foundation Models: From Specialists to General-Purpose Assistants
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa4f27df-5e83-4913-8cb8-e909bf4afdb8 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2246ba8-9403-43fa-9f09-e147d5e42781 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d84ef2d-5fec-424d-922f-c386ee15a067 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71e9bf56-b60f-4038-b075-f2bbf19b93c0 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b678ba0d-8536-4d27-ba67-4011e718dfa2 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Evaluating Object Hallucination in Large Vision-Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b69ddb0-f51d-49e8-ad2a-daef9811f7b1 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c684790d-403a-421b-b6b9-7a70d445709d · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e62c1d3d-15ac-4a15-ab2e-037fa7173fb2 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VILA: On Pre-training for Visual Language Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4667a88a-6cea-4f9a-9748-108961504995 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Microsoft COCO: Common Objects in Context
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9750ee8-4067-4275-8a67-a83b2d563b43 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7572d3a-e83e-4516-be1e-e996279e2ed1 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Improved Baselines with Visual Instruction Tuning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05ed3e91-c99a-4fec-a901-904bf5699223 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training io/blog/2024-01-30-llava-next/
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 003de143-d596-4f17-bc3f-430b4edce97b · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ffe80b4-f047-46b4-9bbf-997312599f30 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccdcdf43-7d9a-4726-9c80-dda1bfa04da9 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MMBench: Is Your Multi-modal Model an All-around Player?
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5cabe2d8-13c7-4438-96c3-ace57effbacc · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training NeurIPS (2019)
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73c880c6-bb6b-4a68-a0fa-272319facf30 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 871a3cd4-4e37-40f2-b055-fdc1f36fbf30 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training NeurIPS (2022)
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e76a5442-ac1e-4876-8b39-6e7a7df635bf · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2019)
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd20772e-827e-4fe0-8507-3e43805db485 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f35af081-a5c1-449f-9c66-9bf0da08ae3b · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: WACV (2022)
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30305ea0-6963-4d35-b7a0-f847132a4b3e · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: WACV (2021)
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8bf52fcc-1aa3-462a-a583-941e5644c6be · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICDAR (2019)
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01d8b1b6-8edb-492f-bc97-5338971b6e4d · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: NeurIPS (2022)
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6827c88e-c9f4-4974-9d28-a097dea10b28 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training DINOv2: Learning Robust Visual Features without Supervision
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbabfbed-a48e-485b-ad9c-e1e65b7a0f60 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 931fbc62-baa8-44de-b477-0f568596d534 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8109ac73-0b16-4720-80b7-f2af40eab767 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICML (2021)
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1bedb279-c508-4f94-b660-122cce42945d · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5123a8ea-01d0-42e5-b3bf-7100642977ba · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training JMLR (2020)
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00ea4575-ceb4-4f8a-9798-008834fa56ae · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICCV (2023)
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f945584-c92f-4092-a28f-6cda591bcfd7 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2022)
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22655654-3f06-49e7-ab3d-77bc325d3960 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f75281d8-deb7-4374-a1fc-6da4e35dfaa4 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50896e14-6c60-41c3-9b5f-547277fcfeaf · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ECCV (2022)
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c87d41e1-e7dd-4f46-99e5-8ae662f85539 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d33cb590-3fef-4e27-b114-29e0d2f12b11 · outbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ACL (2018)
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14b8c1fe-c569-4462-9952-7a6f42276d1b · inbound
A Survey on Multimodal Large Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61c632e2-a024-4ecd-969a-440c2fc20f46 · inbound
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74b488da-e850-4f47-a3e4-a542f4999c6a · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1e3ac38-2099-4fc0-a54e-1fd79265e8e8 · inbound
OpenVLA: An Open-Source Vision-Language-Action Model MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cc19cb7-f707-4a8b-abe7-bba47e7b01d8 · inbound
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cbfdc07-8249-45ed-9f01-533a10a4fe7a · inbound
PaliGemma: A versatile 3B VLM for transfer MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1fda17c4-bc48-4651-9e3e-e1407b6fa0ae · inbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e3cba1e-776f-47fa-bcee-c681ddb079c6 · inbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4bc1dab-bd26-4776-9b60-1e105abdc0bb · inbound
LLaVA-OneVision: Easy Visual Task Transfer MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b45c23d-f818-4743-b9c3-15974fff83ba · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3d25bb7-cf7d-4f64-83e8-2634643781dc · inbound
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9614d6f-21ae-4f82-95a6-fcf0f851adaa · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db67f8a3-8083-40ea-b8f4-c41282c2bd7a · inbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40b0d271-7a01-4da9-a3ec-325d39165583 · inbound
PaliGemma 2: A Family of Versatile VLMs for Transfer MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43f2bb38-b30a-4eb6-b9ae-0c82ada3946c · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ee7314c-6c4c-4dfd-99ef-63d973fe5a0a · inbound
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 172
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 010c6de0-5704-47ef-87d1-90b6876a9286 · inbound
Semantics Disentanglement and Composition for Universal Image Coding with Efficiently LLM Reasoning and Generative Diffusion MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2c3e8b0-a2cb-48c5-8adb-620909f0f6ec · inbound
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ead99e0d-d6a2-4900-aaea-d29ea3561f95 · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c45f6b8-6c9c-4d34-beed-57b63926474f · inbound
MMaDA: Multimodal Large Diffusion Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3d02633-9be4-4692-80de-134a95a7d6fe · inbound
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301aff66-ce55-45c8-91de-c03f5c645ec0 · inbound
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d246c863-0a36-4ae5-a52f-ff926dd5d1d1 · inbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e39d9be-e59f-403f-976d-c0213474736c · inbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e4003eb-1b66-4a0d-bfdf-f24097734015 · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b72702d1-76c2-49ca-ae00-13d7e5d4cdd0 · inbound
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b38f56ce-fee3-45bf-a545-1ee82014b56b · inbound
Synthetic Visual Genome MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd027e17-7925-479b-97f7-7d39f0994751 · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2280f4f3-b2b0-4be1-838b-7c667339728c · inbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26b9cebe-850f-4bff-b75f-82e4460c81a5 · inbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5b53bb2-96bb-493c-b041-ddd1971f8bb5 · inbound
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d879fd-ff47-4680-a577-3c7659f9c642 · inbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 135
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 340c3a01-ac8a-4fda-9dd5-eb288decbba2 · inbound
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb22c86a-fb1b-4d97-8a4e-4a83edbcdedb · inbound
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17f2b8d2-d76c-4729-a0dd-2ed2f6483cb5 · inbound
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e47458c-e6e4-484f-bcf0-2aff0da3f539 · inbound
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f442a76-4533-45d8-b2b6-6159d6599e5b · inbound
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df5b6639-7c56-4154-ab02-f94b25de8d52 · inbound
Compared to What? Baselines and Metrics for Counterfactual Prompting MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6c0c0b3-0151-4b69-b23d-4747d0d5f1bd · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2546657a-2e04-4ad7-84ff-2aa0687b3abb · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c4cf4c1-bc82-485c-a58d-18c69e8190a3 · inbound
An Exam for Active Observers MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.