Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:53:36.381720Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 6 inbound Pith citation observations for arXiv:2412.08443.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:53:36.381720Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:19.033317Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T10:41:04.336073Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 98f326bd-d824-46b8-8c6f-03f5a4358a38 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fffee9bc-21d1-4932-8b30-6199ea5151be · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications InternLM2 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbf6cd1-7f28-4e8e-8fd2-d03e9d4c631d · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a7755b8-4a17-44f0-ba00-f66d6d19457e · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b13e1ea-8a7d-4329-aff7-e811edf72d4b · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56929c1f-128c-4875-8c03-7304f608f8fd · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Opencompass: A universal evaluation platform for foundation models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b610b1-d779-4dda-a0fe-b16a35b5f244 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications NVLM: Open Frontier-Class Multimodal LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63705407-a9b3-45c3-871e-4abcdab7546a · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Flash A ttention-2: Faster attention with better parallelism and work partitioning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa7d6077-5c4c-40ae-aaf9-5149cf7741da · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7540e0-44f8-4006-9bdc-593e85a1777a · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a6064f6-6ffd-448e-bde2-da149ef06fc6 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aa0ed58-c860-4f81-890e-b2152312c996 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8af7fbc-f283-488f-937e-0cb0f76f67e7 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a49dc5aa-2bcc-4e25-a66d-81af93fc73fd · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b1f3a07-e8c8-4486-841f-f83e4bda3359 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Gaussian Error Linear Units (GELUs)
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e42ad58f-133b-4397-9030-1a3c2d1ad19c · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dcf57e1-9e31-480f-8549-2dd20fc0b965 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications A Diagram Is Worth A Dozen Images
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7faf391-27ab-4a72-ae2d-11ac41cfa454 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Visual information extraction in the wild: practical dataset and end-to-end solution
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9e8b26e7-b403-4696-a5c3-cb5433f133bf · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications What matters when building vision-language models?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78734d24-d5fc-4569-a79a-fd35d6b0bb42 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications LLaVA-OneVision: Easy Visual Task Transfer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e10ffa-6e9b-4907-935e-5c3a439df210 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6dc376-ad75-4176-82ef-fe4be5652c06 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66370c34-82cf-4eb5-820f-232b96f1f9b0 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 071b681f-98c4-493c-8cdc-a50cb5a3afe0 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Improved baselines with visual instruction tuning, 2023 b
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f76f56-d232-4a00-9c98-02436a04d336 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972eac2e-bf3e-49cd-834e-783b58a4a2b6 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Visual instruction tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f89b701e-93f5-4293-bae0-c0cfbdd33724 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications MMBench: Is Your Multi-modal Model an All-around Player?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a384600c-cc4b-499b-84dc-3b5cf5d34d46 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Rethinking Overlooked Aspects in Vision-Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7951d8c8-f64f-4ccf-8f2b-5775eec7cce1 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications POINTS: Improving Your Vision-language Model with Affordable Strategies
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50e3ba2d-8416-4470-b73b-fd21746959c4 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f1136b-ad4d-4994-853d-951ba864696f · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications A convnet for the 2020s
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89baec70-f09e-44b5-93b8-14e1d436eaf7 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b327698e-fac0-4c8c-9161-75f40f9c3ca5 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cd4dffa-a152-49c7-bf65-25aef7937ab0 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e226fb37-2f60-4d15-a93c-75819721536c · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbfbed3c-8f3a-4141-b704-b929aba53ed1 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5ad0b438-50a9-4792-af7c-3f1d36dec613 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e8460d5-5ad4-421a-ac79-21eb1a002e51 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Gpt-4 technical report
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fe849fcd-952f-4170-ab54-5f5683f44c81 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f64f42bd-cfde-4509-8d8b-190a3f9fcf55 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Neural Machine Translation of Rare Words with Subword Units
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31dbbec2-e0c7-409c-964f-eb7c2e654e8b · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Fast WordPiece Tokenization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f9a65c3-c770-4e2f-81ef-a00b446ac23d · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Attention is all you need
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14f66df-556f-4525-b716-5a3e9cc4b655 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6318f71a-845c-4003-a7c7-672af8b109eb · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8fc23f-4cbb-4437-b08e-bff5cf8c03ac · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50bf95f2-032e-4064-a534-321dc9d7cfd6 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Emu3: Next-Token Prediction is All You Need
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8b086a8-c3ea-4660-859b-9ba04b5f2c90 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad075e0-148a-4e0e-8a89-62ddac381dfa · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Qwen2 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87c14d9-e5a4-4bf8-8b4d-0a8404ebbba0 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af0f15c6-49b2-4c09-9a44-03c60e20f07b · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications A Survey on Multimodal Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea851cd-f6f8-4c30-92d2-e3df83e2fb48 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Capsfusion: Rethinking image-text data at scale
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 97b5a4a0-ffdf-47c8-8113-9a7a11592997 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48a51f39-2a3d-4042-8e4f-a36e24b4bfc3 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c05cdad-49c4-4650-bc8f-06c059e545cc · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9fabc9-1e7e-4575-8c0a-7e6aa616fac1 · outbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d6fcd753-8a4f-4a5f-ad9b-50ec0fa983de · inbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design POINTS1.5: Building a Vision-Language Model towards Real World Applications
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8995f9b6-2354-4758-aae2-58978eb21bde · inbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model POINTS1.5: Building a Vision-Language Model towards Real World Applications
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ebea4e1-dbe9-44cd-aa48-b66c76f329cc · inbound
Ocean-OCR: Towards General OCR Application via a Vision-Language Model POINTS1.5: Building a Vision-Language Model towards Real World Applications
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f44d6b-23d1-46ab-b8a5-1dc68c20ceb4 · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs POINTS1.5: Building a Vision-Language Model towards Real World Applications
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 978fed47-5263-4088-8de9-d7612181ca27 · inbound
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management POINTS1.5: Building a Vision-Language Model towards Real World Applications
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8bf490b1-1558-4081-a933-75d550175b34 · inbound
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management POINTS1.5: Building a Vision-Language Model towards Real World Applications
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.