Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T19:39:48.245781Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2508.12263.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T19:39:48.245781Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T15:35:37.095627Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T10:11:09.424972Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c6aea20c-74fd-41b4-953c-32687809851b · outbound
Region-Level Context-Aware Multimodal Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df1ea6a-cf53-4abe-8f61-2d24ae8de02a · outbound
Region-Level Context-Aware Multimodal Understanding Flamingo: a Visual Language Model for Few-Shot Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8424158f-e0c5-4b3d-9f5f-c5aab1d835ac · outbound
Region-Level Context-Aware Multimodal Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a73e7ef-8b2c-40c4-b159-6154601ab78b · outbound
Region-Level Context-Aware Multimodal Understanding InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec805674-2e99-49b6-8ff8-56e581a3b0f5 · outbound
Region-Level Context-Aware Multimodal Understanding DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b6ccda-328f-47e0-895c-d34039522811 · outbound
Region-Level Context-Aware Multimodal Understanding MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0acddf-0452-4542-8916-34dbb5c29744 · outbound
Region-Level Context-Aware Multimodal Understanding MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16e4f698-fc0d-4b37-80d7-dc44182c46ea · outbound
Region-Level Context-Aware Multimodal Understanding CaMML: Context-Aware Multimodal Learner for Large Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a816166-f40f-4779-b0a8-40d5702a3811 · outbound
Region-Level Context-Aware Multimodal Understanding Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 064f0dad-04c1-40ce-a4be-87c156d73b20 · outbound
Region-Level Context-Aware Multimodal Understanding Referitgame: Referring to objects in photographs of natural scenes,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7da6b354-52cc-4587-a071-ef4c2aebda7c · outbound
Region-Level Context-Aware Multimodal Understanding Panoptic scene graph generation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f216edd-f8c6-4466-954b-3e5636654f35 · outbound
Region-Level Context-Aware Multimodal Understanding Glamm: Pixel grounding large multimodal model,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2157a260-b549-480a-8763-0fa70a8c963e · outbound
Region-Level Context-Aware Multimodal Understanding Bleu: a method for automatic evaluation of machine translation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56ae1682-2407-4411-8588-67915790807f · outbound
Region-Level Context-Aware Multimodal Understanding Rouge: A package for automatic evaluation of summaries,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d28d5f7c-fb8d-454a-a840-b822c9a5e598 · outbound
Region-Level Context-Aware Multimodal Understanding Cider: Consensus- based image description evaluation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a41a9f92-bcd7-46f8-a4dd-f57476b6cbe9 · outbound
Region-Level Context-Aware Multimodal Understanding Spice: Semantic propositional image caption evaluation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba4e72e0-11cf-4a11-a933-129254c97da0 · outbound
Region-Level Context-Aware Multimodal Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a70b2165-d355-46ca-91c3-e18b54b1079e · outbound
Region-Level Context-Aware Multimodal Understanding Visual Instruction Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 069b9a4b-b512-452b-ab20-7cb3fc479039 · outbound
Region-Level Context-Aware Multimodal Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e50f82b-d859-4e4a-b256-471307846976 · outbound
Region-Level Context-Aware Multimodal Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a10a371-7cdf-43c3-bbfa-86c72a2a00cc · outbound
Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 256390509
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4621501-30fc-44ef-897a-3d9c6edede2f · outbound
Region-Level Context-Aware Multimodal Understanding Mmict: Boosting multi-modal fine-tuning with in-context examples,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5ed4d786-02e6-4e98-bbb7-f525b886ef24 · outbound
Region-Level Context-Aware Multimodal Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac638776-fc03-48f3-85c0-9b918924fde1 · outbound
Region-Level Context-Aware Multimodal Understanding Cantor: Inspiring multimodal chain- of-thought of mllm,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a2cff7f-f92a-4cfd-9d31-a79215644a5a · outbound
Region-Level Context-Aware Multimodal Understanding INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2fcc1d4d-6c3a-42db-a9a9-0668919903c7 · outbound
Region-Level Context-Aware Multimodal Understanding Video-llama: An instruction-tuned audio-visual language model for video understanding,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf21f313-22a6-403a-b487-4054478e175c · outbound
Region-Level Context-Aware Multimodal Understanding Video-rag: Visually-aligned retrieval-augmented long video comprehension,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec322ad7-3d90-4a70-ac13-c2e78f070e97 · outbound
Region-Level Context-Aware Multimodal Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4866031-3d7d-46ce-968e-a6d1cb721dd1 · outbound
Region-Level Context-Aware Multimodal Understanding Video-llava: Learning united visual representation by alignment before projection,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c0506359-bb7d-41f3-a44d-c2e962694888 · outbound
Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 265281544
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66411ede-fb32-46f8-9fc1-6074fe3a8c78 · outbound
Region-Level Context-Aware Multimodal Understanding Manipllm: Embodied multimodal large language model for object-centric robotic manipulation,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d7eda24-ed12-4dfe-90c5-206349a46147 · outbound
Region-Level Context-Aware Multimodal Understanding Rap: Retrieval-augmented personalization for multimodal large language models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 863d828b-bcf9-4ba0-af2f-ac13b716c004 · outbound
Region-Level Context-Aware Multimodal Understanding Yo’llava: Your personalized language and vision assistant,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db9dc312-26fa-4102-9fae-f82b7ab730a3 · outbound
Region-Level Context-Aware Multimodal Understanding Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2631a3bc-7b7e-4b32-bb7d-bb594ec41218 · outbound
Region-Level Context-Aware Multimodal Understanding Palm-e: An embodied multimodal language model,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36873bb4-dd1a-45ee-940a-2d65d5c98aa0 · outbound
Region-Level Context-Aware Multimodal Understanding Llm2clip: Powerful language model unlock richer visual representation,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ae8dd8e-75a3-4d8d-9bc0-668bf4a658cd · outbound
Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 266573457
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f8819b5-6c47-44eb-a65c-fbf2719e09ef · outbound
Region-Level Context-Aware Multimodal Understanding GPT-4o System Card
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea84cffa-cf95-4d85-b959-fcb037b10bb2 · outbound
Region-Level Context-Aware Multimodal Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06819efa-51cd-4240-b180-13c4f8121ed3 · outbound
Region-Level Context-Aware Multimodal Understanding Myvlm: Personalizing vlms for user-specific queries,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e8108fc-7083-4939-85cd-9dcfe042a13d · outbound
Region-Level Context-Aware Multimodal Understanding CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b1a083-761c-416a-9ece-e9b080dc6912 · outbound
Region-Level Context-Aware Multimodal Understanding DeepSeek-V3 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aacd94b3-7518-4400-b327-feb67497b463 · outbound
Region-Level Context-Aware Multimodal Understanding Gemini 2.0: A new era of multimodal models,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a31e000-7b9a-412b-9ce5-ab7d02769624 · outbound
Region-Level Context-Aware Multimodal Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7caa56f4-8a90-467c-939b-7101d79bffb0 · outbound
Region-Level Context-Aware Multimodal Understanding Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f8d553e-db90-487d-b911-114ff3e7f983 · outbound
Region-Level Context-Aware Multimodal Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb47bbe7-f3a4-47da-b8e7-ab3b418008ce · outbound
Region-Level Context-Aware Multimodal Understanding Enabling Large Language Models to Generate Text with Citations
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be0e5648-70cf-4467-88ea-e2d6934d2848 · outbound
Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 248476411
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af0d61fe-6251-47e2-b377-ccfecb726f3d · outbound
Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 258615266
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eba9c55e-cb75-44ff-adae-69bb5301e5f3 · outbound
Region-Level Context-Aware Multimodal Understanding Available: https://api.semanticscholar.org/CorpusID: 266844925
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9865e114-cb2d-446a-9799-6fb45b263c7f · inbound
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Region-Level Context-Aware Multimodal Understanding
Reference 183
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.