Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:09.476244Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 3 inbound Pith citation observations for arXiv:2505.17316.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:09.476244Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T04:35:49.701061Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
52 of 52 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation fadc0ca7-aeda-4650-b84f-c888ac9a3105 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Visual instruction tuning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 29de9b4f-7a6c-45b0-ba66-65863a79d134 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Improved baselines with visual instruction tuning,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 208045f1-7904-41c4-a316-77bcd68edc2c · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70584995-12c7-48a0-89ce-68f725df6f8d · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db061ba-bb92-4291-80e2-ec1c3ab061f2 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98609cdd-071d-48fc-a888-967810b5d0dd · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Flamingo: a visual language model for few-shot learning,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc4869a1-3879-4476-9c1f-272048cb17a6 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Language is not all you need: Aligning perception with language models,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47117c81-1b56-4a5b-aae9-9bfb30599547 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Multimodal Chain-of-Thought Reasoning in Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09fdddc4-ddbc-4e47-8c00-b13a5032e0aa · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Kimi k1.5: Scaling reinforcement learning with llms,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 61a44add-9445-4537-aaa6-470852bb6cc7 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Multimodal transformer with multi-view visual representation for image captioning,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 12569a1d-b954-4dba-951a-aa8085cbd105 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Vqa: Visual question answering,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bb64434-aec4-491f-8c0d-89bf212f3a5b · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc12dd15-7ac6-46de-a797-54b4ce1bddf1 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53446d3c-8f1f-4f37-b759-51d558bf7049 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29156912-f77a-44c8-ac02-3ef445607f04 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Glamm: Pixel grounding large multimodal model,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 099610fd-b9dd-44de-855d-e0271fb59288 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Hallucination of Multimodal Large Language Models: A Survey
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d751b048-6e58-4c5e-9b9d-f1acb8e9589a · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Honeybee: Locality-enhanced projector for multi- modal llm,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3bcaf127-cac2-4f33-aa77-d8a119ec9186 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models The Platonic Representation Hypothesis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29f11f75-9d76-409a-8d86-51f6192b84f6 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models von Neumann,Mathematische Grundlagen der Quantenmechanik
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1cd1b476-c3c0-49af-a2cf-202ca427a365 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Linear algebraic structure of word senses, with applications to polysemy,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ced2b0b2-19d7-456f-875c-eeb48feff724 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba7c44c2-883e-48e9-96ee-9444c5df2e11 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Signal recovery from random measurements via orthogonal matching pursuit,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc3cf9d0-2f91-4990-82cc-56bd5650f420 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Recognize anything: A strong image tagging model,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8fedc27c-4a1b-4068-8304-468498f53fef · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 69cb3100-d9ea-4117-95a9-8016afa821b8 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Segment anything,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5cf6810-da43-4cd8-81d0-acd914e552a2 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a3b850-f28d-414e-a021-fcf65bd986ea · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61da63f-e94a-4644-8fe6-02badd35fdaa · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a07487ff-fff9-47ec-8cc0-69e5657edb1b · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294bd963-4843-4895-a63b-ea8567de710b · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 410e58b5-1c87-450d-9d58-7086b04f6ba4 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78e1c9d-00dd-436e-bc06-0b9b921d9ed9 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Law of vision representation in mllms,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58d0ae74-b066-4fa9-8c35-3430164cc0ef · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Learning transferable visual models from natural language supervi- sion,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation afd9b67a-ee7f-4db6-a5d8-cecc6819e3b2 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5c0e1f4-08fa-4fae-a4d5-f324eebe1557 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 258e03ae-c032-4490-886e-02438a42f2c5 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f7dcb8-de70-4ef8-a31d-a938af74c5c5 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb7bfab-4a76-425f-9254-89705a2460c8 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Honeybee: Locality-enhanced projector for multimodal llm,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4e8b62b-1cb5-45b0-aee7-596158711fb4 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Microsoft coco: Common objects in context,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66925e60-7f15-4ea4-b05c-42d93478d3d6 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Referitgame: Referring to objects in photographs of natural scenes,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 766fa7c7-55c0-4cef-81ca-9e6e07e5deed · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Generation and comprehension of unambiguous object descriptions,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 519e912a-9cd7-48be-99a7-28ddc9d0707c · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b6ad84-73b1-435c-aefb-38348b595dfe · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Introducing idefics: An open reproduction of state-of-the-art visual language model,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation aebb9424-4cd8-4cc6-9db1-d85de7550613 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee107fa-bfd8-4446-bb2a-7f31b967b7e0 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Instruction Tuning with GPT-4
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da6a0ac9-3e8b-4782-b358-5e15b6c0d4cf · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models The Llama 3 Herd of Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55246c22-573a-4a8c-a38f-94e1528d6620 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 717aab0c-5678-414b-b5a8-09c0685b4fdd · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Towards vqa models that can read,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b5e92fd3-23ec-473a-ba67-f160ea063043 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9143f5af-b801-45d4-b5fa-ced2989a23a3 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Ocr-vqa: Visual question answering by reading text in images,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 894462a2-6627-4fe4-a7f7-ad0abbf5ec84 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 81ffb6bc-c061-4c86-96d1-c38e50e59701 · outbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Unresolved cited work
Reference 1932
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 016138de-0e18-46f9-a76b-1f2029ae70c9 · inbound
Latent Denoising Improves Visual Alignment in Large Multimodal Models Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9c925f13-3f9d-4834-b9d5-bd95ebda16bf · inbound
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5774306e-2e1f-44d5-b230-98bb897ca4f1 · inbound
Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.