Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:29:00.857489Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 2 inbound Pith citation observations for arXiv:2505.12884.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:29:00.857489Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T06:21:57.229717Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T08:06:48.286226Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 17423ce2-1c98-46ea-bc8e-c625dc1f694d · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks VQA: Visual Question Answering
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d0877ac-da7a-43fa-98de-8d171c4f7b55 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Qwen2.5-vl technical report,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcfed123-2fa0-4f4e-bbd7-6477c3840119 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9fa5af5-c52d-480a-8a9d-e0567dd8aef1 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c997b720-00aa-4e76-9925-05d62ceae7a3 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5099d1-eb24-4893-9da6-e7a368e8a19f · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f06f3a30-c59a-47f0-8b96-75633590212f · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a037dd18-e019-4312-93dd-6f16cbe51263 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Gonzalez, Ion Stoica, and Eric P
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17d0cfc0-9276-4d20-a138-1a11debebd1f · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Mobilevlm : A fast, strong and open vision language assistant for mobile devices, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 101b1fda-e4a3-45f6-830c-d9d6c6e6a498 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3caa645b-1b7b-4081-bfb4-d1166f4f2118 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 414db313-0a32-49d8-a669-26395d368f6f · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Gemini 2.5 pro
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d6a7ea5e-b73f-4d0f-a3fc-efdfc9442a08 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Textbooks Are All You Need
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78006af3-1525-4be9-8d31-d35166d21960 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks REALM: Retrieval-Augmented Language Model Pre-Training
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f8585da-60de-4eda-ba67-4194516204cc · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d990e4-5bd0-415d-8b23-e28e0f3b4ca3 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c724d8-f52b-4e25-b3d6-1a91f9b97cbe · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Perceiver IO: A General Architecture for Structured Inputs & Outputs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9cf1c9b-9f00-4669-b44e-663f1c5293ec · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Shamma, Michael S
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853b2d3a-5d30-45ca-8faf-b93b928485f0 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd7b7c76-4726-4ae3-ad80-7658c70ac71d · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 764c23cc-2f56-47b9-b429-e88c47322572 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Evaluating Object Hallucination in Large Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72b8153-57be-44dc-84c1-dc9d15de91d2 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Microsoft COCO: Common Objects in Context
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e79af00c-92e7-4e2b-9aab-c4ec141b1565 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Visual Instruction Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879c424e-e4c7-45a4-a6e1-0adc31de3757 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Point- wise mutual information as a performance gauge for retrieval-augmented generation, 2025
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 72fe669f-d46a-4d72-a2d6-9d16ff34f469 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85aa9fa2-cfed-4aa6-99dd-34df8b582aae · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759b2e2e-f58f-4c2d-be5c-1452ae44d473 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks SmolVLM: Redefining small and efficient multimodal models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fdae919-19ba-4e1b-abe4-2ab4e59c52b7 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Ocr-vqa: Visual question answering by reading text in images
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d539910f-4a09-4521-a15c-47d3136685de · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Gpt-4v(ision)
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5ef47174-c9a0-447f-a995-ff7392f94ef9 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Learning Transferable Visual Models From Natural Language Supervision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c30541f-9455-485b-b606-ccaf926af4d4 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks RAVEN: Multitask Retrieval Augmented Vision-Language Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be98cfda-dbbc-475b-8a6e-c0e5178bbff5 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Towards VQA Models That Can Read
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3daacc28-d331-4ac0-b8bc-c2aa4291442e · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks PaliGemma 2: A Family of Versatile VLMs for Transfer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3221ba2-7151-4036-8104-4cd6a51fa574 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Gemma 2: Improving Open Language Models at a Practical Size
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec7adb78-2727-4cd2-9c35-50dfbccecd1f · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee69c63-a57a-489d-8135-9eac5b7e8e90 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Qwen2 Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86a3b35-185a-4c2c-ae1f-3f5ce8acc723 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f60531d-8572-426c-85da-3cecabbd1a48 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da051eb0-143d-4b78-bede-60924f29f9d2 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712c6821-28db-4cc7-b483-c60dca22bc09 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Mm-vet: Evaluating large multimodal models for integrated capabilities,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9c2421-bdfc-4a44-8d2c-c23003f191a2 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457d5fda-bf1a-4599-8837-7db06a02aed1 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98bed36d-456d-43b7-b39b-f44754420ebd · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Sigmoid Loss for Language Image Pre-Training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a5700f-aa9e-4f8a-ab4c-24ed3fb7e4b8 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks TinyLlama: An Open-Source Small Language Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ded0f26c-3587-4a35-8141-f6dcf6067529 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc2613ff-a37e-4026-9bb3-3bfc7e3bd045 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6466768-7743-4596-92a9-18ce65fe1979 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks An information bottleneck perspective for effective noise filtering on retrieval-augmented generation, 2024
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4014949-d4e8-4044-850f-cf6f2fd72180 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks These represent visual information conditioned for the language model
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 36e726cb-f81e-4f96-8a6d-e420dfa1ceef · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks This 2D representation facilitates direct scatter plot visualization
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2a049c74-ba2b-4755-b3c6-a872dc2f4962 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks collapse
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 52b124eb-2aef-4505-bf89-483acc3aa074 · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8f784a9-3a21-41e6-9ef8-5ee0c5934deb · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19ecd5e-7791-4e40-abc4-6c62ef150b4f · outbound
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Qwen2.5-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f75623-e115-441a-a03a-7c14efbec075 · inbound
State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 399f5aae-6630-4ac6-8a9d-6f7ef1c7b06b · inbound
MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.