Pith. sign in

Paper Citation Record · LEDGER

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks

As of 21 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 2 inbound Pith citation observations for arXiv:2505.12884.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12884 v2

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:29:00.857489Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:21:57.229717Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T08:06:48.286226Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 17423ce2-1c98-46ea-bc8e-c625dc1f694d · outbound

This paper cites VQA: Visual Question Answering.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks VQA: Visual Question Answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.619046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.619046Z digest=sha256:d89dd6d90e4b231f200f9d80e39dd27921a18c1f3e6c330860038d458ad2cf4f

Observation 8d0877ac-da7a-43fa-98de-8d171c4f7b55 · outbound

This paper cites Qwen2.5-vl technical report,.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Qwen2.5-vl technical report,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.625655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.625655Z digest=sha256:dc9b5e89d46b8dace7394f9a20f36332e672cce1e0d8961b261a29bbd6f648ea

Observation bcfed123-2fa0-4f4e-bbd7-6477c3840119 · outbound

This paper cites Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.635210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.635210Z digest=sha256:a2541e93b89f3b25cff28c868d29dd066076b36961c758335e4d2ee878a33d19

Observation d9fa5af5-c52d-480a-8a9d-e0567dd8aef1 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.640172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.640172Z digest=sha256:691aaaf4149977494df55222acf78a89691d2ee3a6bb345b49f888d992122487

Observation c997b720-00aa-4e76-9925-05d62ceae7a3 · outbound

This paper cites MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.645102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.645102Z digest=sha256:502dcc08c18689734096a64a1b3629f73e514f9c680d228b1180fb23b5640cbb

Observation db5099d1-eb24-4893-9da6-e7a368e8a19f · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.649803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.649803Z digest=sha256:f863087fe9e021f5f13e761caf972b18326e8109117a36d4f682c02fb44fbf0c

Observation f06f3a30-c59a-47f0-8b96-75633590212f · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.654472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.654472Z digest=sha256:2240a8cd1a2fdab30dbce2fde9b97fc36e13c5e4702bc4d070a2094f84e6376a

Observation a037dd18-e019-4312-93dd-6f16cbe51263 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.659368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.659368Z digest=sha256:43faa6b47d1d13c1575ce2513c6382489f8ec27621a3a2fa42aee2fc774cb612

Observation 17d0cfc0-9276-4d20-a138-1a11debebd1f · outbound

This paper cites Mobilevlm : A fast, strong and open vision language assistant for mobile devices, 2023.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Mobilevlm : A fast, strong and open vision language assistant for mobile devices, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:01.765899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:00.663733Z digest=sha256:6b2636f44b534ce20c12e7bf2310392cb934d1f8ea90496cd52a3a20b56386f3

Observation 101b1fda-e4a3-45f6-830c-d9d6c6e6a498 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.667908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.667908Z digest=sha256:c1e7833fd1a9eb39094e49ed970b3a3312ba1a831cb7a06574ef9149ce675df2

Observation 3caa645b-1b7b-4081-bfb4-d1166f4f2118 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.672424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.672424Z digest=sha256:9410bd55f8a443d2b1fd999e2ac27ebef42204afd293bd62ce2c991b3f83e70c

Observation 414db313-0a32-49d8-a669-26395d368f6f · outbound

This paper cites Gemini 2.5 pro.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Gemini 2.5 pro

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:01.751247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:00.676928Z digest=sha256:3ac46e63a9a2bcbf486455a38a7295404af5ec2081e45157d73b12af244f0135

Observation d6a7ea5e-b73f-4d0f-a3fc-efdfc9442a08 · outbound

This paper cites Textbooks Are All You Need.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Textbooks Are All You Need

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.681459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.681459Z digest=sha256:e4b7455e45ece31b8a0f5c9bd9de4b6852f09ae00350917a3d285c55059c7ba8

Observation 78006af3-1525-4be9-8d31-d35166d21960 · outbound

This paper cites REALM: Retrieval-Augmented Language Model Pre-Training.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks REALM: Retrieval-Augmented Language Model Pre-Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.685893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.685893Z digest=sha256:dddd8d375efdf3d66566b29e6b4f4d45c65ba14a47cad7ee2834b8dabd912da2

Observation 8f8585da-60de-4eda-ba67-4194516204cc · outbound

This paper cites REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.690357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.690357Z digest=sha256:3e6cf9987257f08b669cc85b075ee687fcdf8b49dfa1a949aec531acbd9464b0

Observation 99d990e4-5bd0-415d-8b23-e28e0f3b4ca3 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.694735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.694735Z digest=sha256:9e487dfe9f05977311ed984476c6fb7353d7937ff5c7db0d9dcbf62ee9c0c500

Observation d7c724d8-f52b-4e25-b3d6-1a91f9b97cbe · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.699178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.699178Z digest=sha256:76a02079ff8b94e65b00aac68a554887f5e60ab7001f134c19c9b2e3459ee700

Observation f9cf1c9b-9f00-4669-b44e-663f1c5293ec · outbound

This paper cites Shamma, Michael S.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Shamma, Michael S

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.703771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.703771Z digest=sha256:38169f37f614f68886f488f0d5a55f594c7e279e6c856a1b7b47fc0159d702c7

Observation 853b2d3a-5d30-45ca-8faf-b93b928485f0 · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.713032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.713032Z digest=sha256:77437fec332755dd771f7f8f944f9c0a3b7be39e089989a52caca6e1231968ed

Observation dd7b7c76-4726-4ae3-ad80-7658c70ac71d · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.717547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.717547Z digest=sha256:b2a1a029886414aa5ea8410d6612bb76d2bbe0d5276ba30f3e271eb7eb70c019

Observation 764c23cc-2f56-47b9-b429-e88c47322572 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Evaluating Object Hallucination in Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.721878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.721878Z digest=sha256:a1c11af5a087ad1787c99ac05c4305b41147592090ac889e1c7b35f6fb7a4a20

Observation d72b8153-57be-44dc-84c1-dc9d15de91d2 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Microsoft COCO: Common Objects in Context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.726221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.726221Z digest=sha256:85e3be694b8dc5f078065e67848e4da346416b3d328d53fd35114a1a8b4c59f4

Observation e79af00c-92e7-4e2b-9aab-c4ec141b1565 · outbound

This paper cites Visual Instruction Tuning.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.730662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.730662Z digest=sha256:2ecd0d0f65ab8390640e5bb7e47d40aadc743bc29af959f22bbef170ce3e6126

Observation 879c424e-e4c7-45a4-a6e1-0adc31de3757 · outbound

This paper cites Point- wise mutual information as a performance gauge for retrieval-augmented generation, 2025.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Point- wise mutual information as a performance gauge for retrieval-augmented generation, 2025

Reference 24

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:29:01.368487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:00.735148Z digest=sha256:02e600659780d97b52de0f9856f7ba1972410fd488643278937c734b7b504c92

Observation 72fe669f-d46a-4d72-a2d6-9d16ff34f469 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.739181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.739181Z digest=sha256:f5c747861febfd57de28395dfa6cbdb4198d5ad69b02e16e3e339ff2019e9385

Observation 85aa9fa2-cfed-4aa6-99dd-34df8b582aae · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.743702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.743702Z digest=sha256:75f7480e41ac087b08b4b9b8478801e1752850dd3d7a7d5ce3f81e0f0e8ba5ef

Observation 759b2e2e-f58f-4c2d-be5c-1452ae44d473 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks SmolVLM: Redefining small and efficient multimodal models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.748073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.748073Z digest=sha256:33b42fd211a74a1c5ea8c9cdac13a534c81fdb424d9375f5c649c27fad7f0fd3

Observation 9fdae919-19ba-4e1b-abe4-2ab4e59c52b7 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Ocr-vqa: Visual question answering by reading text in images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.753362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.753362Z digest=sha256:59e075508b3ed394bbbc7a020d0563d5bce4b1e6c039429b7297621821afe3d0

Observation d539910f-4a09-4521-a15c-47d3136685de · outbound

This paper cites Gpt-4v(ision).

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Gpt-4v(ision)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:01.719005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:00.757644Z digest=sha256:277d5964cf5a9e6182900cbd16adb12cc72d98eb50d2ad729149e07d775f91c3

Observation 5ef47174-c9a0-447f-a995-ff7392f94ef9 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Learning Transferable Visual Models From Natural Language Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.761875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.761875Z digest=sha256:3f0866de15feef7ae18a867b8354912f34a4cd33d007fc098e4e5222de414115

Observation 6c30541f-9455-485b-b606-ccaf926af4d4 · outbound

This paper cites RAVEN: Multitask Retrieval Augmented Vision-Language Learning.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks RAVEN: Multitask Retrieval Augmented Vision-Language Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.766553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.766553Z digest=sha256:eaa124a20218e7b3fc44dad91722d8c66538573688e89fb76413238d31681e54

Observation be98cfda-dbbc-475b-8a6e-c0e5178bbff5 · outbound

This paper cites Towards VQA Models That Can Read.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Towards VQA Models That Can Read

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.770958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.770958Z digest=sha256:59888c04a4c2551516c5ccefc817cf46ec8aaef2faab9f5ef7d7b3f6dae33430

Observation 3daacc28-d331-4ac0-b8bc-c2aa4291442e · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.775359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.775359Z digest=sha256:e414c292b4f09790555cb4cb8befd694bf2b650cbc86b8049c91fe0b6fc1c4a8

Observation c3221ba2-7151-4036-8104-4cd6a51fa574 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Gemma 2: Improving Open Language Models at a Practical Size

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.779952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.779952Z digest=sha256:5b9b0315d53648bacbf3c02a197c2f0bd209a337e3c7ff406be1a1cfeed89b13

Observation ec7adb78-2727-4cd2-9c35-50dfbccecd1f · outbound

This paper cites EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.784329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.784329Z digest=sha256:e09ca8eed559cf48379f711a9e68e9e79b489dab337c73492e7c2a816a9112e4

Observation cee69c63-a57a-489d-8135-9eac5b7e8e90 · outbound

This paper cites Qwen2 Technical Report.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Qwen2 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.788737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.788737Z digest=sha256:604f564be8d9d7fb72feb18e0e29685990a00d40619a5b727d71761c2cfe4dba

Observation b86a3b35-185a-4c2c-ae1f-3f5ce8acc723 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.793040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.793040Z digest=sha256:dc1773690ef5f035283bc7879776c3f171cd75cad7ae2ad066b92e0481bcf0f5

Observation 9f60531d-8572-426c-85da-3cecabbd1a48 · outbound

This paper cites Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.797269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.797269Z digest=sha256:724f1f73e0c6940acd097b2a795a30e03779c5f346630337666e7eed8b15e849

Observation da051eb0-143d-4b78-bede-60924f29f9d2 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.802227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.802227Z digest=sha256:53d90fa8f54ad3c929ea72dc384311fe7a9a912f57f1397fdb95a0183ff5b9c0

Observation 712c6821-28db-4cc7-b483-c60dca22bc09 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities,.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Mm-vet: Evaluating large multimodal models for integrated capabilities,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.807099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.807099Z digest=sha256:803113bc0f11689a0a2f502bdbff33e04851a0c984a8b210322d5ebd0babf8ab

Observation fc9c2421-bdfc-4a44-8d2c-c23003f191a2 · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.816370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.816370Z digest=sha256:a93e19337cbd8819c927bf534bf942a6647cc2be305463113aae7d576c4e55c7

Observation 457d5fda-bf1a-4599-8837-7db06a02aed1 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.821907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.821907Z digest=sha256:688ab71c36a1257e32eff35f050a917931ceefdce473e4a4d67b5f2295613daf

Observation 98bed36d-456d-43b7-b39b-f44754420ebd · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Sigmoid Loss for Language Image Pre-Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.826402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.826402Z digest=sha256:a891d6f62942492149be2e1773e7f65a6c804ef88580a09b4835987c330df72a

Observation 64a5700f-aa9e-4f8a-ab4c-24ed3fb7e4b8 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks TinyLlama: An Open-Source Small Language Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.830987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.830987Z digest=sha256:958faf8619be27ad1a22d824d87a0032e16797d03931736030fc777cfdfa1103

Observation ded0f26c-3587-4a35-8141-f6dcf6067529 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.835548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.835548Z digest=sha256:95672038f757ebd5b21a1f357ba8f8adff61e987856d7559cca0721eaa1569ae

Observation bc2613ff-a37e-4026-9bb3-3bfc7e3bd045 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.839816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.839816Z digest=sha256:a16d10ef6a3eef3fe1774580f1976f17036cdfc451d63e1398fa45c84ebfcdb3

Observation d6466768-7743-4596-92a9-18ce65fe1979 · outbound

This paper cites An information bottleneck perspective for effective noise filtering on retrieval-augmented generation, 2024.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks An information bottleneck perspective for effective noise filtering on retrieval-augmented generation, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.844243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.844243Z digest=sha256:e1e8364428def0fea6097394bd2c5bfbf6a26fcd5bae6e660cfd0c30876c4a0c

Observation d4014949-d4e8-4044-850f-cf6f2fd72180 · outbound

This paper cites These represent visual information conditioned for the language model.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks These represent visual information conditioned for the language model

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:01.695970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:00.848639Z digest=sha256:9946c27d4a56d9c48a70a0c8f8cccd1f30f84d4c558d0995d6485bb04438bb2e

Observation 36e726cb-f81e-4f96-8a6d-e420dfa1ceef · outbound

This paper cites This 2D representation facilitates direct scatter plot visualization.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks This 2D representation facilitates direct scatter plot visualization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:01.681750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:00.853070Z digest=sha256:2f6398c53737d50e666f6b4085dc5fa1441c4bff43b41a5e69a41baea7c2b2ad

Observation 2a049c74-ba2b-4755-b3c6-a872dc2f4962 · outbound

This paper cites collapse.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks collapse

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:29:01.666764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:29:00.857489Z digest=sha256:61c3ed4c2b2dd68522772bd94032c23ec6c73614848770fe0df9c8ed799d28cd

Observation 52b124eb-2aef-4505-bf89-483acc3aa074 · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.708324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.708324Z digest=sha256:f7dcd9a09b624e2beffe39d6a3f3db8f66b8f75e4caca4eef1e73834fb3ce088

Observation e8f784a9-3a21-41e6-9ef8-5ee0c5934deb · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.811211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.811211Z digest=sha256:50315edd3765a115f92a441cf21e870c29a1b06c15da2e62e6ea94862f29afda

Observation b19ecd5e-7791-4e40-abc4-6c62ef150b4f · outbound

This paper cites Qwen2.5-VL Technical Report.

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:00.630220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:00.630220Z digest=sha256:fa468661c41e20668ebdef58562b4ec0f19cc1ec62ffe661194d8016ebec33c5

Pith citing papers

Observation 91f75623-e115-441a-a03a-7c14efbec075 · inbound

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading cites this paper.

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:16:26.486082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T11:45:57.291112Z digest=sha256:36f38d62cfd9eeba67e4fef940bee7ddb7cecdb8aec201440e3ed1a4ee0b2a56

Observation 399f5aae-6630-4ac6-8a9d-6f7ef1c7b06b · inbound

MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models cites this paper.

MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:06:48.287812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T06:21:57.229717Z digest=sha256:e86a5c2b66b5ec77d3db5459aa8acdddba16cf67aea3355107b11ffff2974611