Pith. sign in

Paper Citation Record · LEDGER

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput

As of 19 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 3 inbound Pith citation observations for arXiv:2505.09498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09498 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:35:45.733427Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:55.805940Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T17:07:27.421336Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77d02640-a53f-44f4-9eac-958d35160ca0 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.369886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.369886Z digest=sha256:4447cbd0672c32ac282536381f424a1ce4da3980db0db1f5bbd239d95400726b

Observation 9ee49ae1-0b99-4906-878e-c68b37d813e4 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.374869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.374869Z digest=sha256:afd86bc1240a4d7e588019770b66922a5a9d7059f3f26984970dbe03cb8a84b0

Observation 4ca32b92-1f96-4f51-9fd8-6dafaa56b239 · outbound

This paper cites Qwen Technical Report.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.379372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.379372Z digest=sha256:565e78fba9e9113008fb2e926e704b2d87030bd8912895015d841bbb60b2d13a

Observation 6556433b-7743-49a3-b5a7-b3988cf440f5 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.383801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.383801Z digest=sha256:69a0e78bef28fc57468755efd0d7717a78dac9c9cd3d689e134ba7ad32fdb182

Observation 3168a3d7-c33f-4cd3-8d52-59023128b02d · outbound

This paper cites Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.387916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.387916Z digest=sha256:e596e1fd3b23cb3377f041a181cd469315e7a0a707f79240007bdd85ec03a755

Observation 966db0bb-a95a-42a8-b7f0-d3b13b3b6755 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.391928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.391928Z digest=sha256:1594be528ecd034f567c4bfbf4c484f7d43ce5cac0b5323c2e1863c402fb2a80

Observation 66264194-0a48-4bfd-970e-0ddf9524add9 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.396400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.396400Z digest=sha256:f03af0b77a5a82549ad0897d3044af2007d803d86c871d42292f73fa6a6800a3

Observation 4153cdbb-d210-422e-b805-50984e9412c2 · outbound

This paper cites TabFact: A Large-scale Dataset for Table-based Fact Verification.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput TabFact: A Large-scale Dataset for Table-based Fact Verification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.401029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.401029Z digest=sha256:f3f2ac24f5bd09133e526fe95f19cb691665d5fa85ae2cf8510ba43b89604f42

Observation 84a07685-31b8-45be-a663-18ae7d4de02e · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.405373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.405373Z digest=sha256:a9657e182e4842c97d1a718f3083f7bf2a6b04c43270ffffedb18940a22e6ed7

Observation 1f2f4329-4fc6-4a60-aa3f-ec57c3f54919 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.409546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.409546Z digest=sha256:5df9d935f02432b28877c36541bcc1796f513adac7757d0434445cc1177eb7ff

Observation 50bcaa31-2d35-43f4-8bd6-2420879ab19b · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.413527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.413527Z digest=sha256:735b189aed32996292402ad1d28bfd0d15c52ac0aa8676abe741702ed7c2316a

Observation b0624d95-f765-4b31-802c-f244eed3e9f3 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.418722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.418722Z digest=sha256:fd61333e831bba08119e245940ac174245dda2dde02a3a0b5404799c6fbd5014

Observation f08a5d3d-b74c-48ba-803b-8897b421989e · outbound

This paper cites The chinese dataset distilled from deepseek-r1-671b.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput The chinese dataset distilled from deepseek-r1-671b

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.422845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.422845Z digest=sha256:46de2cf82c97fa423ca94156d9611e338386ddf35f6cee46d9913785f4c02b61

Observation 1fbe0308-9ffd-400e-8f8d-06697a7e4357 · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Scalable Vision Language Model Training via High Quality Data Curation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.426783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.426783Z digest=sha256:0b20644d3ad26ab8e8e86ffbadd5aa96e46742c29f7d86effd3192e8f8ffcee4

Observation 7598aa4b-e8f8-4a4d-a38d-0000aad386ce · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.430697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.430697Z digest=sha256:b07dd8226cf30f2accdfe818f4a389adb9965b3b0b5c38c86d2a91e4d84db6a2

Observation abe1dcda-02a8-4c68-a1aa-4639483cc40a · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.435501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.435501Z digest=sha256:9e4fcb12c0d8d3fb646e10f2144fa009f3912d967cb6671915262338e975d91a

Observation 4117b70a-d981-4836-9a08-ddf36b57bc3b · outbound

This paper cites Scalable Pre-training of Large Autoregressive Image Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Scalable Pre-training of Large Autoregressive Image Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.439263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.439263Z digest=sha256:040ae2fd1c60a27e656971c9721fa2249cba5680100b0a9bbcddaedd26316b89

Observation fd30e9ef-8ad2-4394-a2cf-81fcd45e7cdb · outbound

This paper cites Data Filtering Networks.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Data Filtering Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.443439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.443439Z digest=sha256:bed25f7a9d11331628744a19342ae807108b5f99ec6c38d1d0ea43077a2fae38

Observation 5ff1c0eb-cc4e-4745-9acc-859ebcf812c4 · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.447426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.447426Z digest=sha256:39c3a47d50b3a3a994364cc7cc62597d8eedcb253d761c0816389cda6456db03

Observation 95313292-cfef-4e19-bfb4-2d3e34bd3517 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.451439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.451439Z digest=sha256:feb1c5b5eaad01369af235403d9472d7d5760ddb2fc4fe620224208b157a8123

Observation c4bad557-c00a-4824-a46b-b1469325e038 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Datacomp: In search of the next generation of multimodal datasets

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.455776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.455776Z digest=sha256:78d4a296db4b7942555553747e01b18def9cad960f765772f0b3d9c693ad0d32

Observation 1f7bf96f-a08b-4ead-b044-294fa695cd6f · outbound

This paper cites The Llama 3 Herd of Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.459625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.459625Z digest=sha256:cd0e436b6c21167d65fd09fb4d8afd8a7e20c92347eee51b46db69d5ba06eebf

Observation 7f491089-b44c-4db9-889b-1dae43d49693 · outbound

This paper cites Infinity-mm: Scaling multimodal performance with large-scale and high-quality instruction data, 2024.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Infinity-mm: Scaling multimodal performance with large-scale and high-quality instruction data, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.463499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.463499Z digest=sha256:2fe301f8f00decb5b416b3a178dbb89cf9494f5672b280a52341abb3a75db996

Observation 3151c2f8-a06c-4c19-a80b-c08f885a930b · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.467522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.467522Z digest=sha256:ab11be013d724f0241cea6db5cf2012c0261abb11d74adc250d99d76b529eda4

Observation 87b5b437-a964-4a01-9dc5-1f06368e1885 · outbound

This paper cites Textbooks Are All You Need.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Textbooks Are All You Need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.471225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.471225Z digest=sha256:ecffaa01410a5380197ceeafb649c0eb33480aec7637487ffb4ea66089d42899

Observation a18f1126-d174-4fd7-97fb-1955d8d14ca0 · outbound

This paper cites Allava: Harnessing gpt4v-synthesized data for a lite vision-language model.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Allava: Harnessing gpt4v-synthesized data for a lite vision-language model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.475285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.475285Z digest=sha256:b34909fbaf79a7a2139dc74de1aaaef9cfc23bc01ad79e2a760f26321e87b5d1

Observation 3df2794d-c347-4f3f-bc0e-3aedd14f0b7f · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.479009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.479009Z digest=sha256:756f1f5812eb57345e1443eeee6cbdb6d2bea1c670859c99144ae64a4cc8854b

Observation e2a39ce5-b3df-4ad4-a120-3d556f875274 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Lora: Low-rank adaptation of large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.483074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.483074Z digest=sha256:d6a48209df8d6da7f37ee361628c1b0335487137bd3b72ef740f5a93a271db34

Observation 7f12ce08-3dd4-4b59-a640-8737f7d1d128 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.486975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.486975Z digest=sha256:d72aae18ca25339f960afad4a553aca26e3ef43eab4147a79e45a396f7e96669

Observation 4e5c9176-8bed-4123-889d-78f75c21b15c · outbound

This paper cites A diagram is worth a dozen images.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput A diagram is worth a dozen images

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.491054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.491054Z digest=sha256:e4ecd6283c5cf8fba0e83108672cca6b33e22d6f82c59092746162d00a2609ce

Observation 0891194c-555a-4475-9a1e-1c812a7e15b9 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Gonzalez, Hao Zhang, and Ion Stoica

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.494803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.494803Z digest=sha256:c21c48eb5e143f66bfdbcfc59d754f1eb80aa5175dd8233fd75be88d86e8b2c5

Observation 892106fb-1adc-467b-be6f-aa3016d7c989 · outbound

This paper cites Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.498928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.498928Z digest=sha256:da16fc5accc567f30537db45d6ceeaf68799b139d90871d7603a474d16f2ade6

Observation a1607258-610c-4829-bafd-d2514d1672ff · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput LLaVA-OneVision: Easy Visual Task Transfer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.503002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.503002Z digest=sha256:6002147602e96d89df6fa512c77d4d8dd27b2e2de6cfc1b21aa6ca2a4197f753

Observation 14b5bb07-b234-4dc1-adf5-c1cb71bff8e9 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.506882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.506882Z digest=sha256:df53321c31c2dd5f023c8bced82e5997cd71ad5790c95dbc17487ae6bde40466

Observation 4e2b334e-b5ef-4d4b-9f8f-9d9a61a4d88e · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.510960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.510960Z digest=sha256:8e178217497c40d9aba22e5de495c9a842dabfb1a917ce95f990f7f6645e137a

Observation 7879abd3-9b42-4585-84f9-ed87c32b2d35 · outbound

This paper cites VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.514950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.514950Z digest=sha256:d6b9e0ef02a405567beac5119c98b4707d7f10d2edf2de8afa3866790e096c1d

Observation dcbfbcc7-cc48-4171-a781-de9ea4aec315 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.519188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.519188Z digest=sha256:627e48ad17910a21547f37fc7ad1be57b738a9afcb345fb254a684ddd8908057

Observation 285d940d-29c1-468f-92c4-4ec6055204a0 · outbound

This paper cites SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.523117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.523117Z digest=sha256:f3adda1e060940c6dfbc899a4349e4fe78300e27a08e09017f2ad838d0be0871

Observation 262e3dca-f600-488e-a63c-0208e0aac697 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Textbooks Are All You Need II: phi-1.5 technical report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.527057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.527057Z digest=sha256:a2acff67dfc84323016c40527d6b61bf7293e6b39d176fffccfbbcf8a1650df6

Observation 4d3a7fed-2f68-4222-b857-329f2ce4fa42 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Improved baselines with visual instruction tuning, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.531270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.531270Z digest=sha256:6811aa2c2b45e5eb207d4d8dd4d3b69286d2630c1c1c7f475c8c92d3f8618447

Observation f8f2f11e-ca16-4a3d-a265-acb331fbcbeb · outbound

This paper cites Visual instruction tuning, 2023.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Visual instruction tuning, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.534938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.534938Z digest=sha256:d72ad7b7a3874106ad25452c20097e8d7e89ae6654979c426b88505cd6c999d1

Observation 7f4da166-615e-4f69-84a3-3b8c8d9ee509 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.538840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.538840Z digest=sha256:64f1e684cbc3b5df15a8b007e251fb764c0aa1e2b4beea0b84e2c8ed5fb722e5

Observation 0c2b4e1c-e745-4f32-b6d2-9c1a53a866c6 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput NVILA: Efficient Frontier Visual Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.542510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.542510Z digest=sha256:d774408bef4a202f799bd1d9e718d5ebb6092c6f89ffbf069e7b6ac0747e4fcc

Observation cc460974-4449-48d9-81ff-b4efd56d5420 · outbound

This paper cites MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.546581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.546581Z digest=sha256:be6fea948a6705b5bace827eebbcf8a2252995a44ac64a85106b4869e518f885

Observation 85178570-0b19-4208-bc91-0c52c99daa95 · outbound

This paper cites multimodal-open-r1-8k-verified.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput multimodal-open-r1-8k-verified

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:35:46.659649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.550703Z digest=sha256:5fd9ec5be5237d371916f405c7dc6d21ca48b16357d336237beb7df108289ca6

Observation 69cd6e66-a2cb-46b0-9207-4f3f534c86cd · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.554613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.554613Z digest=sha256:c1ee3e6548825e7df26255f5a3a945b96244ab3365f302f7d0fb37eabb677706

Observation 61af8ff3-5098-48b9-afda-61ed107574cd · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Llm-pruner: On the structural pruning of large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.558827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.558827Z digest=sha256:c6e73f10af407637e776bcc13310be7bd4b9ef3c441f34bef4488232cf0421b6

Observation c88816df-f662-4795-919c-7b5eed26718c · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.562885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.562885Z digest=sha256:1cf79ebec4a2c2d7070d20acfd2a32af43b22c62a69dfedf7498cb9b904b3a0b

Observation 18806789-46f5-4ef1-bd4d-4ef203fff0b2 · outbound

This paper cites Infographicvqa.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Infographicvqa

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.567092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.567092Z digest=sha256:ed74acca8f692c12d64f901aeee3c4e8aafffabb5b1bdb249918662768d5f3d2

Observation 03ee9197-4dba-4d6e-a442-0d48d8c33c9e · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Docvqa: A dataset for vqa on document images

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.571076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.571076Z digest=sha256:37ce8116c54ef71d4a7452324c3d761d846de1e11f970ff0d7d9d013f9a80857

Observation bbc42e4a-b329-499d-8a4c-9a1aed72be2d · outbound

This paper cites Compact language models via pruning and knowledge distillation.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Compact language models via pruning and knowledge distillation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.575761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.575761Z digest=sha256:f57ea27159ad192444b674ae0f4f15a6ca4b6c6ea800bc2b694a7852bd2be2d7

Observation c68927e9-7f46-4fcd-84eb-26a8235b62c5 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput DINOv2: Learning Robust Visual Features without Supervision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.579799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.579799Z digest=sha256:cb724972f4224f507633ef5796cf6cf93add0f2a62af9954bc56360b2379d711

Observation b9cd64cd-878a-4a12-b864-eae73003537b · outbound

This paper cites Compositional Semantic Parsing on Semi-Structured Tables.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Compositional Semantic Parsing on Semi-Structured Tables

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.583830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.583830Z digest=sha256:54c3f72a366f9cbeca251ba5b238670ae35e046033142bc5855434ffe518738e

Observation bc7847af-f45d-4050-99b7-985eb969dbb6 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.587831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.587831Z digest=sha256:5f48bfab034776be7c48fe2d1ad92341b83b87217d0f540ba66b06ef1a5c6d2a

Observation dac9ae3e-a777-46fa-bc1c-92f1b2684df0 · outbound

This paper cites Using the Output Embedding to Improve Language Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Using the Output Embedding to Improve Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.591819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.591819Z digest=sha256:a022878db037b61380962fa73e788570f4e1ef74539ee829d7b7c5bd42e65ff0

Observation 077a867c-c008-4641-aacd-4f56210ca3dc · outbound

This paper cites We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.596074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.596074Z digest=sha256:3fe121b3e678f87a164f29927e8ca5658c09f73ee67f757164e0c8eabd99ecff

Observation 245b8ecd-52f7-4b93-bc77-d0d4aa000dab · outbound

This paper cites Learning transferable visual models from natural language supervision.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Learning transferable visual models from natural language supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.599905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.599905Z digest=sha256:d9b945e90328c56f11ecfca9534ba4c344057a155335528c015455c2be00ac9b

Observation a1bb6705-17de-4767-90ff-885d37ef5af7 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Direct preference optimization: Your language model is secretly a reward model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.603741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.603741Z digest=sha256:61c565c351780b8f834b2c5b3acca30fb9cac6d7bb31f07e7337c8a28e5b6f3f

Observation 6d264b0d-e7e8-4b2e-b343-22e19d098de7 · outbound

This paper cites an unresolved cited work.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:35:46.589101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.607754Z digest=sha256:3d170b2f55ad171c66faa5b71a75f7d0efd43194bb3692b0511091ecad2e496d

Observation c43d570e-0951-48ca-b507-58ce26359f51 · outbound

This paper cites Sglang: A fast serving framework for large language models and vision language models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Sglang: A fast serving framework for large language models and vision language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:35:46.574857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.611476Z digest=sha256:c23610e5212d258aa2e180efa7e855ecf305d6498a865c60faa4db7320154ed7

Observation 25f1e04a-f16e-42c6-9a83-c661c422fe84 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.615921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.615921Z digest=sha256:7578c0d81fbfcdcbeb6c164f3a9b6cb666c93c1aebeb8ba420247369758176ba

Observation 03cb215f-f561-49c1-b036-4e63e5c99028 · outbound

This paper cites GLU Variants Improve Transformer.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput GLU Variants Improve Transformer

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.620715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.620715Z digest=sha256:6be13ba58647733346272df541b97a14e0a730614c4590721af0960d375d396a

Observation edf8b53d-c1e9-437a-8b83-10ec97c1628e · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Textcaps: a dataset for image captioning with reading comprehension

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.624878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.624878Z digest=sha256:f44b8f53e9935fa547c54e82af7c50f9a1864e00993ce74e1c5eda304e13ca25

Observation b748258e-7d74-416d-ac81-5517b9d0b1d5 · outbound

This paper cites Towards vqa models that can read.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Towards vqa models that can read

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.629028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.629028Z digest=sha256:97e17d16a1bb7a6d7ea4d05de203851e65f660dd14fef11263e8b606e1b6adc8

Observation 093e15d9-e168-4473-b7fc-e0282d59717a · outbound

This paper cites Kleister: key information extraction datasets involving long documents with complex layouts.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Kleister: key information extraction datasets involving long documents with complex layouts

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:35:46.544289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.632930Z digest=sha256:0777d0712ace05f3a9b31c921ed7c16dc827dcb6466567a24f3ef51bae02bf9a

Observation b88239e0-bec0-4e81-b551-3fbb01b6ae9b · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Roformer: Enhanced transformer with rotary position embedding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.636806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.636806Z digest=sha256:36666f1ce382afed8862adfb32637ad03bb39c9859480e4324cd218232db2d9b

Observation f6fdbc40-7e18-453e-9c7e-54b2c87731e8 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.640709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.640709Z digest=sha256:09f81d9eb7e8659472a67254766923415c93d9ac9a1a1ec6acad8f3458f77b2a

Observation 845123bb-79db-44d0-a423-a125833a7d17 · outbound

This paper cites Svetlichnaya.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Svetlichnaya

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:35:46.522025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.644698Z digest=sha256:b54ede7ae2a0e273707fb897ce8d3bef1a958e8378688ab4168c184e48c5c2f0

Observation b2fbb4fd-2401-4bab-8852-9f39b9a106c6 · outbound

This paper cites Visualmrc: Machine reading comprehension on document images.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Visualmrc: Machine reading comprehension on document images

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:35:46.508219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.648602Z digest=sha256:b7b39114b63f27c4ee7444e23c04a11450debc7d1f016a3997ce5752f85d1c42

Observation 63237fa7-c260-415c-ac05-722c5e703ce6 · outbound

This paper cites Gemma 3 Technical Report.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Gemma 3 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.652445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.652445Z digest=sha256:a268b913a26c7f2e410824e538ec74af51f5cd5fa600b3662e1a526450963002

Observation bbe2ca52-5eef-4f3a-ab2b-8b4ab3e3bf77 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Gemma 2: Improving Open Language Models at a Practical Size

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.656822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.656822Z digest=sha256:2a427d90d7284a8a36120388b5754bb205cdb61fed526a2eb91221556e436528

Observation 8ecabe93-d684-449f-8676-10c4a8636237 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput LLaMA: Open and Efficient Foundation Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.661216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.661216Z digest=sha256:9ca0c78d1b23b8f783567c31fd9fa526ab640dca6cce2dd87104e82c1a82c480

Observation 82d57af6-74d7-43bb-bd6e-db967f63795b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.664975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.664975Z digest=sha256:dc8ebdc2ecf501410f0243b06c2f784bf4f42214e0b9f40336ffeb08b8eff999

Observation 8c17008e-0e6c-49e6-bd1d-83d509ec19a3 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.669169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.669169Z digest=sha256:213d2d1846320b79b9f1fcb7bec0ae5071478a05fe12f73cf8b405e145107e61

Observation 3e8858b4-0a4b-48fc-b715-bebe2018ec73 · outbound

This paper cites FastVLM: Efficient Vision Encoding for Vision Language Models.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput FastVLM: Efficient Vision Encoding for Vision Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.673348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.673348Z digest=sha256:570719aca1e9445159b255f077ea421a84979f586a3b0c74712c25d706df976e

Observation 441fb3ca-af28-48bf-9155-109457acd5aa · outbound

This paper cites Locca: Visual pretraining with location-aware captioners.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Locca: Visual pretraining with location-aware captioners

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:35:46.494028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.677281Z digest=sha256:9102e3b7246df67e6f8c8f8c1e570393e0078b8788f736b5f1e0539dbbd9f699

Observation 02330e0d-aa5e-4092-9c7e-0900d6efe226 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Measuring multimodal mathematical reasoning with math-vision dataset

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.681135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.681135Z digest=sha256:37bb55d3a8e4bdd6ab149130ea99df879a4d80c3c80d91f844cf7f12b70f9574

Observation 795496d8-261f-41c9-a8c4-7b12eb809b67 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.684953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.684953Z digest=sha256:444350d24c8197737f14793d4d6cea8c501b5804965b8f4dcaa3dd5353be5ddf

Observation 2772576b-b40c-4d49-91ca-ae523cf2deb9 · outbound

This paper cites Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:35:46.471476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.688898Z digest=sha256:efc3d80db76323066bbc9ebdf3ecb41f5b7b054f7ed34c6ba66244f9ef28f573

Observation d4b1bfc3-b040-4bd5-9ede-5d75195c726a · outbound

This paper cites Qwen2.5 Technical Report.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Qwen2.5 Technical Report

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.696887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.696887Z digest=sha256:81ae046612c48b695f4ae57e0922b35b54b6b2689c7977bf881d2ea6d8df37f4

Observation 32c3218f-3ff7-4b5d-b036-3b1f78c1d6f1 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.700850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.700850Z digest=sha256:fc096b2eca1bfe6c29096f524241ceb10f7b6badc319e6bf3f2a4079071ef333

Observation cd382ae5-5548-454d-9ddb-22b2d92cdce4 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MMBench: Is Your Multi-modal Model an All-around Player?

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.704869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.704869Z digest=sha256:a98b5d92e1a93db052ac6f868000bbae4b000fb5b4ae7715e2c05cf726fbd51e

Observation 7057affd-be8f-433b-8cc7-7a2916332f2e · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.708826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.708826Z digest=sha256:e50b98acf134116c0be4ce46b923ef802359d84dcaf828ed368aba2b6e090f57

Observation c849d769-fdbd-4c7e-8806-3a8a8ccd5c63 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.712766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.712766Z digest=sha256:210a242e83fcdc63f528ffac506bc8790428e1f8532ec6c30e16728c80a1d764

Observation 550b9d15-b201-4e3b-9f43-274c948f2422 · outbound

This paper cites Sigmoid loss for language image pre-training.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Sigmoid loss for language image pre-training

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.717114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.717114Z digest=sha256:00c7c4be4455e09bbb7e9415a5208715bf145c7a03cd3c26ec505feed99a94ad

Observation 7c530d79-3b25-4e0d-8270-65ec0133eda0 · outbound

This paper cites Root mean square layer normalization.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Root mean square layer normalization

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.720982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.720982Z digest=sha256:fe1caad32d73a6962c792c02fc8d6c0e69ba8bcc925587b56817bf90e2309c71

Observation d5730850-b50d-4102-9db1-35c2b38c49b6 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.725159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.725159Z digest=sha256:d6dbc5462265c3acdf5a982816b469b05186bf47528329455100df70dfea2e50

Observation e1188ee0-77db-412b-9055-c10863d7ea2c · outbound

This paper cites Mavis: Mathematical visual instruction tuning.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Mavis: Mathematical visual instruction tuning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:35:46.429007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.729529Z digest=sha256:767dde09502f28c17c5d09ba054efc0f787daca48f4937a977555b6f585d8466

Observation d3a0d765-9f70-4026-8959-6da6d27734b7 · outbound

This paper cites Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models, 2024.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models, 2024

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:35:46.413678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:35:45.733427Z digest=sha256:9faf4227641656ae6fe6153261316fc9c52ba65862c0301c619df63eb52dae1f

Pith citing papers

Observation 44f5f89c-4f34-486e-be23-80d160b25e10 · inbound

Affordance Benchmark for MLLMs cites this paper.

Affordance Benchmark for MLLMs Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:55.805940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:55.805940Z digest=sha256:cd50b3bd5015c6d1ca9db395d053dd43913c808d999d3f8ef8fc5d91ccd6fffc

Observation c8760376-d25b-4280-92e6-05e754f0d73a · inbound

CLGRPO: Reasoning Ability Enhancement for Small VLMs cites this paper.

CLGRPO: Reasoning Ability Enhancement for Small VLMs Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:46.048132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:46.048132Z digest=sha256:989a6c0b3f11e5fc7f44ab40c5465b9752c6315962cdb7542436bdfc270536ce

Observation 28c9addc-aba3-4d4f-beaa-1c918695bd76 · inbound

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe cites this paper.

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:07:27.424607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T17:07:27.277040Z digest=sha256:0d21e7d98df738383e9622ebaf6527f9c13d04b6a11fc88f9fc5414aca1c347a