Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:35:45.733427Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 3 inbound Pith citation observations for arXiv:2505.09498.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:35:45.733427Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:55.805940Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T17:07:27.421336Z
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 77d02640-a53f-44f4-9eac-958d35160ca0 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ee49ae1-0b99-4906-878e-c68b37d813e4 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca32b92-1f96-4f51-9fd8-6dafaa56b239 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6556433b-7743-49a3-b5a7-b3988cf440f5 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3168a3d7-c33f-4cd3-8d52-59023128b02d · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966db0bb-a95a-42a8-b7f0-d3b13b3b6755 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Sharegpt4v: Improving large multi-modal models with better captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66264194-0a48-4bfd-970e-0ddf9524add9 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4153cdbb-d210-422e-b805-50984e9412c2 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput TabFact: A Large-scale Dataset for Table-based Fact Verification
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a07685-31b8-45be-a663-18ae7d4de02e · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f2f4329-4fc6-4a60-aa3f-ec57c3f54919 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50bcaa31-2d35-43f4-8bd6-2420879ab19b · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0624d95-f765-4b31-802c-f244eed3e9f3 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f08a5d3d-b74c-48ba-803b-8897b421989e · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput The chinese dataset distilled from deepseek-r1-671b
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fbe0308-9ffd-400e-8f8d-06697a7e4357 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Scalable Vision Language Model Training via High Quality Data Curation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7598aa4b-e8f8-4a4d-a38d-0000aad386ce · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe1dcda-02a8-4c68-a1aa-4639483cc40a · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4117b70a-d981-4836-9a08-ddf36b57bc3b · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Scalable Pre-training of Large Autoregressive Image Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd30e9ef-8ad2-4394-a2cf-81fcd45e7cdb · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Data Filtering Networks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff1c0eb-cc4e-4745-9acc-859ebcf812c4 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95313292-cfef-4e19-bfb4-2d3e34bd3517 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4bad557-c00a-4824-a46b-b1469325e038 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Datacomp: In search of the next generation of multimodal datasets
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f7bf96f-a08b-4ead-b044-294fa695cd6f · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput The Llama 3 Herd of Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f491089-b44c-4db9-889b-1dae43d49693 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Infinity-mm: Scaling multimodal performance with large-scale and high-quality instruction data, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3151c2f8-a06c-4c19-a80b-c08f885a930b · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b5b437-a964-4a01-9dc5-1f06368e1885 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Textbooks Are All You Need
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a18f1126-d174-4fd7-97fb-1955d8d14ca0 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Allava: Harnessing gpt4v-synthesized data for a lite vision-language model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3df2794d-c347-4f3f-bc0e-3aedd14f0b7f · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a39ce5-b3df-4ad4-a120-3d556f875274 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Lora: Low-rank adaptation of large language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f12ce08-3dd4-4b59-a640-8737f7d1d128 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e5c9176-8bed-4123-889d-78f75c21b15c · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput A diagram is worth a dozen images
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0891194c-555a-4475-9a1e-1c812a7e15b9 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Gonzalez, Hao Zhang, and Ion Stoica
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892106fb-1adc-467b-be6f-aa3016d7c989 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1607258-610c-4829-bafd-d2514d1672ff · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput LLaVA-OneVision: Easy Visual Task Transfer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14b5bb07-b234-4dc1-adf5-c1cb71bff8e9 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e2b334e-b5ef-4d4b-9f8f-9d9a61a4d88e · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7879abd3-9b42-4585-84f9-ed87c32b2d35 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcbfbcc7-cc48-4171-a781-de9ea4aec315 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 285d940d-29c1-468f-92c4-4ec6055204a0 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 262e3dca-f600-488e-a63c-0208e0aac697 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Textbooks Are All You Need II: phi-1.5 technical report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d3a7fed-2f68-4222-b857-329f2ce4fa42 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Improved baselines with visual instruction tuning, 2023
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f2f11e-ca16-4a3d-a265-acb331fbcbeb · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Visual instruction tuning, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4da166-615e-4f69-84a3-3b8c8d9ee509 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2b4e1c-e745-4f32-b6d2-9c1a53a866c6 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput NVILA: Efficient Frontier Visual Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc460974-4449-48d9-81ff-b4efd56d5420 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85178570-0b19-4208-bc91-0c52c99daa95 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput multimodal-open-r1-8k-verified
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 69cd6e66-a2cb-46b0-9207-4f3f534c86cd · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61af8ff3-5098-48b9-afda-61ed107574cd · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Llm-pruner: On the structural pruning of large language models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c88816df-f662-4795-919c-7b5eed26718c · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18806789-46f5-4ef1-bd4d-4ef203fff0b2 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Infographicvqa
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ee9197-4dba-4d6e-a442-0d48d8c33c9e · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Docvqa: A dataset for vqa on document images
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc42e4a-b329-499d-8a4c-9a1aed72be2d · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Compact language models via pruning and knowledge distillation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c68927e9-7f46-4fcd-84eb-26a8235b62c5 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput DINOv2: Learning Robust Visual Features without Supervision
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9cd64cd-878a-4a12-b864-eae73003537b · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Compositional Semantic Parsing on Semi-Structured Tables
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7847af-f45d-4050-99b7-985eb969dbb6 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dac9ae3e-a777-46fa-bc1c-92f1b2684df0 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Using the Output Embedding to Improve Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 077a867c-c008-4641-aacd-4f56210ca3dc · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245b8ecd-52f7-4b93-bc77-d0d4aa000dab · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Learning transferable visual models from natural language supervision
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1bb6705-17de-4767-90ff-885d37ef5af7 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Direct preference optimization: Your language model is secretly a reward model
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d264b0d-e7e8-4b2e-b343-22e19d098de7 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c43d570e-0951-48ca-b507-58ce26359f51 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Sglang: A fast serving framework for large language models and vision language models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 25f1e04a-f16e-42c6-9a83-c661c422fe84 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03cb215f-f561-49c1-b036-4e63e5c99028 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput GLU Variants Improve Transformer
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edf8b53d-c1e9-437a-8b83-10ec97c1628e · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Textcaps: a dataset for image captioning with reading comprehension
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b748258e-7d74-416d-ac81-5517b9d0b1d5 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Towards vqa models that can read
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 093e15d9-e168-4473-b7fc-e0282d59717a · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Kleister: key information extraction datasets involving long documents with complex layouts
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b88239e0-bec0-4e81-b551-3fbb01b6ae9b · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Roformer: Enhanced transformer with rotary position embedding
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6fdbc40-7e18-453e-9c7e-54b2c87731e8 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 845123bb-79db-44d0-a423-a125833a7d17 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Svetlichnaya
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b2fbb4fd-2401-4bab-8852-9f39b9a106c6 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Visualmrc: Machine reading comprehension on document images
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 63237fa7-c260-415c-ac05-722c5e703ce6 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Gemma 3 Technical Report
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe2ca52-5eef-4f3a-ab2b-8b4ab3e3bf77 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Gemma 2: Improving Open Language Models at a Practical Size
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ecabe93-d684-449f-8676-10c4a8636237 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput LLaMA: Open and Efficient Foundation Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d57af6-74d7-43bb-bd6e-db967f63795b · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c17008e-0e6c-49e6-bd1d-83d509ec19a3 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e8858b4-0a4b-48fc-b715-bebe2018ec73 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput FastVLM: Efficient Vision Encoding for Vision Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 441fb3ca-af28-48bf-9155-109457acd5aa · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Locca: Visual pretraining with location-aware captioners
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 02330e0d-aa5e-4092-9c7e-0900d6efe226 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Measuring multimodal mathematical reasoning with math-vision dataset
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 795496d8-261f-41c9-a8c4-7b12eb809b67 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2772576b-b40c-4d49-91ca-ae523cf2deb9 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d4b1bfc3-b040-4bd5-9ede-5d75195c726a · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Qwen2.5 Technical Report
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c3218f-3ff7-4b5d-b036-3b1f78c1d6f1 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd382ae5-5548-454d-9ddb-22b2d92cdce4 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MMBench: Is Your Multi-modal Model an All-around Player?
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7057affd-be8f-433b-8cc7-7a2916332f2e · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c849d769-fdbd-4c7e-8806-3a8a8ccd5c63 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550b9d15-b201-4e3b-9f43-274c948f2422 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Sigmoid loss for language image pre-training
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c530d79-3b25-4e0d-8270-65ec0133eda0 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Root mean square layer normalization
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5730850-b50d-4102-9db1-35c2b38c49b6 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1188ee0-77db-412b-9055-c10863d7ea2c · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Mavis: Mathematical visual instruction tuning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d3a0d765-9f70-4026-8959-6da6d27734b7 · outbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models, 2024
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 44f5f89c-4f34-486e-be23-80d160b25e10 · inbound
Affordance Benchmark for MLLMs Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8760376-d25b-4280-92e6-05e754f0d73a · inbound
CLGRPO: Reasoning Ability Enhancement for Small VLMs Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c9addc-aba3-4d4f-beaa-1c918695bd76 · inbound
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.