Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T15:27:51.839171Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 39 inbound Pith citation observations for arXiv:2402.03766.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T15:27:51.839171Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:57:51.203800Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T11:37:03.180051Z
76 of 76 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4ac961fb-dca6-47d3-99ad-aef19c168c0e · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model An In-depth Look at Gemini's Language Abilities
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 48f4433e-68f3-436b-b359-702e70248174 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Openflamingo, Mar
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bf6961b-baf2-4dc7-bb68-eb062984e149 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Qwen Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 575954b2-bb8d-4e90-ac67-4b1066b40578 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0cd1f460-d9fd-41d7-9627-acdad45b0736 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Pythia: A suite for analyz- 8 ing large language models across training and scaling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7d66992c-c8e1-4071-93c1-51f8fa90e963 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Lan- guage models are few-shot learners
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c81e777f-0216-4db0-932d-039692b34c9f · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Honeybee: Locality-enhanced Projector for Multimodal LLM
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aca742c5-3139-4668-9931-e8d7e9b14bf5 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df2417ea-5092-4817-bdea-feb8b540397c · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d9c1dc2d-4e6b-4eff-9e9e-417bad7381f6 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ef7bb48-4db9-4437-a7db-99e14d1cd6be · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Lawrence Zit- nick
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a9dfdb84-7400-4b96-9c6f-82231d804e86 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Unifying vision-and-language tasks via text generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 520d1d09-1e61-406e-9b58-a8fdd3ec07c8 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model PaLM: Scaling Language Modeling with Pathways
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ab7ef83-c59f-4a6c-ba2a-b6ab7e23f2c8 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Make repvgg greater again: A quantization-aware approach
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b230a10c-ff3d-4239-9288-978f8af0d072 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5fe84c8-e652-4310-aab8-7611bf2d211a · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Conditional positional encodings for vision transformers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e4940fbc-0982-4ed0-ae52-83bcfbb8aa0e · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Redpajama: An open source recipe to reproduce llama training dataset
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 319b4c68-b983-4269-be0e-fa8dd74a6b49 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8ddaba5-2bde-40d0-a9e7-8c809dc86717 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Visual dialog
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63fa8dd1-b041-4d49-8d28-0c39051f0031 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44b10e6b-9730-447b-b5d2-2ac64cf4cce7 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Glm: General language model pretraining with autoregressive blank infilling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3445fc17-4be6-4ad5-96b3-8cce65eb9628 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Learning factored representations in a deep mixture of ex- perts
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 234ed899-9e02-4f85-ab2f-bb234a2d9d61 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Sparsegpt: Massive language models can be accurately pruned in one-shot
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6994e6a0-6d2b-41bd-9cd1-c5b9536a85fd · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d439e00-5349-4238-8a67-9c212505d4e3 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 026ec8cd-05d3-41ab-9ba2-5bd5371627b9 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 744c4b3b-8881-40bc-83d1-604db3c97a31 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model llama.cpp
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a273e820-2090-4d8f-bf92-569cdde06982 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Gaussian Error Linear Units (GELUs)
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 92eeb03f-c69c-438a-a535-7012e4a99e4d · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b40d2e1b-5a5d-4e5f-84ce-bf0d4469ac32 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Adaptive mixtures of local experts
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da3db281-051c-4ce1-b300-d929e23af9bb · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Segment Anything
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 966822c7-3917-4dac-86a2-54ec9ae3d0f6 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Grounding Language Models to Images for Multimodal Inputs and Outputs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2f92ab5-4202-4cf0-a0bd-328334a5928e · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model LISA: Reasoning Segmentation via Large Language Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db62ab0b-ac30-4e00-a10c-c0e73749223f · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63a8860f-f4a3-49db-ad62-1148186d497b · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a7e1164e-7e3d-4a85-a07f-cbe081b9fe9f · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Align before fuse: Vision and language representation learn- ing with momentum distillation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b57c595-8d08-4ea6-b2a8-6bc299ae5607 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Norm tweaking: High-performance low-bit quantization of large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2eb03ca-d5c7-4e71-9617-76217e8c8511 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model A Speed Odyssey for Deployable Quantization of LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6383d318-68d6-4012-833e-d8df8fa8039f · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Textbooks are all you need ii: phi-1.5 technical report
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e2685a1-7425-432c-bb01-612a73767871 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Evaluating Object Hallucination in Large Vision-Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 489f0fd0-9082-421a-b621-7bebd5eef3b3 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Moe-llava: Mixture of experts for large vision-language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b28d5e23-754b-4146-b66c-2edc1b02d9f5 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Microsoft COCO: Common objects in context
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4fa4aae8-e89c-4984-9d94-c984b02c4190 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Visual spatial reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 393415f0-7035-4ceb-ac53-54562d4a00d4 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Improved Baselines with Visual Instruction Tuning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc8325f2-a57e-4d56-a780-8c7af5cd2464 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Visual Instruction Tuning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b70fd953-5cde-479d-92b6-714a256bc42e · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model MMBench: Is Your Multi-modal Model an All-around Player?
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 845ea1f4-c9ad-4d34-9c49-10df387590df · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Decoupled Weight Decay Regularization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8662f6d-a0d7-469d-88f2-a5c2818da2c8 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83dc7c90-425f-4be1-ac92-cc3f3d854e24 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 934f6e7b-7d22-4a00-806a-38e62fedbe67 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dca8226b-6241-4580-8eb0-c6c5bf74c7ce · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a918778e-858d-49a1-b1d0-d507df791c4d · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Gpt-4 technical report
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d787268-28e5-4a45-a487-cd3db282965f · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Gpt-4v(ision) system card
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46f231c0-bce3-4da2-ba8b-49b1fc79fa88 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb764305-6ec1-4740-a4b9-9cb375379ab7 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Training language models to follow instructions with human feed- back
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ebabdbb0-1260-4f70-b470-19f88da8c8d4 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Tinyllama, Sep 2023
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c9d1cda-54f3-45cf-bcb3-cfaaed5176d0 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model DetGPT: Detect What You Need via Reasoning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f3ca63be-7a1f-4ae7-aea2-d45f1d695eb8 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Learning transferable visual models from natural language supervi- sion
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4084dbf0-30f1-4329-b6d9-41cc6d4e4fc3 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f096a06-d55a-420e-8b3e-e4b810e340f1 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Towards vqa models that can read
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 096c54d1-f988-48fe-9848-86c5be5869b2 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31e89ffa-681b-423a-8e35-2e79facef8aa · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Galactica: A large language model for science
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 98d3c041-a012-4907-92e1-96d8c3ec58fc · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Internlm: A multilingual language model with progressively enhanced capabilities
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8257f077-779f-4112-a334-19dd2658fccb · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model VIGC: Visual Instruction Generation and Correction
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ced92d1-13d4-4817-80a1-a4ac61a6905e · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0bbdbedb-041e-4764-a37c-88270197c999 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3cb97303-4f3a-4cb8-9c4d-7f55523da5bb · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ada21607-cf5d-417c-88d9-6699a4d68a1a · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Smoothquant: Accurate and effi- cient post-training quantization for large language models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e08e007d-314b-4364-b708-bb86230343d9 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Baichuan 2: Open Large-scale Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e779d337-483a-43dd-9583-b1117bc6f29d · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b5169e2-0cd1-430f-9b8a-5eb46314ec32 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model MM-LLMs: Recent Advances in MultiModal Large Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df0ebf9f-713e-4bf4-b58a-de77614a406b · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model OPT: Open pre-trained transformer language models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 60743011-f516-4957-8578-263f3ccdb2dc · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model SVIT: Scaling up Visual Instruction Tuning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2344869-9b88-411a-a554-8a528497e70b · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Lidar-ptq:post-training quantization for point cloud 3d object detection
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d611b88c-bc50-4b44-82b6-e93f7ab976e1 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 238f4ae0-43af-44c0-ac1d-6f6f576eac31 · outbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 263d85c1-f1d6-4364-b24b-a71a956863f9 · inbound
A Survey on Multimodal Large Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1503c723-328b-4cac-a9aa-56a377499b70 · inbound
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46e295f0-9a91-4c7b-bdbb-717cd03e1184 · inbound
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0a91e8e-e316-4caa-96c1-5f67b4ca59b0 · inbound
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 344e15de-5299-43a4-ae83-9383e3b1709f · inbound
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ada2901-c73a-4a18-8701-75ddf83dc74a · inbound
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8ce9c562-fcff-446c-8640-5f2324255869 · inbound
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 10e90e2a-2391-46d1-ba25-090cfe54c8e6 · inbound
Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b12cfc-f810-4ba5-b98f-1a3ff685a3d2 · inbound
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e0de29-c462-44ec-9e5c-e515c51147a4 · inbound
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e53504-3766-4af0-9e9b-137aa3ed3238 · inbound
Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e437234-3ca9-455e-a694-cee88a569e27 · inbound
Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fefafde8-9ebd-465c-bba6-06632c49f513 · inbound
Agentic Services Computing MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a9b26d-283c-4c63-a30d-df97864bca55 · inbound
Mind the Gap: Action Rebinding Attacks against Android GUI Agents MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6fc2fc-95c7-4f4f-a501-fdb244ebeeb3 · inbound
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 48311f00-544d-4618-9e59-65ecc5517e89 · inbound
Vision-aligned Latent Reasoning for Multi-modal Large Language Model MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aef52343-ad2f-448b-8775-3a031ebe0551 · inbound
Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 24be7692-d81e-462a-83b6-fc10da143e59 · inbound
Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d1f4314d-bbc1-4a66-bc8d-f81ca829cfbe · inbound
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f73e3e90-fb04-424f-826d-9ae0b3fd6f3e · inbound
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e560d582-a4c4-4dc7-a0ba-46d3dd9acccc · inbound
UIPress: Bringing Optical Token Compression to UI-to-Code Generation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 66deedc9-2dc7-486d-9b70-61735715618a · inbound
UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 74dba68e-92bd-47b6-bab3-46d01fb03e80 · inbound
Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7ed7f7c2-ee9d-4265-bb9d-a3ecfbc1d01b · inbound
LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed8d7760-2a87-436a-be4b-7cf844f4ee7c · inbound
A More Word-like Image Tokenization for MLLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 638883d0-c9e0-4ffc-bd30-9c572fb44799 · inbound
DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5bcd4345-234b-4e9e-adfd-ee4c9180f3b3 · inbound
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 81f8c5a9-11d6-4e4b-892d-6b6bc179103d · inbound
MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2ea1e56-24c1-4c16-b19b-b10eaeded060 · inbound
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 778d258d-89df-462f-9886-ecf1f382baae · inbound
Efficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a4add3f-3d63-4946-9eed-ffb010edd213 · inbound
Dive Into the Implicit Biases of Low-rank Vision-language Alignment MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 020bc01c-66c8-4479-bebc-9c0864555736 · inbound
Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 334e0b21-23cc-4d92-a6f9-0c5ebd2ea163 · inbound
Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9784806a-cd1b-4c70-b5a9-c2e9048b3ee2 · inbound
PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccfa5a24-c162-4fba-9837-a13979c35d7a · inbound
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eedcb01-ff63-42b9-bef7-2073b8388be6 · inbound
RemiAssist: A Therapist-Supporting System for Photo-Based Reminiscence Therapy in Dementia Care MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fee9ea4-365d-4765-a35f-c423d3ef6543 · inbound
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f98822-efca-476e-a94c-f2048e67ec07 · inbound
Geometric Cross-Modal Token Selection for Latency-Constrained Multimodal Token Communication MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 147f2d56-c4fd-4482-b478-3970180b19b6 · inbound
SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.