Pith. sign in

Paper Citation Record · LEDGER

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

As of 6 August 2026, this Paper Citation Record lists 100 of 133 outbound references and 41 inbound Pith citation observations for arXiv:2312.16886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.16886 v2

Coverage vector

measured 100 of 133 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T16:35:37.937462Z

measured 141 of 141 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:57:51.199560Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 133 outbound references displayed

  • verified exact40
  • verified fuzzy58
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

11
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5eed6081-30a2-47ec-84b1-18e8d3952b82 · outbound

This paper cites An In-depth Look at Gemini's Language Abilities.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices An In-depth Look at Gemini's Language Abilities

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.142322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:c84fbc96e158ca0c29297c205aa8bc0fcc18adb8476600483e693a6100ced077

Observation 9a6c8d66-6592-408e-b290-9bd0d3f1213c · outbound

This paper cites Flamingo: a visual language model for few-shot learn- ing.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Flamingo: a visual language model for few-shot learn- ing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.247234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:09213d4d8173ffab4654374319737f914c456976c4c4125075c10c32a855d1b2

Observation 873f98c9-384f-402a-aed2-31c19988b0d4 · outbound

This paper cites Openflamingo, Mar.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Openflamingo, Mar

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.248955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:30fdecf474d90ccef1777f80ab43fc9ee28616d31d41861110ec354638d1eb66

Observation 82d99c5d-308c-4bd1-a0fe-aa0fca777196 · outbound

This paper cites Qwen Technical Report.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Qwen Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.032133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:86199dfa2dddc4998220db07234427c917ee91e7edf5408dd00e2e5a67038577

Observation 422cb836-9dae-40ac-8b80-4ef6b3a118ac · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.049636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:13bb3be5fd9e819ad6f24942bd0a2f71e0e3443f5158c0d4d69bff25408cd7f4

Observation 50ef19f7-43c4-4251-9ccd-6d36b85f0d17 · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.250741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:fcac9c6d778074da897acb5b7a11bb90ccd028f66b824bfd05ab1a92bd03179d

Observation ffdd2359-c8f1-4b37-bfe7-a3c363be75cd · outbound

This paper cites Pythia: A suite for ana- lyzing large language models across training and scaling.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Pythia: A suite for ana- lyzing large language models across training and scaling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.252559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:e01d84e6d1d5837d1db350a9b7cfc3ec87df061d4b3455575843164e32a0c70d

Observation 3823f6cf-9c2f-4bd3-8fd9-616009c73303 · outbound

This paper cites Piqa: Reasoning about physical commonsense in nat- ural language.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Piqa: Reasoning about physical commonsense in nat- ural language

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.254455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:4e0c0544eeaa4e300dd2ea2e447b03feba8518a9848d2312423cc491290e98a3

Observation 2bb38b1d-1c60-4ea8-a2a8-6063c86ce214 · outbound

This paper cites GPT-Neo: Large Scale Autoregressive Lan- guage Modeling with Mesh-Tensorflow, Mar.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices GPT-Neo: Large Scale Autoregressive Lan- guage Modeling with Mesh-Tensorflow, Mar

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.256566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:53acbd0c547a5cd463e3bbca31766d0f5ea8ffc33f63b46faaa9f75073869285

Observation 0e688f61-389e-4b6d-bd22-50d92130cd59 · outbound

This paper cites A Systematic Classification of Knowledge, Reasoning, and Context within the ARC Dataset.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices A Systematic Classification of Knowledge, Reasoning, and Context within the ARC Dataset

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.064120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:6826b56ed4262761e71ec58d8f75050fa6c74861d64261934dcf7f54b1ff7545

Observation 6d8a0eb0-1608-4def-a6f6-3d96ccd01757 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Coyo-700m: Image-text pair dataset

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.258409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:fbb6fb7b75f81ce37f0191c2e748eef2cdb083fdf4f2dfa9135232abeb68ea30

Observation 0a18f4a3-615d-4dff-89dd-33da347041aa · outbound

This paper cites Once for all: Train one network and specialize it for efficient deployment.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Once for all: Train one network and specialize it for efficient deployment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.260676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:54c9e888e0bc78be586b10e0980dad69540f545181dea899c13eb80b616f638f

Observation 23e72aa8-0b56-4380-8ff1-17df27ab4a85 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.262584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:571ff9fd707429ea8f963bf5814adb6cb6a78cf1d63a264eeda6c2b3108f5e9a

Observation 38c7a0ae-4178-4d44-a838-c0e90b49d1af · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.068532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:e338031ceee15f623406a321af92b8e1a2cbb904726e0d2dd1cd91eb4bdca057

Observation 92d2ae53-8923-421e-90b7-5444a2ee8872 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.076284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:67bbf69851235639a5d43c1343ea134055de6dfdd96a312cf64c343dbbc9e5d7

Observation 328bc085-e260-4e04-8f2f-7dd0eb4a2401 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.125661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:9867ddeb074d46d74dd48a6bf48230da955fbcaff910a08c61c255eb76ab33ca

Observation 1ffb967a-1021-4cee-89a7-9aa628c5f519 · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Extending Context Window of Large Language Models via Positional Interpolation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.132466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:c6d133cc56b63f9339e64108e9d7000a483ec0fa5f3d788ebc8c5904068d4b8a

Observation 2dcbc99f-9354-4f97-8bb3-779a16c5e5fd · outbound

This paper cites PaLI-X: On scaling up a multilingual vision and language model.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices PaLI-X: On scaling up a multilingual vision and language model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.264461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:911503d34026616aac3acc391a4df17d38e2067764ec19d9fe9c3b6b7f8f9aa1

Observation 90927561-6a15-468c-aeca-c9388a5a45e2 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.046449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:549bf6e20cbe5632ee4f5226bcc5e5950299cea7aa2765ad59705ee8ffc53c72

Observation 08696784-6c65-4f33-8a16-260be8c5b887 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Gonzalez, Ion Stoica, and Eric P

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.266220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:6da75d2874bbb1e434d856e4cec8e2605c9411315e187e7e4b33e998b90e1dfd

Observation 28a31722-64f4-438b-b82a-4669d5530dc2 · outbound

This paper cites Make repvgg greater again: A quantization-aware approach.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Make repvgg greater again: A quantization-aware approach

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.268222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:a20b25ecdcc02134a9cbb1194178b85ced07b352fee2f9a09b634d8e5c93386c

Observation df32ecc5-d200-4e25-b607-b1ad62bac922 · outbound

This paper cites Twins: Revisiting the design of spatial attention in vision transformers.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Twins: Revisiting the design of spatial attention in vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.270095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:aaef5bcd0f588830c953b97417d0491b2ba4daab1114c8f18f57dbb7aa928ba8

Observation a1d797ff-8af9-4fe5-8eda-251ce438f4b6 · outbound

This paper cites Conditional positional encodings for vision transformers.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Conditional positional encodings for vision transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.272095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:44a284b81a9f4d36d1af2bbf492642f663a5d2e365d7daff1591a3ad31d0a583

Observation 073406e1-8916-42af-8ab9-7eadd8de39c6 · outbound

This paper cites Fairnas: Re- thinking evaluation fairness of weight sharing neural archi- tecture search.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Fairnas: Re- thinking evaluation fairness of weight sharing neural archi- tecture search

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.274464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:d4862c314825edd004b317e3c594707fed60dfbc492aaa57f0e1ef1a17490aba

Observation 36b2cbfa-e7c2-485f-98d7-5dd4ace69ba5 · outbound

This paper cites Fair darts: Eliminating unfair advantages in differentiable architecture search.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Fair darts: Eliminating unfair advantages in differentiable architecture search

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.276338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:773c7f0407051ad8f2e35829f51e7a7d1943ae9141f6908a1d5dd8a38a756581

Observation 2df0cd91-cc59-4538-8507-2a82a43c1f0c · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Scaling Instruction-Finetuned Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.155666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:f151688f4f2c30c3ca64d265155596494e6ff683e296cb5317ee6f37e17223f2

Observation ad70d21e-b604-4acd-ad03-d529cc1c126e · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.181380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:baaa6c91d3177b6e31c9ceab1e0c091255bf03a682fa54e45c502eb2cc70b7cd

Observation c35004e5-ffe4-4e82-a64d-d900f1b33fab · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.010976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:54da3966cceed258e35d700ae835a176d7a08e2d32a2a3a6154fdb8d7d83fc97

Observation 302f9f3e-d267-4bd8-9c09-b0afa9900015 · outbound

This paper cites Redpajama: An open source recipe to reproduce llama training dataset.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Redpajama: An open source recipe to reproduce llama training dataset

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.278458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:66d954babe07b94dee3182769a59994dc0c2eebdd1f603152ccc47064aaaea74

Observation 8591879d-9bf4-4c6b-ad1e-7f9cdddb7cfa · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.039764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:1bb060e70dd808fd3b01d000523a8cc5f66b1d8256f4c3a4cb9d6bdfdac15f92

Observation 5db8b8e8-b9d1-4bff-8c63-8bb62b17efde · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.043166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:0a94f627741c7db1601f7b24e5e8fcef41bbb8ee501995047f2288159f041b91

Observation 10fb5702-1e3b-433d-9801-311a5d72d28c · outbound

This paper cites Embodied question answering.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Embodied question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.280167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:1bfb7a1a43bc4b0f17323e0d2318295a4d23a61216481b73b0dc25775f1838c6

Observation bfb21d82-0780-48e2-a22d-875f808224cf · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Imagenet: A large-scale hierarchical im- age database

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.282108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:0fdd43e2d67e6180b3df1af4f0cea4affeb9aa2bb47a69d921aa4650255c1e51

Observation 35d22039-aa0c-4796-8d4d-aa09caa2b567 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.052654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:df88857905593df2f6826c8b5336b420913b41d826372f87000b9d56fe5b79be

Observation 0f6e36c1-d041-40e0-a38c-e74bf1b91b45 · outbound

This paper cites GLM: General Language Model Pretraining with Autoregressive Blank Infilling.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices GLM: General Language Model Pretraining with Autoregressive Blank Infilling

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.059819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:53ae25dcec55c2dbf57ff15b7f75c4a1df5be73ddbe492d322877de6749f97ab

Observation 0a2105fa-f271-4aee-b25e-1ce4dde7c76f · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices A survey of embodied ai: From simulators to research tasks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.284070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:acf8155df1b96ce49183ac93e00c5d0507edffa01dac5af7f01d88ac653c2759

Observation 6295a966-5c89-46b3-a756-1ec222594cc7 · outbound

This paper cites Eva: Exploring the limits of masked visual represen- tation learning at scale.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Eva: Exploring the limits of masked visual represen- tation learning at scale

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.285895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:750fb4b4a8f35072141eb2313b3017776e58909e353e61ed9a088fdf5ca66003

Observation a63b2f0c-5f60-4e71-91d0-277d708c92d3 · outbound

This paper cites Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.287680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:260a88dd9afa1d27d2eedf09885418ecfeae4e7c7fa274c583aa69184dc44ad2

Observation 3ac9a826-a65c-4a87-b305-f184f6274905 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.082967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:310b5c462fb071ab34d206578790ff88561b7b5770b2eea7a5e4e9ebd2ec6069

Observation e60859bd-4589-4735-9992-a89e9deddac2 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.103894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:95e2970f96f9dc78d8b9fd8feb318b2290b36840acc7e67a055959d26ee84409

Observation 737d77a7-40dc-464e-9c30-c372d0fa1ddf · outbound

This paper cites A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.122002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:6badd7e879ab53fea65e1aa2aa325d65c41ce0ccdba51cbaae6b40ef3a3ad12a

Observation 9e65d224-ec56-4236-b6b3-4d560624157c · outbound

This paper cites A framework for few-shot language model evaluation, Sept.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices A framework for few-shot language model evaluation, Sept

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.289493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:0e1c312330c99fe30b01a4426c4c55cae577bfba52b15fbb3d538030060c23f7

Observation 03476a30-b7d4-46a8-99b4-43e50c740663 · outbound

This paper cites Openllama: An open repro- duction of llama, May 2023.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Openllama: An open repro- duction of llama, May 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.291252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:7d80867f3c45852c93e03fd97a6b5fc0232ea225bee41a1c00e679c34cd9dce6

Observation dcd0370e-3fb7-48ef-a530-dac1ac6e84a9 · outbound

This paper cites llama.cpp.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices llama.cpp

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.292874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:72c2d26844034fa187c73ebde81ec6c6d3dd8d0bbfbd86521da795c4fd7c3ca1

Observation a3dc4b4c-cd0c-4d35-b5bf-42ac39257f17 · outbound

This paper cites Gemini: A family of highly capable multimodal models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Gemini: A family of highly capable multimodal models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.294761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:d0afd164595b9c5839ba5b6d07d520df947b096df041285ffd31bfe258226d3e

Observation a14c2259-bd3d-416a-a275-586dea7a3d24 · outbound

This paper cites Textbooks Are All You Need.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Textbooks Are All You Need

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.002241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:17587754151fc8d2cdd80c42d8f4f1f6e893fd7f8692d6ffd6f7ead76c934660

Observation 8e036432-74a5-47e1-9e8c-82efbaaafe0d · outbound

This paper cites Masked autoencoders are scal- 13 able vision learners.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Masked autoencoders are scal- 13 able vision learners

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.296658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:0b76256d0425401d0a2587365dea7f029dfbd3043870c546c3b1bdafaa9bcdef

Observation e77f146a-1e49-463e-8746-cdf617c5c504 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Measuring Massive Multitask Language Understanding

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.014977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:05a778d46f353d6ca17fabb902043a51cd31b5310796833545d30d1d163aec2a

Observation b4df69e2-8710-478e-8ae2-a72a8df3a3e3 · outbound

This paper cites Training Compute-Optimal Large Language Models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Training Compute-Optimal Large Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.028539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:6d85c2de7944a7ced16f58bdf9385630f60d37ee3130e5ba36d3219bad8cd3c0

Observation 5f16640c-acd2-4163-a6dc-76cca1fba923 · outbound

This paper cites Searching for mo- bilenetv3.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Searching for mo- bilenetv3

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.298229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:7bc1535fb6fad821da6d391af4dde5e7cbebdec4bebd3763fe53d8e09d7fe978

Observation 933f94f3-ccc7-41ea-aaed-e3320391d3c9 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices LoRA: Low-Rank Adaptation of Large Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.035996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:c27deb87a53941704f83d0b74c668e94d7eb21e0b35a5bdd3eec3f89de62aa4a

Observation a4a53bcd-cb2b-45e7-a417-8e883dbf48d8 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.300025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:dc8a2250fd3c5d74237b7ab189e2e16a4a09328eab69292311acf76a43f70764

Observation 8d6e5942-d8a6-41ec-b681-0b0676532875 · outbound

This paper cites https://huggingface.co/datas ets/Aeala/ShareGPT_Vicuna_unfiltered.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices https://huggingface.co/datas ets/Aeala/ShareGPT_Vicuna_unfiltered

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.301907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:6cfae56b25be607e0e4e4e89ea9505f6e4d8a2d1e57d2981b037c8d003b0b910

Observation d47b5c20-598d-4950-81d2-20fa4755a198 · outbound

This paper cites Open- clip.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Open- clip

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.303517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:4d0ae21c125b8057c78a49f2f04b96ea7275b791e55ad965633bbdb4bde899d6

Observation f69b14af-a604-4e8a-92cd-11deb3ac7976 · outbound

This paper cites Lmdeploy.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Lmdeploy

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.305170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:f6fd03dee9d456c148cbbef430c5a5d2ca5cc8a3aab6cd37475ec1cfd7b217d9

Observation e5a389ee-a58f-4d3c-b6d2-cb26d14a3b3b · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal co- variate shift.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Batch normalization: Accelerating deep network training by reducing internal co- variate shift

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.307044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:6006965122b870c8652c8c139d80143c4d682c7eee5a377259d265b7c60b763f

Observation d4a9c367-42c4-476c-a66d-dc3760b04540 · outbound

This paper cites Perceiver: General perception with iterative attention.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Perceiver: General perception with iterative attention

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.308809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:e643f756d143a6411e300013a7b17472f1cbd640425ddf6215e8115e69758135

Observation a2c1b2ac-18e2-461e-8706-2745423212b4 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Scaling up visual and vision-language representation learning with noisy text supervision

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.310803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:45ee7920b7d46718295af0a5a37a4389e82c4cbd40c9ebd164f97a33f45b0963

Observation e5b420d7-adbf-4aaf-af84-8af249c875f5 · outbound

This paper cites All tokens matter: Token labeling for training better vision transform- ers.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices All tokens matter: Token labeling for training better vision transform- ers

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.312677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:e9e66bd1247491729ae82056a054c2c588bb3230415aee37b8282ee6d7267de2

Observation 2dff6c21-50af-4b6e-a69d-1ee871d077f3 · outbound

This paper cites Scaling Laws for Neural Language Models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Scaling Laws for Neural Language Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.072521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:62f669a0afaf695beb805f4151b2d5b8e20ad09cea98e0594c96aa21a8cf4afb

Observation 49e4294c-b6d2-40a9-a847-65c624f3ab35 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.314444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:a1a778334a0f86a6adf67f94f0ef70ea26aac62c4d20cdd22f027b1d09837a64

Observation d8299f2f-12b9-49f4-b6f0-04d94acd54df · outbound

This paper cites Visual Genome: Connecting language and vision using crowdsourced dense image annotations.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Visual Genome: Connecting language and vision using crowdsourced dense image annotations

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.316229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:80845f44eb5fd801d8fd921f7d8f7342bb1ab9afc6ac7ba1ea93ee913bd8a917

Observation 0c7eda0c-6262-4e7e-acd7-f164e4293e39 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.090406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:56929a03803f2e7642266fe393746fd286c269a54b12840123da1c11d41f3c0c

Observation 5ec7f977-5bef-4bc9-841f-d3a5ff70f55e · outbound

This paper cites OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.099866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:259e6aa9124b79fc93a96fcd719df5b7ab630cd60fe07aded68a5e6d970384ae

Observation 68033dde-e150-45e1-af36-3de428b9d311 · outbound

This paper cites The BigScience corpus: A 1.6 TB composite multilingual dataset.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices The BigScience corpus: A 1.6 TB composite multilingual dataset

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.318061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:328534afd05369f0d25a4e7326f8ba25f60886f27ae71a1f328ff9801bd48d8b

Observation 7b99d196-54b5-41c2-b81e-1d91eb5086d5 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.107789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:3bb8161c17f994079f9d875f1f2d5c0d4bf74e561befd9d1f3977c45c0fab403

Observation 9d1f22ca-9f4f-4c36-b1be-13570d961154 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.319879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:154c17ab1517a202c131dbad68d1e5b68b0ad4e862803f5cd6ed3308163d1a34

Observation beb53d52-047f-4599-9340-e878f1420168 · outbound

This paper cites Norm tweaking: High-performance low-bit quantization of large language models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Norm tweaking: High-performance low-bit quantization of large language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.321618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:f20447da71ecf434796e1cac0c3a4ed2aa1be59b5c1a9fe58a1d9694faffc00a

Observation c6a0b4f5-b26f-4338-96cc-7f572b56764e · outbound

This paper cites Robust Navigation with Language Pretraining and Stochastic Sampling.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Robust Navigation with Language Pretraining and Stochastic Sampling

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.129434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:d54a5f24d336f09b53b1e240befe3718689754bbe9125cd9a9acad939ba92f11

Observation 8f619819-ecf7-445c-a550-cfa079e72352 · outbound

This paper cites Textbooks are all you need ii: phi-1.5 technical report.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Textbooks are all you need ii: phi-1.5 technical report

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.323332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:0b00e64e346cbb7c60b6b76e9a94b06b2b2f4419515b81b25630dea52db92926

Observation d0a8acde-0082-4f6e-8b3f-e1397f65f665 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Evaluating Object Hallucination in Large Vision-Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.135577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:d8b67258fcf70cc8387c5883087f032c4e16e1ff43fca6928534239da005c95a

Observation 865da393-55ea-4b98-8ac0-24f12e83a6d9 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.145468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:645f8307374a7ca6e4c6fc3da422bef76c4492ce2c92a6745d0b74557c46bc3a

Observation 34f16562-3adc-475a-b790-d53553048611 · outbound

This paper cites Microsoft COCO: Common objects in context.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Microsoft COCO: Common objects in context

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.324982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:436ac12f664007465ae5d60683b043789004d0765a9676410f2572f299baa807

Observation 57fd6cf6-0a2d-480a-a375-e2aa56453032 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Improved Baselines with Visual Instruction Tuning

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.165174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:d0e9a76bcbceb5a82183b494e73637695aeda2b4bb63551998598b26ff6599e1

Observation ac10e24e-ec33-4669-936b-9141b317ad2b · outbound

This paper cites Visual Instruction Tuning.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Visual Instruction Tuning

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.168341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:532fdf4c71adbe8799742844ff998d115062cf177ba1df695d6f5c252159e7b6

Observation 9f394269-8989-4ea1-9124-00d27ed7ad6e · outbound

This paper cites DARTS: Differentiable architecture search.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices DARTS: Differentiable architecture search

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.326694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:bbbda5441f1541709d7ff1d5279d0f5f8e5cd35f7b554d350ea5d1a6e84f10ac

Observation b96c18cb-aa7d-4802-81af-1c9b31884c8a · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.185520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:9314d59d8bb5a0f2cdad8b39ad915b91f8d7a79d1ba487c0bbd9f0939ec8e637

Observation 47c53e94-dbcb-405f-a482-606ca0db2f7b · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.188510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:91a2e0e286df37ce9d36a61e3dc8a06cc31fc94f50a272138a9ae0c31d054cb4

Observation 0139974d-b3f5-45f7-ae89-6004eb77eb88 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices MMBench: Is Your Multi-modal Model an All-around Player?

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:37.997679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:5833063aab53aeec8471bd14a9169c0eefe719637b95f2a65b5911b7e1c68439

Observation 0119f562-4fbf-4861-a8e2-ca9ab1224706 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Swin transformer: Hierarchical vision transformer using shifted windows

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.328623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:a6aae404bdbaf05762b63526f6436cf4bf4b3e393afbf973900292c7f5619667

Observation d2c63f41-a7c0-4422-b4dc-adbff138177c · outbound

This paper cites Decoupled Weight Decay Regularization.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Decoupled Weight Decay Regularization

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.006636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:7d1053612d4aaf575e01982f12d2f14a538c699c29bd7a71e0be6776f3569f7d

Observation bbc97eb4-c8e9-4d24-9850-dd14cd9582fb · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.331717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:f243e357c52ad156e5784c885c6fafc826bfbc6c53e7e1ec5978b032cafcc178

Observation e7a12bf6-2719-4ab5-a4a7-53fc837e0db9 · outbound

This paper cites Llm- pruner: On the structural pruning of large language models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Llm- pruner: On the structural pruning of large language models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.333807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:2e9a9a21970a66778c6840516dd9c7ab6a8e281fe1f3312e63635a7b51be5f82

Observation 1ca7654c-62d6-4396-b761-4f444b76aa15 · outbound

This paper cites Point and Ask: Incorporating Pointing into Visual Question Answering.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Point and Ask: Incorporating Pointing into Visual Question Answering

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.019655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:9050cf66b472440e56901a55d07626839a33cb9d7915d080ba11fc0b11932199

Observation e7427b10-da64-4515-bb17-faef5e9396a3 · outbound

This paper cites EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.024396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:aea0c10e1e3eb5b2dfecb5eacbc23a3df74ed51034c3e665f2a513b690301692

Observation 148cc151-c4dc-4e4b-bd3d-06c5ae20a752 · outbound

This paper cites Tensorrt-llm.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Tensorrt-llm

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.336388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:035b3e22bbdff68d52662de31e512725181099bd8799bff20f4ba6ef157f1131

Observation 219821da-8d84-4602-9232-96b1e4c2e5b8 · outbound

This paper cites an unresolved cited work.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:35:38.338160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:1361052943510d903a637fdde5aca51c920da12ddb8bcfb44f04cf33c8169cfe

Observation a6fe276b-6f9b-409c-b73a-51ad5964071d · outbound

This paper cites an unresolved cited work.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-05-16T16:35:38.339857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:bd4c98240652509ffebb95e6183d57e9f7da024e6db1f94fba81a38dddc71a44

Observation c07eaf40-cdad-4231-9321-995fe585770d · outbound

This paper cites Gpt-4 technical report.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Gpt-4 technical report

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.341623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:61703cbee75a6354ebdcab3c241189aa3782ccbb298a607aeefa2d43f9114e55

Observation 071a0a7d-d622-4f73-8740-f238164f9154 · outbound

This paper cites Gpt-4v(ision) system card.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Gpt-4v(ision) system card

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.343282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:4176ceb63d295477308156d0ca1491ae1208891431ec8a14be957e48cb7390cd

Observation a241d2cc-0194-4612-b2e4-45821d294f5c · outbound

This paper cites Im2text: Describing images using 1 million captioned pho- tographs.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Im2text: Describing images using 1 million captioned pho- tographs

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.344986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:c6caa36ed89cff0d80fde9937249629128090746cae4679878aae349e3ee647a

Observation 9e0cdb38-0e3c-46b7-a3a2-d4724feb6298 · outbound

This paper cites Episodic transformer for vision-and-language navigation.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Episodic transformer for vision-and-language navigation

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.346786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:87b447a1e56629ec770a94187ac4bbc1ecf918c5ef369906ca349d35dc4eefac

Observation 57434338-5bc9-4041-984c-35b1187b4727 · outbound

This paper cites Tinyllama, Sep 2023.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Tinyllama, Sep 2023

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.196733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:7ae3d642d33243a430db9115fe03b7ddfc782d6dac8a98a1cf958642926db606

Observation ef5107f2-378f-4f46-8b3f-c2d406fe2129 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.056229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:d372610d80b056bc73a375f8d39fa2dd3a7735766b53b66f5a1ab8246d3758af

Observation b120361c-df5d-463d-9cde-33f82eb435fc · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.199106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:e438a90bc5479c7e5c1c1c2fc87790fb18f122d61c74dfadfb85f6e952159938

Observation 5a5bd5a5-7d98-44b0-aa37-ad611722d250 · outbound

This paper cites Exploring stochastic autoregressive im- age modeling for visual representation.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Exploring stochastic autoregressive im- age modeling for visual representation

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.201294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:76bde64c83605e920651cb862f4b813565b5839baf30e0dc7b4ae31846700299

Observation 23e90195-f4c2-4054-9de8-e3cb20e3aac9 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Learn- ing transferable visual models from natural language super- vision

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.203485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:f8b85be16c91a52a75ed0320ec07ed4450f9bbfb8cb6fa56f80df2a8ec153a76

Observation b2496a81-9dd4-4c1c-b9de-b3e600c5a2d6 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Zero: Memory optimizations toward training trillion parameter models

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.206091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:ffc52f5dd2332b041655d5e9e207fcaf6746a20b6ab419ec4c716766c01d2461

Observation 93002dfd-b06d-4bd0-9ad7-4d1f53fa3353 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion param- eters.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Deepspeed: System optimizations enable training deep learning models with over 100 billion param- eters

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:35:38.209180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:7ec64679ced52eeca0a92e673471e2c2ed137b18a022d4aa8caa8ff2108beb70

Observation bdb84ae6-aeaf-49be-8af4-16da472a8658 · outbound

This paper cites ImageNet-21K Pretraining for the Masses.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices ImageNet-21K Pretraining for the Masses

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.079877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:2b35dc865bdf5eb53f38b4990d4688df9ea715a17d707711f6fe482a21a8735d

Pith citing papers

Observation bdf9622f-563f-4b8a-9dc3-0b14dbe032d9 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:60fb6d66eda650a53dd59b8e3d5ec812d4a23584018e0362a499115062fe4a66

Observation 4077098a-09da-4095-91e3-8d1f95e1a3b0 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:c5ddcc865a4d91906f0baa2f501b099e39f0d162d4af03b3160197a24208d268

Observation b230a10c-ff3d-4239-9288-978f8af0d072 · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:27:51.932747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:27429e9b5645f4985758898c83ee872203bef859644b3678777f849a0712bf56

Observation fa213eda-b313-43d2-b670-e8f7361ca41b · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:55baea7a2afc4a28a1b99f39dd6380192918eb68cd4e8b216c67e2669c2f43f9

Observation a6f2d595-4bbe-4f80-8f1a-d7393d254b97 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:35018c9f57a6b88a626716231a30ef78f9de727764b8985a9c7d58a866811211

Observation 9f623020-5c28-4b1f-bd92-0ba607bd4a3c · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-17T07:44:47.522081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:edf027149f4fcb470273dfd695a013831bb767d52500135bb2378ccc6959f472

Observation 2d8e3930-b40b-40cc-8a51-14a1528b2da5 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:91d4d3ed24fa09743ddf8af3537c6d50da931549374fb793191d528f41b82505

Observation 609ae3d3-49e1-4bf3-95bc-b2c8c79df22a · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:a02b3014c526c084af5fd0e6d4d00ddc3e436f54e2f4158333e68cdc2f991dbf

Observation 3a7116c2-28ef-48f8-b48b-af9dfecf4bad · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:669ec69405352c4ead659276a394509f011bdd2217277322f38c19bab7622c8c

Observation e24af0d0-3eff-4685-b204-4cc8bd709287 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:84be21cef162e4f138f9cd365bb839080a5eac916a906bafacb55cdf63dc26f0

Observation 75d58e03-1e58-44ab-838a-c5e00e0f2415 · inbound

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling cites this paper.

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T08:14:52.890145Z digest=sha256:938d786dcb544b6af2c97d5af097421f59f281853d9df7ec2288ae2ccb7d05f9

Observation 9cb5ae6f-f0ba-420e-ac5c-d1904fbc2810 · inbound

Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance cites this paper.

Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:51.199560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:51.199560Z digest=sha256:069421027bdae6867dab8b2b0ce96c9d5dc2dd0600fa2662958c0d6e7f5181b1

Observation beb49cf1-6177-438f-96b7-487fb8f11626 · inbound

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning cites this paper.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:17.046484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:17.046484Z digest=sha256:0244c26a4e779afc4c171f0fa243a03dc078b86d83eab14499e69d8d72e34b94

Observation c21867b7-2676-48dd-9917-21e5fb1f3d22 · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:53.979850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:53.979850Z digest=sha256:245e7f86b1092e061898782645c4d609811cdeecd2577ce100c0a30b8e2d3dcd

Observation ee357cca-4315-48c2-acb9-af6cfa5369e3 · inbound

AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving cites this paper.

AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:21:49.925595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:21:13.618080Z digest=sha256:3b9b2739b95227386807edfdf5b8a91a6d4830ca0423ad0eb048fa84af7b9523

Observation bf7c652b-a427-46d0-9d49-b5285bf40d4e · inbound

Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching cites this paper.

Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:16:27.818944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T14:13:48.523955Z digest=sha256:332f8a45a003e893e8840d59d820f0ba8083864fc1d5b7fcd6983401364ab286

Observation 758da1d7-98a0-410d-9e64-8ddcd7f9aaac · inbound

Multilingual Vision-Language Models, A Survey cites this paper.

Multilingual Vision-Language Models, A Survey MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:02:37.167925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T13:02:08.000814Z digest=sha256:0fb770b39b22ae4d3e6a6601e80554c9fbfc56315d15799af88622f6ce22d99d

Observation 4e6be4e9-b84b-49bc-aee8-f89e98b47fba · inbound

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models cites this paper.

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T21:31:34.208111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:31:34.208111Z digest=sha256:ea9d4d2f01894e94f3f4cbcd3afbec1cd0e20a8ae47b3936360f94193c07c8d4

Observation 107c6774-e443-4879-901a-287cba4a438a · inbound

MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments cites this paper.

MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T13:51:32.354806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:51:32.354806Z digest=sha256:3dd88fd4b6191cd571ec0dd0a1e9a906b2390ab5c796fed1e1bb0e7ffd3f30a9

Observation 0840b801-c090-4eee-8794-7591892f77bc · inbound

Vision-aligned Latent Reasoning for Multi-modal Large Language Model cites this paper.

Vision-aligned Latent Reasoning for Multi-modal Large Language Model MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:48:23.272002Z digest=sha256:3d499cb8f25f2b8a7088cd4f7f00717835b28bab90d890f296ba57b530e00fbe

Observation 32a198ba-a37e-4d71-81b4-07ff5d605fdc · inbound

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs cites this paper.

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T19:45:33.950587Z digest=sha256:975b32aae4d688b9f1c77355d2a441cb52e6dddecfc2605110ab89c8515f4851

Observation 6aa81487-166b-4bac-a11a-a57a057b8454 · inbound

ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality cites this paper.

ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:28:39.156433Z digest=sha256:470ae6fcd37590c6b45b8b7c667f9a01d6bc32a86d2ccb83e89b9284b70b0759

Observation 6be11098-9650-4c8c-91da-527d7c5d7eac · inbound

Leaderless Collective Motion in Affine Formation Control over the Complex Plane cites this paper.

Leaderless Collective Motion in Affine Formation Control over the Complex Plane MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T09:18:06.363249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:18:06.363249Z digest=sha256:871c56c492df4c6138588c3778c9409c00a5ba95a8fd15ad2981074ee2d31031

Observation c15fe634-21dc-4ee7-8944-4d008ea7cbcf · inbound

Analogical Reasoning as a Doctor: A Foundation Model for Gastrointestinal Endoscopy Diagnosis cites this paper.

Analogical Reasoning as a Doctor: A Foundation Model for Gastrointestinal Endoscopy Diagnosis MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:01:47.592716Z digest=sha256:5981d6d890e2ce7fa9c8a18866ca16095588929a302e5714e48b87723e503ef8

Observation dcf7a7eb-3404-492a-b339-c9dab0d6710c · inbound

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models cites this paper.

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:23:46.371799Z digest=sha256:dc1e479f269009ddaa38126dae1744eaf619a12c6aba2d07462854f3b064e025

Observation bf461f89-2f74-40b4-827d-c1c8a4055f2c · inbound

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency cites this paper.

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:47:38.100037Z digest=sha256:4d24a736ac8460b52b0713bb33669f0ea94d88cf86276d4c5c35854a26bfd87f

Observation 649bd923-e9ca-44a5-82c5-65922f10d106 · inbound

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT cites this paper.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:c59977cf284885a87c0c9579c540faf2e2bbce62a292a55c809094445cd76090

Observation 02e3ff45-293b-4d05-8686-c92064472421 · inbound

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models cites this paper.

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:10:56.991368Z digest=sha256:e7ea2e38264a5238a7c6e8303a72c9475775c3ca2774b20b40836ec22e2b5662

Observation 07461f87-4d39-4091-bb17-31721233a530 · inbound

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models cites this paper.

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-28T20:22:37.626275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T19:53:36.990878Z digest=sha256:9e2c459f1a5132c5c89554cbafc77426f682b10f38374d172cf69669c6c873f3

Observation 0212a6fe-2139-4c8a-8a94-22d46d42a349 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 284

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:39:37.659042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:460ccbc433a7f9a2bdb14296274dbd7f9a8fb4142479f18c719146b6ef695ef9

Observation ce4311c0-fd57-4c09-b2c0-151d8bff0e98 · inbound

Phase Matters: Characterizing Heterogeneous Vision-Language Inference on a Mobile SoC cites this paper.

Phase Matters: Characterizing Heterogeneous Vision-Language Inference on a Mobile SoC MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T18:15:58.582593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T02:16:11.064614Z digest=sha256:ccd41b6fc721d787ab8c8272188a00f3c973a5e51226e04cf5f3ecc3bfe53e1b

Observation 70f44998-7875-4e2f-b2f5-044cd5fb3dc5 · inbound

XS-VLA: Coupling Coarse-grained Spatial Distillation with Latent Flow Matching for Lightweight Robotic Control cites this paper.

XS-VLA: Coupling Coarse-grained Spatial Distillation with Latent Flow Matching for Lightweight Robotic Control MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T21:12:37.909441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:12:37.909441Z digest=sha256:c8e0cbec7bd0b7545ff63a074fc0aa81c15f659e40e47d2e3e9322f8a770b5af

Observation 127e1a24-b813-411f-83a6-41c1d17c4cab · inbound

XS-VLA: Coupling Coarse-grained Spatial Distillation with Latent Flow Matching for Lightweight Robotic Control cites this paper.

XS-VLA: Coupling Coarse-grained Spatial Distillation with Latent Flow Matching for Lightweight Robotic Control MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:44:07.519769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:44:07.519769Z digest=sha256:b050a4db4497e7bc4a4b18ea7b83c12aa9b1c3c954abd45864a5d6bce3b25505

Observation 3b719c83-217d-4fd3-87c7-9729ce56282b · inbound

PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving cites this paper.

PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T16:01:25.344820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:01:25.344820Z digest=sha256:fbf24bbfbf36c4422deb0a77a30924c5e36d4cf0b8c57aac2b9cea4bffbd050a

Observation 1e20b9f0-c704-43bf-9450-8ca3804dcc56 · inbound

Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment cites this paper.

Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T13:27:05.578693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T13:26:41.210163Z digest=sha256:301abad921db1fdf3d868974b54b9e4ac20fd45bee4d28379c251c98bc5570db

Observation dd5adc42-43ab-41aa-9abb-dd4a7f7e8259 · inbound

Large Multimodal Model-Based Environment-Aware Mobility Management cites this paper.

Large Multimodal Model-Based Environment-Aware Mobility Management MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T15:39:36.100888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:39:36.100888Z digest=sha256:3d56330336a60197647218fd56f4681f6f2dcf0d10683fdc4c7e181d102914a2

Observation 9675f9e0-1450-4c98-8d71-2426b9e48ead · inbound

Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs cites this paper.

Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T10:06:48.034588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:06:48.034588Z digest=sha256:d11110f8e60b381f24d5aac00fbc017a575ecc2ce834fc8bc1d11310d7ed92c6

Observation b3d1b04f-6506-415e-83b6-1b300d14a593 · inbound

Offline Vision-Language Navigation with Geometric Goal Localization for Outdoor Environments cites this paper.

Offline Vision-Language Navigation with Geometric Goal Localization for Outdoor Environments MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T05:28:48.437022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:28:48.437022Z digest=sha256:66b55da71a4a9ebf541bb81cf9f3adaf2bbe1e839303bddf5f0ef5c714c655c3

Observation 5f0f4f26-7421-4712-afee-411e59cb818e · inbound

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models cites this paper.

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:49:57.611514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:49:57.611514Z digest=sha256:6a202932283219b4a45a982c6e8f4ccc11103452f2affc51bc86d8b93c724e28

Observation 33549019-2312-4c34-9195-99dedf88e1c9 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.818128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.818128Z digest=sha256:5c5d3401e6789cc8675afc6d2bb378fe76595d9cc3b27a334543c7c675632f9b

Observation 8693891f-9867-4b9e-a551-bceb8e669d47 · inbound

HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving cites this paper.

HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:28.003330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:28.003330Z digest=sha256:72123f301635e87cb4947ce49a4e1a9c754f389cea6e07ea1d29ed0874c8d75e