Pith. sign in

Paper Citation Record · LEDGER

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

As of 5 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 73 inbound Pith citation observations for arXiv:2410.04417.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.04417 v4

Coverage vector

measured 100 of 113 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T14:58:32.303101Z

measured 173 of 173 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 73 of 73 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:38:45.767775Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T21:36:34.338168Z

Reference resolution

100 of 113 outbound references displayed

  • verified exact4
  • verified fuzzy73
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch11

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bbe4c423-ab03-42f4-b5e2-fe252c8a6b48 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.688841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:4e25062f64f7138fd7153e16e736351a9349ece9ddd1f9625c45e9be721b98ae

Observation cd787a66-e539-4c2f-8b84-95613a1b46eb · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:58:32.371498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:9752b54bade3b9a30a8dc47fe9d853d15f3643a2ec2a7b2151a408075015657f

Observation 1a401b2f-33b0-4831-b485-4e2b6f5922a8 · outbound

This paper cites Token merging: Your vit but faster.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Token merging: Your vit but faster

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.694311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:0d9cc262615a75ad8ca8e6828ed83e00f45a6591b8aa9cda0a555714203b3f17

Observation 1bf9eeb3-61b2-475b-9457-47bfafa855c9 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.697775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:40c4fcb03982ef6e87ca9ee9afdf2a4c8d64f7e954606111d6622dfedbc20f90

Observation cafe519d-cde9-4316-b842-b24e6df4374e · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.701238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:fe6892cca675d306ccd54fc0ebae58c20d4b399aec648c9fa5ae239301595763

Observation 7f725070-82e0-43cf-8883-b6dc3f650174 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Honeybee: Locality-enhanced projector for multimodal llm

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.704832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:dec4728296c12bdf2da2f6eaa3a36662e396285f40573a22cb2570b865012d30

Observation 4b3d7932-8a27-4b6e-a91b-f19c9a050d7e · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.708547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:b20d6b251dc18752ee36abf2fc5ed4950918fa879cb8507b13c094ad87b7e2f4

Observation 76c7ad83-e08f-4570-a827-5f62e76e5301 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.711855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:d6452773e1cabee9453b7992330ddede4dcb9d4ef2269095a0adcc8ae30b7f13

Observation a57b32f8-ea2a-4b29-b11b-3d09a8ca36f4 · outbound

This paper cites Instruct BLIP : Towards general-purpose vision-language models with instruction tuning.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Instruct BLIP : Towards general-purpose vision-language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.714936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:33c6f11008792cf95319fe9cd5d6cc1d1c4f36c5b75ac6201d746b19065619b9

Observation 254cd483-b4ca-46a6-847b-d0ff915f4d1b · outbound

This paper cites Flash A ttention: Fast and memory-efficient exact attention with io-awareness.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Flash A ttention: Fast and memory-efficient exact attention with io-awareness

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.718501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:fdab56dc1029fb954ef5214c223b34933d83d034f4e3c68a3d4e3c0ae1a3f4b5

Observation ff3b4942-7816-4ee2-a132-6ac3117703a5 · outbound

This paper cites Glm: General language model pretraining with autoregressive blank infilling.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Glm: General language model pretraining with autoregressive blank infilling

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.722087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:66bf3df675afde7eb29712c695cf82d1811ab25e5fd71d34d5a963273642d979

Observation 92526777-bc9e-47d2-a87d-a1b12af27062 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:58:32.395041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:d0080440efaa36706dcc9923cdec1e1301cb1d3f2a40e3a096f082e7b3ea5b12

Observation f9c66094-2fab-4ff3-a26d-645bd5c45850 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.725577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:5c730d03ca0a54f0d13d73974f536688bff59c7a71d56a389b45b39bec554785

Observation 8aa3d3cc-adde-4128-b499-10de1ef2cafa · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.728853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:6f0df20a08b1dd9ab1b1a4b0257d2bc373810f493c9c13fad9489d1ed5720524

Observation f4a288f2-9cfa-4c01-b2c2-7f927990e616 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.732110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f5df79bcd9af5aa298f3f58a084f95dce0a20223a73795fbbc0f841b269988b6

Observation 641ea025-3580-4247-ad2e-b03e7e3498d5 · outbound

This paper cites Videopoet: A large language model for zero-shot video generation.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Videopoet: A large language model for zero-shot video generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.735220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:127399ba8fe77d101785b51fc10461269bd65b99074719473deaa36aa116a4e7

Observation c5880045-832e-43ec-8c78-0c13e6d2cb5b · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Seed-bench: Benchmarking multimodal large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.738307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:fc4ade1c1bca169ead2b43807a3ddeb442dd4d2975b629d76223d49ea6ff7225

Observation 52163e20-d853-4853-b04c-3f3b9d4b32b8 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.741445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:dcc5f28bc4bdb5c5cc5c8803893d90dd4e61ede459a161951b6bf30eabab707a

Observation 04fb6005-a86e-4d52-8934-9e354e6c6090 · outbound

This paper cites X., and Wen, J.-R.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference X., and Wen, J.-R

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.411376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:aabc5946dc0bc98934150de39b584a51c95ddbf59fdb1b7cddc15b7bf5fd6838

Observation 546aa8b8-bc21-4b7b-ba8b-9a0e60644508 · outbound

This paper cites LLaMA-VID : An image is worth 2 tokens in large language models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference LLaMA-VID : An image is worth 2 tokens in large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.414752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:67ab2e9c82920290dc38706d27351d61ef5b154138b531097af9f2c6f3d04294

Observation fc677043-36e8-499e-b951-f40ebe3202d3 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Video-llava: Learning united visual representation by alignment before projection

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.418273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:9aab58ba43ddb93abf402f37bcbc6e45d75790d1fb6a180136e8490598d1bca2

Observation 1e1169e2-7dc8-4c42-977e-28830054fe61 · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.421767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:a1cb77c9cbbd3379b7b15991a77c5a62368a96d7bd0d78efa47ff770e86a5dda

Observation 1670af51-a156-45e5-a905-798c7dd2a347 · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.425288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:50ee0390dea6839173656dd673527c62f798ae4c01c61ea0154ee74fc54cc91d

Observation e716eee0-0727-4e68-8209-d2badfe2593d · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In Proceedings of the European Conference on Computer Vision, 2024 c.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Mmbench: Is your multi-modal model an all-around player? In Proceedings of the European Conference on Computer Vision, 2024 c

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.429539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:8b44573cce35c7a7f08ef6e24ffe8466997d91618ab5b6fa5ebe1c3c544e4568

Observation 0b880c23-be0f-497e-8c54-10518daf84cd · outbound

This paper cites A convnet for the 2020s.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference A convnet for the 2020s

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.433274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:70e9dc01c343d1ff8a1098b4af4ad8717528b9a7c6574d4768476a9b9ab6833d

Observation 1384f337-b89e-4f90-b802-c3f1b397ba53 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.436749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:3ccc3112a083b3045803018e7892342247896366ecf1c7eb5cceb4c34b02d7d7

Observation 96097744-1bae-417d-bb5c-5b618c1f49ba · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.440476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:c3b83140b896a2b0327cbbe52a1f8213b19aa69aeded0c9cfbcd1c93850f4186

Observation 40912adb-d435-4732-b824-833bcb3049bf · outbound

This paper cites Vision: A computational investigation into the human representation and processing of visual information.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Vision: A computational investigation into the human representation and processing of visual information

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.444161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:587e8918ca8310cdca597e3a9b3b55a661753155d9bfad70c01440fa7b41ffa1

Observation 1324da50-7d94-4321-afd7-6cebe2d1146e · outbound

This paper cites Language models are unsupervised multitask learners.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Language models are unsupervised multitask learners

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.447691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:13181d1c6420b5a057098d0884820c07d8f25806de36a3cf21c13d4c3b3c63b0

Observation 39cbe1d0-fb5e-4aa6-a2a0-fe81294ae6bc · outbound

This paper cites Clustering by fast search and find of density peaks.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Clustering by fast search and find of density peaks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.450803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:93587152ba761269c882aaf30428e98604e3ff870ff3d965771cdf0e37a1622c

Observation 4cf49f61-7616-452a-ac68-1acacc84ba21 · outbound

This paper cites Towards VQA models that can read.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Towards VQA models that can read

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.454218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:8ecb6d1010a03f642af5028bee2fd0c03ab6df06334dbaa4445946f83d3a157a

Observation ab59e97b-3e02-4df7-bc2a-78daf15a3f6c · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.457610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:df819d5611024304c26e84a3115695883da1ca2e141ea3f415da60bbe6cdc2fa

Observation 0f25f953-2a84-4b5d-8723-49d9b3ec9fe5 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference N., Kaiser, ., and Polosukhin, I

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.461069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:274c0dbe5b64d055ea06a6be973b7629bf13a78c6ca178e336faa8dd02c8380c

Observation 1a0f098d-100b-4a31-bd7a-cf6bd86f49b7 · outbound

This paper cites Q., Wang, Q., Gao, Y., Xu, Q., Xu, T., Hu, Y., Chen, E., and Shou, M.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Q., Wang, Q., Gao, Y., Xu, Q., Xu, T., Hu, Y., Chen, E., and Shou, M

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.465147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:b3779694102bf0cdedc78d6e20b01ffc815762548695046ab13fda9b0ba63bb3

Observation edc6858b-c696-4eaa-b4eb-28514079b42b · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.469235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f4d8ca0d9c85f87cfbc843348b5fe704dca752ba1b1ec7155bf3ffa644fae3a5

Observation 4c61b5a2-368f-4dc9-8d3a-ac6c510359b5 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Video question answering via gradually refined attention over appearance and motion

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.472826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:6af8091c7847ee83768d78b55bf46733f161377e12b7487f8e168f6d2235a37a

Observation 139efa7c-ec3a-43f4-b657-741794c2b013 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.399310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:b4794fa917f26532a2fb417cfa31f6af9c6ee327ca1c31f37125a3d5f5506905

Observation 20e3e900-2492-4fc6-94a2-4a48e6faa88a · outbound

This paper cites VoCo-LLaMA : Towards vision compression with large language models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference VoCo-LLaMA : Towards vision compression with large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.477231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:63b22d5a5457b41d49e3e9f077ca127752aea646ae26b82c789cf93712e0e256

Observation cf1e51f3-4088-4d9c-a1ea-a29d20ea99b3 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.482247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:36ee04a256b9b1c781403d5162e3d2d779907a1c97094b9f39c3c1fa0007fc1a

Observation 9d820b12-87ab-4189-8888-df4f77de4dea · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.487154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:8b7581730b0c6ae9dc037135394aa3bd0cdccda37f43a6b1e5a439785c856f5d

Observation afc66c46-bd6b-40b2-9fda-f8f806fd62b7 · outbound

This paper cites Unveiling the tapestry of consistency in large vision-language models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unveiling the tapestry of consistency in large vision-language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.492153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:1022f464680f4938b85a7da8c6de7cfce5474f7d2563889daaac2a6b4048a794

Observation e7e7e5a1-7f14-4504-905d-66405628a671 · outbound

This paper cites Freekd: Knowledge distillation via semantic frequency prompt.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Freekd: Knowledge distillation via semantic frequency prompt

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.496612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:b2743375d3e9798fd1ea5e6fb811c135d21680a104e029cb6af57e3221457b50

Observation b6b3b280-f33c-47d6-ba60-6b2a9b2c8f5c · outbound

This paper cites Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.500572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f4773412e93c35c94caba7295b25cbdcfcb70a201dd729f78735d5b77f64bc77

Observation 9dbea14a-1625-433b-94d0-474443087fb7 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.504948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:817ceac65dbe804b5a4f85f0ff8ae330a0c0771a062617f5641c50a271d8b726

Observation 62c34109-145b-421b-aafc-86aa2e22a468 · outbound

This paper cites Scaling Learning Algorithms Towards.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Scaling Learning Algorithms Towards

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.509259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:77176639548b6d880a6fe5c4d6515ee221f4d0da042958b062372953c73c9648

Observation 9ecf8bcc-c241-4b63-814a-1dbe9498f065 · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference and Osindero, Simon and Teh, Yee Whye , journal =

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.516390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:77b2a8a031372d8c8fef0f3626e2a258978b9436ffeee3ad1e710691227bc14b

Observation 185d2150-c75a-4db9-9081-0ac4aa0fffc2 · outbound

This paper cites 2016 , publisher=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference 2016 , publisher=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.520258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:d76cb5a5b6c4cda83aa3b7b5f9eb1bb35ac18257e34914362c7363b0b0c342b7

Observation 1d4ce623-25d2-4c26-b28d-886bfc5eaeab · outbound

This paper cites Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part IV 14 , pages=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part IV 14 , pages=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.523886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:2635bd8665ca99a8e8ce62b8318d77d6a498b977ac2a184b54627d7baa918474

Observation d402cc9e-6924-48d8-9f7c-268614322812 · outbound

This paper cites International Conference on Machine Learning , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Machine Learning , year=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.527701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:e97123a4c4635e08cf52560787ecd7781e30fd52630facc10fb7879ab8bb20d9

Observation 27a38d1e-fb27-4d06-aa45-86c2f34827d4 · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.530801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:5e0e3d6f0681b989d005ca2a4b00d0d50ff9c695f425a51833c3ea26a907ca1c

Observation f56280d5-867e-4094-8128-348abe91a2a6 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.534089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f3511a61888ff86facaf780d810605625fe2ac773786355f878e1f177fe7b92a

Observation c5973323-64d2-4a50-9985-f79741b4a740 · outbound

This paper cites OpenAI blog , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference OpenAI blog , year=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.537061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f0433fef0ace71a00eb5e0a90a827996832cb0feb68d10265994154f2631a287

Observation 3c1983c8-290a-4a6f-80c5-10ed6eea9450 · outbound

This paper cites GPT-4 Technical Report.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference GPT-4 Technical Report

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:58:32.344640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:6e0e6a342b92ea983cf3ba0840e79ef50751ec2edb60100b58555f376b58e86b

Observation 82166124-1841-4514-8bd9-a1f346978f1a · outbound

This paper cites Deepseek.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Deepseek

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.540566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:ce59d3e22f9ca0735831300c70e4fce7cdac99ab93f208b2690e6872afd218b7

Observation cb085220-253d-4d0f-9c7f-a94f8ee1f1a9 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference LLaMA: Open and Efficient Foundation Language Models

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:58:32.375873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:873c4b7c70c992a86cb5c89faf6f90bd1d821cb56e40793422fbe3d070a2cd2d

Observation 73df4974-5e08-46fb-bc36-d5f822fb6dbe · outbound

This paper cites Instruction Tuning with GPT-4.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Instruction Tuning with GPT-4

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:58:32.379830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:acbc1a7b279cfcb9c2edc91e16373bd0d619a758cb1559466f4495c8da8d1b1f

Observation 38f1d31b-8b4f-45ca-a2cf-465ef4bdd60d · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.543823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:6927de9c50dabb759ceba4746c13d09fa3fdc215628fce51957244ef2823f377

Observation d9bc5d36-7143-40cb-95ea-5256385aa9a7 · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.547628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:53d8f145bd31fd7d58cbe2a268d902218b61f1b67023d1ece7635866ea76159f

Observation bd1f434b-ac0a-4523-83c1-3a224f41f5ed · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.552587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:32a3a068bfb051766b5701c5f4c3ed3d910a50328ac7f7225e875a00691cf530

Observation f358fb2f-dda9-4cc7-86cc-12fe91e0de39 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:44:47.683285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:fbcaca223fdfdfe3d087ff3abe1734949245eae96ee0244d1bbdb03afdcf9df0

Observation 9a415b51-90f2-4a9c-a614-397ca336a455 · outbound

This paper cites Proceedings of the Annual Meeting of the Association for Computational Linguistics , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the Annual Meeting of the Association for Computational Linguistics , year=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.556362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:400bd9941e5c3ebb7b13aed474fd604c2e917621251d9a8f76f50dc2f6d53672

Observation 392c2795-8ad4-402a-9ede-dffd78bbd27b · outbound

This paper cites International Conference on Learning Representations , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Learning Representations , year=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.559685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:b6f0a728ebcbdf00f1dfe316a4b24892052f8fe1b032f58ac24a04ae86fc3c58

Observation 3a6be316-e1a7-4ae3-b40e-c1f41febb8f5 · outbound

This paper cites Proceedings of the Conference on Empirical Methods in Natural Language Processing , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the Conference on Empirical Methods in Natural Language Processing , year=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.563762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:ce7425c3898b8d7f2d798c001c3f191a6019fb27cf9a48bee0929f0bae34be4b

Observation e06f6570-92e6-4b4b-88b7-f28a537fbc80 · outbound

This paper cites Qwen Technical Report.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Qwen Technical Report

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:58:32.384055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:bd054a7b7384087582b34a6b0e29696271e3cf912db0c4e65ba83467512cbc0a

Observation a160a8a3-5b59-41a7-ad44-a867a787e175 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Gemini: A Family of Highly Capable Multimodal Models

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:58:32.391559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:fc473d6f8ea6a3d7bc404f75e79ba34bf13709247c4085b39a950a11485d9684

Observation 360ab59c-956e-4d05-8d77-280c5bdab154 · outbound

This paper cites LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.566826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:af3616dafd684604ccdd7a608f6cdbd05f6ca3c6b4b040755b8b34057362ba5b

Observation 8dea3b02-ea03-421a-8155-ebd80b743fa2 · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.569559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:b190de8bf56e9bf46a44aff30f02a2e7c23a37730041f83f5090c1bd35e1686f

Observation da95dcf1-1602-4433-8d9e-d8128bcd57fa · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:58:32.403358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f460aa9051cc8d13129631f4a0f479f4204d492060fba9d08393b7053f4ea770

Observation 724ddaa9-7e39-4b2a-adf6-984c04aa1b4f · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference CogVLM: Visual Expert for Pretrained Language Models

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:686ddbaf01af607861da2ce1a94948170761c7941b514118c1d97ab6fefcbc3d

Observation cbee58ce-9561-45fb-8ed3-bd9709de813d · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.572136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:ce0314c2f1b71c93d728b4978928843a28143fed8a8e38b139895259578a1d21

Observation 824b2051-9475-457d-bb27-6fbce43df43d · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.575118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:853ac3b1efb5df7a6e04cf63c489e38b4694fbe3ccd791f0225b5b9d28de247c

Observation dca7e261-89ab-422a-802c-6b3872e4f7e1 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:58:32.359216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:69c58dde25cc01869f23944c94289f7c9e82ac0f263b967f921dceeb8c91d8e7

Observation cc6fb78a-8f51-4d87-a3fd-bbabd8d802cd · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:58:32.367631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:5359485f68714d3941242ff0fb792c816fd4266d07155b19bcd9e707264d29dc

Observation 25168446-00a5-41c8-a819-b503e95ef52e · outbound

This paper cites International Conference on Machine Learning , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Machine Learning , year=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.578134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:9308464896acae90dd2a5bd6b1382a8802eeb5b30972f78cc6a60e15f8484875

Observation ed1294c9-16a7-4cdb-b083-74ffbc8090bc · outbound

This paper cites International conference on machine learning , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International conference on machine learning , year=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.581052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:ba47c4c6ed92c1cfe458ae870b7d4b98f3874df166376c039d29921f6933ace5

Observation 65b3a6bd-ec2e-448a-84a5-253ae723cf52 · outbound

This paper cites International Conference on Machine Learning , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Machine Learning , year=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.583629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f177a7b065a513bb115b3a87d4ca386d7fecef886e83ad20e6bff25542a02724

Observation 092dafaa-0a0c-4637-8314-1f7a742092b9 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.586511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:c3db0b1b72d248829f88c4e0a76a29f574d3d165c849f2adecb56bfb52104c77

Observation 727270f8-54fb-486b-99bd-d1865ebaef68 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:58:32.387959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:8d781c869b5c7c19f03afb20744b44acbd4c85dcd9b6b3b25c15669b3fcc940f

Observation af9b45bb-16e3-4637-92a1-c2a7ddcd15c2 · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.589099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:0a8085e1dcf2b569753f13b7d34e3ccec69afdaf375acd983aa76581d06e0387

Observation c9aefdaf-9374-44ea-ac70-d913f6aca89f · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.592158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:4aeba4fd32d732b39e85ff96242de3026dcac6b6001c9c88963351bdd01cc390

Observation a4415560-5b8b-4497-8b74-8ea0dab43bd4 · outbound

This paper cites Instruct.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Instruct

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.595080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:82b28a68d7da034911de28dfb4f7af780f566eda60f28a1a74dba72cf006656f

Observation 579f9052-de24-4511-85f3-737e6c5bd333 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.598133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:d3f1a323f17f7c71e7740a99ffbba32f9eb96e0cc369fcef9ccd3ebdda6ab9b8

Observation c9388f0e-fb0a-42db-a212-7db617d249d8 · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.601739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:a9f3d00b5f59ac931b35641e970ac1230e924a02d28355bbeb1c4da5d20d7449

Observation f6c79a3c-e83d-433f-9635-e69183b1a5a9 · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.604473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:a53bc88088f49efddd5369949267f1293433321c387fcffe6723b3b0f7f46ee9

Observation 67fd3450-cbfa-43b7-af75-589f7ceba8e8 · outbound

This paper cites Proceedings of the European Conference on Computer Vision , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the European Conference on Computer Vision , year=

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.608289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:01b037c1dbf31354e32d6382706a741fefd1bc2d917a98eb924b359126912d28

Observation 2f5b6e13-0fb0-4797-a9b3-5570608c0b17 · outbound

This paper cites International Conference on Learning Representations , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Learning Representations , year=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.611928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:4254828f4eb5db27504378da6830c4a7bdebc9076321a554a6aff4d2ebb3c51e

Observation 5bb5baf0-fa3a-4890-9701-37e117997244 · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.614793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:8f4f437ab38b6b4ede8efa048e0b2262473595a03e13db5b2e3a3beef61ab413

Observation 895ebd9d-7261-4327-8059-dbe2325b8ad9 · outbound

This paper cites Proceedings of the European Conference on Computer Vision , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the European Conference on Computer Vision , year=

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.618542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:82a72e80bbafda960c19e5004fb10961b1db775aadd97064b60c956ccf727a24

Observation 971c3880-3323-4c27-989e-90a383869f70 · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.622366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:0a191a7c53cd1f5dfca92bef7dccdbe6f9a932b5cbd2f9ee7820738df5df2a20

Observation 443542ba-84bf-4bca-b0ba-2782d17ca12b · outbound

This paper cites Proceedings of the Conference on Empirical Methods in Natural Language Processing , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the Conference on Empirical Methods in Natural Language Processing , year=

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.626015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:774c0359e2911e9591fd88d5afa063896ef5af2b472cbb8ca83361eb871cf0e9

Observation b92328e1-e9ed-4a72-9a21-0705373d1543 · outbound

This paper cites an unresolved cited work.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:58:32.629510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f2cf352d56c8c565a52fec17677e2fd2e3733d99f01487d1c332d763033b8b64

Observation 23da06c3-f225-4ac8-8525-454173a2408f · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.633110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:2d4151c96becb613631c93aa61d21166fe6efbe246da4fe12a41e2d35f31b058

Observation c4dd97ed-1855-49f1-99ba-c2b2ee2db692 · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.636166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:82b822f47949ff22f0ce627e342351913e6b75d0e627d1b67f5839cb1261b696

Observation ff772063-fda5-404e-9a10-ffd34664517d · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.639498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:a5cdefea4965e19b78677ff33ef03fca861ec560adbc94abd66db9698fa38a03

Observation 297c5267-576c-4530-b850-20752a3401d7 · outbound

This paper cites Journal of econometrics , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Journal of econometrics , year=

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.643002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:a8c4567d2c95ba878f2b3c7a801d7a37e0b8a8a45a75f5d21e2cc928181cb3d1

Observation 422cef44-6284-4bbe-93fa-ab0c24ddb7a6 · outbound

This paper cites Knowledge-Based Systems , volume=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Knowledge-Based Systems , volume=

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.646219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f84b7a603ef5cf66896eb3ad85b8bf4391c817f3358ba46263f5fa88b0f56ff7

Observation 27e989b9-fb16-4f46-937d-0c8a7a92436d · outbound

This paper cites Science , year=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Science , year=

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.649741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:f4d28c2f395c75263fdc46611de8071ec6aa2f42265946d7d1c5be50cc0f7501

Observation 580920e4-29cd-4e00-9ad2-2b577764e17e · outbound

This paper cites 2010 , publisher=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference 2010 , publisher=

Reference 103

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.653084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:d177c91949922a766e08975beb2557288658d9d7611a0c384c711fa72ad33105

Observation 35c8eff8-72f2-4ac6-bb04-8381b3b7f6f7 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , volume=

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:58:32.656479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:8a04871aabfe5fddd86377363cf3c9a7df0a8575403581b7bfeac4600ca20605

Observation cddc2a3c-ddc3-4838-9aa0-49b43fdb7468 · outbound

This paper cites Efficient Visual Transformer by Learnable Token Merging.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Efficient Visual Transformer by Learnable Token Merging

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.354566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:5cf609572f1a3f23373981b01e1519d3f24db988eea232fabd39dea32a03d162

Pith citing papers

Observation 713e1eb2-e512-46fc-a25b-15e954a44785 · inbound

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction cites this paper.

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:12:14.613620Z digest=sha256:a22fa10f16290f34eeb82b0eb9c488402291a57d409731467c82f5e491375e38

Observation 412b627f-29b0-4e73-a2d6-7fb859981462 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:02:17.816265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:1b31ae85f84d4a876a11d8b5e73ff480898d75af838c1b63d79ca25b1ab4c05b

Observation f5730463-af67-46f2-9b2d-3ce57c9be9e1 · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:48.044552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:48.044552Z digest=sha256:21387a55cb4a8401a6d1ff126982a10f6903ef4a7f3ae511d13c8745f9d5ee5c

Observation 73899475-a3f9-4248-aae6-fe606603b262 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:45.767775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:45.767775Z digest=sha256:1426d6123c1d6c480411e22af1d5549174e913ca726fc3fca9beb7a1155f8a97

Observation d3ffd459-c0a8-4ae1-9314-4b28c2d20b69 · inbound

Adaptive Token Merging for Efficient Transformer Semantic Communication at the Edge cites this paper.

Adaptive Token Merging for Efficient Transformer Semantic Communication at the Edge SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T18:26:55.083957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:26:55.083957Z digest=sha256:884e2fda241c1ef4152a40beebfbae84f0b004be3739a147d9776aaa1d033eda

Observation e97e06fd-23d2-496c-8b08-07dde644d45a · inbound

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models cites this paper.

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T21:31:36.995321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:31:36.995321Z digest=sha256:5a09916494c9e0fa966aacec91cf2b4d50510f30b83c5e39829ed240822301c1

Observation 6a3ef22d-aed3-4c3e-a224-a53f6ea7211e · inbound

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference cites this paper.

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T21:14:01.887983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:14:01.887983Z digest=sha256:602173a9542ca7f6b14e8fd2186224ecd2dc3ec495958ece418b7bcd5340d0a5

Observation 61e242bf-540c-4220-82f1-071e7282745e · inbound

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention cites this paper.

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:29:09.984286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:28:22.652509Z digest=sha256:944db7acc33d14e1aa0d1128b5d9604f184f213baea6b4ee7bfbfbbb5c3f4da4

Observation 65c3f1c3-6151-4691-ade4-43c804cb5f76 · inbound

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model cites this paper.

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:19:00.623483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:17:07.534291Z digest=sha256:a85118e9e6a8a9af46dfd399149583ad0d4aa41cf655bd919675b315ed02e470

Observation b2a4a894-25fc-4cd2-aa43-f872688b19bc · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:43.164201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:43.164201Z digest=sha256:9e262f21287cc6f0515736f7536442dd6b9f826f3784c040cbd0198103b0c216

Observation 04b546e9-49f7-49cd-81bc-2c5d19cb03eb · inbound

Selective LoRA for Visual Tokens and Attention Heads cites this paper.

Selective LoRA for Visual Tokens and Attention Heads SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:18:23.671109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T20:17:07.884892Z digest=sha256:68682a8d29c43104cd50b82db6187c6d8ccf75732fb28d1f244f4875a0a04d0d

Observation 5cf2c86d-e2b8-450e-8eed-a61145211edf · inbound

LinMU: Multimodal Understanding Made Linear cites this paper.

LinMU: Multimodal Understanding Made Linear SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:33:15.158927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T18:32:23.951566Z digest=sha256:bb51eedd554b28f31e4f361f766ba6c2eccbf2e5db55760be4201d94352b80b6

Observation 6d510b01-19d4-4d2f-8829-9cf4b29286e2 · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:26:27.029877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:920bb6aa6c0a977a4a4774a82788068b55816bb0c5e205b66a3e77a0e839b4a5

Observation 086f0fae-5e5d-4141-b654-1a80b552f07c · inbound

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling cites this paper.

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:56:47.841355Z digest=sha256:bf1e7a6ede55ffcd9a40749243c484401ea43cdcba65b8bbef0ba326f0cc57a1

Observation af3f1071-10f5-430e-8db7-95601dffa92a · inbound

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies cites this paper.

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:40:55.343169Z digest=sha256:0e27a1b8657f90c4a2e6d73f65cbae4873d3f642b44894289ff9c30d49c2334c

Observation 735cf995-5c67-4147-b990-2d8d30107ffb · inbound

DINO-VO: Learning Where to Focus for Enhanced State Estimation cites this paper.

DINO-VO: Learning Where to Focus for Enhanced State Estimation SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T17:05:52.493660Z digest=sha256:c7a1942398df60a547da46c53953e241a6cb1892d594aaa33f93cf046c907593

Observation dd444398-ac9c-420d-b0e3-919fc72a05d0 · inbound

Do Vision Language Models Need to Process Image Tokens? cites this paper.

Do Vision Language Models Need to Process Image Tokens? SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:26:37.371488Z digest=sha256:009e8b518f09e4c2d92d155df1e36accda41d915d9ac3edaa57f84a5a4a74f84

Observation a1fb4d6a-f906-41c9-87a0-b760b9169ac0 · inbound

Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding cites this paper.

Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:49.681399Z digest=sha256:968beedc509517a5c63d0669f022b507dcb179d9afcd4dbc26aae03618ed7efd

Observation 320c0715-28f5-4777-9509-f4ee8fa70cc1 · inbound

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models cites this paper.

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:15:04.261855Z digest=sha256:2f9d39a0cc0a86c683634b518c3918bd08c63a0ae19b85d5c0fa8a0209f913d6

Observation bc3789bb-c4a6-4641-b63b-6f2628f92e74 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:f673b73e5d6c5bc786db8692cc4fb96f8be4552e3d775743dfcb1b9140419dc8

Observation 134db799-d7df-4f55-8a5f-579f23486f11 · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:a59b6d600c7529e1da89a15c60061b786240f2a07825315a0ee48032ada69921

Observation f837c46d-02e0-434c-9567-753c4fcfbe8d · inbound

VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models cites this paper.

VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T11:28:44.565497Z digest=sha256:ba7c82335eb3fe3665c37a12b02c4d1d322b7011b0039835efa7aca0ec94ad4b

Observation 754973df-c953-4670-8a06-12fe6d8d8439 · inbound

Geometry-Guided 3D Visual Token Pruning for Video-Language Models cites this paper.

Geometry-Guided 3D Visual Token Pruning for Video-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:49:38.346274Z digest=sha256:38d67d3cf4d307406e405df3756d875ae658f4d905ba0f9bb67dd01cd2e91d23

Observation e7ef7a39-2192-43b5-9fb8-8c64b9d114ec · inbound

Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment cites this paper.

Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:04:46.460069Z digest=sha256:d00dc7614cf73666ddd36881e19d6650c78e8ad69c0f034ab1f4847a54b5580f

Observation d4e97c48-e2d8-48f7-8b7b-5bf39c6e6403 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T19:49:28.591419Z digest=sha256:ddb8da69d9a2ee5257af78b6bb67a20e80e662ede864dba9a5a26f782fa0f7e4

Observation 9697ceca-8532-4cdb-8746-e0026decc93b · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:e7c810c67ba50972b9ed025f86f8499b59fbece3cbd61b76d416f64ba8e899de

Observation 3ac9958a-fddc-4ea2-9fe8-a387932cee44 · inbound

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks cites this paper.

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:16:31.820718Z digest=sha256:1d1822349b2c50ec954e7a4dd7e1d5b44ae8b58fd0aadf80f35315af72eba2ef

Observation 4ad340de-855a-4f43-aa49-0b316c4f93ba · inbound

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks cites this paper.

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T05:19:56.337244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:19:56.337244Z digest=sha256:516d9e0b2ffa2a035ee5664f11734cbfeef59af92f00ec1c20d8f9176d204c9c

Observation 1084ac3a-e8bf-44b0-a80e-b48a1701b455 · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:48:39.444933Z digest=sha256:ba0a8c3c3d2ddb2b985b9d6595be84f2ea853a9f887d8eccbd606f8fb87cb8f7

Observation 57c41584-e60e-4225-aeee-465649f1db51 · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:57:42.822121Z digest=sha256:df83fc3ef304d30d4316a18b09f0661cdd473bf480412b7f1042c74d9d0ee347

Observation 334343cb-bb53-4a99-99d3-68afa25870c6 · inbound

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? cites this paper.

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:06:45.858231Z digest=sha256:3d1b1270507d9bcc41604621dfa2d03213b2d70b5b67b24bf6bc52a55792055a

Observation 14a8b709-b2b2-40bd-a491-717c9768bcae · inbound

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction cites this paper.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:945bd91723a2dc45bebb6c52b9326f7da633ec4b0ef9d3fae2e7b1d522f7bb1d

Observation d24a6c1c-2161-45ac-a059-0ae1af1dccff · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:43:08.029165Z digest=sha256:12c1ad7386034952c488408b6d50fbb0bf8092e2b327c1f77f788560ad26d2fd

Observation 3893fcc5-32df-4f4f-8ed2-997afa4ca99f · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:45:05.706588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:42:05.592911Z digest=sha256:564a7632bd92af26609a51bc639c6605da324b18bc61c43c97b15f32b5c54182

Observation c78ba4e1-e20f-44b1-8ce3-30f301c0c87d · inbound

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs cites this paper.

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:08:54.511241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:04:15.273932Z digest=sha256:dfb0512548ec6f7cf7fae198e61b492ca40868656a3e61843decc89e58799bb8

Observation ac6f9ec6-0c85-4544-bbca-76559b1fc031 · inbound

Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference cites this paper.

Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:43:23.787783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T07:43:18.828740Z digest=sha256:58afe08cefd80ad58d1b6fff9cf6776a89077bb678e75e3ee5a1be6c716d62b7

Observation 266c9dd0-45cd-4d21-958f-87e6bdc2d019 · inbound

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs cites this paper.

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:58:05.980263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T06:55:01.619441Z digest=sha256:a4f6b4884f7c61147cc1513389b69d9467d907ca0d9a3d94dfc4d0e611a72443

Observation 6b45e725-3b0f-45c1-9897-3a46cd7eaad8 · inbound

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models cites this paper.

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:23:58.599149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T05:20:55.448430Z digest=sha256:be1f692695be1e446ef4cd91f896f29cadb62b4ce7038d62955252ac3efacdeb

Observation bfd0ed5f-dd96-457d-bf19-e09a8301c35f · inbound

ASAP: Attention Sink Anchored Pruning cites this paper.

ASAP: Attention Sink Anchored Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:14:45.587099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:12:13.451406Z digest=sha256:38b93d6afd0b4b612704a9dc483cb089a6e67b637fea64bdce194efc0e4c337b

Observation ba0f3934-791d-4b2f-8526-3ec3a0a4f16e · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:04:52.633378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:20aec2c329be87a9e4ef1c54e70735bbab36ad167d11e70ceb870898557576c9

Observation 776a3956-fbb5-403c-b2de-79f7e62b6484 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 203

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.697865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:3a3a8944d6842e916202e779fc652ed2ca2aad97f15911936bffe8543ba95d6e

Observation 7826f094-5a2a-4f23-a81e-3986287b16e0 · inbound

AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference cites this paper.

AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:53:15.983445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T08:48:58.583657Z digest=sha256:8a5d2cf140c6a55b90b5cb33302f514c50f95eea93f5581aaa123fa8de62aab9

Observation b63ea36c-b162-49fd-890f-cdc457fa99c5 · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:14.963744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T08:40:10.152344Z digest=sha256:97bc09635ecba6a78ecfadbdfb0568b9078aed4a546382c7e4f12712b91f447f

Observation 10bbd4fd-3501-4b67-ab9c-d96639cccf2d · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T12:54:43.337210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:54:43.337210Z digest=sha256:8b64fbf2c9e4235635896c0262cb9562bd694ba2dff866e8d208ae45e847b423

Observation f2657672-9bc6-45fe-b72b-356543277b0f · inbound

Differentiable Efficient Operator Search cites this paper.

Differentiable Efficient Operator Search SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 32

Resolution
malformed identifier
local_arxiv, observed 2026-07-02T06:06:41.789534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:30:30.355657Z digest=sha256:37b03b73bab0ddc6c024f76e858748178ca3f6fab8ca6039dd11fc2aab6bf694

Observation 328daec8-ed63-4a77-b572-2251d15884ec · inbound

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models cites this paper.

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:28:04.130474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:35:24.118536Z digest=sha256:65773c6542bde53829d9833d3a59acca5c002c6c6d95a3d6902b2a04e131767c

Observation 6774e7e9-bbe3-41a6-ae2a-da04a729711b · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:19:29.929448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:66e3951c770a42d4a4600b60ab3df5bc71a3e07cc77b9d374b579ed043ec8333

Observation af33357b-3b81-403b-b59b-c9fa619735b7 · inbound

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference cites this paper.

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:09:53.801383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T04:24:38.917137Z digest=sha256:aee83bfc1ce015dafa3497d507f7e9c3b1f98cf558ecb56ee0c41d240f9c6cae

Observation 9b66dbaa-2730-4f61-bf2a-4af54dbdbe83 · inbound

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving cites this paper.

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:23:51.150255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T05:08:46.145616Z digest=sha256:5047920a66d149def1714245f4e293c13ba827de958d45566ed42cfdf715f67e

Observation ef58a4c4-f0bc-40ca-8c83-d40a021c1849 · inbound

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving cites this paper.

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:37:24.219111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T21:35:33.280591Z digest=sha256:ae30a1e09bfda23382a5f0e896fcb8256fa7b7d0d1756f5bd09f2942c909edd2

Observation 728ac82e-1ddb-43be-a05f-01f81d20643b · inbound

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs cites this paper.

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:15:44.599724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T05:41:04.184461Z digest=sha256:7cb8ae9922959eebfbcdfb5ecb86c9145c0a1ab17a5e3eb53073ea8d13c08002

Observation e141fd5b-5c54-4ece-bf48-58ff3d61c78c · inbound

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning cites this paper.

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:15:45.204019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T05:36:43.609602Z digest=sha256:781f63b075b3b24b7371047a5ae2e5d57bfc081af3dd586e2c3d9d092b6b1555

Observation 393c4af5-5d93-4b98-8030-a697617cb543 · inbound

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning cites this paper.

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:48:32.367832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T14:47:35.377391Z digest=sha256:632ad27a76c82087f449a5cf3097352109b0861b9ca8584794fca653e9c253a6

Observation 18157f53-17f8-435c-8620-248a523d780a · inbound

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering cites this paper.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:793c16f871cd0082943a3233acfab322e6ee06dc4ca08477e6e325434dcb7860

Observation f908bd59-8cb1-41b8-b779-8b2e0a937fef · inbound

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring cites this paper.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.598163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:0a2dc5e14a57bf4d8e01b033828b249fb3e8e183be2f3bc7a3d29c998c617c01

Observation f7f7c72c-5c6a-4234-a012-70093eb9f5d6 · inbound

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning cites this paper.

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T21:36:34.339662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T21:28:27.693908Z digest=sha256:c8e8f46566c731286a9c5add64f6e4d958f2f52b70bfb46f4adad07cc1b827e4

Observation b1469d2f-cd50-4194-8faa-eb3157bf131a · inbound

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning cites this paper.

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T15:54:18.404613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:54:18.404613Z digest=sha256:5e0c7990295cf5248a87e8c8a7343cd1d25930aab49c1efe6401d0313e69d4db

Observation 483171d1-7c26-4052-b98e-86a3922af1af · inbound

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning cites this paper.

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T08:12:31.386718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:12:31.386718Z digest=sha256:a4afbda439f14d2fb943493ca050d5654a4d8482e586a11398a4f93497924034

Observation 96b53720-c364-4532-aaf9-35cc61dc0c44 · inbound

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models cites this paper.

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T10:17:43.230980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:17:43.230980Z digest=sha256:fdb28fce8a95248e546d4dfe9d001118694dfaa3239817c24a5c90d71ca4ea80

Observation 3f50efab-5fd9-46a0-8cee-eeeba7b76617 · inbound

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models cites this paper.

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:15:25.071270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:15:25.071270Z digest=sha256:a098e6b6083d968b74d774298f12532bc402eda8498d03b913058f7756383b58

Observation c428b69e-f27f-497b-8076-78de25232b7b · inbound

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation cites this paper.

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T01:48:57.761577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:48:57.761577Z digest=sha256:e5eed05e344f924996045915886930c063edf1fce039a65de9984d36adc55a4f

Observation f31faea6-6422-447e-ae2f-ab364101c218 · inbound

Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models cites this paper.

Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T19:13:28.370100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:13:28.370100Z digest=sha256:4299ac36623d216407743b9f9816fba159bcebb441912a1bb2423f7326e39b0e

Observation bdc79aab-85ee-45c5-9f0b-ac6b65b5dca2 · inbound

Visual Token Compression Enhances Robustness of MLLMs cites this paper.

Visual Token Compression Enhances Robustness of MLLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T13:10:37.711024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:10:37.711024Z digest=sha256:baee0676e75f8d64dfdf4af1c641613c3a0b29bf0e3915000de41f1420f224bc

Observation 647515c7-7869-4297-9099-aa17db97fb8b · inbound

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models cites this paper.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T03:25:31.650756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:25:31.650756Z digest=sha256:859e0398d293065f613981e9bbc400a7f89ba937dff86b08e36740dac6b4b902

Observation dd4ec0f3-2ddf-4830-b903-7a0a7a708e1a · inbound

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models cites this paper.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.522646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.522646Z digest=sha256:14d01774439320b223bc3b9c04aabb1d6424cc9c954fdb88ac5c35bed1040dc8

Observation 3814d9e8-3e61-434a-a671-c35ad689f654 · inbound

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models cites this paper.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.654463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.654463Z digest=sha256:3b11d5d044d671588e4f360ffe7bcb81a7804caa2ce89994f1f0fac383632296

Observation e305149a-9e40-4727-a2fa-94e72f60e8c1 · inbound

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs cites this paper.

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 1057

Resolution
unresolved
no resolver link, observed 2026-08-01T01:54:22.000078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:54:22.000078Z digest=sha256:4c89240b00ccf45680d1f005164cc1dfedf8871bb8bdd5251598e3d1f3dd7a57

Observation f836949a-bf70-4340-a192-55de70e61a77 · inbound

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models cites this paper.

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T01:26:50.499048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:26:50.499048Z digest=sha256:fefb1ee3f021bc9d0e76ebf614b31819e96fa29fd2e60489b7d71947877aea51

Observation b76c417c-7df6-4cef-976a-4adddbfb8e78 · inbound

MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models cites this paper.

MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T13:31:50.031383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:31:50.031383Z digest=sha256:4e818ea293e4c65242bc2aea3bfba8895c36fd6c4cc85500acc7268ef9811a1d

Observation a85a632c-e75e-472f-bcc1-42eb7aecc286 · inbound

LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference cites this paper.

LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T22:24:19.216556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:24:19.216556Z digest=sha256:4a3d0aa28604f8ab518d8bb14db5f7d1d887dfc54eda31831358661140c30561

Observation a7fe473c-2c51-41f6-9fe4-6c91d74d1e75 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.635563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.635563Z digest=sha256:a9b30ee763f75375453bb05fac83779b553f2036f1cf31277592db16c41b7303

Observation 9fce38af-f563-4acc-9a48-9305c6f2779f · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:00.303154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:00.303154Z digest=sha256:6e61aae36b7b85384c1bf64ff040876fcc90a0d3e0c55cc548aa1772b0b903a3

Observation 4a8e9594-1b6f-4d93-94d8-584de078c6d5 · inbound

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware cites this paper.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.848690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.848690Z digest=sha256:0796a7e7c134d2dea1f7b16ad49d812bf9c4222ff802442d473582b0204c0711