Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T14:58:32.303101Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 73 inbound Pith citation observations for arXiv:2410.04417.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T14:58:32.303101Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:38:45.767775Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T21:36:34.338168Z
100 of 113 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bbe4c423-ab03-42f4-b5e2-fe252c8a6b48 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd787a66-e539-4c2f-8b84-95613a1b46eb · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a401b2f-33b0-4831-b485-4e2b6f5922a8 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Token merging: Your vit but faster
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1bf9eeb3-61b2-475b-9457-47bfafa855c9 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cafe519d-cde9-4316-b842-b24e6df4374e · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7f725070-82e0-43cf-8883-b6dc3f650174 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Honeybee: Locality-enhanced projector for multimodal llm
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4b3d7932-8a27-4b6e-a91b-f19c9a050d7e · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76c7ad83-e08f-4570-a827-5f62e76e5301 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a57b32f8-ea2a-4b29-b11b-3d09a8ca36f4 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Instruct BLIP : Towards general-purpose vision-language models with instruction tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 254cd483-b4ca-46a6-847b-d0ff915f4d1b · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Flash A ttention: Fast and memory-efficient exact attention with io-awareness
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff3b4942-7816-4ee2-a132-6ac3117703a5 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Glm: General language model pretraining with autoregressive blank infilling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92526777-bc9e-47d2-a87d-a1b12af27062 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f9c66094-2fab-4ff3-a26d-645bd5c45850 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8aa3d3cc-adde-4128-b499-10de1ef2cafa · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4a288f2-9cfa-4c01-b2c2-7f927990e616 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 641ea025-3580-4247-ad2e-b03e7e3498d5 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Videopoet: A large language model for zero-shot video generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5880045-832e-43ec-8c78-0c13e6d2cb5b · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Seed-bench: Benchmarking multimodal large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 52163e20-d853-4853-b04c-3f3b9d4b32b8 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04fb6005-a86e-4d52-8934-9e354e6c6090 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference X., and Wen, J.-R
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 546aa8b8-bc21-4b7b-ba8b-9a0e60644508 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference LLaMA-VID : An image is worth 2 tokens in large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc677043-36e8-499e-b951-f40ebe3202d3 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Video-llava: Learning united visual representation by alignment before projection
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e1169e2-7dc8-4c42-977e-28830054fe61 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1670af51-a156-45e5-a905-798c7dd2a347 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e716eee0-0727-4e68-8209-d2badfe2593d · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Mmbench: Is your multi-modal model an all-around player? In Proceedings of the European Conference on Computer Vision, 2024 c
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b880c23-be0f-497e-8c54-10518daf84cd · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference A convnet for the 2020s
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1384f337-b89e-4f90-b802-c3f1b397ba53 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96097744-1bae-417d-bb5c-5b618c1f49ba · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40912adb-d435-4732-b824-833bcb3049bf · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Vision: A computational investigation into the human representation and processing of visual information
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1324da50-7d94-4321-afd7-6cebe2d1146e · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Language models are unsupervised multitask learners
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 39cbe1d0-fb5e-4aa6-a2a0-fe81294ae6bc · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Clustering by fast search and find of density peaks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4cf49f61-7616-452a-ac68-1acacc84ba21 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Towards VQA models that can read
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ab59e97b-3e02-4df7-bc2a-78daf15a3f6c · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f25f953-2a84-4b5d-8723-49d9b3ec9fe5 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference N., Kaiser, ., and Polosukhin, I
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a0f098d-100b-4a31-bd7a-cf6bd86f49b7 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Q., Wang, Q., Gao, Y., Xu, Q., Xu, T., Hu, Y., Chen, E., and Shou, M
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation edc6858b-c696-4eaa-b4eb-28514079b42b · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c61b5a2-368f-4dc9-8d3a-ac6c510359b5 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Video question answering via gradually refined attention over appearance and motion
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 139efa7c-ec3a-43f4-b657-741794c2b013 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 20e3e900-2492-4fc6-94a2-4a48e6faa88a · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference VoCo-LLaMA : Towards vision compression with large language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf1e51f3-4088-4d9c-a1ea-a29d20ea99b3 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Mm-vet: Evaluating large multimodal models for integrated capabilities
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d820b12-87ab-4189-8888-df4f77de4dea · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation afc66c46-bd6b-40b2-9fda-f8f806fd62b7 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unveiling the tapestry of consistency in large vision-language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7e7e5a1-7f14-4504-905d-66405628a671 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Freekd: Knowledge distillation via semantic frequency prompt
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6b3b280-f33c-47d6-ba60-6b2a9b2c8f5c · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9dbea14a-1625-433b-94d0-474443087fb7 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Minigpt-4: Enhancing vision-language understanding with advanced large language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 62c34109-145b-421b-aafc-86aa2e22a468 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Scaling Learning Algorithms Towards
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9ecf8bcc-c241-4b63-814a-1dbe9498f065 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference and Osindero, Simon and Teh, Yee Whye , journal =
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 185d2150-c75a-4db9-9081-0ac4aa0fffc2 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference 2016 , publisher=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d4ce623-25d2-4c26-b28d-886bfc5eaeab · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part IV 14 , pages=
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d402cc9e-6924-48d8-9f7c-268614322812 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Machine Learning , year=
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 27a38d1e-fb27-4d06-aa45-86c2f34827d4 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f56280d5-867e-4094-8128-348abe91a2a6 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5973323-64d2-4a50-9985-f79741b4a740 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference OpenAI blog , year=
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c1983c8-290a-4a6f-80c5-10ed6eea9450 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference GPT-4 Technical Report
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 82166124-1841-4514-8bd9-a1f346978f1a · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Deepseek
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb085220-253d-4d0f-9c7f-a94f8ee1f1a9 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference LLaMA: Open and Efficient Foundation Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73df4974-5e08-46fb-bc36-d5f822fb6dbe · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Instruction Tuning with GPT-4
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 38f1d31b-8b4f-45ca-a2cf-465ef4bdd60d · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9bc5d36-7143-40cb-95ea-5256385aa9a7 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bd1f434b-ac0a-4523-83c1-3a224f41f5ed · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f358fb2f-dda9-4cc7-86cc-12fe91e0de39 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a415b51-90f2-4a9c-a614-397ca336a455 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the Annual Meeting of the Association for Computational Linguistics , year=
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 392c2795-8ad4-402a-9ede-dffd78bbd27b · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Learning Representations , year=
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3a6be316-e1a7-4ae3-b40e-c1f41febb8f5 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the Conference on Empirical Methods in Natural Language Processing , year=
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e06f6570-92e6-4b4b-88b7-f28a537fbc80 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Qwen Technical Report
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a160a8a3-5b59-41a7-ad44-a867a787e175 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Gemini: A Family of Highly Capable Multimodal Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 360ab59c-956e-4d05-8d77-280c5bdab154 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8dea3b02-ea03-421a-8155-ebd80b743fa2 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation da95dcf1-1602-4433-8d9e-d8128bcd57fa · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 724ddaa9-7e39-4b2a-adf6-984c04aa1b4f · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference CogVLM: Visual Expert for Pretrained Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cbee58ce-9561-45fb-8ed3-bd9709de813d · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 824b2051-9475-457d-bb27-6fbce43df43d · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dca7e261-89ab-422a-802c-6b3872e4f7e1 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc6fb78a-8f51-4d87-a3fd-bbabd8d802cd · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25168446-00a5-41c8-a819-b503e95ef52e · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Machine Learning , year=
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed1294c9-16a7-4cdb-b083-74ffbc8090bc · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International conference on machine learning , year=
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65b3a6bd-ec2e-448a-84a5-253ae723cf52 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Machine Learning , year=
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 092dafaa-0a0c-4637-8314-1f7a742092b9 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 727270f8-54fb-486b-99bd-d1865ebaef68 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af9b45bb-16e3-4637-92a1-c2a7ddcd15c2 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9aefdaf-9374-44ea-ac70-d913f6aca89f · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4415560-5b8b-4497-8b74-8ea0dab43bd4 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Instruct
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 579f9052-de24-4511-85f3-737e6c5bd333 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9388f0e-fb0a-42db-a212-7db617d249d8 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6c79a3c-e83d-433f-9635-e69183b1a5a9 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67fd3450-cbfa-43b7-af75-589f7ceba8e8 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the European Conference on Computer Vision , year=
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f5b6e13-0fb0-4797-a9b3-5570608c0b17 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference International Conference on Learning Representations , year=
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5bb5baf0-fa3a-4890-9701-37e117997244 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 895ebd9d-7261-4327-8059-dbe2325b8ad9 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the European Conference on Computer Vision , year=
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 971c3880-3323-4c27-989e-90a383869f70 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 443542ba-84bf-4bca-b0ba-2782d17ca12b · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the Conference on Empirical Methods in Natural Language Processing , year=
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b92328e1-e9ed-4a72-9a21-0705373d1543 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Unresolved cited work
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23da06c3-f225-4ac8-8525-454173a2408f · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE conference on computer vision and pattern recognition , pages=
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c4dd97ed-1855-49f1-99ba-c2b2ee2db692 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , year=
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff772063-fda5-404e-9a10-ffd34664517d · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , year=
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 297c5267-576c-4530-b850-20752a3401d7 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Journal of econometrics , year=
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 422cef44-6284-4bbe-93fa-ab0c24ddb7a6 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Knowledge-Based Systems , volume=
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 27e989b9-fb16-4f46-937d-0c8a7a92436d · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Science , year=
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 580920e4-29cd-4e00-9ad2-2b577764e17e · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference 2010 , publisher=
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 35c8eff8-72f2-4ac6-bb04-8381b3b7f6f7 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Advances in Neural Information Processing Systems , volume=
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cddc2a3c-ddc3-4838-9aa0-49b43fdb7468 · outbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Efficient Visual Transformer by Learnable Token Merging
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 713e1eb2-e512-46fc-a25b-15e954a44785 · inbound
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 412b627f-29b0-4e73-a2d6-7fb859981462 · inbound
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f5730463-af67-46f2-9b2d-3ce57c9be9e1 · inbound
Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73899475-a3f9-4248-aae6-fe606603b262 · inbound
Leveraging OS-Level Primitives for Robotic Action Management SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3ffd459-c0a8-4ae1-9314-4b28c2d20b69 · inbound
Adaptive Token Merging for Efficient Transformer Semantic Communication at the Edge SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e97e06fd-23d2-496c-8b08-07dde644d45a · inbound
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a3ef22d-aed3-4c3e-a224-a53f6ea7211e · inbound
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61e242bf-540c-4220-82f1-071e7282745e · inbound
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65c3f1c3-6151-4691-ade4-43c804cb5f76 · inbound
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2a4a894-25fc-4cd2-aa43-f872688b19bc · inbound
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b546e9-49f7-49cd-81bc-2c5d19cb03eb · inbound
Selective LoRA for Visual Tokens and Attention Heads SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cf2c86d-e2b8-450e-8eed-a61145211edf · inbound
LinMU: Multimodal Understanding Made Linear SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6d510b01-19d4-4d2f-8829-9cf4b29286e2 · inbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 086f0fae-5e5d-4141-b654-1a80b552f07c · inbound
ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af3f1071-10f5-430e-8db7-95601dffa92a · inbound
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 735cf995-5c67-4147-b990-2d8d30107ffb · inbound
DINO-VO: Learning Where to Focus for Enhanced State Estimation SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dd444398-ac9c-420d-b0e3-919fc72a05d0 · inbound
Do Vision Language Models Need to Process Image Tokens? SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1fb4d6a-f906-41c9-87a0-b760b9169ac0 · inbound
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 320c0715-28f5-4777-9509-f4ee8fa70cc1 · inbound
Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bc3789bb-c4a6-4641-b63b-6f2628f92e74 · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 134db799-d7df-4f55-8a5f-579f23486f11 · inbound
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f837c46d-02e0-434c-9567-753c4fcfbe8d · inbound
VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 754973df-c953-4670-8a06-12fe6d8d8439 · inbound
Geometry-Guided 3D Visual Token Pruning for Video-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7ef7a39-2192-43b5-9fb8-8c64b9d114ec · inbound
Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4e97c48-e2d8-48f7-8b7b-5bf39c6e6403 · inbound
RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9697ceca-8532-4cdb-8746-e0026decc93b · inbound
VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ac9958a-fddc-4ea2-9fe8-a387932cee44 · inbound
Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ad340de-855a-4f43-aa49-0b316c4f93ba · inbound
Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1084ac3a-e8bf-44b0-a80e-b48a1701b455 · inbound
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57c41584-e60e-4225-aeee-465649f1db51 · inbound
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 334343cb-bb53-4a99-99d3-68afa25870c6 · inbound
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14a8b709-b2b2-40bd-a491-717c9768bcae · inbound
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d24a6c1c-2161-45ac-a059-0ae1af1dccff · inbound
AttenA+: Rectifying Action Inequality in Robotic Foundation Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3893fcc5-32df-4f4f-8ed2-997afa4ca99f · inbound
AttenA+: Rectifying Action Inequality in Robotic Foundation Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c78ba4e1-e20f-44b1-8ce3-30f301c0c87d · inbound
LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ac6f9ec6-0c85-4544-bbca-76559b1fc031 · inbound
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 266c9dd0-45cd-4d21-958f-87e6bdc2d019 · inbound
DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6b45e725-3b0f-45c1-9897-3a46cd7eaad8 · inbound
Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bfd0ed5f-dd96-457d-bf19-e09a8301c35f · inbound
ASAP: Attention Sink Anchored Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ba0f3934-791d-4b2f-8526-3ec3a0a4f16e · inbound
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 776a3956-fbb5-403c-b2de-79f7e62b6484 · inbound
Toward Native Multimodal Modeling: A Roadmap SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7826f094-5a2a-4f23-a81e-3986287b16e0 · inbound
AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b63ea36c-b162-49fd-890f-cdc457fa99c5 · inbound
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 10bbd4fd-3501-4b67-ab9c-d96639cccf2d · inbound
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2657672-9bc6-45fe-b72b-356543277b0f · inbound
Differentiable Efficient Operator Search SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 328daec8-ed63-4a77-b572-2251d15884ec · inbound
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6774e7e9-bbe3-41a6-ae2a-da04a729711b · inbound
ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af33357b-3b81-403b-b59b-c9fa619735b7 · inbound
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b66dbaa-2730-4f61-bf2a-4af54dbdbe83 · inbound
MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ef58a4c4-f0bc-40ca-8c83-d40a021c1849 · inbound
MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 728ac82e-1ddb-43be-a05f-01f81d20643b · inbound
MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e141fd5b-5c54-4ece-bf48-58ff3d61c78c · inbound
Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 393c4af5-5d93-4b98-8030-a697617cb543 · inbound
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 18157f53-17f8-435c-8620-248a523d780a · inbound
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f908bd59-8cb1-41b8-b779-8b2e0a937fef · inbound
AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7f7c72c-5c6a-4234-a012-70093eb9f5d6 · inbound
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b1469d2f-cd50-4194-8faa-eb3157bf131a · inbound
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 483171d1-7c26-4052-b98e-86a3922af1af · inbound
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b53720-c364-4532-aaf9-35cc61dc0c44 · inbound
Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f50efab-5fd9-46a0-8cee-eeeba7b76617 · inbound
Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c428b69e-f27f-497b-8076-78de25232b7b · inbound
Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31faea6-6422-447e-ae2f-ab364101c218 · inbound
Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdc79aab-85ee-45c5-9f0b-ac6b65b5dca2 · inbound
Visual Token Compression Enhances Robustness of MLLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647515c7-7869-4297-9099-aa17db97fb8b · inbound
OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd4ec0f3-2ddf-4830-b903-7a0a7a708e1a · inbound
OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3814d9e8-3e61-434a-a671-c35ad689f654 · inbound
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e305149a-9e40-4727-a2fa-94e72f60e8c1 · inbound
OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 1057
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f836949a-bf70-4340-a192-55de70e61a77 · inbound
SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b76c417c-7df6-4cef-976a-4adddbfb8e78 · inbound
MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a85a632c-e75e-472f-bcc1-42eb7aecc286 · inbound
LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7fe473c-2c51-41f6-9fe4-6c91d74d1e75 · inbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fce38af-f563-4acc-9a48-9305c6f2779f · inbound
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 126
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a8e9594-1b6f-4d93-94d8-584de078c6d5 · inbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.