Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T05:03:30.238634Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2607.13500.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T05:03:30.238634Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f4ab3642-ce73-454d-b2c0-611e1b14a6c4 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9847a69-a0ba-41e3-afb5-90c2ec4a0642 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226cf9b0-c893-49a1-bc4f-2e0510e1024a · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea241e75-2bb5-4ef9-b833-a0bc3a062b41 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models BLIP- 2:bootstrapping language-image pre-training with frozen 11 image encoders and large language models,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b2ade47-0381-4662-9bb0-0a2d01c718ef · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Cost-efficient and secure federated learning for edge computing,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dbdf4ac-b3dd-40f0-9467-1989ec35030a · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models To- wards online privacy-preserving computation offloading in mobile edge computing,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c110d243-9905-40b3-a029-e4f09786feb6 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models AFLoRA: Adaptive Federated Fine-Tuning of Large Language Models with Resource-Aware Low-Rank Adaption
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54509875-4356-453a-92fd-965ad18dd9bd · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Towards efficient edge learning for large models in heterogeneous resource-limited environments,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62585956-b6f2-4224-b033-c57ae6dbf983 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e4c373d-8840-48a8-b27f-9bdcb9d49091 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Learning transferable visual models from natural language supervision,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e52ff1-b723-42a7-a6ba-b6613a029b58 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Sigmoid loss for language image pre-training,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c0b6c8-aa10-4928-92c3-7aefb0c0232c · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Tap-vits: Task-adaptive pruning for on- device deployment of vision transformers,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d088e23-6195-4a16-bcc8-5d894330ded2 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision- language models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b2543d-5b97-41a7-82d5-84b2e311bc4e · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models VScan: Rethinking visual token reduction for efficient large vision-language models,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 064b0bc9-7083-443c-8844-053fe9866142 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Adaptinfer: Adaptive token pruning for vision-language model inference with dynamical text guidance,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fc6fe26-451d-44fc-84b4-94e73926579c · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Variation-aware vision token dropping for faster large vision-language models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a58740-0054-4df7-bbdd-43b100638dd8 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Sparsevlm: Visual token sparsification for efficient vision-language model inference,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3d6150-0858-4287-b4a9-ec64c80658be · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Pyramiddrop: Accelerating your large vision-language models via pyra- mid visual redundancy reduction,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13493d25-1882-4ac2-9c22-5251d778b29b · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d13bbfb4-1f49-43db-b16c-80d33edd3730 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ff3eba6-3f51-43f7-a7fb-97b6a2b91979 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Visionzip: Longer is better but not necessary in vision language models,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b084ba45-397d-4a07-89af-fc30c640744a · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 546ee96d-be61-4ad7-84d2-45d3f3cbd165 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Flashat- tention: Fast and memory-efficient exact attention with io-awareness,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2deb5366-dbe8-4b57-bbb0-61b467999fb5 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Flashattention-2: Faster attention with better parallelism and work partitioning,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc87243f-39eb-46ac-bcda-da8422f6b7ce · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Flashattention-3: Fast and accurate attention with asynchrony and low-precision,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59915c0f-7492-4139-8a10-0ad3127b019b · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Efficient memory management for large language model serving with pagedattention,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9774a8cb-bf6f-4697-814e-07c4097cb51e · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Similarity-Aware Token Pruning: Your VLM but Faster
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22517f44-5ec4-418c-aa52-cf96dcdb49a2 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d64b4bf-0c8f-4592-af83-ac3c21952221 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Holitom: Holistic token merging for fast video large language models,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5205495b-aa63-4cf8-b571-2f04b3603e2c · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Dycoke: Dynamic compression of tokens for fast video large language models,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87784a5b-205f-46f8-b004-c1470cc9e687 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and- language tasks,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be0da18c-1d61-449f-9964-be0b68b85197 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Uniter: Universal image-text representation learning,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8089c62-c843-492d-ad37-af8967cc2be0 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Vflowopt: A token pruning framework for lmms with visual information flow-guided optimization,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dd0f4b1-427a-4618-b40c-39e40a799b4a · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Hidrop: Hierarchical vision token reduction in mllms via late injection, concave pyramid pruning, and early exit,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8aa322-7b15-44e8-bcba-029479783c31 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Todre: Visual token pruning via diversity and task awareness for efficient large vision- language models,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c635a1e2-7d3a-4d3b-8cb3-6c77aa6b9dab · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Tamp: Token-adaptive layerwise pruning in multimodal large language models,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19994dd-9808-4d3f-8bd0-6836fccd94db · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Swiftvlm: Efficient vision-language model inference via cross-layer token bypass,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132bf8bb-d0dd-40d9-96a0-259f560bd421 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Fit and prune: Fast and training-free visual token pruning for multi-modal large language models,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b61b77-33b6-4232-bd3b-dfeaa1c807e6 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Llava-mini: Efficient image and video large multimodal models with one vision token,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda51eae-d9b7-4f65-9f58-6e8067fb0f7a · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Lvpruning: An effective yet simple language- guided vision token pruning approach for multi-modal large language models,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1edaa9e9-2c2a-4a06-bcef-5a0127102bd0 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Matryoshka multimodal models,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a558b01d-b723-454a-b8f6-6bbf7e44ee96 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Matryoshka query transformer for large vision-language models,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089ab825-792e-428a-9c62-f70ae9564a87 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Improved baselines with visual instruction tuning,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3626fb5f-c1ba-4f9a-8255-a056036adf9c · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2bec4f0-55dc-4e30-8613-937675819682 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Mme: A comprehensive evaluation benchmark for multimodal large language models,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00ddbee-7c28-4e7f-ae45-562ce7929f69 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Towards vqa models that can read,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773bc19c-4e06-4a80-a624-ee3a1a702f9c · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf231ed-aef2-4a71-b883-5cee9fa5a70d · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11c6af6-3425-4468-a3a1-a874fad95616 · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Qwen2.5-vl,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ae835d-ec1b-4842-9824-74af18c345bd · outbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Pytorch: An imperative style, high-performance deep learning library,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.