Pith. sign in

Paper Citation Record · LEDGER

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2607.13500.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13500 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:03:30.238634Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4ab3642-ce73-454d-b2c0-611e1b14a6c4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.436267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.436267Z digest=sha256:f3dfeb1937629e898301567d39cb4e73b4c7f82574644c12ad63729c03edab42

Observation c9847a69-a0ba-41e3-afb5-90c2ec4a0642 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.536860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.536860Z digest=sha256:5e77db9a4d292cfd207da1481958b34006a43a56ed923e5ebdf481e2e8fda3ac

Observation 226cf9b0-c893-49a1-bc4f-2e0510e1024a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.689482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.689482Z digest=sha256:e88af65330689757645246927fce6afddf3be0a7d3c1fc536eaf8ea57dae64ea

Observation ea241e75-2bb5-4ef9-b833-a0bc3a062b41 · outbound

This paper cites BLIP- 2:bootstrapping language-image pre-training with frozen 11 image encoders and large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models BLIP- 2:bootstrapping language-image pre-training with frozen 11 image encoders and large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.814141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.814141Z digest=sha256:bd03c13c7bb81017498e07d572feab21f66f36145ffb813f60502d3d97dfc41e

Observation 9b2ade47-0381-4662-9bb0-0a2d01c718ef · outbound

This paper cites Cost-efficient and secure federated learning for edge computing,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Cost-efficient and secure federated learning for edge computing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.897278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.897278Z digest=sha256:4e8f1f87d2025bbf7dfeb57c47291f1288dae64dd50758f08deb80f7524de21f

Observation 0dbdf4ac-b3dd-40f0-9467-1989ec35030a · outbound

This paper cites To- wards online privacy-preserving computation offloading in mobile edge computing,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models To- wards online privacy-preserving computation offloading in mobile edge computing,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.037156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.037156Z digest=sha256:56dc23954dc8b0a933448723b51d8dbf9d2264c2c88140aa3817e1f1fb64056c

Observation c110d243-9905-40b3-a029-e4f09786feb6 · outbound

This paper cites AFLoRA: Adaptive Federated Fine-Tuning of Large Language Models with Resource-Aware Low-Rank Adaption.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models AFLoRA: Adaptive Federated Fine-Tuning of Large Language Models with Resource-Aware Low-Rank Adaption

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.152278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.152278Z digest=sha256:10ac76754c024483c9ebdb01c10e853b1bb1e8dbd9e480e64e6925b04a1f83a1

Observation 54509875-4356-453a-92fd-965ad18dd9bd · outbound

This paper cites Towards efficient edge learning for large models in heterogeneous resource-limited environments,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Towards efficient edge learning for large models in heterogeneous resource-limited environments,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.325447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.325447Z digest=sha256:14313ea23c333fbfe1fe10b2ab12f9378a20e953cad0f85b8751f8fafab84ad0

Observation 62585956-b6f2-4224-b033-c57ae6dbf983 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.434759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.434759Z digest=sha256:53f31c8dec3cd29f26a25fdce45204353988d1114462643d7d9c1ab1d938b573

Observation 1e4c373d-8840-48a8-b27f-9bdcb9d49091 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.501494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.501494Z digest=sha256:7eef700383a81c41694e3459de35ba4e14238896ce9b53419200ae5ac438b312

Observation b0e52ff1-b723-42a7-a6ba-b6613a029b58 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Sigmoid loss for language image pre-training,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.611352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.611352Z digest=sha256:81c21e67981d3bad3ac55593daf2c016e627f9e4fe1e0856f447adacb7949a64

Observation b3c0b6c8-aa10-4928-92c3-7aefb0c0232c · outbound

This paper cites Tap-vits: Task-adaptive pruning for on- device deployment of vision transformers,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Tap-vits: Task-adaptive pruning for on- device deployment of vision transformers,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.794944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.794944Z digest=sha256:a1deb05fb1a6829b476df8429687539d73735390dfdcf4411a30ef6fe4112779

Observation 4d088e23-6195-4a16-bcc8-5d894330ded2 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision- language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision- language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.924755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.924755Z digest=sha256:fcc17c7a033f39faf84588b368d200fb6a0a1dabc8abb7c13e182113d992344b

Observation f5b2543d-5b97-41a7-82d5-84b2e311bc4e · outbound

This paper cites VScan: Rethinking visual token reduction for efficient large vision-language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models VScan: Rethinking visual token reduction for efficient large vision-language models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.059229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.059229Z digest=sha256:44ecce330f4dc6d0db6c0be87559481aad26716b17de5c38b6ef56261510ca64

Observation 064b0bc9-7083-443c-8844-053fe9866142 · outbound

This paper cites Adaptinfer: Adaptive token pruning for vision-language model inference with dynamical text guidance,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Adaptinfer: Adaptive token pruning for vision-language model inference with dynamical text guidance,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.144959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.144959Z digest=sha256:acf671c68df55f1316fd663ccbb4d97cc33ca2423ac84cb5f6753c2f9b59a5f9

Observation 7fc6fe26-451d-44fc-84b4-94e73926579c · outbound

This paper cites Variation-aware vision token dropping for faster large vision-language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Variation-aware vision token dropping for faster large vision-language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.344745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.344745Z digest=sha256:d814dd2e5c49550d6c5369f0965b804d9cf725b08f7bda6b9d46b8a4e1842daf

Observation b7a58740-0054-4df7-bbdd-43b100638dd8 · outbound

This paper cites Sparsevlm: Visual token sparsification for efficient vision-language model inference,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Sparsevlm: Visual token sparsification for efficient vision-language model inference,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.444741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.444741Z digest=sha256:cf134aa17bd9dee7fd80c2ea7d66058b013b56714e6d7b2e0e231b300aaac21a

Observation 6e3d6150-0858-4287-b4a9-ec64c80658be · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyra- mid visual redundancy reduction,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Pyramiddrop: Accelerating your large vision-language models via pyra- mid visual redundancy reduction,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.552698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.552698Z digest=sha256:525b7efdb9440756bff8a75d6d3252f613fe2fa1d7cf7ed49c66ac630380853e

Observation 13493d25-1882-4ac2-9c22-5251d778b29b · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.664752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.664752Z digest=sha256:1dab3b29ec961d032ddcd4d48631f6df361702859d479817ebde7341f8a22b45

Observation d13bbfb4-1f49-43db-b16c-80d33edd3730 · outbound

This paper cites LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.782904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.782904Z digest=sha256:0e8d0e53448f5f369b086ff75dae07ae103962219a37de17dce4ee25bf0fa1bb

Observation 0ff3eba6-3f51-43f7-a7fb-97b6a2b91979 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Visionzip: Longer is better but not necessary in vision language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.870445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.870445Z digest=sha256:f237a86010f1ab09d080d03cde6269af6921df6a65601fdbfe8b8664d61c3d78

Observation b084ba45-397d-4a07-89af-fc30c640744a · outbound

This paper cites Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.948681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.948681Z digest=sha256:8b9750d3288f3e8d1098dc63694049032a89a5a309cf41769de2b5b727624b88

Observation 546ee96d-be61-4ad7-84d2-45d3f3cbd165 · outbound

This paper cites Flashat- tention: Fast and memory-efficient exact attention with io-awareness,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Flashat- tention: Fast and memory-efficient exact attention with io-awareness,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.092271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.092271Z digest=sha256:bba54f2eff2c78e74c28dfdb104087c07f0db00c88631ac09175178063ddeeac

Observation 2deb5366-dbe8-4b57-bbb0-61b467999fb5 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Flashattention-2: Faster attention with better parallelism and work partitioning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.234826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.234826Z digest=sha256:43222b7ee6726654534a62d4f3635517347bc9fb2580cbd02c4362744aecf1c4

Observation cc87243f-39eb-46ac-bcda-da8422f6b7ce · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Flashattention-3: Fast and accurate attention with asynchrony and low-precision,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.403994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.403994Z digest=sha256:28007606028a3bf08a02834b46608c8545cd88b7bee647f53e4ab717ec30444e

Observation 59915c0f-7492-4139-8a10-0ad3127b019b · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Efficient memory management for large language model serving with pagedattention,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.538342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.538342Z digest=sha256:8af0fb1ca08af13e9bd452a1ea462eb9ed42cbc45f440e5b961303c76f93f961

Observation 9774a8cb-bf6f-4697-814e-07c4097cb51e · outbound

This paper cites Similarity-Aware Token Pruning: Your VLM but Faster.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Similarity-Aware Token Pruning: Your VLM but Faster

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.630859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.630859Z digest=sha256:44b41c73f8c87af1f2611447113c861d1b042bb271fff7f11ec8ceee776e3a2c

Observation 22517f44-5ec4-418c-aa52-cf96dcdb49a2 · outbound

This paper cites DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.729468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.729468Z digest=sha256:28db199382f6028a96a6d2a9d806d70f8f1d9f5b3cd3e7f7041fb0dc10fcda4b

Observation 3d64b4bf-0c8f-4592-af83-ac3c21952221 · outbound

This paper cites Holitom: Holistic token merging for fast video large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Holitom: Holistic token merging for fast video large language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.825305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.825305Z digest=sha256:815b75b93a1d14b1b12387124f30912923a90d1fcb8db39f1661ded17d6c2046

Observation 5205495b-aa63-4cf8-b571-2f04b3603e2c · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Dycoke: Dynamic compression of tokens for fast video large language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.926783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.926783Z digest=sha256:5062f2956e9ef18802c4d9d17ee91fff311e47d225fd34097031e26965ac03b1

Observation 87784a5b-205f-46f8-b004-c1470cc9e687 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and- language tasks,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and- language tasks,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.013830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.013830Z digest=sha256:81ab518414c01ae9e22445c4afbbd5b2a8e05afa12863fa063b976410309080e

Observation be0da18c-1d61-449f-9964-be0b68b85197 · outbound

This paper cites Uniter: Universal image-text representation learning,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Uniter: Universal image-text representation learning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.121689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.121689Z digest=sha256:283834adc020255e7bbb2e331001accbfb58e080bb1ab4cdb826822c1a6149f0

Observation d8089c62-c843-492d-ad37-af8967cc2be0 · outbound

This paper cites Vflowopt: A token pruning framework for lmms with visual information flow-guided optimization,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Vflowopt: A token pruning framework for lmms with visual information flow-guided optimization,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.226288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.226288Z digest=sha256:3a5e2dedd8d43018ef36a6afe24e75b1c3f7b3fac07e0a637363bb5f24bd260f

Observation 6dd0f4b1-427a-4618-b40c-39e40a799b4a · outbound

This paper cites Hidrop: Hierarchical vision token reduction in mllms via late injection, concave pyramid pruning, and early exit,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Hidrop: Hierarchical vision token reduction in mllms via late injection, concave pyramid pruning, and early exit,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.371847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.371847Z digest=sha256:1da8f2d55db0a47d2bb9c852b0c9b07913050836efbef4804e0bc4c28274a4bf

Observation 4e8aa322-7b15-44e8-bcba-029479783c31 · outbound

This paper cites Todre: Visual token pruning via diversity and task awareness for efficient large vision- language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Todre: Visual token pruning via diversity and task awareness for efficient large vision- language models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.468713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.468713Z digest=sha256:d6989a27d0b000477e5c4c32960f7d1c842a983d8d7b220eb467e398bdab0709

Observation c635a1e2-7d3a-4d3b-8cb3-6c77aa6b9dab · outbound

This paper cites Tamp: Token-adaptive layerwise pruning in multimodal large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Tamp: Token-adaptive layerwise pruning in multimodal large language models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.594759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.594759Z digest=sha256:a517f25fd31f63d7bd174e4e153481f9e698e9c691d2b7d12246a6df850c8677

Observation c19994dd-9808-4d3f-8bd0-6836fccd94db · outbound

This paper cites Swiftvlm: Efficient vision-language model inference via cross-layer token bypass,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Swiftvlm: Efficient vision-language model inference via cross-layer token bypass,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.694745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.694745Z digest=sha256:65103658b546587e9f02ecad5681f22acbfa9a42f8de0500cd1d9a4cf6d22a12

Observation 132bf8bb-d0dd-40d9-96a0-259f560bd421 · outbound

This paper cites Fit and prune: Fast and training-free visual token pruning for multi-modal large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Fit and prune: Fast and training-free visual token pruning for multi-modal large language models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.821960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.821960Z digest=sha256:c6b3b01f6e519994974d48340401ddca55b484ed8841c5c66c1a18d47c9c2561

Observation f5b61b77-33b6-4232-bd3b-dfeaa1c807e6 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Llava-mini: Efficient image and video large multimodal models with one vision token,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.915880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.915880Z digest=sha256:c468e2ec5ea5980274de18bf4f2a7936d6ce6af189a47dcdf8ef6a7b9a00b4c5

Observation cda51eae-d9b7-4f65-9f58-6e8067fb0f7a · outbound

This paper cites Lvpruning: An effective yet simple language- guided vision token pruning approach for multi-modal large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Lvpruning: An effective yet simple language- guided vision token pruning approach for multi-modal large language models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.080528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.080528Z digest=sha256:157f89b5a388b79893763b071c71d5a5781bb76179c8db02fe92f83961f729f0

Observation 1edaa9e9-2c2a-4a06-bcef-5a0127102bd0 · outbound

This paper cites Matryoshka multimodal models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Matryoshka multimodal models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.165018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.165018Z digest=sha256:f2a66eb351c459cf84bc48df2f6916e94520842d6772126ed060ebcc652f6859

Observation a558b01d-b723-454a-b8f6-6bbf7e44ee96 · outbound

This paper cites Matryoshka query transformer for large vision-language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Matryoshka query transformer for large vision-language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.310156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.310156Z digest=sha256:ca4b8d45d998b1839b141e880851c61612d376aa8a78b49c2ad76e699776050d

Observation 089ab825-792e-428a-9c62-f70ae9564a87 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Improved baselines with visual instruction tuning,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.446041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.446041Z digest=sha256:f1c7d06fe58cb0d0fb32d58d814146c372a9f111a16a1fba2ae05c824dbea6d1

Observation 3626fb5f-c1ba-4f9a-8255-a056036adf9c · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.497826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.497826Z digest=sha256:69fb2125ee1d381b2fa5ce28b61726d23e2f475742dd2757e98e58674533727e

Observation d2bec4f0-55dc-4e30-8613-937675819682 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Mme: A comprehensive evaluation benchmark for multimodal large language models,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.612676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.612676Z digest=sha256:39e51fca588f6a15b0fec67ad6e2799811eff2d0cbc1f9f9a5864e09599ab1db

Observation b00ddbee-7c28-4e7f-ae45-562ce7929f69 · outbound

This paper cites Towards vqa models that can read,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Towards vqa models that can read,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.779055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.779055Z digest=sha256:05f3726a3aa3a77de933834bbe1af533a00e0658c474e9f28c19c76687843a02

Observation 773bc19c-4e06-4a80-a624-ee3a1a702f9c · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.912039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.912039Z digest=sha256:77b07449026bb912a0846f659a9957fa3eaf22e8b0094957a72e2cfb15b38c0b

Observation ddf231ed-aef2-4a71-b883-5cee9fa5a70d · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:30.032616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:30.032616Z digest=sha256:43a63ccfd7ccf43f7b7597f8ab349765adc97018a139cd9d926ff30e965253b6

Observation b11c6af6-3425-4468-a3a1-a874fad95616 · outbound

This paper cites Qwen2.5-vl,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Qwen2.5-vl,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:30.141140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:30.141140Z digest=sha256:1f6c77c7824fb99b4802e8dd3d48e78ba57d1e9dd29aa1f5cf1b311c65220616

Observation 71ae835d-ec1b-4842-9824-74af18c345bd · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Pytorch: An imperative style, high-performance deep learning library,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:30.238634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:30.238634Z digest=sha256:50e5793b4983b4454e60b70e82a7f440f1cfa108a1472e1cbb444e1aae2fc6f8

Pith citing papers

No inbound Pith citation observations are available.