Pith. sign in

Paper Citation Record · LEDGER

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective

As of 5 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2506.01097.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01097 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T10:59:31.772378Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T12:39:57.398423Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-16T12:40:54.865986Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c26d6faf-e54f-4511-974e-3daee6bb58b8 · outbound

This paper cites Gemini: A family of highly capable multimodal models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Gemini: A family of highly capable multimodal models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.792368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:9cdec2a27e310b74cb94a1f273651b90b35f0ed1271f5741f5ef39a9b8cad0fe

Observation 463d9ccb-700d-4596-b93c-2951ae06dece · outbound

This paper cites Qwen-vl: A frontier large vision- language model with versatile abilities.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Qwen-vl: A frontier large vision- language model with versatile abilities

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.781798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:602d99c084acddc8f735932bda5d5be6c6ff21fdb9184d5c5657b83e1ba1913e

Observation a323e9bf-818f-4022-a267-0d28f2477c79 · outbound

This paper cites Deepseek LLM: scaling open- source language models with longtermism.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Deepseek LLM: scaling open- source language models with longtermism

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.758489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:998fa2d9f26fbfbf2ccbb7d6375e23ec1218462b3e9541b8431db61b20bd3679

Observation d0bb35c0-7361-4084-93c4-ec94a29beff8 · outbound

This paper cites Token merging: Your vit but faster.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Token merging: Your vit but faster

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.814716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:fc1ebe577f7ec51551047ba96a46e974fddf35bed6e3806092756b38251fdef3

Observation ece57648-e3e0-4a50-b0f7-dc302d43fb59 · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, et al.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.754726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:bf353621b05c88196f0395cf9fceaf783f4838b0abf30a8333197e1c134e3cbd

Observation 6740cd9d-f69a-4f5b-8176-c8e17631b05e · outbound

This paper cites Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.742900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:95cf7f3be35a494a894611945f6ad4b953437b0bc05e84e94c50a4be0373cba1

Observation 2cde2d05-9939-4d8a-9057-0ec4f26d4d11 · outbound

This paper cites Transformer interpretability beyond attention visualization.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Transformer interpretability beyond attention visualization

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.826915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:56d2d5d416f1e7a81e811e0e22e3b35d2670eef0cc982790feee5cfd0222774d

Observation 4e0f0848-5fc8-4fb4-8234-529094968439 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.788918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:f2b25c8e5976e47897ee91d762babf46c4ecd6ddf24704552dc2ceb7e83091cd

Observation b5bdd3da-90c3-4771-9a40-c63b2eb27b24 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In NeurIPS.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Are we on the right way for evaluating large vision-language models? In NeurIPS

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.752110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:b5b2c3dcc86fc44e6036c34ace14d2b917bd5041beb7aeaf3971a2888d0a976e

Observation f7e59611-31b1-449d-98fa-affd21fb87f2 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.794180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:b3bd25250600fd0d801319d2d15bf6454647aa275075bdc8c59dfa097caa25fa

Observation 0e328d32-f596-44a6-81e3-e5f416e4ba38 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.796017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:034e20f1c80c23f9ef88fcda156d22b86634e5213802efd2707432d34e60cef3

Observation 00612136-5f73-457e-aa88-91a2678a5f67 · outbound

This paper cites Xception: Deep learning with depthwise separable convolutions.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Xception: Deep learning with depthwise separable convolutions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.785921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:0f7cf6ca9668f1c586e758b3c5e8933d462ff887fb5b8e085c3c6bf4559b974b

Observation 2d765c91-b966-4ecd-8915-1aa63d15f7f4 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, et al.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Fu, Stefano Ermon, Atri Rudra, et al

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.783782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:621aeb3e6d94e710b5c5f1b62d4aa770ee7f2bfa568129b5288944f36c1c2179

Observation 0cb89053-8c1b-45be-9d08-ac9d3ab4c87f · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.797439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:08e301d692596a322bf42e060971d66031218b035b09f2b4a07170aefa3dfa7f

Observation d9895cb2-3f3b-4a1a-987f-374f3390de16 · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.795616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:8295a64b77918f817a8c1920940170413a8ade2bff07800ab9b1e47aa4683f6e

Observation fb67f627-4052-42aa-8d68-6ddff2000fa0 · outbound

This paper cites MME: A comprehensive evaluation benchmark for multimodal large language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective MME: A comprehensive evaluation benchmark for multimodal large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.764808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:73150e2141759f786799c7bbd16831cf6bdb8e124f12205ae0b2bbfc706edbe9

Observation ad3832a9-14a1-4c00-bc25-7233fbdacf92 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.761076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:664e9e7bf49c78aee4544a1495ae094e84e6b0af11096cebdc0eb664cf7c89cf

Observation e51329c3-b83b-41a7-8c10-4ef82f5c4412 · outbound

This paper cites Exploiting behavioral consistence for universal user representation.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Exploiting behavioral consistence for universal user representation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.799856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:6127e94c990f41b086511fba366fef8a447bcceb3164ae1a8083936638f5e2f4

Observation 8970b9f6-7091-4178-85b3-16e6b834d9d9 · outbound

This paper cites Infinity-mm: Scaling multimodal performance with large-scale and high-quality instruction data.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Infinity-mm: Scaling multimodal performance with large-scale and high-quality instruction data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.801323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:ab5fecc1e53312add57fe8429877cdb315e22db09e6c4a6f9df2082cd68d9763

Observation 690c090c-ba6e-440c-8ce0-b47c12d64d2d · outbound

This paper cites Prunevid: Visual token pruning for efficient video large language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Prunevid: Visual token pruning for efficient video large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.781537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:4c0a0b36635d7300b8546e6143a2c35d6a630ed463d705c196eaced983301544

Observation f8cbbb8b-cd20-4d4a-801f-31034972905c · outbound

This paper cites Kingma and Jimmy Ba.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Kingma and Jimmy Ba

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.812538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:54aad07070b69fc8cb2ebc1a07a3c48e200b7dce79cd975eb0dafe569685c608

Observation b6608df9-224a-4c09-81ee-708bb4512b43 · outbound

This paper cites Llava-onevision: Easy visual task transfer.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Llava-onevision: Easy visual task transfer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.807554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:38f52c9c27b21218ce52a4ef8235634f775b91ec67639bd40bbf067bd4149eb8

Observation cbd1a489-469f-4add-8b0c-db3ab4be9a2a · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Seed-bench: Benchmarking multimodal llms with generative comprehension

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.790507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:2f00fa3fe81bb0256658f504d4b7212840c78b511d67e1587bac938ef0762b7f

Observation 1af1265d-7bdd-44f0-b884-71b1a87b5de1 · outbound

This paper cites an unresolved cited work.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-19T11:03:02.793888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:d0e755b559635c63d4912179b102a882c2ad0607c1afa3dc704da5f4d1b60503

Observation cd308669-d817-4768-bd6a-aaed0e7405d7 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.806751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:26f65b0784c6f1a5659e7e34dace582a182e70f8a4b878ea1228648e4c1a686c

Observation f7ac3d09-63c1-4b81-b858-728cf3317b29 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Llama-vid: An image is worth 2 tokens in large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.768182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:0156cc39d79b031f4d8b0a920a2a3b009c5bc4d44e79a655ecfd4230debcadb3

Observation f6346074-1edd-4d08-b476-0a4ab9f6f317 · outbound

This paper cites Visual instruction tuning.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Visual instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.753035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:a9472d781bdc372cb1b3a3e1552a89f60a12d0bb032385aa36fa8b2fa6bd8f4b

Observation 6813d43a-fa9b-47e9-aec5-6827d169e46a · outbound

This paper cites NVILA: efficient frontier visual language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective NVILA: efficient frontier visual language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.822153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:b431fee7cd4475bc3d31a04026e3fa6259884ebece7bcb3a44f858cd7a939349

Observation 58415689-e9a5-4df8-b349-afac520ab9de · outbound

This paper cites GPT-4 technical report.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective GPT-4 technical report

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.779785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:9fe64bc4acdcde6dc8a15327c8d725d021f2051d9ffb2a6bfb259805f64be9b2

Observation 2990951d-f910-48b0-ba3e-91b1fc4af379 · outbound

This paper cites Instruction tuning with GPT-4.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Instruction tuning with GPT-4

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.772214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:775127cc9e89bccf8da2c4d78e77067b9551d61c0f15250a365e8dbbee5a817f

Observation 77d09e1f-76d6-41f2-beaa-6dc1859e6f94 · outbound

This paper cites Efficiently scaling transformer inference.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Efficiently scaling transformer inference

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.773174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:64fa8c4204a856fcb6e6638093555cf3da3b71a4de3c2aca993d9f882d9cc0ba

Observation c9b3ba46-07eb-4746-8fa7-5492e1d8c84c · outbound

This paper cites Fastvid: Dynamic density pruning for fast video large language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Fastvid: Dynamic density pruning for fast video large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.766521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:19c547a4403b0855387db2123677a4ae7b4f8b8d579c3e75949029db7073f68b

Observation 15270be9-8b46-4eca-9e6b-583616f5ac96 · outbound

This paper cites Tempme: Video temporal token merging for efficient text-video retrieval.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Tempme: Video temporal token merging for efficient text-video retrieval

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.776491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:78e970da2624c18d81ccfb3a4b6f5e8de08e90fdabcc12d3b058863873531b7a

Observation 8616e710-937f-46e3-8e98-0089ceeff31d · outbound

This paper cites Tokencarve: Information-preserving visual token compression in multimodal large language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Tokencarve: Information-preserving visual token compression in multimodal large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.778158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:671c20c04d90f0c307ca6b8c0fd6c87e209fc488f7298a602a7a8eb74fbb29a6

Observation aaf2ffd8-0043-444a-b009-c2b428ab7e78 · outbound

This paper cites Llama: Open and efficient foundation language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Llama: Open and efficient foundation language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.785219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:d35354baef06cf76cb73bc02928564427b3c3adf421d8488824df93747973835

Observation 5e16955e-c6c4-41f6-b026-883d180d7eba · outbound

This paper cites Analyzing multi-head self- attention: Specialized heads do the heavy lifting, the rest can be pruned.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Analyzing multi-head self- attention: Specialized heads do the heavy lifting, the rest can be pruned

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.745696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:587cc5d84263b66c384537379aa0b7eaf594c7ac7d21c913740b941c780c7d8a

Observation ec626fb4-423c-4056-962c-dd5c4d03d1e4 · outbound

This paper cites FOLDER: accelerating multi-modal large language models with enhanced performance.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective FOLDER: accelerating multi-modal large language models with enhanced performance

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.759247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:36f8c91efb9f169c8f3fb80a908b0d52fb166c0e91449fb33db04650efc219fb

Observation 6559ec26-13b8-4cc5-869f-0cfff9678a56 · outbound

This paper cites Dynamic-vlm: Simple dynamic visual token compression for videollm.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Dynamic-vlm: Simple dynamic visual token compression for videollm

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.810944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:d370d198634b0987405fa4e41371974bde85ca1dbe48fac4d1d8fe9f7adb63a0

Observation c99d1974-4863-417c-ab43-8b7d91e57182 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.799208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:7f7413ad59dfe4ba91ebe6b5c35048fe6fb162296ad8c7578f6bfabcbffbcc61

Observation 022417ea-8118-44a6-ab5a-03354d0ef972 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Next-qa: Next phase of question- answering to explaining temporal actions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.764600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:bf1846cc689a6eb3ad0accd59ec29f59f6c1a50616a06dbde4a821d84ed63e82

Observation 2c7549e5-d1c8-4840-b59d-da3afa6246b8 · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.750211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:2f2f5b0f8bf960fcf96e6e62883ea060718e69c7db7305ac626531cbcac0ca63

Observation d49da3a2-9929-4bc0-8305-10911d2105fd · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Visionzip: Longer is better but not necessary in vision language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.819923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:c63d0c9dd8729a2b9e7fd79b932151513bb99e3160b7f05dd5b1f13d95a68732

Observation 58662a3c-35b6-426f-bef0-7298ac8cd680 · outbound

This paper cites Deco: Decoupling token compression from semantic abstraction in multimodal large language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Deco: Decoupling token compression from semantic abstraction in multimodal large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.792193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:0122ff30bdd19373b72e9abe37d3226129ef16391ec70fb58d309e07dec30ad0

Observation b58b009b-4828-49c0-a284-4a3254153a1f · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.748335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:e96d1d047f0e68ab84f25aa4d867f59b3fbb312a55c8fa4ee8a478bc21490607

Observation 3ee5a0b4-d8ce-4390-91fd-54dd017d4d72 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.786937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:90e6c670d79291a8ccbe6e2bc89627eaa2d54be7d2514e864895fdb85e0ee08e

Observation 4567bb8e-21c6-422a-ac61-a76ed20c83cb · outbound

This paper cites Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.747425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:f1a951f453c6b248eb42dc6789dac85c99a0e3b94aabd78b30ce4f96da7cab52

Observation 54c26a27-b30f-47ac-8406-54c80c763bb4 · outbound

This paper cites Sparsevlm: Visual token sparsification for efficient vision-language model inference.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Sparsevlm: Visual token sparsification for efficient vision-language model inference

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.817470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:eed05da418eb60fd2d531dc474e90afbf188d242de4d040ccaa50b48b2b237a6

Observation 50f7d372-fe02-4c2b-8dc2-4f14dbe6c130 · outbound

This paper cites Video instruction tuning with synthetic data.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Video instruction tuning with synthetic data

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.774821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:4b616a5d877f02fcd31c524c736d04299d0fe9c4a3a34c866499b8812cae1638

Observation 93308d3c-b335-4ceb-a3c3-f9bc2230dbf2 · outbound

This paper cites A stitch in time saves nine: Small VLM is a precise guidance for accelerating large vlms.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective A stitch in time saves nine: Small VLM is a precise guidance for accelerating large vlms

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.824360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:bcaf2497027c6e4f08a2ffd6ee0b24b08b6c13b0bb92a0f2cbebd7fa4b306ba9

Observation 0f6733d5-8ba6-46b7-9af9-0c0bc74f516f · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:03:02.808694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:852b40740ee023dd8883761a0a7f41322c8866ca2d1006c52d61a92ad67a714e

Observation a3835aac-8831-463a-911d-3c2293451c7f · outbound

This paper cites Focusllava: A coarse-to-fine approach for efficient and effective visual token compression.

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Focusllava: A coarse-to-fine approach for efficient and effective visual token compression

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-05-19T11:03:02.814781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:59:31.772378Z digest=sha256:2f97989e96107ad795119c5d11241e822901b0d182ecf897bbec3b1d39d2d769

Pith citing papers

Observation bae3f6c3-5cdd-4409-b8f2-c31e6040e976 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective

Reference 165

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:40:54.867214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:8087976396e88e322effb00d98ff9b1a8bb85791c5418748354f2b198a2d5fd7