Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T10:59:31.772378Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2506.01097.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T10:59:31.772378Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-16T12:39:57.398423Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-16T12:40:54.865986Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c26d6faf-e54f-4511-974e-3daee6bb58b8 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Gemini: A family of highly capable multimodal models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 463d9ccb-700d-4596-b93c-2951ae06dece · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Qwen-vl: A frontier large vision- language model with versatile abilities
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a323e9bf-818f-4022-a267-0d28f2477c79 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Deepseek LLM: scaling open- source language models with longtermism
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d0bb35c0-7361-4084-93c4-ec94a29beff8 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Token merging: Your vit but faster
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ece57648-e3e0-4a50-b0f7-dc302d43fb59 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, et al
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6740cd9d-f69a-4f5b-8176-c8e17631b05e · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2cde2d05-9939-4d8a-9057-0ec4f26d4d11 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Transformer interpretability beyond attention visualization
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e0f0848-5fc8-4fb4-8234-529094968439 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5bdd3da-90c3-4771-9a40-c63b2eb27b24 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Are we on the right way for evaluating large vision-language models? In NeurIPS
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7e59611-31b1-449d-98fa-affd21fb87f2 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e328d32-f596-44a6-81e3-e5f416e4ba38 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00612136-5f73-457e-aa88-91a2678a5f67 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Xception: Deep learning with depthwise separable convolutions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d765c91-b966-4ecd-8915-1aa63d15f7f4 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Fu, Stefano Ermon, Atri Rudra, et al
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0cb89053-8c1b-45be-9d08-ac9d3ab4c87f · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9895cb2-3f3b-4a1a-987f-374f3390de16 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Mmbench-video: A long-form multi-shot benchmark for holistic video understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb67f627-4052-42aa-8d68-6ddff2000fa0 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective MME: A comprehensive evaluation benchmark for multimodal large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad3832a9-14a1-4c00-bc25-7233fbdacf92 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e51329c3-b83b-41a7-8c10-4ef82f5c4412 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Exploiting behavioral consistence for universal user representation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8970b9f6-7091-4178-85b3-16e6b834d9d9 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Infinity-mm: Scaling multimodal performance with large-scale and high-quality instruction data
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 690c090c-ba6e-440c-8ce0-b47c12d64d2d · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Prunevid: Visual token pruning for efficient video large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f8cbbb8b-cd20-4d4a-801f-31034972905c · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Kingma and Jimmy Ba
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6608df9-224a-4c09-81ee-708bb4512b43 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Llava-onevision: Easy visual task transfer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cbd1a489-469f-4add-8b0c-db3ab4be9a2a · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Seed-bench: Benchmarking multimodal llms with generative comprehension
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1af1265d-7bdd-44f0-b884-71b1a87b5de1 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd308669-d817-4768-bd6a-aaed0e7405d7 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7ac3d09-63c1-4b81-b858-728cf3317b29 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Llama-vid: An image is worth 2 tokens in large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6346074-1edd-4d08-b476-0a4ab9f6f317 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Visual instruction tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6813d43a-fa9b-47e9-aec5-6827d169e46a · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective NVILA: efficient frontier visual language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 58415689-e9a5-4df8-b349-afac520ab9de · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective GPT-4 technical report
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2990951d-f910-48b0-ba3e-91b1fc4af379 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Instruction tuning with GPT-4
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 77d09e1f-76d6-41f2-beaa-6dc1859e6f94 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Efficiently scaling transformer inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9b3ba46-07eb-4746-8fa7-5492e1d8c84c · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Fastvid: Dynamic density pruning for fast video large language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15270be9-8b46-4eca-9e6b-583616f5ac96 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Tempme: Video temporal token merging for efficient text-video retrieval
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8616e710-937f-46e3-8e98-0089ceeff31d · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Tokencarve: Information-preserving visual token compression in multimodal large language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aaf2ffd8-0043-444a-b009-c2b428ab7e78 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Llama: Open and efficient foundation language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e16955e-c6c4-41f6-b026-883d180d7eba · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Analyzing multi-head self- attention: Specialized heads do the heavy lifting, the rest can be pruned
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec626fb4-423c-4056-962c-dd5c4d03d1e4 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective FOLDER: accelerating multi-modal large language models with enhanced performance
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6559ec26-13b8-4cc5-869f-0cfff9678a56 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Dynamic-vlm: Simple dynamic visual token compression for videollm
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c99d1974-4863-417c-ab43-8b7d91e57182 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 022417ea-8118-44a6-ab5a-03354d0ef972 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Next-qa: Next phase of question- answering to explaining temporal actions
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c7549e5-d1c8-4840-b59d-da3afa6246b8 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d49da3a2-9929-4bc0-8305-10911d2105fd · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Visionzip: Longer is better but not necessary in vision language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 58662a3c-35b6-426f-bef0-7298ac8cd680 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Deco: Decoupling token compression from semantic abstraction in multimodal large language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b58b009b-4828-49c0-a284-4a3254153a1f · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Mm-vet: Evaluating large multimodal models for integrated capabilities
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ee5a0b4-d8ce-4390-91fd-54dd017d4d72 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4567bb8e-21c6-422a-ac61-a76ed20c83cb · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54c26a27-b30f-47ac-8406-54c80c763bb4 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Sparsevlm: Visual token sparsification for efficient vision-language model inference
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50f7d372-fe02-4c2b-8dc2-4f14dbe6c130 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Video instruction tuning with synthetic data
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93308d3c-b335-4ceb-a3c3-f9bc2230dbf2 · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective A stitch in time saves nine: Small VLM is a precise guidance for accelerating large vlms
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f6733d5-8ba6-46b7-9af9-0c0bc74f516f · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Minigpt-4: Enhancing vision-language understanding with advanced large language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a3835aac-8831-463a-911d-3c2293451c7f · outbound
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective Focusllava: A coarse-to-fine approach for efficient and effective visual token compression
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bae3f6c3-5cdd-4409-b8f2-c31e6040e976 · inbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.