Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:55:50.162416Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2608.03112.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:55:50.162416Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6aa95c04-c4fd-424c-bf4e-5f080bb33dc9 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Divprune: Diversity-based visual token pruning for large multimodal models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42b15421-c221-4477-b3a9-e3a5bfdbbe09 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ec23800-fd5b-4476-8bc9-bb6b9304a4b7 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a00e5f9-79d9-4cc1-b3a2-4b63c60a11b1 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6aaaff18-9677-42cd-9ac8-98930dc07965 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04ecb77-5faa-4c3d-a1c5-ab31c2c5dd2b · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Instructblip: Towards general-purpose vision- language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43fb8f65-9a80-46fd-9c21-fc66ef589dfb · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126021bf-1271-4442-a0ca-2f14caebe633 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97536288-417c-4bb4-a920-90d2bcc1e097 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Filter, correlate, compress: Training-free to- ken reduction for mllm acceleration.arXiv preprint arXiv:2411.17686, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e115841-d1f7-4f52-aae3-970b7392d232 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Efficient Multimodal Learning from Data-centric Perspective
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b88ced9-c6c7-4e39-8b1f-f92e6d2fd0e6 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Ivtp: Instruction-guided visual token pruning for large vision-language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 475ef900-2adf-4dfa-aa48-3fb750537e2c · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Fast pruning using principal components.Advances in neural information processing systems, 6, 1993
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 17d30c23-5d1b-4b62-b66b-5d3912721494 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2efe416-7b05-43a0-938b-57e1118d90c7 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Llama-vid: An image is worth 2 tokens in large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2fbb7ef-1df6-4de6-a5ce-e13c9d235442 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0860654-ac66-4fdc-9b8d-f70b7d351ac5 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Video detail caption
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c51946b4-8667-4d8d-8930-2d124f6c9292 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 08fe127c-bdcf-4157-b0ee-2d503c0556d9 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 581c3805-4d45-4d1e-bd00-1602d8e60b17 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Sys- tems, 36:42748–42761, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30bbac77-b536-4487-b583-6426e3c42043 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cad450f-7298-4506-beb1-ab398a6982d8 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Imp: Highly capable large multimodal models for mobile devices.IEEE Transactions on Multime- dia, 2025
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 346c318a-9fc1-49a1-8dd2-d59fbe86ab65 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb7bdd5-f0d4-4314-9070-1171eb84c3f4 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636baae2-9a92-4e79-84c1-c37fdd4ea012 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5449c859-83cf-4de4-87c8-6b5172039234 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10cb08f1-eda7-487a-89fd-613ca4a412ab · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 637995de-1493-4280-9a81-4a9e02102bf0 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Next-qa: Next phase of question-answering to explaining temporal actions
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c40ef14-ccdd-40e2-ba5e-56e629d10789 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Topv: Compatible token pruning with infer- ence time optimization for fast and low-memory multimodal vision language model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f98ddbd-2456-4d52-bcca-2b69b771b108 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Atp-llava: Adaptive token pruning for large vision language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ed439344-4c14-4165-ae58-c54fe0510026 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b898cd-650a-4113-89c2-8fac10f34d16 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89dabbd2-75bb-43a6-9a2e-664507f1caa0 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e854eb82-62e2-4e56-8079-a8cc27b8091f · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Also, we use beam size of 1, and the number of maximum new to- kens is capped to 1024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f25dd4a7-d912-47b3-91ad-b1909b26b579 · outbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.