Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:39.894521Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.21184.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:39.894521Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0bf89560-8927-4e62-a325-3ac8e3dc034c · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Flamingo: a visual language model for few-shot learning.NeurIPS, 2022
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89ab8ae2-2bac-47c4-9175-999e557b1902 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c9cd646-4ba3-44ce-8828-ca0f1aeea3ea · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636967fc-0fe2-4008-8d22-6ac784cd54cc · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df79a092-af4f-40f0-8f24-efedd0d61324 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9a38c5-4775-4a78-91b0-82d2885203f9 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding NVLM: Open Frontier-Class Multimodal LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29079ea3-7f5a-45e3-b673-b96d566aed81 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad67feca-2ef2-4227-929e-8d25852f59fd · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2800f197-06c6-49eb-bd7d-99a1b40740be · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2140fb08-a1a6-43e6-a721-ec067c175a23 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a744d3-7a27-4fd3-be55-82d104ffe9b1 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a35b2b65-cf45-416e-b0d5-41cc98cb33d9 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52d6f597-9a2e-40e1-8491-2f15f6a944e9 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e90a2db6-5f6f-4e3e-a0ed-6b1928ac8116 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de1bfdc-6e02-41d3-98f1-9a17c5964c27 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa3dd12-bc54-41ac-8ae1-f4f037053790 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Llama-vid: An image is worth 2 tokens in large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8481452-cfd8-4211-b57b-d2ce89db70bc · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d511ea-937f-486f-8f0b-2f0166c69cb9 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1f090d4-785c-47d6-aa1c-6eb6d28552d2 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Visual Instruction Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699307ac-b544-4401-aa11-366d012c7f29 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Visual instruction tuning, 2023
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e75c140-9583-4936-995e-a5ad9f726e48 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b1cea3d-234c-40a5-8664-57c158420a68 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a50540c1-0041-4ace-8dd5-f98bb031f217 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 450847d4-afc5-4fe9-9ad1-f79661ed5286 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0589e541-1a20-440f-a926-855b8c3b8573 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Gpt-4 technical report, 2023
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b2c5357-e7ff-4183-bfbd-6f2f9be7bf05 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92e400c0-ffa5-4f75-ac26-451ca4a7d5f8 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f4681f-c7bd-4a2d-b978-8acdd833a70a · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b21a56-e1ad-4778-a6ce-c322c3a42042 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b25ca4dc-5cdc-466e-a600-f35ebdf1f7d5 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55d743eb-1d52-4b25-92e5-2c110b521247 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0fe3f2f-7ee5-458f-9016-4aeb9d543755 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be7b36a6-61cb-4357-be28-6e974b9a6c70 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad14b21-a422-405e-9593-c7614134390f · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69461d4-c78c-4c67-801a-af0204f757d2 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Internvideo2: Scaling foundation models for multimodal video understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 333a114b-7802-40da-aee7-bd656b079e76 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc1e9d7c-528b-49e6-9fda-5a5a44fc868c · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6124ce6e-9411-4742-a532-b56e838e24cf · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d34835d6-5ad3-46c6-b379-d68751029d8a · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ed76ec-112b-4f43-92a6-a1c7b91af129 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48fd399-8528-405c-9c19-0202d1e113de · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122a666a-3a86-4f3c-a925-119f5f0a1358 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Long Context Transfer from Language to Vision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ddb7bf0-c48a-4da6-8128-73469e3d4e19 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Llava-next: A strong zero-shot video understanding model, April 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5360d085-0391-455f-b137-d4d1dc11a146 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52688802-2f3f-46f2-8c31-bdd3aad996d9 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef244d8-c636-4f1a-8998-377c3e8507fa · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03bdf067-a8a2-409b-ace8-51dffb6b97b6 · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e76717b-f6d5-4173-ba86-00d49d6b5f0d · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349009e5-9f3b-4f1c-9ff6-80ab18e6d2ae · outbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.