Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:13:11.715462Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 13 inbound Pith citation observations for arXiv:2506.19225.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:13:11.715462Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:15.479594Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:58.312479Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2ac8c17c-ef6b-4bb5-b661-b76f7c694d18 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b639ce59-0d16-4689-a0d5-feccc88d13f4 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1378af-d07c-4fa4-a737-f20479b34dcf · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Claude 3.https://www.anthropic.com/news/claude-3-family, March 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4fadd24-7603-4e0a-99b1-1f7d1ceb31ef · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee6a7a4-e6d5-40f7-8eb9-7aafec2469ab · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf7c58a-339b-43ed-9661-ce81ff93e999 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 700229ed-b13f-4ceb-9151-932718b384b3 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Nvila: Efficient frontier visual language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97be42b9-d35b-4687-a316-74b8b543f544 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaVA-OneVision: Easy Visual Task Transfer
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d6f43c-d283-44e8-8993-c24db3336a95 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0a7ef3-0626-44b4-8c0a-303153eea28b · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38fdb9fe-d313-4570-8f2b-565d97b42853 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Transfer from Language to Vision
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9bd112b-a33b-408d-b03c-4849703a8fb1 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42939bff-b41f-400e-a552-b7378b9660f8 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6344937-afcc-4870-aefa-9a1db9dde81f · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat: Chat-Centric Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ed5029-2983-493e-ab0a-c3a858ba68dc · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7cc081-46f7-4693-b683-5c2b592d35d0 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d543b0-28aa-4420-97a8-ca925f1e9c4a · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad403f34-6964-4c60-8b85-f105b25d7c35 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Snapkv: Llm knows what you are looking for before generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44361d5-4a69-4341-adfa-b6cbd04d5bfb · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Efficient Streaming Language Models with Attention Sinks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b84e53-46fd-45f3-9053-5d382c196328 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69586b48-fd67-484f-b4bf-51fc7821089c · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7182a99b-e30f-4d2d-8389-98e59e56bd28 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8140ec18-9ef6-4d59-a03d-33427aef3e36 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Visual Instruction Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57b1f109-c7b6-46c0-a018-99850cf9d76f · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21a8a70-8e6a-4077-9023-624668c7e886 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Flamingo: a visual language model for few-shot learning.NeurIPS, 2022
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78962752-4e28-4664-8d71-945f60701170 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892779bc-8de8-467a-9c2c-783d8d95eb7b · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46f4c0f-ea63-49b5-b8b9-e65a78283748 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Vidtext: Towards comprehensive evaluation for video text understanding.arXiv preprint arXiv:2505.22810, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde2b3d5-766a-4eb8-98d8-c27cdfbdd032 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Vid-SME: Membership Inference Attacks against Large Video Understanding Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ecbc304-40b5-428f-8234-871d5a1d251c · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e61a192d-bfd2-4f6a-a674-b10c7a491ad3 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Llama-vid: An image is worth 2 tokens in large language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d632ca7f-a4c0-4de8-9889-4d3d58abaa42 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc11cca0-5f31-4481-8d6d-72be0407d318 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Token Merging: Your ViT But Faster
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c1400bf-89e9-49e6-8860-b49ebd31ec38 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Compression with Activation Beacon
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae348ee-dcfc-42ba-8eb1-4c7c705acd17 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Lighter and better: Towards flexible context adaptation for retrieval augmented generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71970c1b-8b03-4058-a55e-e5b57db8986c · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be77acfd-9103-4b5b-8419-9107f89383fb · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4fe11e0-a9d9-4a45-832a-8ac68d19e073 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7987dfb0-bcf0-4639-9fc3-122cf89bc1c1 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb51bb97-22e8-4030-9768-4f0a8f198555 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-xl: Extra-long vision language model for hour-scale video understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa52dc7-5d5a-4b51-8d3d-97c898209423 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff13dda0-1cb8-4aa8-a1d6-dbe5b2d43db2 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96e6cf11-c3aa-4c13-b51b-147e6c544cae · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Qwen2.5 technical report
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ef0ce03-ebad-4ebd-8418-b74ff83892fe · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36d57352-bccc-4850-955c-2236ca62a090 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 606b8738-b551-45d9-b74b-f46224ce850f · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Internvideo2: Scaling foundation models for multimodal video understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 817a48cc-2a7b-4f0d-a292-666c9c209a86 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b144f7-02c7-4305-af42-64cee1aca6f3 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Memory-enhanced Retrieval Augmentation for Long Video Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a75ceb8b-10b7-4ebc-bc28-60f157a9b6ee · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MLVU: Benchmarking Multi-task Long Video Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed07ec05-ffd7-4a34-8c26-c952a8137a10 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21397d08-1c0d-4164-9c4f-9539c132815f · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51b153dd-a8c3-4786-8d0b-f0d8d967a8e5 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LVBench: An Extreme Long Video Understanding Benchmark
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7fc560d-b42f-45f3-a437-4be1ade2a1d1 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741a5211-f14f-4cc4-adce-e004e9b0feae · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Tall: Temporal activity localization via language query
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c4697dc-4d72-452b-95f8-863bae8de435 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a909771-f00d-4697-acda-ccf2c2a7a875 · outbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af469056-2597-4916-802c-3e9a4e6eb44a · inbound
Infinite Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb56a904-e9af-4aa6-bde5-e33822a78dcd · inbound
Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6eec0db-231d-4e97-a023-493219000c2a · inbound
$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1ac90b-963c-42f6-aea2-6103032d8c14 · inbound
cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddef1e6c-2631-451c-9f2b-f16c6b743166 · inbound
CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b87dd21e-68ef-4313-9a82-2ef1a7024db7 · inbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae7501ad-991c-4e6e-8fac-cec4f55cf572 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55fdb6c8-43b3-4cec-8dc1-257732f2f102 · inbound
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 605428c9-ced5-4e26-b65c-746de9d130bc · inbound
FOLIO: Focused Semantic Memory for Streaming Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e95a7ad-809f-46f5-8825-6b314a59a3ec · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e370153-360e-4d06-9b8a-c84ad2d9ff37 · inbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f8a689-5d3e-49c2-b6f5-3ecbe11ffdf0 · inbound
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bcfcad6-5456-4d17-b706-deb69eae8560 · inbound
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.