Pith. sign in

Paper Citation Record · LEDGER

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 13 inbound Pith citation observations for arXiv:2506.19225.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19225 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:13:11.715462Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:15.479594Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.312479Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ac8c17c-ef6b-4bb5-b661-b76f7c694d18 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.005716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.005716Z digest=sha256:498ec27116be71f143a40fda24c5d1b025ac2155e097be17cc2ff27907516da1

Observation b639ce59-0d16-4689-a0d5-feccc88d13f4 · outbound

This paper cites Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.063644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.063644Z digest=sha256:72b1ef9a0d1fa08dfabbda8c9216d07b88693550834ae7e8bf4999dd5d88dae1

Observation 7e1378af-d07c-4fa4-a737-f20479b34dcf · outbound

This paper cites Claude 3.https://www.anthropic.com/news/claude-3-family, March 2024.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Claude 3.https://www.anthropic.com/news/claude-3-family, March 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:14.477127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:13:06.126866Z digest=sha256:d009d928c26c60e089e1330da1ae9076dc8a172ebbe82a9a62b00b9f1f52b440

Observation f4fadd24-7603-4e0a-99b1-1f7d1ceb31ef · outbound

This paper cites Qwen2.5-VL Technical Report.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.296080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.296080Z digest=sha256:a0063179758e426e3f7264c238410de8367389fd54e09e14c8bb39fd84d8415a

Observation fee6a7a4-e6d5-40f7-8eb9-7aafec2469ab · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.349713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.349713Z digest=sha256:574052369d026b641848c702d0e012ac2d3bb63872dbeffed929f3bba17802da

Observation bdf7c58a-339b-43ed-9661-ce81ff93e999 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.438610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.438610Z digest=sha256:0e5130e8ef03cea85169c0ea15d516468819843b002d7669ffa8833be64de7a5

Observation 700229ed-b13f-4ceb-9151-932718b384b3 · outbound

This paper cites Nvila: Efficient frontier visual language models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Nvila: Efficient frontier visual language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:14.244169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:13:06.555749Z digest=sha256:40bd52673e74d599fd453b1783f2ac90100d5f6cd0a14a40f8306a821d1b3568

Observation 97be42b9-d35b-4687-a316-74b8b543f544 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaVA-OneVision: Easy Visual Task Transfer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.733535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.733535Z digest=sha256:158fd58a06e767145df208e468458d5461df49b9b6bc4e263b51b8e6978e2bfe

Observation 15d6f43c-d283-44e8-8993-c24db3336a95 · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.878329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.878329Z digest=sha256:261881a77a7ac9426f4ceba81be993db5fcb5c0c08824599f755af4719cf5a17

Observation db0a7ef3-0626-44b4-8c0a-303153eea28b · outbound

This paper cites Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.962953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.962953Z digest=sha256:e1211f41f099f4d218e3ad03d51f171a88b3cf53618fdc92f1f2f3f034509c8f

Observation 38fdb9fe-d313-4570-8f2b-565d97b42853 · outbound

This paper cites Long Context Transfer from Language to Vision.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Transfer from Language to Vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.042420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.042420Z digest=sha256:be0cf555e415bf078698a8e3581793d4f56ba65cd7ffffa20bce5b53157ba710

Observation f9bd112b-a33b-408d-b03c-4849703a8fb1 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.193817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.193817Z digest=sha256:8ac869cb587ff6b8789ab8d7febd75fe96a13600f003e5bdc73eb0ed10e72145

Observation 42939bff-b41f-400e-a552-b7378b9660f8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.311670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.311670Z digest=sha256:81924ee7626e83247efc3be10a0017369cee842497a590300c89588b2c60f4e0

Observation a6344937-afcc-4870-aefa-9a1db9dde81f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat: Chat-Centric Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.440858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.440858Z digest=sha256:a6d365f0fe3098c5de72108112169f5209e46cb7d298ac02490021c66215822d

Observation 82ed5029-2983-493e-ab0a-c3a858ba68dc · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.554817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.554817Z digest=sha256:79f4786b7b7d3706986a3b9ebc72dd52366d149e41641c59113c5ec9d1c71d2b

Observation eb7cc081-46f7-4693-b683-5c2b592d35d0 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.690732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.690732Z digest=sha256:e5d1d85a9ef44e9efe3ce2c732495a3f0172e48ebeaba4381854dc4581e883fd

Observation 15d543b0-28aa-4420-97a8-ca925f1e9c4a · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.794754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.794754Z digest=sha256:9336230be0a9b704bcc12b718757bf300d72723a3cd25348b26d913e519bea54

Observation ad403f34-6964-4c60-8b85-f105b25d7c35 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Snapkv: Llm knows what you are looking for before generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.926631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.926631Z digest=sha256:8c8495cee3234ede189cff43f85dac46f34d94aa12cdf22c5b0a36394fb81264

Observation a44361d5-4a69-4341-adfa-b6cbd04d5bfb · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Efficient Streaming Language Models with Attention Sinks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.030081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.030081Z digest=sha256:c8d9a82af16e492355eaddda0ede0c481513f9a0517e976a450fa041398191b4

Observation 36b84e53-46fd-45f3-9053-5d382c196328 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:14.063101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:13:08.152487Z digest=sha256:4d4fb63e2fe357e4bba27dd1d92594b12fad8c6d68ac238b6d925a6f3de64fc7

Observation 69586b48-fd67-484f-b4bf-51fc7821089c · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.220997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.220997Z digest=sha256:d55dd8631ecc9343ca703cd2f0458899e80d660caaaa0f65447487c83219e3f0

Observation 7182a99b-e30f-4d2d-8389-98e59e56bd28 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.301692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.301692Z digest=sha256:3c06daeb9bcd655f17bff6a17e9e39787c3ea4ebb748f1b856a6ecd9eeec2ce1

Observation 8140ec18-9ef6-4d59-a03d-33427aef3e36 · outbound

This paper cites Visual Instruction Tuning.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.356194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.356194Z digest=sha256:b8d87bf9aae34eea7e46c61bc4e61be2385d59578d7c73052dc4bb8f55789600

Observation 57b1f109-c7b6-46c0-a018-99850cf9d76f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.443107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.443107Z digest=sha256:20fd8793a78ae07faff819c6700f4ef2e0d6325419d4219f3424ad90c4bd39e2

Observation c21a8a70-8e6a-4077-9023-624668c7e886 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.491853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.491853Z digest=sha256:2bd504320288ea6e8813b30b47ddd5991b7595282c64d2e3112a4e941f6056ba

Observation 78962752-4e28-4664-8d71-945f60701170 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.548099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.548099Z digest=sha256:cef24c3c792a50aca45d170ab7e55756d4e7e2efd6081625cd7282ca25468228

Observation 892779bc-8de8-467a-9c2c-783d8d95eb7b · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.631395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.631395Z digest=sha256:4781901ae862adc2c5db32d6afb67eea7c3e944008c81094c6763ef8234ce0ed

Observation b46f4c0f-ea63-49b5-b8b9-e65a78283748 · outbound

This paper cites Vidtext: Towards comprehensive evaluation for video text understanding.arXiv preprint arXiv:2505.22810, 2025.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Vidtext: Towards comprehensive evaluation for video text understanding.arXiv preprint arXiv:2505.22810, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.728844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.728844Z digest=sha256:4191f4cd2621cc7945d06efee5a3ed6e1432d3527ba024bd9363baaf69b8d8c6

Observation cde2b3d5-766a-4eb8-98d8-c27cdfbdd032 · outbound

This paper cites Vid-SME: Membership Inference Attacks against Large Video Understanding Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Vid-SME: Membership Inference Attacks against Large Video Understanding Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.790871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.790871Z digest=sha256:82e46a9146bca7114345c0d5feefc46ec971abfc02694b7eaa4b1d8b3d8f7fe8

Observation 8ecbc304-40b5-428f-8234-871d5a1d251c · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.857202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.857202Z digest=sha256:b711ae62d34dabbb8d63981196bafb1711cf0f7a4a82ae0bc6ae190138a73ba0

Observation e61a192d-bfd2-4f6a-a674-b10c7a491ad3 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Llama-vid: An image is worth 2 tokens in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.928962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.928962Z digest=sha256:886e46278199e24cc335ceb2f87785c59cb39ff6158b041976ec6a89faedf328

Observation d632ca7f-a4c0-4de8-9889-4d3d58abaa42 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.989110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.989110Z digest=sha256:4892770cc40244d090ade7f1efc615c448c9987513a30e1106eeb55c5bae829c

Observation dc11cca0-5f31-4481-8d6d-72be0407d318 · outbound

This paper cites Token Merging: Your ViT But Faster.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Token Merging: Your ViT But Faster

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.033586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.033586Z digest=sha256:001ae53d8c14f479d9feb022699efb33ed7e5907447dc0e7f5776b84016874fa

Observation 0c1400bf-89e9-49e6-8860-b49ebd31ec38 · outbound

This paper cites Long Context Compression with Activation Beacon.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Compression with Activation Beacon

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.127706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.127706Z digest=sha256:11439a8f9fd290cb0b86b8c50fc6461de1affb09de0d299ea94ffc003e427861

Observation 6ae348ee-dcfc-42ba-8eb1-4c7c705acd17 · outbound

This paper cites Lighter and better: Towards flexible context adaptation for retrieval augmented generation.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Lighter and better: Towards flexible context adaptation for retrieval augmented generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:13.797703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:13:09.293482Z digest=sha256:b6ebdc840f89393e7411990c5a0d8ccd1cb335593b742031073b95ec7475a65c

Observation 71970c1b-8b03-4058-a55e-e5b57db8986c · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.457172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.457172Z digest=sha256:9b557a3e597e9549aa575f20d5bdb278d7a2bfd3342f2a90409ded0983bb03bd

Observation be77acfd-9103-4b5b-8419-9107f89383fb · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.567257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.567257Z digest=sha256:025488738889c4cbc30ff1b7cb725d421f4df1e44b98b9c391f5ddf073d9c5c4

Observation d4fe11e0-a9d9-4a45-832a-8ac68d19e073 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.732084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.732084Z digest=sha256:d55d992bcb5782ce47050d746b12702a9febf6d877f0160080123ad91e400f7a

Observation 7987dfb0-bcf0-4639-9fc3-122cf89bc1c1 · outbound

This paper cites MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.877936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.877936Z digest=sha256:de914bb7a72dc1a45488b57d6a2f8f66a6b32e481c86af87f5d44d0cda36689b

Observation cb51bb97-22e8-4030-9768-4f0a8f198555 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-xl: Extra-long vision language model for hour-scale video understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.996456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.996456Z digest=sha256:b8e2f6a6abbae1c64d8e7b849c0c063855104ad0fbad5a82dd5f294af2a59819

Observation afa52dc7-5d5a-4b51-8d3d-97c898209423 · outbound

This paper cites ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.091131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.091131Z digest=sha256:7de55dd4cc94c02050ad980657cb4f84b092b8da78248130ec607e2f176d8410

Observation ff13dda0-1cb8-4aa8-a1d6-dbe5b2d43db2 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.165153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.165153Z digest=sha256:5cf8b26f9236f472c6ab0be5ee3b485ec37c92c03d9fe4639e9a62600c44a9e9

Observation 96e6cf11-c3aa-4c13-b51b-147e6c544cae · outbound

This paper cites Qwen2.5 technical report.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Qwen2.5 technical report

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:13.570856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:13:10.225687Z digest=sha256:4bf4d575929236236bdc71c781f19872cc5ad5360beeec5c085df787d46140a0

Observation 4ef0ce03-ebad-4ebd-8418-b74ff83892fe · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.317216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.317216Z digest=sha256:24ba02d2ec89feb1cf64e354c892dcdff00c6a98af294857ee55b54e242ccf7b

Observation 36d57352-bccc-4850-955c-2236ca62a090 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.420868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.420868Z digest=sha256:449c9b0b76df783ac8cd7169419da5452886018f0d124ac5aa0d4c27d167366f

Observation 606b8738-b551-45d9-b74b-f46224ce850f · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Internvideo2: Scaling foundation models for multimodal video understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.526472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.526472Z digest=sha256:ee55992862c2291a58a1b17f0935115a9103aa3943225e17088fe1c1efaa7908

Observation 817a48cc-2a7b-4f0d-a292-666c9c209a86 · outbound

This paper cites MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.664784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.664784Z digest=sha256:75c72fa919854ef456ee3328bfd1a784901265c7021fc794bdba10324b91ada2

Observation b3b144f7-02c7-4305-af42-64cee1aca6f3 · outbound

This paper cites Memory-enhanced Retrieval Augmentation for Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Memory-enhanced Retrieval Augmentation for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.804795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.804795Z digest=sha256:d15e63fd9b2a9617c8eba1c6763132df176bc3da74441a6495e816d15dd86125

Observation a75ceb8b-10b7-4ebc-bc28-60f157a9b6ee · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MLVU: Benchmarking Multi-task Long Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.893556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.893556Z digest=sha256:d24cd23d46efb24bcc39c1bdccf250b2c7d26e88064c4d64f865fd12308c0b81

Observation ed07ec05-ffd7-4a34-8c26-c952a8137a10 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.007680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.007680Z digest=sha256:d80788b748b46f732ef6dd92aad629d6892deacde1ea4cc507fdea28b6d362be

Observation 21397d08-1c0d-4164-9c4f-9539c132815f · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.123012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.123012Z digest=sha256:8efcd890e3063f44c421b8f982ac50a84a17aaebf369176ef3e22aa503f7fe1e

Observation 51b153dd-a8c3-4786-8d0b-f0d8d967a8e5 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LVBench: An Extreme Long Video Understanding Benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.200327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.200327Z digest=sha256:d1589fd3863403985e9529f3d9c8131cf97f646243935d4659df9c792fac6b57

Observation c7fc560d-b42f-45f3-a437-4be1ade2a1d1 · outbound

This paper cites VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.304314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.304314Z digest=sha256:ce3554dc414e86a2e0e71f7739e691062f2e754746d2110c08f5e0469d4d5dff

Observation 741a5211-f14f-4cc4-adce-e004e9b0feae · outbound

This paper cites Tall: Temporal activity localization via language query.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Tall: Temporal activity localization via language query

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:13.313796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:13:11.448061Z digest=sha256:693d81f4415cbe343474e051586bb1dd64bbf8cfc3d7655ff1e2304e226a0101

Observation 0c4697dc-4d72-452b-95f8-863bae8de435 · outbound

This paper cites V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.605015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.605015Z digest=sha256:facea586e3f44f48a282442de8bcb932b0444c79c1e5ab85de99fcb0247b007b

Observation 4a909771-f00d-4697-acda-ccf2c2a7a875 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.715462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.715462Z digest=sha256:f0ad90eac23ae8b940543227f8d34c2a5e7972156780c5a8988e4aca3b863aba

Pith citing papers

Observation af469056-2597-4916-802c-3e9a4e6eb44a · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.072860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.072860Z digest=sha256:f04ac4a4a4c74773c95a48c4490ed9d8d7d72cce3c163f3e45725e3228e1b7fa

Observation bb56a904-e9af-4aa6-bde5-e33822a78dcd · inbound

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models cites this paper.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.477147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.477147Z digest=sha256:3a00f68c1d5a65ed64927631095fc69855034c43a82f1032c87b8cec761a0d9d

Observation c6eec0db-231d-4e97-a023-493219000c2a · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.149183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.149183Z digest=sha256:c7789d64e125da12ddea0a4d280a88a1eae2db460c762f69f1d8585e15e3ab17

Observation 7e1ac90b-963c-42f6-aea2-6103032d8c14 · inbound

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations cites this paper.

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T09:09:32.932051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:09:32.932051Z digest=sha256:38db1c59c567bc21865f3c443738266fadb4b828de62cc3ac495358470328fd9

Observation ddef1e6c-2631-451c-9f2b-f16c6b743166 · inbound

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference cites this paper.

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:52.769031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:45:26.555395Z digest=sha256:dc694da526148a42090fd025687884f873022439692ac2fe7744c1ee83f8e003

Observation b87dd21e-68ef-4313-9a82-2ef1a7024db7 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.566701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:126335a193fcfbb292537f25be87ce91db8bfbfb9f89e4cfc8449aae9abff72f

Observation ae7501ad-991c-4e6e-8fac-cec4f55cf572 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.628146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:f3ba98f50e393844fb7f67bc25c1b9cef72b9d99b1615684bf2b764a369fb095

Observation 55fdb6c8-43b3-4cec-8dc1-257732f2f102 · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.314625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:5908df2cefc3619e9ec6949908645cbdf64de24309536dc6999a67278991784b

Observation 605428c9-ced5-4e26-b65c-746de9d130bc · inbound

FOLIO: Focused Semantic Memory for Streaming Video Understanding cites this paper.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.344843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.344843Z digest=sha256:74f31cf93c6caa32f39d00dc8e24709535aaab5fdb3efcf2879fd1b073a482e3

Observation 7e95a7ad-809f-46f5-8825-6b314a59a3ec · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.162135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.162135Z digest=sha256:d43fef46728b35ab8ffe42038025caea67794f9adb6280978633978d3daa0057

Observation 3e370153-360e-4d06-9b8a-c84ad2d9ff37 · inbound

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding cites this paper.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.782850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.782850Z digest=sha256:c345aac606921f8b52000df6db7469b98bfecc762f9bb34cf09fae3a83fb73a4

Observation 64f8a689-5d3e-49c2-b6f5-3ecbe11ffdf0 · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:46.830446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:46.830446Z digest=sha256:ed340ce367465c7f935c67e3e1ea902ef62596789ab977799a20ae479d347a36

Observation 2bcfcad6-5456-4d17-b706-deb69eae8560 · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:15.479594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:15.479594Z digest=sha256:99f13941b426b5559d4f35d8214e4091cc6c70b452fc23f3847163d2340c5d4d