Pith. sign in

Paper Citation Record · LEDGER

Task-Aware KV Compression For Cost-Effective Long Video Understanding

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.21184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21184 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:39.894521Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0bf89560-8927-4e62-a325-3ac8e3dc034c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:41.772060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:36:36.457741Z digest=sha256:520869aa35dd801a97f192412b32cbcd530a2a803a7395666ca5bb9b49e6589a

Observation 89ab8ae2-2bac-47c4-9175-999e557b1902 · outbound

This paper cites Qwen2.5-VL Technical Report.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.521460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.521460Z digest=sha256:6cd4020685ece5fbc2597a38d6227461228205cbb0422ede83e3a6d9b204e7cf

Observation 3c9cd646-4ba3-44ce-8828-ca0f1aeea3ea · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Task-Aware KV Compression For Cost-Effective Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.613827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.613827Z digest=sha256:469481bf3442e68b6d34664de8119e668c989105a9fd1820589cdd387fb6f3e9

Observation 636967fc-0fe2-4008-8d22-6ac784cd54cc · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Task-Aware KV Compression For Cost-Effective Long Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.671820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.671820Z digest=sha256:b9aa2886a7648e07856384eb3d0f04775777e2af520828720c526bdb569ed0e1

Observation df79a092-af4f-40f0-8f24-efedd0d61324 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.711180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.711180Z digest=sha256:50ed1b7b9d19ee147c8f51deb1faec771d335373bce1f1c558d08fec04f01cc7

Observation 2e9a38c5-4775-4a78-91b0-82d2885203f9 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Task-Aware KV Compression For Cost-Effective Long Video Understanding NVLM: Open Frontier-Class Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.783376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.783376Z digest=sha256:f697578670c99d227c5b07837362c97dceb084a2c6a1f1d16fd21ae07fca685c

Observation 29079ea3-7f5a-45e3-b673-b96d566aed81 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.868717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.868717Z digest=sha256:34d8faed40160238df183ba546117b562f72f15f77c11ece4b97cec144846432

Observation ad67feca-2ef2-4227-929e-8d25852f59fd · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.961899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.961899Z digest=sha256:626768bd88221c97c877e5155b2fea627ff740393d55b93ce5831ce0644da153

Observation 2800f197-06c6-49eb-bd7d-99a1b40740be · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.020202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.020202Z digest=sha256:6e126f63390f45c29cbde683e58e7296c131a99c1816100c938cb93362d2fe91

Observation 2140fb08-a1a6-43e6-a721-ec067c175a23 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.156162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.156162Z digest=sha256:69cebd1948f8ca95d3fc493ef168a65acab91dc9082c5d9b2b2d338deff89927

Observation f7a744d3-7a27-4fd3-be55-82d104ffe9b1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.209127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.209127Z digest=sha256:3075de2f3b9327f3464d3000b738245c2838820a3404d64d5feb192283426570

Observation a35b2b65-cf45-416e-b0d5-41cc98cb33d9 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:41.378601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:36:37.288226Z digest=sha256:addebb10bb0b8f2cad87ab630f87ae1b70f9eb2fd87a7147f2561e9f5a9f997a

Observation 52d6f597-9a2e-40e1-8491-2f15f6a944e9 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.359101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.359101Z digest=sha256:314637a6e7787b852a505120004eeab0abaf7ac06e783759d5b96ae230809246

Observation e90a2db6-5f6f-4e3e-a0ed-6b1928ac8116 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.422766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.422766Z digest=sha256:ff0e41dc0fe7087128f2ac3f59f1c69e324736a6c22c105eb2fa790ea1078e1b

Observation 3de1bfdc-6e02-41d3-98f1-9a17c5964c27 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.514961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.514961Z digest=sha256:aa982e7a158c0bc3d3a6858d8c1ac16677767596870ad44a806fa2a7bc7c3ab6

Observation 2fa3dd12-bc54-41ac-8ae1-f4f037053790 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.563665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.563665Z digest=sha256:87b4c521682d3acedc8c7bafb181f764b8de0aebc2d50e561284415476bd4b00

Observation f8481452-cfd8-4211-b57b-d2ce89db70bc · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.622776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.622776Z digest=sha256:5579bc7cde18807445171fed547a30bbd344fa189096e47ab51154d1e291e6f7

Observation 98d511ea-937f-486f-8f0b-2f0166c69cb9 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.715027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.715027Z digest=sha256:efdfcccf6d3b5cef3ac69e8e931eb5094b44a09616493265024e1204bce20b91

Observation b1f090d4-785c-47d6-aa1c-6eb6d28552d2 · outbound

This paper cites Visual Instruction Tuning.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.790351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.790351Z digest=sha256:3fe326a3f546526e422ef36779c8250cf604df815fb2cac13e7c945fe178f062

Observation 699307ac-b544-4401-aa11-366d012c7f29 · outbound

This paper cites Visual instruction tuning, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Visual instruction tuning, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.899741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.899741Z digest=sha256:48d3f6200f6ac4aa1d55dcdd3a0e9ff60502716ced2b3db426ed10d823103085

Observation 8e75c140-9583-4936-995e-a5ad9f726e48 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

Task-Aware KV Compression For Cost-Effective Long Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.979162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.979162Z digest=sha256:872aca68c02c44b8f73b7268c8ca699b5b6a25d3a29def7a6656e78ac69544c2

Observation 7b1cea3d-234c-40a5-8664-57c158420a68 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.084306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.084306Z digest=sha256:c3cf3cf847b58bd9caf45c7b032372fb946bffa639d1f37e38ac61695775bd7a

Observation a50540c1-0041-4ace-8dd5-f98bb031f217 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.176899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.176899Z digest=sha256:7e1b13f8e904731260459afbf62609fba6bf63543f8418af7f21637cb7408d6e

Observation 450847d4-afc5-4fe9-9ad1-f79661ed5286 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.221582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.221582Z digest=sha256:00d7e6a9b1c3c7c709312a0492c9abb5a6431c37b55615b61df88a0672006201

Observation 0589e541-1a20-440f-a926-855b8c3b8573 · outbound

This paper cites Gpt-4 technical report, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Gpt-4 technical report, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.276888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.276888Z digest=sha256:cd0cb16d7eac5d5d996a1ea42224d8ecd1961eba946930c8fef6ea6d22501546

Observation 6b2c5357-e7ff-4183-bfbd-6f2f9be7bf05 · outbound

This paper cites Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:41.041465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:36:38.352178Z digest=sha256:f4568499f0f76f9102f6c0415a1a8ade6e9c0911ff656bb471a28cdc5a08201b

Observation 92e400c0-ffa5-4f75-ac26-451ca4a7d5f8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.387129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.387129Z digest=sha256:16376c42c3850da6cc6caf6de3da41d7fd4c836051189547ec93c41a292710a6

Observation 69f4681f-c7bd-4a2d-b978-8acdd833a70a · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.446461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.446461Z digest=sha256:ec7adbe185fc1f7e3a96d06a8fba1e47ac9d6aa8afd030b4d63433aa8f7452b7

Observation b3b21a56-e1ad-4778-a6ce-c322c3a42042 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.507321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.507321Z digest=sha256:501a237f69830202c0ccf4126e24bc4cb2c56f0eb0f218c2a27bc60d109e3a40

Observation b25ca4dc-5cdc-466e-a600-f35ebdf1f7d5 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.567659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.567659Z digest=sha256:f3ee610d64191864408440052f955c509e57a78e9fd152f22522de8120548f4f

Observation 55d743eb-1d52-4b25-92e5-2c110b521247 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.677891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.677891Z digest=sha256:e80969615a753807a24a3439017ba35cb2fc9a64755b9a468eb6f049a912f3e5

Observation c0fe3f2f-7ee5-458f-9016-4aeb9d543755 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Task-Aware KV Compression For Cost-Effective Long Video Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.709295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.709295Z digest=sha256:cec819d66d66ac07cc3781c27be6e39dafdf8df315b8b63989cf696b1bee777f

Observation be7b36a6-61cb-4357-be28-6e974b9a6c70 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.757096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.757096Z digest=sha256:80033bdcf66118abb6ef6001b3e1ef50b354666feefad2898e7ab3f882b75d9f

Observation aad14b21-a422-405e-9593-c7614134390f · outbound

This paper cites Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.823039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.823039Z digest=sha256:1ae0d79963b2a4bfe44db71891898273a2638384ffc641fd2aa67249a1a0f348

Observation b69461d4-c78c-4c67-801a-af0204f757d2 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Internvideo2: Scaling foundation models for multimodal video understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.922661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.922661Z digest=sha256:b57cd22af4dc8f22caf6bd887254720978cfc777982c3f44f77ea89de8bfb775

Observation 333a114b-7802-40da-aee7-bd656b079e76 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.968114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.968114Z digest=sha256:2974dfbcd04edd415d38dff6fffd679864e174e8f5e7b6849716642e18cc8e7b

Observation dc1e9d7c-528b-49e6-9fda-5a5a44fc868c · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.027331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.027331Z digest=sha256:7fb9e926960ef0ed62ec449687f070d56071922df7a6f87b611b2da698fecb17

Observation 6124ce6e-9411-4742-a532-b56e838e24cf · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.126615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.126615Z digest=sha256:851a379b5ee1b1c543f7c11990b99808ddc7f6fedd8dc783ac935ba553d08ce9

Observation d34835d6-5ad3-46c6-b379-d68751029d8a · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

Task-Aware KV Compression For Cost-Effective Long Video Understanding InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.216006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.216006Z digest=sha256:170685df65881894390a578b1bdaed0d2c683266f888658b1c12bd4eebbae7c8

Observation 68ed76ec-112b-4f43-92a6-a1c7b91af129 · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.267392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.267392Z digest=sha256:d54d34a79762eeeeb57d4b8e2459ff367b8c1da94ff0bfdef1622942e05bf97d

Observation a48fd399-8528-405c-9c19-0202d1e113de · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.329386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.329386Z digest=sha256:70ea58130d6e49bf695c92510a1def201fa53c28231d73c1b8556afdb3be9d95

Observation 122a666a-3a86-4f3c-a925-119f5f0a1358 · outbound

This paper cites Long Context Transfer from Language to Vision.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.426259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.426259Z digest=sha256:6c3f2a4b5635a37db4ce58654245bc92170c88ce18ddd1d5992c6c22fd7b0092

Observation 2ddb7bf0-c48a-4da6-8128-73469e3d4e19 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Llava-next: A strong zero-shot video understanding model, April 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.503476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.503476Z digest=sha256:0d5091aa47cc86a8dd7c4c6bd051f596e18796be44e56f7aa3b992eba2eeddf0

Observation 5360d085-0391-455f-b137-d4d1dc11a146 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.551825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.551825Z digest=sha256:a540b8ba06909960d5dbe100a8883e7e5a4d6982a43b1787a5219c6749f09210

Observation 52688802-2f3f-46f2-8c31-bdd3aad996d9 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.601082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.601082Z digest=sha256:785abd64f592c133dae2b52129179d6222b8f2bb5323a14a114b90dcbdf33e0e

Observation 7ef244d8-c636-4f1a-8998-377c3e8507fa · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.692207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.692207Z digest=sha256:2074e3dce893768e790bd251ce6244c2676099663c6ec81da2c5b0686cdbb353

Observation 03bdf067-a8a2-409b-ace8-51dffb6b97b6 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.772187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.772187Z digest=sha256:e84feb2533114dbb9bd6c668fa670503c8fcfaa67faa64dce333589e45b58244

Observation 0e76717b-f6d5-4173-ba86-00d49d6b5f0d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.829071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.829071Z digest=sha256:3f770c7452edf54ea951093b2caa233e952b2cf3af6cd40d88d8c584aa2aeca4

Observation 349009e5-9f3b-4f1c-9ff6-80ab18e6d2ae · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.894521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.894521Z digest=sha256:6331af12fc317e08d474d1d58109546b5788469908f35bab8b45ecc45c31e48e

Pith citing papers

No inbound Pith citation observations are available.