Pith. sign in

Paper Citation Record · LEDGER

Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2408.14023.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.14023 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:31:16.907077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.765108Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8bbff39-2518-420d-8b3f-584957ae697c · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.536953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:807ff89c65b898276ca41ecdcb7fc11e36f5bcabae2bbaae9f06bfc3e0e00f19

Observation 8ce1ff7b-9ec1-483c-9524-5056486179ab · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.501093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:2aa9ca3f4ec9b2088e2ab7e8c433924044943828f03bfebd67e61f9a1d9453bb

Observation 23d85641-bcfe-4bb7-b319-6c101585e709 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.716663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b4335e7f4d920226e0ce268fb84e8bb9eef5660e11e3cfaddc318753a7d6816d

Observation a1e3ce7b-6cf8-4833-9414-91be823de78e · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.714773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:91af34f4400f6ddd26a7280ebefbe7d6d5f9e14673460b4b1359e4f3aece7cba

Observation fe4d3805-a816-4272-b56a-86137f4f33ec · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.255829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:66b68a3dc1afa77a3f81e63b9036becc6fc0228f8ddd5f9c8368eb9ae4f465aa

Observation e3cfbcff-8b4d-487f-a123-e2ae1da5721e · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.907077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.907077Z digest=sha256:6c93c68eb4456a3d74c78ebf1348a2a71cdde266bbd2c5d7ac15bef90df933b9

Observation ce17acaf-da91-44b8-a33a-65f9f5f53e0c · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.981605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.981605Z digest=sha256:39a6d086de03298ee315b5d5906b6cbe083ff4c17e575bc157bc4ed7cdf478e1

Observation 4c265e10-9dcf-4610-9ba4-825985f4bb4d · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.775593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.775593Z digest=sha256:ff4760772cf43b339eb787d3b5705c55cfc5c1f70b09943e32db34b3d1e01613

Observation 889779c6-2937-4cab-bb15-295b9ff8aeea · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.421786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.421786Z digest=sha256:da6c8f6c7a199a165b7e1bdf14326b6f639085738a227c63fce1d8e068d7b415

Observation 97b1818a-42de-4fdb-820d-985568e79bc4 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:51.431837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:51.431837Z digest=sha256:b14469885c0f0878c1b3f3eea7f2f0187b9258fd3ffa754dd499a29475a9e690

Observation 7952014f-ff33-40b3-bd24-7d4a7bfe2962 · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.491781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.491781Z digest=sha256:d0d64de069b9cd22f7c8facf9acf96d679f1b7aeef2424db8d2df697aa987162

Observation d632ca7f-a4c0-4de8-9889-4d3d58abaa42 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.989110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.989110Z digest=sha256:57fba7f69d7310b7dcd40e53052dd4666f7c577fa7d00aaa66af608cc34f0cf3

Observation 29079ea3-7f5a-45e3-b673-b96d566aed81 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.868717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.868717Z digest=sha256:e946d70d846db99d87c8ccf0a6cc6c14abf0efef9e25557e63145d880e87ecdc

Observation 55a5ff53-44d3-4273-8475-82950400326c · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.340057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.340057Z digest=sha256:9966026075b5f1499d6ea5e6f69471aca0f8d26fee7881b9b2b2cf2705c8cb81

Observation bb93b653-36b1-4e29-a744-433a9651ab04 · inbound

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory cites this paper.

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:00.121528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:44:00.121528Z digest=sha256:c9584246c34c55817a65283082c1d313f3bb2768948f35de8997beb71bc82f96

Observation d8855dac-cc08-4c80-8e64-ec425fe59203 · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.951565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.951565Z digest=sha256:5de3e169218de39db78dd0c66f1b1636c61712717986efbded8439785bd66e8b

Observation 444ed85b-6e62-4477-8c1d-a427b44b15fa · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.727170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.727170Z digest=sha256:e4635c987f0b5c3a5347163777422ecf7af872839639abede21acc2e32a55d5f

Observation f875cfc8-e833-497e-87d3-73a3ddd9fb79 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.524954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:2ce417910089ae57ddbf2fe2ac268b41257e946a69d021f2e2c5c2833f148f8f

Observation ed57979b-9d67-4aca-9efa-d64112ba6835 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.419248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:8c88405392eb960f27ea43e0a9259df3d2010837ca9275e3b18af43a5af53ae4

Observation 1c61831f-6a19-4f9f-a654-df1feb208d45 · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:04.334732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:08bfc0790cbb6cdfc362c85c50c7820f60b51312462a2b32ba8edcf2e554cc49

Observation f2f9886c-ae29-49aa-9743-76d23ee107a3 · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.648482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:b0e85cd854a382a3fe9b344ee7ac752a3fa120a360a969f6b4f5886507d93020

Observation fdd603bd-a208-4c6a-bfda-351152e3382f · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.039662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T14:48:39.444933Z digest=sha256:2e6bfe2bd2cb434f3f330c458fa977326920a8ce2a3bf11236d7bbd57354f7e9

Observation e0758a7f-de1a-4979-95b9-fb764a784d0b · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:56.556078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:57:42.822121Z digest=sha256:cf562cb213d0cfbd1d942ef818dd3ff4b1323da0dee482edcc2f9422865a7d1b

Observation f8cf6d9b-4edd-435d-8e9e-65ffefa10bf4 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:20:56.454668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:1e47ae06bf92a1931a01c1bc64cfe6f41cc65bb7b17dfefc98ee946eb868a87a

Observation 0a27ff8e-79be-4bdd-9476-ea62c810ea37 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.707528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:83be7f5f7edefb1513424a83ee49eff2e31a7044fef29f9a332daca0986a46d4

Observation 353b1242-ee1c-4f1f-8d1c-6fcd19ab717f · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.438058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:8cd1cda48eb5aac90fa251ea494aa3317a645e0a19507450ba83990c1df08ccb

Observation 4fd44adf-2d77-4b56-b812-061e54bb2403 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:48:14.924683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:170094a6434acb81f669e72ac1964936574f2b18ac4a8b824974b1777d58e2ac

Observation 60c46325-5641-4837-a59b-d9537c6f88d3 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.748177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:2954844c46c2918236326bc98f22018d2098dde2aecf3b18f3b006583c6767c0

Observation 22a9fa5d-4c61-43d1-8f22-ee450655e785 · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:08.028902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:803256bd173b1b5618e6d168b958a595115ce09dac7a5f2684f4a903020e8c3b

Observation 433ae2e0-4eaa-49c3-a4f7-36f6b24270f2 · inbound

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models cites this paper.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.202123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:f934f92d5ae0f959e7421d0554d185171c42607e2004de696a9cd4a4a1a2c79d

Observation 172a8c6a-851a-49fb-8c07-0e34cf02566d · inbound

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining cites this paper.

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:02:50.352612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:16:42.001361Z digest=sha256:caec429f97cbcd3a3b98cd72c01cad7bcb44ab356d07cb783b6bd07a60b00938

Observation a60abe6e-cf91-4903-b0b6-abe486c2ae41 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.358625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:66b7a021bbc92a19d6ad99504abe734ea0421b96f26cf7f65e59a233fefddaab

Observation b06a113c-fe5c-436d-904e-14032ee29aaa · inbound

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding cites this paper.

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.056825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:04:29.739632Z digest=sha256:6d5a72fd7a76b81fb30842d82cf51ca3cfe7e473c6c401f81fa5dc8cb2cfbf3d

Observation 612534ff-5173-43db-934e-1bda7fb6300c · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 249

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.974132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:eae66e2fd0a7dcf090f8a45d9dd245c0f341a3a1bebd68929bb30c5945c2a4fb

Observation 1cefe5d0-9b92-4e34-afdc-c43f7e417913 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.452455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:3212ea2b3d4f3341a9f0bf806bcee8d417fb1d8d4752b939390423c74350fe45

Observation dbaf52bc-d04f-4e4f-a079-f604648b05b8 · inbound

CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams cites this paper.

CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:39:46.766518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T09:42:58.861599Z digest=sha256:c900fb361a38d35d9ffd3772ea1c694450be459f9870a2fbddbaae8358211812

Observation 8611e1e8-588d-4426-9146-fe6284f1944d · inbound

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs cites this paper.

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:21.335724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:25:38.593423Z digest=sha256:9388c7c2669db5365e43996434cd36f7aa120a4a26151dd1794eeba1983a461c

Observation 440a1036-f026-47e1-8c1a-cf37f656f759 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:fc4690dc7dedb4ca2eeb704c04481b5c69aae40a137689ec97d2ce4780eab761

Observation 2bbef040-3c0b-47ca-98da-46922382d900 · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.181503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.181503Z digest=sha256:d5fb7e114822f4d365669b724b71674b6fde0cbe02e98e4ecac7d1e684ff0ba4

Observation dc306c73-3c39-4073-9433-30caa848ad63 · inbound

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding cites this paper.

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:00.541193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:00.541193Z digest=sha256:cd67e9c79e36c89efc9e208e3315eb2ba8d28a54a588546d6ea6b98f4095a5ef

Observation 2adaf088-b632-4f60-a978-61943e5b8a5f · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.215228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.215228Z digest=sha256:0bf3cc9eb09f2afbe8a021caa4ddbe682b031eb8a567caa7ad03803a36ad9c72