Pith. sign in

Paper Citation Record · LEDGER

Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2409.14485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.14485 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:43.251039Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.682244Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 60d92558-0dc5-40b6-b968-b9685ac8c8b2 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.433404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:090f1bb2a4d1025c28e54177e76aeacaebd1e365722f742f09dbf6ac75f73a1c

Observation 1cf91e4a-058b-46e0-825a-d4ae68ca3742 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:42.987016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:266a35197dad455add68ce6cfed4f3338e11f5505e3f4535f60f612dd0987122

Observation 731ff6fb-eb7b-48c7-b925-bcf2ff8d0414 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.503594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:af41057ca223dc4cd6272a147dabc0aefac3b76ecfd33241d316b3893f8b9d37

Observation 8467a30b-0658-47f3-8e17-081184e9bdde · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.770114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:91dc76369bc4dc67f186ecf6b03bd523f6c735ebe6d3396fb036bf0c52db3cff

Observation 5523b6de-9a63-4e83-8221-c14dd59cc6bc · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.251039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.251039Z digest=sha256:3de4e5aa223864533b8d68d2649d4f0b85ea58c079c61f6b54a42e06c6d06c26

Observation 8d9c289e-2d5a-4082-a8ec-fcc72a4a1192 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.584507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.584507Z digest=sha256:bd31dc22658ad2679d56bc3c50f48702459513db26b6f16c75089507ca62c6ed

Observation f19b5cbc-5429-425a-8010-6d84e9d4ed61 · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.806587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.806587Z digest=sha256:b46eb4ee68d81fff9d7dd7c024b5b2db77cffa952ff4d8cbcb0a015afe6cd23f

Observation 358355fd-9998-444c-aae7-b87c1418de3a · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.637236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.637236Z digest=sha256:66c8aea1750be9b1a440324d8cad771d15eb402215e643b0567701328f72a11b

Observation c8017646-7848-455d-855d-be4096fd1aca · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:04.246155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:04.246155Z digest=sha256:d9f34da0a7c3db4635dea456a5ff642b5e11f6a60da6801f611e2e182ee32232

Observation e951b058-651e-41cb-82f6-f9ca968c09dd · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.226160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.226160Z digest=sha256:d7d886cae204c72a7dfdbbfee82dfa790537f58fac00a9fb1454942d1872f20f

Observation b3b21a56-e1ad-4778-a6ce-c322c3a42042 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.507321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.507321Z digest=sha256:501a237f69830202c0ccf4126e24bc4cb2c56f0eb0f218c2a27bc60d109e3a40

Observation 51b1b8f9-be18-4d35-bc41-a649916768d0 · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.755141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.755141Z digest=sha256:168038c62d30b8936bfd1cf0eacf87d7546cc8d3e649bd795a65cf69591bfb89

Observation bee69751-5848-4f41-9434-e3ee6d7c25b8 · inbound

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding cites this paper.

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:00:07.069669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:00:07.069669Z digest=sha256:ef3b6ab88e570147ae953920d66431db935e969698ddabd6fe4251f39520f684

Observation 402060c6-11ee-452e-afb6-8d77ad4cb01e · inbound

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory cites this paper.

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:02.472353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:44:02.472353Z digest=sha256:45f3c304937b7338bb65aad5971c54e75fc1269f09b5e02dee2a1aeee7c21e35

Observation 11789708-fe6c-41c0-b5a3-1c112e649775 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.223612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.223612Z digest=sha256:1f8fb9150cbbf62e81b2214168f47e19de4a2eea4441e4ee9463f30733028a73

Observation a03ecc67-4006-4690-9bf6-bdae02f2ad6f · inbound

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding cites this paper.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.899236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.899236Z digest=sha256:d7a1c5dbdd2a965ac0fb282b12f64c467104bd43bc0909f609522947c880942f

Observation c2df738d-2eb6-4bef-b7fd-a788422984d2 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.296306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.296306Z digest=sha256:3bf44d118a916e86401405413b380bed12d382ba31993d6b0eee8b2a6f63fff6

Observation e6347225-408e-4c8d-a637-f6576a7f0245 · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.402521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.402521Z digest=sha256:1278b398de1a5e07b49663a40abe4bb9c7c506634851cfadc21bdc2ee9848fc1

Observation dab9472d-2449-4130-acf1-d9b6d776d216 · inbound

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding cites this paper.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.738153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.738153Z digest=sha256:92bbfd4d470130036b4b07456be56554dd21ca219fdb4065e919ded7d826ce47

Observation e27b375b-5317-4540-b42d-d6fe98846f67 · inbound

VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding cites this paper.

VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:46:19.505318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:46:19.505318Z digest=sha256:6cb2f430168fd56353a06cdb0da5006ad6c95bb7e1de42ccd2b510b0aac8fba6

Observation b7efc35b-e33d-4610-9fea-8e2297ab618c · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.818052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.818052Z digest=sha256:a98593f5ec7269f0f55f8fa46e1852ecf7ed4ca14985d7dd303ff18e8095c665

Observation b0b2081b-5f53-4d21-8b42-08d2cc83303b · inbound

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding cites this paper.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:29.150095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:29.150095Z digest=sha256:e2c1550fdf4d10ee51c888eba4b1e39fcdcdce4e775a03ec1f24643aa08ab219

Observation be7f6567-0ae3-4f92-81cc-1cc6f5346e46 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.322261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:bf81882fb848c5bd78554ae15b8dad0fadd3459eb4217e85f1e26f9ee825aa00

Observation 32f16455-d6c1-4bea-9e65-a38aff7fcc5c · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:04.170343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:bd79fa320fcf9eed3aeb68c72d84f71c7c0af8692bdfc60aa814762e26523644

Observation eb07d3ab-fd95-49e8-babd-079361566bca · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.608836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:c4fa8844d8171acab45fbcf329ef1a5ab350b3f52dc5fadc3633b6a334c01ec3

Observation 396f883b-4be0-4ea3-8768-009ca23e2b43 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.379342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:c6cb151ebcb9d0e442b9572e19b793343fa55a677d06d764580e9e257bb3c9a8

Observation 6dc7a3a2-5fda-4960-a6a6-90e408212562 · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 111

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.401036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:f9b3812559995aefe1837634bac805f1caaedcd120f0dd51cf3752e426e80e4a

Observation a4d5c3f9-8265-4d0c-8731-2cf740e7b81c · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 253

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.928923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:ea8275ee57b42a7481da0a588edefcbdb7ad2e2e98122d2cd0f9c55dc6682009

Observation 73de4bda-be24-4798-8bc7-9b58a48d0114 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 139

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.683787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:f01e0e6bfd25ac7168da28f682e1e303c0e4aeba2aef5b4711fce92582e07a5b

Observation 7dc63e75-60e7-488f-827d-82080ef78b28 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.608318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.608318Z digest=sha256:7c8adbc87b6a43b2a32f5170bebd9962414c2b3d5084423d239f58840fe3a62f