Pith. sign in

Paper Citation Record · LEDGER

From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2409.18938.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.18938 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:50:43.987304Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T16:41:37.883000Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e6a75d1a-adf9-44d3-b2a1-a4922ef02ac5 · inbound

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs cites this paper.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.845769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.845769Z digest=sha256:3442f6f9f7b8dc5276a94f4a0bfe0c0a687bbf89d21792883af3c66181d97235

Observation 9f1186a5-ad1c-4a60-93fb-6907e9032452 · inbound

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding cites this paper.

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.387574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.387574Z digest=sha256:d10b475a46c3c1d240935df6aa8a06f59a647cec9451452caf3c49f94d59c66f

Observation db1551b9-ffdd-478d-b7fb-7678a1a290c6 · inbound

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes cites this paper.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.987304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.987304Z digest=sha256:1c2fd9b4bcef84636444ffda623a283287c03c195aa30f315447eb402fd96ac4

Observation be76c2b1-c408-4058-bf43-522ab05a4069 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.602440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.602440Z digest=sha256:f3430bc492311ef286e9a5019260819e1d78c71dd6977a2606d0883bc56f048c

Observation ff6bf875-6dfb-403e-aa65-9bd3504869d2 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.919272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.919272Z digest=sha256:d4ad278ea052db6a7b8f09e1ce5a855e55d22edf5f76bcb5a60b03489fb7a732

Observation ea62f121-061b-4707-9caa-a1925014c700 · inbound

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? cites this paper.

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:41:37.886557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T16:40:16.630602Z digest=sha256:b58696aa72f4e238225a17c0261240772ad6bdaa8cf73d3dea9bc316e3dd7a86

Observation 8e778e17-56dc-4dc4-bba4-d48ba65db18d · inbound

NeMo: Needle in a Montage for Video-Language Understanding cites this paper.

NeMo: Needle in a Montage for Video-Language Understanding From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:28.374737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:28.374737Z digest=sha256:664757d8fef069cd965ebbf8d2217ce7ea5f8f5b0956c64de9727565f271547b

Observation e29c6fd0-4f3e-4ee1-848f-d899feec41dc · inbound

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs cites this paper.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.792970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.792970Z digest=sha256:c4d66cb5637c243f97bba05ca33c394529832c77e74346015493042e87ef6d33

Observation a0514b3d-a8e3-4c6c-868a-06a09409553a · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.464586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:562df7771d85d7df327581db85765a98cb65cddd8170ae31bced3d78be791e2b

Observation 8f04618e-5571-46ee-b2e7-ef91a1f20067 · inbound

OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling cites this paper.

OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:30:16.509688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T19:29:55.075825Z digest=sha256:ce4410969bd62808013ce9b2685a8a9d88500f39c3683d30fa24b47009ff7d7c

Observation 31ca61a2-8b6d-458e-8126-1548b3b032be · inbound

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios cites this paper.

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.297229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T10:15:15.129358Z digest=sha256:7cf044763c5150678fe960d7a25eb43d719f261fb8c98fa96df69ead53dbfd3b

Observation 361ab0a2-12b2-4b95-b360-4dc8aaf572d5 · inbound

Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning cites this paper.

Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T09:31:22.762570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:31:22.762570Z digest=sha256:fb7cd49633a7475b5626ba4c86bed3e2b9fef12709cb5b95476d58284fefdcb8

Observation 9e709361-2ea4-4002-9f1d-ab988ae2cf75 · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:0814f16251b3d1a1f6adefeb0df5fb277a29eeb98e0c84393fb0a907b3166f70

Observation 5f7ffbed-55c0-485d-86de-2473a4e3220b · inbound

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding cites this paper.

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:41:18.339565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:41:18.339565Z digest=sha256:9b81e33f0bd6926528f37175ff619455c66d9b3939bdbeb2fcff9f6a72f6adfe

Observation 476af6ef-3f96-4c9b-9408-e552cf9c8741 · inbound

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video cites this paper.

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:22:31.850547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:22:31.850547Z digest=sha256:03999b34fd82df9523217c1a330769c165ebef406bb43e6384e2615bd5a032e7