Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2403.14622.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:50:21.390322Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T10:48:03.032393Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 203814b7-9f3a-4acc-9ebb-9953a9e5bd9f · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Language Repository for Long Video Understanding
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 129af2f7-9a6b-4b7e-bc82-8b43005c2458 · inbound
On the Consistency of Video Large Language Models in Temporal Comprehension Language Repository for Long Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776b7648-a119-449b-8d8f-2d49fa3dc781 · inbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Language Repository for Long Video Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e49bfb55-3144-4a2f-b33b-59b8d134b026 · inbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Language Repository for Long Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9a8cb9-4666-4f03-af54-b49098b2ee9b · inbound
VidCtx: Context-aware Video Question Answering with Image Models Language Repository for Long Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1416b0-dba8-4eb8-a22a-a93777d61210 · inbound
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment Language Repository for Long Video Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc972c75-5884-407d-9b57-524c466e2845 · inbound
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process Language Repository for Long Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 690f6774-ef24-4d7d-a30c-aeba5e17c8c8 · inbound
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding Language Repository for Long Video Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9fd451f-0f61-4df7-9b37-e04adb55aa37 · inbound
ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning Language Repository for Long Video Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0731a994-2002-4fc0-9026-e61e684b35e5 · inbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Language Repository for Long Video Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378dba24-81bd-4297-9f4f-974a00cd3418 · inbound
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Language Repository for Long Video Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5da05aa-da0b-4824-b946-801df598799f · inbound
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering Language Repository for Long Video Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02bd78e0-fb21-4a3a-aa91-739025b936dc · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Language Repository for Long Video Understanding
Reference 300
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93e3cbc7-3329-48b7-8fe3-9b30deb89d17 · inbound
Towards Sparse Video Understanding and Reasoning Language Repository for Long Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae420d3-4fca-4572-a022-1281112ef2c7 · inbound
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding Language Repository for Long Video Understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0b07c839-57e8-4a8f-ad5a-13c7299fd366 · inbound
Why Do Vision Language Models Struggle To Recognize Human Emotions? Language Repository for Long Video Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d689b70b-4670-4ce6-9430-3a91d56b3fd9 · inbound
Why Do Vision Language Models Struggle To Recognize Human Emotions? Language Repository for Long Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd8ee88-1ac9-4f25-8b22-94e9d820cc0c · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Language Repository for Long Video Understanding
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.