Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2401.08392.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:50:21.415193Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T17:27:15.774204Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c8cbdc7d-43b5-4171-a8d2-750e90176b8d · inbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 082b019d-a421-4657-8511-afcfe992b43e · inbound
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f292242d-640a-437c-86e3-790d00c30bbe · inbound
AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4c62be1-79ee-47ba-8f63-e7094acf4d73 · inbound
Augmented Vision-Language Models: A Systematic Review DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 123
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bf1b2bd-85a8-40ca-a03c-b01c871af1b0 · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 268
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53bad080-355c-4aa5-af1b-41d431437a13 · inbound
AdsQA: Towards Advertisement Video Understanding DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d20c6347-fe73-4431-8d86-cc57c0648904 · inbound
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dbe15880-f783-41a7-a4e9-942021e05449 · inbound
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f4dfe0ad-4679-4848-a92d-a0a34013d476 · inbound
A Domain-Specific Language for LLM-Driven Trigger Generation in Multimodal Data Collection DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2487c2a2-dcdf-4047-9b39-392c67cee788 · inbound
HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d5562b21-57cd-4344-b9c2-f6584486ec0e · inbound
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6a3faa29-a56e-42ab-8467-fef25f847e09 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 177
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.