Pith. sign in

Paper Citation Record · LEDGER

DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2401.08392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.08392 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:50:21.415193Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.774204Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c8cbdc7d-43b5-4171-a8d2-750e90176b8d · inbound

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering cites this paper.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.610768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.610768Z digest=sha256:e67f42610b913fe5e871ec22ca175258855a710aa6980579becf3674203dbe0a

Observation 082b019d-a421-4657-8511-afcfe992b43e · inbound

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding cites this paper.

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:21.415193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:21.415193Z digest=sha256:10e9725c84182c6da26da63869f771f7fdbebc655e176a57a7095a6545afc377

Observation f292242d-640a-437c-86e3-790d00c30bbe · inbound

AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction cites this paper.

AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T18:51:52.995364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:51:52.995364Z digest=sha256:5688d904ec5b4bd7dc8bf3e0678cbd6e0ab908aa8bcfadb7d1939b8eee021207

Observation f4c62be1-79ee-47ba-8f63-e7094acf4d73 · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:43.846351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:43.846351Z digest=sha256:acf8652f16895db267f93b821c921bf95ac4a18d028144452e8468c445197d12

Observation 0bf1b2bd-85a8-40ca-a03c-b01c871af1b0 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 268

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:09.648164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:09.648164Z digest=sha256:51f917d6eed69c5a92389e750ded283e7114c2ec185f80751ac03d3846ec50c9

Observation 53bad080-355c-4aa5-af1b-41d431437a13 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.928527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.928527Z digest=sha256:642d36368f76cf59e8d2e443afcee2aee124935da71c4531633d7960a3b7f73c

Observation d20c6347-fe73-4431-8d86-cc57c0648904 · inbound

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents cites this paper.

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.597271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-18T12:40:19.544260Z digest=sha256:cd1f43a3c46593ea6529a24ac0c87a40744c5c6693567cd883e8481691805cb6

Observation dbe15880-f783-41a7-a4e9-942021e05449 · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:17:51.837626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:2dd5b1b9af216687c38f330177f567f9fba25421140dfdff6961d99b35579b99

Observation f4dfe0ad-4679-4848-a92d-a0a34013d476 · inbound

A Domain-Specific Language for LLM-Driven Trigger Generation in Multimodal Data Collection cites this paper.

A Domain-Specific Language for LLM-Driven Trigger Generation in Multimodal Data Collection DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:49:58.941893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T11:45:32.281685Z digest=sha256:775e1b58f72911e84648da1eaeec1ff4468127ea5b4a20d53b44f5c85eee5b20

Observation 2487c2a2-dcdf-4047-9b39-392c67cee788 · inbound

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration cites this paper.

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:26:02.209980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T21:55:35.699057Z digest=sha256:58e0e890d60984de65e20b0cbacc5590280a9829a0be7b4b689aaa62aac4f120

Observation d5562b21-57cd-4344-b9c2-f6584486ec0e · inbound

ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding cites this paper.

ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:26.470770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:08:13.214803Z digest=sha256:99752dc92dec74080b4587de3e42a50969f08f8fe54fcf02586e361ce336b88b

Observation 6a3faa29-a56e-42ab-8467-fef25f847e09 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.775729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:762b5dd4df15eb7d534ae9945f8efda72db1eb9ad71ea138b2c13262344ef91a