Pith. sign in

Paper Citation Record · LEDGER

VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2403.11481.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.11481 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.422447Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T20:58:45.440572Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8e46ba88-1b9a-4c66-8fb1-33fd76ae6a0a · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.400997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:91dfbbf6cf9646fbefc1cde4c69b1bae5ec081282878142d689647c671fafbd4

Observation 1eca1a7e-e2ae-4794-9fb8-a2cf4f1428ac · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.422447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.422447Z digest=sha256:0ba5e7c9562b7604224cd7a1b79f4ff07f7dcf5508e774279a5b155d35022e70

Observation c61625f3-b6a4-4f2a-888a-2195dd747f97 · inbound

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 cites this paper.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:56.536060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:56.536060Z digest=sha256:970606b82a4735d4f53c905c57e7ca640f805fe7ad4b210208acef0de2fb8abf

Observation 11cde63a-8d9e-4665-aecd-aac9877d3456 · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 265

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.442734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:8efae698897f48d333e083884aefe0cabdebc770b4b85087ff837ac8b318e29b

Observation 13f78976-f146-4f2e-acd1-a6bf30837265 · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.198959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.198959Z digest=sha256:9447100a72e369ef866e084a32092f28511d8903831628040edd8b79b41320c7

Observation eaee1a83-20c3-4af4-ab39-0dad63a6d941 · inbound

Towards Sparse Video Understanding and Reasoning cites this paper.

Towards Sparse Video Understanding and Reasoning VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:09.537352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:09.537352Z digest=sha256:2f35544440553eef2657951551d6916501ee94d7d283aa3ff93a09038017bd27

Observation 3cc202b1-1d58-4256-b5b2-f7658fa03e0f · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:2c75b5a3cf6d4e939cb327ecf990c103e272d63fecf9fca7a37a9d578f878db8

Observation 908d5d2a-9124-46fb-90df-20662016a6a9 · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:49.142583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:14fc97a80ab50b50cb2a0feb6e1f200790d9bfbff8b00e9d8642db6c47d71c41

Observation 4c8a43a9-a33f-41e1-bdc1-6dc4586347d7 · inbound

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models cites this paper.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.659733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:a207183f3df561e4bc71a750bb5444d4ac06830f29ee59f58a67f023e1fa169b