Pith. sign in

Paper Citation Record · LEDGER

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.21891.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21891 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:20:13.504824Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c61625f3-b6a4-4f2a-888a-2195dd747f97 · outbound

This paper cites VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:56.536060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:56.536060Z digest=sha256:970606b82a4735d4f53c905c57e7ca640f805fe7ad4b210208acef0de2fb8abf

Observation 2a8ba645-1d35-41e4-89f6-481e20ca63be · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.499956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.499956Z digest=sha256:f45181e3a1b87b11a021abcb7abb85ae81d1a0df1bb1f464776bd18e14c760d0

Observation 9cde9d80-fc8e-42e3-8db5-5d172aee904d · outbound

This paper cites VDMA: Video Question Answering with Dynamically Generated Multi-Agents.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VDMA: Video Question Answering with Dynamically Generated Multi-Agents

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:20:13.720355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:12.669954Z digest=sha256:24fb98f0fddaf8c5514652db191d436900e1200ae5c0c02d976f532e1a6ebce1

Observation bc3b5004-dfc5-4e02-a6dc-af4e8fc96a92 · outbound

This paper cites VideoMultiAgents: A Multi-Agent Framework for Video Question Answering.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.739322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.739322Z digest=sha256:ca5bd1a9d0669743bb28bd3f47575fe463d3318ebaf8d4c525908120ffcc6d9e

Observation fb16a04b-c044-4c56-a2d8-b371f7fb96a6 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:15.304002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:12.784643Z digest=sha256:7e9c550038b8aa933c37780c0f4bac055872da9f2fbfd7b7b0968c4761e15432

Observation a087af98-aff9-463e-8a79-5b22ead6f354 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Llama-vid: An image is worth 2 tokens in large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:15.067116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:12.823501Z digest=sha256:449b040081a9a3c60d80c6cdbf645b510c53ad739b6412cccecb4def1d8294fd

Observation 7e5ea4b5-c1c8-4bc9-b052-589030d25ee9 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Video-llava: Learning united visual representation by alignment before projection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.893626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:12.861049Z digest=sha256:0b9a7471ce9a27cfc3d1a51d29df95b83d1dbfeddfd601ef4cd9f7bcc19fd30a

Observation a4b88318-1e65-4027-b18d-84a66ec415d2 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.887325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.887325Z digest=sha256:07755a6ad8549882e07acf016628c3e1bee83788a3bd7d920d6376d3cc265822

Observation 2d71e20c-afd2-494e-9538-9cc555bf9bea · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.709409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:12.904269Z digest=sha256:2d5dd7aeb1be9821e2e16d44c47f6f3ddc8c5646799672f65c94eba98eab626f

Observation d5b3873f-d0f6-4d18-a067-566e4192ea4c · outbound

This paper cites Introducing deep research.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Introducing deep research

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.561944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:12.925037Z digest=sha256:2d9a90aa2d5f97961d24a5d87970335af068ef7ad621b8b59848df440942e0cd

Observation c76d8af5-6ecf-405e-8499-77ad6f20ccc0 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.971888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.971888Z digest=sha256:ce76b91b2ee86066cf08ac9b784a575d801a323040fdd07385708bd75e8cd166

Observation aaa0365b-6b29-41ec-be0a-3520fe7119fc · outbound

This paper cites Traveler: A modular multi-lmm agent framework for video question-answering.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Traveler: A modular multi-lmm agent framework for video question-answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.338089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:13.023972Z digest=sha256:1ab77739523227a6eb6a63cfdcf7b5c4d6a5994813f1696d169860507abf063b

Observation 65570e96-d72a-4e44-9caa-adc8fba857e2 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Videoagent: Long-form video understanding with large language model as agent

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.180597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:13.078536Z digest=sha256:332be11db82bf33c3c5de916e7def8b98f59b3a16607b50bee8190a316550a70

Observation a2fe6278-5c86-4b72-9064-55b7470acee5 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.123695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.123695Z digest=sha256:38564ee3524a88ef901cba739ebbb4efcb1da73899dacd7096cad050b13993bd

Observation 697143f2-d66c-4bd1-8ec6-42d50ff53d0d · outbound

This paper cites Lifelongmemory: Leveraging llms for answering queries in long-form egocentric videos.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Lifelongmemory: Leveraging llms for answering queries in long-form egocentric videos

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.068571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:13.161396Z digest=sha256:91befd6b8c5331c3cd719d566e4a2699af1a96290a7741e592307180af2930ca

Observation 283646a8-8b70-457c-b779-f6eb04f50ef7 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.236841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.236841Z digest=sha256:7da6a9bd92da47a5a3fdb4afc57e325a7db7ce3e9a3b07574e440a38db000ce7

Observation ec97086f-91a3-4ef2-95fd-5df53212c9d5 · outbound

This paper cites Zero-Shot Video Question Answering via Frozen Bidirectional Language Models.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Zero-Shot Video Question Answering via Frozen Bidirectional Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.284415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.284415Z digest=sha256:1c269be80c5f66e7c13b2afd0b66d7722068c60930dd89288db36e26313e5b64

Observation b8a684b2-0020-46df-8aab-ea856fc2d9bb · outbound

This paper cites A simple llm framework for long-range video question-answering.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 A simple llm framework for long-range video question-answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:13.971016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:13.339297Z digest=sha256:6cccc89beffdcc058ae862430fc300b4e46426a4f057d93e1485ada1797f00bd

Observation 34503728-fbe3-4366-9fa0-4d0be677738a · outbound

This paper cites Hcqa @ ego4d egoschema challenge.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Hcqa @ ego4d egoschema challenge

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:13.870767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:20:13.389497Z digest=sha256:f887a46bbb84f03bf6a4b1ba0d43b5ea56867a6bac1db783570627203651c867

Observation 1bfc5510-31fb-46b2-a18e-27f3ea93819f · outbound

This paper cites VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.504824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.504824Z digest=sha256:3bdd36946c27fb7adade12e970ec589f1c9b0c73c5bf932a188ae4025f0427ba

Observation 573b2a95-b937-4130-a765-1ad2530963c4 · outbound

This paper cites HCQA @ Ego4D EgoSchema Challenge 2024.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 HCQA @ Ego4D EgoSchema Challenge 2024

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.460205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.460205Z digest=sha256:31e940b58eb25f527193f10417446792d213e443f09cb3747ca2f91c9227316e

Pith citing papers

No inbound Pith citation observations are available.