Pith. sign in

Paper Citation Record · LEDGER

VideoRAG: Retrieval-Augmented Generation over Video Corpus

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2501.05874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05874 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:02.359068Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 54eae12c-d336-4b32-99a9-95fd043322e0 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:26.672299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:26.672299Z digest=sha256:c659eaa385466ddc6bf2efb5c92494a169ddb2374ff696d2c4faba307e5f28da

Observation d46fa311-2b27-4b04-a6ab-e182a8e94d36 · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:02.359068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:02.359068Z digest=sha256:28c07a204f3c1df418358295cf59258d5172a478dc8a8a5547394e36279f8b68

Observation 8590ccea-a560-4f54-a566-bd67cde45dde · inbound

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding cites this paper.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.801689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.801689Z digest=sha256:8c7e0ee5df30f69da4d0477ddd191794977d86b66494eecbb18aa7be7ca57098

Observation aeaa6e0c-f7df-432c-8d78-712798e27b10 · inbound

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification cites this paper.

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T01:04:08.928124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:04:08.928124Z digest=sha256:40bbde1fb6bf5a52e2373c76db0c4400ea1bf9beb8ede007343e8950f2eece9e

Observation 2d83ecfe-48f0-4669-aeb5-e8d9310fb367 · inbound

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation cites this paper.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.433930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.433930Z digest=sha256:2395818d616f9635c852ba8b4d4f7569f06180321683498927da2a45465438ac

Observation 8bc95e34-6faa-4aaf-8c14-fbb97c7c3366 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.503874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.503874Z digest=sha256:26c42f2dd026efe2c845ec6674ead2bf7f147176198e8cb07780ceec5ade13e4

Observation 745eaa08-5760-42de-b4b3-bbab4e0afe39 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:52.533572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:7d86b28b640c18239fb1925445a057a397427587ea27ff040933ecad213f55d0

Observation 6fba3a16-47c4-4041-843e-2a1cdccdcdf4 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.251652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:fd5adc1b7f6e3a4ced63d8fc83bf93533138c2d80c93c0e6a8fa46e86d555655

Observation b2a4c83f-20a5-404b-ae7f-ef26dd1707f6 · inbound

UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos cites this paper.

UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:04:06.093362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:00:16.944922Z digest=sha256:0ba2c4262011cb2e4a6e1974d7f20480c4a66555c6f189434a58810e22f0e0f9

Observation 54deaee5-99ef-4616-a056-ee59baee9b5a · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.367481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:e9a487fd7783c23c62f6b37295ff067eda70da109b70718db54776c637880335

Observation 7c2ee8c2-0367-4fbe-b531-b3d3b71668f7 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.602804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:be150b1896fd4890ed32afa078b0b9bb927926a8a0c083f9e69707d2f9dce728