Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:01:14.798928Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 2 inbound Pith citation observations for arXiv:2505.03173.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:01:14.798928Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T14:53:56.693464Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T15:28:33.980713Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dd674685-4fb2-4151-9ace-18582f4983fa · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0c62299-891b-4494-b823-7904943be452 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph An image is worth 16x16 words: Trans- formers for image recognition at scale
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c8811903-3d78-4809-98d6-23b0db109b08 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Assistgpt: A general multi-modal assistant that can plan, execute, inspect, and learn,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fcb1de7c-c632-4199-8614-e6fac6925367 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Ma-lmm: Memory-augmented large multimodal model for long-term video understand- ing
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a18529be-24f9-42ea-ba7e-91f92a28d9d3 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Video recap: Recursive captioning of hour-long videos,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5021188d-ee2d-49a0-8938-43d3f977888a · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Videorag: Retrieval- augmented generation over video corpus,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 27a4157c-9b7a-497b-b809-96b8663b1a76 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Egoschema: a diagnos- tic benchmark for very long-form video language under- standing
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0768b8ad-b363-4f7a-be89-a2fa77856ff3 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Verbs in action: Improving verb understanding in video- language models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d65c5189-8114-4652-9ab9-9b54a5a3328a · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Learning to compress prompts with gist tokens
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8161181f-2bb3-4950-b669-e0351003646b · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph A simple recipe for contrastively pre-training video-first encoders beyond 16 frames
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 03ee98dd-2565-4305-9f96-bf84ea34ef64 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Learning Transferable Visual Models From Natural Language Supervision
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3599a215-5317-44ef-b9b7-3d0ce8bc8f25 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph SAM 2: Segment Anything in Images and Videos
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf75a49d-0623-40bf-91e3-fe5ad7a0b0ea · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Sentence-bert: Sentence embeddings using siamese bert-networks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation be8a1184-f330-438a-a3c1-c83e71f1ebe7 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b73d6246-93f3-4368-88c4-970bc6c5cfa7 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 437d1270-31e3-4984-8e90-bd84155314d3 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Vipergpt: Visual inference via python execu- tion for reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a5b72c30-35c9-4022-8741-26c3e5a64d9c · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88415a62-705a-428f-a11b-fe8a0863cb71 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Ada- coder: Adaptive prompt compression for programmatic vi- sual question answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7488f1a5-dbe9-43d4-99f5-b28335ffd11a · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph ViLA: Efficient Video-Language Alignment for Video Question Answering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d416d066-b936-416a-be9e-7c95e9de30f6 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Multi-object event graph representation learning for Video Question Answering
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff6182f9-1f46-4c46-a885-50cc455b62bc · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Lin, and Shan Yang
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a304a8b1-b8f2-4402-b061-e1be1a3df203 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Next-qa: Next phase of question- answering to explaining temporal actions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 66fa0fe5-85ee-4cbd-a34b-6af5a826998b · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Panoptic video scene graph generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 59da04ba-dce6-43e0-bb27-f2689270641f · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Hitea: Hierarchical temporal-aware video-language pre-training
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cf82f1df-f90a-417e-820c-0252b135159b · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Self-chained image-language model for video localization and question answering,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a46a9c46-72b0-4d6f-a180-cf40ee98ea68 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph A simple llm framework for long-range video question-answering,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4a9a83e-32cb-450f-9419-6bacb76822ad · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Scene Graph Generation: A Comprehensive Survey
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 601ee9d7-c920-4086-bbd6-fa27ee0eed10 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Blip-2: bootstrapping language-image pre- training with frozen image encoders and large language models
Reference 2011
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d8918c0-6144-4c31-8156-6cac6447526c · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Annotating ob- jects and relations in user-generated videos
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec729ce4-c5cc-4b2d-ae53-ea614241bbc9 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph The Llama 3 Herd of Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f17f6d06-23d4-4578-a86f-5c7997e98c73 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph H ´enaff
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9e38ea69-44dd-4392-bf0b-ca863ef50213 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Video scene graph generation from single-frame weak su- pervision
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f9b52198-9885-4c3c-9785-495002c2ed05 · outbound
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph Thinking, Fast and Slow
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a2985342-f94f-4a93-8025-16aeb2d2fd8d · inbound
Rethinking RAG in Long Videos: What to Retrieve and How to Use It? RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ef80f3f6-6590-4486-9016-d493673eabde · inbound
Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.