Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:51.010785Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2506.07600.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:51.010785Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T19:15:02.124035Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T20:06:13.270265Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0f78863e-e058-499f-8271-8b3df8218a37 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63a5283-71c9-45e2-a079-fc9ef20f04ec · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35: 23716–23736, 2022
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b4c3222-d80a-40e0-9ff5-6713bb3c1325 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Vivit: A video vision transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4049a5bd-46e6-4991-a9e3-2b1c5d84afce · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unified graph structured models for video understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3c25b8a-437b-4dc0-9001-eca88c2a9f25 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Self-rag: Learning to retrieve, generate, and critique through self-reflection
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa681f7-bfa0-43cc-9892-b054964886a7 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0bd1e4f-7665-4a89-852b-f07dce6a7fa9 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 787bbd23-4312-46c8-a59a-1bdd76534d1b · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d00498a-0c60-4167-9e19-2234800adfc9 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5ce122e-a137-4c49-a7ca-877690d956bb · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Sscformer: Push the limit of chunk-wise conformer for streaming asr using sequentially sampled chunks and chunked causal convolution
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 353333cb-08da-4433-80ec-c3f2e8d06b12 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Event segmentation and seven types of narrative discontinuity in popular movies.Acta psychologica, 149:69–77, 2014
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53b133fe-8f81-4f76-ac9a-d7e0e3ebc3b4 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Large-scale narrative events in popular cinema
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42d938a9-4bf6-435d-ad1f-ca8e98c3aa5d · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding From Local to Global: A Graph RAG Approach to Query-Focused Summarization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16cffcf3-1e1b-45a6-83ef-a3abc8e033f4 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Videoagent: A memory-augmented multimodal agent for video understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2616490e-330d-4164-ad63-ef68cacedb31 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1cab8fa-f191-40aa-9038-3dd0efdbfaf9 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8acc01-3b82-452f-92dd-3ba6f04c347f · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Imagebind: One embedding space to bind them all
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c4d502e-b1bd-405a-8c0c-27d89342ab2e · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Lightrag: Simple and fast retrieval-augmented generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8590ccea-a560-4f54-a566-bd67cde45dde · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoRAG: Retrieval-Augmented Generation over Video Corpus
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fa4fc56-0f84-4871-994c-fa8357254b9c · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Learning temporal video procedure segmentation from an automatically collected large dataset
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3c713945-46a5-408d-acb2-a0ecf0d7a903 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Diffusionret: Generative text-video retrieval with diffusion model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ed6b7bfa-2166-4a9b-93b6-5f8e8ff72f6e · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c089fd8-d280-4021-97cd-dbdd99f5ce81 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Multimodal Reasoning with Multimodal Knowledge Graph
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 351fcb6e-6148-4978-b96e-70ca201633ab · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98855b0-59d2-4594-862a-32d7436c411f · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19432be-2e3c-44ab-82f5-9ad99d1453ae · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5542f591-d7eb-4a35-9ee4-6c374c36ef95 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Llama-vid: An image is worth 2 tokens in large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bde4f57-de10-42fb-a65a-61b852cbd1d7 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13c512e3-d06d-4697-80a8-633f478325e7 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding The impact of continuity editing in narrative film on event segmentation.Cognitive science, 35(8):1489–1517, 2011
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06cd2825-1733-4f1c-b7ce-5c6ed7b1e87f · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Learning joint embedding with multimodal cues for cross-modal video-text retrieval
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc233881-faed-4732-8fd9-8fe0c8ade76d · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Boundary-aware Self-supervised Learning for Video Scene Segmentation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ffcb655-c661-401a-9aee-a54baf56b1ce · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Video shot boundary detection: a review
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 62a16ffe-9666-4ddb-8c19-34f9ce839849 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding What is a multi-modal knowledge graph: A survey.Big Data Research, 32:100380, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e5e5683-a5e7-4451-a79b-9e0c6bf79174 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Learning transferable visual models from natural language supervision
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2e3e7c-3ea5-4db7-ba22-9e6a489e87bc · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Robust speech recognition via large-scale weak supervision
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 569ffada-7c69-455a-ba47-b05e7ad7e021 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf78b05-a0ad-4f6c-bd63-585137abeb79 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Temporal video segmentation to scenes using high-level audiovisual features.IEEE Transactions on Circuits and Systems for Video Technology, 21(8):1163–1177, 2011
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dfd66a0f-e71e-4f30-b259-61b9958a83e3 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Videobert: A joint model for video and language representation learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a62d48e-553a-420c-b665-e452fc124e47 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Temporal scene montage for self- supervised video scene boundary detection.ACM Transactions on Multimedia Computing, Communi- cations and Applications, 20(7):1–19, 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d489f82-4221-48ca-93fe-4945e8906294 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Videoagent: Long-form video understanding with large language model as agent
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 539bc9a7-3595-46ab-8a8e-871564008167 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Scene consistency representation learning for video scene segmentation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bbecc0ac-db34-4231-bf89-b3a95d52059a · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417c172f-13e5-4862-81ef-0eff5edd6e6a · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Retrieval- augmented egocentric video captioning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b844bde7-bfd0-42a9-b90a-90514f00b513 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7a23ea1-dcb5-4a6d-a212-4b6c0bf8ac72 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Momentseeker: A comprehensive benchmark and a strong baseline for moment retrieval within long videos.arXiv preprint arXiv:2502.12558, 2025
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d972c0f-14f9-4ff4-87b8-55fc80973776 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding A formal study of shot boundary detection.IEEE transactions on circuits and systems for video technology, 17(2):168–186, 2007
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8cdd8dc1-7a4b-429b-b457-13c683723b0f · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding The brain’s cutting- room floor: Segmentation of narrative cinema.Frontiers in human neuroscience, 4:168, 2010
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94805a66-2b73-4012-b5ab-e94a4f8e294d · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Merlot reserve: Neural script knowledge through vision and language and sound
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 74bf6da5-efe4-415c-8d58-fe897a7d3590 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding RAKG:Document-level Retrieval Augmented Knowledge Graph Construction
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6722b5b0-1c8d-4a75-a47b-60ab28d58aa8 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff9f0a4-ca16-4e0f-a52a-b64249c31166 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Please maintain the required format in your response
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66e0bc2a-77d6-4210-af0d-cb3bf0ebfd8c · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5b9ec369-cc81-4154-bc82-1272e2439b41 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Please maintain the required format in your response
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7353939f-dbf5-4101-88c7-32311ade0fd3 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39d875e6-f1c9-4291-9ae0-0c50abad866e · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bf189050-0e90-4dd1-835d-073c1343bfd4 · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dcc53b3c-47fd-46e1-8b8d-b7856737744b · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56d795ad-6a2c-412d-b5f2-5fdaba24c93f · outbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Is This the End of RAG? Anthropic’s NEW Prompt Caching
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 421e6316-0ac4-469d-b56b-9ab6d9809818 · inbound
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 47a983ba-dc27-4128-a3a7-eec95b5f3594 · inbound
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.