Pith. sign in

Paper Citation Record · LEDGER

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2506.07600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07600 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:51.010785Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:15:02.124035Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:06:13.270265Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f78863e-e058-499f-8271-8b3df8218a37 · outbound

This paper cites Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.716342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.716342Z digest=sha256:0450f821248ab07dfce24907e546573d0fcb2c891f53642ca3ab19080793e565

Observation b63a5283-71c9-45e2-a079-fc9ef20f04ec · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35: 23716–23736, 2022.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35: 23716–23736, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.722098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.722098Z digest=sha256:6bbf5e070af3d00937502c377bdc759f59ec21ca8ba4936903d0877c789002c3

Observation 2b4c3222-d80a-40e0-9ff5-6713bb3c1325 · outbound

This paper cites Vivit: A video vision transformer.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Vivit: A video vision transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.726908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.726908Z digest=sha256:fd6d7a111a0f65f5723ce990a8b65b4d18a1d1f8232a2e8c02815fbf09080d14

Observation 4049a5bd-46e6-4991-a9e3-2b1c5d84afce · outbound

This paper cites Unified graph structured models for video understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unified graph structured models for video understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.098046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.732034Z digest=sha256:9c90cb5f9bd35e07fa12ab8a8f8e6a587d29a7c375b97d90323cec6df6cde541

Observation b3c25b8a-437b-4dc0-9001-eca88c2a9f25 · outbound

This paper cites Self-rag: Learning to retrieve, generate, and critique through self-reflection.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Self-rag: Learning to retrieve, generate, and critique through self-reflection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.736730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.736730Z digest=sha256:f8127d90d1a957e6f889693b61cb125c0f4d4b6d839e552a7e728fd5d6f2803c

Observation 4aa681f7-bfa0-43cc-9892-b054964886a7 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495, 2024.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.070182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.741138Z digest=sha256:d2d0a257afc158c9b6484be42059134d96c18a760c745563f4daa21bfe4274f0

Observation d0bd1e4f-7665-4a89-852b-f07dce6a7fa9 · outbound

This paper cites WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.746422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.746422Z digest=sha256:834f2882fd6d96c7a0aa6bcba473fb3e89fd87affcdf9b4d0286df375d8f7eb2

Observation 787bbd23-4312-46c8-a59a-1bdd76534d1b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.751331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.751331Z digest=sha256:6aeee5ed8f83fa7ad503f40e64c734242d702a53fcab18bcb41b1b9c1b2bf5b6

Observation 1d00498a-0c60-4167-9e19-2234800adfc9 · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:52.052581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.756117Z digest=sha256:53bca441a0fac1d87297ff08553fbf31635a8eb206f1552f38baec84f3c08b75

Observation a5ce122e-a137-4c49-a7ca-877690d956bb · outbound

This paper cites Sscformer: Push the limit of chunk-wise conformer for streaming asr using sequentially sampled chunks and chunked causal convolution.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Sscformer: Push the limit of chunk-wise conformer for streaming asr using sequentially sampled chunks and chunked causal convolution

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.034449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.762345Z digest=sha256:e1fabc8a2b53aa8268bd620b0c41a209cba827234a6b9f880d289b8a0f9d3e4e

Observation 353333cb-08da-4433-80ec-c3f2e8d06b12 · outbound

This paper cites Event segmentation and seven types of narrative discontinuity in popular movies.Acta psychologica, 149:69–77, 2014.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Event segmentation and seven types of narrative discontinuity in popular movies.Acta psychologica, 149:69–77, 2014

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.017494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.766627Z digest=sha256:c5c7ca173495f00e9f474506cb7302b102e4638e3aa9398885004c41d2b73448

Observation 53b133fe-8f81-4f76-ac9a-d7e0e3ebc3b4 · outbound

This paper cites Large-scale narrative events in popular cinema.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Large-scale narrative events in popular cinema

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:52.000990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.770754Z digest=sha256:836c9fbf27a14c7fcf8da901656b6023f96c2910e4ac0c80219a2d5c411d52f0

Observation 42d938a9-4bf6-435d-ad1f-ca8e98c3aa5d · outbound

This paper cites From Local to Global: A Graph RAG Approach to Query-Focused Summarization.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding From Local to Global: A Graph RAG Approach to Query-Focused Summarization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.775173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.775173Z digest=sha256:1691e825cf3bd8cd63c25d27acc2038876f2057890aad9668a7e45ffc27a0070

Observation 16cffcf3-1e1b-45a6-83ef-a3abc8e033f4 · outbound

This paper cites Videoagent: A memory-augmented multimodal agent for video understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Videoagent: A memory-augmented multimodal agent for video understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.985406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.779943Z digest=sha256:d7db4a082740950d42c9d13b199e19484e64a27b9b589bf7335d73d516b0ef54

Observation 2616490e-330d-4164-ad63-ef68cacedb31 · outbound

This paper cites Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.784040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.784040Z digest=sha256:8b2d115d62df17759214dd0027eb44bc188618e60f769c79e2e95d81f6ee9dc8

Observation c1cab8fa-f191-40aa-9038-3dd0efdbfaf9 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.788520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.788520Z digest=sha256:b5dcbb903b0e432d970a672d7da82f6e0944179200cc98c6011f53bb7720d580

Observation de8acc01-3b82-452f-92dd-3ba6f04c347f · outbound

This paper cites Imagebind: One embedding space to bind them all.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Imagebind: One embedding space to bind them all

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.793225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.793225Z digest=sha256:e2b2462a29281c144b079c10f7b60e8818f40c36ec4df0d022517f4b54554e3e

Observation 1c4d502e-b1bd-405a-8c0c-27d89342ab2e · outbound

This paper cites Lightrag: Simple and fast retrieval-augmented generation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Lightrag: Simple and fast retrieval-augmented generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.797523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.797523Z digest=sha256:1154abd5299ad3633a7a03c51b56eaff653126240a42e695e2aad7bd6ff7cb89

Observation 8590ccea-a560-4f54-a566-bd67cde45dde · outbound

This paper cites VideoRAG: Retrieval-Augmented Generation over Video Corpus.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.801689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.801689Z digest=sha256:8c7e0ee5df30f69da4d0477ddd191794977d86b66494eecbb18aa7be7ca57098

Observation 5fa4fc56-0f84-4871-994c-fa8357254b9c · outbound

This paper cites Learning temporal video procedure segmentation from an automatically collected large dataset.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Learning temporal video procedure segmentation from an automatically collected large dataset

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.948243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.806583Z digest=sha256:38b0efad5b903db644ba5077dd53dc01a89f868b7bf2e9b91865ebf9127b72f8

Observation 3c713945-46a5-408d-acb2-a0ecf0d7a903 · outbound

This paper cites Diffusionret: Generative text-video retrieval with diffusion model.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Diffusionret: Generative text-video retrieval with diffusion model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.932126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.810986Z digest=sha256:2f900c688ae841d37344619b6e60a7909fa10c862182db5c3d9d638fe416485f

Observation ed6b7bfa-2166-4a9b-93b6-5f8e8ff72f6e · outbound

This paper cites Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.815034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.815034Z digest=sha256:3d097504caebd0edaac97c4efe024072b051f68fa8c8bbe3b6c7bd98333e1d0e

Observation 6c089fd8-d280-4021-97cd-dbdd99f5ce81 · outbound

This paper cites Multimodal Reasoning with Multimodal Knowledge Graph.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Multimodal Reasoning with Multimodal Knowledge Graph

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.819345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.819345Z digest=sha256:657de390924ee59dffb433c10a5290c626f95408051486d169b4c08221b5d848

Observation 351fcb6e-6148-4978-b96e-70ca201633ab · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.823707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.823707Z digest=sha256:73ff368cf5978be8d4eda7c34293711cd1b08fffcf58f9c44b1e53df6ba2eeb3

Observation b98855b0-59d2-4594-862a-32d7436c411f · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.828025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.828025Z digest=sha256:0f04b460fec29112f7a8d6a68c947a9a791ccece2b04510ae8eb84c963ce9505

Observation c19432be-2e3c-44ab-82f5-9ad99d1453ae · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.832358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.832358Z digest=sha256:3c169eb6c7aff0b58b2de0eca63ae4adb7f8066ad3230a5517bde1a2a8c9929d

Observation 5542f591-d7eb-4a35-9ee4-6c374c36ef95 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.837354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.837354Z digest=sha256:9d4f342a666114ab08e0e7a129ff0925c43162dfb9da242d0ec61b3f07d173b1

Observation 8bde4f57-de10-42fb-a65a-61b852cbd1d7 · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.842483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.842483Z digest=sha256:1823e6602436a194b8a7f8edd5ba79892a343d3979bc4387b36c167d24120232

Observation 13c512e3-d06d-4697-80a8-633f478325e7 · outbound

This paper cites The impact of continuity editing in narrative film on event segmentation.Cognitive science, 35(8):1489–1517, 2011.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding The impact of continuity editing in narrative film on event segmentation.Cognitive science, 35(8):1489–1517, 2011

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.883271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.847978Z digest=sha256:bd09ef3b14955cf8e396d74338ba9f0635db1f341dc2806c236c9993b05a3871

Observation 06cd2825-1733-4f1c-b7ce-5c6ed7b1e87f · outbound

This paper cites Learning joint embedding with multimodal cues for cross-modal video-text retrieval.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Learning joint embedding with multimodal cues for cross-modal video-text retrieval

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.866744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.852386Z digest=sha256:0472d06c0dd7ad884ad1156c5e79be5e2eb0e43879b293e5992f96c03d1589b8

Observation bc233881-faed-4732-8fd9-8fe0c8ade76d · outbound

This paper cites Boundary-aware Self-supervised Learning for Video Scene Segmentation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Boundary-aware Self-supervised Learning for Video Scene Segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.858475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.858475Z digest=sha256:6a7104a7816fa434395bd31ffb635a6370c8eaac3161e9da85f6cd3eaea29861

Observation 4ffcb655-c661-401a-9aee-a54baf56b1ce · outbound

This paper cites Video shot boundary detection: a review.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Video shot boundary detection: a review

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.849862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.863274Z digest=sha256:01db1a6759038c95b11d72247fc23fd496f998a26d79f6fd8f6e24369548c99d

Observation 62a16ffe-9666-4ddb-8c19-34f9ce839849 · outbound

This paper cites What is a multi-modal knowledge graph: A survey.Big Data Research, 32:100380, 2023.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding What is a multi-modal knowledge graph: A survey.Big Data Research, 32:100380, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.832449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.867733Z digest=sha256:5fd09a831ebd8e34cba6ec6643c58a1698ab244b3a3430a86384f09a93cc8485

Observation 0e5e5683-a5e7-4451-a79b-9e0c6bf79174 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Learning transferable visual models from natural language supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.872103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.872103Z digest=sha256:22f1bc761abc0ab6e0f3823873773db678d708fc2b37055b57e19c6a263f2e0e

Observation 0c2e3e7c-3ea5-4db7-ba22-9e6a489e87bc · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Robust speech recognition via large-scale weak supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.876870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.876870Z digest=sha256:9e1bf15879b3d386d3569ffbeab2f12b38c458974b477ec80866a846d8fe9fa8

Observation 569ffada-7c69-455a-ba47-b05e7ad7e021 · outbound

This paper cites VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.881284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.881284Z digest=sha256:fdb3f70b63a28913eddb21c35f54185983fd8ce58b41e63fe3eef1baf7211d6b

Observation 3bf78b05-a0ad-4f6c-bd63-585137abeb79 · outbound

This paper cites Temporal video segmentation to scenes using high-level audiovisual features.IEEE Transactions on Circuits and Systems for Video Technology, 21(8):1163–1177, 2011.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Temporal video segmentation to scenes using high-level audiovisual features.IEEE Transactions on Circuits and Systems for Video Technology, 21(8):1163–1177, 2011

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.792701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.885872Z digest=sha256:b1580ecd91137a7792274260205f7f78ace2adb31e48f4f296842e824456d137

Observation dfd66a0f-e71e-4f30-b259-61b9958a83e3 · outbound

This paper cites Videobert: A joint model for video and language representation learning.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Videobert: A joint model for video and language representation learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.775488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.890403Z digest=sha256:b59f8542e4eb91ad5bbe92a8ac4a8ef96a850822c0165d9752fe1556e5e9639e

Observation 9a62d48e-553a-420c-b665-e452fc124e47 · outbound

This paper cites Temporal scene montage for self- supervised video scene boundary detection.ACM Transactions on Multimedia Computing, Communi- cations and Applications, 20(7):1–19, 2024.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Temporal scene montage for self- supervised video scene boundary detection.ACM Transactions on Multimedia Computing, Communi- cations and Applications, 20(7):1–19, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.758787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.894622Z digest=sha256:5019a16753e6b0c2cbfe00ef163b88b25d8be53165d2ec9b357a89554061a4fd

Observation 7d489f82-4221-48ca-93fe-4945e8906294 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Videoagent: Long-form video understanding with large language model as agent

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.898518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.898518Z digest=sha256:efe98cfa975c3a2080a90e28f85b68cbd83aebdc30fc699f4777f32f2a0088ce

Observation 539bc9a7-3595-46ab-8a8e-871564008167 · outbound

This paper cites Scene consistency representation learning for video scene segmentation.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Scene consistency representation learning for video scene segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.731995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.902567Z digest=sha256:d1bce4962228534a946426807894316ea1b7d032511835d6542b0d0d8bcf71f6

Observation bbecc0ac-db34-4231-bf89-b3a95d52059a · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.906801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.906801Z digest=sha256:ad074fe27ab25e111a303ddaefce1c7cc71f0ebb89a307e3aac420b1af5cd9f7

Observation 417c172f-13e5-4862-81ef-0eff5edd6e6a · outbound

This paper cites Retrieval- augmented egocentric video captioning.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Retrieval- augmented egocentric video captioning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.715997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.911215Z digest=sha256:cfd182c008079998228a0016197cbe8a9e4679ab53f78ee8debda80738024d93

Observation b844bde7-bfd0-42a9-b90a-90514f00b513 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.915326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.915326Z digest=sha256:d31162a451f430adf105524d3c4823e9ce4ace4fc4a81675f795aec6fa2bd9be

Observation e7a23ea1-dcb5-4a6d-a212-4b6c0bf8ac72 · outbound

This paper cites Momentseeker: A comprehensive benchmark and a strong baseline for moment retrieval within long videos.arXiv preprint arXiv:2502.12558, 2025.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Momentseeker: A comprehensive benchmark and a strong baseline for moment retrieval within long videos.arXiv preprint arXiv:2502.12558, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.920388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.920388Z digest=sha256:bd8e7ca9324631038df39834ac8fc40cb776c6da4138893e8fbc04a046efa80b

Observation 3d972c0f-14f9-4ff4-87b8-55fc80973776 · outbound

This paper cites A formal study of shot boundary detection.IEEE transactions on circuits and systems for video technology, 17(2):168–186, 2007.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding A formal study of shot boundary detection.IEEE transactions on circuits and systems for video technology, 17(2):168–186, 2007

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.699263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.924862Z digest=sha256:52eed53ee3c077cb025345f7bfaefee959146cb28a7a3989d6134cd09fb60728

Observation 8cdd8dc1-7a4b-429b-b457-13c683723b0f · outbound

This paper cites The brain’s cutting- room floor: Segmentation of narrative cinema.Frontiers in human neuroscience, 4:168, 2010.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding The brain’s cutting- room floor: Segmentation of narrative cinema.Frontiers in human neuroscience, 4:168, 2010

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.683472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.929441Z digest=sha256:bb3f46087fa28fd144e85e86593e3a51b326450050c3aedcf3710f4cedf5e31c

Observation 94805a66-2b73-4012-b5ab-e94a4f8e294d · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Merlot reserve: Neural script knowledge through vision and language and sound

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.667190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.935121Z digest=sha256:5aaf18fb674db00e70e56b852578910739ed3a102f5e76e87a73c6fdc4911451

Observation 74bf6da5-efe4-415c-8d58-fe897a7d3590 · outbound

This paper cites RAKG:Document-level Retrieval Augmented Knowledge Graph Construction.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding RAKG:Document-level Retrieval Augmented Knowledge Graph Construction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.939358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.939358Z digest=sha256:bfe92111b394e5604668e5ddb2dd7a31a2dc6b6e744a0c797b2d2a230af0916b

Observation 6722b5b0-1c8d-4a75-a47b-60ab28d58aa8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.943960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.943960Z digest=sha256:7a23aea962dbe2a73cd5a7c53fc5562b5cf84daaf60c1ba41c98928a4ae8331b

Observation 8ff9f0a4-ca16-4e0f-a52a-b64249c31166 · outbound

This paper cites Please maintain the required format in your response.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Please maintain the required format in your response

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.595091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.967095Z digest=sha256:552fc47d1f77c18dc0720d5243636b723a49a22e8f6ab0fc2897925d3527b5e9

Observation 66e0bc2a-77d6-4210-af0d-cb3bf0ebfd8c · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.650621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.971971Z digest=sha256:fd63cda303fc3a02ba400bfcf5902d2e1d831897ed4428bd427c4aa58de2929e

Observation 5b9ec369-cc81-4154-bc82-1272e2439b41 · outbound

This paper cites Please maintain the required format in your response.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Please maintain the required format in your response

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.560866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.989912Z digest=sha256:6180e17df974ab00663746917c65370bc6de7efded0ed8a6a9a369dd194048b6

Observation 7353939f-dbf5-4101-88c7-32311ade0fd3 · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.536324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.994061Z digest=sha256:da8aadcb44037ed5e21a0686721fc992db161d0fd46a2b99029077ea42825f9a

Observation 39d875e6-f1c9-4291-9ae0-0c50abad866e · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.629732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:50.998519Z digest=sha256:389e14e3bce980ee0f40eecef89db8b1ba3ac419ffb4223d615dc0f3f69ca418

Observation bf189050-0e90-4dd1-835d-073c1343bfd4 · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.612920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:51.002426Z digest=sha256:52802a93b801fd03905697b5c3c9dde4a2bf253157fe83f9ae3c7e2a82ab0239

Observation dcc53b3c-47fd-46e1-8b8d-b7856737744b · outbound

This paper cites an unresolved cited work.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:35:51.578428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:51.006523Z digest=sha256:d2e13819d2101853b84751f9e9f3e5384fa9c558305c7cba4f203fb91d0be3df

Observation 56d795ad-6a2c-412d-b5f2-5fdaba24c93f · outbound

This paper cites Is This the End of RAG? Anthropic’s NEW Prompt Caching.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding Is This the End of RAG? Anthropic’s NEW Prompt Caching

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:51.519683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:51.010785Z digest=sha256:1d9dd2244479005459086f80eaee860ec0a30c9fb0a221c2a1264eff1de88484

Pith citing papers

Observation 421e6316-0ac4-469d-b56b-9ab6d9809818 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.244866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:37e4c56f985e61beefefa8729aeee2cb05ed57bf35b8f034dc46e1956431b575

Observation 47a983ba-dc27-4128-a3a7-eec95b5f3594 · inbound

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios cites this paper.

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.272901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T10:15:15.129358Z digest=sha256:95c69603d64f7dd40e71f8daa064893ce0ee7cad782898cfe9215e49676a0046