Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:43:56.260106Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.01938.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:43:56.260106Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 43c302e5-3448-4025-a67e-3f2addf75b44 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ebc2bbc-2b1f-4875-8e90-0bf649427181 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4802eb36-4927-402e-9cb4-ba5de35f8b21 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bacf8874-ee68-4c67-b252-4550bb47cc04 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Video generation models as world simulators, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dc6d1286-0035-4aa1-9ff1-5e25eab8a452 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Muse: Text-To-Image Generation via Masked Generative Transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 893d855c-e8c7-4912-934f-7574eb969633 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Sharegpt4v: Improving large multi-modal models with better captions
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b16689d6-c1ab-4694-956c-5fefb917da8e · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ae716e0-f550-4984-a7e7-c66ce3d745bf · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c26b38-ff68-4f8c-9868-bfce4f575a80 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Yolo-world: Real-time open-vocabulary object detection
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 54f45424-b34c-4608-89b0-aff910a20f3f · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Autoregressive Video Generation without Vector Quantization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9778989b-e92b-49d3-868a-b5a4894899e0 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Long video generation with time-agnostic vqgan and time- sensitive transformer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation af402e0c-d85b-4ea5-bd6b-07dec89b34af · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Imagebind: One embedding space to bind them all
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 914b15d5-6e58-4f7d-a702-6ff1ec805cbd · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Language is not all you need: Aligning perception with language mod- els
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7786ac6-bae9-48bf-9099-ed46134fe665 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Vbench: Comprehensive bench- mark suite for video generative models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c2f98bf3-07c7-456a-ab32-346e99bbbc1e · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset GPT-4o System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78cac45b-6862-4a70-85a1-77d031eb7518 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Phi-2: The surprising power of small language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e987e9a4-bb83-4785-bf3c-4dbd864715b2 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Miradata: A large-scale video dataset with long durations and structured captions
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4c6a182-0564-4b9c-844f-bb400bca1984 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce8d505-8d75-4718-a23e-d4d7ae723baf · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Obelics: An open web-scale filtered dataset of interleaved image-text documents
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f68c233-efe3-4ba6-af5f-2241af2a0f34 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Autoregressive image generation without vec- tor quantization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e87fad39-fdeb-4355-a567-ee6e2f2406ed · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Open-Sora Plan: Open-Source Large Video Generation Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7658d01-f6fa-45a2-afa7-233a783f36ac · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Decoupled weight decay regularization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 89f81468-e333-496d-89e9-bb1d430facfe · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798a4ad3-fd62-4e98-bbfa-b8dbaddf5a59 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Improved denoising diffusion probabilistic models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa9d88cf-5aae-4e94-ac11-49ca387fde33 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Learning transferable visual models from natural language supervi- sion
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9189915f-074b-4403-b8f3-ffb67b1a3b52 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2d364bf-b8a7-4e32-a9bc-c6e60bf5a320 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 53aeb84d-0c09-40bc-a54b-87155aaad337 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Raft: Recurrent all-pairs field transforms for optical flow
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b811d9-46b6-4faf-b471-aa4e413b4c5b · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset VideoTetris: Towards Compositional Text-to-Video Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation affb843b-73a5-4689-9c36-cc35776635b7 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Mcvd-masked conditional video diffusion for prediction, generation, and interpolation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c2bdba8a-bd9b-481b-8228-7e8c00c13748 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Emu3: Next-Token Prediction is All You Need
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f0879d8-0e8e-4004-86de-263b2740616c · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a67842d9-e018-4f27-a2d6-f5a6c0c93625 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Ad- vancing high-resolution video-language representation with large-scale video transcriptions
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae47c4a9-105c-400c-bcc5-1530b30d7e8f · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Vript: A Video Is Worth Thousands of Words
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c2dbd5-ad44-4826-90c9-d8dd9ada56d2 · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf65951-163e-4df2-8e47-77145e20c97e · outbound
CI-VID: A Coherent Interleaved Text-Video Dataset Multimodal c4: An open, billion-scale corpus of images interleaved with text
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.