Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:41:46.943342Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2501.07647.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:41:46.943342Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4b6e5f49-3636-464c-85e6-6c87130acc2b · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b22be7-b64b-49b6-bb56-d74b7c0df6e8 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Align your latents: High-resolution video synthesis with la- tent diffusion models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 631277c9-2b1b-4913-9743-51ef281ec728 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 982eb44b-35c5-4c2f-bacf-68319494733f · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1c833fd9-240e-4cdc-b39b-776bf5163bc5 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Training-free layout control with cross-attention guidance
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db8ca74-d25a-4698-b9b4-01745dd7cbac · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 62167b68-ea64-405b-95cd-b3d5386b4fcd · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491fb305-9892-4bae-a1d2-045effcab649 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aef14e19-ec4e-4cf6-a00e-429c19d20942 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Layoutgpt: Compositional visual plan- ning and generation with large language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6f1c1714-89cf-4df0-a3f1-5849036a0d2a · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations On the content bias in fr ´echet video distance
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6e6a0b24-f824-4afb-b223-96d9176a029d · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f14832ef-e18d-4d32-a077-8c94087197b6 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Denoising dif- fusion probabilistic models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a509f378-ab9a-4a66-b701-430e417f8341 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b512f7f3-1b04-4d41-99d7-6e620314bd92 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Perceiver IO: A General Architecture for Structured Inputs & Outputs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c315cf22-2268-4589-8112-96615752f09f · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations A style-based generator architecture for generative adversarial networks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0ea4ed25-7e74-4691-ac5a-8e870a1ef29e · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Open-sora-plan, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f326661a-2d8c-4cec-935a-a19fc11da731 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Dense optical tracking: Connecting the dots
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4f48b987-b94a-4382-80d7-4a9ce3047b56 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28659284-8aa8-4c06-ab70-20953d6fe2ce · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2205b1eb-5f02-41b9-bafa-78dd3044cfba · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Gligen: Open-set grounded text-to-image generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 63ba12ee-8134-4f6c-821f-c4423f069a4f · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 98d54c78-f21b-462e-af0b-3d79843d08ec · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llm-grounded video diffusion models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 79d97a22-8d0e-4683-a42c-1cbf14c23562 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca446ca-1ebd-4616-90ab-485c5903ad16 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations MotionClone: Training-Free Motion Cloning for Controllable Video Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cd566d9-8960-433c-84e1-b57251091fef · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations 9 Blobgen-3d: Compositional 3d-consistent freeview image generation with 3d blobs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8486c0f-a363-4208-a11c-460b454bed70 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e5f15a6-bbf0-4b47-9d48-ff02b26a31f0 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 725a678a-e2aa-42c1-bcc8-e348297e1adb · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Evalcrafter: Benchmarking and eval- uating large video generation models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 935fbd6d-8b40-454a-9c1a-3a63c044764a · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 60011944-a7c1-480e-9db9-e76d8c7bf9bd · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05fafea7-0cf5-4277-9f27-4771e99299a0 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Compositional text-to-image gen- eration with dense blob representations
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2d717236-49e6-4c5e-ba08-857d2cbaa3eb · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Scalable diffusion models with transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13761c41-a74a-434e-8e8f-72247be31bd4 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations SAM 2: Segment Anything in Images and Videos
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6873bc93-5efe-403c-87a4-62149957d438 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations High-resolution image synthesis with latent diffusion models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62089b14-6a46-4765-b1c9-f1f6b602a0d2 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa572ee-3d91-417c-9f34-4d0026793aa7 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14c488e3-40ff-4233-a5a9-a8794e84021b · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Fourier features let networks learn high frequency functions in low dimen- sional domains
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac317e7e-1472-4e97-b502-6692c27daa47 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VideoTetris: Towards Compositional Text-to-Video Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fa053de-4929-4f99-af7b-6c46cfe7c4d7 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f5a13b-4a3e-41a6-840e-c7eb99a4043d · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Score-based generative modeling in latent space
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8e568363-7a99-4ff4-88c9-4ab15332b295 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations ModelScope Text-to-Video Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ce706c-3b10-4e44-bc29-965b62ff289c · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Boximator: Gener- ating rich and controllable motions for video synthesis
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 237149c0-cfe4-4a60-b919-c6db8fe2eca8 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b1cfec-798c-4e18-b41d-b42376453dfd · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Msr-vtt: A large video description dataset for bridging video and language
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d7aa9d-038b-4415-9fa3-c9886b3c3a7f · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Open-vocabulary panop- tic segmentation with text-to-image diffusion models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 29c53d94-d326-43b2-bc10-7c6520361e72 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Ad- vancing high-resolution video-language representation with large-scale video transcriptions
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff31b245-ad15-466e-ac4c-245d8d1ddc02 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1c8caf67-ceb5-4d66-87b9-2a5e3a4fed17 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Compositional Video Generation as Flow Equalization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b362f654-ab1f-418b-a79a-7247e9727bb1 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d215340d-085f-45be-b3e4-e7d42fc77633 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Scannet++: A high-fidelity dataset of 3d in- door scenes
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ae457b4-859d-4308-aebf-f18053923296 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Show-1: Marrying pixel and latent diffusion models for text-to-video generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 127e9784-ed54-430e-b1b0-ff3e5fc55929 · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Adding conditional control to text-to-image diffusion models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51f1e579-ef71-4ae9-9ec9-8cbd61e9839d · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llava- next: A strong zero-shot video understanding model, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5c3d4a12-2050-450f-81be-34389c94427f · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations both foreground and background
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87f0d80a-6b97-42d8-b5aa-e73017baeb4f · outbound
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Frame0”: “Object2
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.