Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:01:02.423170Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2505.22980.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:01:02.423170Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6c51cfaf-ecdf-40d7-ad8c-681e9defe6fb · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d82a762-a827-434e-bece-cc68cfb7d6f8 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Frozen in time: A joint video and im- age encoder for end-to-end retrieval
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be240a50-c82f-4585-9e23-60e5a9698fa5 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Lumiere: A Space-Time Diffusion Model for Video Generation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d0e7af-36c5-4746-99e5-dd6195be5d4a · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Multidiffusion: Fusing diffusion paths for con- trolled image generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 842663ce-81b1-4b8c-8aa1-fa07edf8a6ab · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36d37ad9-256f-4803-b3a0-9f2618fed3e0 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Align your latents: High-resolution video synthesis with latent diffusion models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7ccfdea-b328-485b-9aa7-2f6d59470f66 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b6071c2-fd71-4667-990e-1f617c379fd0 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Videocrafter2: Overcoming data limitations for high- quality video diffusion models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b335be6a-f2fe-4445-abf3-c8838c6ffc80 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86692de2-767f-4608-8f4a-42cd3c20b68f · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Sora as an agi world model? a complete survey on text-to-video generation.arXiv preprint arXiv:2403.05131, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccaa405-75ef-4749-abaa-783b7ae51a71 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b25e274-37d8-4f14-9df1-b3ea4c4abc93 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation DiffSynth-Studio: Enjoy the magic of Diffusion models!https://github.com/ modelscope/DiffSynth-Studio, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3935d16d-c838-4c6e-a3c0-37efa2f38078 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.International Conference on Learn- ing Representations, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37fc4c8d-8a6f-458c-a9a0-ec81da22db58 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f32561d-1157-4416-8631-f864ad6099ec · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation CLIPScore: a reference- free evaluation metric for image captioning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1933484-4e79-4734-841f-0027a4878843 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Imagen Video: High Definition Video Generation with Diffusion Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33db6231-caa2-4980-810c-6cb240c89537 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Denois- ing diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a3b2a17-96da-4b8b-b8e0-5a45bb38fc4c · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Video Diffusion Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa1bd13-a401-4165-ba1b-c5537d101575 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50e588c3-2169-4d3c-acd8-eae01d532af9 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Vbench: Comprehensive benchmark suite for video generative models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6015e7c6-9ffa-4ff7-b1aa-9ed77f66ca46 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation High-quality Text-to-video Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b974e219-4fb0-4c1e-ae84-95e7459971bd · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation KLING AI: Next-Generation AI Creative Studio.https://www.klingai.com/, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 272c985e-138a-496a-b615-57a1b4732816 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Multi-concept cus- tomization of text-to-image diffusion
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2233dae1-8b17-4b73-b21d-c9aca587af4f · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be3c788d-8ee8-4078-a85b-caa23787f32f · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Gligen: Open-set grounded text-to- image generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80044332-1829-4001-9a49-611707efe2e6 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Movideo: Motion-aware video generation with diffusion model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff403e82-cb7a-4ac2-a7ee-489674d20e87 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a634ebcf-1c1d-4882-9502-b3558c670a96 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Detector Guidance for Multi-Object Text-to-Image Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be93cda6-eecb-4039-958a-99d6463f79db · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ab781fb-589e-4f9b-93fb-8ee640cf5507 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Lumaai.https://lumalabs.ai/,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32ae22fd-4d49-4c6d-ad5c-14872d010344 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Vidm: Video implicit diffusion models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37de2084-99d7-4691-bb74-cfadb942a24d · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Hailuo AI: Captivating AI Videos Gen- erated with Hailuo AI .https://hailuoai
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00d20cf3-5a94-4aab-a214-7fbe8a35b8a0 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Dreamix: Video Diffusion Models are General Video Editors
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed0700fb-40f1-46c2-b9df-8dc13412dc53 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation WorldSimBench: Towards Video Generation Models as World Simulators
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd886c3-2fdb-4825-9188-ee47f941fb16 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b61f693c-65c2-4189-8e45-07b895997203 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation High- resolution image synthesis with latent diffusion mod- els
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 077b15c3-383c-4714-9e5e-ec3e29d03024 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Gen-2: Generate novel videos with text, im- ages or video clips, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d5210f5-31b9-4dd7-9af0-34f0764c78a1 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55510d57-fa72-422a-81f8-f760346fd4d4 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d567d424-2a2d-472e-ab42-d20348480375 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Denoising Diffusion Implicit Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b94524fd-15c5-4b14-a04b-404da42a760b · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf967497-233e-4b07-818c-d3b7f7e4bfe7 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Raft: Recurrent all-pairs field transforms for optical flow
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d42c452d-9d30-4735-b207-c73d02727e5c · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation LLaMA: Open and Efficient Foundation Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c91d916-d50f-466b-ae54-23f87b483035 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Fvd: A new metric for video generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9b98d0c-7c9f-4cf4-8689-727ac648b204 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Vchitect 2.0: Embark on a Visual Fan- tasy Journey.https://vchitect.intern- ai.org.cn/, 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7728518b-2d64-4a24-aa08-5fe07fb488c1 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Phenaki: Variable Length Video Generation From Open Domain Textual Description
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92d328c-fdd7-4fc2-8881-e6dc1d5cc576 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation ModelScope Text-to-Video Technical Report
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76bd7851-616f-47fe-a686-038e71947e91 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Boximator: Generating Rich and Controllable Motions for Video Synthesis
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd7da98-74db-4435-9e17-d1bf73de581c · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVLM: Visual Expert for Pretrained Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77514227-2a4a-46c9-a9b4-dbef5600b2ae · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Emu3: Next-Token Prediction is All You Need
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5527bad4-bd6c-4d2a-a004-2ab86e607bdb · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b73f522-f386-48ad-8fa6-f7b34042c99f · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation LaVie: High-quality video generation with cascaded latent diffusion mod- els.IJCV, 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0969d8ed-1fb6-4f3b-99cf-9b277420cd86 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Customvideo: Cus- tomizing text-to-video generation with multiple sub- jects.arXiv preprint arXiv:2401.09962, 2024
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25ef5394-9984-4e2d-ba69-6861e8466566 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Grit: A generative region-to-text transformer for ob- ject understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5b9fce2-6191-40d8-9afd-ef79d6c2815a · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation FreeInit: Bridging Initialization Gap in Video Diffusion Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81dc84f9-3b18-4b54-ba9f-d8405d3a5837 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Dynamicrafter: An- imating open-domain images with video diffusion pri- ors
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 591e9272-6616-4805-bcaa-020b06632107 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation MSR- VTT: A large video description dataset for bridging video and language
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dd678f3-210e-4785-8b6f-57f08c563b28 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Advancing high-resolution video- language representation with large-scale video tran- scriptions
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bdfe9b52-8ebb-4196-94bb-615dfa0b7647 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Video in- stance segmentation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc07aed5-accc-4f00-9c2c-b2af7ecc7e2a · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e031dc-3a43-4f2c-95a1-ce678daf7a23 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1095eefb-eac3-49e4-a793-9812637b7ad5 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Show-1: Marrying pixel and latent diffusion models for text-to-video generation.Interna- tional Journal of Computer Vision, pages 1–15, 2024
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 274b38f8-d867-4814-bdab-22d674d6b567 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a86a191-9976-4ebc-9f37-0abae1793df8 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Real-time vehicle detection based on improved yolo v5.Sustainability, 14(19):12274, 2022
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2111eff1-6b07-4239-8e8b-1b32f1cb52c7 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation Open-sora: Democratizing efficient video production for all, March 2024
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 808ab046-08ce-426e-8557-ccebdeab53d0 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53331444-f74f-4ea2-b685-554cd1e63306 · outbound
MOVi: Training-free Text-conditioned Multi-Object Video Generation 2, 5, 7, 8
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.