Pith. sign in

Paper Citation Record · LEDGER

MOVi: Training-free Text-conditioned Multi-Object Video Generation

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2505.22980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22980 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:01:02.423170Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c51cfaf-ecdf-40d7-ad8c-681e9defe6fb · outbound

This paper cites GPT-4 Technical Report.

MOVi: Training-free Text-conditioned Multi-Object Video Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.424420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.424420Z digest=sha256:5b1053c571753f36bd625a4c7f05b932edbf0a7486f329d3f14e653deb9b2b7d

Observation 1d82a762-a827-434e-bece-cc68cfb7d6f8 · outbound

This paper cites Frozen in time: A joint video and im- age encoder for end-to-end retrieval.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Frozen in time: A joint video and im- age encoder for end-to-end retrieval

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:11.203309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:56.521401Z digest=sha256:dd92a9308f98281aadccf0b0d8ad79c7758362a249c5fd41e05cdee7a4f2a56e

Observation be240a50-c82f-4585-9e23-60e5a9698fa5 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.611108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.611108Z digest=sha256:9349d94911d715fb475174c03956817aa6f65ce5a153c44ed8f6a35c8d76c44d

Observation 43d0e7af-36c5-4746-99e5-dd6195be5d4a · outbound

This paper cites Multidiffusion: Fusing diffusion paths for con- trolled image generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Multidiffusion: Fusing diffusion paths for con- trolled image generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:10.845129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:56.724522Z digest=sha256:730efbca3cf0df7b6036f42a70605330dcdd94e5288ca44021850d8e70207c9c

Observation 842663ce-81b1-4b8c-8aa1-fa07edf8a6ab · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.872776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.872776Z digest=sha256:df8203fc0f697305341f2c4e72bba01b0ad7e5ec619e109221dc2a564cab16ee

Observation 36d37ad9-256f-4803-b3a0-9f2618fed3e0 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Align your latents: High-resolution video synthesis with latent diffusion models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:10.495466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:56.970597Z digest=sha256:7e32bbb9ec2245b96009282d0c9d4ef53bb60d18b5a809669abbce95f373aa9c

Observation a7ccfdea-b328-485b-9aa7-2f6d59470f66 · outbound

This paper cites VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.039138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.039138Z digest=sha256:7e3344fa13b0c831bcac5d6869612726c881e4c58b91a3f57b26eb5c90c87109

Observation 9b6071c2-fd71-4667-990e-1f617c379fd0 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high- quality video diffusion models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Videocrafter2: Overcoming data limitations for high- quality video diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:10.125256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:57.143508Z digest=sha256:17211ec1d819f3cf9bad64cd216694ca3f80a20bd112d2974177833832539f07

Observation b335be6a-f2fe-4445-abf3-c8838c6ffc80 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:09.819009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:57.188341Z digest=sha256:e2240b733b66c50f1c43168d3d118adf8d9f00ecf1e6b82d727e03d675e09c0a

Observation 86692de2-767f-4608-8f4a-42cd3c20b68f · outbound

This paper cites Sora as an agi world model? a complete survey on text-to-video generation.arXiv preprint arXiv:2403.05131, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Sora as an agi world model? a complete survey on text-to-video generation.arXiv preprint arXiv:2403.05131, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.269517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.269517Z digest=sha256:ddc3186a2543516e2a681e3e7e0bcf284dfbd05f395283ed087eff1e3c7eceed

Observation 9ccaa405-75ef-4749-abaa-783b7ae51a71 · outbound

This paper cites Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:09.551068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:57.433196Z digest=sha256:8cc59864bf4e51c92853385173ce455de41a369131f4444503aa3016b232076b

Observation 9b25e274-37d8-4f14-9df1-b3ea4c4abc93 · outbound

This paper cites DiffSynth-Studio: Enjoy the magic of Diffusion models!https://github.com/ modelscope/DiffSynth-Studio, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation DiffSynth-Studio: Enjoy the magic of Diffusion models!https://github.com/ modelscope/DiffSynth-Studio, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:09.166354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:57.582602Z digest=sha256:da27eeff2a43d42b0c2c691b909c1f38b04278b9a8c20df883bbb4c71c60c3ba

Observation 3935d16d-c838-4c6e-a3c0-37efa2f38078 · outbound

This paper cites Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.International Conference on Learn- ing Representations, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.International Conference on Learn- ing Representations, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.832656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:57.717505Z digest=sha256:59a8f081ccd904b562606dfb6fc7178020893aa57a303077ce0ba5313585fedc

Observation 37fc4c8d-8a6f-458c-a9a0-ec81da22db58 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.842598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.842598Z digest=sha256:e389c17ca4221564da7a360fcd93b5ef334f4df62184c2e2c0563b4d9bda41e4

Observation 6f32561d-1157-4416-8631-f864ad6099ec · outbound

This paper cites CLIPScore: a reference- free evaluation metric for image captioning.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CLIPScore: a reference- free evaluation metric for image captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.484965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:57.985309Z digest=sha256:da5482e3b14d6198e564a4e2eb855634d7ee018f9e1aa277f38993c11a332332

Observation d1933484-4e79-4734-841f-0027a4878843 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.093105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.093105Z digest=sha256:6d9e21f5ed4c065c8d160c29ac9b034ac74fedc327ef5087e6d638d22115f495

Observation 33db6231-caa2-4980-810c-6cb240c89537 · outbound

This paper cites Denois- ing diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Denois- ing diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.197779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.197779Z digest=sha256:2d59542cb8f5ae15a6bbc9302c3a00c9de59f1cbffae3c30b439c62498adb980

Observation 9a3b2a17-96da-4b8b-b8e0-5a45bb38fc4c · outbound

This paper cites Video Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Video Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.322441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.322441Z digest=sha256:e5d379e256255a5fbac80a2637dbdcd3d7f2c40d17e96959579ed39b5a29a7ad

Observation caa1bd13-a401-4165-ba1b-c5537d101575 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.427430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.427430Z digest=sha256:d549fd0e9102d8fde075638d75390ceeff25938223d28e722c0ffc4d0723d63d

Observation 50e588c3-2169-4d3c-acd8-eae01d532af9 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Vbench: Comprehensive benchmark suite for video generative models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.315535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:58.539695Z digest=sha256:40a61e39d5477d8170f553a31f74e4a73f240263e4184e856aea97312adda081

Observation 6015e7c6-9ffa-4ff7-b1aa-9ed77f66ca46 · outbound

This paper cites High-quality Text-to-video Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation High-quality Text-to-video Models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.307585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:58.658289Z digest=sha256:c80586b11fb7e6c629d1c3082fe87cf076d7f00d7046a77a133bd38fb9b760ea

Observation b974e219-4fb0-4c1e-ae84-95e7459971bd · outbound

This paper cites KLING AI: Next-Generation AI Creative Studio.https://www.klingai.com/, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation KLING AI: Next-Generation AI Creative Studio.https://www.klingai.com/, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.180073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:58.748386Z digest=sha256:6a86661160ad28d5f2a43f8ba3b9c7aa9a726c82dec69a922f162448e7742885

Observation 272c985e-138a-496a-b615-57a1b4732816 · outbound

This paper cites Multi-concept cus- tomization of text-to-image diffusion.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Multi-concept cus- tomization of text-to-image diffusion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.862669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:58.858770Z digest=sha256:75b1724f27425f21d2a6fc06de7377df3f4cf185efea7aeb50762221fa23e59c

Observation 2233dae1-8b17-4b73-b21d-c9aca587af4f · outbound

This paper cites TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.952626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.952626Z digest=sha256:5f6b618906484c7e05e08f2c2a3a4e94f86e54c829e9e7883791c895a9aec1d7

Observation be3c788d-8ee8-4078-a85b-caa23787f32f · outbound

This paper cites Gligen: Open-set grounded text-to- image generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Gligen: Open-set grounded text-to- image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.638758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.025274Z digest=sha256:a62a1219ca5c53c6bfac6e3eb234eafcc0fe760c23d9bcea3f605cc1c0a7335a

Observation 80044332-1829-4001-9a49-611707efe2e6 · outbound

This paper cites Movideo: Motion-aware video generation with diffusion model.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Movideo: Motion-aware video generation with diffusion model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.496035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.091672Z digest=sha256:a0cb37f9ccb00af17bd9cdafd62737da7a62f683e403cafa13af1844017c81be

Observation ff403e82-cb7a-4ac2-a7ee-489674d20e87 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.343576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.172626Z digest=sha256:d624f3222abf237b3af69ce4028b9d625d0f08cc8b59409d21d68d55db0e2630

Observation a634ebcf-1c1d-4882-9502-b3558c670a96 · outbound

This paper cites Detector Guidance for Multi-Object Text-to-Image Generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Detector Guidance for Multi-Object Text-to-Image Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:01:03.014410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.268187Z digest=sha256:0bb8f3c3f4d580a45444e6ddec3d389d44cac769e01562cade297023204ddb99

Observation be93cda6-eecb-4039-958a-99d6463f79db · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.360313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.360313Z digest=sha256:0fe9ea3ade019d7572458d2745d43c9dc42c6f7b22323a5c12c2b50b15269b4f

Observation 4ab781fb-589e-4f9b-93fb-8ee640cf5507 · outbound

This paper cites Lumaai.https://lumalabs.ai/,.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Lumaai.https://lumalabs.ai/,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.181827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.480131Z digest=sha256:b94b5e42a334ec5660e39a79090873d185c648de0600e555df5f2d47fd83eabb

Observation 32ae22fd-4d49-4c6d-ad5c-14872d010344 · outbound

This paper cites Vidm: Video implicit diffusion models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Vidm: Video implicit diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.675171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.677588Z digest=sha256:c38c3afcebe6a229329aaa06430ec6618eabca31255fffaebceb390948634484

Observation 37de2084-99d7-4691-bb74-cfadb942a24d · outbound

This paper cites Hailuo AI: Captivating AI Videos Gen- erated with Hailuo AI .https://hailuoai.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Hailuo AI: Captivating AI Videos Gen- erated with Hailuo AI .https://hailuoai

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.512260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.744083Z digest=sha256:45f39ca60cdd7fd05d210e5c918c41d89ee5c2a40bc8abf598a1c2d609f9ffd1

Observation 00d20cf3-5a94-4aab-a214-7fbe8a35b8a0 · outbound

This paper cites Dreamix: Video Diffusion Models are General Video Editors.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Dreamix: Video Diffusion Models are General Video Editors

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.801514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.801514Z digest=sha256:4aab56c2e292089e9f94ad4f43b361db8b6854db615a8bbfd9c50c433038fcca

Observation ed0700fb-40f1-46c2-b9df-8dc13412dc53 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

MOVi: Training-free Text-conditioned Multi-Object Video Generation WorldSimBench: Towards Video Generation Models as World Simulators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.855458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.855458Z digest=sha256:2d9fa795f5a076c9f868c89ba9614ca877871b0f8ed37a9ff13eb08c63a8bdd8

Observation ccd886c3-2fdb-4825-9188-ee47f941fb16 · outbound

This paper cites FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.932469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.932469Z digest=sha256:5f9dc928e54886174c4a27019ca5ae07314fbabd81df501b8af9ee4e15075228

Observation b61f693c-65c2-4189-8e45-07b895997203 · outbound

This paper cites High- resolution image synthesis with latent diffusion mod- els.

MOVi: Training-free Text-conditioned Multi-Object Video Generation High- resolution image synthesis with latent diffusion mod- els

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.991815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.991815Z digest=sha256:0a2f8806ccef2ba04b4643a86c274ebf82d30c6494d59e5d213a59c827e39164

Observation 077b15c3-383c-4714-9e5e-ec3e29d03024 · outbound

This paper cites Gen-2: Generate novel videos with text, im- ages or video clips, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Gen-2: Generate novel videos with text, im- ages or video clips, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.325296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:00.054526Z digest=sha256:393f8761bc36542c317aec75d0db65237a2e3c33b181b6b4e235e5a28d262820

Observation 0d5210f5-31b9-4dd7-9af0-34f0764c78a1 · outbound

This paper cites Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.166910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:00.104985Z digest=sha256:325d48a7374b16ff7c30db70ec84ffc64425db5af24b27aadddd755f2e7d546c

Observation 55510d57-fa72-422a-81f8-f760346fd4d4 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.153544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.153544Z digest=sha256:08672e40960e671cd8f905271f9359ffd0fd6d80e3108047ec74ad72c6675d2c

Observation d567d424-2a2d-472e-ab42-d20348480375 · outbound

This paper cites Denoising Diffusion Implicit Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Denoising Diffusion Implicit Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.214571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.214571Z digest=sha256:78108c43a3029905ade343f7cf98849d16cd00a31345e87c611b7dcde7062ca6

Observation b94524fd-15c5-4b14-a04b-404da42a760b · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MOVi: Training-free Text-conditioned Multi-Object Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.283843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.283843Z digest=sha256:bf8cc5856f38da199076501cb283b0feff75932915aa4f5d234a2c5dc07f099a

Observation cf967497-233e-4b07-818c-d3b7f7e4bfe7 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Raft: Recurrent all-pairs field transforms for optical flow

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.011982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:00.331605Z digest=sha256:fe225007c9af2d604635621ede1d00d84130b0e4d5c97c37b2a8892c1dd86197

Observation d42c452d-9d30-4735-b207-c73d02727e5c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.389339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.389339Z digest=sha256:57fefce64348fceb3f521f1d7d24671b07ff3c5073d087ab062c26032493545c

Observation 9c91d916-d50f-466b-ae54-23f87b483035 · outbound

This paper cites Fvd: A new metric for video generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Fvd: A new metric for video generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.828168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:00.436115Z digest=sha256:b9c5744468881dce5c8640d5f562a6a36f85dede326cad80591bff96185c839f

Observation c9b98d0c-7c9f-4cf4-8689-727ac648b204 · outbound

This paper cites Vchitect 2.0: Embark on a Visual Fan- tasy Journey.https://vchitect.intern- ai.org.cn/, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Vchitect 2.0: Embark on a Visual Fan- tasy Journey.https://vchitect.intern- ai.org.cn/, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.693951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:00.504131Z digest=sha256:ddf07fedc2b4ed6cfacce46e6676aab6ea22b21b1ba34aa7aa35381aa1964181

Observation 7728518b-2d64-4a24-aa08-5fe07fb488c1 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.551776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.551776Z digest=sha256:0324f203d6c9c1ed37b736a15a50625cfe9a5db3579f57418e32dd5426428e59

Observation b92d328c-fdd7-4fc2-8881-e6dc1d5cc576 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

MOVi: Training-free Text-conditioned Multi-Object Video Generation ModelScope Text-to-Video Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.629908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.629908Z digest=sha256:a853b18596fb3809817e555efa842aa4a4f9576c55fbb596b0bb69352c187e25

Observation 76bd7851-616f-47fe-a686-038e71947e91 · outbound

This paper cites Boximator: Generating Rich and Controllable Motions for Video Synthesis.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Boximator: Generating Rich and Controllable Motions for Video Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.698486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.698486Z digest=sha256:facf641bdf1debf621c028cde6eda279a4ac2496f3716c6f94b05197ea2bee3b

Observation 0bd7da98-74db-4435-9e17-d1bf73de581c · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVLM: Visual Expert for Pretrained Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.755454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.755454Z digest=sha256:a60049a0f2f1f075a38e4db329bcd619da17fa035e192c509ba587f1e9c37e4f

Observation 77514227-2a4a-46c9-a9b4-dbef5600b2ae · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Emu3: Next-Token Prediction is All You Need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.803646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.803646Z digest=sha256:8a4ee04b62a568dc6acac6be798665456b187a2173708cdc9f8ad33546c4d2e1

Observation 5527bad4-bd6c-4d2a-a004-2ab86e607bdb · outbound

This paper cites WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens.

MOVi: Training-free Text-conditioned Multi-Object Video Generation WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.850247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.850247Z digest=sha256:fff9e006037130a14bc5794a480182d5e81fb521a934b919b93a82951b575873

Observation 5b73f522-f386-48ad-8fa6-f7b34042c99f · outbound

This paper cites LaVie: High-quality video generation with cascaded latent diffusion mod- els.IJCV, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation LaVie: High-quality video generation with cascaded latent diffusion mod- els.IJCV, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.494450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:00.900024Z digest=sha256:c78c8298dbf23eda4f89858e4245ced5acd937a107ab64b4418386792db4c0d3

Observation 0969d8ed-1fb6-4f3b-99cf-9b277420cd86 · outbound

This paper cites Customvideo: Cus- tomizing text-to-video generation with multiple sub- jects.arXiv preprint arXiv:2401.09962, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Customvideo: Cus- tomizing text-to-video generation with multiple sub- jects.arXiv preprint arXiv:2401.09962, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.965204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.965204Z digest=sha256:94a1c0aff267a8bad7e93ccfd2ea1b9f0db2c0caf7a421cf7b91ca40b6403c43

Observation 25ef5394-9984-4e2d-ba69-6861e8466566 · outbound

This paper cites Grit: A generative region-to-text transformer for ob- ject understanding.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Grit: A generative region-to-text transformer for ob- ject understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.242780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:01.016503Z digest=sha256:3453335d89d47d58f4687c3d04c5490fcdaeddb636ef978f0ab4e242d9b49d38

Observation a5b9fce2-6191-40d8-9afd-ef79d6c2815a · outbound

This paper cites FreeInit: Bridging Initialization Gap in Video Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation FreeInit: Bridging Initialization Gap in Video Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:01.085135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:01.085135Z digest=sha256:9790e4a9da49e846d8f80c1392034e349e0ae224f3ab2044a1b933257ba40b90

Observation 81dc84f9-3b18-4b54-ba9f-d8405d3a5837 · outbound

This paper cites Dynamicrafter: An- imating open-domain images with video diffusion pri- ors.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Dynamicrafter: An- imating open-domain images with video diffusion pri- ors

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.971631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:01.159014Z digest=sha256:8b4d19c9c36d99d666b163ef808f08a963df7a24a9bfc116e3f035962a78390e

Observation 591e9272-6616-4805-bcaa-020b06632107 · outbound

This paper cites MSR- VTT: A large video description dataset for bridging video and language.

MOVi: Training-free Text-conditioned Multi-Object Video Generation MSR- VTT: A large video description dataset for bridging video and language

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.732407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:01.271572Z digest=sha256:de2306cd3423c679f6d7c970e4ac14e9e90e5f175b65f66d5d1153ea11d72650

Observation 5dd678f3-210e-4785-8b6f-57f08c563b28 · outbound

This paper cites Advancing high-resolution video- language representation with large-scale video tran- scriptions.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Advancing high-resolution video- language representation with large-scale video tran- scriptions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.458872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:01.363298Z digest=sha256:35ba559641859a07ed73ea468a4fc8a12debb027065941d289d9d740d7916114

Observation bdfe9b52-8ebb-4196-94bb-615dfa0b7647 · outbound

This paper cites Video in- stance segmentation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Video in- stance segmentation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.185001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:01.481638Z digest=sha256:3ec6b4e814af727b6670480a5aeb7156663b1fc2916a5ba27f7fcd810b3d220a

Observation dc07aed5-accc-4f00-9c2c-b2af7ecc7e2a · outbound

This paper cites EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing.

MOVi: Training-free Text-conditioned Multi-Object Video Generation EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:01.607138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:01.607138Z digest=sha256:5b64ae8ec01811e721e4eadf83ccf012bfed08d9f3b82ccd7027dfd8ff9036f3

Observation b7e031dc-3a43-4f2c-95a1-ce678daf7a23 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:01.724034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:01.724034Z digest=sha256:d7013423e40ad46886cd953cdfbf025a852cae1fb1b4568caf4f33103f5e1114

Observation 1095eefb-eac3-49e4-a793-9812637b7ad5 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.Interna- tional Journal of Computer Vision, pages 1–15, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Show-1: Marrying pixel and latent diffusion models for text-to-video generation.Interna- tional Journal of Computer Vision, pages 1–15, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:03.849630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:01.876054Z digest=sha256:ccac4759004b994d4d6f8c73b8870342b69e35f5dafb3d91d640b50b8d2e9e48

Observation 274b38f8-d867-4814-bdab-22d674d6b567 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:02.040690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:02.040690Z digest=sha256:00fc89c4913f1c5f5d0c94f5ec78c35511b9175c8d108ebbc69889eb3b245b66

Observation 4a86a191-9976-4ebc-9f37-0abae1793df8 · outbound

This paper cites Real-time vehicle detection based on improved yolo v5.Sustainability, 14(19):12274, 2022.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Real-time vehicle detection based on improved yolo v5.Sustainability, 14(19):12274, 2022

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:03.624415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:02.166347Z digest=sha256:b0fce247477ff76b6dc7b6fe313d2a2e2c8bf716f23b736e5ec88330982a9723

Observation 2111eff1-6b07-4239-8e8b-1b32f1cb52c7 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, March 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Open-sora: Democratizing efficient video production for all, March 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:03.398534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:01:02.307369Z digest=sha256:e55aab45f685f5fd82366a11fb16ac94fd0f3f812e6bbdaeb2fe6d5be1569b54

Observation 808ab046-08ce-426e-8557-ccebdeab53d0 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:02.423170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:02.423170Z digest=sha256:3d376f1d972424ee4097ac7420d37776c955c764970fbb3192817e72e40f5b25

Observation 53331444-f74f-4ea2-b685-554cd1e63306 · outbound

This paper cites 2, 5, 7, 8.

MOVi: Training-free Text-conditioned Multi-Object Video Generation 2, 5, 7, 8

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.966840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.590268Z digest=sha256:42f28695cc511f347f4b26ad4bdb843a6261f675173be2f2f878629e7b7d3f29

Pith citing papers

No inbound Pith citation observations are available.