Pith. sign in

Paper Citation Record · LEDGER

MOVi: Training-free Text-conditioned Multi-Object Video Generation

As of 9 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2505.22980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22980 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:01:02.423170Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c51cfaf-ecdf-40d7-ad8c-681e9defe6fb · outbound

This paper cites GPT-4 Technical Report.

MOVi: Training-free Text-conditioned Multi-Object Video Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.424420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.424420Z digest=sha256:7b3b197a3dc12c03d911fefe34699ece6288c52280ecbe5bb2881382fc1b4ed2

Observation 1d82a762-a827-434e-bece-cc68cfb7d6f8 · outbound

This paper cites Frozen in time: A joint video and im- age encoder for end-to-end retrieval.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Frozen in time: A joint video and im- age encoder for end-to-end retrieval

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:11.203309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:56.521401Z digest=sha256:c4027d4ff0e477a2168a60555cc1048d45d076e53f25505a514c847a09fffd76

Observation be240a50-c82f-4585-9e23-60e5a9698fa5 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.611108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.611108Z digest=sha256:1848c0a403875f83746d682926b2fb9b7fbbbd346e325750e16dae0ff5048885

Observation 43d0e7af-36c5-4746-99e5-dd6195be5d4a · outbound

This paper cites Multidiffusion: Fusing diffusion paths for con- trolled image generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Multidiffusion: Fusing diffusion paths for con- trolled image generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:10.845129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:56.724522Z digest=sha256:26d2d79bec8bd0ccd9a73fb9d86aba49fdb65c880cd3ca0aabea7785d5990771

Observation 842663ce-81b1-4b8c-8aa1-fa07edf8a6ab · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.872776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.872776Z digest=sha256:646db918c99d54897443e6fb15225b80234f254745a380ae080aa87aa9b3da5a

Observation 36d37ad9-256f-4803-b3a0-9f2618fed3e0 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Align your latents: High-resolution video synthesis with latent diffusion models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:10.495466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:56.970597Z digest=sha256:587b989242c714b71142ea91157289349222815fb6633a1a93cec8c710e41a86

Observation a7ccfdea-b328-485b-9aa7-2f6d59470f66 · outbound

This paper cites VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.039138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.039138Z digest=sha256:1316934025d78959de8d85ababb08220bf9b3f9df0c8f9534a6a68627d85088a

Observation 9b6071c2-fd71-4667-990e-1f617c379fd0 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high- quality video diffusion models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Videocrafter2: Overcoming data limitations for high- quality video diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:10.125256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:57.143508Z digest=sha256:49b770b8ca7b77e36e411ebf44677f8a14ca8cfb37f76d406128e63169fe30a1

Observation b335be6a-f2fe-4445-abf3-c8838c6ffc80 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:09.819009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:57.188341Z digest=sha256:a79f4d910a14b3ec466701baf47c3437f7daef6d9a042fcef6a2cbeb714b90b9

Observation 86692de2-767f-4608-8f4a-42cd3c20b68f · outbound

This paper cites Sora as an agi world model? a complete survey on text-to-video generation.arXiv preprint arXiv:2403.05131, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Sora as an agi world model? a complete survey on text-to-video generation.arXiv preprint arXiv:2403.05131, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.269517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.269517Z digest=sha256:d04fc37231809ba2c37783e1f3b13e671b195cdfa8d6ad722915b1c690490fc6

Observation 9ccaa405-75ef-4749-abaa-783b7ae51a71 · outbound

This paper cites Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:09.551068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:57.433196Z digest=sha256:7a4d7d3da3454ecc2e6c2edf6297781a9cbb8d166fbe2533beaa10d5fd85e9bb

Observation 9b25e274-37d8-4f14-9df1-b3ea4c4abc93 · outbound

This paper cites DiffSynth-Studio: Enjoy the magic of Diffusion models!https://github.com/ modelscope/DiffSynth-Studio, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation DiffSynth-Studio: Enjoy the magic of Diffusion models!https://github.com/ modelscope/DiffSynth-Studio, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:09.166354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:57.582602Z digest=sha256:33955ac7dbf91f5c97ab5fefffce962977d1c7b6d6766e13f407d4e72f8b3949

Observation 3935d16d-c838-4c6e-a3c0-37efa2f38078 · outbound

This paper cites Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.International Conference on Learn- ing Representations, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.International Conference on Learn- ing Representations, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.832656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:57.717505Z digest=sha256:ce53a4e8d8256ad9ea833c07fec15c6b51ea294c03c2d3940e69a3b7a23fd721

Observation 37fc4c8d-8a6f-458c-a9a0-ec81da22db58 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.842598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.842598Z digest=sha256:ab9b42a62f26f105b3efeeaa0a9c72c10ce37cdca8de34cdf81f9f07a0a7468d

Observation 6f32561d-1157-4416-8631-f864ad6099ec · outbound

This paper cites CLIPScore: a reference- free evaluation metric for image captioning.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CLIPScore: a reference- free evaluation metric for image captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.484965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:57.985309Z digest=sha256:4823bd62942a243fa9c67dd5d06ae4718ede129c76be3c8916b6d6d384e1aeee

Observation d1933484-4e79-4734-841f-0027a4878843 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.093105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.093105Z digest=sha256:54a33597cede4c572407896065bdbbcf62b68e0a5354728a6c7a1d4962425d5b

Observation 33db6231-caa2-4980-810c-6cb240c89537 · outbound

This paper cites Denois- ing diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Denois- ing diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.197779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.197779Z digest=sha256:d1955f35cde3134781e77b3735e8615b8c9d9884e860d12c4f4fed9029ab4570

Observation 9a3b2a17-96da-4b8b-b8e0-5a45bb38fc4c · outbound

This paper cites Video Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Video Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.322441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.322441Z digest=sha256:b884243f608c0136631acaa63085c40163fe7822fc880ee0394f775a08e74fdb

Observation caa1bd13-a401-4165-ba1b-c5537d101575 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.427430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.427430Z digest=sha256:5d7a10d07460509521e37aed23090cd577cba2a9c91bfe047bb106d8c0a33a3f

Observation 50e588c3-2169-4d3c-acd8-eae01d532af9 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Vbench: Comprehensive benchmark suite for video generative models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.315535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:58.539695Z digest=sha256:af176fd70a0f84e12bd77739279b7024f84cfa0be403579995d47c38ebd3d9a6

Observation 6015e7c6-9ffa-4ff7-b1aa-9ed77f66ca46 · outbound

This paper cites High-quality Text-to-video Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation High-quality Text-to-video Models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.307585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:58.658289Z digest=sha256:d347851116044e4fae1af4da342edf8c8765d5fa8671e3aa266c85295cc185e1

Observation b974e219-4fb0-4c1e-ae84-95e7459971bd · outbound

This paper cites KLING AI: Next-Generation AI Creative Studio.https://www.klingai.com/, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation KLING AI: Next-Generation AI Creative Studio.https://www.klingai.com/, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.180073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:58.748386Z digest=sha256:b1cd7f08c1ab84128a674a4a6aed35be62a336327c6594fff29df91a8b9493d4

Observation 272c985e-138a-496a-b615-57a1b4732816 · outbound

This paper cites Multi-concept cus- tomization of text-to-image diffusion.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Multi-concept cus- tomization of text-to-image diffusion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.862669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:58.858770Z digest=sha256:a1fab797f4a9a3acb0e98c2546eecc1552c6b9a74c59568aadbaaac5fedb07d7

Observation 2233dae1-8b17-4b73-b21d-c9aca587af4f · outbound

This paper cites TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.952626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.952626Z digest=sha256:d2edafa22b913da7845a6fbd1ee33745149a86e71416792d01da1ab8dcb8d3e5

Observation be3c788d-8ee8-4078-a85b-caa23787f32f · outbound

This paper cites Gligen: Open-set grounded text-to- image generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Gligen: Open-set grounded text-to- image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.638758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:59.025274Z digest=sha256:b95560990eba1b0b88dc971c1bcd1bd5f5b9beb26a8566d4bbc826511b4dcf99

Observation 80044332-1829-4001-9a49-611707efe2e6 · outbound

This paper cites Movideo: Motion-aware video generation with diffusion model.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Movideo: Motion-aware video generation with diffusion model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.496035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:59.091672Z digest=sha256:4d2258e62da44a8f5db6117c22442942ddc58ff87e3a760bffa09fcb32e7249a

Observation ff403e82-cb7a-4ac2-a7ee-489674d20e87 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.343576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:59.172626Z digest=sha256:fdd020ba1b0b5c12d07691a95f8ccafb7e4568b93570396405ebc9e0a5ce46c1

Observation a634ebcf-1c1d-4882-9502-b3558c670a96 · outbound

This paper cites Detector Guidance for Multi-Object Text-to-Image Generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Detector Guidance for Multi-Object Text-to-Image Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:01:03.014410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:59.268187Z digest=sha256:741e4d7c897ee4b43f307aef5927ddef73c1cd4701e9848f22d31f1f61c9c8d0

Observation be93cda6-eecb-4039-958a-99d6463f79db · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.360313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.360313Z digest=sha256:5449bf76b753ab752f976085faef9cba395ecc989a690177a3cb4b193482c218

Observation 4ab781fb-589e-4f9b-93fb-8ee640cf5507 · outbound

This paper cites Lumaai.https://lumalabs.ai/,.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Lumaai.https://lumalabs.ai/,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.181827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:59.480131Z digest=sha256:4fe93284199a1d5bb68512fcecf026c27799137941b5245befa14e000287a747

Observation 32ae22fd-4d49-4c6d-ad5c-14872d010344 · outbound

This paper cites Vidm: Video implicit diffusion models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Vidm: Video implicit diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.675171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:59.677588Z digest=sha256:fbf82742711066aac9b12d5a6728d5ead18d7f90ffd7e14b9c6244d93b741161

Observation 37de2084-99d7-4691-bb74-cfadb942a24d · outbound

This paper cites Hailuo AI: Captivating AI Videos Gen- erated with Hailuo AI .https://hailuoai.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Hailuo AI: Captivating AI Videos Gen- erated with Hailuo AI .https://hailuoai

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.512260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:59.744083Z digest=sha256:698c7659ca98b942e6f4cee8b83befb14cfa686584caaff9821366db3fdef331

Observation 00d20cf3-5a94-4aab-a214-7fbe8a35b8a0 · outbound

This paper cites Dreamix: Video Diffusion Models are General Video Editors.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Dreamix: Video Diffusion Models are General Video Editors

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.801514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.801514Z digest=sha256:fd2d4443043b305750cb6ea814a44a829daea5f5f6c032781aee1e414fb8b051

Observation ed0700fb-40f1-46c2-b9df-8dc13412dc53 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

MOVi: Training-free Text-conditioned Multi-Object Video Generation WorldSimBench: Towards Video Generation Models as World Simulators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.855458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.855458Z digest=sha256:7e57410b705b5bdfdfbec9fc5947841fc8169b35f39dd482c6e0de921637db8b

Observation ccd886c3-2fdb-4825-9188-ee47f941fb16 · outbound

This paper cites FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.932469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.932469Z digest=sha256:7af648d9e5eabf306cb0525a7e0421595839fbae34d5dc55cc605e8ce389a3ca

Observation b61f693c-65c2-4189-8e45-07b895997203 · outbound

This paper cites High- resolution image synthesis with latent diffusion mod- els.

MOVi: Training-free Text-conditioned Multi-Object Video Generation High- resolution image synthesis with latent diffusion mod- els

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.991815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.991815Z digest=sha256:88c3aba80fae4e04e60a1415fc8f44a4008c320188a6576438ed98372648e2ce

Observation 077b15c3-383c-4714-9e5e-ec3e29d03024 · outbound

This paper cites Gen-2: Generate novel videos with text, im- ages or video clips, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Gen-2: Generate novel videos with text, im- ages or video clips, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.325296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:00.054526Z digest=sha256:ae64916be64b6f1910700c48033fa58747b9060881b523f4ee20e5408782734d

Observation 0d5210f5-31b9-4dd7-9af0-34f0764c78a1 · outbound

This paper cites Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.166910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:00.104985Z digest=sha256:39addcce0d6ddad2beb34edc60eb38472dc7383c7a5dae962553443c8ae5f7d1

Observation 55510d57-fa72-422a-81f8-f760346fd4d4 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.153544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.153544Z digest=sha256:c67c07c14d6f9a66c40b097a3176688c41059b55dee0f4bd37560c0f1eef180b

Observation d567d424-2a2d-472e-ab42-d20348480375 · outbound

This paper cites Denoising Diffusion Implicit Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Denoising Diffusion Implicit Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.214571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.214571Z digest=sha256:b091c7571921b81f368bd3ce4209599bb99bc4d949642d754d2c172838feca55

Observation b94524fd-15c5-4b14-a04b-404da42a760b · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MOVi: Training-free Text-conditioned Multi-Object Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.283843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.283843Z digest=sha256:05cc75236f75baffaaf2c8b6eec1c0a7db55d1837c395a15039496bd157a0798

Observation cf967497-233e-4b07-818c-d3b7f7e4bfe7 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Raft: Recurrent all-pairs field transforms for optical flow

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.011982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:00.331605Z digest=sha256:e837d96b4810dd235698eab829f4927b1c05b9666d178c11adb214e26256391f

Observation d42c452d-9d30-4735-b207-c73d02727e5c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.389339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.389339Z digest=sha256:9f223a1e3b6c991a4c83af1b688307e8d4431680f2bbb6985a47c5e9601f09e2

Observation 9c91d916-d50f-466b-ae54-23f87b483035 · outbound

This paper cites Fvd: A new metric for video generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Fvd: A new metric for video generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.828168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:00.436115Z digest=sha256:199f5061c90ef51490198179b9c1af1f5600a634a6aee7efe14c9793ed3857e9

Observation c9b98d0c-7c9f-4cf4-8689-727ac648b204 · outbound

This paper cites Vchitect 2.0: Embark on a Visual Fan- tasy Journey.https://vchitect.intern- ai.org.cn/, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Vchitect 2.0: Embark on a Visual Fan- tasy Journey.https://vchitect.intern- ai.org.cn/, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.693951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:00.504131Z digest=sha256:b298f4e89dc5d3cbef0bf3ac7ef48e62ddfd4b6cec4d2c31000fd9e9cf664891

Observation 7728518b-2d64-4a24-aa08-5fe07fb488c1 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.551776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.551776Z digest=sha256:30fccd32d104c3523dd625c1582867b9173da90c3a8409e4979e9820403691fa

Observation b92d328c-fdd7-4fc2-8881-e6dc1d5cc576 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

MOVi: Training-free Text-conditioned Multi-Object Video Generation ModelScope Text-to-Video Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.629908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.629908Z digest=sha256:be5bc2f93af93fa591def63d9880b9b9ca24477589d7f545565b025787cb525d

Observation 76bd7851-616f-47fe-a686-038e71947e91 · outbound

This paper cites Boximator: Generating Rich and Controllable Motions for Video Synthesis.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Boximator: Generating Rich and Controllable Motions for Video Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.698486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.698486Z digest=sha256:e691f2e334c9d69e05d20417161f7568046ebeb3e7bdfea082682954f9305722

Observation 0bd7da98-74db-4435-9e17-d1bf73de581c · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVLM: Visual Expert for Pretrained Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.755454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.755454Z digest=sha256:67f77fc7f77c1f5ad563faebd6bfcae57956a53df051e2787533f22277b274b1

Observation 77514227-2a4a-46c9-a9b4-dbef5600b2ae · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Emu3: Next-Token Prediction is All You Need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.803646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.803646Z digest=sha256:a6138dc571f3dc1f0529512454cbf4ffd72b2368547e16c4172b2d7ea2b91114

Observation 5527bad4-bd6c-4d2a-a004-2ab86e607bdb · outbound

This paper cites WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens.

MOVi: Training-free Text-conditioned Multi-Object Video Generation WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.850247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.850247Z digest=sha256:f33103ea27c9a12320af1d1e5739e8df5ad1a9de62c15d75ea243abfa69fa6a6

Observation 5b73f522-f386-48ad-8fa6-f7b34042c99f · outbound

This paper cites LaVie: High-quality video generation with cascaded latent diffusion mod- els.IJCV, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation LaVie: High-quality video generation with cascaded latent diffusion mod- els.IJCV, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.494450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:00.900024Z digest=sha256:b1ba3367b8169458e0d5abd49e2e0ae2544c8245c1f62014ebd52286d8c4f1b3

Observation 0969d8ed-1fb6-4f3b-99cf-9b277420cd86 · outbound

This paper cites Customvideo: Cus- tomizing text-to-video generation with multiple sub- jects.arXiv preprint arXiv:2401.09962, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Customvideo: Cus- tomizing text-to-video generation with multiple sub- jects.arXiv preprint arXiv:2401.09962, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.965204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.965204Z digest=sha256:64b8809ea18a5b3e947773a3d7ea7bf6c0a36863892ac6580ce5392d03ba1b2a

Observation 25ef5394-9984-4e2d-ba69-6861e8466566 · outbound

This paper cites Grit: A generative region-to-text transformer for ob- ject understanding.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Grit: A generative region-to-text transformer for ob- ject understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.242780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:01.016503Z digest=sha256:90b91f85b56aa6a49711dcb35ec54345f330f68d278cf1f358afe37095ea6df1

Observation a5b9fce2-6191-40d8-9afd-ef79d6c2815a · outbound

This paper cites FreeInit: Bridging Initialization Gap in Video Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation FreeInit: Bridging Initialization Gap in Video Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:01.085135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:01.085135Z digest=sha256:6d0a8da19ec55048efe4e9c767e4b4d07486ae9f1a737228de172fd8791e9a7a

Observation 81dc84f9-3b18-4b54-ba9f-d8405d3a5837 · outbound

This paper cites Dynamicrafter: An- imating open-domain images with video diffusion pri- ors.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Dynamicrafter: An- imating open-domain images with video diffusion pri- ors

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.971631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:01.159014Z digest=sha256:ae85f1fbb576b6d97944ef3618ecc8093f7a8cc129fdfb7d5504dc95b58da169

Observation 591e9272-6616-4805-bcaa-020b06632107 · outbound

This paper cites MSR- VTT: A large video description dataset for bridging video and language.

MOVi: Training-free Text-conditioned Multi-Object Video Generation MSR- VTT: A large video description dataset for bridging video and language

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.732407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:01.271572Z digest=sha256:880193bb88cddd8fe0ca2838779de5c2bd3b96c57b6efc2e84b62aa2d1397b8e

Observation 5dd678f3-210e-4785-8b6f-57f08c563b28 · outbound

This paper cites Advancing high-resolution video- language representation with large-scale video tran- scriptions.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Advancing high-resolution video- language representation with large-scale video tran- scriptions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.458872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:01.363298Z digest=sha256:204750857155b1b799ce6a8306888b2557a0fd989f8f6f70055ee8362b559950

Observation bdfe9b52-8ebb-4196-94bb-615dfa0b7647 · outbound

This paper cites Video in- stance segmentation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Video in- stance segmentation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.185001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:01.481638Z digest=sha256:3217ae13f7df2b4e9ec884380235bff1b4a1284f06da8fd9a9d9a089ec5431bc

Observation dc07aed5-accc-4f00-9c2c-b2af7ecc7e2a · outbound

This paper cites EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing.

MOVi: Training-free Text-conditioned Multi-Object Video Generation EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:01.607138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:01.607138Z digest=sha256:77285dd3c503d8cd29fa0913036bf4d33847c70043a484de32495f791373570d

Observation b7e031dc-3a43-4f2c-95a1-ce678daf7a23 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:01.724034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:01.724034Z digest=sha256:784cd1566a3f556c9d63f96d0c71af6e05142e08af232b19cac8d33aea4820cd

Observation 1095eefb-eac3-49e4-a793-9812637b7ad5 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.Interna- tional Journal of Computer Vision, pages 1–15, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Show-1: Marrying pixel and latent diffusion models for text-to-video generation.Interna- tional Journal of Computer Vision, pages 1–15, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:03.849630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:01.876054Z digest=sha256:6dc96b9eec0c676f725d67354454d916e231aaa71bbdca1f1e0648ff399986b4

Observation 274b38f8-d867-4814-bdab-22d674d6b567 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:02.040690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:02.040690Z digest=sha256:4bc60106667cdba3fed70a887073a60f4caed9be8b82fe89aa7ed56d20bc1fe5

Observation 4a86a191-9976-4ebc-9f37-0abae1793df8 · outbound

This paper cites Real-time vehicle detection based on improved yolo v5.Sustainability, 14(19):12274, 2022.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Real-time vehicle detection based on improved yolo v5.Sustainability, 14(19):12274, 2022

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:03.624415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:02.166347Z digest=sha256:072f6582c33dc31e462d6b9cacc5949588acae96c6b9baa5a808ff05cf08fcf0

Observation 2111eff1-6b07-4239-8e8b-1b32f1cb52c7 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, March 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Open-sora: Democratizing efficient video production for all, March 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:03.398534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:01:02.307369Z digest=sha256:b1d9b6d9a14158150a545b43730c34f34343c20af9a192827a44cb9d9a76d8a9

Observation 808ab046-08ce-426e-8557-ccebdeab53d0 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:02.423170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:02.423170Z digest=sha256:a6675383771b3be69a483394a80846951652678e9387474e5767155cfc2babdc

Observation 53331444-f74f-4ea2-b685-554cd1e63306 · outbound

This paper cites 2, 5, 7, 8.

MOVi: Training-free Text-conditioned Multi-Object Video Generation 2, 5, 7, 8

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.966840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:00:59.590268Z digest=sha256:a5c658ee29e354c648f428565dad9648f7e7c4e100f2627101cdb7333d8e55cb

Pith citing papers

No inbound Pith citation observations are available.