Pith. sign in

Paper Citation Record · LEDGER

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

As of 17 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 8 inbound Pith citation observations for arXiv:2412.19645.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19645 v2

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:10:16.837253Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:11.282581Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:17:45.922106Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a38a04e0-f354-48cd-9c60-c74140685150 · outbound

This paper cites GPT-4 Technical Report.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.369176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.369176Z digest=sha256:bb61017e05c98a51692007ecb776c29cb3cb6ac1147fbead4cb4d5b87a8c0a56

Observation 5cc5ff6c-a0ad-4513-8942-cf887edf3699 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.374902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.374902Z digest=sha256:89e5e4edd98a98ab5741fbd6664230d50bddba9bf578956202c8f72819e09a13

Observation 0eeaea9c-1fc0-4dfc-8760-af1d379e6729 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.379761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.379761Z digest=sha256:4ddb0668b9370568c060a1db6861106ad8ad5716ac3ab2169661d3d059534d0f

Observation bbef9055-5c60-4e8e-81d9-b99265240545 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.386193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.386193Z digest=sha256:9bfb300709f17ffbbd2d9bd823b05645b2411ea6d649eec8e2c9715b77d2ada1

Observation 09283e1c-c635-4dd0-9158-78e16a594a45 · outbound

This paper cites Video generation models as world simulators.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Video generation models as world simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.391513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.391513Z digest=sha256:a76a34505a75f957efb9c5a0229fadfbff54fd988f2d67de052d2c5d4711b4a9

Observation c5c21e6c-514e-4ef1-9c27-383ad2395849 · outbound

This paper cites Still-Moving: Customized Video Generation without Customized Video Data.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Still-Moving: Customized Video Generation without Customized Video Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.397962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.397962Z digest=sha256:db84e855cfaa4b4ee87ee337369bcdcc7388fc4a8dab7c67eec863e5c714c0b7

Observation b4b17713-a035-4d86-9e7b-2af503e57094 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.403929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.403929Z digest=sha256:b1890a3bb5ce0d13ecf31bfa94af0b1c696e60f449f44bcf29ee661a8d707176

Observation 7f3f8124-c933-43db-b495-d4d4111e68b2 · outbound

This paper cites DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.410279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.410279Z digest=sha256:71b9d761f035bc4ae36376ee4dcdfcfd1abdc9ad188169900722f0a61208c01c

Observation 0e963efa-4742-4f6e-8f56-21a474c1b87c · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.416327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.416327Z digest=sha256:dc3ad469aa4420320198787473180edbc41434870b08bb879de85103050f6f09

Observation 810b07c9-614e-45aa-b5aa-aa1976bf3f79 · outbound

This paper cites Anydoor: Zero-shot object-level im- age customization.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Anydoor: Zero-shot object-level im- age customization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.606762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.422629Z digest=sha256:d4eecbce4c30a0a31425ed0631284b57e7df0f0ee1c04d8dc08399a506956f09

Observation 149af399-253d-475c-9a6b-949dc728daba · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Arcface: Additive angular margin loss for deep face recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.586386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.429397Z digest=sha256:52c36c1b1af3016819ddbb4fccd3f177cfb23a3ef6d6ceeb97a397de7e610910

Observation 2f8feda1-2a7a-4a50-8020-3974f688445e · outbound

This paper cites Gvd- iff: Grounded text-to-video generation with diffusion mod- els, 2024.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Gvd- iff: Grounded text-to-video generation with diffusion mod- els, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.567425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.434687Z digest=sha256:8ad6079f8b976bf2dde19e52689c57a6d62c79dfbfc976ab1ca8f45b462509cb

Observation 680ea694-98e0-4ac7-ae54-f81443dbb7d9 · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Structure and content-guided video synthesis with diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.544382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.439067Z digest=sha256:24e1986d42d1985658ca3f6067090ab4980a04961cd58a30f6600f9315f2cbe4

Observation 39e1dbc9-5d55-4b2d-8f01-7ad52784260e · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.444114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.444114Z digest=sha256:e7864a6e4d88af9aab0282e2dd5256b5d6509cb8fc1c58953dcba11ff9e86ca3

Observation 357bea52-1f0e-4b7f-bfbf-139406aeef81 · outbound

This paper cites Bermano, Gal Chechik, and Daniel Cohen-Or.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Bermano, Gal Chechik, and Daniel Cohen-Or

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.450325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.450325Z digest=sha256:98d1626853d2309efba3710d3b07dfbb99c18e2fc500d4222105ed9594411dbf

Observation d45d7ecf-ddff-41aa-86bf-6857b4160546 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Preserve your own correlation: A noise prior for video diffusion models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.515354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.455737Z digest=sha256:71d3bc49a898fc865087e0dd3560b50c5dd104c4fde174f5fa39a3d44083ecd0

Observation b742eb0c-580f-4f5b-b34d-c306fa973399 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.462262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.462262Z digest=sha256:5ca24cf9dfcb20fe69fd66690a764328e2d1385e394fad04c76375468f8e726c

Observation c2c74933-5148-4c10-b647-e1e55b9d7bed · outbound

This paper cites Mix-of-show: Decentralized low- rank adaptation for multi-concept customization of diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Mix-of-show: Decentralized low- rank adaptation for multi-concept customization of diffusion models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.495769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.469158Z digest=sha256:0108b0a9ac5b2b693b2d672d5fd81b7fb851d4d0c3bdbe6218b2e62283966f5d

Observation e54ea6dc-7614-4b13-8399-12c7e5efca3a · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.474649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.474649Z digest=sha256:79b93f0d7a4c99f999b994895be2124eb2f8bd428342e03e37b72563c85e969a

Observation 79a97b69-301f-4943-a09c-bfb1445076d0 · outbound

This paper cites Photorealistic video generation with diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Photorealistic video generation with diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.476856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.480605Z digest=sha256:4b7e690b8ef46b1e06a8d936098f6b64882d9ff42e40b902f8b926510499a1a6

Observation 256e437f-d703-41ce-bc4b-7359c5c675a7 · outbound

This paper cites Svdiff: Compact param- eter space for diffusion fine-tuning.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Svdiff: Compact param- eter space for diffusion fine-tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.457882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.487020Z digest=sha256:fc3384c0e5ac2b1a804cfa946ad83ae96a294f6148a1e53d680d1f0312c89952

Observation 36ad9b32-11f0-42c7-9b7b-271cd8988dde · outbound

This paper cites Face Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute Control.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Face Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute Control

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.491471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.491471Z digest=sha256:1f37d029aacb05ba2027ecfbdf5046994e0278df32f0c1b11e0da88c281523e3

Observation 89cf3de4-2fd4-415f-a5f1-3813eb1e59d8 · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.497037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.497037Z digest=sha256:85cde0f640472757ca2bbe266e70437d5834e6a3c69ceca7ec411343757d5f14

Observation 562d4900-e5a3-48f0-9859-26e534783bbf · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.502286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.502286Z digest=sha256:68e3c13a98a9059002dfe7c6c6a8d2686c9fd6ac688552260771951bff17ce67

Observation 0be9a833-05ef-4175-8883-5645e61dc531 · outbound

This paper cites Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.508087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.508087Z digest=sha256:3673ec136e3a58959e41925cafa8dfde0c79b72b72c68d6a6db18b9825993a0e

Observation c5a6ba51-b072-4f95-bdde-d43bde072159 · outbound

This paper cites Animate anyone: Consistent and controllable image- to-video synthesis for character animation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Animate anyone: Consistent and controllable image- to-video synthesis for character animation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.439302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.513653Z digest=sha256:ccbaed271f7ac6f6f6e985b1470a5689cdf2ecf352e5749aa4a6c12945be1de5

Observation 97ba6677-e3d4-4739-a10f-0f3793b04760 · outbound

This paper cites DreamTuner: Single Image is Enough for Subject-Driven Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models DreamTuner: Single Image is Enough for Subject-Driven Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.518225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.518225Z digest=sha256:a96c1bffb33cfa0530c0fe1997cc42368bb995eaa84f57593842197ff49001e4

Observation def76366-29bb-412b-b45e-c84a03e26662 · outbound

This paper cites Group diffusion transformers are unsupervised multi- task learners, 2024.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Group diffusion transformers are unsupervised multi- task learners, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.419607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.523193Z digest=sha256:f3da8a5923d9f5e19d7386b8b325ac1db1d008c41fd65f4ad3fc40972262dfbf

Observation 794760b9-5956-4352-a6a8-f21423322cdf · outbound

This paper cites In-context lora for diffusion transformers, 2024.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models In-context lora for diffusion transformers, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.399798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.528528Z digest=sha256:2b92982803a10eb8054de79e0c08719057ea1878ae78cf35eb2d8134952b7db6

Observation 06b51ad7-85df-49c8-bec0-c12397158ea0 · outbound

This paper cites Chatdit: A training-free baseline for task-agnostic free-form chatting with diffusion transformers,.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Chatdit: A training-free baseline for task-agnostic free-form chatting with diffusion transformers,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.381478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.533625Z digest=sha256:f26ce6b2c4d4d58c918382405ff01de2372fa2b48eb2e973db6fcd1f9cebbe5c

Observation b5132bb8-628b-4da7-b8d7-7505ef586365 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Vbench: Comprehensive bench- mark suite for video generative models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.363034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.539273Z digest=sha256:7ba1db9237ae2ef4c68892cacdda45cec6ee2951d9a2c21c32267a056fc9ad4a

Observation b4d5fa38-8bc0-413b-8678-5c951cce31ee · outbound

This paper cites Videobooth: Diffusion-based video generation with image prompts.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Videobooth: Diffusion-based video generation with image prompts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.342382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.544760Z digest=sha256:d72710f3b9057a6ccfdfe3f4f2aec2ee0bf112564af7b91f4e65d3c538414dea

Observation 249d3d11-45b4-4f11-af20-b6aca5a58216 · outbound

This paper cites MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.550221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.550221Z digest=sha256:b4e8a5d4ec72d66da6e8b5851cd2492468f088a23f22838a98e379dbfb485c91

Observation aa6cc8ac-03ed-4cdd-b0d2-b88e8e8ca161 · outbound

This paper cites Segment any- thing.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Segment any- thing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.317882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.557660Z digest=sha256:8351675faa62f03fb4911e05c058fe369df87019d122f80849c794139703d609

Observation a3c7af00-2e8d-4f6b-a47c-3d7bc0691486 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Multi-concept customization of text-to-image diffusion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.299344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.562588Z digest=sha256:143bfb096ac39fcc9d74b437d3b1bfeb31f99d39b03a7de9b69d141cea89c338

Observation fd6eb863-16ab-4acf-850b-fb5b95e54390 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.280239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.567717Z digest=sha256:02a56f1bca6a27337c741840955d74482d4c59cf4950beb4306fae931c67187e

Observation bc04fe08-215a-4bb3-940e-535541747bc8 · outbound

This paper cites When stylegan meets stable diffusion: a w+ adapter for personal- ized image generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models When stylegan meets stable diffusion: a w+ adapter for personal- ized image generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.259738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.573355Z digest=sha256:d2cfaed13f070ecbc7a647d439a336bc297299cf0b41cbf995230d28701cceda

Observation e413ad6e-6294-41bf-aaa7-b8f6c1d6391a · outbound

This paper cites Photomaker: Customizing realistic human photos via stacked id embedding.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Photomaker: Customizing realistic human photos via stacked id embedding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.241936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.578124Z digest=sha256:6586e155df84eef66d4e81126bb01702622f72f4989e2f5ea7dabd2c43444c85

Observation e6ef939f-4e92-489a-a8da-09a9f4568853 · outbound

This paper cites Idea-bench: How far are generative mod- els from professional designing?, 2024.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Idea-bench: How far are generative mod- els from professional designing?, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.224854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.582901Z digest=sha256:eec05d2c2b32ab638ac118b9d83a07b92f3931890752a3add06ceb5e9378ff91

Observation 1a6e50b6-ab13-48f8-8311-63c084574991 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.587649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.587649Z digest=sha256:fff43e49b9e97cce250f5715192d2e3a3868fe93c988138739adf30f25d1582a

Observation 604d3b37-8cda-48da-aad0-85f9fbd6beee · outbound

This paper cites Cones: Concept Neurons in Diffusion Models for Customized Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Cones: Concept Neurons in Diffusion Models for Customized Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.593565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.593565Z digest=sha256:952931d6048e0950086be8787fa7202d74fb90c9e996a05a7bc25a1a934bef0d

Observation f05394d2-ca55-473f-9f22-0b428b6e2ccb · outbound

This paper cites Cones 2: Customizable image synthesis with multiple subjects.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Cones 2: Customizable image synthesis with multiple subjects

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.208393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.599313Z digest=sha256:85f094e3b7a7bb313570bdc2c423396b91a4e9256f084cb13ea9b2dc04a5e5cd

Observation 7e9312c9-5fa5-49ec-a56d-cd62a191dcaf · outbound

This paper cites Mode: Clip data experts via clustering.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Mode: Clip data experts via clustering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.188620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.604623Z digest=sha256:a318a0be929aed4f33239991a40bb1a7b3b7412ea43fac040bd142c247dcb891

Observation ab395715-156a-4d5f-9a0a-ffc21c01d429 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Latte: Latent Diffusion Transformer for Video Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.610201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.610201Z digest=sha256:632e68dbedaa15a0a840968b763891ea5583ffbf80efb2f4b597896f38fe3ae8

Observation 5b159497-e6ac-4f33-bad7-7ff9865ff93e · outbound

This paper cites Ovmr: Open-vocabulary recognition with multi-modal ref- erences.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Ovmr: Open-vocabulary recognition with multi-modal ref- erences

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.169650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.617145Z digest=sha256:520423fc63bb738d2bfcbeaecd99a280bde7bc47f1d12b98bcfb7c43f40d27c8

Observation c5e0d985-5e75-4044-ae06-f78bdf36865f · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.148654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.622508Z digest=sha256:2aa6d35474a75710d1fb5913501042b5ee8260c444570033337cb664f193cbc7

Observation 7c2a8efb-b864-486c-9092-7963903c9a90 · outbound

This paper cites Glide: Towards photorealis- tic image generation and editing with text-guided diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Glide: Towards photorealis- tic image generation and editing with text-guided diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.132135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.628017Z digest=sha256:94db90b5de89f6473bc16166599eafe6a6845ff3a7118a9a9e4f5dca6a1a9504

Observation 67014fe7-77b1-43a4-8de4-478565529300 · outbound

This paper cites Scalable diffusion models with transformers.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Scalable diffusion models with transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.115283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.633509Z digest=sha256:947ad2cbfe5618ab1f5e90c7932bc33652689313fcdf09dc9a09642320e44919

Observation 9c0da423-d28e-4640-8c95-6fec925e651e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.639042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.639042Z digest=sha256:c04fa2448026f3d152c938c46677a9dc398a5f378a6c51334de1a529f4d0d60b

Observation e34be218-5326-4563-b72f-f52127aa1db0 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Learn- ing transferable visual models from natural language super- vision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.092363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.644863Z digest=sha256:cd98e01f185ce1192610b0e5987cbb904f10ddcb8e2ed01f6ae51cd5b3464bbb

Observation 7e134820-6933-45ab-8af9-44fdb1e0ec4d · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.650712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.650712Z digest=sha256:ee236b0aec239edd5c4ae3845d003946a99959f7c2c0d44d9bac7d1e49ec64da

Observation fa6d7b36-3381-4045-9e19-9992f8f8acb6 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models High-resolution image syn- thesis with latent diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.074114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.656328Z digest=sha256:15359b142f47488cb3228f9a0dd8b94c61fb37b099bb6e967ca19b52ea012833

Observation 82402dc4-8265-4c45-ba9b-62fdba98cb23 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.055618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.661537Z digest=sha256:9218f050fb6a393ad6ae1a05009a0c74aa56c76b41e6dead43cb6943715fe17e

Observation 1be00dee-eaf7-4555-b104-238ba83e3a99 · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.031125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.666548Z digest=sha256:ba9759ad10c2093a69b97fa53e204f1f6cec0a9936380f8f25f977cfcf4c793a

Observation 166d3413-1ddd-4b39-8e6b-d15e7e36ae85 · outbound

This paper cites Instant- booth: Personalized text-to-image generation without test- time finetuning.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Instant- booth: Personalized text-to-image generation without test- time finetuning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:18.014165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.671643Z digest=sha256:0aabf4e28713f5df08db2e354a2bedd0b7e0d639ae1c593843106da076c90b1a

Observation 58dba9a0-72d4-4e36-9fc3-e01d2cd845c9 · outbound

This paper cites Denoising Diffusion Implicit Models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Denoising Diffusion Implicit Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.676989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.676989Z digest=sha256:a257981682748768953a26042dc3a73ceaf4d62af69f4e93281f69816ebd3858

Observation d8d414da-4f5c-4ac6-b21a-50bc77205b10 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models ModelScope Text-to-Video Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.683672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.683672Z digest=sha256:ccad29244a44ace5b498a94a127c61dda03f01571df058d46e4676fbc86d0ba5

Observation d95836d5-dc1e-4f8e-b845-c39e9c46e375 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.690369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.690369Z digest=sha256:97def698f656212afde86164424082938e2d806f6edff9aceee21daec3c7d114

Observation fed664e7-9746-4563-81b5-26a7dbc4ec9d · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.696450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.696450Z digest=sha256:0ae72999e280cf6cd695892e706f704fe90b9ad476b79a5e79fa0e56e8aa930b

Observation 82fcf46f-a514-46fe-ae53-87faaa2ced75 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Videocomposer: Compositional video synthesis with motion controllability

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.992854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.702427Z digest=sha256:58817f8db026e7d7c1cae2ca46b36f5bb24cc2eeec494506f792f4bec2ea1552

Observation 12fa1013-217c-4f8e-8c1d-91b4e0755ad4 · outbound

This paper cites A recipe for scaling up text-to-video genera- tion with text-free videos.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models A recipe for scaling up text-to-video genera- tion with text-free videos

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.972894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.708068Z digest=sha256:be0667bdbcc898a2d5cb5344f4e313bee838dece1a32f1d6baf3ffbae51ae7b3

Observation a8e37e7a-9459-4575-a3c9-36e1dfeb177b · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.713313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.713313Z digest=sha256:7e7648bb78633ea9d4d80d3470cc3b754a01a2f145a84910e54042a9b1faa42f

Observation aaf3c0a0-e3aa-49f1-929e-cbe52da254bd · outbound

This paper cites Customvideo: Customizing text- to-video generation with multiple subjects.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Customvideo: Customizing text- to-video generation with multiple subjects

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.718378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.718378Z digest=sha256:5fc045e26d8cf07f63b9985ea586ca0d463ecde6bfdc683268d1f442afb845e1

Observation 4cbc0782-f118-4416-8840-1c4a9e4b3985 · outbound

This paper cites Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.950850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.722942Z digest=sha256:454ad2d06f92a2f9523848c8256e4cc36627fe8a7270fdd8e8600eca53555b88

Observation 6f22435f-91d4-42e1-8d20-82d498494dfb · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Dreamvideo: Composing your dream videos with customized subject and motion

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.929501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.727925Z digest=sha256:ff7c8dca091369e852a38af95c2b06dc072908801058232f0b4343d0ae551da9

Observation 4812631a-f4f0-4492-a8e8-3120fa62d133 · outbound

This paper cites Spherediffusion: Spherical geometry- aware distortion resilient diffusion model.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Spherediffusion: Spherical geometry- aware distortion resilient diffusion model

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.909499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.733633Z digest=sha256:03b833076fa5964428f740a2a2e252f2a6300bda706f0b2cce15f4422aa81817

Observation 402eb5ac-632a-4958-9c2d-63bbcbce4a51 · outbound

This paper cites CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.739129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.739129Z digest=sha256:54ba33f869621634cd4c2f15461e74aaf88df6799df6b875d1b153ce8f7e42b9

Observation 143526f4-a79b-4b0a-b7a7-67dfb6eda3e4 · outbound

This paper cites X-portrait: Expressive portrait anima- tion with hierarchical motion attention.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models X-portrait: Expressive portrait anima- tion with hierarchical motion attention

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.891524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.744405Z digest=sha256:26743dab6f9a1fa3588c727617076070aa38b1b13e21c8302166f1a4a8c1229a

Observation bb07be18-246b-4771-b638-ab03d4915360 · outbound

This paper cites Demystifying clip data.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Demystifying clip data

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.874493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.749867Z digest=sha256:341b153026bfd16c9955a5295fa92a41ee05c1af77d1b8b17603e83f3b70b6e7

Observation ce9db863-7eda-40ea-af97-0ef8f78675e6 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.755038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.755038Z digest=sha256:8df5691c49b9f02b869d04ef11ab319bb5383921bbe7e1dc1ecf4d0b4927fcba

Observation 1f9da4a8-e9bc-4673-8d0a-44a07c75ac35 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.760656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.760656Z digest=sha256:fe60c9bbdaa3b08afee923b533b53cbe21940a741f7c7b859359d3fe17562eee

Observation 3aee5cad-39a8-44c3-9977-3a122958ffb1 · outbound

This paper cites Celebv-text: A large-scale fa- cial text-video dataset.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Celebv-text: A large-scale fa- cial text-video dataset

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.766562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.766562Z digest=sha256:6c700ef3354a00e50f1ce45adc78dbfdc53aea4f3e21e4266996cc99e986a1ca

Observation 4ae66ad6-4d87-4416-a314-ee597bba4843 · outbound

This paper cites Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.772204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.772204Z digest=sha256:50de581dc7efefde278c0614530707803266e9e48259f65c4e3ba22d3a708e1e

Observation e71b0634-a6c2-4f63-a80b-c07d35238af7 · outbound

This paper cites Inserting anybody in diffusion models via celeb ba- sis.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Inserting anybody in diffusion models via celeb ba- sis

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.844470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.778056Z digest=sha256:0e20d06b20658edaef054279548afe6352392087d79ed3b0350eec8cb4425ce3

Observation b42aebb2-ecc8-4a45-ad82-668a8eee4981 · outbound

This paper cites Instructvideo: instructing video dif- fusion models with human feedback.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Instructvideo: instructing video dif- fusion models with human feedback

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.827084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.784743Z digest=sha256:2f0900b468fd865d82f59df374e287490a8aceaf1c305ca089f47dd88b4b08ff

Observation 63522520-1978-4b69-a9d6-83c249d8fdff · outbound

This paper cites Se- manticmim: Marring masked image modeling with seman- tics compression for general visual representation, 2024.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Se- manticmim: Marring masked image modeling with seman- tics compression for general visual representation, 2024

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.808809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.790432Z digest=sha256:e5373d2dfd37eaffb59fb1fb67336b51b6735d31afb93651e54f016824e08d26

Observation c37df1d1-d6dd-436a-a651-bbf31ec42fb8 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.790905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.797213Z digest=sha256:c88c44750303fd7c6218b2bcdd894c42fe2cefd809484e4de799fc778fbab2a3

Observation d65aa335-33f9-47fa-bbd5-80114771d055 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.802501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.802501Z digest=sha256:1b2c014ea192bb8361dc2b110792c62071483ea20db63558456e37b58e9ddad3

Observation 9df6535c-529e-4e5c-b892-17a515a19d83 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Adding conditional control to text-to-image diffusion models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.773205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.809329Z digest=sha256:bfab4c2d38bebb43e8d717e77b45c1738b3e1487e624f9eb91d1c5f703aaa0f7

Observation a083c5bb-aa99-499f-a3b9-a05be84ba76a · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.814722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.814722Z digest=sha256:e9680d25802b77e831ab560dd470a98cd3de17e77fa0cc23da46ed0abfab25cc

Observation ceb8bec5-3173-4077-a39c-8c5bb497d928 · outbound

This paper cites Videoassembler: Identity- consistent video generation with reference entities using dif- fusion model.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Videoassembler: Identity- consistent video generation with reference entities using dif- fusion model

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.753931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.820723Z digest=sha256:3956fcb3c11e93b78195ec2b585e14eea3ae925ea0d65ec10f09f8700e4c75fe

Observation 2d9db86a-cc03-444b-b52d-b218b78270f1 · outbound

This paper cites Towards fine-grained hboe with rendered orientation set and laplace smoothing.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Towards fine-grained hboe with rendered orientation set and laplace smoothing

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:10:17.733600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:10:16.826406Z digest=sha256:13bc5d2520fdaa01fef3daf056326019b0a86a52261995e3e5bb0ab6ee64289e

Observation 1b125d9c-8362-4838-b666-9ef21ecf1623 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.831784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.831784Z digest=sha256:e0591644fa5252fb8cb961e53e129b7b923eed8a0d7e24ad51677ef41734379b

Observation 3619a73a-f57a-4d64-9f1c-32d5a0585e98 · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.837253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.837253Z digest=sha256:012aaeb4b783755be737f360f4c6cb90c6210f35360d8924fe41bb8c88cac165

Pith citing papers

Observation de7a006c-226f-4bc8-9be9-c46371fc1674 · inbound

3D Object Manipulation in a Single Image using Generative Models cites this paper.

3D Object Manipulation in a Single Image using Generative Models VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T16:40:47.679958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:40:47.679958Z digest=sha256:6e12a3125ed983659729b9771a505949453faa0f1de0a48e4c17bf6c340e1a65

Observation 3032288c-c5d5-45b8-9240-017192d0beeb · inbound

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control cites this paper.

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T19:40:21.394334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:40:21.394334Z digest=sha256:39a2953ac7060e0d66d80d907364c0b3764df9a4b44c5c7bcfb9d432623ad3d1

Observation dcc6d590-eba2-4688-b6e3-daffc7ced655 · inbound

Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform cites this paper.

Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:11.282581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:33:11.282581Z digest=sha256:1da2f4ed37580ce24ef4670a5bc2eb0f84e6976f591397aab6a34aecb24b1fd7

Observation 826819c3-4153-418f-983c-d7476cac4c33 · inbound

DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition cites this paper.

DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T10:46:42.382695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:46:42.382695Z digest=sha256:c9769dba117ffcbc0aa7ab1f3d29d071089f998b87b0ba19d679904cf01600df

Observation 4cbccfbe-3158-4999-8727-8ff4197e943d · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.772784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.772784Z digest=sha256:7a90ead3b7759f5b667c369b719dc6399d0b7cff809792171970054f590cc7d3

Observation 85200cf7-3b60-4892-8817-72461ee0b9e9 · inbound

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation cites this paper.

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:10.490458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T07:59:49.398271Z digest=sha256:e228e5ff7e47d7b06b745657d080d028d13479e9dc4d7af071575507d536fb8c

Observation b3e9c20b-5861-4e0c-82a3-f64d54955d99 · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.047227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:6f803d0b9306fe08bf95557724a89724057c7309a026be6e8f17625edb4a3da6

Observation 1f722949-2bec-4b18-b24c-d5afb16ece11 · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.923807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:82ae87088c38afb8c58de13254385ac936d353ebb8c0e5090452b2b6d2f37593