Pith. sign in

Paper Citation Record · LEDGER

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner

As of 13 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2412.10533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10533 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:56:59.884771Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T07:59:49.398271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T08:02:10.321954Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59bbd1ae-a2aa-4e20-a7d3-d2373a329493 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.708772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.708772Z digest=sha256:267c79846ab3ce8a3d012cb7a3bafb41f639d97c07ac8c191f35378c183901de

Observation 2da0f429-201e-4bbc-b52f-3584b8011976 · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.712836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.712836Z digest=sha256:7940ee0a653d791d523c8057818e93f48e667cc23a6ab67959a9fe2af12d8ec9

Observation 47ccf589-3e8c-4a81-a250-70e485f63dcb · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.716910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.716910Z digest=sha256:7e28251c68a645ab7144dcefe9157a064a6d802d5ceadddef707aa544a5f9f38

Observation a36975d0-24fe-4705-82b5-f943b630e39b · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.386952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.721156Z digest=sha256:fb8b82feb6f212c9469133dc19722046ffc980ceab3991cd55dde4969be3cd72

Observation 39352f1e-6020-4f50-93e1-e8d983bde9e4 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner In- structpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.724691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.724691Z digest=sha256:b4333dd29f37bb93add12f65ad47691ecafc3aa49c5c9e54f51098321ae749b9

Observation 0c507006-2423-4db7-a323-9863727f7743 · outbound

This paper cites Video generation models as world simulators,.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Video generation models as world simulators,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.728227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.728227Z digest=sha256:3af8d666842bfbe4c5aa282fe8181559897848d5f5b3d2eba763876b260f5c1f

Observation 19a02f93-75f1-43c3-a073-e295672d3991 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Emerg- ing properties in self-supervised vision transformers

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.363853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.731701Z digest=sha256:da0839fbcf865224a4b9599c4aa9dbf9ee55d3cfd766cc35c97c54a8f90f1b1d

Observation b90e1452-3a61-48ec-8aca-a7dcaaf71de6 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Panda-70m: Captioning 70m videos with multiple cross-modality teachers, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.352581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.735334Z digest=sha256:d5668c9267ea082c0348d3ad3e542f619bed65dd696d3e445652c160f7033195

Observation 1ad3085a-95b8-41f1-b6ef-dbf825b5011b · outbound

This paper cites Subject-driven Text-to-Image Generation via Apprenticeship Learning.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Subject-driven Text-to-Image Generation via Apprenticeship Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.739929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.739929Z digest=sha256:6ebdc63ba5d48d62a965bdd5b86de07a5b2f9216b27ef227af74521ade9022ee

Observation de93377e-56ca-4595-afec-b447c7392da1 · outbound

This paper cites AnyDoor: Zero-shot Object-level Image Customization.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner AnyDoor: Zero-shot Object-level Image Customization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.743465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.743465Z digest=sha256:17cbac1c88553ac61745d589660333a1e319ea90965834e75464f049263f20a8

Observation 98d16545-fab3-48ce-9904-a5a0de3c467c · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.747211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.747211Z digest=sha256:43984f732829830569247d1a906b34ea2e1c808cd176c813f81356645df8158e

Observation b9daad35-7f56-459e-83d2-4af4696b8dda · outbound

This paper cites An image is worth one word: Personalizing text-to-image gener- ation using textual inversion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner An image is worth one word: Personalizing text-to-image gener- ation using textual inversion

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.334812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.750662Z digest=sha256:6a8571289115e322b21c8d2afaa79de4d31812391bdbac924d3ea4bf28b4fded

Observation 1c6d2022-1679-4b6c-ad95-00f8dd4db59d · outbound

This paper cites Photorealistic video generation with diffusion models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Photorealistic video generation with diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.323368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.753971Z digest=sha256:3d7b88246925ba3c1fa4d09df281ab981c870b90e03e84e3f1032cb8d2fb2491

Observation 1a2aa6c9-2160-4eae-9379-631fb3845d01 · outbound

This paper cites Classifier-free diffusion guidance.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Classifier-free diffusion guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.757231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.757231Z digest=sha256:d1f5a5a4644f39648bc0705987f11be3a4c00363ed109240677e1a0e7184737e

Observation 87563d6d-011a-4ccf-b66f-8fceb9f4dd53 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Imagen Video: High Definition Video Generation with Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.760051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.760051Z digest=sha256:896f6d12e0186caa24c3e7b06730e3e9c2f6cf5b21a2c79814878534e1df87f6

Observation c5225c12-bf47-4636-b644-ba3ba1acf683 · outbound

This paper cites Vbench: Com- prehensive benchmark suite for video generative models,.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Vbench: Com- prehensive benchmark suite for video generative models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.763446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.763446Z digest=sha256:97296b63cff43fb4a6c80d3417a69427adf59b983b337ebcb95573121cff6626

Observation 6448f32a-29d8-44cf-a834-14cb6e323c87 · outbound

This paper cites Videobooth: Diffusion-based video generation with image prompts.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Videobooth: Diffusion-based video generation with image prompts

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.300238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.766888Z digest=sha256:70cc311eeffb331a2917c2b72cbe3d9ec911828ec0c9991ab8024179f5f90b45

Observation 9b7c4109-6273-4269-836a-ad3c8baaf761 · outbound

This paper cites Auto-Encoding Variational Bayes.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Auto-Encoding Variational Bayes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.769753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.769753Z digest=sha256:91511a81918007e298733c8c11ff6a9837b2b455ec10109e9e43129b15cc7b43

Observation 50558660-b6e4-4882-ba4c-733cb2f67264 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.773238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.773238Z digest=sha256:829f31bd5efdae825f97f49a716ea1bd6573bf0344a878eb788678c29697c624

Observation 633e40d1-d54a-4ff4-9af5-814284aae0de · outbound

This paper cites Multi-Concept Customization of Text-to-Image Diffusion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Multi-Concept Customization of Text-to-Image Diffusion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.776657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.776657Z digest=sha256:2ff9e78c29c76a83da11d7041de0f2ebc97802ef47e93e29f206ce76d85076cd

Observation 6f922097-c920-4097-bf7d-78921a88ae39 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Multi-concept customization of text-to-image diffusion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.289833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.780665Z digest=sha256:bc1fbc4c76d430171758bcfc2695955378df2862615158a4828d8991e1670a26

Observation d11d5300-00fa-4a67-be32-82579b415bdd · outbound

This paper cites Flux, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Flux, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.784156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.784156Z digest=sha256:a1c0993459b0493e84b5cb775f0641f3e61a6efc61ef73941d184498c1ee48da

Observation d95009a3-d759-4854-8e37-0299f77eafaa · outbound

This paper cites Luma dream machine, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Luma dream machine, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.272606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.787688Z digest=sha256:f584b30662beeeca97ebb96c4c679c779622cca3545a0f06dda37a39fec9e353

Observation e1e03782-b62d-45ab-b5e6-2e7bb7fc09c9 · outbound

This paper cites Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.791130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.791130Z digest=sha256:a7d649ee7c62e8692e6761991e426acf6ed9f9585cab29864b80c1bd37a5654e

Observation d6e1cda3-4448-4b27-a6da-8eb19ecb3b51 · outbound

This paper cites Dinov2: Learning robust visual features with- out supervision, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Dinov2: Learning robust visual features with- out supervision, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.261606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.794848Z digest=sha256:3bccc0c83f435dca6aba6eb8922e86d225a2ad35565c52a6f9389dfd0a1f31fa

Observation 4b479e08-728f-4663-b945-eddc8005300b · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.798517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.798517Z digest=sha256:51f4c46f77eb6263ca8141119bb24b8a614b50741288a5f637e5ffddfd93d92e

Observation 40c450c9-ddfe-44be-afe6-d3d52000a588 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.250736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.802465Z digest=sha256:1fac28ea1cce7a8f6e455492361926c0d56b1bcd32f025b87f8b8253566691df

Observation 5e08c81f-3c88-4e61-860c-d37d401936bd · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Movie Gen: A Cast of Media Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.806149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.806149Z digest=sha256:97b87d89aa01780e935a0583b049da9e2a6b137a69f95b4336f502bcae1d74dc

Observation 158342f1-584f-4a33-843f-9cb45ecf18c3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Learning transferable visual models from natural language supervi- sion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.240382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.810182Z digest=sha256:630cd11ebe2f6d828ee145d0ac251c3ace8ea08b58fbf4b41087db7c3c80c9f6

Observation b73b69ad-680a-4ed4-9775-650c7037ff28 · outbound

This paper cites Zero-shot text-to-image generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Zero-shot text-to-image generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.230002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.813440Z digest=sha256:0c7fa8b62cf062cee52eea4f96c435bdb1509e084b21fdf7d76592cb2aad0a96

Observation 559d896d-e333-4552-ba5c-dea380914e72 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.816893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.816893Z digest=sha256:ab7590826951cc4951f0cb97e481a73daad53097323c4b401550df3fe8b99ede

Observation 8a12b6b9-2fdc-411d-8db7-aa32eda2861e · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner High-resolution image synthesis with latent diffusion models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.820341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.820341Z digest=sha256:92769cdac85c6a87d9d51b5afa389599ec5557b4a418e2031b5ece7c0c986463

Observation a8414ba3-8c42-49b0-b472-3d24bee5d2c6 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.207888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.823699Z digest=sha256:097d0a6253689abead1b3d575bf78ae851a98e8ef463eabff0590c2d4aa9e5c0

Observation 21ce158a-a063-4d99-b97b-da9aa34c843b · outbound

This paper cites Gen-3 alpha, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Gen-3 alpha, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.198475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.827139Z digest=sha256:30c3553839963b4cfaebd440e02e6eb98d1cdbd0ad5a1ea60d824ca9865019ee

Observation 5491f2eb-6d4a-4dbb-bd1b-04c6b0d53833 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Photorealistic text-to-image diffusion models with deep language understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.830494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.830494Z digest=sha256:72adb7694cbdc7f9c64009bc6e63e58290df3a8e14ce0430919c106db59f0a2a

Observation db47afd9-34be-4b37-9adb-275cf0cb64e5 · outbound

This paper cites InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.834021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.834021Z digest=sha256:428f9eb42953d6487a15c3ce7218abb19ccebc9d1f415ffd4d4480636beb4214

Observation 2571f636-ea87-4d93-a27c-5cc3282ffd47 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Make-a-video: Text-to-video generation without text-video data

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.182057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.837693Z digest=sha256:7cdbe0dcf90132e91986a95b2f6f49a454cc835896bad653536a5b11fde4eda2

Observation 140dc4fc-bdbf-4cee-8409-0842352a484e · outbound

This paper cites Vidu-1.5, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Vidu-1.5, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.171942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.841160Z digest=sha256:f6dddeb5a3332816fee6cb857264efe46f27cb069057d1c67db495ba8ee44c74

Observation 74d9c82e-6976-4601-90e8-36461d81ec3a · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Raft: Recurrent all-pairs field transforms for optical flow

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.844534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.844534Z digest=sha256:bae50275d046efc5386d841698b5b8b7e05216e9700311e13c000d8232a184a3

Observation b31fab30-08df-4b24-a023-d72695fb7a04 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.847951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.847951Z digest=sha256:fb01fcc69131b355bf57abcec0910881a156cf68caacf45ffd7298a0f1ad3f79

Observation 8cbef5c0-658a-4364-9608-70e79551bd21 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.851385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.851385Z digest=sha256:af8a52209451fa5da4eabcafa08653a58080294d2a6c4e8bba50a357842f7ecd

Observation 2773f20a-be5a-40e7-b5f4-45f16055ae43 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.155337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.854770Z digest=sha256:6a6aff5c87ed66664c13e7455f07109d67dde232bdd7850071f9e93008698b3f

Observation db47e8a1-1982-4929-bca5-aa81b5b3b573 · outbound

This paper cites ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.857766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.857766Z digest=sha256:c19f3e227fea9248398e3417a987d430eb375d28ee8b87b3a77d25c969db8a1a

Observation aca6d5a6-5ab2-4d32-9675-11172a53646f · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Dreamvideo: Composing your dream videos with customized subject and motion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.144278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.861277Z digest=sha256:a7d296f5ce9c00fa0a9b5a3611e98c491daa088ac532150bf20566a547047bac

Observation c9b27005-25cd-49eb-8902-43ac6720a73d · outbound

This paper cites DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.864206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.864206Z digest=sha256:f8fe5dd897f5806f11b75279d89b110e89e91f64c1c8b8d76a2179f917171f0e

Observation 5621611e-500a-407e-ae53-d9bc6eaf7c1a · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.867818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.867818Z digest=sha256:a245bd807b6f01942972e3521f11f6cab91adbc2fd2c275f066c5fbe02bdc4d7

Observation e5ab495c-cb78-4a2b-9de3-372923b1b56c · outbound

This paper cites DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.870792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.870792Z digest=sha256:901127a52b5b983f893473ea153c9f56fba75f9d2bd5ed45ab26e6e47fc1bbe6

Observation 4f4d8f0a-e63a-4cdb-ba46-260d45622ead · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.874312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.874312Z digest=sha256:1b79635b4f5f94a1fe37153cffae32d82c63c381f552b88c48a3512e820b5f8d

Observation 1e2374a1-e3d0-4798-a079-795cecbda668 · outbound

This paper cites Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.877924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.877924Z digest=sha256:a7397fe0513deee6592b8577d2450525054f7af3d8ba8628fb40088c6fe02195

Observation 3a35599c-b16c-48cb-aa01-978009ec18a0 · outbound

This paper cites Cus- tomization assistant for text-to-image generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Cus- tomization assistant for text-to-image generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.127398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.881408Z digest=sha256:6c7121e514ce4825ce131711e281cbece997adf84885ea356d870e4598412944

Observation 0690ad41-8ad0-4b17-a153-70a55b4b9120 · outbound

This paper cites pencil drawing drawn by a hand.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner pencil drawing drawn by a hand

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.116346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:56:59.884771Z digest=sha256:8e81ddc846c64530d618c1fd2d92e618f5df72956378195233cf0e27234427cc

Pith citing papers

Observation 4fa29253-1ce0-4f96-9a5e-93e646f1a396 · inbound

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation cites this paper.

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:10.325181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T07:59:49.398271Z digest=sha256:cdd8980f62ec9e0b2517f61451b2a99c1df9eeae12d808df69c1c8fa47d09e75