Pith. sign in

Paper Citation Record · LEDGER

Identity-Preserving Text-to-Video Generation by Frequency Decomposition

As of 20 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 23 inbound Pith citation observations for arXiv:2411.17440.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17440 v3

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:10:27.446268Z

measured 123 of 123 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:29:36.698279Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:42.422443Z

Reference resolution

100 of 109 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved84
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2d0e0ec-1a23-492c-8683-b8ff4ccc6bc8 · outbound

This paper cites BoT-SORT: Robust Associations Multi-Pedestrian Tracking.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition BoT-SORT: Robust Associations Multi-Pedestrian Tracking

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.147151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.147151Z digest=sha256:a9edacbde20f1c9194f38b35028f63df4ece9e141b330454e693886ad9ad9e2e

Observation 3b01816a-a9b4-48ed-ae0f-b1d93d77f301 · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.151095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.151095Z digest=sha256:76fd960f93e3e40f6aa65e226370d33ad35f8c1c6e520903169e95fa79de6d37

Observation 11e9c834-ac41-47e1-8ccf-46e76d5e3859 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.153999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.153999Z digest=sha256:63845af88bc33501650732e9b526874d48b13a5f6433ba4f7320885aa86ef8c6

Observation a9304edf-a19f-4be8-820f-f7c9781cd1de · outbound

This paper cites Improving vision transformers by revis- iting high-frequency components.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Improving vision transformers by revis- iting high-frequency components

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.156942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.156942Z digest=sha256:7ecbe1a0f15f5eacb7b8dc66713d8d4b814eabd4e418922218f0da1e4a13a94c

Observation 721cd4b0-9dc0-4a75-b00f-88288aec8286 · outbound

This paper cites All are worth words: A vit backbone for diffusion models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition All are worth words: A vit backbone for diffusion models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.159959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.159959Z digest=sha256:4ad0e92a595e06a1c3b41cee836e610afef0ed3c20f0eb6885bbb97a1d594ea2

Observation 723be1c0-4869-42cf-b611-e11687c59717 · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.163774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.163774Z digest=sha256:493c40ef374fafd86d71cb1eaca37373973d3f6bbc8907a30dcec22fa00c4519

Observation 5a8391a1-cc4b-442d-8d99-b9af9bfdf7a0 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.167230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.167230Z digest=sha256:0b0a024c228f3f343037e0f5c3e15158c2edf793fcdaa073af0a7fe047d686d2

Observation 11854595-7c91-4e08-ad11-48ff98a0570d · outbound

This paper cites Video generation models as world simulators.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Video generation models as world simulators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.170191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.170191Z digest=sha256:1d2650770ffbe86efa2584fe1fcbf94be1b92b896696cc48f8f86877650fae63

Observation 214b5091-34aa-4cc0-ba00-20fbd4b754ff · outbound

This paper cites Vggface2: A dataset for recognising faces across pose and age.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Vggface2: A dataset for recognising faces across pose and age

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.174130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.174130Z digest=sha256:c707a49db359f9952039eb153e920b24e2933a0b7094608b690c98365dea4f14

Observation 368d66d3-70ec-4b0e-94dc-8116dbc690f9 · outbound

This paper cites Still-Moving: Customized Video Generation without Customized Video Data.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Still-Moving: Customized Video Generation without Customized Video Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.177019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.177019Z digest=sha256:61b27c5d9e9607db5cc64b896f140cd793fdc1d449b5bf54c5c882e943cd1236

Observation ab41c5f0-1efa-4138-bcda-1b37fcc5db25 · outbound

This paper cites PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.180024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.180024Z digest=sha256:f46cae29e385a272e2024f858876b97f9be2650aca7d418fe43723029f6fd718

Observation e7e0707c-a7ca-47d7-b2a9-4e3f20cf528f · outbound

This paper cites Boosting Camera Motion Control for Video Diffusion Transformers.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Boosting Camera Motion Control for Video Diffusion Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.183197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.183197Z digest=sha256:a310b330dcc62bd53f5665f86109a1bc4b30eec10f5d9109141f56488ceadb5c

Observation 460b06dc-c729-45d4-8c39-1d27f9fda29c · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Arcface: Additive angular margin loss for deep face recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.186346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.186346Z digest=sha256:910bfbdfb08d0c7c3529c63fd75de6fcc2bd8ed76de0cdb9ed6d31f25a2c101f

Observation 03475cd6-06a5-4412-9d8d-b84c305963f5 · outbound

This paper cites DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.189537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.189537Z digest=sha256:e91990dc09f50a460979ec522c2edc533d9f911a7ae2d635b9324eab255357eb

Observation 1cc5c8f9-4f50-4a02-8604-fe8a5b1b0761 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.192501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.192501Z digest=sha256:fb0746e965294468f207a728aa64399044c8b6d9293a4e47a3f3b8b389036d35

Observation 1125e7cb-1f32-4d58-93b7-8edfa8158e43 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Animatediff: Animate your personalized text-to- image diffusion models without specific tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.195578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.195578Z digest=sha256:7bcb7c12cd823ce3f631f1c2bbd102c83ccfe554e8fce518bcf00cd84af82b33

Observation 59868d19-d236-4225-8ac2-125bb57ce12f · outbound

This paper cites PuLID: Pure and Lightning ID Customization via Contrastive Alignment.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition PuLID: Pure and Lightning ID Customization via Contrastive Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.198608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.198608Z digest=sha256:3e9e3625d1a170b31ac1312626897023f1d06d4587edf30a3483718f4ab7d9b2

Observation 9f6a9c2f-3607-4bbb-884b-6f68dddb49f8 · outbound

This paper cites UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.201466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.201466Z digest=sha256:16f3feb158b11d577672f1eb8a01ec373355e0423d00ec3c3fb9255875c91f64

Observation c3f99e9b-69c8-4810-add0-a2c7cb8feec5 · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.204401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.204401Z digest=sha256:e6f0707a8c53ab3e48b4e100d08b0d3c758dc92f6ad48ab531bb83b28fe4746a

Observation be4cbf3f-6460-4e4a-adb1-66951930f25c · outbound

This paper cites Imagine yourself: Tuning-Free Personalized Image Generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Imagine yourself: Tuning-Free Personalized Image Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.207265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.207265Z digest=sha256:173ad63af7781ec6b137746fca03620c90b9767060a627b7afbee0d98b642f6d

Observation 2a70c6db-af35-4834-8628-962d20f2c769 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.210128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.210128Z digest=sha256:0c73f1143b563b163c4fd7d942af310a17ad5d25ea892436de89b0db914c6b39

Observation 2783bbce-ae3c-435c-984a-91b68ef028c0 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.213056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.213056Z digest=sha256:36ab79c0d1ded04ba96a5a48fcc797078285d56a1cbde9552c5d92ed0bbb0e62

Observation 94a468ab-0a64-4879-87c5-31f6fa45c1b4 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Classifier-Free Diffusion Guidance

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.216945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.216945Z digest=sha256:11ab2cf26ca436e86091804b94715fbed70e9e87ce1b38fa047b3b9f70ad652d

Observation 25908296-956f-4e5c-8630-52d1eb7c6ca6 · outbound

This paper cites Denoising diffu- sion probabilistic models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Denoising diffu- sion probabilistic models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.219812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.219812Z digest=sha256:b46af66068bac492449baeacc3f5c525f42c377af1fb8f3e607d007be89e26ee

Observation 6792fc7b-cdc7-4f79-9c4b-185c7a95bd47 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition LoRA: Low-Rank Adaptation of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.222643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.222643Z digest=sha256:bf70cda23ad119426d7c6426c089c38e540407714ee770f3dcec8e7565dd024d

Observation cd01a0ac-281e-49f5-8a20-4de8d2b3a8df · outbound

This paper cites Deep networks with stochastic depth.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Deep networks with stochastic depth

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.225563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.225563Z digest=sha256:f00047da57d177127c8db39292d957b67ade7fc22a4f3635c0204ab9db1345af

Observation dc4ae9da-bfed-4094-9cc6-5c987530d06a · outbound

This paper cites Curricularface: adaptive curriculum learning loss for deep face recognition.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Curricularface: adaptive curriculum learning loss for deep face recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.228758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.228758Z digest=sha256:9976e0f25cb3bd16fe1630d19a9297e9069a931d492229e256ca20c4160bee8a

Observation a846910f-7ff8-462a-81c9-8e89fcf3e053 · outbound

This paper cites Real-time intermediate flow estimation for video frame interpolation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Real-time intermediate flow estimation for video frame interpolation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.231443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.231443Z digest=sha256:a4865919397790e2f02b75106e3c3afd477f997fb40d0d7cd4f0230d2e4f256d

Observation f5eca017-db05-4e46-b140-46f5b1eb1afd · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Vbench: Comprehensive bench- mark suite for video generative models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.234050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.234050Z digest=sha256:9df388ab9150ac4ce03e89c4d65f45ffa9dc7e3690e38f0db65483586f77df76

Observation 945f2c14-a021-4d75-87f5-534a92c910b4 · outbound

This paper cites ultralytics/yolov5: v7.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition ultralytics/yolov5: v7

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.236736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.236736Z digest=sha256:ff5634884ab7afe3401b9106db3ec3e5ee700347ffa04ac3666816932dfd9a2a

Observation 05eb1106-bb5d-4eef-afac-c7e0e1578847 · outbound

This paper cites Progressive Growing of GANs for Improved Quality, Stability, and Variation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Progressive Growing of GANs for Improved Quality, Stability, and Variation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.239404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.239404Z digest=sha256:15d8a4f3d503f6191bc745df99d241ebe7ef84ac0c3bb82e5833c82d9a380e3d

Observation 5c833d34-2396-422a-867b-bdc66102ef7d · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.242981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.242981Z digest=sha256:1408c9d65f74dde157c587939579994cc252f33d6e178447e2d220475c962d5c

Observation acc08e78-ac1f-4a44-80aa-51f7a4b8e724 · outbound

This paper cites Multi-concept customiza- tion of text-to-image diffusion.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Multi-concept customiza- tion of text-to-image diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.246120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.246120Z digest=sha256:760547ec8290d2253b7ac9daff813cf201ebab82bcf7200228bfb475a5b036ef

Observation c6e3dea8-633a-45ec-b833-4c73096f853f · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.248925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.248925Z digest=sha256:9d5682ef7869e0b29494c0fcc087bf7859e7bd86a5a5a4e9ebe4f71bf62a0f76

Observation 890166e4-fbda-49d3-9fbd-dd41597cbecf · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.251514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.251514Z digest=sha256:2166ea3c86de25ea0850b298ceea05504e566c93bbdbe5d9ff6495771b4f3885

Observation 800e02b5-ad1a-43f1-a255-64ef9138bbc5 · outbound

This paper cites Photomaker: Customizing re- alistic human photos via stacked id embedding.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Photomaker: Customizing re- alistic human photos via stacked id embedding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.254330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.254330Z digest=sha256:76c9083c06e6a2cd30e6877ebca47c4c45c918e18807540006455bdf8d9d70cf

Observation 2c4de337-d5d3-42da-84c1-f8de135de424 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Open-Sora Plan: Open-Source Large Video Generation Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.256976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.256976Z digest=sha256:cf1edcaa1f76ee7e014e3732c5140f4b71c7a658b95f8267881d1cb00d9b4565

Observation 5bb877a6-12e6-4fb8-adc0-6608a723ee47 · outbound

This paper cites Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.260005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.260005Z digest=sha256:2649fcbaf9ed4dc9d6e13d3d42da754001b582dd50c12cd75664937b516591fb

Observation bf7357a8-f1aa-40ab-a956-1856d518fad4 · outbound

This paper cites EvalCrafter: Benchmarking and Evaluating Large Video Generation Models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.262744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.262744Z digest=sha256:ce3120fabd392ca6921ef7fb58ce467f9b6821a2c192217561810e3e38e7a6e8

Observation b366be98-ae72-4d63-a80d-79ee0e587352 · outbound

This paper cites Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.265599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.265599Z digest=sha256:ae37ed47aac572824d031b1febab9b06e789ae22e934b3db9e642e6558a37d40

Observation 75dafaa2-adbb-40e2-99f4-60bb54e30105 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Latte: Latent Diffusion Transformer for Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.268441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.268441Z digest=sha256:aeca38a4015fd2e0caa7bc7f9fb02682889e4e60418a506dd79df3fad1076656

Observation 2a2e3827-9be0-4234-8c31-dafce71c4619 · outbound

This paper cites MagicStick: Controllable Video Editing via Control Handle Transformations.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition MagicStick: Controllable Video Editing via Control Handle Transformations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.271821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.271821Z digest=sha256:3fd45c1c62d12e9135171e56fa01dcda8e18bc3c243e51717fb731a0b2b3b162

Observation 3afb0d6b-c836-4e11-8c7c-2aaffa160b5c · outbound

This paper cites Follow your pose: Pose- guided text-to-video generation using pose-free videos.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Follow your pose: Pose- guided text-to-video generation using pose-free videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.274875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.274875Z digest=sha256:9214e90cd623d123fee826b96d8d354898df4271aea765f3bd0b8874a36f90f1

Observation 1922c170-7ae4-43c6-8d6e-b242ca9438cd · outbound

This paper cites Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.277604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.277604Z digest=sha256:d14f2e3089823c401bc007e76de800781ed2fb9ff3df084c0742ea27a3e5fdd3

Observation 0b647d22-0fa6-42ca-a3ec-4c40d4d21e84 · outbound

This paper cites Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.280542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.280542Z digest=sha256:e86f60111ae472f766a05cdc9e0e866ee4d85f359bbbd9be61fd1467273edb6e

Observation 5629f967-a2e3-4960-a258-fb87152d2e3f · outbound

This paper cites Magic-Me: Identity-Specific Video Customized Diffusion.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Magic-Me: Identity-Specific Video Customized Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.283252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.283252Z digest=sha256:92db27277304f17ca29dda5d30f57bc88c568e3497e3dae6514921940dafadee

Observation 7b7768e7-5bfd-4036-b0ee-afd2f5b75a39 · outbound

This paper cites T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.286112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.286112Z digest=sha256:7d9658a6d8e1dd2d90197009322c493890098effbdaa0c0a90b187819867bab0

Observation 7ae029cf-1607-405d-8b4f-4fdc1f92bf6b · outbound

This paper cites V oxceleb: Large-scale speaker verification in the wild.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition V oxceleb: Large-scale speaker verification in the wild

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.289015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.289015Z digest=sha256:313d1b1132c3fbf01888996f846e98bb4e6108c7e685a8a63d2fa8be271e0e48

Observation 718e76bb-b221-449d-8500-17b13c27d0c5 · outbound

This paper cites Toward verifiable and reproducible human evaluation for text-to-image generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Toward verifiable and reproducible human evaluation for text-to-image generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.291695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.291695Z digest=sha256:4b7aa911c7ac20390cf6e0f0986d515dcf0ea4b2a05a47b100ec33ede051ac3e

Observation e8c1dfa9-d079-4cef-a2ae-9a467e3b30fd · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Movie Gen: A Cast of Media Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.295231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.295231Z digest=sha256:d420834fb80ea2bd802603513571b90bc3892a737df7512529603a69e9bce74f

Observation 5a193d71-9f60-47f1-b3e3-c1d72e60e2f4 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Learn- ing transferable visual models from natural language super- vision

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.298356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.298356Z digest=sha256:ffd2a8f97522e5a30215b1646f8bb6551aa54b7301a2e2fab664bca93fe02e49

Observation 9894d1b5-045d-45ee-891e-cdf0d97035ce · outbound

This paper cites Zero-shot text-to-image generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Zero-shot text-to-image generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.301090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.301090Z digest=sha256:6f6a5243355dcf3f85a759cd108a4696e3f012a1e86e8115076d603f19bf2fa0

Observation bcc49a71-8317-4efa-ab22-343620b348e4 · outbound

This paper cites Rethinking Video Deblurring with Wavelet-Aware Dynamic Transformer and Diffusion Model.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Rethinking Video Deblurring with Wavelet-Aware Dynamic Transformer and Diffusion Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.303694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.303694Z digest=sha256:13a7e1f25094a29e987550c91c5fa25935a3ef5ad9650c1dcf7fedc4d6522cf3

Observation 6725d220-8b9a-4408-9c47-801ea0ba63bc · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition SAM 2: Segment Anything in Images and Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.306512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.306512Z digest=sha256:b2457cfd56879635a7b531a88a06838d0fe24166b7d0f46a17d17457c101f1da

Observation bc71d83e-778c-4940-8d7a-4b12d41690e5 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.309340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.309340Z digest=sha256:414d3a57b5756505553c22883f3ae3bc294135412f5113f18e15280918774644

Observation b8be134f-7a51-49aa-a531-c8802c49aba9 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition High-resolution image syn- thesis with latent diffusion models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.312118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.312118Z digest=sha256:42d041d3234204689b31fbc62455b9c309b57100d99e7209191852dc2d5f1bdf

Observation 4d01fc48-4e47-4083-bcb7-f49b72fd6535 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.315179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.315179Z digest=sha256:d1759e4b07dd24546bdc5b0851c4fab30b7bc35cc6cdd0c1579474334e9a7f72

Observation 16b6a115-2731-4bb7-bb90-2e93616ca0bc · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.317872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.317872Z digest=sha256:65f34a1709d9bb3caf0498c3311890b1cf2a647c19a0a22eea1442e548342bf3

Observation 34962758-460c-4aca-9b3a-85791136ef97 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.320544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.320544Z digest=sha256:0e32d3e67b799de5bb01f226ed7986e927e26a1ad7aa65c64dbe7545508cce7f

Observation 25b6ff77-ff17-42ac-9c63-8e89a8484903 · outbound

This paper cites LightningDrag: Lightning Fast and Accurate Drag-based Image Editing Emerging from Videos.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition LightningDrag: Lightning Fast and Accurate Drag-based Image Editing Emerging from Videos

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.324292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.324292Z digest=sha256:71991c802c2d9beb4e398f734f140e3bd4ac1cdfa0574f22c0a7856b1f65ff1c

Observation 6bce8180-8e55-4f97-939d-eb0df03d68b0 · outbound

This paper cites Freeu: Free lunch in diffusion u-net.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Freeu: Free lunch in diffusion u-net

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.328108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.328108Z digest=sha256:e790bcb55c54f6b8141637cfd6da14a07b27c89471cf6538de0b04ecc8a5b248

Observation afaac3b4-2a48-4ae1-a679-e05d2ca7f8d5 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Deep unsupervised learning using nonequilibrium thermodynamics

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.330802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.330802Z digest=sha256:26fc99e324112f9f16a01c184b04ce7b7dc51c7e18a4abc203d4ee1653489e75

Observation 355ca527-c795-4339-b96b-d2ba52084856 · outbound

This paper cites Denois- ing diffusion implicit models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Denois- ing diffusion implicit models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.360714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.335184Z digest=sha256:7b81a416edcf66b55d9dc14f920186fad4cbe7accf71c8f951e0aa387463ae41

Observation 2f75b500-4eea-4385-9a84-891b45ab226a · outbound

This paper cites Rethinking the inception ar- chitecture for computer vision.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Rethinking the inception ar- chitecture for computer vision

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.352176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.338597Z digest=sha256:7b41439492c436a5f348cfd24119111f2cff449221b9a31c2da7934b33dd0459

Observation 8e132efe-df6d-4f35-bdbf-155665106af7 · outbound

This paper cites Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.341396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.341396Z digest=sha256:718d49f2eaf0e12d1b4753a5cfabbca0345d5bfd97af1d36caea4dea8cc6180e

Observation 2b53f5f2-5606-4b45-a4f0-ed911523ecbe · outbound

This paper cites U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.344306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.344306Z digest=sha256:775bc2bcd5c4335ca3fabf407cb55591227a3ebd1b0a6b2b37d49e8b356a2783

Observation 93931778-6e20-42af-836a-81b311793404 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.348230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.348230Z digest=sha256:4155fc676af76398c483df7300943688837bab271415e15ad0bd362ed933518d

Observation a3da74bf-4d3b-4a39-a803-3c5307d73654 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.351254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.351254Z digest=sha256:d82a1771fa6b3d814d8b89760bfe7fe4fb2534cd4abe04c57d7eb3a27ce1ad74

Observation a0fc3ef0-fe18-46c4-8089-0723dfd9201e · outbound

This paper cites One-shot free-view neural talking-head synthesis for video conferenc- ing.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition One-shot free-view neural talking-head synthesis for video conferenc- ing

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.343107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.355058Z digest=sha256:0cc1f69d1096f829eb047aca465e7551d2a1d9fdd43dc30482a13bc852a43433

Observation 86c2a531-d16f-4958-a020-f60add71ee80 · outbound

This paper cites Customvideo: Customizing text- to-video generation with multiple subjects.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Customvideo: Customizing text- to-video generation with multiple subjects

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.357880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.357880Z digest=sha256:3f5f410722a89cba764971f3629e544799630409c03309bbb50e25d9423ca8c2

Observation 2ed1ce0c-614c-48c8-a84c-66a17cd16a0c · outbound

This paper cites Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.333942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.360611Z digest=sha256:93914761014e7fafd0068e3b6cb3787f00d5cfe0fc25982b2849cfbc146ce2f7

Observation cf792dd2-35c9-4846-aa7a-9f6720b453fe · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Dreamvideo: Composing your dream videos with customized subject and motion

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.325155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.363326Z digest=sha256:0bb72dae569e0b4963a1fb4b0766d84ecec7a355c32a2c3e4ecf38c81bc861de

Observation 07aa1095-693d-4394-87ec-0dc9f0d5d9a8 · outbound

This paper cites MotionBooth: Motion-Aware Customized Text-to-Video Generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition MotionBooth: Motion-Aware Customized Text-to-Video Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.366011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.366011Z digest=sha256:f3ca1a5b7f242054db67f96c48efa5114f5d3b1c5e1299dd872a4d147d76e75e

Observation ff9151ad-01c8-4f3c-b5d4-055ac1b25727 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.315316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.369992Z digest=sha256:be83cdcb8563a86ad27274cefdcff5f35916a69549ccffe92189e9ca4f208a99

Observation 3b0035df-556a-42d0-a748-2c8e81ab871e · outbound

This paper cites Autoregressive Models in Vision: A Survey.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Autoregressive Models in Vision: A Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.372712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.372712Z digest=sha256:cb6a54fbd493be26283650fc3171fe8f92b28e9e6d6f7130a1685e328a4b3ad4

Observation ee086857-eb1b-4928-8bc0-849e57061636 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Easyanimate: A high-performance long video generation method based on transformer architecture

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.375564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.375564Z digest=sha256:a4f94ce2ef0fa677f6bd886a1470f596f5cdaf2c0d8dd014a21780e2cb8801ab

Observation 0b29d11b-317b-45ad-ae07-3e8fc19ab70e · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition xgen-mm (blip-3): A family of open large multimodal models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.378139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.378139Z digest=sha256:f3b32079db30f334f4e4f7cec9697076c2a87b85fea2c9dd155b169a770cd210

Observation 7815a96d-65db-4063-870a-45442c7e6602 · outbound

This paper cites Ucf: Uncovering common features for generalizable deep- fake detection.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Ucf: Uncovering common features for generalizable deep- fake detection

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.306976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.381013Z digest=sha256:8433b036cf8d43e2dff5162749df995e2b66b174d9930920621e6adaaa2d5bdd

Observation d2d24c38-733d-4568-8187-8cde5a5a31a9 · outbound

This paper cites Transcending forgery specificity with latent space augmentation for generalizable deepfake detection.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Transcending forgery specificity with latent space augmentation for generalizable deepfake detection

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.297487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.383586Z digest=sha256:431176bec5f3009d76e19ac6719c5d9c5b45b710e992215e24a925949468504f

Observation 77912a41-63ef-4572-902f-85246471f1d4 · outbound

This paper cites DF40: Toward Next-Generation Deepfake Detection.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition DF40: Toward Next-Generation Deepfake Detection

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.386281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.386281Z digest=sha256:3a0b500930304a3f8187ca01997163262e0ada79ec1f3b6aeb58026a4de64fd9

Observation cd4b482a-eed3-44f6-8e4d-fa20f8210918 · outbound

This paper cites Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.389333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.389333Z digest=sha256:5f5a1511c2a8a2cbce8b4846a0629fce45fac19f254cc28e28d92c1a71db3660

Observation 61d4ac17-7192-40df-8b76-e794b3aa3036 · outbound

This paper cites Is Parameter Collision Hindering Continual Learning in LLMs?.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Is Parameter Collision Hindering Continual Learning in LLMs?

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.393339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.393339Z digest=sha256:715730cf82632ad0a59e2132fb9db0d0bf899d5031e73fbedaa39f200aa25fdc

Observation 292c531f-d710-4a65-8307-20a55105d023 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.396136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.396136Z digest=sha256:765f6601c753acecb793a104bf434f6276d1db5d2ea6f46726861a4cac446092

Observation 45ab19fa-3548-4a39-963c-788094c1d185 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.398916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.398916Z digest=sha256:b243998b613acaff18edc6bff65c5337114cb7bc7833d113ae3f3e3caab47080

Observation 8e9dd844-0b8f-45ec-a99b-a3eee5efe5f7 · outbound

This paper cites Celebv-text: A large-scale facial text-video dataset.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Celebv-text: A large-scale facial text-video dataset

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.289067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.401705Z digest=sha256:3a0e1daf7f59cbf025f971b208aecc94567c24b3601f17aa1afffce2c09dd4b5

Observation c96b841a-4d5b-4f3c-996e-6d86321146f1 · outbound

This paper cites EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.404398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.404398Z digest=sha256:e734ade7e452e73bb46d861aa9abb6496e6f9cbe9ab53cafbd5de33e912bb10c

Observation 138b37f9-405a-4511-953b-4bf4be403990 · outbound

This paper cites ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.407385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.407385Z digest=sha256:c0f787e09aa1900c2fa410085e2d2ce94b134494f36d4c198658bcf408f6d906

Observation 9bde9494-b3ab-4ea8-8af4-e0215d332ee6 · outbound

This paper cites PromptFix: You Prompt and We Fix the Photo.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition PromptFix: You Prompt and We Fix the Photo

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.410195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.410195Z digest=sha256:417ee85b83191cac80eb162c71d88af0505aa7142c37bba80ed1ef8467ba66bf

Observation 1c999eb4-805d-497a-a8bd-d6268108168b · outbound

This paper cites MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.413348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.413348Z digest=sha256:2bf1e0d787feff6640f25e34feb60f06122e836d17d5b75690cb8e169cfa50e1

Observation 3f0041b0-430c-4dc4-86f2-86272fb14e4d · outbound

This paper cites Chronomagic-bench: A benchmark for metamorphic evaluation of text-to-time-lapse video gen- eration.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Chronomagic-bench: A benchmark for metamorphic evaluation of text-to-time-lapse video gen- eration

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.279923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.416786Z digest=sha256:3d503b0c7e18bc718b4de874f28240aac485a45489314306be5e5b111587ad4d

Observation 1815ef45-f2c1-48f7-a3de-dafa8659d9e8 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Adding conditional control to text-to-image diffusion models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.419443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.419443Z digest=sha256:5cfd98953e6fa2d7309ec9c3d37fe44fe91b23486d94b8a678aa2a08d62fbbb0

Observation 24417da9-9ee4-43bd-b3ea-f64ef3ee9f8c · outbound

This paper cites Bytetrack: Multi-object tracking by associating every detection box.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Bytetrack: Multi-object tracking by associating every detection box

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.266966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.422097Z digest=sha256:35c7d5c42e3fbe11643a50cd65b0d4392031d8dac563c3430fa1574c09f6abae

Observation de62994f-ac24-4243-a4a4-89378a38a8ff · outbound

This paper cites Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.258359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.424825Z digest=sha256:cc27a60c1d1e56d74afec52599d8d6515a7e6721957afb393740fe2e490cdc59

Observation 3e25e70c-5698-4d2d-98e1-d3f937668bb6 · outbound

This paper cites Tora: Trajectory-oriented Diffusion Transformer for Video Generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Tora: Trajectory-oriented Diffusion Transformer for Video Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.427864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.427864Z digest=sha256:6a81436d3c0129d6466a6902ef6c764e019b421a5e147f52814349c8e7013ebb

Observation 309f281a-43bf-4e63-9485-20a473ebb7cd · outbound

This paper cites Videogen-of-thought: A collab- orative framework for multi-shot video generation.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Videogen-of-thought: A collab- orative framework for multi-shot video generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.430917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.430917Z digest=sha256:f197552b038566a18c7c5df6341b7efda37cdcdd6c571efb0bf41be732a9887b

Observation 44332149-83a2-4357-a81b-d6ffbd579a35 · outbound

This paper cites Open-sora: Democratizing efficient video production for all.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Open-sora: Democratizing efficient video production for all

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.250426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.434101Z digest=sha256:21d83b610bae9b05537d7358e7f89767f340a66dc6b85fb9ad8129d6c47365e8

Observation f647344a-a72b-4487-9aaa-b17a7f8bb391 · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.437705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.437705Z digest=sha256:aebb08098246effbab16a8334fa9c11ca70fc4da10f99acb13a82c30c40ab602

Observation d7c612bd-a6f2-4b27-858b-33b58e9784c8 · outbound

This paper cites Celebv- hq: A large-scale video facial attributes dataset.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Celebv- hq: A large-scale video facial attributes dataset

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.242501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.440592Z digest=sha256:beceed73ae8f7196cc52fbe92cff064389bebd88e8d974f6e8eed1717e51316b

Observation f07f4994-ab72-4c30-9e73-6f2532861887 · outbound

This paper cites Comparison with Closed-source Method.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Comparison with Closed-source Method

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.234373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.443204Z digest=sha256:3b2ab67efaf5f196188556bc972bc4480c3693595aeb20b8bcb8afcf79a96aa0

Observation c1cbe94f-536e-4e77-8ece-ca07c01e635e · outbound

This paper cites Visualization of Different Injection Methods 6 3.2.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Visualization of Different Injection Methods 6 3.2

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:28.224595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:10:27.446268Z digest=sha256:2541baa6059ce69c5c6c0468b111724c1a62b748817e63e931819b164b71a749

Pith citing papers

Observation d379e5df-9720-4c55-a0bf-0781afdd4dbf · inbound

PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation cites this paper.

PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:38:48.352472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:38:48.352472Z digest=sha256:e66573b58b20890c2a11dd5ee4da1ca7d4332ea92305b3238db1cf047b48c929

Observation 7ce45a66-28a6-4abc-9a71-afef92ef4e3d · inbound

WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model cites this paper.

WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T12:11:26.900571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:11:26.900571Z digest=sha256:b4ff8dcfd5b6bcdb4cb2373866b37942e13f514609261e66be645bb5cc3c85e0

Observation 94c1d87a-70f9-4ab6-ae4d-14ec73a6f39a · inbound

Ingredients: Blending Custom Photos with Video Diffusion Transformers cites this paper.

Ingredients: Blending Custom Photos with Video Diffusion Transformers Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:26.056014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:25:26.056014Z digest=sha256:6c3bca84a7c7ea4785e65c43467993cf93b26ebb757f075c1675343dc38f3fa0

Observation 654c9a00-8f6d-4039-9593-bbe9cf27e3e1 · inbound

EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion cites this paper.

EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T16:11:59.963511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:11:59.963511Z digest=sha256:76c2f69cd8a95928243959b7ef0396487c9893e3d3125a261b20c596c1bf03b1

Observation 157cdf74-9c98-48e7-9520-9f6312c91447 · inbound

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute cites this paper.

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:51:54.418911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T17:51:41.947939Z digest=sha256:79ed1111d3425d77045ac6b2f5522f1ca771f52aee043c9ec1cd4f8941b344dc

Observation 29436255-9733-4d60-a3ad-87f0001dae02 · inbound

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation cites this paper.

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:29:36.698279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:29:36.698279Z digest=sha256:cc91374a223cb8d2e81c2d808003391e1e9f4584266538effd9b597538dde551

Observation a0f306ac-9e16-44a7-9521-3256679dafe1 · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:17:45.442169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:a934ec8b30694ba5f557300d942869f13b6f9799fbb95d41b24074d8fed95cb3

Observation ab3fa46f-dd86-49fc-afe4-7e1a6d1947b2 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.479084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.479084Z digest=sha256:2d28e5e11cffebbbdd128076acebc1b01dbbb8153cb22ddf86679f85b7002f64

Observation 46a1ad32-721b-48e2-a3d0-d42d1fac1549 · inbound

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation cites this paper.

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.070137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T17:34:26.951644Z digest=sha256:2f56628acf7526e6d8c4e0bfddeae347f0af582b7ce8aa82e435a85184c87a95

Observation aaadd1b3-ad36-4399-88df-ec3aae29cbb2 · inbound

UNIC: Unified In-Context Video Editing cites this paper.

UNIC: Unified In-Context Video Editing Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:42.872478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:42.872478Z digest=sha256:a1494b81efa06ccd20b044c5e0f6ca9224a47b2bbab0fa4e7f383d05051e9d9e

Observation fef7acec-22d0-4247-bda7-769ee6ae4dbd · inbound

Follow-Your-Creation: Empowering 4D Creation through Video Inpainting cites this paper.

Follow-Your-Creation: Empowering 4D Creation through Video Inpainting Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:26.982019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:26.982019Z digest=sha256:91709ce5e22e3076f0285014438516cfa13b4869b776042790d72534a9d4a7cb

Observation c3761732-b039-47c0-a26f-1e59f3b149d6 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.052194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.052194Z digest=sha256:045cc74f9cf3a635097349b4c2e10c1fa185d647f1ccdd14c3e9d37414237628

Observation 241d36e5-0297-4afd-97b0-cd103d00fd43 · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:39.711058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:39.711058Z digest=sha256:afbc30a4182ab6c3a1416e297fcdd74f14aea25e4d449b0f25a227ee69cc97fe

Observation 75be7455-bf0e-4288-ad58-2f692377874b · inbound

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation cites this paper.

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:10.514105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T07:59:49.398271Z digest=sha256:48e161844396447beffed05fdd5bc3b05b42879e91bce9c2f0840cc49a48f4a4

Observation 9a45fe12-994c-4630-b9ce-94e3317f18ec · inbound

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation cites this paper.

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:19:32.838413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:19:32.838413Z digest=sha256:276566e1b894ac089b35fe24f7186a60043f6535c8eb0bd5f5b8022658f6a565

Observation 368fc191-8da6-4ce5-9ed0-67e9b7be3445 · inbound

A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea cites this paper.

A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T04:44:59.127465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:44:59.127465Z digest=sha256:c4fc76089043abdf561681e3322a523f77f41e654a290be64bd220a8bd4e171c

Observation 1d4434e8-b090-4277-b226-97d0e06fddcc · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 196

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.584551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:6c534d087b5465c940cf533de9320e043c0b4d8c5bc472e3ba98d49ece7f1b58

Observation 0975307d-bb4c-460d-ae3a-0e738b08c3fc · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.483661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:cfec959d2cb4b1ba7c70a1b2c53163f5614adbd447f6f645dbcf00c092de40bc

Observation 609aa3f1-5b4b-42f1-a784-2f903f394879 · inbound

MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation cites this paper.

MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:55:10.368181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T06:53:13.152677Z digest=sha256:bc9d239c90ad64b42f34906b4a38771824e95586ae16b30d12a5f7d45d5b6dec

Observation beaaa26e-a4a8-4af6-88e3-1f09932a983f · inbound

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices cites this paper.

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:43.795138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T20:00:27.987481Z digest=sha256:8a9e69c459a0bb78f4a0d3765530dbc1e16d1f8aeb0a6b0966371c35a8329d9c

Observation 22a7e0d9-e635-464f-ba6d-a50f53319213 · inbound

Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy cites this paper.

Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.308647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T22:41:06.600948Z digest=sha256:30241e8ff1b2e6ab85609dd61d7fdb6f9e69b6951d8a128e773c6ad12cdbe75a

Observation 5a9005f6-bf64-40aa-a5a7-0e02b26ee519 · inbound

A Comprehensive Ecosystem for Open-Domain Customized Video Generation cites this paper.

A Comprehensive Ecosystem for Open-Domain Customized Video Generation Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.902889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T10:06:33.572182Z digest=sha256:a800baca4f9e0d492aeeba7fd3e97017c5c0c569072da85f584a43921cf7e55c

Observation 2819a73a-0c3c-428b-8e6b-a5cfa6e6ec2d · inbound

Customizing Video Portraits via Identity-ActionDecoupling cites this paper.

Customizing Video Portraits via Identity-ActionDecoupling Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.423833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T10:56:11.553110Z digest=sha256:5d2a32caad374693cff629b9d36c31445c001cf2bee84d716626c45d40917f07