Pith. sign in

Paper Citation Record · LEDGER

Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 85 inbound Pith citation observations for arXiv:2405.04233.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.04233 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 85 of 85 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:55:41.631468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:11.362384Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 18bbe2b2-819f-4c89-97c9-ff43643b7d4e · inbound

Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data cites this paper.

Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:06:09.126266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T12:06:09.047780Z digest=sha256:42e4b1b5d0ee0ef0ab3c80f8ffba40634d85cdcc077b941ed8ab78164ba59fe9

Observation d83fc5a1-b65c-4edc-b103-26e11cd5cd2f · inbound

High-Resolution Image Synthesis via Next-Token Prediction cites this paper.

High-Resolution Image Synthesis via Next-Token Prediction Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.631468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.631468Z digest=sha256:8b0ae25e38a433de7ada540a971945c756c2e85d5896454dc5b7ff63c9b88638

Observation 723be1c0-4869-42cf-b611-e11687c59717 · inbound

Identity-Preserving Text-to-Video Generation by Frequency Decomposition cites this paper.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.163774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.163774Z digest=sha256:08da6995e2e88b244296e1155d35266c2f1e6ff691306d761375212289ef3630

Observation f5b28f82-0f24-47b4-aade-d7b4d6c1a2cd · inbound

WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model cites this paper.

WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:11:26.698538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:11:26.698538Z digest=sha256:6018e70a91799155b57c1cdb52cda6d8462f1d38e271148911655337d37c6b43

Observation 427fb57a-4732-4d77-be2e-eb4973b5f6d0 · inbound

StableAnimator: High-Quality Identity-Preserving Human Image Animation cites this paper.

StableAnimator: High-Quality Identity-Preserving Human Image Animation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:56:28.104758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:56:28.104758Z digest=sha256:f1249986ddbd197a665d826258f070b4df9ee258923bb96d0c4df2f0fe24eb25

Observation 113b5598-9276-4780-9ba0-f607cb288791 · inbound

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook cites this paper.

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:47.617707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:08:47.617707Z digest=sha256:a1016b022789294c7ddbc57d3826dd2abcc9072f7313c9e40abea023d27d3459

Observation eecc9c12-7171-4315-9c38-f4f58459c13c · inbound

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation cites this paper.

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:44:52.385078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:44:52.385078Z digest=sha256:b3cce6357f70c09e95b67b8ffa28853e3c4c5c9aaa59142438c8f52717ef19a6

Observation 2931710b-9f08-4443-9809-b6230cff4857 · inbound

Human Action CLIPs: Detecting AI-generated Human Motion cites this paper.

Human Action CLIPs: Detecting AI-generated Human Motion Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.604699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.604699Z digest=sha256:e1b0c90f919b65a927d53d9f43f67c446c4f17c0631aea2be3bef0bb77d18da1

Observation 46889338-b496-4e84-bcd5-aaebe0604597 · inbound

Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer cites this paper.

Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:06:43.778769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:06:43.778769Z digest=sha256:dee1f855b21b7148e0bf949dc253c9071c21053fd3aadbbbf0f55e110c429cfb

Observation 18176a46-c1fd-4228-b5f5-1d1e321a218a · inbound

CPA: Camera-pose-awareness Diffusion Transformer for Video Generation cites this paper.

CPA: Camera-pose-awareness Diffusion Transformer for Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:27:23.817811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:27:23.817811Z digest=sha256:91e6060b8656cd999720423b4dcd2c4a3c54c837854672d538076470bfbdf3f2

Observation c7101dd3-137c-4bdb-9ff6-816ca157ce1f · inbound

Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis cites this paper.

Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:34.294210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:34.294210Z digest=sha256:ff84fc3683fd15349fbe1d71155c6d1c9133e008fb76ec8b02516decddb391e4

Observation 2da0f429-201e-4bbc-b52f-3584b8011976 · inbound

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner cites this paper.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.712836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.712836Z digest=sha256:7940ee0a653d791d523c8057818e93f48e667cc23a6ab67959a9fe2af12d8ec9

Observation e95ef856-03be-42cd-9fd0-226ac38cddbf · inbound

Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM cites this paper.

Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:29.286042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:29.286042Z digest=sha256:20875b9bfb733425179f5bbc8376997fe95cad29fb5f7a2a2a0540cfe5f485c2

Observation ac241a6e-8ae8-46ba-be9b-5709a63748c2 · inbound

Large Motion Video Autoencoding with Cross-modal Video VAE cites this paper.

Large Motion Video Autoencoding with Cross-modal Video VAE Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:13:47.282586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:13:47.282586Z digest=sha256:aba02cc0101809dc3b1ec56a4a41add3e05eb5c225eb7498c95c7cc968eb929c

Observation 8f3b9f71-58d3-4ee1-b45c-787a4af160a9 · inbound

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution cites this paper.

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:04:48.570035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:04:48.570035Z digest=sha256:cd086151c1e8b7e6a883843177fbdca1bb64a2d77fb1147f2084a26e44f2a15a

Observation 283a21c8-50e2-4fb2-a1ff-9040b0af1c90 · inbound

RepVideo: Rethinking Cross-Layer Representation for Video Generation cites this paper.

RepVideo: Rethinking Cross-Layer Representation for Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:16:58.768754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:16:58.768754Z digest=sha256:e4c4eb27d956ff82515eb0e86b4c5b351cc00f3a498a293b0f4cff18d007ca10

Observation a22fa437-87f1-49b6-bc8b-c3aad126c97a · inbound

VideoShield: Regulating Diffusion-based Video Generation Models via Watermarking cites this paper.

VideoShield: Regulating Diffusion-based Video Generation Models via Watermarking Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:21:05.933448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:21:05.933448Z digest=sha256:7ab0fbefe0f7a6ff3f4a46be97e102b883d343016ef9be394887d689e776f2b5

Observation 3f1a83bc-06f2-4c58-bf62-a40c1d175703 · inbound

Elucidating the Preconditioning in Consistency Distillation cites this paper.

Elucidating the Preconditioning in Consistency Distillation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T10:42:54.348203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:42:54.348203Z digest=sha256:57290cbb062acbf27683b2399595aa105796df71b72a13380dfe0012e2a46d62

Observation e12fadcf-51da-4a5c-b584-4ab1e0be4858 · inbound

Seeing World Dynamics in a Nutshell cites this paper.

Seeing World Dynamics in a Nutshell Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T04:42:05.761410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:42:05.761410Z digest=sha256:6d56c0fe8c266494b934704986f7ecbec7d6512b727407730acbf1de40eff625

Observation b5b4c075-a60f-453f-8102-d9a4c7949249 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.168829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.168829Z digest=sha256:ec720af803aadb72673ae723f82bc5be1bcac2e02fc18f506406a8129eb4730d

Observation ff15710d-5d38-4a39-9939-51666f77cd8c · inbound

AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance cites this paper.

AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T10:09:43.574143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:09:43.574143Z digest=sha256:eba6d34b94173ef31bace1b2465685c2925d38009d4eda60d2d2ecb087eef277

Observation fc8f9a49-89b8-4ed5-918f-0115d6e9b55d · inbound

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control cites this paper.

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T19:40:21.146056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:40:21.146056Z digest=sha256:cacf21f51bbb29d9cd0b5fd48c082284c149712441f85e602befbb8a724e95dc

Observation aefb1f88-d662-4015-8ba2-8f0d9d14bed6 · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:07:14.436097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:47fa7ac98778bab2b3c6cc0e56ce4a44497acb76deb04b803b0607d4512a685f

Observation aabdfd97-d1db-47a9-a6b5-03483b016066 · inbound

MotionPro: A Precise Motion Controller for Image-to-Video Generation cites this paper.

MotionPro: A Precise Motion Controller for Image-to-Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:02.311605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:02.311605Z digest=sha256:a510fea7a1fdf73f7c01929849815411f6a6bb59c1874db428017f69598ac33f

Observation e211f830-c275-4c09-812e-117d517cad8a · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.048350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.048350Z digest=sha256:6c4ec3a0965cf814551eefa1706e1c563a93b8af5fd5d9e30728c53d9380602e

Observation 3f76c878-9e6d-4e47-9621-6bc259ad158c · inbound

Versatile Cardiovascular Signal Generation with a Unified Diffusion Transformer cites this paper.

Versatile Cardiovascular Signal Generation with a Unified Diffusion Transformer Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:00.981355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:00.981355Z digest=sha256:7cf87e72a9880d46db04e2976d61dd86f925a94451b2bedf4d2cd2ffc8476e0d

Observation 108ea320-d8ba-4fa5-945f-d231ff320481 · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.021385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.021385Z digest=sha256:655d069d2fdcfb11f870152f1f21ec6a782cf7c29a937f44beeb41774798fab6

Observation bbf9ac7e-42af-4848-945d-545a386a0e00 · inbound

Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution cites this paper.

Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:18.834828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:18.834828Z digest=sha256:28ab375f6bc5b3cd1786691b82878f9f53569033365ff15a00a1ad6b5cb3f6f2

Observation 5adca448-c3bb-461f-997c-2bfb5c6fbb01 · inbound

LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model cites this paper.

LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:59.897735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:59.897735Z digest=sha256:82779aff073f0b1e11dcd172caae378283050d79535ceec5bd2ee4cc20f4e855

Observation b48ef145-51f6-43df-867a-cd77217e437f · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:55.744800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:55.744800Z digest=sha256:866b7cfa1861b1068b64639e77646d861716b777627b6cb009320b174a70f851

Observation 96500013-c643-49cf-bbf2-11d951ccd5f1 · inbound

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval cites this paper.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.523287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.523287Z digest=sha256:b01e5de93ab0d9c1a7deed2c82b948648539a075379b5a4a2599168c782780e5

Observation a6c3ec74-a015-42f8-9241-a2e75cc0d662 · inbound

FastInit: Fast Noise Initialization for Temporally Consistent Video Generation cites this paper.

FastInit: Fast Noise Initialization for Temporally Consistent Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:53.402851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:53.402851Z digest=sha256:e5294cac2e288d1277d182094aadb7377807d31d625a95ce0bb0a80cb2e4fe68

Observation bf5f1056-59cb-49c0-84c8-15847e188534 · inbound

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion cites this paper.

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:18.123749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:18.123749Z digest=sha256:cac2e076c3ffdcd79a9cf2dd7dea50aaf36755c0a1c198784e0f007ffad8b184

Observation e7c00d3a-a334-42de-98db-1afb1b1e3d4b · inbound

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation cites this paper.

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:19:32.682892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:19:32.682892Z digest=sha256:2679928e62b2ab4c72f25b9c4227568ebb8fec3850c192f48e2268d4b47a78f9

Observation 545dcd23-e751-45b1-8d86-22b24255c284 · inbound

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation cites this paper.

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:52:57.579444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T03:52:18.984005Z digest=sha256:b7ed44c1a5f7823f493f9d2f5d41fc2b0548c76d4cf777a6e81f9453965f0413

Observation 313702cf-e86a-48ed-895f-f83b9e3c3c4f · inbound

Vidar: Embodied Video Diffusion Model for Generalist Manipulation cites this paper.

Vidar: Embodied Video Diffusion Model for Generalist Manipulation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:54:28.303206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T09:54:28.271928Z digest=sha256:8e3f3c15598a76fe5aa29e687fb3621ee140353309386767c1b5e76fbea5e94c

Observation 19fcc1b4-f978-469b-bae8-ef51b579d778 · inbound

StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation cites this paper.

StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:24.337873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:47:24.337873Z digest=sha256:390aea6e9c6792749b0983b24a91f186a1f4d9f750d47e81e2a7f986a1f5de6c

Observation 36bc65c8-c9c2-4d2d-af15-bb7aad95c0f7 · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.068815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.068815Z digest=sha256:9fa7414efb8ac6c32ab85194a562628311929ad366fba01acc26380f9f4f02f9

Observation a7ef3343-70ed-4b77-9ca3-e3a861e050d7 · inbound

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement cites this paper.

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:42:39.584121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:42:39.584121Z digest=sha256:9b201a7131beb4d4be89f3da62520a5ff6016fd048a4a9b84fdca11180fadfb1

Observation 86905ea5-c022-43bb-8b50-c37819f87f6f · inbound

RewardDance: Reward Scaling in Visual Generation cites this paper.

RewardDance: Reward Scaling in Visual Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:08:55.396895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:08:55.396895Z digest=sha256:6bd72b36d75951a87bdeb1c48281279f0c1c2cee1be9ec91ffacd70f04d8a46b

Observation 98940530-59ab-4151-b584-309fb0b43ae5 · inbound

Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency cites this paper.

Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:51:09.187561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T08:46:16.541104Z digest=sha256:f075f2324348f344584ec970af2a5cb27f58df4554cd3ff09c5068e1c46d91c2

Observation 5c7c8dea-5d39-4e18-bf50-10a8b8b27477 · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:50:13.046825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:a8849f968a77c4367199f186d4bca4b9b80a9a53451b9362255429649cd4a14f

Observation c91e4821-a5a2-4b44-be0c-2c2daf15a2e8 · inbound

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation cites this paper.

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T08:16:22.346891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:16:22.346891Z digest=sha256:8d1edd0161725d19cce0da3193d4e8b467e59563666ad053561d2ec3b7b8a428

Observation 96da0831-e0f1-4f25-ac88-8c26074f3cca · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:32:01.738632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T17:32:01.642256Z digest=sha256:dd18cf97a2bbbc2ba5d206a7b0081b447cf8521f05ad95dc2546839c561eece9

Observation 6b6b3205-e9d1-4614-aa4b-8fcb451277c5 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:51:29.779880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T11:48:00.633421Z digest=sha256:2b9652098b1357fe80fe75af813d0569fc16adecede0063c2e4d40e99aff73b5

Observation dc357c4b-d5fa-4e89-a476-b8612bdc3696 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T05:30:23.868383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:30:23.868383Z digest=sha256:3cd92b06c57b28e04c0d6d13df15d86855d19bb7fc771294f19141fe40aca007

Observation 29457579-037c-4ef8-9875-4c6349a85fd6 · inbound

Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation cites this paper.

Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T00:03:22.328973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:03:22.328973Z digest=sha256:4a67940ea989dccc94b1b27dcf693a21074312fa5576a2487ab4402d9a235c65

Observation 6dad6221-0e31-46cf-9a7e-fa9422ad42ad · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.449120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:faf1ad52178ce4ae1051d80849b567df611758336825d1ea17c9e2e09706c9b4

Observation 28a8d85c-9609-4c7b-9cc7-c36e8b190dbf · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:61dfaa392203fe6d90f7cadcf639dca0e55c949b502b817782404f3fda92d879

Observation edd8279e-e4f4-43cb-93fa-14dc68f654bd · inbound

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation cites this paper.

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T12:23:05.876881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:23:05.876881Z digest=sha256:7cb76cd4d1b6a2998c114f827e8f2fa47db5f236a73626ced6076ba56ae66260

Observation 906edc20-62e6-4da1-ac04-1acada01650a · inbound

ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos cites this paper.

ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:00.846669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T17:06:28.335731Z digest=sha256:8e48409ceda1e7b364eb4e60a8290cd7648192ae8ccc66409a18a73c5c0f0fc7

Observation 820dcff3-2566-44df-90ac-38a0a4466f14 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.962535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:d4918528a397bfb890ab7a3b1a28dbee0b44366397ecaa19147a077408bf4e0e

Observation 99af4c32-bd3b-4f48-814e-e8f49e99c200 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:b1b14f29c31f2b75c94e53d582105f9f6ff568cc844f2dcbefe91d220e6be6ae

Observation a2680bb9-fda6-47c3-b0d7-9a66a13d4cd0 · inbound

ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception cites this paper.

ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:02.297387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T14:56:00.679543Z digest=sha256:fe92996ac4a6e7ebea034434df1e93f0dca08ad4fba1881a937168cd6c33ee1a

Observation 366aa9e4-766e-44fe-b0c0-2613841ac62d · inbound

ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception cites this paper.

ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T05:32:16.069091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:32:16.069091Z digest=sha256:c0f5269e80a55ffd0d257a7a21dd71e82d18977d68a22e3ed56bb2f2030b8bcb

Observation 73c73cbf-a0dc-43f8-92b5-adff13846f15 · inbound

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement cites this paper.

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:05:22.877393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T04:41:45.903419Z digest=sha256:34a8c7b956f952005ffa4593748875e68c96b23d2672eb07c2812efa067558bf

Observation c9450f4f-9750-4f84-98a9-c895ccfc8914 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.506068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:ec9ff4b05536215a097ef587f2eca42f8dc8fa801dbf4fdc47d33b6f487a3841

Observation 1ac8e706-a259-490e-ac8c-eb808e75e14f · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.854281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:c8ff759ba05133e2d2f7725c8d4e1ed8eb58e6b846db0d88143d72bbe5e2fded

Observation 315b9a4c-bc63-4560-9263-9f3aeb421230 · inbound

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE cites this paper.

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:25:47.686249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T18:26:58.696936Z digest=sha256:e85236b538938db1008a346d9d6175d22c39f7cac878ef03cf6b4248840b6d1a

Observation e00cfbf3-fd1a-4057-9b84-cc756a447b16 · inbound

AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics cites this paper.

AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:56:32.177838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T01:23:27.990472Z digest=sha256:b59b199b63f749309c956fae6df7a6877b09e8bce302b226b57c16473c4b64d4

Observation 5fb29813-602e-4cc0-a541-577187e41b7f · inbound

AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics cites this paper.

AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:34.461197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T19:04:02.414172Z digest=sha256:5823c3463603ea0102781ef0c12e4b8e52ea33d7f80fc0e23731aba17777ce2c

Observation bdf2acaf-39ee-4f07-a0e0-1d4eb3148872 · inbound

AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics cites this paper.

AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.248845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:39:47.335022Z digest=sha256:3ef95b107a8c682ec737436747b2dac4d68d51121e14de0c017becb80859047a

Observation 4b7632f0-21ad-41f7-a338-9b3ef276d05a · inbound

FreeSpec: Training-Free Long Video Generation via Singular-Spectrum Reconstruction cites this paper.

FreeSpec: Training-Free Long Video Generation via Singular-Spectrum Reconstruction Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:07.733835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T13:11:46.079614Z digest=sha256:a0c4f070a9b62e96e09b82506617a758926a62ecd3f0efef9c6e9ce46ab72ac7

Observation bacb9b99-890a-47b1-9cfa-119a0473e048 · inbound

Advancing Reliable Synthetic Video Detection: Insights from the SAFE Challenge cites this paper.

Advancing Reliable Synthetic Video Detection: Insights from the SAFE Challenge Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:57.484674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T00:53:18.735506Z digest=sha256:73629eb95056de8dd301763f89a037bc5a475ed1b8e858638b7f91b9429eee1c

Observation 1c0305ce-01e1-41d5-b4ce-9f1a3937770d · inbound

Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers cites this paper.

Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:56.284351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:58:25.291448Z digest=sha256:c091a0b17a701308dfedfe1f03799a81097a751280cf3f7185a0dd7530a07088

Observation 6ece8cb1-63da-427f-bf15-1e745f273244 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.797687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T01:52:14.874049Z digest=sha256:f90cafba3de22e0710da430b38c9bd6e77382ce33260c1d1c627d3da5088ac91

Observation e0883180-b879-40c3-a4f4-a89f27e48976 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.448785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T21:54:33.902256Z digest=sha256:684eda63d53664504fe52cab3f0cd058d3bc9d3b76d9acf0538866bea2711d7c

Observation d30ec527-f7ec-401b-a07e-30a280e6b492 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.673014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T09:12:35.777810Z digest=sha256:50212edc47220d3ac12583c7d64d03f454a3f584a6ff3ca490ea7d38e8675017

Observation 276a741a-360a-44a7-a800-be7d3fe1e9f0 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:45:05.499291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:44:21.427570Z digest=sha256:04e188a59055e798f11facd5c9684b3f784b697f74883d1156bdd4f95521cc80

Observation 5d5d664b-2c8d-4fa7-b654-8779667e78a5 · inbound

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation cites this paper.

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.406243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T20:59:34.847496Z digest=sha256:635f3f82873102f05b963ca8629fd306e65158efa651ddae69e7f0140091899c

Observation caa404be-6042-495a-af23-765adadb3387 · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.963760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:bf20691f793200760bcb08776a9bb3ea6a7367b21048ac058fd59ed4b33a361c

Observation 1d7cd6f7-7a09-469a-ae35-a5751b533910 · inbound

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models cites this paper.

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:14.795387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T08:22:48.074724Z digest=sha256:c761b62249f9df904c91b403b1d7261a8ee63cc5d40395229357a1e2f27d1869

Observation 56fda64e-c019-4d5e-b6af-9dc56e19871c · inbound

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation cites this paper.

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:25.196547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T17:05:57.685728Z digest=sha256:095e3720132b7e1226da18d7d11522ae8e410d8af9f75696083dadd6101c8bc9

Observation 46c5aad6-4827-4c51-9342-8c0eab89d6ad · inbound

Explainable Forensics of Manipulated Segments in Untrimmed Long Videos cites this paper.

Explainable Forensics of Manipulated Segments in Untrimmed Long Videos Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.987409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T14:49:11.015838Z digest=sha256:6456d37c164ab484eef899506cde0011a749661961637485be7685acf41d785a

Observation b72fe840-58c4-4a2e-a7ce-0ada4fa6ed0b · inbound

BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression cites this paper.

BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T11:59:12.867166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:59:12.867166Z digest=sha256:b675d2453f067189617eb31f12ffbfd418693f77c758cd095da1e717cf86684b

Observation 3f0cf397-87ba-478d-ad4a-7d92ecddd226 · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.973433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:690e30a8c368105e62661b2da658917951700601f3c2da638b15fb05754e0a60

Observation ca3c9ce2-55ce-4f24-8cad-efdb17ea1565 · inbound

Pulse: Training Acceleration for Large Diffusion Models with Automatic Pipeline Parallelism cites this paper.

Pulse: Training Acceleration for Large Diffusion Models with Automatic Pipeline Parallelism Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:39:24.664412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T19:36:14.248003Z digest=sha256:80baa4c3839bef905e9b11ccfbad7b48d56c9ba5f5546b91688faab066b01392

Observation 8070024f-e6c1-40b7-a7e0-4f4c5f45b318 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:30.253059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:15:51.862192Z digest=sha256:ad8266ebf1450e1ed0517fa51b52f8ed8a70e3da2479649d5aa870c095a1ded1

Observation 63236064-2202-4cfd-b6a6-927e9602c6e8 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:04.099603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:04.099603Z digest=sha256:6142b5f7fbc0d4beec71770a6e885dea08352e806aa73a282cb2f6143bfa9541

Observation f1e06150-11b4-4595-9c6d-9954ef5f4660 · inbound

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation cites this paper.

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:05.888697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T21:33:38.643889Z digest=sha256:f1315b407e3214aae632470d4e91bcc3530e115ce4694d4e0673e3d31fce8189

Observation bfb82141-da1c-4737-86d2-ac420b481bab · inbound

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models cites this paper.

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:11.364149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T20:57:30.765802Z digest=sha256:2518e5ba9894067228eb0e1ebf3bbc2a8bc54301fec856da069d9fa994833f29

Observation 23814f50-bb84-46bd-bfab-beb691e994f2 · inbound

MemLearner: Learning to Query Context memory for Video World Models cites this paper.

MemLearner: Learning to Query Context memory for Video World Models Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:41.915950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T05:30:56.140465Z digest=sha256:d2b5a082f4b87cc4155174bffcf05ef4e5e27c4e09f2053c2892bf5cec6fd115

Observation f9db2bae-7545-4317-a7c5-37aff2e01a3e · inbound

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation cites this paper.

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T03:12:33.331799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:12:33.331799Z digest=sha256:9c4c440719cc8f4321c9ef6a21fc8ed2383899b7ea44159c9fadeff52607ebcf

Observation fec10dbc-a174-4198-95bb-6b713a840974 · inbound

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation cites this paper.

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T20:13:31.342466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T20:13:31.342466Z digest=sha256:8876282c46c66627e48b76c4b258e06e22a474b91289b3cc77e32d15202a31a6

Observation 8c14f1c3-cab0-45e0-b77e-7aa0ca7795d9 · inbound

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity cites this paper.

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:53.366492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:23:53.366492Z digest=sha256:75d654889edf2e30e5dbf09a2404540eb5ba5b096a6136a219d6aedc08a11798