Pith. sign in

Paper Citation Record · LEDGER

Wan-S2V: Audio-Driven Cinematic Video Generation

As of 21 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 32 inbound Pith citation observations for arXiv:2508.18621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18621 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:24:59.444766Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:05:47.808077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.703864Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4968f6d7-9c97-4a4f-b775-bba1ef82af99 · outbound

This paper cites Qwen2.5-VL Technical Report.

Wan-S2V: Audio-Driven Cinematic Video Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.584884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.584884Z digest=sha256:911c844bb316ce606ce2db1e7ccab30f8c766c03eabb42f48787f37340f4854e

Observation 7fb76f94-98a9-4e52-b57a-a60ac840fe62 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.824740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.824740Z digest=sha256:2f6c272044d11f9cc79c8f31c25405d86188ff4522854aeb7a6ce9ffb69642ca

Observation 04bda69f-6eb3-4468-bb05-9d351ea780ed · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Wan-S2V: Audio-Driven Cinematic Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.874741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.874741Z digest=sha256:f1cf7a35e2f5913d65a18585bd8a4f201a24603d026eceefa45c813cf06d648f

Observation 6897fdfb-9ade-43ff-a776-5b1078b0f010 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

Wan-S2V: Audio-Driven Cinematic Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.025790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.025790Z digest=sha256:eca9264d2b3350f1ffc8153466cf3792fa15998bd44b1d00fc96cf48eecd9871

Observation 3ae816ba-4882-42a5-9c5f-4b5e41715aa9 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Wan-S2V: Audio-Driven Cinematic Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.084742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.084742Z digest=sha256:6bde2332555f7ad6194f51e9cd53bc88a1717c5c338412e2276629cdd5b213af

Observation 54d8d8a5-a960-4da1-b302-d210086694cd · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

Wan-S2V: Audio-Driven Cinematic Video Generation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.254749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.254749Z digest=sha256:592bbe829fc45b470586552f816ee925641714d30c3c3c13e6b5d58b5bc0d5e4

Observation 163e6c2f-8abb-466f-b17d-b21ee7be9c84 · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

Wan-S2V: Audio-Driven Cinematic Video Generation Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.301375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.301375Z digest=sha256:e05f4bd797722deee5ad3849e335b6a5f1d7808d09a52d829009574bd77cc856

Observation 0ac0e266-00bd-4c76-9f87-ba7e83315f05 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

Wan-S2V: Audio-Driven Cinematic Video Generation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.444766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.444766Z digest=sha256:09b6bfe4d410a399b3bdd9360b03e2a3b59c06a2d210465b982174e8efad677d

Observation 13af52ce-5fa9-44ae-8acb-0f90b09b0c9c · outbound

This paper cites Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin.

Wan-S2V: Audio-Driven Cinematic Video Generation Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.354746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.354746Z digest=sha256:33208b07b1b6cfc86437ff62936cace0c2d64935025f30cdec9e4b36ec2f0cc5

Observation ca1bd54f-4053-41a3-b5f2-245a5bca4ded · outbound

This paper cites an unresolved cited work.

Wan-S2V: Audio-Driven Cinematic Video Generation Unresolved cited work

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.774990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.774990Z digest=sha256:06339306f0dfca948c2760dbb6670ef30a764fa1b8fa82c47cfd66c024b00cc2

Observation b7516b7e-1033-429c-b33c-0060dd4bfa65 · outbound

This paper cites Christoph Schuhmann.

Wan-S2V: Audio-Driven Cinematic Video Generation Christoph Schuhmann

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.974740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.974740Z digest=sha256:3547b0c677e52968dc9cb3bf3858bf55852f43bbe50d124706c1641f3ca6a5aa

Observation b29b68cb-f0fc-44d7-9b16-1208154a8507 · outbound

This paper cites Image quality metrics: Psnr vs.

Wan-S2V: Audio-Driven Cinematic Video Generation Image quality metrics: Psnr vs

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:25:02.079683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T16:24:58.724742Z digest=sha256:f047e7fce77f5ba13bae09627a0952801ae137984e353ca75071341a7da4f526

Observation d77ea705-6962-4e0f-b2a1-7babb4f8852b · outbound

This paper cites ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation.

Wan-S2V: Audio-Driven Cinematic Video Generation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.404743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.404743Z digest=sha256:7f37be62215ec037ae13c09cd23d1bb6ef5d2fa4d7d4413a1c4d2e19f6c0621d

Observation 90c58207-0cbf-4150-bb93-e2f38d6b30c1 · outbound

This paper cites Flow Matching for Generative Modeling.

Wan-S2V: Audio-Driven Cinematic Video Generation Flow Matching for Generative Modeling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.924743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.924743Z digest=sha256:d7e2293a8712a29491c212f1067a24e97718b1a4056dc80c18ed48cff02d7a6f

Observation d377db72-a42b-49dc-a21b-bcd2603004dc · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

Wan-S2V: Audio-Driven Cinematic Video Generation USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.674741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.674741Z digest=sha256:022d7aae30a66bcb7779cf328f12d77498e7e3b9cf2e8f56892b11cc96229923

Observation 3e091873-51f3-4079-b666-ccbce3d342bf · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.625126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.625126Z digest=sha256:9f578a79b5d188cb22476631da58b2693ea337b3112cfb7cabcdb607c562e94d

Pith citing papers

Observation 823551b6-d303-469f-a713-b7cff82979e0 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.808077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.808077Z digest=sha256:44dae6ecf824390a42bae4ebcaed04917989ea67642e4a64b116ed6815bf1275

Observation 31fba0ea-5d3d-4267-8785-c127026cd2ba · inbound

ASTRA: Let Arbitrary Subjects Transform in Video Editing cites this paper.

ASTRA: Let Arbitrary Subjects Transform in Video Editing Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:16:14.277804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T10:13:10.141426Z digest=sha256:b3efb203272a068ef9e51ff9938de6b67a582e0d926fc2700de2b4e7c0e1beb8

Observation 6e6f1077-bff7-42c6-91f1-b22c9c3940db · inbound

Understanding, Accelerating, and Improving MeanFlow Training cites this paper.

Understanding, Accelerating, and Improving MeanFlow Training Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:39:16.078524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:39:16.078524Z digest=sha256:bb60b3df57ea335ff0051b670a1800efd1832e4d14efb71e42b715ca561b3f3a

Observation d3775d6b-28f1-4e2d-8dd4-07c78e054cc5 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.463818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:53dcb62c1b598e4a5e69a460bf1d82e0afa6a00faac6d3c3c81048e6773ad6ff

Observation 81c585a6-fcc6-41ec-8429-171803c4fc2b · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:41.295310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:41.295310Z digest=sha256:5c287ec8ea9ef8827bb47a10efe93a90cdf6315d6fbef2c5e410689a98f64b6b

Observation f1e3d810-2745-4a83-8ad8-805a95a82067 · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.546710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.546710Z digest=sha256:46c5f8d5d1816aaa0b358ed70e81131ba1eeafe70152ed82cb1570117eb2688f

Observation c39098de-37ea-4f0a-9dee-cc030938ab7f · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.610336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:9cd268fdbcefdae16571e9a25fdff27f9a4f34e6bf38315331881a97c38db954

Observation aeee2f62-f7d9-4b5a-93e4-43ba92da5b75 · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:50:40.238327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:21b5bb237bd2a879ece6a7a3cd9248ab325c8f723bfd89e7c34644f0c5be0dd2

Observation b2b5bc18-be68-4dac-98bf-d3478a9c8bab · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.182089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:e7b3b61143cf600f8bc96f8037a520647fa8e8d445d912ffbe152ecdfa1c63e2

Observation ed9b6ca1-feca-4e93-8d37-466dc6ccce21 · inbound

LPM 1.0: Video-based Character Performance Model cites this paper.

LPM 1.0: Video-based Character Performance Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:59.901752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:09:17.360697Z digest=sha256:1418f9b597e7e6fda67c84bc5d9493857e9c3ca476b60bc5a2f4b32261a94259

Observation 53a221fe-6b10-4d3c-91ec-02fc5e31bc78 · inbound

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence cites this paper.

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.586177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:13:06.005185Z digest=sha256:86e046bac42f20a6ad0e0001f6404cd2911476d958150ac12f9183eefd7b3a8b

Observation 915f1fcc-8e73-4951-b8ff-7eabc64a9178 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:04.221942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:12bdc1acd855fb717ba97ed5e58280ff00ac4196e9f32f3063239a4c30673331

Observation 3ea3c7b8-d52c-47e6-9c48-0973c96485da · inbound

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation cites this paper.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.635424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:5c7a2e7442e972812bee7f54996631aa4012232352a45ddfb7614b2a5dac8561

Observation 0615e64f-6599-41bd-83d8-5bcb33534e7c · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.776769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:c55a952ea429fb3ea6c87012bd45f9e5c8d4f94243249c91a054cb6cd61ac1d0

Observation a3dc51c6-bb87-40c8-b871-35b2a91dfdc1 · inbound

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields cites this paper.

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:06.456958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T13:50:37.638842Z digest=sha256:4047fa350954da666be276460eff8aefaf2615388679c93c23f5c4cd97c630d4

Observation 37293b50-a45c-4a27-a122-eaa91e372b17 · inbound

VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors cites this paper.

VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:06.244777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T02:00:29.091220Z digest=sha256:0935cf0cd3c36c8714785d97ff7ec55bc80a45adb168de1c6add77069e7dd850

Observation 85ee8bdd-7e3f-4f54-b055-863779596bf1 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.362381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:2bc27a856b0b822c769ece7fa8b6e1961cd6a07ffe92a813a6d2102326ff15c8

Observation fe6b0a7c-2970-472b-a76c-05962dcf8ef6 · inbound

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration cites this paper.

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.909558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T22:15:17.804936Z digest=sha256:8a407748f67be6b73e756e4077cb91c20e1912075a08ca6ad83d61559f784442

Observation 39e56e96-9933-4318-b211-5bfd45fd9511 · inbound

LongCat-Video-Avatar 1.5 Technical Report cites this paper.

LongCat-Video-Avatar 1.5 Technical Report Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.607869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T18:28:16.968339Z digest=sha256:056cbbab6ced8997e5580706b73fdb33c296f3e8c5ed632f84953d6e5773e253

Observation 802dd96e-2a18-45e5-94a0-861643019afd · inbound

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance cites this paper.

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.705373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T06:56:35.931176Z digest=sha256:46eb6a5114784c2fedae91eaaef6048ff5d380bfd541317f8bfd5638f2a33ad6

Observation 9adfd245-c7ea-4461-bb45-c772ba35572c · inbound

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data cites this paper.

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:44:19.552832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T06:35:58.943585Z digest=sha256:51dab5016a0a01fb43d1dae068b17de7a0d92ec316f3a482de251984a54627c7

Observation abda8adc-4997-457f-95ab-4eb69b4bc66c · inbound

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation cites this paper.

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.338694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T06:13:11.504526Z digest=sha256:fbbb4cb3578ab034635cd125b2d9effa536a846db6a7a618039442f40e1a9ce0

Observation e55d4727-a3f8-4397-a63d-f09874dc90a8 · inbound

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models cites this paper.

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.093129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T13:16:16.676244Z digest=sha256:3c89da4c2f850f0db3c5033b2e4de32d49e603d124b457ecb3ce1f4d96194e76

Observation 6500ddfb-d75d-4892-a023-fe235c500c26 · inbound

Vidu S1: A Real-Time Interactive Video Generation Model cites this paper.

Vidu S1: A Real-Time Interactive Video Generation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T04:46:57.581402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:46:57.581402Z digest=sha256:c62f148d6d7c1424f41ff58aad601bb148b2a43f5698a524497a727a08860e77

Observation f9925a27-de13-409c-9838-6f113ba0d756 · inbound

Vidu S1: A Real-Time Interactive Video Generation Model cites this paper.

Vidu S1: A Real-Time Interactive Video Generation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:31.525012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:31.525012Z digest=sha256:40244d23d9443853e7ad30a65a611d6c96ceefa631a637843777069213ab0718

Observation 8d80c2af-be36-458f-a467-b2f597e4d5ab · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T01:59:43.167178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:59:43.167178Z digest=sha256:efa7f35c62a7e16831105dbe2616752f244030bdb9aa2e5b7bac5bc10c2736b1

Observation 96151756-fd25-4fe9-a567-e7ac1b8a8321 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:10:13.757131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:10:13.757131Z digest=sha256:f3c97555432049f6fd64c6a2ed9ac1357794f886464c6ecd61d36c40d2438a4c

Observation 0db1233f-e7bb-4a10-bb85-3d0ff1a05a65 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T07:35:22.880155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:35:22.880155Z digest=sha256:e9219d2bf091e62e64077242963184a60a68bfb3b8941f28cf19bc33b36ec022

Observation 51f7c75e-916d-4188-b3d9-db5c41a5be5a · inbound

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars cites this paper.

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T23:18:59.468371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:18:59.468371Z digest=sha256:78aecc768f242be00abc7c40bbb34efefe3951004e4358a17e888fd4aa578946

Observation 32ba9d0c-e283-4059-9009-9515cc4f2452 · inbound

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation cites this paper.

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T17:08:11.944416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T17:08:11.944416Z digest=sha256:c0cd90f55e1dd10115707930dd573d69bbc95f6c0efa293b8a83ba4a5fb63750

Observation cbc871b1-6b14-47db-b6b1-2c0f3c82aaef · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:13.104187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:13.104187Z digest=sha256:873192defe8b0f0463db4837af49b6e76101cd702d31444dad88f807033af620

Observation afd2025c-a931-4640-a65e-51dd3c8b3371 · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:46.515828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:46.515828Z digest=sha256:437c69f0dcbf78dd72376e9900dbb4a38a67c50710313823149e728d5748c79f