Pith. sign in

Paper Citation Record · LEDGER

Wan-S2V: Audio-Driven Cinematic Video Generation

As of 6 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 32 inbound Pith citation observations for arXiv:2508.18621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18621 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:24:59.444766Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:05:47.808077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.703864Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4968f6d7-9c97-4a4f-b775-bba1ef82af99 · outbound

This paper cites Qwen2.5-VL Technical Report.

Wan-S2V: Audio-Driven Cinematic Video Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.584884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.584884Z digest=sha256:665b48c10871712f8fdafbd7b6cec7ae0dfda7e491837f72c6d4e05431aa5706

Observation 7fb76f94-98a9-4e52-b57a-a60ac840fe62 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.824740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.824740Z digest=sha256:98d4f9c5ac94cb3d3532f1b5e9575150ca24795402e91e656cc7567167601452

Observation 04bda69f-6eb3-4468-bb05-9d351ea780ed · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Wan-S2V: Audio-Driven Cinematic Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.874741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.874741Z digest=sha256:496b609cd5b65fe45d45766767d5b9d9a17dc8565e5204d2c1fd3c01d396ba99

Observation 6897fdfb-9ade-43ff-a776-5b1078b0f010 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

Wan-S2V: Audio-Driven Cinematic Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.025790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.025790Z digest=sha256:d400ceb332a758e8257c62d7211fd82aea11eda3a4217eb99311f4cc54662530

Observation 3ae816ba-4882-42a5-9c5f-4b5e41715aa9 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Wan-S2V: Audio-Driven Cinematic Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.084742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.084742Z digest=sha256:656929e816fc64e822b436ee5162487d4647e65c83e2e01a60da33a71a53e169

Observation 54d8d8a5-a960-4da1-b302-d210086694cd · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

Wan-S2V: Audio-Driven Cinematic Video Generation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.254749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.254749Z digest=sha256:853549f6ff89f09388bfd6e0dfd496328b3cdea0b1df00b610c14fc34761c4a2

Observation 163e6c2f-8abb-466f-b17d-b21ee7be9c84 · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

Wan-S2V: Audio-Driven Cinematic Video Generation Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.301375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.301375Z digest=sha256:74d869a5fc00885ee749fcc062f97dd5fb2dc17af268b72338fdba4dc46abd55

Observation 0ac0e266-00bd-4c76-9f87-ba7e83315f05 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

Wan-S2V: Audio-Driven Cinematic Video Generation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.444766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.444766Z digest=sha256:2fb33367c95501ef83b264d6cfa430db8f6af4c5bdbdb48c83e261fb2cb890d5

Observation 13af52ce-5fa9-44ae-8acb-0f90b09b0c9c · outbound

This paper cites Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin.

Wan-S2V: Audio-Driven Cinematic Video Generation Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.354746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.354746Z digest=sha256:b85411d787ac01af5828c6b37e5cdb7cc2a9f35f06debf81e75721c8a4f9ab03

Observation ca1bd54f-4053-41a3-b5f2-245a5bca4ded · outbound

This paper cites an unresolved cited work.

Wan-S2V: Audio-Driven Cinematic Video Generation Unresolved cited work

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.774990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.774990Z digest=sha256:4f99a89fe3a3a864a3d26601d731a5e141e1fbbee487206f654d2b8e05f5f809

Observation b7516b7e-1033-429c-b33c-0060dd4bfa65 · outbound

This paper cites Christoph Schuhmann.

Wan-S2V: Audio-Driven Cinematic Video Generation Christoph Schuhmann

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.974740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.974740Z digest=sha256:e179b5d4b7b7285ce4e583ca631ad95d2db04ea712c717e6edc44bf03dac93b5

Observation b29b68cb-f0fc-44d7-9b16-1208154a8507 · outbound

This paper cites Image quality metrics: Psnr vs.

Wan-S2V: Audio-Driven Cinematic Video Generation Image quality metrics: Psnr vs

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:25:02.079683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-05T16:24:58.724742Z digest=sha256:ddef6143078b1b6c24be70c74cac485a664e15732e98a58d3cba888dfebe0c7b

Observation d77ea705-6962-4e0f-b2a1-7babb4f8852b · outbound

This paper cites ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation.

Wan-S2V: Audio-Driven Cinematic Video Generation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.404743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.404743Z digest=sha256:b81d5dabc5ceee2b58d1223b422cc022636073c7895c06af02b6bba769988e20

Observation 90c58207-0cbf-4150-bb93-e2f38d6b30c1 · outbound

This paper cites Flow Matching for Generative Modeling.

Wan-S2V: Audio-Driven Cinematic Video Generation Flow Matching for Generative Modeling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.924743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.924743Z digest=sha256:495ae33423d853e2512bbce1decffa96e07eaaf4b6dda39b9bd312f691fda9da

Observation d377db72-a42b-49dc-a21b-bcd2603004dc · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

Wan-S2V: Audio-Driven Cinematic Video Generation USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.674741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.674741Z digest=sha256:d7b67d4eb8af9ec6393c7ae1265ecdc7a412d8c1131390b97c61025267716e94

Observation 3e091873-51f3-4079-b666-ccbce3d342bf · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.625126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.625126Z digest=sha256:e3c141dc02eab48332125a703ef890c494c0a320ab5ce7fbeeb0a2b74b662a5e

Pith citing papers

Observation 823551b6-d303-469f-a713-b7cff82979e0 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.808077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.808077Z digest=sha256:124025ef0cfdc4735c1057fe1cae7da774e402d027bb6f6116acffafa1733a70

Observation 31fba0ea-5d3d-4267-8785-c127026cd2ba · inbound

ASTRA: Let Arbitrary Subjects Transform in Video Editing cites this paper.

ASTRA: Let Arbitrary Subjects Transform in Video Editing Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:16:14.277804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:13:10.141426Z digest=sha256:f51b42ba1c9447938a3b6ad3786ccdd10af5d23cc9c4b158acbf1fe77b682974

Observation 6e6f1077-bff7-42c6-91f1-b22c9c3940db · inbound

Understanding, Accelerating, and Improving MeanFlow Training cites this paper.

Understanding, Accelerating, and Improving MeanFlow Training Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:39:16.078524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:39:16.078524Z digest=sha256:8a4056b9fc4cfe9db3280b1dadcd62c179cde8df32390c2b6cee7c5d1f649ae3

Observation d3775d6b-28f1-4e2d-8dd4-07c78e054cc5 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.463818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:4a0e447fed118625885f42dbec3da320b6a7e9a58079fb92b7a21adab0a3407b

Observation 81c585a6-fcc6-41ec-8429-171803c4fc2b · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:41.295310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:41.295310Z digest=sha256:7422210bd2ddcdd20bd6f76cfadf8905f862c0710aa86b7c6a25ad81968053ca

Observation f1e3d810-2745-4a83-8ad8-805a95a82067 · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.546710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.546710Z digest=sha256:da0a9b976496d1c351310aae37180e80917f8ae49824b4f5b56b9e1c99f458ff

Observation c39098de-37ea-4f0a-9dee-cc030938ab7f · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.610336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:8e89d6d1ba98d1ae099f67cffa2998806117b4e535f1836156428fb98f4ebf88

Observation aeee2f62-f7d9-4b5a-93e4-43ba92da5b75 · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:50:40.238327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:fbabd49d8f288d24ac6c365846c2f2ce7b5f019521dc76c8ef90d7bd3e8854d2

Observation b2b5bc18-be68-4dac-98bf-d3478a9c8bab · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.182089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:c96a48fbf43ca09a32d451035f260d2e376fc58e61394409492090c6f449c609

Observation ed9b6ca1-feca-4e93-8d37-466dc6ccce21 · inbound

LPM 1.0: Video-based Character Performance Model cites this paper.

LPM 1.0: Video-based Character Performance Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:59.901752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:09:17.360697Z digest=sha256:0f941e70191b617b6070883c700a7f75f48fad8a3c57da964699689ccd1545f5

Observation 53a221fe-6b10-4d3c-91ec-02fc5e31bc78 · inbound

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence cites this paper.

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.586177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:06.005185Z digest=sha256:271b674b7ec904af4442fc3c596e0c96d45f500c36fdd7eb2e96aeac26d64471

Observation 915f1fcc-8e73-4951-b8ff-7eabc64a9178 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:04.221942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:ce57e30412042d0e846d12b706429b39ec6e724114e1333e8c9f7f46cf31352c

Observation 3ea3c7b8-d52c-47e6-9c48-0973c96485da · inbound

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation cites this paper.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.635424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:1ab21f6c9584d0d1bb1ba9e815ed69261a15ca3a73d97fa801b712f9a94a4792

Observation 0615e64f-6599-41bd-83d8-5bcb33534e7c · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.776769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:7f945622940e447ba8d8a20ae1c356eea25303eb164caa4372b1418b5f3b7eb6

Observation a3dc51c6-bb87-40c8-b871-35b2a91dfdc1 · inbound

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields cites this paper.

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:06.456958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T13:50:37.638842Z digest=sha256:ef9f87c7b2d825ae3e6feaefc96f499ed2c26e6370b211f475606d12b0f02771

Observation 37293b50-a45c-4a27-a122-eaa91e372b17 · inbound

VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors cites this paper.

VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:06.244777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:29.091220Z digest=sha256:8bc4a339dbfecb7547ec0ea6965e8f2c549b30d9c70325aed62f8e81ffba6b77

Observation 85ee8bdd-7e3f-4f54-b055-863779596bf1 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.362381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:a63babcaf1d1e6db968f46f28857b0c8ed1e639a9bf3289b2250aefad74b4257

Observation fe6b0a7c-2970-472b-a76c-05962dcf8ef6 · inbound

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration cites this paper.

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.909558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:15:17.804936Z digest=sha256:e0bef7ab7c49de1d2ac40e5e2281613dd8d219df8cd3701e543ca9595830a0e1

Observation 39e56e96-9933-4318-b211-5bfd45fd9511 · inbound

LongCat-Video-Avatar 1.5 Technical Report cites this paper.

LongCat-Video-Avatar 1.5 Technical Report Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.607869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T18:28:16.968339Z digest=sha256:a018b6a586232710147fa9caf089f024be226d7084ed6f57c20b6c1b1bc239ee

Observation 802dd96e-2a18-45e5-94a0-861643019afd · inbound

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance cites this paper.

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.705373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T06:56:35.931176Z digest=sha256:8fc3cc6240aa0be0806227ebe4f82570147c537a8ff918f8e76ba56b03156d1e

Observation 9adfd245-c7ea-4461-bb45-c772ba35572c · inbound

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data cites this paper.

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:44:19.552832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T06:35:58.943585Z digest=sha256:1f6c326fa2e4c2c7cee7fb76fa66bea9440fc4a887d14e55049a65a56437f65b

Observation abda8adc-4997-457f-95ab-4eb69b4bc66c · inbound

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation cites this paper.

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.338694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:13:11.504526Z digest=sha256:c7fc2ed22142ad5e15aa8d408569e1ea997f5937471b56012a89d453e572df1e

Observation e55d4727-a3f8-4397-a63d-f09874dc90a8 · inbound

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models cites this paper.

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.093129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T13:16:16.676244Z digest=sha256:7c659c7d2c0d8a247a75f8aa403983ffe9eb48d8ca16c6552690f2054952bd89

Observation 6500ddfb-d75d-4892-a023-fe235c500c26 · inbound

Vidu S1: A Real-Time Interactive Video Generation Model cites this paper.

Vidu S1: A Real-Time Interactive Video Generation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T04:46:57.581402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:46:57.581402Z digest=sha256:8b4dafb28fbdccfd83064675834de0f55c8f2c37ecaee3d1102067c9de0bd287

Observation f9925a27-de13-409c-9838-6f113ba0d756 · inbound

Vidu S1: A Real-Time Interactive Video Generation Model cites this paper.

Vidu S1: A Real-Time Interactive Video Generation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:31.525012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:31.525012Z digest=sha256:8198d92b325c6369fac5bfc4989af344e0be731cc525f40715a607dce8492025

Observation 8d80c2af-be36-458f-a467-b2f597e4d5ab · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T01:59:43.167178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:59:43.167178Z digest=sha256:fee834cdb4fa2f4a42382ea9951cfbd15f534d7ead310ef3ee71c8000af58ff5

Observation 96151756-fd25-4fe9-a567-e7ac1b8a8321 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:10:13.757131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:10:13.757131Z digest=sha256:8060524b9f225319a7203dff3876595ff3d0531fbfd73db91a3e3607ecde671a

Observation 0db1233f-e7bb-4a10-bb85-3d0ff1a05a65 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T07:35:22.880155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:35:22.880155Z digest=sha256:63a22c3d41c86aa7d414c8ee2f9c9150f7c494fa29c20dcf332450df6fe57a7e

Observation 51f7c75e-916d-4188-b3d9-db5c41a5be5a · inbound

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars cites this paper.

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T23:18:59.468371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:18:59.468371Z digest=sha256:01a7bb1ca5b9483039c695fdbccf68058b4ef46b9c42d56937dedd8a57f426fb

Observation 32ba9d0c-e283-4059-9009-9515cc4f2452 · inbound

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation cites this paper.

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T17:08:11.944416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T17:08:11.944416Z digest=sha256:119d3e24c8aabbd7adec55adc3a3ace3c8b1d683146618c48b4648366c76a82a

Observation cbc871b1-6b14-47db-b6b1-2c0f3c82aaef · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:13.104187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:13.104187Z digest=sha256:2ae75f907a7f3611fda83fb0c2b63a93f9feca1ab47edc826339b2d555a3a403

Observation afd2025c-a931-4640-a65e-51dd3c8b3371 · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:46.515828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:46.515828Z digest=sha256:871a9888f8b6de85e435538ee8e0e331304bd43670232f3b429757a93c5fc15f