Pith. sign in

Paper Citation Record · LEDGER

Wan-S2V: Audio-Driven Cinematic Video Generation

As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 32 inbound Pith citation observations for arXiv:2508.18621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18621 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:24:59.444766Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:05:47.808077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.703864Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4968f6d7-9c97-4a4f-b775-bba1ef82af99 · outbound

This paper cites Qwen2.5-VL Technical Report.

Wan-S2V: Audio-Driven Cinematic Video Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.584884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.584884Z digest=sha256:dbcd7650fc2c032420696206caaf0bd9f6c0cfc96b08649aee8ce606b8674a31

Observation 7fb76f94-98a9-4e52-b57a-a60ac840fe62 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.824740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.824740Z digest=sha256:e6bce49aa3a6850aef6ecfa1f9bc12b9fd8c654c2d35a3ccb822735371819180

Observation 04bda69f-6eb3-4468-bb05-9d351ea780ed · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Wan-S2V: Audio-Driven Cinematic Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.874741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.874741Z digest=sha256:c36d600cceca31871e68c00d3aa1785378d6f6603d193921239ba3fb5028eeca

Observation 6897fdfb-9ade-43ff-a776-5b1078b0f010 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

Wan-S2V: Audio-Driven Cinematic Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.025790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.025790Z digest=sha256:ecb0d539893bcd4b7d417c901221d1c54edfd44def53fdadfc195f2ac0e0c78f

Observation 3ae816ba-4882-42a5-9c5f-4b5e41715aa9 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Wan-S2V: Audio-Driven Cinematic Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.084742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.084742Z digest=sha256:19c81dffbc3aae5ad3c35811cef8ff5622a51834beb99222ede54ca909ac99b7

Observation 54d8d8a5-a960-4da1-b302-d210086694cd · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

Wan-S2V: Audio-Driven Cinematic Video Generation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.254749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.254749Z digest=sha256:26332e4236c429cfd9cbb3f2d6bb74e1b38ab898d62573c87c1bb02ea4822e48

Observation 163e6c2f-8abb-466f-b17d-b21ee7be9c84 · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

Wan-S2V: Audio-Driven Cinematic Video Generation Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.301375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.301375Z digest=sha256:5dfebbcd3d3b419f573c9e5804b6b583777fd8d029c027a2f753bdde467a93d1

Observation 0ac0e266-00bd-4c76-9f87-ba7e83315f05 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

Wan-S2V: Audio-Driven Cinematic Video Generation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.444766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.444766Z digest=sha256:3dc7e8a14a86e1b723879c3b94954bc0a107d61d226dbb23ffd85dbf0f80acd3

Observation 13af52ce-5fa9-44ae-8acb-0f90b09b0c9c · outbound

This paper cites Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin.

Wan-S2V: Audio-Driven Cinematic Video Generation Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.354746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.354746Z digest=sha256:c8464a6bd74958d787e253f8d309a5796635588568ac4835697ae4d6f2f62b51

Observation ca1bd54f-4053-41a3-b5f2-245a5bca4ded · outbound

This paper cites an unresolved cited work.

Wan-S2V: Audio-Driven Cinematic Video Generation Unresolved cited work

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.774990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.774990Z digest=sha256:a7d97c9c6d1cd88c9322b50d7ba3238a426733dea4df028070b7a84c1d247146

Observation b7516b7e-1033-429c-b33c-0060dd4bfa65 · outbound

This paper cites Christoph Schuhmann.

Wan-S2V: Audio-Driven Cinematic Video Generation Christoph Schuhmann

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.974740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.974740Z digest=sha256:0003848099ff3b97c42c57aaabe68f46ccf821c4ccd89f4c8ab3e517e17118bb

Observation b29b68cb-f0fc-44d7-9b16-1208154a8507 · outbound

This paper cites Image quality metrics: Psnr vs.

Wan-S2V: Audio-Driven Cinematic Video Generation Image quality metrics: Psnr vs

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:25:02.079683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:24:58.724742Z digest=sha256:1e7990c6b4d020e922968d23e8612a287a32f8a315ae0020ae85cf1ed38a71a3

Observation d77ea705-6962-4e0f-b2a1-7babb4f8852b · outbound

This paper cites ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation.

Wan-S2V: Audio-Driven Cinematic Video Generation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.404743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.404743Z digest=sha256:fcff10b1f6197b0c69ef9819af955bdb83ce4f22b3d676fc53714e8a91156e85

Observation 90c58207-0cbf-4150-bb93-e2f38d6b30c1 · outbound

This paper cites Flow Matching for Generative Modeling.

Wan-S2V: Audio-Driven Cinematic Video Generation Flow Matching for Generative Modeling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.924743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.924743Z digest=sha256:1651e85eae7f6e05a01e6a41eb2ca5a04169357b0ec80c9dec1425666bffb9f9

Observation d377db72-a42b-49dc-a21b-bcd2603004dc · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

Wan-S2V: Audio-Driven Cinematic Video Generation USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.674741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.674741Z digest=sha256:71c5d1bdfa7403261ab3427f15918a2684ff13942e2ab8c10a48636e24ec6d5d

Observation 3e091873-51f3-4079-b666-ccbce3d342bf · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.625126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.625126Z digest=sha256:2c1be9e05fa3d77217958b2e5e6e43ee05adcf443229be887cf142727f6eb8b1

Pith citing papers

Observation 823551b6-d303-469f-a713-b7cff82979e0 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.808077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.808077Z digest=sha256:998b2ecd1212653fb11c9c8f5d248e0e872026aceebe6a01cad3528982b5cb36

Observation 31fba0ea-5d3d-4267-8785-c127026cd2ba · inbound

ASTRA: Let Arbitrary Subjects Transform in Video Editing cites this paper.

ASTRA: Let Arbitrary Subjects Transform in Video Editing Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:16:14.277804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T10:13:10.141426Z digest=sha256:e91f3778436a8a501aaaf4f2a2e5f4eb6a29dfc5bf19ca5295a61e902c9cfe75

Observation 6e6f1077-bff7-42c6-91f1-b22c9c3940db · inbound

Understanding, Accelerating, and Improving MeanFlow Training cites this paper.

Understanding, Accelerating, and Improving MeanFlow Training Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:39:16.078524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:39:16.078524Z digest=sha256:c05c05ba4813df48201737d8d141941c139e5d5240643a7e9d70276b04c3b672

Observation d3775d6b-28f1-4e2d-8dd4-07c78e054cc5 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.463818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:6aab0a9b8bf75352a7838514d097644d809b2860acb436b982171986b333fbd8

Observation 81c585a6-fcc6-41ec-8429-171803c4fc2b · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:41.295310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:41.295310Z digest=sha256:81d8244dd3464c50fc619191690a0369fa4a594791e3e551e8879561196eb2cf

Observation f1e3d810-2745-4a83-8ad8-805a95a82067 · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.546710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.546710Z digest=sha256:318fc69f10aeb66872529b5456d5de44862a838bb6f7526c5b2fb9d7dbff577c

Observation c39098de-37ea-4f0a-9dee-cc030938ab7f · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.610336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:831888f487272a59843def34da77e5d44f9e097ad68ea4697dceca3479710516

Observation aeee2f62-f7d9-4b5a-93e4-43ba92da5b75 · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:50:40.238327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:ed5ee1df9f0afd6d44eafbd34953cb233dcb3aa999a2e7583313da1efe451c87

Observation b2b5bc18-be68-4dac-98bf-d3478a9c8bab · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.182089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:1b11d5b7170b902ca57f2f2582ba3bfbd9804bb4b605e578cb0e0e4fb63916ed

Observation ed9b6ca1-feca-4e93-8d37-466dc6ccce21 · inbound

LPM 1.0: Video-based Character Performance Model cites this paper.

LPM 1.0: Video-based Character Performance Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:59.901752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:09:17.360697Z digest=sha256:9c20a7168fcbc5a713b7d4b56c455602d78942ee6f7b4ef3710801e8dccc986f

Observation 53a221fe-6b10-4d3c-91ec-02fc5e31bc78 · inbound

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence cites this paper.

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.586177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:13:06.005185Z digest=sha256:01b3de9adacd7a5e978b0eaacefe6745fcf3ca5f85ab24f996fb281cdcd7e62f

Observation 915f1fcc-8e73-4951-b8ff-7eabc64a9178 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:04.221942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:1a488b9cae9e7e8315f0172bc1b537ccd3fde9ed2134a0009944fcf62d91b524

Observation 3ea3c7b8-d52c-47e6-9c48-0973c96485da · inbound

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation cites this paper.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.635424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:f480735e4025da553d24137dec7266eef524cf348d09bdd0ac2cc8d0e430f046

Observation 0615e64f-6599-41bd-83d8-5bcb33534e7c · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.776769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:fb3249aed43cbef4d67fb0ba30b9d2e2325bc89104b4d3fdad479c0b78dd61c9

Observation a3dc51c6-bb87-40c8-b871-35b2a91dfdc1 · inbound

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields cites this paper.

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:06.456958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T13:50:37.638842Z digest=sha256:99ea8042f0607d75d22cafc57451e3bd28b7c656beff1c0dcabea3ea38adf44e

Observation 37293b50-a45c-4a27-a122-eaa91e372b17 · inbound

VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors cites this paper.

VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:06.244777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T02:00:29.091220Z digest=sha256:93e160a78127fc1379f3edfdba6d72732d489621614498ea16831a424bacb645

Observation 85ee8bdd-7e3f-4f54-b055-863779596bf1 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.362381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:de3c0548cd527335a68865f2eb7e0d6e4f1b8abf5cf5ff3bf4f6efa0fb151c76

Observation fe6b0a7c-2970-472b-a76c-05962dcf8ef6 · inbound

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration cites this paper.

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.909558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:15:17.804936Z digest=sha256:10cc617476d9abb37b4fa29f57a7bdf67239d011bd725a6acbcd9156f7537a04

Observation 39e56e96-9933-4318-b211-5bfd45fd9511 · inbound

LongCat-Video-Avatar 1.5 Technical Report cites this paper.

LongCat-Video-Avatar 1.5 Technical Report Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.607869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:28:16.968339Z digest=sha256:80748d8c1eafaa9a327f4aa2c45bfe05c6059ceabebfc9154407e85edcdd586f

Observation 802dd96e-2a18-45e5-94a0-861643019afd · inbound

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance cites this paper.

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.705373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:56:35.931176Z digest=sha256:1d9de0f1e0048858ac1245a65ed2752ad2b6f3e973c167828d03a5a992ab94af

Observation 9adfd245-c7ea-4461-bb45-c772ba35572c · inbound

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data cites this paper.

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:44:19.552832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:35:58.943585Z digest=sha256:fa1329db789117227e122195e72d134df47b54df81f020c48eb8bff5c0ecf525

Observation abda8adc-4997-457f-95ab-4eb69b4bc66c · inbound

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation cites this paper.

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.338694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T06:13:11.504526Z digest=sha256:6e98c9e1e3e8a0192ec9d4fb5f4ee9a86963da39af35b155c053dc60f2adf155

Observation e55d4727-a3f8-4397-a63d-f09874dc90a8 · inbound

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models cites this paper.

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.093129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T13:16:16.676244Z digest=sha256:0339e34b41112dfc179cd98faacd402a16a7aac4d35d2d98149b155479039b29

Observation 6500ddfb-d75d-4892-a023-fe235c500c26 · inbound

Vidu S1: A Real-Time Interactive Video Generation Model cites this paper.

Vidu S1: A Real-Time Interactive Video Generation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T04:46:57.581402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:46:57.581402Z digest=sha256:c417630e9274797a30fe9dbe7fb29deaf5b18a863717a344952680c4d598a1dd

Observation f9925a27-de13-409c-9838-6f113ba0d756 · inbound

Vidu S1: A Real-Time Interactive Video Generation Model cites this paper.

Vidu S1: A Real-Time Interactive Video Generation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:31.525012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:31.525012Z digest=sha256:e435f77321a9c1232db2fec26cb444c05c33cbf1795fd85e2ad70a4ed32ac1f9

Observation 8d80c2af-be36-458f-a467-b2f597e4d5ab · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T01:59:43.167178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:59:43.167178Z digest=sha256:90e09d0d901ab11828210d8e0eac6b113776eb16ae8331b2de2b953e6149eedb

Observation 96151756-fd25-4fe9-a567-e7ac1b8a8321 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:10:13.757131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:10:13.757131Z digest=sha256:7a0dda8f869fd99203269472977890a63fd2aacc1a6efbe30a98223a352cae7c

Observation 0db1233f-e7bb-4a10-bb85-3d0ff1a05a65 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T07:35:22.880155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:35:22.880155Z digest=sha256:5360e63d16ea9054ddbe61970804464444c11bcc4e6e101b58b50af91c12521f

Observation 51f7c75e-916d-4188-b3d9-db5c41a5be5a · inbound

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars cites this paper.

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T23:18:59.468371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:18:59.468371Z digest=sha256:57fa796575972a2fc84182b90d20e53bc67dffb0d414fbe174c4cc58621e8085

Observation 32ba9d0c-e283-4059-9009-9515cc4f2452 · inbound

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation cites this paper.

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T17:08:11.944416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T17:08:11.944416Z digest=sha256:3460daa611d2bba1dc205f1358ec9576b58f8c319623d23ce7ca6189e0aaf311

Observation cbc871b1-6b14-47db-b6b1-2c0f3c82aaef · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:13.104187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:13.104187Z digest=sha256:9095a374cc66617e9842d57f6a096d80d4004d21b8875b05c90165d3fea607e9

Observation afd2025c-a931-4640-a65e-51dd3c8b3371 · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:46.515828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:46.515828Z digest=sha256:88b1ae8055e9e6b2cf0fa698ea048480ce05c26cb22af19d9a1b31f711edf9eb