Pith. sign in

Paper Citation Record · LEDGER

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

As of 7 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 2 inbound Pith citation observations for arXiv:2506.00830.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00830 v1

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:09.561127Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:16:52.478464Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T18:50:16.683799Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2df14923-be37-4abe-ac00-5d3e78b821e4 · outbound

This paper cites https://www.prnewswire.com/news-releases/deepbrain-ai-delivers-ai-avatar-to empower-people-with-disabilities-302026965.html.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers https://www.prnewswire.com/news-releases/deepbrain-ai-delivers-ai-avatar-to empower-people-with-disabilities-302026965.html

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.255827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.255827Z digest=sha256:2c1d89a0cd8d4e510738dfc7e48a2ed511ca85a97d80b67b8033f7db7863ca0e

Observation 135e397a-f7e2-49ce-9051-87c42fd9b728 · outbound

This paper cites Efficient 3d implicit head avatar with mesh-anchored hash table blendshapes.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Efficient 3d implicit head avatar with mesh-anchored hash table blendshapes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.296619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.296619Z digest=sha256:3b96c42cc329f015faa36b3737166f907cf7d96cc8ec32c0c596207b65a919cb

Observation ea566ae7-4703-4d94-9f1b-107b32f69767 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.358842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.358842Z digest=sha256:b4bbc8d767c49b14f795f13dbcfa8872bbbe0f2fa65f9a32cfded7ee0e96e7de

Observation f662ac27-511f-4fe5-a720-da09c6b99c49 · outbound

This paper cites Lipnerf: What is the right feature space to lip-sync a nerf.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Lipnerf: What is the right feature space to lip-sync a nerf

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.454450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.454450Z digest=sha256:eeaff9cba0183d6f2be4b478656965a0b8fd42363b650bc25f2db785a05945d3

Observation 1dc9ce03-21bb-42dd-b406-8b9551be6ea4 · outbound

This paper cites GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.496571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.496571Z digest=sha256:1b7e18e618ab8dd89b9cad28b57a256a256a1901e9ab41d0d7511e5f0f470bb5

Observation 0ce23d57-43e9-458b-aca8-e89518896901 · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers SkyReels-V2: Infinite-length Film Generative Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.582389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.582389Z digest=sha256:6e4768305df901263f29b1142885853ce06cf459a34d042422a31fe76b63bb05

Observation 82d2b146-147b-456e-ab9c-c15ede9407d4 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.662450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.662450Z digest=sha256:8e18a19d553c99e3416b713f1908bfe1e4d81ef0ebabe2bba19bc8dc5ef0da81

Observation 088bdfa4-1de9-482d-9b3e-bce0028c595a · outbound

This paper cites EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.758873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.758873Z digest=sha256:bde5690c433bb013437495c0e3ca57f810383c0ebe0ff66df62fc330adb507db

Observation e395d580-5434-479f-ade6-c58c3185fddf · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Yolo-world: Real-time open-vocabulary object detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:15.136306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:04.843546Z digest=sha256:97ae3a29bad6cad365b2123a0638de3c1b61df005cdf27e58219c1a9724b539c

Observation 7807034e-597f-4eee-ade7-54a915c1d5b0 · outbound

This paper cites GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.901517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.901517Z digest=sha256:3c348c4e5ee328faccd715d733423d63116d2894b29e7f6d6a81d33a0031c1e1

Observation 38667ac5-dbeb-451d-8e18-efaa6597b4ef · outbound

This paper cites Out of time: automated lip sync in the wild.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Out of time: automated lip sync in the wild

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.988523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.988523Z digest=sha256:37801ea708895c8794875085340ab4e3ad0afa17afc19f934cd991f92846461f

Observation c430a401-a2e7-4075-9517-77f260d9c4df · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.049606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.049606Z digest=sha256:5ff69bbcb6e1c7f7bf0c7bc3c60dffeecf6c5585d154d078c957b7870ad9c730

Observation 64299477-3f87-489c-8ad7-41a7079f339a · outbound

This paper cites Adam: A method for stochastic optimization.(No Title), 2014.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Adam: A method for stochastic optimization.(No Title), 2014

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.131395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.131395Z digest=sha256:5a791154b51a2aa97542308e55a80bda546bb53031211f00008f66b8f115afe6

Observation a9673042-2025-4f01-b54d-8f8ab5cee2ef · outbound

This paper cites Scaling rectified flow transform- ers for high-resolution image synthesis.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Scaling rectified flow transform- ers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.180625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.180625Z digest=sha256:15089c791fb48cbf9983c8df54ee5213c7372cc75d9077024f8680ee4f443e69

Observation 5292dd83-0062-4f74-ad13-1b1c5d45a7cb · outbound

This paper cites Motioncharacter: Identity- preserving and motion controllable human video generation.arXiv preprint arXiv:2411.18281, 2024.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Motioncharacter: Identity- preserving and motion controllable human video generation.arXiv preprint arXiv:2411.18281, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.262820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.262820Z digest=sha256:69749ded73d2482c765455e3aedd3bcc0835d70fbc33cf0cef3c127396971fab

Observation a7a0ba40-9daf-45c9-82b3-89ca956fac7d · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.332091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.332091Z digest=sha256:3691cf3010e3c20705608f076e2435c7a66543a2dda34c420da87c9f5874fcfa

Observation 4a16b846-7325-46de-8529-136300117e68 · outbound

This paper cites Scalable Diffusion Models with State Space Backbone.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Scalable Diffusion Models with State Space Backbone

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.394751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.394751Z digest=sha256:ae45e961790bc01caf6029f55baa2efb34d486d09a47ddfb7f2eee4512bfc815

Observation e3b93b7c-1e8c-476d-8216-53b62b7e32aa · outbound

This paper cites Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.470046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.470046Z digest=sha256:738d35dce54cf181122e5bdf8c858c3ea2e676273ee77ac7255f0f3da6003469

Observation ab65c0ea-f8a6-4c34-b126-694ca2bb456c · outbound

This paper cites SkyReels-A2: Compose Anything in Video Diffusion Transformers.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.579680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.579680Z digest=sha256:41723f7efd4a472dbc0e1177af6b91e276e18145147935abbae1144cab3a55ee

Observation 297cb0a0-1f64-46ef-a439-4906f905835b · outbound

This paper cites Ingredients: Blending Custom Photos with Video Diffusion Transformers.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Ingredients: Blending Custom Photos with Video Diffusion Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.662040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.662040Z digest=sha256:1d4bea2fd5c12d74547685920111345b5395537d8978f6545ba1823e5f38f963

Observation b804ddc7-a3e7-4ade-96e2-4494385e97c7 · outbound

This paper cites Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.733194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.733194Z digest=sha256:45ff47c2994da5c9d6130d8250fe99720f671561b8dbb3ab206106e7eb1b1aa0

Observation e19b237a-6533-420b-9d9b-0649773bb5ee · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based generator.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Stylesync: High-fidelity generalized and personalized lip sync in style-based generator

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:14.966252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:05.836728Z digest=sha256:8caf2ce58823290fb119c41c02a8a4434f9009e1a6f66c962b80bab253c5b2bb

Observation d2f9ac78-4775-44dd-8155-7965204d2f2a · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.900408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.900408Z digest=sha256:002ecd1b305997739c42acf6ac7ecc286620db1771fb9ad9ad7befdb0bc95ed8

Observation 905819a1-f1e7-4e62-b054-d4b53bc336ae · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.008788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.008788Z digest=sha256:e7a0b978576317dea0a25ba610e6ab56ff93c6cb386b742995b4503be53ccd69

Observation a4796f77-70ec-4dae-b235-7a8b14eb59a0 · outbound

This paper cites Sonic: Shifting Focus to Global Audio Perception in Portrait Animation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.053086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.053086Z digest=sha256:5f49eaa40a8dab9591d2cc7b848a9f3758a13fb8ba25fbc7093b9291ff966d09

Observation 66b01852-f46e-48d8-a37a-77f2ae21e8a8 · outbound

This paper cites Audio-driven emotional video portraits.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Audio-driven emotional video portraits

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.134499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.134499Z digest=sha256:e008f6fc14243f210a3e6d2e855d0daa55eb2b87103c60453f188a6dce95e974

Observation e5ab0ec4-5864-45ea-a4f2-1c29fda9eda8 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.223848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.223848Z digest=sha256:4742641a184405c8d9003eb60fe58b8de395910847deb929c11d1b5cf9a4dad1

Observation bc4f37c4-a317-4b99-8a35-2d9277195d7b · outbound

This paper cites Assessing empathy and managing emotions through interactions with an affective avatar.Health informatics journal, 24(2):182–193, 2018.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Assessing empathy and managing emotions through interactions with an affective avatar.Health informatics journal, 24(2):182–193, 2018

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:14.818039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:06.343672Z digest=sha256:7954fb969298b8e6abac6681e1834668a7cb26a1431650d3d6d19e5e7072a676

Observation 0413af8e-d553-412d-bdd7-1a50bdef7b88 · outbound

This paper cites Educational virtual reality game design for film and animation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Educational virtual reality game design for film and animation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:14.669546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:06.410579Z digest=sha256:f7de234f4448f459a27ade0db42d7e05b33781243e694016e787808938fc7926

Observation f5c80228-bd0d-453a-ae91-4cb97050f5f4 · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.ACM Trans.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers 3d gaussian splatting for real-time radiance field rendering.ACM Trans

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.487818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.487818Z digest=sha256:ac36e85105a03834103b22146c03777ea49d012394e9227d84892ce13c4e626d

Observation d26bead3-9f11-4848-96f9-3026309ff20f · outbound

This paper cites Auto-Encoding Variational Bayes.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Auto-Encoding Variational Bayes

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.530018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.530018Z digest=sha256:8acbf988752c0083d7ccc35637dc4b5c7fdc73fdb1fe84f20c7a36423186d7aa

Observation ddb1a76e-3e1e-411e-afe9-2ac66e000fd0 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.596226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.596226Z digest=sha256:b44b9c9bda3aa659dce0c9f9e82e1c545b158a86c72a7b4d90345fcb081c823e

Observation b0ba2e37-e88a-4c9f-adc8-9341d5a5b77a · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.693590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.693590Z digest=sha256:ca2a9e9f163808006da41b609f2ce4897fa0622d008659811a82b8fda98335e1

Observation bfcec5b4-063a-40e0-8512-198b8c24f05c · outbound

This paper cites Ae- nerf: Audio enhanced neural radiance field for few shot talking head synthesis.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Ae- nerf: Audio enhanced neural radiance field for few shot talking head synthesis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:14.546838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:06.754688Z digest=sha256:8bd39cb97a2b19e8bd5431b6ab18c86b42dac519d2a7a9117f419e21a027b644

Observation c73b9690-e4ec-40d0-9335-48afd3d47c78 · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.821951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.821951Z digest=sha256:f3582a1fac33bd5e8c90e42aefe493e6346c3215a6f80b3d872b5b18949dc2cc

Observation e7811ccb-b8bc-4c31-9af8-acce66108131 · outbound

This paper cites TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.890022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.890022Z digest=sha256:0d8dd2599883afd529ec6de613cf1fd7f8ff1a42cc1f532889a47ca3d5fba4be

Observation 1b47b76c-b73d-4402-87bb-2eea3da8d6cf · outbound

This paper cites Learning a model of facial shape and expression from 4d scans.ACM Trans.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Learning a model of facial shape and expression from 4d scans.ACM Trans

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:06.956820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:06.956820Z digest=sha256:471b701916318fbd5add0327c9465f22f446aadf775a936d40f00fe9f06f2277

Observation 14975a58-5155-44b6-901d-3338eb87cc2e · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models, 2025.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:14.360029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:07.042283Z digest=sha256:7096cc58c2f4955d201a8940477eaa7c0895a6789171e5cd8dce11b0d558df4f

Observation fe03e516-e660-49f5-b0ab-28728adcb9ac · outbound

This paper cites Flow Matching for Generative Modeling.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Flow Matching for Generative Modeling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:07.094496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:07.094496Z digest=sha256:739e9606a813f13c709acf35dd2781048057a8a9e3990674ff8b1e761de55a26

Observation 800eea90-808a-45e7-b61e-bcd224ae6bfd · outbound

This paper cites Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:07.173781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:07.173781Z digest=sha256:b6e8fb36d7328365d214ec96947fe5cdc3b879253e14a084ad77b51c1186411f

Observation 3a80e1e8-b2fc-4b43-aa6f-13d79eb19bff · outbound

This paper cites Moda: Mapping-once audio- driven portrait animation with dual attentions.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Moda: Mapping-once audio- driven portrait animation with dual attentions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:14.190565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:07.249929Z digest=sha256:741e22a5d833060f5cfafc88460dd78bb7d3fefdc8498de687db6847e0f7bb78

Observation 74a95d57-eb40-48c1-8976-88b356343155 · outbound

This paper cites Live speech portraits: real-time photorealistic talking-head animation.ACM Transactions on Graphics (ToG), 40(6):1–17, 2021.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Live speech portraits: real-time photorealistic talking-head animation.ACM Transactions on Graphics (ToG), 40(6):1–17, 2021

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:13.952818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:07.331183Z digest=sha256:a35ec168cbcfd1d44a7735b20a94519a9195b587e51eb2af597b5f644659d2d4

Observation df620f67-8d1b-4205-98dc-3ddcfe928ae3 · outbound

This paper cites Pixel codec avatars.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Pixel codec avatars

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:13.722317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:07.418254Z digest=sha256:08d5a0f1c0f980f77a03fc81c76287d2b78a6f62ee1afe8c26286834e28c7e6d

Observation 7cbe88cb-eafd-47a0-82a9-be069f4087ca · outbound

This paper cites Styletalk: One-shot talking head generation with controllable speaking styles.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Styletalk: One-shot talking head generation with controllable speaking styles

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:13.445427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:07.475665Z digest=sha256:35f57d648eae07bd3c2ed37c7fe1b14773a51d210e0fbb7693c6edc89cf5a71c

Observation 5b524bad-e05e-4c5b-82cf-4f52c311a11b · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:07.552280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:07.552280Z digest=sha256:5bd74ff882772aabdb384ce915985e358eb4338a72258d1dc6cd6f479abd574d

Observation b491905f-e470-4cb2-a8d4-ab33c83b2819 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:07.669803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:07.669803Z digest=sha256:0ccc4b049fb6f176b13f430ea89c8445d21fa92dbfc1ee049fdddb5653720b75

Observation 89d54c32-20a1-4d3f-8165-98039b15531d · outbound

This paper cites Scalable diffusion models with transformers.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Scalable diffusion models with transformers

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:13.233371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:07.720769Z digest=sha256:bd61cd736672a9c3784d22a8601aceef0e2887340affe4835d75e14b8c820c53

Observation 038adc5b-50c4-4171-a42a-02f6a61dfaf7 · outbound

This paper cites Synctalk: The devil is in the synchronization for talking head synthesis.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Synctalk: The devil is in the synchronization for talking head synthesis

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:13.042259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:07.764696Z digest=sha256:79cf1635988eea8c01a88b35170c98de92572a7341f89414bb9e78456d391de1

Observation 47f55df0-4e7f-40be-b0c3-5216353557df · outbound

This paper cites Emotalk: Speech-driven emotional disentanglement for 3d face animation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Emotalk: Speech-driven emotional disentanglement for 3d face animation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:12.895509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:07.812253Z digest=sha256:1366cadc6bee3ac4799f454de56ffab80af4448d74e1d9c387a52da00dbed83a

Observation 16317c3c-7c29-483a-ab92-48d8f5d94aa1 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Movie Gen: A Cast of Media Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:07.866515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:07.866515Z digest=sha256:7cb9693f8b8634ff0c771a75dfce6c597e21d58f3b1acce09684b7d53cc61455

Observation 480e6963-8300-4990-b826-fe3a48def1ac · outbound

This paper cites MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:07.916907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:07.916907Z digest=sha256:d96724315f4bb1021bcfba64440beb0a9e317a811d5569d3f9c1aa961dd1187b

Observation 8697d019-ad12-4a05-8c31-75557fe4b10d · outbound

This paper cites SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:07.948041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:07.948041Z digest=sha256:213de0f36e20103e0d911651e88128b3109b51c8ab3df50a58c79719898dbf00

Observation 564fe11f-0334-4575-9643-06fef90f614d · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Robust speech recognition via large-scale weak supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.022078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.022078Z digest=sha256:69331db94576168e58a5cab6b1e1a87a70816d1195ec3273119ad06a3d4edd8e

Observation 67cddcfa-9b75-43f7-81b9-640aa62b749e · outbound

This paper cites What role can avatars play in e-mental health interventions? exploring new models of client–therapist interaction.Frontiers in Psychiatry, 7:186, 2016.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers What role can avatars play in e-mental health interventions? exploring new models of client–therapist interaction.Frontiers in Psychiatry, 7:186, 2016

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:12.658036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:08.066203Z digest=sha256:8afc91836a95e4f3fb24f9841c1b0d2c3b354feb1a11611a7ebe03dd0a3c5b71

Observation 83bc4b14-d6cc-472d-932e-9acd5622bbec · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers High- resolution image synthesis with latent diffusion models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.118208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.118208Z digest=sha256:71560f1d14558b6c433571cfd82169a8323c8bea3716984cceeeaa237870093c

Observation 4e4907f8-4b5e-4997-9001-5d198d5a6821 · outbound

This paper cites Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.IEEE Transactions on Circuits and Systems for Video Technology, 33(3):1247–1261, 2022.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.IEEE Transactions on Circuits and Systems for Video Technology, 33(3):1247–1261, 2022

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:12.413991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:08.178368Z digest=sha256:f0a54d86df3e08d9fa3f770dd7ade08cb7966649a48c4789ba4ca9d8b225dd4f

Observation 35f01443-daac-45b5-8b08-a7dc58d68c7f · outbound

This paper cites Audio-driven High-resolution Seamless Talking Head Video Editing via StyleGAN.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Audio-driven High-resolution Seamless Talking Head Video Editing via StyleGAN

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:01:10.171923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:08.213996Z digest=sha256:d1430eacf45ac67b264c8a7719c9ed9d8fc8ac91c92a38ada96ff7cbbaa52e8a

Observation 64137b0e-2f89-4712-8445-429e7625541b · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.261036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.261036Z digest=sha256:ba9e405329a21ded3d25c7cc646d59375b3e961a6c177d1668b62226b5418c47

Observation da075765-8bee-4534-a681-bf727410007c · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.319003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.319003Z digest=sha256:1cf7354289fdd7ed83fef0659da68ffc82aebb3e8b9524b59328e8550ba0707b

Observation 506bcbd1-e816-4aa5-b6b7-125fb1d175d8 · outbound

This paper cites Nonlinear 3d face morphable model.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Nonlinear 3d face morphable model

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:12.052762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:08.365172Z digest=sha256:d68c20892569a6c9ced08ebd21440aba01dcac835f6d8734b117f74b84c2fde4

Observation 2d1ebd9a-f3b2-4fd0-abeb-7ed6dcd73987 · outbound

This paper cites Fvd: A new metric for video generation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Fvd: A new metric for video generation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:11.720900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:08.427473Z digest=sha256:466e870e45ce47426cb98cc348c5ad1754fd2a092806a3cc34b9b4cf4463583e

Observation 789024b5-367d-4b6b-ac4d-5409abc2e1e7 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Wan: Open and Advanced Large-Scale Video Generative Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.462691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.462691Z digest=sha256:ad3c3ee15c2a5cdeca7f8eac18359bdefa4757194fe527ad63be012aa6395d98

Observation d072b7de-f82f-4ae4-b929-62e7ced42bc0 · outbound

This paper cites V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.482037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.482037Z digest=sha256:e5db955795856c81195432101396b380ba8b6d1bf77b065ecd19b56442cffa08

Observation 046874be-8b90-4980-bbf8-202b867cb7a3 · outbound

This paper cites Seeing what you said: Talking face generation guided by a lip reading expert.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Seeing what you said: Talking face generation guided by a lip reading expert

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:11.391506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:08.513318Z digest=sha256:5bca2842a0308b12620acb808f791464202140e3d6f9a4c0ca7eb2d753d3a220

Observation 005dad01-d72c-4c8f-a3ec-0c0f2d47335a · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.579080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.579080Z digest=sha256:965afb7a7240b3601f6c8b9124047504df9aa31aafdb4dfb6eabab9af3ebb1e3

Observation 45a057ed-6cca-4212-b8c6-a5f4821867dc · outbound

This paper cites One-shot free-view neural talking-head synthesis for video conferencing.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers One-shot free-view neural talking-head synthesis for video conferencing

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.629999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.629999Z digest=sha256:a1ce1a91112e3320c2c4263ffc8018d51b971f19edb1fdf306eb5d272698c3d1

Observation 0465191f-b30f-46b5-aa31-73b4e528087b · outbound

This paper cites MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.685760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.685760Z digest=sha256:32d76f04a7f24f41bc47296045948a234c76c9f9b8cfbd36c341c4abefb23b43

Observation fb09a09f-5380-400b-b2e2-df521df8918a · outbound

This paper cites Panda: A gigapixel-level human-centric video dataset.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Panda: A gigapixel-level human-centric video dataset

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:11.058377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:08.738549Z digest=sha256:0ea4918779459a14298b60b1c09632306d0da02a46828504a58824561fcfa1e3

Observation 67cedca9-14dc-44f2-b242-aff70eb235b5 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.767656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.767656Z digest=sha256:e1c96ac5ae99277143ab6996791b89035bb5645c4e244c90f5c5580956c5430c

Observation b0558abf-a37e-4cfd-bcad-e2474b2d8609 · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.824060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.824060Z digest=sha256:021c3aa5f2a6fed78eee6fb593e862ad312f29cdd6c59fa28d2bc59b450abcdd

Observation 4b103765-461f-4928-83fb-877d4d0092b0 · outbound

This paper cites Codetalker: Speech-driven 3d facial animation with discrete motion prior.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Codetalker: Speech-driven 3d facial animation with discrete motion prior

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.881537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.881537Z digest=sha256:b51e67dc25eafb63cce76433b63b68fef040a8b3235590ac987ddfe1605468d5

Observation baffbf9b-c211-4082-9dd4-470662f66343 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991, 2024.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991, 2024

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.918756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.918756Z digest=sha256:41cd5c6c98f994bf67c6af55d0efff59915b619027c179e08e9be8ffa1bea3cc

Observation 8d643f5b-4a3d-4e4a-aa01-157731aff9db · outbound

This paper cites MegActor-$\Sigma$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers MegActor-$\Sigma$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.976755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.976755Z digest=sha256:ec943bd3b2055e7982e69486228d56d7c83a202de8332d41bc67ff763df4c601

Observation 23f9d9b7-859d-4575-9660-e8b056d0648d · outbound

This paper cites Effective whole-body pose estimation with two-stages distillation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Effective whole-body pose estimation with two-stages distillation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:09.030034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:09.030034Z digest=sha256:91cea208f6f0a6726032d9dda744a30a0ab9a4725328688196f128862bb53ef4

Observation a472f972-a72d-429c-813d-7718562ddb21 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:09.064934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:09.064934Z digest=sha256:400236033002a650736e878d3a254b4de4cafb07455028eb884df79fcb67778c

Observation f7ed62cd-0952-48cc-a2e7-53fa20bc31c0 · outbound

This paper cites Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:09.114112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:09.114112Z digest=sha256:5716253cabdc018d7dcd3402c63d5be4aafbdf200a485d0d6f926c64eae15f19

Observation 6948a4df-cda3-4398-8093-8a432a7d9b6c · outbound

This paper cites MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:09.189058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:09.189058Z digest=sha256:e12d75479e7e33d28c4a092190219bd10b0d483b93a2544258def730d763407b

Observation b2abdc3d-d2eb-4e01-9ac0-bdbff7effd07 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:09.233037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:09.233037Z digest=sha256:88314380efb5669051383eb63776851b599bc0fe32495a7ffc9a8da77fd684eb

Observation 4ebb5b4c-f4f4-439a-8f30-e00783eeb534 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:10.709000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:09.303110Z digest=sha256:9c456443a95b20efcd9f9016064b1ac8f25d613f32a8ce8b3ddd036cf7960063

Observation 9c612247-a0e5-44b8-83f3-b9743f5235ed · outbound

This paper cites MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:09.337775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:09.337775Z digest=sha256:1771875715d83d4ce0bc437ba9b7bf87fc706581cacda058d2f4efd9be8c9468

Observation 46f733c2-0eb1-4197-9ca4-6e096bada95d · outbound

This paper cites Guided Score identity Distillation for Data-Free One-Step Text-to-Image Generation.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Guided Score identity Distillation for Data-Free One-Step Text-to-Image Generation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:09.387556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:09.387556Z digest=sha256:9ff8e6f0e04dec09e11cb6553183bd1f6aa39fab5293d4b3f271593c6e2570b2

Observation 0f82b039-061e-4c9a-ab90-560cdd0c3a34 · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:09.480529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:09.480529Z digest=sha256:6f08cf9c48b0a9b4d46e3737518016bd739f8134ef58ce0078780425a77b75ba

Observation 1b6aa078-8e17-4763-873c-33eaaa507789 · outbound

This paper cites Learn2Talk: 3D Talking Face Learns from 2D Talking Face.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Learn2Talk: 3D Talking Face Learns from 2D Talking Face

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:01:09.744867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:09.561127Z digest=sha256:057434b88bc25fca2f2a660c71921a144c8e64cf216c51340b2bcc27d41e64b3

Pith citing papers

Observation ab09bcec-d3a2-40f7-9f11-4c6229fdc586 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T18:50:16.689954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T18:50:16.222228Z digest=sha256:2287eca3c80fb6473f41f5c43d886599d9bc6e28deecbec822581817e8ab563c

Observation 085f7559-aa81-4101-b8b4-4d893902df8a · inbound

Vorch-Omni: Multi-Task Orchestration of Sight and Sound cites this paper.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.478464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.478464Z digest=sha256:c9287c5482c8c19edc2b71770e31af2a43f19be823117a2b1b70aaff387e63e1