Pith. sign in

Paper Citation Record · LEDGER

AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 58 inbound Pith citation observations for arXiv:2403.17694.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.17694 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:11:25.341871Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:16:26.575398Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 01855afa-757d-419b-9cb5-40e207cc27d3 · inbound

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation cites this paper.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.558326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:eaa31120e0b91908b32c9f3522f36edca31940add370c6ab267819e2dc3999b4

Observation a0475f1b-34d4-4481-a77b-500538431ca7 · inbound

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation cites this paper.

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:23:15.361627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T17:19:51.411937Z digest=sha256:b08454f2bd3d321bf8a9dc0edd25837aea6d43372f7d57eb890ccbfa84de3a47

Observation 88ec1c50-ff40-442b-8334-1dada8c4bcc6 · inbound

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model cites this paper.

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-07T21:11:25.341871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:11:25.341871Z digest=sha256:88afa82781ea52944915c7897145abc6396993da535dc774a0246a090b03067c

Observation 1f0bcfd1-d675-4f87-ae55-6c336198b589 · inbound

Exploring Timeline Control for Facial Motion Generation cites this paper.

Exploring Timeline Control for Facial Motion Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:52.852981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:52.852981Z digest=sha256:ab0ef451219d63a556d51dc605e91493224d457fe41b440ebeb531c20033004b

Observation 0d178024-8051-4ec1-a070-bf1c516af756 · inbound

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing cites this paper.

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:46.497816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:46.497816Z digest=sha256:ec11e9223def77b56140121cc226fb77c1b6bdb8bed4e7c643cc4f8286c4a6f0

Observation b146ccef-7509-4d14-9277-f7205658eaec · inbound

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation cites this paper.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.059612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.059612Z digest=sha256:38656b6271c532a66005e5e058e1b473a34a0a268733fd2b5a8a12e4f42bdc05

Observation 67cedca9-14dc-44f2-b242-aff70eb235b5 · inbound

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers cites this paper.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.767656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.767656Z digest=sha256:4eb33d89c050df1f232e4fcb614bf19478e2e54c000ff0d7288ceadbe6f28716

Observation 9ac92218-4b33-48cd-bef9-b7b07723083f · inbound

Speaking images. A novel framework for the automated self-description of artworks cites this paper.

Speaking images. A novel framework for the automated self-description of artworks AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:56.065494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:56.065494Z digest=sha256:8e197e96907e3b3075313ad1d4a2bd0a53db15e0183a998e9fb9298120b3d5aa

Observation 8dce6ce4-191d-456a-94f1-149be2ffcbab · inbound

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models cites this paper.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.450551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.450551Z digest=sha256:963ce44915fdf9a4223b5908c4909f8443c06a1634003bb48ab0bf002a1c3d10

Observation c3a1929b-d64c-4a7e-b2a7-33342cbf61a2 · inbound

Audio-Sync Video Generation with Multi-Stream Temporal Control cites this paper.

Audio-Sync Video Generation with Multi-Stream Temporal Control AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.408110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.408110Z digest=sha256:c1ea5bee8aaa75ce63493d53e7ef4b32392931bfd0d2c7435436d5dff731f091

Observation e7768a50-d62b-4394-bd18-c1521d85f809 · inbound

Controllable and Expressive One-Shot Video Head Swapping cites this paper.

Controllable and Expressive One-Shot Video Head Swapping AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:37.820190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:37.820190Z digest=sha256:18a1627ebbe151c38732652b92ae813950a0e1f47e1f8e4e6a31ddbce1126d09

Observation 042642b1-40a0-4eb1-a81b-f1ab250bac94 · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.067493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:835a05bd68eea909452ea9e82d243a2ac7f2cdba88bb1a7f5bc58333a488fe0b

Observation b3eff734-9988-4b39-90e7-632f17a4d897 · inbound

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases cites this paper.

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:30.336835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:30.336835Z digest=sha256:383f8f54e4a00ce13a723f98eae30d9a087c7e597b70e0e1ecd511937990dcce

Observation 4654cc85-ca5c-4d2b-8181-5c1c68606216 · inbound

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation cites this paper.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.556209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.556209Z digest=sha256:f040334b61e62cd784c14bde2a78c328df7df9d37cc9a379cc5df2053be8cb38

Observation 2b7a5166-c530-4d93-928f-bb09d9a70d22 · inbound

Democratizing High-Fidelity Co-Speech Gesture Video Generation cites this paper.

Democratizing High-Fidelity Co-Speech Gesture Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:00:03.247582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:00:03.247582Z digest=sha256:3362825de82767b10122e246ae206cb0ca66be2de8d87cc147a3e420b3ff1e1d

Observation 8e5fc1b8-01d5-4dad-9028-01b91495ccec · inbound

HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation cites this paper.

HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:38.865416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:44:38.865416Z digest=sha256:f4aadf04ba428062e2d540dfb82a0baf7afea284d8950efcce15aac77f050b58

Observation 451ecb1f-5204-42bb-bf13-a7b8433af352 · inbound

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation cites this paper.

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:09.367614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:09.367614Z digest=sha256:25fa361301fdb608bbd1280efbf1d5d9e3feaa20659c62720bf4232a07163451

Observation 57b2f6d9-df91-483c-b7ec-79d5c4258cc8 · inbound

X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention cites this paper.

X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:03:44.884845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:03:44.884845Z digest=sha256:6d859fb1e59d3056d4e56131c1a072e0537b8961dcfc895e9baf243b768ffd36

Observation 247b5a6b-0495-44f0-bebb-21bf2c18ce6e · inbound

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering cites this paper.

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:49.133791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:05:49.133791Z digest=sha256:f50dedcbdfe2b7f4a8ee10bdda4b8dbb995eab0a16685ec6fe1d7c90bc6c316a

Observation f0271651-de71-42ed-baa9-1453bd3231fa · inbound

PoseGuard: Pose-Guided Generation with Safety Guardrails cites this paper.

PoseGuard: Pose-Guided Generation with Safety Guardrails AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:04.490234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:02:04.490234Z digest=sha256:ca86ca7b89139daa8e6882a32faf2cb51a2fd3f4c89a540d89b4d5c5bbe89fd2

Observation 499289a0-6f66-421d-af3e-61575593d220 · inbound

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation cites this paper.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.684599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.684599Z digest=sha256:9beda31f46becc79c01c37a0f4ccd7df53b552b6fe5e809166ee7dbd1b5191fd

Observation b0f0a0fb-6853-4d42-82b3-bdd80fcd9c44 · inbound

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation cites this paper.

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T22:04:35.670348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:04:35.670348Z digest=sha256:51e1d4641327513ad5be5a38a7316eb5f7827182d0c6ed45f58a58d7db273e7c

Observation e80c8543-d238-47bf-8d0b-1784c00321a0 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:49.448579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:49.448579Z digest=sha256:9417ddad1c361c741e0b83b44f215196a7e4ea0c0641768a784257fc946fcbaa

Observation 4023f325-6310-4ae0-81b6-c88c7ac8add9 · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:36.310450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:36.310450Z digest=sha256:5d50dc8f199ebb9bf039613a36318dece0e05ff400d4d0c9aa51c3628ea8e59a

Observation 579ea692-c4df-4f52-9229-a7b404e6c94e · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.290012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.290012Z digest=sha256:b5af2a2a9b41bab9512399fac42c7c70a5b027c29e740a77c3d32956d35d6a1a

Observation 2728a245-1ee2-4b8b-9c40-ab9ca16ade0a · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.244894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.244894Z digest=sha256:2b3f8a639ab5dcb8b4d3e53d592435dc1bbb606d2dc76d8741f259cb93cedf91

Observation d021f543-f425-431b-8fc4-9e227d662046 · inbound

Human Motion Video Generation: A Survey cites this paper.

Human Motion Video Generation: A Survey AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:57.303214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:57.303214Z digest=sha256:b42272b5997e87cf9f9847e7c4719d26a93eb797199b324d53047fd2d43e2d84

Observation e89d7844-db38-42da-8e1b-1700a9abf86a · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:42:43.813294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:f249cf3b6895e8fbc42dbe3f72743769955255b3be4bcaeb2f269d8c572a54ba

Observation ff743ba2-f1ba-4f2e-a92c-58fd02cb41be · inbound

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits cites this paper.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.321612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.321612Z digest=sha256:3b9ccd3e1cfa11b440497b57ccaf09ac6690cd31c9cf78cc7ac2059b361c0964

Observation be3101ce-5b9a-497a-a0f2-5d72983df5a3 · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.327070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:b028a4b222be3887bd19fd82d85c3e43c6ca492c902cfcb5d69efb651a71c74d

Observation 01552f3e-7d80-4027-934f-60e289d09be6 · inbound

Instant Expressive Gaussian Head Avatars at Over 100 FPS cites this paper.

Instant Expressive Gaussian Head Avatars at Over 100 FPS AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:59.211811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:59.211811Z digest=sha256:4b1959457aa20ae80f94d3fa2eb21cf3ee81434f80a41e36d62a2d935843bfb9

Observation 75c95618-e92d-4942-898b-04cdcaf3db5c · inbound

UIKA: Fast Universal Head Avatar from Pose-Free Images cites this paper.

UIKA: Fast Universal Head Avatar from Pose-Free Images AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:04:51.573998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T12:02:27.828968Z digest=sha256:52dd9ebe174c0f37d6b0971c5ec7746891a4aee0040fb2bb6ac08fd9a47d4d75

Observation ef2d74b1-b261-492d-934c-d8f159487fab · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.464159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:aab9a293391a6a67a043125ad401533aa1913963efb363684b7bff20a76ff22f

Observation 1b11bd3d-be73-40cd-916d-4bf807868b56 · inbound

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors cites this paper.

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 129

Resolution
unresolved
no resolver link, observed 2026-07-13T22:49:03.259461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:49:03.259461Z digest=sha256:e1e16e1fe17d6dbdc0b8bac6d720e42fcdc9d29c8938e6fd1b1b374dbf9b6dbe

Observation 6b77c536-6fe4-4752-b714-7494b8aa2b33 · inbound

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization cites this paper.

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:51.059816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:47:37.605996Z digest=sha256:d06ded35cee9ed95afcb5406660f441a895d8645b8d14516418957b9a0104810

Observation 869f6de3-8dd9-4338-8b44-166d6aa071da · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:25.219400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:30:36.462567Z digest=sha256:93aff058e7af4193fdad755d12df609f31e4c16da361c4bf27394d841f5fb127

Observation 84eddbc5-c482-4812-a9d6-7c416a6dbf5b · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T16:38:16.290404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:38:16.290404Z digest=sha256:5628ec4f840d63fc5492f17b72cc548d1b2281a78c6a6d7b0b2885ed864f61de

Observation c194592f-4807-4c2f-931d-091a5e279976 · inbound

PianoFlow: Music-Aware Streaming Piano Motion Generation with Bimanual Coordination cites this paper.

PianoFlow: Music-Aware Streaming Piano Motion Generation with Bimanual Coordination AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:00.065751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:54:30.725820Z digest=sha256:54bd15e6de20bee98165d7227e101aa44bd1b3f549497573b7e8486e84625ddc

Observation e4cc6ec4-2ecb-41f5-a7de-4b26bff0e8cf · inbound

Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head Avatars cites this paper.

Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head Avatars AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:50:20.281786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T11:47:42.039240Z digest=sha256:ee65f4b7bf0c5196aeb3c5e282246b045befed63575c6ea790509e6675da09bf

Observation edf47450-dea3-49f0-ac3b-54d3b7f004f0 · inbound

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation cites this paper.

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:15:10.370873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:13:27.689539Z digest=sha256:06241104a9516db6f5ede520938b9e1f1fb2eb5cabbf10cc327797ee556518d9

Observation 00aa68bb-6a1b-4202-a546-18bad7f72798 · inbound

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment cites this paper.

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:04.133105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:30:18.645845Z digest=sha256:ad5edb40d24ff93e74fff8d474bf53ae088611b90866f3ff8093c3e32bc53d0e

Observation 9d3dab7b-01be-4261-b5db-b5f24d0e5bc0 · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.066848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:389d01409bae75f4760906cad437b875f33a25c673fad1827754a29227cdac74

Observation 9f3a0d5e-ad8e-414e-bf73-ac6bf6b8856e · inbound

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection cites this paper.

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:06:04.288003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:04:52.065878Z digest=sha256:ecef7f2b13d1f13315fb775b0bf42833add0398490cb2867d316c3cd71533e4a

Observation b429ade0-fbb9-496e-8c84-a4209b84db4a · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:11.234958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T20:19:39.565157Z digest=sha256:9c6f177a45cef40f4711e03eba3feb4c45a0ed9a0fdf5dd231a3d303b6f5019f

Observation 8671e9e2-b04c-4b2b-89e0-90fa9bde142e · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:57.546592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:18:58.996355Z digest=sha256:e1f88e7ef677ec6e14d52a908d1cd25699f332a70a6c99be9d0dc86b2097436c

Observation 3c1edea4-37c8-4276-bb89-6a154abf2069 · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.949010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:cf797edf5c985bb560aed30752eadd677ebdb432537b884f9749a678b5e64206

Observation ebfd67df-04cf-4a8c-b3bc-87646700981d · inbound

Loki: Representation over Architecture for Diffusion-Based Portrait Animation cites this paper.

Loki: Representation over Architecture for Diffusion-Based Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:24:49.856815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T15:23:48.899214Z digest=sha256:62a25f317079e979f440f84f11359da6f8521bb4be99f137d7edd0f2e32509b4

Observation 4d174dec-3c35-407f-95c6-8f7806df1f53 · inbound

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation cites this paper.

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.748177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:36:53.138354Z digest=sha256:a09bc925ac68bf21d93ab913cba819227869c918a71dfc0758346656592ec1d1

Observation 190bd956-2db7-4ee8-a375-ab9eeb6c1bfa · inbound

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning cites this paper.

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:13:27.648129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:03:27.312548Z digest=sha256:6511731d0b5865cc9cfed5e2329936a68bf16786fb063768982812742e4456eb

Observation 35c1c419-e08c-47c2-a075-f516d9e9a118 · inbound

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation cites this paper.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.369443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:9f416cd17eac1fc2f7301bd9bc02ea575a1eeadb8d396c5ed1d2123e8646c287

Observation 8294845f-7655-4acb-a80a-8ef13ad32a3b · inbound

Archon: A Unified Multimodal Model for Holistic Digital Human Generation cites this paper.

Archon: A Unified Multimodal Model for Holistic Digital Human Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.783826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:03:04.294439Z digest=sha256:ddf2ef97c644447c6d88b3b143ea18960b8b4c55c71c2c957e15d8e3bce0c3f1

Observation 33df6ba3-c176-445b-8002-068b531b6d8c · inbound

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation cites this paper.

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:46:15.580499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T16:19:55.048831Z digest=sha256:5b82bb3d76369506044c85e9ad442a41d1d26d17e0e9fc91ca5becb5ff8e36e8

Observation 679e79a2-4c15-4b52-8c6a-2414065c442f · inbound

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs cites this paper.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.010720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:a59a64fd340978c63c3cb100ff0b9b8673ee9f9ce8b2af53ead24b709da26018

Observation eba3d6fb-668c-474d-a03d-982536240543 · inbound

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation cites this paper.

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.577129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:07:20.577007Z digest=sha256:76b2744b7de429f1126ae1ff71e3bebeb87a990e7263e9022a8f37edfebd01a6

Observation f3d16bdd-0726-4ba9-91db-abde66eb85a1 · inbound

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation cites this paper.

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.320788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:13:11.504526Z digest=sha256:089a51289ad04a20584cd03d9be6761131a4bb55186290df0868ce15fb51c6a4

Observation 2656af68-61a0-4795-8024-6653e9887fee · inbound

ViDS: Video Diffusion Shader using 3D Face Tracking cites this paper.

ViDS: Video Diffusion Shader using 3D Face Tracking AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-31T23:01:31.435147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:01:31.435147Z digest=sha256:27d1d2f320b086b9a9e98112d5c0bc421fc45f0304e7a8c1b855ef5b4d986720

Observation 4f7baa9b-77d8-4a3e-8cf2-10a4f0f8b209 · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:12.829607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:12.829607Z digest=sha256:9281e6d7413d01618332f32cead87965f69b59ff52ffe2349f2bd4c2e458f2bd

Observation 5056342d-4155-41e5-8f60-43ca33456263 · inbound

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation cites this paper.

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T01:03:12.618881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:03:12.618881Z digest=sha256:9774811cb76a7fd161d8c1520ed4a7bfa3a84617d5860d151e7d086682944054