Pith. sign in

Paper Citation Record · LEDGER

EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2407.08136.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.08136 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:11:25.264649Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T17:23:15.330561Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 459c315e-f939-4f26-89a3-5e93e0005029 · inbound

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation cites this paper.

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:23:15.333451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T17:19:51.411937Z digest=sha256:e7734eced8f988bb6d35326c2cc54ae401e03d233755082488a8ecf836581f83

Observation 1c22e319-d0ba-4505-95bf-b858d41972fa · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.511917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:d21b3efc5f7320eab6c5e69c40cf6052106f9a5a406061991638e03d192d66ff

Observation bd35958b-9046-4297-a994-217df4ce15b3 · inbound

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model cites this paper.

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T21:11:25.264649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:11:25.264649Z digest=sha256:c1d336a5651879f661ac0bb199a666488867bc2166774ab99258845e8491b565

Observation 17f683c4-da09-48fb-9451-d598bf33cdd2 · inbound

Exploring Timeline Control for Facial Motion Generation cites this paper.

Exploring Timeline Control for Facial Motion Generation EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:48.623861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:48.623861Z digest=sha256:c976b50a64ef3f3791d3c4847d928e6cfd358cafc133a5b845244d72c9f86cbd

Observation 088bdfa4-1de9-482d-9b3e-bce0028c595a · inbound

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers cites this paper.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.758873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.758873Z digest=sha256:bde5690c433bb013437495c0e3ca57f810383c0ebe0ff66df62fc330adb507db

Observation fefc2380-e9f2-4f0e-84ae-c59b2f8e36c4 · inbound

SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting cites this paper.

SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.536887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:44.536887Z digest=sha256:8bb9502678aa5a6ae1c49682029bbe45299fc9ae37d50a3e4f5acb0f25234898

Observation 72ca8740-86ea-4c9c-8211-1b18966d0519 · inbound

GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation cites this paper.

GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:18.025643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:28:18.025643Z digest=sha256:663c9b5ab30fdf4d9044db7f0bfd8a62a8c2ecae781568a870e6f4747eb429d7

Observation 793f0781-8f72-4cc1-8ea0-f91ba008a5f9 · inbound

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations cites this paper.

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:12.435775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:12.435775Z digest=sha256:a7294c576e2bdf3c9665cc52e6652dfc01144d6d2db98eb7b6f3c013a1a3d319

Observation d2afcd9e-6efc-48d3-97ef-7e7d6fc4ee67 · inbound

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases cites this paper.

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:22.315210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:22.315210Z digest=sha256:dee3c42b3216f2bec2de61896d44a1b010ed55a92f45918e739052679b2d5e95

Observation f10498d9-b198-4772-8a64-fa468374c71e · inbound

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation cites this paper.

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:36.485347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:36.485347Z digest=sha256:5dd6a34cb9e35e0cb16d6a2fb0dbd2f52cb2395536ecb6405a4e179ebba9ab75

Observation 9987ca77-76db-4641-ba9a-b2f7287e4d49 · inbound

Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model cites this paper.

Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:16:58.094767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:16:58.094767Z digest=sha256:8266bcac7efa698059b8e4c86664ba1cfd453fb2341a22957fef1238350b25c1

Observation 208cfa65-5d95-4f6b-96ab-6aa09243dc8c · inbound

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation cites this paper.

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:10.781894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:10.781894Z digest=sha256:bd4524db73734a8c98ddf7c3a0c0b8354b515764f418f19989a8da331d65f426

Observation 8db2d14d-2303-49f8-b42c-0436688b80a9 · inbound

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads cites this paper.

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:18.930569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:18.930569Z digest=sha256:8d2a9867e0d71defc7f43ba9a49416374e4eda8e221b1942a648439db8a27ebe

Observation 04945ce8-166c-4c18-b9a1-af83f84ca9a4 · inbound

Multi-human Interactive Talking Dataset cites this paper.

Multi-human Interactive Talking Dataset EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.037670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.037670Z digest=sha256:49fbcf5b73f7afb45b66564c0e758d14a0f1a6fd13f8eba95aca20174369dab6

Observation 70b94a3a-6139-4b0a-a4e8-08319be5b1de · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:41.978583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:41.978583Z digest=sha256:e82c3ae6d789d70ffe2b5b5a9956156f079507f169212bc2b7312bc22c0c596b

Observation e47c006d-73d1-47e7-8e90-e6c2ecf9bc31 · inbound

Human Motion Video Generation: A Survey cites this paper.

Human Motion Video Generation: A Survey EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:52.470098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:52.470098Z digest=sha256:b5c08a3481252de1ec0c985f559fb00293c901904718639137e1c5aae0d8c728

Observation 748011d4-950d-4699-a4fc-47c388e0fae5 · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.425605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:eccf9c00ed13ee6fdfa3f1d5e597f33f4f1f8a57c0c87480a365766999f43697

Observation 189a8263-52db-49b3-8304-6ec9a721669a · inbound

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation cites this paper.

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:12.859273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T06:56:19.795651Z digest=sha256:58b0ac418d0c05b69b05fda5c2772b9b74c96c929b4d5537c5e98288a8baf8d7

Observation eb16d071-8429-4de8-9c1f-96a92f4a9664 · inbound

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models cites this paper.

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:17:48.701892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T21:14:25.606097Z digest=sha256:a39c6c39d3e7976be592a0551aaf70a139d86bab96393b730311ea7cde3e403a

Observation 2f11b0c3-665f-492c-a3e9-e6b505c18d59 · inbound

Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection cites this paper.

Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:03.459550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:48:03.459550Z digest=sha256:145431c6867743ccc6a80e3774610f14cd728a470807ef72b2716169d9e0efbe