Pith. sign in

Paper Citation Record · LEDGER

Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2107.09293.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2107.09293 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:00:25.336103Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:46:15.581615Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f00ab899-3afb-4a75-92b6-fdef16094941 · inbound

EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion cites this paper.

EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:34.727893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:34.727893Z digest=sha256:8aba09766150e25e687eee28bcf62db8ad3cdf3cf0b78528084f1dcb0eb5901c

Observation e54efc39-3ba1-4a82-93b9-1890e0c57b89 · inbound

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model cites this paper.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.031619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.031619Z digest=sha256:aecaab122d9df893a908b09eda3d2bbc12acbcfa6a75c5c3754ed6d760d7d59a

Observation c14eba74-4fe0-4f7f-ac29-7aa1bf51a4c0 · inbound

UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control cites this paper.

UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:02.564160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:02.564160Z digest=sha256:344c204c63356c8399f6622cc01b001d10df15ae9b50cdc2a034dde240b30cd0

Observation 4fafed32-fa17-4481-810f-3c079d939cc3 · inbound

Identity-Preserving Video Dubbing Using Motion Warping cites this paper.

Identity-Preserving Video Dubbing Using Motion Warping Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:33:20.028036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:33:20.028036Z digest=sha256:345074e6de21d78238e7c1ccd8df0fc067bed1d0890840f05a0a8342f3d42d72

Observation bfe44a18-122c-4c00-bf34-d3f42b90037c · inbound

Joint Learning of Depth and Appearance for Portrait Image Animation cites this paper.

Joint Learning of Depth and Appearance for Portrait Image Animation Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T20:25:06.193141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:25:06.193141Z digest=sha256:60e9048a9818100c35bd0060a9a913e6d5189238640e3988847743a8dda2542c

Observation cb04df4c-a84f-4ec1-93a3-b2fd16aa261e · inbound

EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis cites this paper.

EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T18:16:40.933559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:16:40.933559Z digest=sha256:7cd6ade0b8fc3714859bcca363d7895bed4a02a213f2ca0401c7e5c86095cc9d

Observation c6111573-4bb5-4e1a-a5fb-e502009c9c16 · inbound

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities cites this paper.

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-09T18:23:42.444845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:23:42.444845Z digest=sha256:19dd31f49ce1a7631fe97bc814b0f8d64d9c5a762b42e81b328d1d7380a229b8

Observation 4990ae01-00a6-440d-94d2-4aee3b75ffcb · inbound

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model cites this paper.

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T21:11:25.336787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:11:25.336787Z digest=sha256:2d0b929250345b11e912178acbd339e4880b317353e2373893924cbb36cdd146

Observation 49565245-a006-4168-b27b-68b18c85b1c1 · inbound

NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results cites this paper.

NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:51.176183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:51.176183Z digest=sha256:37dbde088a56afeb2c2a1f89c03551db48e252dcfe0ea8870cc9e182f2392488

Observation 5f3739da-5e95-40cc-ad27-7280efcb2de9 · inbound

SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting cites this paper.

SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.518622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:44.518622Z digest=sha256:7d015a32ddb0c85a7fc7256f79911e1e3238da926f2e77c1f858ef646636cc9c

Observation 528e5442-e7fe-46db-9ce9-340f80441f27 · inbound

Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation cites this paper.

Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.336103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.336103Z digest=sha256:8d89ac515e2cb5f0ee0353fb476c91c9dc2d79d780d9727e18452018b28b0107

Observation b549b02b-96c3-43dc-8314-61caf0adf92d · inbound

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads cites this paper.

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:19.120381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:19.120381Z digest=sha256:e33390ab7ce7bcf0e7e7538c9c79291165ffc36f775067b074ef737ba3df351e

Observation 1b9799a9-e6e2-4e32-af47-421865ccbb43 · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:50:40.246667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:f54b4ce7f02d51e0c5111b42dd25fecad69aa3000eb767eed1923be770276597

Observation 707c910a-ce2d-4375-ba78-2e5ec50a3b58 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:04.425900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:b109996c15e0ce172acad659af06071e698584ba057071012b7f9c685ca75d6d

Observation 98ed7b23-2b3a-4cd5-9d47-00028f0b7d70 · inbound

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation cites this paper.

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:46:15.583006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T16:19:55.048831Z digest=sha256:4cf1910f50ec5d632e63f3f762832683fed5043fdc01ecc0ab3fd5dbbe6bba75

Observation dbf346d2-dc65-4f91-9949-97b7be03f722 · inbound

Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation cites this paper.

Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T17:26:50.150881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:26:50.150881Z digest=sha256:98f2151990c632d6308a9c18ff9b79bc589576a4f547899eeac8c68606ec170f