Pith. sign in

Paper Citation Record · LEDGER

CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2409.01876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.01876 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:47:48.969366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T22:20:22.389622Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6d7f08c3-fe76-4136-905a-fa240c155c0a · inbound

FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation cites this paper.

FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T06:04:29.441503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T06:04:29.441503Z digest=sha256:19c7a6f56908597c4af2bf40b8d508f054638b0f91b5b8600f6a466dc75852f1

Observation 82600da0-044b-4226-bcbf-aad73e472a8b · inbound

Joint Learning of Depth and Appearance for Portrait Image Animation cites this paper.

Joint Learning of Depth and Appearance for Portrait Image Animation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:25:06.057752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:25:06.057752Z digest=sha256:8d264ca6bafe7be1742a5f80af684e39fcd0442fa29e751bbbccf6e6ab91993c

Observation a1c41a43-2780-4aa2-8759-0477c9e9f580 · inbound

EMO2: End-Effector Guided Audio-Driven Avatar Video Generation cites this paper.

EMO2: End-Effector Guided Audio-Driven Avatar Video Generation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T19:07:45.078919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:07:45.078919Z digest=sha256:fbcc0f891d19a856d071a4794602b70904e10ea07cb26a5c86793dac0f7f1c56

Observation 95f0d104-63d7-4cfe-b340-b57b9e71ad3d · inbound

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation cites this paper.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.518388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.518388Z digest=sha256:1c2306ae1193f09b1fd035fa62cc99b3f690c2a63b2fc415ed4431ee985ae2e0

Observation e7535f10-8dfe-4ede-9af2-d9dcacc654a6 · inbound

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters cites this paper.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.276147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.276147Z digest=sha256:6660ae6d14ef430804e863cb6a74203e4437c585bfc67caac737d79360a4236a

Observation 4a76ebc3-cbf2-4246-8763-529b4b596b24 · inbound

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models cites this paper.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.579865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.579865Z digest=sha256:2dbc3dc143098c2b281e7c9341a9b57710cb09f49f14763cb57c0d71c4025418

Observation cfd25029-0362-4163-8bca-e628478d122c · inbound

HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation cites this paper.

HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:15.069517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:07:15.069517Z digest=sha256:42f7953eb8be393d00469397a04cb4987b1a54a6445a175ae938acd115afb99d

Observation bef286cd-91bc-46eb-a97f-07d4dd01e09e · inbound

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation cites this paper.

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:47:48.969366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:47:48.969366Z digest=sha256:72cc2e758e1205b653e2feeccb489b7003d655bfb2245054fb0739204bca16ba

Observation 116419c8-75ac-4a35-a5cc-d9c71e95eac8 · inbound

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases cites this paper.

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:25.954676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:25.954676Z digest=sha256:492973fd165b90dea2999477ca6fc5d53aae79a9120ca7308367e3e67ff1b126

Observation ef699690-c396-468d-82c0-b52ebe43f6e2 · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:35.849232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:35.849232Z digest=sha256:f87681e115c996f027fd5e4fceca62e6f281bfa5f462e4fd8408196e230e8fcc

Observation 3022cb00-cc3f-403b-b724-35d80d10d2ff · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.173088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.173088Z digest=sha256:788b45e50260847534cafcd4b4d31f70355329f441c7fe71614a688aef619376

Observation 494aa347-bf7a-4c9f-9295-dede8215a7ae · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.393304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:ab0d644991eccbe96d5b243a1280434079b8397ca2ecc9ebec7b4e518997db98

Observation ee7d7163-f085-4500-bac4-0b306139681c · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:04.452982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:20a847c71dccef003209a0efbab20a6e6baf34ee58acc8c2e66a1d078b7c3713

Observation 8ac299f8-45cc-4570-b2e1-bb876963de7d · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:17.814966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:17.814966Z digest=sha256:9c81a460ae9ded08bf9192291af61b2baeaffac5c38c2e70ca75c19c19bd74cc

Observation 6396fe44-4f99-4e34-a201-139ff0df4eda · inbound

Multi-View Face and Gesture Animation with Dynamic Gaussians cites this paper.

Multi-View Face and Gesture Animation with Dynamic Gaussians CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:19.841581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:10:19.841581Z digest=sha256:66246fd56b6934514bc6beb494387b621eb5dac4a3a559879d1d7692460dc7a5