Pith. sign in

Paper Citation Record · LEDGER

Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2412.00733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00733 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:27.124310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:47:38.051126Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ed5ff8a9-c3c7-4bfe-9597-dfb59813d059 · inbound

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters cites this paper.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.124310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.124310Z digest=sha256:7adcc7219d7838131cce801ac70aa608eff057b42df91e3ccd80844d7d2a1ac6

Observation 70d63217-d6c4-46e1-bdd2-99b5f5f11c18 · inbound

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing cites this paper.

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:48.736504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:48.736504Z digest=sha256:a86f2ad8b30d85d38259b77b25c0cad57305f2c3c2d85f6c521a5fdbb95e9d7d

Observation 126822ad-cdb2-4732-b1c0-b087d9131350 · inbound

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation cites this paper.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.596042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.596042Z digest=sha256:495736abd84da264099a1ff1a8a037d211a400261baa5e161c06eb77a185bf2a

Observation c430a401-a2e7-4075-9517-77f260d9c4df · inbound

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers cites this paper.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.049606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.049606Z digest=sha256:c29acc0e6a71729f805c22d6eef7bc23cc091b5abf7640d17c405bbdbc4c9429

Observation 788de9a0-d16f-45c6-8538-c159f581c7c2 · inbound

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models cites this paper.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:55.731959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:55.731959Z digest=sha256:87ccba08ee524a474e12e65cbafe2edb5889ddddc81ddaa1c3a56de12b2f97d9

Observation e92e6c34-72c1-457f-94e0-37c71a4b7bc5 · inbound

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models cites this paper.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.346992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.346992Z digest=sha256:46fea5c3348dd4feb6ac4a42c30dc421182d236acab7e5fbecaa9df104afa9d6

Observation 6ab8f872-6a90-4bdd-b6ee-a89e061ba897 · inbound

Seeing Voices: Generating A-Roll Video from Audio with Mirage cites this paper.

Seeing Voices: Generating A-Roll Video from Audio with Mirage Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:08.242965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:08.242965Z digest=sha256:47d7e9dacaa094cd3999749a845ce7b510e58c1f26f5c6b94e03aacaa6ac96bf

Observation a1c3c98c-3551-4b07-8025-db60c0ef2481 · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.030141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:77af593b14da3b5c31579e22fd0e1c85e05c8426e76c514130a5220be26afdb8

Observation 02d17b77-7cb3-43cf-9141-75f12639f65e · inbound

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases cites this paper.

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:22.787467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:22.787467Z digest=sha256:5d67860ae58a9583535c30aefef47cce7caa874df530e9439b91fc36d56f49bd

Observation 827a515c-bd40-4c02-88f5-187defd38238 · inbound

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation cites this paper.

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:36.738210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:36.738210Z digest=sha256:b8089aa20f49925d0338b45af74053820f5e02e95b083214a3654a933f6ec025

Observation 0d167413-f8a8-4448-a2ba-0bbd79f161a2 · inbound

Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching cites this paper.

Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:13.223768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:20:13.223768Z digest=sha256:04616a96f26f82b8c8ce47ea32d5a2d31ce9b937839891a8b0939f016545c654

Observation 46495ded-d986-48ff-b43c-3b6fce0f7863 · inbound

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 cites this paper.

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:28.276988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:28.276988Z digest=sha256:9d3f2c24365e14e64d793c30d35b87ceab6f2f94d819cf0124b16344a21b7355

Observation d160ab63-8d60-43af-b7dd-56f2ed06878c · inbound

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos cites this paper.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.723175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.723175Z digest=sha256:5b7aa3a209bfac44ce27432051cf7d552f32df344bb6298a22997088762658b7

Observation 386640b5-47f1-4cba-a1f2-9d32c87c48cf · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:35.468621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:35.468621Z digest=sha256:cdab7b9db7079257483d06e249423e958b980021a8c65e5ac6fe706452e28cca

Observation 3ba92e8b-0db8-49f0-99a5-6d3c0095b1eb · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.217430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.217430Z digest=sha256:5c37fdec14f9f4f46108c0f9b96ab093685a01e1889c468764d2d8ede444a45a

Observation 9079251d-3480-4c4e-a654-a288346b7d44 · inbound

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation cites this paper.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.743448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.743448Z digest=sha256:928e912b4dd7cd0675fe4c207622503e79dd4100a69027c426a8971ed67c1042

Observation c85cec5c-d3de-468e-b6cd-8fe99dcf3647 · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.110895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.110895Z digest=sha256:a9f8560ba4a39b68832f7dedb50b790d358e25726639feab25d125a25dd92f35

Observation fe3aeb0e-de97-4d4d-9356-a6d32479a1b2 · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:27.217770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:30:36.462567Z digest=sha256:9567bf01d0f1b9855bd469d4d21020e398902fcdf4a4ca0a88a1ac11aa6e4517

Observation d3e81ac7-1f87-4158-be19-1e92286e6a08 · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T16:38:16.047826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:38:16.047826Z digest=sha256:638df7dbb9ad465c6a4c5e11dba4df82d3f392a45e40993ebb4909569ac7b311

Observation 36be6018-cd74-4d97-ad99-7dd1b6af3e07 · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.234694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:56df0b97df4fb2ab82c9c427d67654406f7ce9cc53e3c22b225edc8a1d0cb4e1

Observation c355550e-f74e-45b4-85b2-2d69bc87576e · inbound

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization cites this paper.

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.052617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:40:14.506288Z digest=sha256:b9a9f35552962b2ae066218b1a572863ba0ad6dd81eaafa10f99b335423a2e38