Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:38:20.590974Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.05092.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:38:20.590974Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8cee1c3d-4d3b-4c04-bcf2-587a78f80c85 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A morphable model for the synthesis of 3d faces
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4f8edd8-0cc6-47df-afb1-b065500818bb · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc95a1b9-b64f-41a4-bf41-efd907fe8f5c · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation You said that?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4d553b6-7fa0-4490-9a37-b2ebf5c3621f · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Lip reading sentences in the wild
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23f4c44d-469a-4e9a-879e-8adeea596953 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d9fa577-2ed1-4238-a56d-6653311e1b01 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Dae-talker: High fidelity speech-driven talking face generation with diffusion autoen- coder
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d3848b8e-6786-49dc-92a1-bafbcdb3aedf · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Generative adversarial networks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe11f5c-a2a4-4837-a253-a6c786e08c48 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Ad-nerf: Audio driven neural radi- ance fields for talking head synthesis
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 03a3c0e8-84a5-4ecf-810d-a754cbdf8bd8 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7c1bd97c-5501-4a64-aa92-1ca5b3f01f88 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation GAIA: Zero-shot Talking Avatar Generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f1a09b4b-2371-44f8-b2bc-3d9669cd06b6 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0328e1ad-3b23-491a-b9fe-563c3d8b4f84 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Implicit identity representation conditioned memory compensation network for talking head video generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe1ef8fa-8c66-4cda-a6bb-dd4350f8f3f8 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 167f1b13-1462-49b2-9c92-1bc5f0dac928 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Depth-aware generative adversarial network for talking head video generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa640803-062e-4a2a-9704-f5eec816209f · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f58e79-35e0-4779-a77f-f2f8913b2b09 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-driven emotional video portraits
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8243d0ef-8a90-4b24-a7cf-e347d3ca4fa7 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Transformers in vision: A survey
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4d0c738-bfa5-4595-a262-3ad63505262d · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Deep video portraits
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b6204282-38eb-45d7-8ddf-72ca548fb991 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Auto-Encoding Variational Bayes
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7716d64-440d-4cf7-8da5-33d4e7156fef · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6230e1cb-f0f8-44d4-910c-4e30cc395eec · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Moda: Mapping-once audio-driven portrait animation with dual attentions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc2d224c-570b-4ae8-9600-2000bb2cc771 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Live speech por- traits: real-time photorealistic talking-head animation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 602d3749-d771-43e3-9f6e-6b43b772e56d · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Training strategies for improved lip- reading
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3d446572-f400-4b64-92f0-d6db6a9cd57a · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c2e43db-8b8a-431b-b31b-810aef76b4fc · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0049d14a-a49a-4224-91ae-ef785403e837 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Librispeech: An asr corpus based on public do- main audio books
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 66454920-a701-4407-b670-4873c58a61a8 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A lip sync expert is all you need for speech to lip generation in the wild
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4d84c517-e4cb-4281-bbe7-e67d57a618de · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation High-resolution image syn- thesis with latent diffusion models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90859342-df51-45a5-affb-1005a4294afc · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation pytorch-fid: FID Score for PyTorch
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 182a4583-fc6d-40a7-98af-a6d4a49ca6cb · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Difftalk: Crafting diffusion models for generalized audio-driven portraits animation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eab7c841-fa7b-4c58-96b0-e37a38906e01 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation First order motion model for image animation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b328f7e-e45c-473a-b45c-e6a49c2fb524 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Denoising Diffusion Implicit Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db8b93ee-57c1-445e-8a74-f7e50d88e5a7 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking Face Generation by Conditional Recurrent Adversarial Network
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e5556e3-48c2-4283-9fd3-23c6d9d927bd · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Diffused heads: Diffusion models beat gans on talking-face genera- tion
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d31bcbb6-b8f8-42cb-b8a4-c4d34335d420 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c3df29d-292f-4f0d-a801-1f09a062b89d · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Masked lip-sync prediction by audio-visual contextual exploitation in transformers
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3d2b0f69-728b-419e-8371-57b07ec3b469 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synthesizing obama: learn- ing lip sync from audio
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e12a29a5-f076-4969-97b5-4afc1ff8e784 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Human-centric founda- tion models: Perception, generation and agentic modeling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 595754d1-34d3-431e-bb73-9d6e071cbc06 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Neural voice puppetry: Audio-driven facial reenactment
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a77f82c-db0b-40b5-b1ce-92998ccadec3 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19dcfa08-89cb-4ef2-88dc-3f3cc904297e · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4654cc85-ca5c-4d2b-8181-5c1c68606216 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 507c5c7c-e994-4ad0-91d6-a0c91649c2de · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Photorealistic audio-driven video portraits
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc2f2c89-b3ef-4d21-8fc4-bdd1eb374986 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Monocular depth estimation using multi-scale continuous crfs as sequential deep networks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4f8d4c40-5d73-4182-855e-64329f5e6bcd · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12e33c3-f04c-447c-9d90-d17c40df500b · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Jointly attentive spatial-temporal pooling networks for video-based person re-identification
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dca78684-f3f4-463c-bd87-e6cd7a31d531 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff6ae6e-ce79-4f35-a1a7-bbd36534cfdc · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 75e3770d-6dcf-4c39-a418-8e21e67f5dbd · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b552dea-6b52-4284-a90d-f2bb8f8158d8 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Thin-plate spline motion model for image animation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation be9364e6-806f-4294-96bd-d1152f147fb2 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synergizing motion and appearance: Multi-scale com- pensatory codebooks for talking head video generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8869ee41-c585-4503-9ad9-5b28eb0f11cc · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking face generation by adversarially disentangled audio-visual representation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 689b50cd-62c7-4faf-8a0e-6f3f39699a20 · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 16ce3c0d-c8d6-4417-98a8-ce28ee3543ca · outbound
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Makelttalk: speaker-aware talking-head animation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.