Pith. sign in

Paper Citation Record · LEDGER

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.05092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05092 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:38:20.590974Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cee1c3d-4d3b-4c04-bcf2-587a78f80c85 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A morphable model for the synthesis of 3d faces

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.736538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.430133Z digest=sha256:91f3dd13d33e2ef8dc198bda904ae9167f34558c814a8f4d2f5682a56f23c923

Observation b4f8edd8-0cc6-47df-afb1-b065500818bb · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.513004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.433575Z digest=sha256:4da8704a26a0856fbbe0d324b90599fc24bd829c46af4657356050dc74b07d3e

Observation fc95a1b9-b64f-41a4-bf41-efd907fe8f5c · outbound

This paper cites You said that?.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation You said that?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:21.116132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.437079Z digest=sha256:2f28c375b13b76f4876f9ededf0445362fa1cd4d4e737890e5377b862d42f0d3

Observation b4d553b6-7fa0-4490-9a37-b2ebf5c3621f · outbound

This paper cites Lip reading sentences in the wild.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Lip reading sentences in the wild

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.189584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.440540Z digest=sha256:7ef9e88873bf6c14de03b2079ececb5d6044651a76bcb7c5cc90a14292ebf533

Observation 23f4c44d-469a-4e9a-879e-8adeea596953 · outbound

This paper cites Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.858375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.443919Z digest=sha256:b46c1e852bbf4d20efb5948f9a0a88ee89f822f538a810cd731d4eb2843f49b6

Observation 1d9fa577-2ed1-4238-a56d-6653311e1b01 · outbound

This paper cites Dae-talker: High fidelity speech-driven talking face generation with diffusion autoen- coder.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Dae-talker: High fidelity speech-driven talking face generation with diffusion autoen- coder

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.574230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.446762Z digest=sha256:2c5847f8ca9de9cb08ac5f1ef9f87072df62f722ffa20efdcfb6faefd7e1ff60

Observation d3848b8e-6786-49dc-92a1-bafbcdb3aedf · outbound

This paper cites Generative adversarial networks.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Generative adversarial networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.449788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.449788Z digest=sha256:738e2fb344a02e70854d51264cfc62733eebc548b138406a4a92bea7718bcdf2

Observation 5fe11f5c-a2a4-4837-a253-a6c786e08c48 · outbound

This paper cites Ad-nerf: Audio driven neural radi- ance fields for talking head synthesis.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Ad-nerf: Audio driven neural radi- ance fields for talking head synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.242496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.452970Z digest=sha256:a396076ce5804016401bf5b532c860fdeb7a428a9db8ccd005696c669d14b9a3

Observation 03a3c0e8-84a5-4ecf-810d-a754cbdf8bd8 · outbound

This paper cites Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.887210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.455770Z digest=sha256:6c3c50b6ca29e672575a60ffa4a41507ffcc2e9950aa643c3debeb0129ac845a

Observation 7c1bd97c-5501-4a64-aa92-1ca5b3f01f88 · outbound

This paper cites GAIA: Zero-shot Talking Avatar Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation GAIA: Zero-shot Talking Avatar Generation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:20.866922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.458871Z digest=sha256:cbfd922f2b4f0dcff527da0860899ac9b948c1337a3a693dccea8a3cddf333ce

Observation f1a09b4b-2371-44f8-b2bc-3d9669cd06b6 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.593769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.462422Z digest=sha256:5684ca8277e9333de3065fb12542fc2039de5edbfc189bc1a8b4b31b571f6bf1

Observation 0328e1ad-3b23-491a-b9fe-563c3d8b4f84 · outbound

This paper cites Implicit identity representation conditioned memory compensation network for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Implicit identity representation conditioned memory compensation network for talking head video generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.281921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.465191Z digest=sha256:b0b6aac8788b48ab381a7819848588c3479c22d11b1263a86ca9bf55797b8688

Observation fe1ef8fa-8c66-4cda-a6bb-dd4350f8f3f8 · outbound

This paper cites DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:20.756341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.467907Z digest=sha256:06c9cc801dee59bd0492f2d64399034483f0dc64e9eac11d067a7b9fd6e83e95

Observation 167f1b13-1462-49b2-9c92-1bc5f0dac928 · outbound

This paper cites Depth-aware generative adversarial network for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Depth-aware generative adversarial network for talking head video generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.009063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.471010Z digest=sha256:61434c6c020a5511902f4ffa31169783a4787bcc65faa8f41f7c5dab6bf7e079

Observation aa640803-062e-4a2a-9704-f5eec816209f · outbound

This paper cites Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.474478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.474478Z digest=sha256:a96081bb8dcab2557e3e8eba784db299bf73088a26d9cdc7fc11d6e09d7429cf

Observation 84f58e79-35e0-4779-a77f-f2f8913b2b09 · outbound

This paper cites Audio-driven emotional video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-driven emotional video portraits

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.781678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.481151Z digest=sha256:713f1f491795742f2795dc3320e5809ed2613e7613745ae78031ebd6bcfb64e4

Observation 8243d0ef-8a90-4b24-a7cf-e347d3ca4fa7 · outbound

This paper cites Transformers in vision: A survey.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Transformers in vision: A survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.506549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.484001Z digest=sha256:59f7a79291e30aff4ba8091b1b8cbc6e1f43b9880916b1e664103ea933f43920

Observation d4d0c738-bfa5-4595-a262-3ad63505262d · outbound

This paper cites Deep video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Deep video portraits

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.275743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.486816Z digest=sha256:c6e2796f248585867928be9d1c5b8b2d6daf443ec4ab513a56bf50c19b68e993

Observation b6204282-38eb-45d7-8ddf-72ca548fb991 · outbound

This paper cites Auto-Encoding Variational Bayes.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Auto-Encoding Variational Bayes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.489513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.489513Z digest=sha256:d891672536e604548683dc958a4eac472705a4535fe8cf9fca4f6c7e0064fad3

Observation e7716d64-440d-4cf7-8da5-33d4e7156fef · outbound

This paper cites AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.492669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.492669Z digest=sha256:1bcc8cf0776d11a9086ce93045b0a576f55a642e7e1d7375984acc03d5208acd

Observation 6230e1cb-f0f8-44d4-910c-4e30cc395eec · outbound

This paper cites Moda: Mapping-once audio-driven portrait animation with dual attentions.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Moda: Mapping-once audio-driven portrait animation with dual attentions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.994177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.496048Z digest=sha256:2e38078a9e08ca3633dce6b5c2153cee0320d62344df33381694dc7d55e2ff51

Observation cc2d224c-570b-4ae8-9600-2000bb2cc771 · outbound

This paper cites Live speech por- traits: real-time photorealistic talking-head animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Live speech por- traits: real-time photorealistic talking-head animation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.776730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.498870Z digest=sha256:69114ba9fe69cd0829ffc16e696e661e112ad119640d596988fa8285f3c02ab6

Observation 602d3749-d771-43e3-9f6e-6b43b772e56d · outbound

This paper cites Training strategies for improved lip- reading.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Training strategies for improved lip- reading

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.437494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.501957Z digest=sha256:1fc041684c839428daad0f0251358d895a0fd1986a25c74e4ea17dc495978317

Observation 3d446572-f400-4b64-92f0-d6db6a9cd57a · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.504641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.504641Z digest=sha256:daf1ef62232e60214426baa521266619c320568ccb8c4d88c427a1218e64c352

Observation 9c2e43db-8b8a-431b-b31b-810aef76b4fc · outbound

This paper cites DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.507706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.507706Z digest=sha256:eafca5b76a28808f55ef9d0623cac4e12dc9d1a2a41f178dede2e11e3afd644e

Observation 0049d14a-a49a-4224-91ae-ef785403e837 · outbound

This paper cites Librispeech: An asr corpus based on public do- main audio books.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Librispeech: An asr corpus based on public do- main audio books

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.182893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.510785Z digest=sha256:7c41f985ccae38243cbdc0c0db6c3e724901bee0193c40a9a7bf3f96e8a9080d

Observation 66454920-a701-4407-b670-4873c58a61a8 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A lip sync expert is all you need for speech to lip generation in the wild

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.916830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.513555Z digest=sha256:c7dbdbef56cb544313e7a726366df80edd63e1e1c038951c6d0b1e8b97028fe4

Observation 4d84c517-e4cb-4281-bbe7-e67d57a618de · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation High-resolution image syn- thesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.657706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.516395Z digest=sha256:3b60067a77646f532d9dabb7c9ed053b040c0556b93dc4349afb13efa83a535e

Observation 90859342-df51-45a5-affb-1005a4294afc · outbound

This paper cites pytorch-fid: FID Score for PyTorch.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation pytorch-fid: FID Score for PyTorch

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.519303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.519303Z digest=sha256:6b120b4c80d0008b75e3faf62a802e6108f94c5e526ebdf6c95c454f17d1cc91

Observation 182a4583-fc6d-40a7-98af-a6d4a49ca6cb · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.321835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.522395Z digest=sha256:e2bdcbaf4baaa2cefb5c53f19402c7b873b5df451a423d87e7b4059abb512932

Observation eab7c841-fa7b-4c58-96b0-e37a38906e01 · outbound

This paper cites First order motion model for image animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation First order motion model for image animation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.005868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.525008Z digest=sha256:ceb1d189bfee1d69fb5e77e52767f426d82aac6d533e679d9c29c9a2520b002f

Observation 6b328f7e-e45c-473a-b45c-e6a49c2fb524 · outbound

This paper cites Denoising Diffusion Implicit Models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Denoising Diffusion Implicit Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.527573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.527573Z digest=sha256:d1703bc6d83c92f682fec8bac36139c379c1e366481089335b29dd5fb1572733

Observation db8b93ee-57c1-445e-8a74-f7e50d88e5a7 · outbound

This paper cites Talking Face Generation by Conditional Recurrent Adversarial Network.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking Face Generation by Conditional Recurrent Adversarial Network

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.530316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.530316Z digest=sha256:e8aefc8b64c0861324c417c1f1713dadb425bbcaeb904733a9b9c328b6b59baa

Observation 5e5556e3-48c2-4283-9fd3-23c6d9d927bd · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.766718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.533486Z digest=sha256:156c2ea8f29f5488db5262ab490e135b0533233d1887c246a83ad7e7139031af

Observation d31bcbb6-b8f8-42cb-b8a4-c4d34335d420 · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.536182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.536182Z digest=sha256:7fd6d1b598111d9fcdee0dda3831b640dfd61850d6ec08d29b8d726179c56ec9

Observation 3c3df29d-292f-4f0d-a801-1f09a062b89d · outbound

This paper cites Masked lip-sync prediction by audio-visual contextual exploitation in transformers.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Masked lip-sync prediction by audio-visual contextual exploitation in transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.596939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.539264Z digest=sha256:50b72c39e7ae6436689e38cb2f7b02c692d48761239c7f0b52fa1994a67807eb

Observation 3d2b0f69-728b-419e-8371-57b07ec3b469 · outbound

This paper cites Synthesizing obama: learn- ing lip sync from audio.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synthesizing obama: learn- ing lip sync from audio

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.416179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.542272Z digest=sha256:ccd344421949d7e508f90623d355189285e62de548b2bc638c15d01401167edb

Observation e12a29a5-f076-4969-97b5-4afc1ff8e784 · outbound

This paper cites Human-centric founda- tion models: Perception, generation and agentic modeling.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Human-centric founda- tion models: Perception, generation and agentic modeling

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.155157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.545160Z digest=sha256:98af7e551b21a6d7fa502c77c26c30cfc8e15dc1cb4020365add6cba853c9b85

Observation 595754d1-34d3-431e-bb73-9d6e071cbc06 · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Neural voice puppetry: Audio-driven facial reenactment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.933754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.547678Z digest=sha256:ddd4a020c4628825897e90754ee080919b7a5019bea88227627a0b25dc264245

Observation 4a77f82c-db0b-40b5-b1ce-92998ccadec3 · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.550356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.550356Z digest=sha256:61424598257f787669da0cf904eaf2ed99ee3f53da718a1b5740b80775fdbab9

Observation 19dcfa08-89cb-4ef2-88dc-3f3cc904297e · outbound

This paper cites Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.698180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.553474Z digest=sha256:29523cc665d57b4c1b2a4a421359f8de70dad1b5625e7730b2185d0034ab660f

Observation 4654cc85-ca5c-4d2b-8181-5c1c68606216 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.556209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.556209Z digest=sha256:747e57470c63710ddf48e185428c0c2b81ae1fd06370050c0e6d5398ae2a9c9e

Observation 507c5c7c-e994-4ad0-91d6-a0c91649c2de · outbound

This paper cites Photorealistic audio-driven video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Photorealistic audio-driven video portraits

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.520340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.559258Z digest=sha256:635a5b5482436eef14150a4fa2d304c397c32d4c2fa2dffd20bf0f52090083ab

Observation bc2f2c89-b3ef-4d21-8fc4-bdd1eb374986 · outbound

This paper cites Monocular depth estimation using multi-scale continuous crfs as sequential deep networks.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Monocular depth estimation using multi-scale continuous crfs as sequential deep networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.248574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.562097Z digest=sha256:e1e7d3cecee1a62c00427c16e4ac82855c481ba57ec17fc79c49706124fce936

Observation 4f8d4c40-5d73-4182-855e-64329f5e6bcd · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.565284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.565284Z digest=sha256:e35517532502825ae9d60065c6f678ccaabbb5c8f0d99e2a687e43085bc22531

Observation d12e33c3-f04c-447c-9d90-d17c40df500b · outbound

This paper cites Jointly attentive spatial-temporal pooling networks for video-based person re-identification.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Jointly attentive spatial-temporal pooling networks for video-based person re-identification

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.054479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.568145Z digest=sha256:41d5ffaa882975af5f9dc05cdfe6fae9002b708d2aba39f577a5d2aa62d2df9d

Observation dca78684-f3f4-463c-bd87-e6cd7a31d531 · outbound

This paper cites VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.570734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.570734Z digest=sha256:5b9c4285c45fb5dff868374f591d4fb398f13899a01886555d713e9f328d7588

Observation dff6ae6e-ce79-4f35-a1a7-bbd36534cfdc · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.788464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.573841Z digest=sha256:54ddecf12831bc3cf2b4bf190bc942696a1f27496344be0b638d72380dbbcada

Observation 75e3770d-6dcf-4c39-a418-8e21e67f5dbd · outbound

This paper cites Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.536446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.576418Z digest=sha256:a7684f038d5bc872a1cc00d7c7130981343911fdcf0422de136ade1feec2a008

Observation 5b552dea-6b52-4284-a90d-f2bb8f8158d8 · outbound

This paper cites Thin-plate spline motion model for image animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Thin-plate spline motion model for image animation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.300827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.579388Z digest=sha256:a137bfa40cccc335118d955a4b5ded70d65dda162c4d3eeb61423d5f00e196e6

Observation be9364e6-806f-4294-96bd-d1152f147fb2 · outbound

This paper cites Synergizing motion and appearance: Multi-scale com- pensatory codebooks for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synergizing motion and appearance: Multi-scale com- pensatory codebooks for talking head video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.071391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.582069Z digest=sha256:5d86c7e44f2addff3da92c3663bd5056fb27e8b73025448b4f670bdc4faa3f6b

Observation 8869ee41-c585-4503-9ad9-5b28eb0f11cc · outbound

This paper cites Talking face generation by adversarially disentangled audio-visual representation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking face generation by adversarially disentangled audio-visual representation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.866644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.585041Z digest=sha256:438206696a52fad2f14518ca483d05a409bc7474c1b1154d775848bd014db606

Observation 689b50cd-62c7-4faf-8a0e-6f3f39699a20 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.588387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.588203Z digest=sha256:5f475a972da87a64b355d03d0a3307b2421334ae96becad2e61180299d792699

Observation 16ce3c0d-c8d6-4417-98a8-ce28ee3543ca · outbound

This paper cites Makelttalk: speaker-aware talking-head animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Makelttalk: speaker-aware talking-head animation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.405568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:38:20.590974Z digest=sha256:9717ac1d0233f156be06f1444ddff436a3d19d53c1715ed84fab8a8306754802

Pith citing papers

No inbound Pith citation observations are available.