Pith. sign in

Paper Citation Record · LEDGER

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.05092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05092 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:38:20.590974Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cee1c3d-4d3b-4c04-bcf2-587a78f80c85 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A morphable model for the synthesis of 3d faces

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.736538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.430133Z digest=sha256:63b111e049dee9801d2be2f3111b134b0bca7dc097a5750ce8780207aa07a89e

Observation b4f8edd8-0cc6-47df-afb1-b065500818bb · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.513004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.433575Z digest=sha256:a327313a13d24fba28ad96ce22c9380e2d047e451e7743a51e1d514647baf7db

Observation fc95a1b9-b64f-41a4-bf41-efd907fe8f5c · outbound

This paper cites You said that?.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation You said that?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:21.116132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.437079Z digest=sha256:f8ac5b433b0b7df833bf98c367f3aaf7f76dd898f70492b4a1caee80847d4798

Observation b4d553b6-7fa0-4490-9a37-b2ebf5c3621f · outbound

This paper cites Lip reading sentences in the wild.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Lip reading sentences in the wild

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.189584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.440540Z digest=sha256:0d755218082254d64cfbd959241bc6732b7abcb721749fe7f553d814943530c3

Observation 23f4c44d-469a-4e9a-879e-8adeea596953 · outbound

This paper cites Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.858375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.443919Z digest=sha256:688e9fcc3aaa2cfb81a10d98cb5ffe5683f16e047470e3124dabad0aa031fd8a

Observation 1d9fa577-2ed1-4238-a56d-6653311e1b01 · outbound

This paper cites Dae-talker: High fidelity speech-driven talking face generation with diffusion autoen- coder.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Dae-talker: High fidelity speech-driven talking face generation with diffusion autoen- coder

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.574230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.446762Z digest=sha256:6393498a97e52fceeef047a755c5d734f59389a02d42ffdaacf9490a64762a6d

Observation d3848b8e-6786-49dc-92a1-bafbcdb3aedf · outbound

This paper cites Generative adversarial networks.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Generative adversarial networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.449788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.449788Z digest=sha256:738e2fb344a02e70854d51264cfc62733eebc548b138406a4a92bea7718bcdf2

Observation 5fe11f5c-a2a4-4837-a253-a6c786e08c48 · outbound

This paper cites Ad-nerf: Audio driven neural radi- ance fields for talking head synthesis.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Ad-nerf: Audio driven neural radi- ance fields for talking head synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.242496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.452970Z digest=sha256:12b503cd3a9bbb7c2734f9b64bd93d1473c1f83e522f89c3d7952b1f2e8082d4

Observation 03a3c0e8-84a5-4ecf-810d-a754cbdf8bd8 · outbound

This paper cites Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.887210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.455770Z digest=sha256:b0eafe57a91915c3c63f59404827055bc5caecdd4b1b75d145e96062d9da78ab

Observation 7c1bd97c-5501-4a64-aa92-1ca5b3f01f88 · outbound

This paper cites GAIA: Zero-shot Talking Avatar Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation GAIA: Zero-shot Talking Avatar Generation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:20.866922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.458871Z digest=sha256:c8ceafb6325a2e0dfe9c56c32ef838b2496aa705652c0a16c891bbfd222d062b

Observation f1a09b4b-2371-44f8-b2bc-3d9669cd06b6 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.593769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.462422Z digest=sha256:ffa83eedbba37749dfd314b36d5eadc364ba11bbeafedc6207ee376cfcece6b5

Observation 0328e1ad-3b23-491a-b9fe-563c3d8b4f84 · outbound

This paper cites Implicit identity representation conditioned memory compensation network for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Implicit identity representation conditioned memory compensation network for talking head video generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.281921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.465191Z digest=sha256:3c81589a21418fe01c398130c98718613e4b355a5f0fb4c663541ac5f42294f4

Observation fe1ef8fa-8c66-4cda-a6bb-dd4350f8f3f8 · outbound

This paper cites DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:20.756341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.467907Z digest=sha256:0731afba4f95944fa9ef6662d1f02fb99db8f214797a3d0e18c98cef825d8a88

Observation 167f1b13-1462-49b2-9c92-1bc5f0dac928 · outbound

This paper cites Depth-aware generative adversarial network for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Depth-aware generative adversarial network for talking head video generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.009063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.471010Z digest=sha256:d74f2d70c2e49f430d406499d4cfb3e1f45939165b487d94609ddadbee585e91

Observation aa640803-062e-4a2a-9704-f5eec816209f · outbound

This paper cites Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.474478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.474478Z digest=sha256:a96081bb8dcab2557e3e8eba784db299bf73088a26d9cdc7fc11d6e09d7429cf

Observation 84f58e79-35e0-4779-a77f-f2f8913b2b09 · outbound

This paper cites Audio-driven emotional video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-driven emotional video portraits

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.781678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.481151Z digest=sha256:3d35aeb06f37a294f20b29e34ab0cfde5250ea7d933fb8e3cde94a3ccb04be3e

Observation 8243d0ef-8a90-4b24-a7cf-e347d3ca4fa7 · outbound

This paper cites Transformers in vision: A survey.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Transformers in vision: A survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.506549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.484001Z digest=sha256:310ee3724b12b830f68d57d12187d2c6bb22aae61777d931e5c482cd006bff50

Observation d4d0c738-bfa5-4595-a262-3ad63505262d · outbound

This paper cites Deep video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Deep video portraits

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.275743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.486816Z digest=sha256:86785250d01da1c87380c9c621b16a673b0520ba15ad180310e2873f4a5ed210

Observation b6204282-38eb-45d7-8ddf-72ca548fb991 · outbound

This paper cites Auto-Encoding Variational Bayes.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Auto-Encoding Variational Bayes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.489513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.489513Z digest=sha256:d891672536e604548683dc958a4eac472705a4535fe8cf9fca4f6c7e0064fad3

Observation e7716d64-440d-4cf7-8da5-33d4e7156fef · outbound

This paper cites AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.492669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.492669Z digest=sha256:1bcc8cf0776d11a9086ce93045b0a576f55a642e7e1d7375984acc03d5208acd

Observation 6230e1cb-f0f8-44d4-910c-4e30cc395eec · outbound

This paper cites Moda: Mapping-once audio-driven portrait animation with dual attentions.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Moda: Mapping-once audio-driven portrait animation with dual attentions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.994177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.496048Z digest=sha256:218a7a45bcb0d1a163e2a82b591c9b2870d3fdae457857ac3803d02760e99b34

Observation cc2d224c-570b-4ae8-9600-2000bb2cc771 · outbound

This paper cites Live speech por- traits: real-time photorealistic talking-head animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Live speech por- traits: real-time photorealistic talking-head animation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.776730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.498870Z digest=sha256:49a2916f841e1338e4547508390dd849927780480e9e46dd557478d5688f6ba4

Observation 602d3749-d771-43e3-9f6e-6b43b772e56d · outbound

This paper cites Training strategies for improved lip- reading.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Training strategies for improved lip- reading

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.437494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.501957Z digest=sha256:648502620769fb8c0d868a5e2bbc73a6ffb378d7288b0070dac77d416fb5c235

Observation 3d446572-f400-4b64-92f0-d6db6a9cd57a · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.504641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.504641Z digest=sha256:daf1ef62232e60214426baa521266619c320568ccb8c4d88c427a1218e64c352

Observation 9c2e43db-8b8a-431b-b31b-810aef76b4fc · outbound

This paper cites DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.507706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.507706Z digest=sha256:eafca5b76a28808f55ef9d0623cac4e12dc9d1a2a41f178dede2e11e3afd644e

Observation 0049d14a-a49a-4224-91ae-ef785403e837 · outbound

This paper cites Librispeech: An asr corpus based on public do- main audio books.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Librispeech: An asr corpus based on public do- main audio books

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.182893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.510785Z digest=sha256:e877c8fa5b1b764d914cfe81a90be717d28c3ccf741ac73c142fb085bf4495f1

Observation 66454920-a701-4407-b670-4873c58a61a8 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A lip sync expert is all you need for speech to lip generation in the wild

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.916830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.513555Z digest=sha256:c4355fd8c15537159b635d7cdd4f209a9ef02435699d7e9c64adac302ca1f8b8

Observation 4d84c517-e4cb-4281-bbe7-e67d57a618de · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation High-resolution image syn- thesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.657706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.516395Z digest=sha256:c437bde4ebf3dbef14dcd1f42237051ede1be653641c2dbe7c449b9682a72a8b

Observation 90859342-df51-45a5-affb-1005a4294afc · outbound

This paper cites pytorch-fid: FID Score for PyTorch.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation pytorch-fid: FID Score for PyTorch

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.519303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.519303Z digest=sha256:6b120b4c80d0008b75e3faf62a802e6108f94c5e526ebdf6c95c454f17d1cc91

Observation 182a4583-fc6d-40a7-98af-a6d4a49ca6cb · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.321835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.522395Z digest=sha256:9b42c7c13202a28b657f0762086cad9d24608a3baaf05441c6b78c83aeeb221d

Observation eab7c841-fa7b-4c58-96b0-e37a38906e01 · outbound

This paper cites First order motion model for image animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation First order motion model for image animation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.005868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.525008Z digest=sha256:7d9a87d4de62418e3134b83e36823a1d134f9dac5e373a8e3a956f804eb665f8

Observation 6b328f7e-e45c-473a-b45c-e6a49c2fb524 · outbound

This paper cites Denoising Diffusion Implicit Models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Denoising Diffusion Implicit Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.527573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.527573Z digest=sha256:d1703bc6d83c92f682fec8bac36139c379c1e366481089335b29dd5fb1572733

Observation db8b93ee-57c1-445e-8a74-f7e50d88e5a7 · outbound

This paper cites Talking Face Generation by Conditional Recurrent Adversarial Network.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking Face Generation by Conditional Recurrent Adversarial Network

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.530316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.530316Z digest=sha256:e8aefc8b64c0861324c417c1f1713dadb425bbcaeb904733a9b9c328b6b59baa

Observation 5e5556e3-48c2-4283-9fd3-23c6d9d927bd · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.766718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.533486Z digest=sha256:5a69be1cc86794ee0d181b9b1a4b105d3c73f2570bcb971b0e45d430f3b0b7a8

Observation d31bcbb6-b8f8-42cb-b8a4-c4d34335d420 · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.536182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.536182Z digest=sha256:7fd6d1b598111d9fcdee0dda3831b640dfd61850d6ec08d29b8d726179c56ec9

Observation 3c3df29d-292f-4f0d-a801-1f09a062b89d · outbound

This paper cites Masked lip-sync prediction by audio-visual contextual exploitation in transformers.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Masked lip-sync prediction by audio-visual contextual exploitation in transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.596939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.539264Z digest=sha256:451dde754d759269f6ba89b57c4a2b0ef7276c9daec8e1db92f10a2398600540

Observation 3d2b0f69-728b-419e-8371-57b07ec3b469 · outbound

This paper cites Synthesizing obama: learn- ing lip sync from audio.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synthesizing obama: learn- ing lip sync from audio

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.416179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.542272Z digest=sha256:b825ef4e51ea63c404264a5a231de56c96025adf9fd5f2b1a1a02d5f1255d83e

Observation e12a29a5-f076-4969-97b5-4afc1ff8e784 · outbound

This paper cites Human-centric founda- tion models: Perception, generation and agentic modeling.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Human-centric founda- tion models: Perception, generation and agentic modeling

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.155157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.545160Z digest=sha256:22c0b9bedce0653649c08792762c7e671e6375ee0779ed1768d2beece6e792d3

Observation 595754d1-34d3-431e-bb73-9d6e071cbc06 · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Neural voice puppetry: Audio-driven facial reenactment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.933754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.547678Z digest=sha256:1e80aac854ee66caf595a52446af7e05e7da63c3b6d8014e23cc975f0ae10f25

Observation 4a77f82c-db0b-40b5-b1ce-92998ccadec3 · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.550356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.550356Z digest=sha256:61424598257f787669da0cf904eaf2ed99ee3f53da718a1b5740b80775fdbab9

Observation 19dcfa08-89cb-4ef2-88dc-3f3cc904297e · outbound

This paper cites Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.698180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.553474Z digest=sha256:7959124d13e20ddeccbddfd4d64f1376ffbdc8743b8729000cb840f0c01b06c9

Observation 4654cc85-ca5c-4d2b-8181-5c1c68606216 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.556209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.556209Z digest=sha256:747e57470c63710ddf48e185428c0c2b81ae1fd06370050c0e6d5398ae2a9c9e

Observation 507c5c7c-e994-4ad0-91d6-a0c91649c2de · outbound

This paper cites Photorealistic audio-driven video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Photorealistic audio-driven video portraits

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.520340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.559258Z digest=sha256:8eec596b4fd5ff5c796286318d80e7896ab2f2ff88932d64d05abf85820aa662

Observation bc2f2c89-b3ef-4d21-8fc4-bdd1eb374986 · outbound

This paper cites Monocular depth estimation using multi-scale continuous crfs as sequential deep networks.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Monocular depth estimation using multi-scale continuous crfs as sequential deep networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.248574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.562097Z digest=sha256:c444905eed11c27c91e97ab5d7b2410585acf924f6dc6173a21e27fd8247a04e

Observation 4f8d4c40-5d73-4182-855e-64329f5e6bcd · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.565284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.565284Z digest=sha256:e35517532502825ae9d60065c6f678ccaabbb5c8f0d99e2a687e43085bc22531

Observation d12e33c3-f04c-447c-9d90-d17c40df500b · outbound

This paper cites Jointly attentive spatial-temporal pooling networks for video-based person re-identification.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Jointly attentive spatial-temporal pooling networks for video-based person re-identification

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.054479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.568145Z digest=sha256:ea578a05dd652616881a8de41ad8c52468d3f47bc509f8df3558e8e689d939ff

Observation dca78684-f3f4-463c-bd87-e6cd7a31d531 · outbound

This paper cites VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.570734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.570734Z digest=sha256:5b9c4285c45fb5dff868374f591d4fb398f13899a01886555d713e9f328d7588

Observation dff6ae6e-ce79-4f35-a1a7-bbd36534cfdc · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.788464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.573841Z digest=sha256:86536e5bec4a1429f107844080416032ac0169d8f1b4d910dcd696a8545a59f2

Observation 75e3770d-6dcf-4c39-a418-8e21e67f5dbd · outbound

This paper cites Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.536446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.576418Z digest=sha256:5138c94b811fb050ae0267dd62b9ece8e778e09d87b63c3209808e9274e1c142

Observation 5b552dea-6b52-4284-a90d-f2bb8f8158d8 · outbound

This paper cites Thin-plate spline motion model for image animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Thin-plate spline motion model for image animation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.300827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.579388Z digest=sha256:e2f868bb7ae6ecaea249209176fc4849853e7559a428c9648185376c8b6ef0c1

Observation be9364e6-806f-4294-96bd-d1152f147fb2 · outbound

This paper cites Synergizing motion and appearance: Multi-scale com- pensatory codebooks for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synergizing motion and appearance: Multi-scale com- pensatory codebooks for talking head video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.071391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.582069Z digest=sha256:f47418dca8eb624fa0786051580fe19fe1ee17c1486169d1b9483b890b707f36

Observation 8869ee41-c585-4503-9ad9-5b28eb0f11cc · outbound

This paper cites Talking face generation by adversarially disentangled audio-visual representation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking face generation by adversarially disentangled audio-visual representation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.866644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.585041Z digest=sha256:166936832d86e781263a8f716c5003ddac36380cbe23ad99593b0fb78982367c

Observation 689b50cd-62c7-4faf-8a0e-6f3f39699a20 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.588387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.588203Z digest=sha256:486b6d3d341a6caf564a4a9e0956a3fdca0923c1a5717a1910b9befbf571b82b

Observation 16ce3c0d-c8d6-4417-98a8-ce28ee3543ca · outbound

This paper cites Makelttalk: speaker-aware talking-head animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Makelttalk: speaker-aware talking-head animation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.405568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:38:20.590974Z digest=sha256:574bf35e65ed1145ab28e0e4dcc3d3129b399e9ec0169b17b1c299bdd3a740c8

Pith citing papers

No inbound Pith citation observations are available.