Pith. sign in

Paper Citation Record · LEDGER

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation

As of 22 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 0 inbound Pith citation observations for arXiv:2507.20953.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20953 v1

Coverage vector

measured 100 of 112 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:10:19.735725Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 112 outbound references displayed

  • verified exact7
  • verified fuzzy49
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9836d1c-40de-4247-93fa-1901ad46d875 · outbound

This paper cites Deep audio-visual speech recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Deep audio-visual speech recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.321025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.321025Z digest=sha256:b473780c8e6930b00b7a113df621364c8ce60241ec7a9435c027600cdb2534e8

Observation e44250eb-0e52-4ba4-9828-3f35e7705e3a · outbound

This paper cites Self-supervised learning of audio- visual objects from video.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Self-supervised learning of audio- visual objects from video

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.394341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.394341Z digest=sha256:5cac63b4a3cd7ffcd71bfd26315767f93e60af6f886ba9a0bb766d204fc5df97

Observation a26a3b85-3e13-4604-9d4e-300a40fe3a18 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation A morphable model for the synthesis of 3d faces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.435036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.435036Z digest=sha256:6874fbf85ae27535408ebb12460e3ba38823a9551ae482d9c1c30ffa663796c0

Observation 47cec664-b9c1-47b1-8ea0-2516f105dff5 · outbound

This paper cites Large scale 3d mor- phable models.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Large scale 3d mor- phable models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.457608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.457608Z digest=sha256:7041fb897154bea424b8ce4e1d4bb70f52f41819356e01f4a1efc315b43faa28

Observation aa7ec033-a8cd-4eb7-9a73-a9b8dc04ff0c · outbound

This paper cites V oice puppetry.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation V oice puppetry

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.673658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.673658Z digest=sha256:c35e8a921cbdba31aafc27df0274276d27e1786b0c06db939277ca5c3c8290b6

Observation 1247365f-f787-4d8f-ad42-c20fb3f75153 · outbound

This paper cites How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks).

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.825656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.825656Z digest=sha256:d62a394c4641042a103140b5b8045499b3fa2ae723b9d653f3c01eb749b196a6

Observation 48638b90-8a57-4852-ad8b-26736454c962 · outbound

This paper cites JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.966755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:18.912385Z digest=sha256:5003450a7f6ac8e80736689918e690a30a41cf933dc63466be922f931db31b95

Observation a2f7b28a-1a6f-4e2e-a351-f81ad5ff73ed · outbound

This paper cites TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.954691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:18.981498Z digest=sha256:5bfd78b762be947a6755eb028f6b65fc24bfde4e903e3051f5b8e6da4ad2482c

Observation 2295c8f3-91b3-4243-955b-983ef24f91a1 · outbound

This paper cites Implicit neural head synthesis via controllable lo- cal deformation fields.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Implicit neural head synthesis via controllable lo- cal deformation fields

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.010196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.010196Z digest=sha256:8cffa4c4c630667e713e9c3914022cb59338f16828a0d401cdc0915486213f60

Observation 73213bc5-db62-49a3-96f6-7311a9ae1884 · outbound

This paper cites Audio-Visual Synchronisation in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-Visual Synchronisation in the wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.013378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.013378Z digest=sha256:3671a0915f503e1e97a3bdf0a3d9d752e36ffc169e2b29b3ca541d568a842eb7

Observation 3a41c1df-a95c-4af2-b709-2002971afbc2 · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.016638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.016638Z digest=sha256:759bafe8c8e04ed2850869c84b4330376d28a8d95357bdbcfff5690e896d50f2

Observation 64367a9b-4d71-4470-8663-8f071082d597 · outbound

This paper cites Videoretalking: Audio-based lip synchronization for talking head video editing in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Videoretalking: Audio-based lip synchronization for talking head video editing in the wild

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.039371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.039371Z digest=sha256:4911cb954b2452ac56c5977df9a689b894945ffbda67e45948f63b40a6abec1b

Observation bff3e6aa-a911-44ba-a68b-1396f64f1d27 · outbound

This paper cites GPAvatar: Generalizable and Precise Head Avatar from Image(s).

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation GPAvatar: Generalizable and Precise Head Avatar from Image(s)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.117914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.117914Z digest=sha256:5c04f0abea8a5691462e723692a5e69d21c8c31dc10c53771980f63dfd7e3be6

Observation ee007793-6166-4ab5-9bfc-17b23c410ea5 · outbound

This paper cites Lip reading in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Lip reading in the wild

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.197057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.197057Z digest=sha256:e845d8b66052765dff78c0420bc0688342c9acf9b9f36e681cfaf245c29a6db7

Observation de5759f6-daf1-4d47-ba93-af9702d9fe3b · outbound

This paper cites Out of time: auto- mated lip sync in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Out of time: auto- mated lip sync in the wild

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.251086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.251086Z digest=sha256:a8e2e069498224de8aacb2b725086bc0d48ef380db5010e016915dcf1216c890

Observation d7077656-b694-4b0e-a0b6-485bdf88205b · outbound

This paper cites Perfect match: Improved cross-modal embeddings for audio-visual synchronisation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Perfect match: Improved cross-modal embeddings for audio-visual synchronisation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.326692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.326692Z digest=sha256:91b565b059e5611750a70e97634f948a9116f38a7eb39866af6721a67e995381

Observation 7e629fbd-0271-4e8c-b7ea-2228696613d8 · outbound

This paper cites Speech-driven facial animation us- ing cascaded gans for learning of motion and texture.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Speech-driven facial animation us- ing cascaded gans for learning of motion and texture

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.385049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.385049Z digest=sha256:a5d47445b6ab68dda3d89e57b8fa18c2e4845740eb6d8eabfc67621f2d9dc5fd

Observation c63e7167-c726-4d70-b92c-2039a471566e · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imagenet: A large-scale hierarchical im- age database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.459990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.459990Z digest=sha256:0deeba1dbe6497072d28b9889662275306e53f414ffaf21c7324d57db1063428

Observation 947f87c2-5421-482a-8e89-119d03e4e5b5 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Arcface: Additive angular margin loss for deep face recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.502108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.502108Z digest=sha256:b63eb23beff8eff20bef60f72fb5d2e58335c63e19f41941ca3fc200ff1431ec

Observation 8c0b5163-0e1f-4848-9af1-704a4430b9f4 · outbound

This paper cites End-to-end generation of talking faces from noisy speech.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation End-to-end generation of talking faces from noisy speech

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.508637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.508637Z digest=sha256:d56b49c8da237be22c4df94a58da593af38a48453c1089eb5092033c5ae3cc2f

Observation b4dce184-be8c-4a22-89d1-ff2bb1490e0d · outbound

This paper cites Efficient emotional adaptation for audio- driven talking-head generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Efficient emotional adaptation for audio- driven talking-head generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.511638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.511638Z digest=sha256:ab72fb271c01f0146d30e7a146d6865e41631854a3a3ffcbe419847cc9eed1c9

Observation 1b9b07f6-4d3c-4ce0-859b-12423c898c07 · outbound

This paper cites Generative adversarial nets.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Generative adversarial nets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.514379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.514379Z digest=sha256:632666a0a7fd75e732e9a8473fd0c3bbc6c27fa2d9deee6b950c99337846919f

Observation 9595d764-85bd-42be-9a1e-da5ec0d9a827 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based gener- ator.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Stylesync: High-fidelity generalized and personalized lip sync in style-based gener- ator

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.517398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.517398Z digest=sha256:2f948792a04ed3c8ccabbaa952c21a310dcb095b5dec61769ce46168fb602754

Observation 7908e6e8-4050-4c77-9977-7b8a4c4f7175 · outbound

This paper cites Ad-nerf: Audio driven neural radiance fields for talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Ad-nerf: Audio driven neural radiance fields for talking head synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.520614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.520614Z digest=sha256:8d789a675d96f4be18e2bd8547580b3e36b172c60acb33b83749c9f048fa3a93

Observation 93ac78b8-73f8-4c80-b831-ffbf666981d6 · outbound

This paper cites Audio vision: Using 9 audio-visual synchrony to locate sounds.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio vision: Using 9 audio-visual synchrony to locate sounds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.523372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.523372Z digest=sha256:b888cb8a0311b40ce760a0b8e25c94aeafcecc80866d03f4c4bb450de35ee7fc

Observation 7b289745-c745-4903-9d2d-1fbc51f4e216 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.525963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.525963Z digest=sha256:79a330813dd561f72107b18de8c65528041175fc77f5490b4926dac1cf729e03

Observation d8848ea0-459e-4dc8-b06b-cf5cf66b28ee · outbound

This paper cites Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.528609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.528609Z digest=sha256:752fcfe65d864ad5eb1c047cca9e245eb8824d609feb7070d9df95d20a0d033b

Observation c59b91eb-fd0a-40c5-abaf-ef8fbfac12ae · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Arbitrary style transfer in real-time with adaptive instance normalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.531664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.531664Z digest=sha256:5749d7abeeee9dd1cbe7460747182c4d2cac72167ec65fb96aab808a4bee09d1

Observation 04ca43ee-d2a3-4adf-b2d8-13e0320596ca · outbound

This paper cites Discohead: audio-and- video-driven talking head generation by disentangled con- trol of head pose and facial expressions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Discohead: audio-and- video-driven talking head generation by disentangled con- trol of head pose and facial expressions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.534433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.534433Z digest=sha256:f1af9fd73ba83e537938386e82f760e77d261f17e02720ea639a345aef790b34

Observation 96076c35-38b5-4dcd-b088-ec50f6f83c1c · outbound

This paper cites Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.927064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.536817Z digest=sha256:fd308b950c9262a8f227893c4ad6c1e6e3bd3a68dfc8940a44867860cf6164dc

Observation 2ac0bac6-07cc-42c9-b2b6-78c4779ac027 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal co- variate shift.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Batch normalization: Accelerating deep network training by reducing internal co- variate shift

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.539510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.539510Z digest=sha256:4bbac9f9c86ae02cf7b734991c374b022215e6736ac171f45a0ef1a14bb92f37

Observation bf1738ef-787b-4d2b-9561-0d3ae75af01b · outbound

This paper cites You said that?: Synthesising talking faces from audio.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation You said that?: Synthesising talking faces from audio

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.542094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.542094Z digest=sha256:22d5ac10782766d512a5def758efb1c3afa8637d212f3300652ab590c211412f

Observation f0bb7da1-c2ff-4dd5-afa3-d1d15ebcb4ea · outbound

This paper cites Audio-driven emotional video portraits.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven emotional video portraits

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.544758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.544758Z digest=sha256:323341a701ee06956d45b2d4d0faea1a9e918bb17a34a981c701ea7e67300345

Observation 54b7fb74-5b2d-4f1e-bdc7-1389719fc14a · outbound

This paper cites Eamm: One-shot emotional talking face via audio-based emotion-aware motion model.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Eamm: One-shot emotional talking face via audio-based emotion-aware motion model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.547228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.547228Z digest=sha256:36cfaacc8de33221ea89af334e98be23d905dccb88a0dfdcf8358642f045b5f1

Observation 50611682-95c7-4edc-a162-fca7cf7d985a · outbound

This paper cites Audio-driven facial an- imation with deep learning: A survey.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven facial an- imation with deep learning: A survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.549789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.549789Z digest=sha256:439403ceb1fef2e97ce652b41f935e3db4789d475520609a42e3ce340da9edc7

Observation 74a4bae3-77ca-42db-85e4-5f3bc5a3a944 · outbound

This paper cites Percep- tual losses for real-time style transfer and super-resolution.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Percep- tual losses for real-time style transfer and super-resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.552282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.552282Z digest=sha256:25a5e32ef4c3bd20915a4695024e1c9089860fd18c444813b24f2feac78b66d9

Observation 603d623b-998b-49fd-91cd-1f36294bb2f4 · outbound

This paper cites VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.916020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.554988Z digest=sha256:2d6d136bbb51971816b3b5452e7b03c5e7b6dcb229e63ca9714f2ff0d12f6490

Observation 63d3afe6-400e-415d-9453-343bf0cd2f9c · outbound

This paper cites NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.904549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.557924Z digest=sha256:186a6068a2db00b667161b9a25fe567df2c9dd7c0fa868c715449ffccab33a9e

Observation 7d5ff702-a208-41f0-982e-a75fda6385cf · outbound

This paper cites End-to-end lip synchronisation based on pattern classification.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation End-to-end lip synchronisation based on pattern classification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.561409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.561409Z digest=sha256:85751c114391f9e74b17a2db30076505fb0dd84ffeb01e9084840ba90892ef9e

Observation 4fda48fa-ff56-435a-8872-805a678b6d6b · outbound

This paper cites Towards automatic face-to-face translation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Towards automatic face-to-face translation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.461263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.564369Z digest=sha256:00b6dfbc3495a30cd858496f68cd9ce871e56fce3552acf43a6c01e978e00fb5

Observation 13883951-8b0d-44a1-9327-244bf9e0391e · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imagenet classification with deep convolutional neural net- works

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.453293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.566904Z digest=sha256:00333db6f4cce63209acb93e7c4c2334320f8e6269f79d034b0263062a258db7

Observation 1d741386-e3a3-445c-960f-b667e71f8da9 · outbound

This paper cites Layer normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Layer normalization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.445443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.569647Z digest=sha256:584913b73b3708fb795668d4d30d821f5aca54b1929bef571949ba1066cfe8da

Observation 707b5b6f-89ec-46c4-8d69-dea8ab395aed · outbound

This paper cites Efficient region-aware neural radiance fields for high- fidelity talking portrait synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Efficient region-aware neural radiance fields for high- fidelity talking portrait synthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.437092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.572738Z digest=sha256:1185121752d5f9bbc3f981da6a23378e0227c3b583a510c2924f26353bbb487d

Observation cadf50fa-9c70-4cc5-a319-565e6e90dc60 · outbound

This paper cites One-shot high- fidelity talking-head synthesis with deformable neural ra- diance field.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation One-shot high- fidelity talking-head synthesis with deformable neural ra- diance field

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.429530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.575782Z digest=sha256:d5f259af677f1e3291afb46df427458c2134b461b7bf7dec77f9c05f0ed0aba0

Observation b0690c38-67e4-4175-99a8-2eb20b746046 · outbound

This paper cites Expressive talking head generation with granular audio-visual control.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Expressive talking head generation with granular audio-visual control

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.420942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.578661Z digest=sha256:8899d1ef6fdfbc0bc1d53ded2e7c24e8ff9a8c343e6a7845d83eb86512550540

Observation 54f85be3-39e3-498a-930b-e85474827d81 · outbound

This paper cites Font: Flow-guided one-shot talking head generation with natural head motions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Font: Flow-guided one-shot talking head generation with natural head motions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.412736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.581524Z digest=sha256:9589a84149b62d1286485cfb0163b2006065c9f4547125e50890ff08f43f59f6

Observation 9320ef49-2fa4-4e36-a7b6-dfa1cd2b1997 · outbound

This paper cites Opt: One-shot pose- controllable talking head generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Opt: One-shot pose- controllable talking head generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.404441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.585066Z digest=sha256:6662103b1a412e6bfec34ba841172cd31c82c581bc98038b50700da6dfd8396c

Observation 34290a6a-be79-4b19-ac6f-e795ca15481a · outbound

This paper cites Semantic-aware implicit neural audio- driven video portrait generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Semantic-aware implicit neural audio- driven video portrait generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.396341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.587601Z digest=sha256:3577d586acf90229dc3a2a59ee7a96ead0dd7357215ec69af88b0cce70fda4ff

Observation 926e99f2-d6c7-4f41-81ca-136480043d3d · outbound

This paper cites Moda: Mapping-once audio-driven portrait animation with dual attentions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Moda: Mapping-once audio-driven portrait animation with dual attentions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.389007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.590321Z digest=sha256:2b9efbf0e4f3eb8d261555f8831e7b553d8fab93e68f4a896426edfbbc669937

Observation 94b78c77-e869-430a-a66f-388a4e2f3f1d · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation MediaPipe: A Framework for Building Perception Pipelines

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.592928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.592928Z digest=sha256:1e0aaf9b8fb5a1e7fb9168a2dc7ad3f096cc059899c60924139d7a2ffd44fddf

Observation 884beea5-1ed5-4ba5-8f21-4e6dd446f871 · outbound

This paper cites Cvthead: One-shot controllable head avatar with vertex-feature transformer.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Cvthead: One-shot controllable head avatar with vertex-feature transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.380403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.596067Z digest=sha256:69e3375c7a96fef8da7d3217e3bff3146442a886d8df2855187e866b63a566d5

Observation ba0a46d6-2ca8-4a44-85cf-0847073d3a51 · outbound

This paper cites Styletalk: One-shot talking head generation with controllable speak- ing styles.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styletalk: One-shot talking head generation with controllable speak- ing styles

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.372221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.598971Z digest=sha256:29c3c42e2e1abf041321cf31b72e73b3ffc84f64a91fd1598095968203865a8c

Observation 288b8791-b9ae-48dd-b171-d59d1f9503f7 · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.601406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.601406Z digest=sha256:1cf6169e5a8f040d841ce356aa12387f88a2cb17e79f8789154ee992851b6310

Observation 523113b3-b116-4e6e-9da3-443bf7f20502 · outbound

This paper cites Otavatar: One-shot talking face avatar with control- lable tri-plane rendering.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Otavatar: One-shot talking face avatar with control- lable tri-plane rendering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.363238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.604026Z digest=sha256:151d201ceddc4c28d5c93288df74be3b1237b41e7475ad3170a8727e48937abf

Observation f4e159a3-aadf-4dd3-8212-a8334a1fd1c0 · outbound

This paper cites Sidgan: High-resolution dubbed video generation via shift-invariant learning.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Sidgan: High-resolution dubbed video generation via shift-invariant learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.355365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.607458Z digest=sha256:134470c5fb63fb96d9606e28f012239dab58479914e9d9b7db17db0d6fbc61fc

Observation 97f21d94-2da1-4794-be23-23bb058d756f · outbound

This paper cites Diff2lip: Audio conditioned dif- fusion models for lip-synchronization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diff2lip: Audio conditioned dif- fusion models for lip-synchronization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.347539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.609925Z digest=sha256:758029de6ca118390e18a730d20e7a88f6cbdeaff02b756eb4ae2686df89d965

Observation e2764bf5-eb83-40be-88ca-e845cc200340 · outbound

This paper cites Rectified linear units improve restricted boltzmann machines.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Rectified linear units improve restricted boltzmann machines

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.339828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.612889Z digest=sha256:dcd28a577ffddc3260ecfddb8d26a543d38c530d0efff79103dec6d7b205d9a2

Observation dd662cb7-fed5-4e7c-a4bf-d28150096b28 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-visual scene analysis with self-supervised multisensory features

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.332060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.615652Z digest=sha256:eccc0608baa8762a6739e4cae9e40a64c298843e256c8511df4da8088b9da29b

Observation ef00998f-538b-4937-add0-0ec3cf2e446e · outbound

This paper cites in-the-wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation in-the-wild

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.324159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.619193Z digest=sha256:4112a82965b71bd0acf7c4e24a3eefc692570e4d9e8bf98de5a252fe6748405d

Observation 2f85de8b-b5c0-4c33-9255-2a73ed77ca1d · outbound

This paper cites Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.315937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.621947Z digest=sha256:9ac064d066a483075a1f1db46b6f997c2beef2447fa8d5f777b26f025b2c9440

Observation fd35efd7-792e-44ca-9eba-57fd936d6b3b · outbound

This paper cites Semantic image synthesis with spatially-adaptive normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Semantic image synthesis with spatially-adaptive normalization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.308577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.624656Z digest=sha256:a581f3112be575dd009f7475acc2c045b07de52b2f1ff172f322c9b954bd0a54

Observation ade94ba0-b672-461c-b057-199276a2257f · outbound

This paper cites Synctalk: The devil is in the synchronization for talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synctalk: The devil is in the synchronization for talking head synthesis

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.300514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.627651Z digest=sha256:9f2f0acc700f33d8dfd9a1a06eceeb3b0acee84c660a036841f1634ad6876016

Observation 9cadb96b-feec-4e4b-b72f-224f52d8bfae · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation A lip sync expert is all you need for speech to lip generation in the wild

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.292496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.630189Z digest=sha256:3a4fdb6ba235e4035f2cbe4310a8fdbfb2156312e7ac0310976210f8abf3cd13

Observation 2da4e927-122b-47b4-b155-b6c22f891c9b · outbound

This paper cites U- net: Convolutional networks for biomedical image seg- mentation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation U- net: Convolutional networks for biomedical image seg- mentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.284326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.632941Z digest=sha256:697dec8227979aa9a2a5d625c06d88c99cf5f3b0d07003e7caed1e80e458ed58

Observation ad693093-af0d-40a4-bfb5-f85717dbeadc · outbound

This paper cites Learning dynamic facial radiance fields for few-shot talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Learning dynamic facial radiance fields for few-shot talking head synthesis

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.275948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.635850Z digest=sha256:482b76eb1408fa6539935b50ed9afdbb72c5f07a5ad70d100ef9404de99bfbfa

Observation a28a1ad9-22f4-44ef-8aae-b48360df52ce · outbound

This paper cites Difftalk: Crafting dif- fusion models for generalized audio-driven portraits anima- tion.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Difftalk: Crafting dif- fusion models for generalized audio-driven portraits anima- tion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.268356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.639035Z digest=sha256:675055c3281a0c25b47b9ca965287e42ee806c2a299362429a53ccde46c822d6

Observation 31a4aa87-2b76-4f62-ae70-d3c3653e449b · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.260481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.642120Z digest=sha256:2c25f694017a6c42b2e74d65c784c1a73ba35b080d530d8f43c45f617f1fe22a

Observation fcc24d0e-7f31-4adc-8fb3-e90e3d0b7940 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.645004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.645004Z digest=sha256:99008d85cfd0e608be8d7bb659ed0c591de1b4fa46f601a8a16b5cbaa9696cef

Observation 4f7efac3-601e-40f6-be5a-1d29a0daefba · outbound

This paper cites Facesync: A linear operator for measuring synchronization of video facial im- ages and audio tracks.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Facesync: A linear operator for measuring synchronization of video facial im- ages and audio tracks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.252056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.647521Z digest=sha256:ada8503bbeae9e3257b56e393eee19793dcaa83defe537332a592e7873ca5bb7

Observation 56d1ce50-6a9a-4143-9da0-33072d0e0c4e · outbound

This paper cites Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.244244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.650253Z digest=sha256:93bb9b3c5a40cecd8a4d3dec05378c8254196898e03d36603e32e141c1f7a99f

Observation 13db41f6-5ae3-4c5c-986f-d5b61442860f · outbound

This paper cites Everybody’s talkin’: Let me talk as you want.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Everybody’s talkin’: Let me talk as you want

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.236266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.653102Z digest=sha256:099e7f870ced85b0b6e27ecf48047ef8f94df2026be3b6f8cec16f114391a188

Observation 3a95139c-8a27-4893-82b8-0682e13e607a · outbound

This paper cites Talking Face Generation by Conditional Recurrent Adversarial Network.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Talking Face Generation by Conditional Recurrent Adversarial Network

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.869488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.656203Z digest=sha256:04f861585c1ee9d1ef2ce8382244097876c3deca7dcfa4f8ac1a26930f5f95dd

Observation 9c298006-4cc7-430b-9dd0-4d3cf4ff1767 · outbound

This paper cites Diffused 11 heads: Diffusion models beat gans on talking-face gen- eration.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diffused 11 heads: Diffusion models beat gans on talking-face gen- eration

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.227469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.659807Z digest=sha256:46b4b7fca2f1d0874e67d4319291ad354241487b0502c29652ba9e4111a9b318

Observation fa57b284-95aa-4720-8122-48f2245aff9c · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.662786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.662786Z digest=sha256:8a676850224230517d4e6743b3386ef81f904a3c9f2aed2cfa973ba963f50795

Observation 0af2d784-07f0-4183-afb5-af0fa7269a19 · outbound

This paper cites Masked lip-sync predic- tion by audio-visual contextual exploitation in transform- ers.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Masked lip-sync predic- tion by audio-visual contextual exploitation in transform- ers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.218728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.666211Z digest=sha256:41676f9671a536c6f1d1a8644580f516af0b6aaeae316fddd4729639dd1fdd5d

Observation 71220db6-5298-4d11-b709-face2ee95da0 · outbound

This paper cites Synthesizing obama: learning lip sync from audio.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synthesizing obama: learning lip sync from audio

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.210090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.668698Z digest=sha256:786722a30cbe59997e022de4a5e40d44ccc98eb6ac48fe38bc9cca5a5c0c7e55

Observation 89baec36-0a0d-465a-9ae3-b5a4fbc037db · outbound

This paper cites Rethinking the inception architecture for computer vision.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Rethinking the inception architecture for computer vision

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.200831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.671548Z digest=sha256:6297aa4ad9667faa77970c2fbd01efcdafab74fab198781faadf00339e355bf0

Observation 43a03a3b-b351-4e10-a744-7273bc44ee51 · outbound

This paper cites Emmn: Emotional mo- tion memory network for audio-driven emotional talking face generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Emmn: Emotional mo- tion memory network for audio-driven emotional talking face generation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.192518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.674028Z digest=sha256:86876dfce6c3bdf803f83d43bf59189f0f405e81f3f845a722c4ab02ab2c740a

Observation 77cbb8cc-050a-41ed-9803-552063712921 · outbound

This paper cites Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.676863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.676863Z digest=sha256:cc0e3eeb0a496ae52a2983dd57c738f693fa5086294a695a810893e368a74574

Observation ce345719-f4e4-4368-ab1e-eaafe44a5c34 · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Neural voice puppetry: Audio-driven facial reenactment

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.184100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.679847Z digest=sha256:64a51a1a215333539cc034a735a83c5bf6f196e5d41d394ba01e8c033baa4a3b

Observation ffec64f7-ca81-4640-bf4b-13c69139ce6f · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.683024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.683024Z digest=sha256:f33850065dc4c412c30a5d143c62f11c5a76cdf810df17607645eb9dd4881551

Observation 5390878b-824a-4c0a-9588-953d1edbec11 · outbound

This paper cites Attention is all you need.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Attention is all you need

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.686259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.686259Z digest=sha256:ba368bcdde73f51d64e6cecb9e839bd434edec995e7b735466026809144264b8

Observation 2c1cdc65-6158-44dc-8663-cc88fc4c5d81 · outbound

This paper cites Realistic speech-driven facial animation with gans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Realistic speech-driven facial animation with gans

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.170908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.688989Z digest=sha256:c71b2eed7e27985935fd761dac32c1df884b55db364b93cecc1131baeedd72fd

Observation 67bbe2e9-9fd3-4856-b015-21f45598c9ab · outbound

This paper cites Progressive disentangled representa- tion learning for fine-grained controllable talking head syn- thesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Progressive disentangled representa- tion learning for fine-grained controllable talking head syn- thesis

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.162335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.691728Z digest=sha256:31eb770aebedd16355303eeccfb1cdad28c08ad2a2fa5ea6c09457f46f96a41e

Observation e85864e7-bb3a-437a-8389-fef662dabc56 · outbound

This paper cites Seeing what you said: Talking face gen- eration guided by a lip reading expert.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Seeing what you said: Talking face gen- eration guided by a lip reading expert

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.154256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.694660Z digest=sha256:0243576665d055107ae80c6bccf53c0bb302f17032f209df102f3a3f6f490eb8

Observation 65ea565c-68cf-404b-a3ef-15deae8292cb · outbound

This paper cites Lipformer: High- fidelity and generalizable talking face generation with a pre- learned facial codebook.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Lipformer: High- fidelity and generalizable talking face generation with a pre- learned facial codebook

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.146253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.697320Z digest=sha256:bde7b52fd116d8ae28ccb0aa10fc33ec6f3c79b79250b586a37d6692afdbbabd

Observation 8bdd2035-f607-4f0a-a3be-173ac2636cab · outbound

This paper cites Styletalk++: A unified framework for controlling the speaking styles of talking heads.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styletalk++: A unified framework for controlling the speaking styles of talking heads

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.138768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.699987Z digest=sha256:f2b54c7d85725747f89656b6d598d1d92beb232788058132a216f779d073e0dc

Observation 0a3835a7-da50-4a24-8ade-78355e5da511 · outbound

This paper cites High-resolution image synthesis and semantic manipulation with conditional gans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation High-resolution image synthesis and semantic manipulation with conditional gans

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.131528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.702818Z digest=sha256:f6f3eb237e03fd074640149fb575b35e80ba7dfc9a8a4ffeaf1b0d11ebe0f469

Observation 4aba5721-85e4-468a-b946-2292b91a0e6b · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Image quality assessment: from error visibility to structural similarity

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.123405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.706016Z digest=sha256:0927cb29cd823e244cb0f3e80ceda075f652be6c485c9d2869ff395f34d3faa7

Observation 5c688292-ed0f-45b3-8365-a39e22ecab81 · outbound

This paper cites Imitating arbitrary talking style for realistic audio-driven talking face synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imitating arbitrary talking style for realistic audio-driven talking face synthesis

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.115487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.708385Z digest=sha256:d6aed2207060aef97aa6a1dabf8a153e97fc6c1e944605c021da01c9315bca8a

Observation e2f6bcc4-c996-4de0-894e-64a8dfd382c9 · outbound

This paper cites Ganhead: Towards generative animatable neural head avatars.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Ganhead: Towards generative animatable neural head avatars

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.107368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.711283Z digest=sha256:eaeec4cb524f07ca61c21d890a2c01f42b0e97d7f276b4d0017faa4d78e26f04

Observation 29b07864-bf89-421b-b1bb-15f1d78bcb0d · outbound

This paper cites High-fidelity generalized emotional talking face generation with multi-modal emotion space learning.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation High-fidelity generalized emotional talking face generation with multi-modal emotion space learning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.099524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.713845Z digest=sha256:9ad1cf28bc8e80e39332eda033543c94a9a16adeabcdb06d5a4ef22b0d65a1c6

Observation ad5563d8-07ce-40b8-a753-308cf54b5653 · outbound

This paper cites Audio-visual speech representation expert for enhanced talking face video generation and evaluation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-visual speech representation expert for enhanced talking face video generation and evaluation

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.092115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.716184Z digest=sha256:62cf195230c673b4bde4ecef6d88c83b0a03c2bcda42d580832191f21c3b7d77

Observation 2342329e-7e22-4ced-8090-c566a2e7bfde · outbound

This paper cites Audio-driven Talking Face Generation with Stabilized Synchronization Loss.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven Talking Face Generation with Stabilized Synchronization Loss

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.832891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.718719Z digest=sha256:6b306bdee9cc51c4542f8b612724e410ba21e4740cce69fa12cbe2269c9fefbf

Observation 354d3f85-f09d-4b74-8eef-cf5b8850e70e · outbound

This paper cites DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.721450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.721450Z digest=sha256:dc3296846bff5e735491b6e67f804740b2bd99a42d209e80023f2b80505f116d

Observation 47656887-e25c-4bb2-be3e-a6f50dc87109 · outbound

This paper cites GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.724040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.724040Z digest=sha256:713ee624939fb5528c65d3872fa9fa40af559ac0984bd17472123c683fc31d5d

Observation 06b9f2ff-c777-4e0c-b011-a9a0bec6255e · outbound

This paper cites Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.726995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.726995Z digest=sha256:79db605b974ab7ceab77ff5ee2fd40f65a39b628837c3bc1c14b88bd77f9468e

Observation d27549de-85a5-4933-be4e-b151ba5d8dcd · outbound

This paper cites Quantitative association of vocal-tract and facial behavior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Quantitative association of vocal-tract and facial behavior

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.083847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.729858Z digest=sha256:17b3b06f670c4ded49e6549ee64b1246f5a8637234f99d91215866ded7c60d02

Observation 16907925-7056-447c-990a-e7fc25da1dee · outbound

This paper cites Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.076042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.732934Z digest=sha256:3846483b45eaf27ffce210d06c9ce7ddb2a8d91a1833a81e47567b6a3d929c5e

Observation cdd47e1d-af3e-44e5-bc7b-6df6cbaf5e97 · outbound

This paper cites Multimodal image synthesis and editing: A survey and taxonomy.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Multimodal image synthesis and editing: A survey and taxonomy

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.068461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:10:19.735725Z digest=sha256:842113fb5ab5c59e274ecc4a152dd09b14b13ac8c1e955472036645647e868de

Pith citing papers

No inbound Pith citation observations are available.