Pith. sign in

Paper Citation Record · LEDGER

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation

As of 7 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 0 inbound Pith citation observations for arXiv:2507.20953.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20953 v1

Coverage vector

measured 100 of 112 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:10:19.735725Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 112 outbound references displayed

  • verified exact7
  • verified fuzzy49
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9836d1c-40de-4247-93fa-1901ad46d875 · outbound

This paper cites Deep audio-visual speech recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Deep audio-visual speech recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.321025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.321025Z digest=sha256:629eb2dc86317733d4d431b23d01ab7631c9c2d898624c2f643ed167b0e3d235

Observation e44250eb-0e52-4ba4-9828-3f35e7705e3a · outbound

This paper cites Self-supervised learning of audio- visual objects from video.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Self-supervised learning of audio- visual objects from video

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.394341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.394341Z digest=sha256:994a16bb0f0c7b7762e07cd181ff6a0e76fba4b9f247669018087f8e56986992

Observation a26a3b85-3e13-4604-9d4e-300a40fe3a18 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation A morphable model for the synthesis of 3d faces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.435036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.435036Z digest=sha256:21813675d46a107056ca61dd90b9bcb6910b83eaf05850cd01c5350a25eb97f0

Observation 47cec664-b9c1-47b1-8ea0-2516f105dff5 · outbound

This paper cites Large scale 3d mor- phable models.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Large scale 3d mor- phable models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.457608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.457608Z digest=sha256:57149155205fb6a629fb036d6e82bafa8457e307cb5b53aadec7b074c62da70b

Observation aa7ec033-a8cd-4eb7-9a73-a9b8dc04ff0c · outbound

This paper cites V oice puppetry.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation V oice puppetry

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.673658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.673658Z digest=sha256:9ede0c1bc7fea4c1b6369cce5901a769e318e812d8bab02cca1c2c0d9adbdf51

Observation 1247365f-f787-4d8f-ad42-c20fb3f75153 · outbound

This paper cites How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks).

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.825656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.825656Z digest=sha256:a69e1eaf7921ef4fbe1cf1aced5e8a6a8325a2fe061c3234b3bd5bfb75f31052

Observation 48638b90-8a57-4852-ad8b-26736454c962 · outbound

This paper cites JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.966755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:18.912385Z digest=sha256:a0b010dd36a1ebe5cac3f71df14b5e11bfdf84c835a87094e726a163a6362a52

Observation a2f7b28a-1a6f-4e2e-a351-f81ad5ff73ed · outbound

This paper cites TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.954691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:18.981498Z digest=sha256:832bcb72a06748f9e88feb23d62b43f598fc6bc1d856bba800c40927c15f8ab9

Observation 2295c8f3-91b3-4243-955b-983ef24f91a1 · outbound

This paper cites Implicit neural head synthesis via controllable lo- cal deformation fields.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Implicit neural head synthesis via controllable lo- cal deformation fields

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.010196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.010196Z digest=sha256:0035889e362d7a95c4830ec02f8a1aac813757e8ab078a786922851950c1197f

Observation 73213bc5-db62-49a3-96f6-7311a9ae1884 · outbound

This paper cites Audio-Visual Synchronisation in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-Visual Synchronisation in the wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.013378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.013378Z digest=sha256:d9b3e4e6583b7304bf48a4842e1636e6f7357e7eecb9580d7086113257b5ddf5

Observation 3a41c1df-a95c-4af2-b709-2002971afbc2 · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.016638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.016638Z digest=sha256:a56b8fc466ff1091dc8cd6320d8371cd2f16bf3e8cf0e38edb5b5b9e871c62ae

Observation 64367a9b-4d71-4470-8663-8f071082d597 · outbound

This paper cites Videoretalking: Audio-based lip synchronization for talking head video editing in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Videoretalking: Audio-based lip synchronization for talking head video editing in the wild

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.039371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.039371Z digest=sha256:3543c66a283b769f35c5b414c477e73d009eeea93be42b28cc2a072dee65a6b3

Observation bff3e6aa-a911-44ba-a68b-1396f64f1d27 · outbound

This paper cites GPAvatar: Generalizable and Precise Head Avatar from Image(s).

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation GPAvatar: Generalizable and Precise Head Avatar from Image(s)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.117914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.117914Z digest=sha256:6a3bd4ed32d4135eb4eaf8bb22cb1ad81f257a02b2ebe876129a434be30dbe71

Observation ee007793-6166-4ab5-9bfc-17b23c410ea5 · outbound

This paper cites Lip reading in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Lip reading in the wild

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.197057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.197057Z digest=sha256:8ffa4653044bc2ba05c079a825555130c49a8f84f9982616ca29c06d80bb7db3

Observation de5759f6-daf1-4d47-ba93-af9702d9fe3b · outbound

This paper cites Out of time: auto- mated lip sync in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Out of time: auto- mated lip sync in the wild

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.251086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.251086Z digest=sha256:a651a34a1fbb0f4bec1f085cedc5a8bce0f31eb6b1848abc4424a1d2f4b9d4ba

Observation d7077656-b694-4b0e-a0b6-485bdf88205b · outbound

This paper cites Perfect match: Improved cross-modal embeddings for audio-visual synchronisation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Perfect match: Improved cross-modal embeddings for audio-visual synchronisation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.326692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.326692Z digest=sha256:da74d567e33a93a25f4159afaa30364ef35bf23ce6df97e1c039ba0a0e9b4cd4

Observation 7e629fbd-0271-4e8c-b7ea-2228696613d8 · outbound

This paper cites Speech-driven facial animation us- ing cascaded gans for learning of motion and texture.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Speech-driven facial animation us- ing cascaded gans for learning of motion and texture

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.385049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.385049Z digest=sha256:2dafac5c1cded5586e28db07a74b1ec7387a408f0676864125d21a3ac0b4068c

Observation c63e7167-c726-4d70-b92c-2039a471566e · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imagenet: A large-scale hierarchical im- age database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.459990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.459990Z digest=sha256:2b2f95b86e8f73557a18fd7068c578683f7729bbe3e27543b62d0541092cb33d

Observation 947f87c2-5421-482a-8e89-119d03e4e5b5 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Arcface: Additive angular margin loss for deep face recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.502108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.502108Z digest=sha256:0a25b9b2011005b3ecdd04373a44aa2054d1b9d88eb50483007a6c3d2e8c6fd2

Observation 8c0b5163-0e1f-4848-9af1-704a4430b9f4 · outbound

This paper cites End-to-end generation of talking faces from noisy speech.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation End-to-end generation of talking faces from noisy speech

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.508637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.508637Z digest=sha256:bd111534c4244ec39188876a56b28d56ef1bf085d2418d383367c7ea33675190

Observation b4dce184-be8c-4a22-89d1-ff2bb1490e0d · outbound

This paper cites Efficient emotional adaptation for audio- driven talking-head generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Efficient emotional adaptation for audio- driven talking-head generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.511638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.511638Z digest=sha256:2ed94c2af9370fbdf64065e91678b1dcd9d01fba64743b92ae16d059a70f83ee

Observation 1b9b07f6-4d3c-4ce0-859b-12423c898c07 · outbound

This paper cites Generative adversarial nets.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Generative adversarial nets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.514379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.514379Z digest=sha256:d66480c7cfbf6bce8555a1c90784ac4a6b7527ee395cdbaab29f306076fe1dbf

Observation 9595d764-85bd-42be-9a1e-da5ec0d9a827 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based gener- ator.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Stylesync: High-fidelity generalized and personalized lip sync in style-based gener- ator

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.517398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.517398Z digest=sha256:c15a59af25b34dd90b65d143c02d9622def20285329fe0253f285b2f44d252a5

Observation 7908e6e8-4050-4c77-9977-7b8a4c4f7175 · outbound

This paper cites Ad-nerf: Audio driven neural radiance fields for talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Ad-nerf: Audio driven neural radiance fields for talking head synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.520614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.520614Z digest=sha256:5743f140d33a07808adfb8978d40203fb4fd3ddd8bff0e22c93c44ecd185c06a

Observation 93ac78b8-73f8-4c80-b831-ffbf666981d6 · outbound

This paper cites Audio vision: Using 9 audio-visual synchrony to locate sounds.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio vision: Using 9 audio-visual synchrony to locate sounds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.523372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.523372Z digest=sha256:4c6c1e922a9d15a5a898ab9a3b51ddae7347ad1a0cf81ebb2444216c412d698e

Observation 7b289745-c745-4903-9d2d-1fbc51f4e216 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.525963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.525963Z digest=sha256:aa459f0e6caa0df18bc89ef879df183e67389a357a664da71c3ae1472f3c857f

Observation d8848ea0-459e-4dc8-b06b-cf5cf66b28ee · outbound

This paper cites Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.528609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.528609Z digest=sha256:796542959111abb8d361a00b2de337d73fac349f9cf359640d9a5adf196ca69d

Observation c59b91eb-fd0a-40c5-abaf-ef8fbfac12ae · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Arbitrary style transfer in real-time with adaptive instance normalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.531664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.531664Z digest=sha256:62f0f48e3daa6e6e1ea2665eabf99fe6bc1c618c4d2288cb510be51c5eed7f51

Observation 04ca43ee-d2a3-4adf-b2d8-13e0320596ca · outbound

This paper cites Discohead: audio-and- video-driven talking head generation by disentangled con- trol of head pose and facial expressions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Discohead: audio-and- video-driven talking head generation by disentangled con- trol of head pose and facial expressions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.534433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.534433Z digest=sha256:8443820ff4929d767c89f51fc5ba27f50f3d075bede81d5d5537176ac615b7dd

Observation 96076c35-38b5-4dcd-b088-ec50f6f83c1c · outbound

This paper cites Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.927064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.536817Z digest=sha256:74e9d76e7bb13392eb725371beff9dad6c034d55902abacd97279464095a34f2

Observation 2ac0bac6-07cc-42c9-b2b6-78c4779ac027 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal co- variate shift.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Batch normalization: Accelerating deep network training by reducing internal co- variate shift

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.539510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.539510Z digest=sha256:aee298bad1d6d707843ba60aaafcdf341f15dc2f8357bc657217ca37542bd525

Observation bf1738ef-787b-4d2b-9561-0d3ae75af01b · outbound

This paper cites You said that?: Synthesising talking faces from audio.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation You said that?: Synthesising talking faces from audio

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.542094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.542094Z digest=sha256:98efd6adf5135f737e5ef8e1ab8e1b1852842a8c69a6674576db4389dfa4611c

Observation f0bb7da1-c2ff-4dd5-afa3-d1d15ebcb4ea · outbound

This paper cites Audio-driven emotional video portraits.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven emotional video portraits

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.544758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.544758Z digest=sha256:b0fd23095e58df2730c1902de9581163781491a63ecf1f9b4e1b7a28384f4455

Observation 54b7fb74-5b2d-4f1e-bdc7-1389719fc14a · outbound

This paper cites Eamm: One-shot emotional talking face via audio-based emotion-aware motion model.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Eamm: One-shot emotional talking face via audio-based emotion-aware motion model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.547228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.547228Z digest=sha256:132308fbac7159b0434b905b65cadb2d8eaac2b3687842055b066522b05773a8

Observation 50611682-95c7-4edc-a162-fca7cf7d985a · outbound

This paper cites Audio-driven facial an- imation with deep learning: A survey.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven facial an- imation with deep learning: A survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.549789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.549789Z digest=sha256:5b87c568fe7e4baf0f944b7f3f5a020b2310497299b27628948ab8847b4f2ada

Observation 74a4bae3-77ca-42db-85e4-5f3bc5a3a944 · outbound

This paper cites Percep- tual losses for real-time style transfer and super-resolution.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Percep- tual losses for real-time style transfer and super-resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.552282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.552282Z digest=sha256:3a8d09adb6d5cd0c949b73fb0de96d8f2805bf14d6dbc46034259bc8aa64c9b4

Observation 603d623b-998b-49fd-91cd-1f36294bb2f4 · outbound

This paper cites VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.916020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.554988Z digest=sha256:5b613078d4c8db72277eb5a461fdc2c10a12aeaf623f172c50cac6ae4876e7ce

Observation 63d3afe6-400e-415d-9453-343bf0cd2f9c · outbound

This paper cites NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.904549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.557924Z digest=sha256:97fcbae1bd84acc897c97c564b4d0f98c2a72219a3b34a5d79c057434d0bfd22

Observation 7d5ff702-a208-41f0-982e-a75fda6385cf · outbound

This paper cites End-to-end lip synchronisation based on pattern classification.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation End-to-end lip synchronisation based on pattern classification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.561409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.561409Z digest=sha256:357b3af1ad120fbcd4f9358c8de8dba02bf51567f2190b13b8f3d1b49844b71c

Observation 4fda48fa-ff56-435a-8872-805a678b6d6b · outbound

This paper cites Towards automatic face-to-face translation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Towards automatic face-to-face translation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.461263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.564369Z digest=sha256:c41efb560e166e8ecb4730560788efb682deaa3fc181063205f032f0b980efa3

Observation 13883951-8b0d-44a1-9327-244bf9e0391e · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imagenet classification with deep convolutional neural net- works

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.453293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.566904Z digest=sha256:9fa66a3ae62ed1a314ec787a9cc5c40e5bdacb6ad88b7c6e1d1a744b309aa595

Observation 1d741386-e3a3-445c-960f-b667e71f8da9 · outbound

This paper cites Layer normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Layer normalization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.445443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.569647Z digest=sha256:9df90ccd90d3cf8fbb75d17259b9eabb69653771898d98eb16c0aeef0654789a

Observation 707b5b6f-89ec-46c4-8d69-dea8ab395aed · outbound

This paper cites Efficient region-aware neural radiance fields for high- fidelity talking portrait synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Efficient region-aware neural radiance fields for high- fidelity talking portrait synthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.437092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.572738Z digest=sha256:f1abe2fce37343fffe1c239c981f260944cf746e565056229a5e43d0256a7ea2

Observation cadf50fa-9c70-4cc5-a319-565e6e90dc60 · outbound

This paper cites One-shot high- fidelity talking-head synthesis with deformable neural ra- diance field.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation One-shot high- fidelity talking-head synthesis with deformable neural ra- diance field

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.429530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.575782Z digest=sha256:a9d69948d3079ef092f769d54b3a728efc0787351f5216a5aea429de71e5c76b

Observation b0690c38-67e4-4175-99a8-2eb20b746046 · outbound

This paper cites Expressive talking head generation with granular audio-visual control.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Expressive talking head generation with granular audio-visual control

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.420942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.578661Z digest=sha256:5dd54305af011002ebb89bbc0b9651e86184f019aa8913a94948ac84ccee85bf

Observation 54f85be3-39e3-498a-930b-e85474827d81 · outbound

This paper cites Font: Flow-guided one-shot talking head generation with natural head motions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Font: Flow-guided one-shot talking head generation with natural head motions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.412736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.581524Z digest=sha256:59dc5b205faf4de997f363297832fabc1ad4a10f7dc07dda05d96413f12f40d6

Observation 9320ef49-2fa4-4e36-a7b6-dfa1cd2b1997 · outbound

This paper cites Opt: One-shot pose- controllable talking head generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Opt: One-shot pose- controllable talking head generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.404441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.585066Z digest=sha256:0405b352b7395373bb49daf506223b861c353c97459779294c74fcd506a88afa

Observation 34290a6a-be79-4b19-ac6f-e795ca15481a · outbound

This paper cites Semantic-aware implicit neural audio- driven video portrait generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Semantic-aware implicit neural audio- driven video portrait generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.396341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.587601Z digest=sha256:cb796a8b4e33821c6205e13216e2bc547ac6ce88f1a72f46ae7c4766b28f31cd

Observation 926e99f2-d6c7-4f41-81ca-136480043d3d · outbound

This paper cites Moda: Mapping-once audio-driven portrait animation with dual attentions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Moda: Mapping-once audio-driven portrait animation with dual attentions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.389007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.590321Z digest=sha256:f69beceda8c9846b456f22125889749ef76976e550ab56cee5e986050d574725

Observation 94b78c77-e869-430a-a66f-388a4e2f3f1d · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation MediaPipe: A Framework for Building Perception Pipelines

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.592928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.592928Z digest=sha256:0ed89c7bacf6178b053e7375e67cd15b18cb3ac8de03c787e876ba072ed65f11

Observation 884beea5-1ed5-4ba5-8f21-4e6dd446f871 · outbound

This paper cites Cvthead: One-shot controllable head avatar with vertex-feature transformer.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Cvthead: One-shot controllable head avatar with vertex-feature transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.380403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.596067Z digest=sha256:3170f6acc045f84a4bb40143fb01f2923f1836f4503d6b5da92d74ee04c5308a

Observation ba0a46d6-2ca8-4a44-85cf-0847073d3a51 · outbound

This paper cites Styletalk: One-shot talking head generation with controllable speak- ing styles.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styletalk: One-shot talking head generation with controllable speak- ing styles

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.372221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.598971Z digest=sha256:a0e2fbce44c49281549521495519b35a5cb99aad15cd1fb5715bfb97143ac5fd

Observation 288b8791-b9ae-48dd-b171-d59d1f9503f7 · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.601406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.601406Z digest=sha256:a67107040282d78b750c2d43a31aaf0741ce2f0032d5c1adf85e78b2aba7e52a

Observation 523113b3-b116-4e6e-9da3-443bf7f20502 · outbound

This paper cites Otavatar: One-shot talking face avatar with control- lable tri-plane rendering.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Otavatar: One-shot talking face avatar with control- lable tri-plane rendering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.363238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.604026Z digest=sha256:c0ea28c220ea9134714415de1baad29d0dbb604230f234ef5d9bcbceb98d4837

Observation f4e159a3-aadf-4dd3-8212-a8334a1fd1c0 · outbound

This paper cites Sidgan: High-resolution dubbed video generation via shift-invariant learning.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Sidgan: High-resolution dubbed video generation via shift-invariant learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.355365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.607458Z digest=sha256:b87ca76dcd49c1794127ca54e20e273a01597e53ba56586675f4af540b1d6d7c

Observation 97f21d94-2da1-4794-be23-23bb058d756f · outbound

This paper cites Diff2lip: Audio conditioned dif- fusion models for lip-synchronization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diff2lip: Audio conditioned dif- fusion models for lip-synchronization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.347539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.609925Z digest=sha256:311eff9b709216b5b49ca76522a5acf1f14a334450f0a82aef71c288ab38dbba

Observation e2764bf5-eb83-40be-88ca-e845cc200340 · outbound

This paper cites Rectified linear units improve restricted boltzmann machines.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Rectified linear units improve restricted boltzmann machines

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.339828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.612889Z digest=sha256:1df98f65ae08c25ab4a2d118dc874222ec958dffd43f3f2fbce3167f8fea7fcd

Observation dd662cb7-fed5-4e7c-a4bf-d28150096b28 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-visual scene analysis with self-supervised multisensory features

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.332060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.615652Z digest=sha256:a15e059d5dc6885e154b3e1570298857307255d9ccddfead2d9531119e9fcd97

Observation ef00998f-538b-4937-add0-0ec3cf2e446e · outbound

This paper cites in-the-wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation in-the-wild

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.324159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.619193Z digest=sha256:8c88a212fa59949d4999f291b1ba3ded1fe175d8163edf44aa90c0b1fe37a909

Observation 2f85de8b-b5c0-4c33-9255-2a73ed77ca1d · outbound

This paper cites Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.315937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.621947Z digest=sha256:f9c5cb00b1055bfda89edb938f12c9e8a00da51c3f176e7cb0146faac983916b

Observation fd35efd7-792e-44ca-9eba-57fd936d6b3b · outbound

This paper cites Semantic image synthesis with spatially-adaptive normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Semantic image synthesis with spatially-adaptive normalization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.308577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.624656Z digest=sha256:8e3b26319923362e8503d5725856bd9b97d67ea10d747df5f2273391bfbe0bba

Observation ade94ba0-b672-461c-b057-199276a2257f · outbound

This paper cites Synctalk: The devil is in the synchronization for talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synctalk: The devil is in the synchronization for talking head synthesis

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.300514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.627651Z digest=sha256:d8e07a4c4a0e1c671d54de21902fb91080bbf960c6142eb4c36f4ab276194043

Observation 9cadb96b-feec-4e4b-b72f-224f52d8bfae · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation A lip sync expert is all you need for speech to lip generation in the wild

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.292496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.630189Z digest=sha256:971a91c9294832b18d78391ae0e2a926465a03c790040fea94c40b4602ce9a0c

Observation 2da4e927-122b-47b4-b155-b6c22f891c9b · outbound

This paper cites U- net: Convolutional networks for biomedical image seg- mentation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation U- net: Convolutional networks for biomedical image seg- mentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.284326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.632941Z digest=sha256:fee727bd23aabac22b6a06f3a85d7cb353537d3a3b5c106e3f41663aed86ec34

Observation ad693093-af0d-40a4-bfb5-f85717dbeadc · outbound

This paper cites Learning dynamic facial radiance fields for few-shot talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Learning dynamic facial radiance fields for few-shot talking head synthesis

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.275948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.635850Z digest=sha256:0cfa5eb8afbd177391e8fc03663bbe57f52184bf6609ce86c32c6aefcea775ca

Observation a28a1ad9-22f4-44ef-8aae-b48360df52ce · outbound

This paper cites Difftalk: Crafting dif- fusion models for generalized audio-driven portraits anima- tion.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Difftalk: Crafting dif- fusion models for generalized audio-driven portraits anima- tion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.268356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.639035Z digest=sha256:02c9a43ee3223a1b111150759c1d2b18e2f2f390bee9d86518e3020c0fb54b66

Observation 31a4aa87-2b76-4f62-ae70-d3c3653e449b · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.260481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.642120Z digest=sha256:7fdd73da3241a160fce2e3c7dab3c935f8700d00da3365995dace32cb79b191b

Observation fcc24d0e-7f31-4adc-8fb3-e90e3d0b7940 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.645004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.645004Z digest=sha256:8b8aea8e018ab4a7e1168d05d8f2295ef1709085ea43a744cf694d2fc40d2d4d

Observation 4f7efac3-601e-40f6-be5a-1d29a0daefba · outbound

This paper cites Facesync: A linear operator for measuring synchronization of video facial im- ages and audio tracks.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Facesync: A linear operator for measuring synchronization of video facial im- ages and audio tracks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.252056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.647521Z digest=sha256:39a389e1efe03f6c6c49f5f843c3c03887f2e85ad53f836f6f7363cbf7aa7d3f

Observation 56d1ce50-6a9a-4143-9da0-33072d0e0c4e · outbound

This paper cites Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.244244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.650253Z digest=sha256:b48053e715ffa6675ad4b840e4ce0540aae1a6c47d99dd497f18adaaf8216060

Observation 13db41f6-5ae3-4c5c-986f-d5b61442860f · outbound

This paper cites Everybody’s talkin’: Let me talk as you want.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Everybody’s talkin’: Let me talk as you want

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.236266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.653102Z digest=sha256:175c990e8fc7849d25d8b9c723bc7f93e9928727dc7cbcd0c8e7c7c85ce3fceb

Observation 3a95139c-8a27-4893-82b8-0682e13e607a · outbound

This paper cites Talking Face Generation by Conditional Recurrent Adversarial Network.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Talking Face Generation by Conditional Recurrent Adversarial Network

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.869488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.656203Z digest=sha256:97115767c45ed4baf817b74b77e0dd4cfa50a9194160aedba30613c64364678f

Observation 9c298006-4cc7-430b-9dd0-4d3cf4ff1767 · outbound

This paper cites Diffused 11 heads: Diffusion models beat gans on talking-face gen- eration.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diffused 11 heads: Diffusion models beat gans on talking-face gen- eration

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.227469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.659807Z digest=sha256:7bffd89c7638b2d12b9ad703ee88d86922a5e25bad419c550db9ce97d2e281cd

Observation fa57b284-95aa-4720-8122-48f2245aff9c · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.662786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.662786Z digest=sha256:9f765e5248d57132ce2d534e158f0c7229014de9bf03c80ee3b93f304c94c1fa

Observation 0af2d784-07f0-4183-afb5-af0fa7269a19 · outbound

This paper cites Masked lip-sync predic- tion by audio-visual contextual exploitation in transform- ers.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Masked lip-sync predic- tion by audio-visual contextual exploitation in transform- ers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.218728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.666211Z digest=sha256:3fc0015680bc3fd7dca04e8bc2ca60e5fc7f6b16575546e88ef8b390a68261a2

Observation 71220db6-5298-4d11-b709-face2ee95da0 · outbound

This paper cites Synthesizing obama: learning lip sync from audio.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synthesizing obama: learning lip sync from audio

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.210090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.668698Z digest=sha256:d82de58623f34882f288b45f69a4b662d8e37f66ba367e0fe1bc12017e6a6783

Observation 89baec36-0a0d-465a-9ae3-b5a4fbc037db · outbound

This paper cites Rethinking the inception architecture for computer vision.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Rethinking the inception architecture for computer vision

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.200831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.671548Z digest=sha256:81b3717e734181af410391ea2f1347d6164b96f9004b12e2bca0aa10961fc86d

Observation 43a03a3b-b351-4e10-a744-7273bc44ee51 · outbound

This paper cites Emmn: Emotional mo- tion memory network for audio-driven emotional talking face generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Emmn: Emotional mo- tion memory network for audio-driven emotional talking face generation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.192518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.674028Z digest=sha256:ebf683b50dbd84153c34e3f6c635f41f8f83bffcdad2169c6f45cd4539317559

Observation 77cbb8cc-050a-41ed-9803-552063712921 · outbound

This paper cites Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.676863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.676863Z digest=sha256:5d04ffdd787ce8dc2e84b4596b4cf8cc8c6b0a63fc3ac88549dc540d1cc6378b

Observation ce345719-f4e4-4368-ab1e-eaafe44a5c34 · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Neural voice puppetry: Audio-driven facial reenactment

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.184100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.679847Z digest=sha256:7c23663ca9f5a8b2c94ba1ad589fc3ac6440063daf31141baf59682eab1c60fe

Observation ffec64f7-ca81-4640-bf4b-13c69139ce6f · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.683024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.683024Z digest=sha256:9434f343ef1842474dc02e1629e4e5a77bfbe309d3597d37efc70b028f45bc3f

Observation 5390878b-824a-4c0a-9588-953d1edbec11 · outbound

This paper cites Attention is all you need.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Attention is all you need

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.686259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.686259Z digest=sha256:168b481e407e42df4187c3ff7bedc3044dddc7547d9efbfc9b3cb594d2a85a8a

Observation 2c1cdc65-6158-44dc-8663-cc88fc4c5d81 · outbound

This paper cites Realistic speech-driven facial animation with gans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Realistic speech-driven facial animation with gans

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.170908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.688989Z digest=sha256:98334e961fb59bee59e15f32b437e48e0735ce5d3cfa5c90bb97389922212690

Observation 67bbe2e9-9fd3-4856-b015-21f45598c9ab · outbound

This paper cites Progressive disentangled representa- tion learning for fine-grained controllable talking head syn- thesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Progressive disentangled representa- tion learning for fine-grained controllable talking head syn- thesis

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.162335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.691728Z digest=sha256:f2000773ced96b7bb056a6e4633681244c00c62e5d4f2854d76a5bc531f9f3d8

Observation e85864e7-bb3a-437a-8389-fef662dabc56 · outbound

This paper cites Seeing what you said: Talking face gen- eration guided by a lip reading expert.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Seeing what you said: Talking face gen- eration guided by a lip reading expert

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.154256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.694660Z digest=sha256:daabc41f7e737e8ad3dd9c528dcfd0ee97708d792c5f62673dfb258e84289623

Observation 65ea565c-68cf-404b-a3ef-15deae8292cb · outbound

This paper cites Lipformer: High- fidelity and generalizable talking face generation with a pre- learned facial codebook.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Lipformer: High- fidelity and generalizable talking face generation with a pre- learned facial codebook

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.146253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.697320Z digest=sha256:7bfc177bc14253d9b0aa8183a6bbd141ae6dc06ecf7eca79806374ea5626612f

Observation 8bdd2035-f607-4f0a-a3be-173ac2636cab · outbound

This paper cites Styletalk++: A unified framework for controlling the speaking styles of talking heads.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styletalk++: A unified framework for controlling the speaking styles of talking heads

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.138768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.699987Z digest=sha256:2fc0ef25c98edbf5c0c232649784d93e1f46d17d1a9e03c8304d9df935b1bac6

Observation 0a3835a7-da50-4a24-8ade-78355e5da511 · outbound

This paper cites High-resolution image synthesis and semantic manipulation with conditional gans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation High-resolution image synthesis and semantic manipulation with conditional gans

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.131528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.702818Z digest=sha256:2981c036f628e119cb36632c58904bf2c18839057749a08b8941e9a38ab153da

Observation 4aba5721-85e4-468a-b946-2292b91a0e6b · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Image quality assessment: from error visibility to structural similarity

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.123405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.706016Z digest=sha256:6d2dc231af69d1aea11e2051ebaf9caae4641c7d2a109e5699dc108fb00cba47

Observation 5c688292-ed0f-45b3-8365-a39e22ecab81 · outbound

This paper cites Imitating arbitrary talking style for realistic audio-driven talking face synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imitating arbitrary talking style for realistic audio-driven talking face synthesis

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.115487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.708385Z digest=sha256:93d002ac82dfe8cec6e3b77c69783132cbbe6c5e1cbac1ddb6b327e0ccf87fca

Observation e2f6bcc4-c996-4de0-894e-64a8dfd382c9 · outbound

This paper cites Ganhead: Towards generative animatable neural head avatars.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Ganhead: Towards generative animatable neural head avatars

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.107368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.711283Z digest=sha256:9e75894d57672cc92ca153420ea86fb84b0509ddbd165bfee217e1974928d93c

Observation 29b07864-bf89-421b-b1bb-15f1d78bcb0d · outbound

This paper cites High-fidelity generalized emotional talking face generation with multi-modal emotion space learning.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation High-fidelity generalized emotional talking face generation with multi-modal emotion space learning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.099524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.713845Z digest=sha256:1d9d07f4c604c4ccda2cbd46a17c3a37c36a68dd4424dcd35c73a32d76755b08

Observation ad5563d8-07ce-40b8-a753-308cf54b5653 · outbound

This paper cites Audio-visual speech representation expert for enhanced talking face video generation and evaluation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-visual speech representation expert for enhanced talking face video generation and evaluation

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.092115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.716184Z digest=sha256:1cb90bcdd211c8ed003d234999d041e8592163cefe61262c98fbb1866f31440d

Observation 2342329e-7e22-4ced-8090-c566a2e7bfde · outbound

This paper cites Audio-driven Talking Face Generation with Stabilized Synchronization Loss.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven Talking Face Generation with Stabilized Synchronization Loss

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.832891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.718719Z digest=sha256:a78866ee988df721dde8106bcdae243a883aab3fd4a6ef49fc5832feeb76841f

Observation 354d3f85-f09d-4b74-8eef-cf5b8850e70e · outbound

This paper cites DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.721450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.721450Z digest=sha256:548a67ca4fb19f7181a16e5f7265612929bb53b63f1ded10dd7e6445eb7c4df0

Observation 47656887-e25c-4bb2-be3e-a6f50dc87109 · outbound

This paper cites GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.724040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.724040Z digest=sha256:2587f5e6d8200285c4838880fd2931342bff54416ffb3c10877ad5fb4358570b

Observation 06b9f2ff-c777-4e0c-b011-a9a0bec6255e · outbound

This paper cites Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.726995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.726995Z digest=sha256:4d9d3c61b957c55198b1db8b7cc065642b09d592681059111ea9fc2dfb9314a8

Observation d27549de-85a5-4933-be4e-b151ba5d8dcd · outbound

This paper cites Quantitative association of vocal-tract and facial behavior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Quantitative association of vocal-tract and facial behavior

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.083847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.729858Z digest=sha256:177af4ce1dd788689f00dccd25219de22277768d3cfd7287a37cbd984c3bb801

Observation 16907925-7056-447c-990a-e7fc25da1dee · outbound

This paper cites Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.076042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.732934Z digest=sha256:dc4f7cc3834f54d4d68447c189b25484712580687953187b745485c7eed78d51

Observation cdd47e1d-af3e-44e5-bc7b-6df6cbaf5e97 · outbound

This paper cites Multimodal image synthesis and editing: A survey and taxonomy.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Multimodal image synthesis and editing: A survey and taxonomy

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.068461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:10:19.735725Z digest=sha256:125b1368ffa9c6c60c81abe1d325a1a23cb7931fb1c7e17e314af606bc1f6298

Pith citing papers

No inbound Pith citation observations are available.