Pith. sign in

Paper Citation Record · LEDGER

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2505.23290.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23290 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:07.692944Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T08:44:49.087457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T08:45:19.106137Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy39
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb9c18dc-9ec7-48ec-964e-53082ce003f5 · outbound

This paper cites Facetalk: Audio-driven motion diffusion for neu- ral parametric head models.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Facetalk: Audio-driven motion diffusion for neu- ral parametric head models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:16.526076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:01.558841Z digest=sha256:b75f488da62f5de2d9ce257a565e7d4ff3475de30fc66d72b3048cea339272ca

Observation 880fe63a-bdeb-4098-ac64-810dc5327e99 · outbound

This paper cites Gesturediffuclip: Gesture diffusion model with clip latents.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Gesturediffuclip: Gesture diffusion model with clip latents

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:16.329111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:01.674615Z digest=sha256:7921702929a7399b6efb321f594bf74d7cd631e6e6c8160254627ac77162747d

Observation 7d4a71ea-d624-4e29-a067-de88585b9d1d · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:16.131668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:01.833993Z digest=sha256:7dd7f7f8433fea7fb99821d2327f66af2879635f64e2b07d4fae9e16bbedcae2

Observation 977188a8-b5d4-4ae9-bacd-4406c897afa0 · outbound

This paper cites Expressive speech-driven facial animation.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Expressive speech-driven facial animation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:15.953223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:02.000493Z digest=sha256:60dafa281412ef582c3e47d2739d62451c772d94619222f43dd30404f4ed9fb0

Observation 75aeae6e-a4ff-41c2-9523-564d9d338a9d · outbound

This paper cites Talking head generation with audio and speech related facial action units.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Talking head generation with audio and speech related facial action units

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:15.722871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:02.167821Z digest=sha256:31bf1897d4552c3da487dee335d95b887ccd6b9fb1a3dfa237a7ffa5a68059f4

Observation 34bf5d04-6b3c-4cac-89c6-f140472b8119 · outbound

This paper cites Peters, Michael J.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Peters, Michael J

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:15.545363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:02.329610Z digest=sha256:52cfea2c0d7732346acfd71b08a8015be7067f99c6729220a042d415bdeb14e5

Observation 2db9501e-b2c4-40c2-9685-533be715a0f3 · outbound

This paper cites an unresolved cited work.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:54:15.341782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:02.505572Z digest=sha256:58236ab256f3c1aeff9ef606105f67943735bba8a821e9cf3c14e936962a4f09

Observation b861fe77-0988-4cfc-84c1-3d4d68035ae6 · outbound

This paper cites Emotional speech- driven animation with content-emotion disentanglement.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Emotional speech- driven animation with content-emotion disentanglement

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:15.080161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:02.674341Z digest=sha256:9fa0516143a6ec4dcd66e0a10b376f85bc1bf1d2ec16f986bdf5180b53a8bde2

Observation 3f361a3e-6cdf-4c5b-996e-ae2be75dc88c · outbound

This paper cites BERT: pre-training of deep bidirectional trans- formers for language understanding.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation BERT: pre-training of deep bidirectional trans- formers for language understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:14.921585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:02.812228Z digest=sha256:bbbe06bb99c28183995d47566cedd678ef77cae338c647d6a636b8b270acb37d

Observation a4e84034-5fcc-4be4-88bd-f59e57c348cf · outbound

This paper cites Cross modal audio search and retrieval with joint embeddings based on text and audio.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Cross modal audio search and retrieval with joint embeddings based on text and audio

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:14.714250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:02.978900Z digest=sha256:ddbdd097606a88833566cf3b7d6a83765a13d82880100d1119c29074bb75b452

Observation c3c3ab6f-bd83-452c-8259-e903f434c39a · outbound

This paper cites Unitalker: Scaling up audio-driven 3d facial animation through A unified model.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Unitalker: Scaling up audio-driven 3d facial animation through A unified model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:14.534129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:03.177986Z digest=sha256:05e5e4460944f5d8a4f43a0281f5813f9347c752cf62f267ed1cdc335da1ff46

Observation 4275f2ce-7cbc-487b-91c7-c2bb3693ddf8 · outbound

This paper cites Faceformer: Speech-driven 3d facial anima- tion with transformers.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Faceformer: Speech-driven 3d facial anima- tion with transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:14.290319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:03.337151Z digest=sha256:45d03903f68bb38934ed3b0282a8f6cc031717f7deabf16dc587a8763cc81147

Observation f35f5ab2-2349-4993-b5cb-30f6c1d2c79e · outbound

This paper cites Joint audio-text model for expressive speech- driven 3d facial animation.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Joint audio-text model for expressive speech- driven 3d facial animation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:14.053194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:03.441886Z digest=sha256:70cd117353367aa998b901bf99654b37ddfbcf60f2b6f7044c3deabc8b0e28d3

Observation 7358179a-b5e2-466d-9ac3-41e30f57c03f · outbound

This paper cites Mimic: Speaking style disentanglement for speech-driven 3d facial animation.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Mimic: Speaking style disentanglement for speech-driven 3d facial animation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:13.867242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:03.588034Z digest=sha256:bb4cfcc0a6073ebba68f4acb65a16fff5378fed8b663be3966774fb7960a3229

Observation 65055018-5466-4ab9-9fed-4aa94ed05350 · outbound

This paper cites Deep Speech: Scaling up end-to-end speech recognition.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Deep Speech: Scaling up end-to-end speech recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:03.725530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:03.725530Z digest=sha256:45ab132d11722e23bab488c10f7d299a20ad63f7fd347c4bd69d07c372a66685

Observation 5e964826-975d-4e95-b6c5-7fe2b55c6323 · outbound

This paper cites Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:13.711516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:03.884744Z digest=sha256:33c42229c760f686518081d9f5219c881b081aa0430dc6789f3dd09ed5859cd9

Observation f512baf2-847a-4f41-9713-c95797b3b3b8 · outbound

This paper cites Phonemic similarity metrics to compare pronunciation methods.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Phonemic similarity metrics to compare pronunciation methods

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:13.534159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:04.032531Z digest=sha256:64302e599ff3ae69ad00a32c43cc1ae38ae0c1e440ffc2a42f6c476c3bd97a3b

Observation a46928a1-0165-48dc-adc8-8f7941280d8e · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:13.330358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:04.210627Z digest=sha256:7780026b1173e4803b046a51987f20880fb9a0fef0dc2a1cbc0deeaa6f1b90d3

Observation eecda33e-2527-47ba-9aae-186f8dbaabf7 · outbound

This paper cites Kingma and Jimmy Ba.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Kingma and Jimmy Ba

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:04.387362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:04.387362Z digest=sha256:39006fa80a24e46510382c8caaebeeb0baf72e4d278a349bcfd547ed70e82138

Observation 9fec146e-fe99-423b-97d8-873b698ca677 · outbound

This paper cites Mask-fpan: Semi-supervised face parsing in the wild with de-occlusion and uv gan.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Mask-fpan: Semi-supervised face parsing in the wild with de-occlusion and uv gan

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:13.166424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:04.564888Z digest=sha256:605c30c739f0589e2ddffad691ec0e6246fa78c52fde256e90efbdb27e3773b7

Observation c087cc0c-1812-41a4-8c06-d83571ade16d · outbound

This paper cites Omg: Towards open-vocabulary motion generation via mixture of controllers.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Omg: Towards open-vocabulary motion generation via mixture of controllers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:13.009303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:04.723876Z digest=sha256:7c668cf08618ad6d0e67db4fc7cbf69e906227d135d76b820af233a2eff9a22d

Observation 2732a15f-42c3-42d0-afa1-c77dc26619d7 · outbound

This paper cites an unresolved cited work.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:54:12.818499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:04.899820Z digest=sha256:6321102f8d2fafafa681365bf43bd414aea29896d41404e279db0738253c49c6

Observation b09fbcc8-af9a-42df-9bea-3a8ed2adcf6c · outbound

This paper cites Convofusion: Multi-modal conversational diffu- sion for co-speech gesture synthesis.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Convofusion: Multi-modal conversational diffu- sion for co-speech gesture synthesis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:12.349161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.069941Z digest=sha256:0ea762697e897f047d539be70669ec74538be12303ad60f9f62832b9dcdf7e34

Observation 8a7ae7cd-79d9-47c0-8b7a-f82bf9c98c67 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Librispeech: An ASR corpus based on public domain audio books

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:12.088950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.237049Z digest=sha256:3c790389db2d0e51d02d6de0a0d9df9219065b20cd3f8ab2dc3c6a7fbf93a709

Observation e573cb65-3a46-4b16-99f9-2d03cfab0138 · outbound

This paper cites Emotalk: Speech-driven emotional disentanglement for 3d face anima- tion.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Emotalk: Speech-driven emotional disentanglement for 3d face anima- tion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:11.911769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.408691Z digest=sha256:403a33fbfdac4d0b70b6b7394b98f2f871c3ea55cf1213bbc08fb4c0e21efc28

Observation 61c77f8b-5e1c-4a0b-be0f-fd65413d8565 · outbound

This paper cites an unresolved cited work.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:54:11.757569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.550901Z digest=sha256:7ad318a936b8fd6158292201b5a7b93e0cd546b7a6c8224854b8f34fab956405

Observation 8266a45f-18e7-4439-856a-1e54ffee48e6 · outbound

This paper cites Meshtalk: 3d face an- imation from speech using cross-modality disentanglement.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Meshtalk: 3d face an- imation from speech using cross-modality disentanglement

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:11.519201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.684758Z digest=sha256:9b80d124cfc495be0483f41856f0107fcbb719badee3ef6562ae9a9c5472bfe3

Observation cd2f1206-42e4-4598-9c91-74258ffc8680 · outbound

This paper cites Expressive 3d facial animation generation based on local-to-global latent diffusion.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Expressive 3d facial animation generation based on local-to-global latent diffusion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:11.347172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.744599Z digest=sha256:6e8625ce14f01f97cfb96af68107737497d9e2d05d16557bed4582b003f17ca3

Observation b98f197b-bb53-44d7-9993-02bbaadae7f7 · outbound

This paper cites Talkingstyle: Personalized speech-driven 3d facial animation with style preservation.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Talkingstyle: Personalized speech-driven 3d facial animation with style preservation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:11.131632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.855418Z digest=sha256:bb5165a90f83af9ee6a238fd315b6b688387e59b4ddee18d6e1194cb83cd3dfc

Observation 06772a52-8deb-4a45-bc51-303597147eef · outbound

This paper cites Facediffuser: Speech-driven 3d facial animation synthesis using diffusion.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Facediffuser: Speech-driven 3d facial animation synthesis using diffusion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:10.968012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.897428Z digest=sha256:2dbb281e164e94f95d7c71f8413589c7d1ea0a16feb2d80187c3a7087e02b8d6

Observation 791ea625-63bf-4195-96fd-ad4b096e68fd · outbound

This paper cites Sun and Li Deng.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Sun and Li Deng

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:10.785640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.923633Z digest=sha256:aba1c7f8cccaca15294563546e6bcc6711eef62d4ef63b9e68ca5992160f2fb9

Observation 7e96ea9f-ce27-47dd-ba78-6515fd3d71dd · outbound

This paper cites Diff- posetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Diff- posetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:10.574651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:05.955195Z digest=sha256:a990d95b9e70ec417f29195e19ff4eae9420904653fda591274e10d299eabd12

Observation df5437b6-5604-4c32-abf7-e877c91d6c41 · outbound

This paper cites Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica K.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica K

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:10.380201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:06.058357Z digest=sha256:8ef386a82902fbd18ad01ea9486e93d1b52c56fde356e43dd15bf76c385b72c8

Observation a486deaa-0431-46a5-947c-fd63a0f7ca75 · outbound

This paper cites 3DiFACE: Diffusion-based Speech-driven 3D Facial Animation and Editing.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation 3DiFACE: Diffusion-based Speech-driven 3D Facial Animation and Editing

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:06.174746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:06.174746Z digest=sha256:ed5f5aa2eed4b4866fd9f3384012685c775a7bc62357625cbc71b973f5c622ff

Observation 5d4d6f0a-9126-4816-8bc8-bcc575928f07 · outbound

This paper cites Imitator: Personalized speech-driven 3d facial animation.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Imitator: Personalized speech-driven 3d facial animation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:10.215324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:06.304332Z digest=sha256:5d5238ace3a4fded2be11a7a4f00fff192f680e0c62a0f6a44ff2c1219d39328

Observation ea41a362-d842-4fd4-98df-b3789ecea87e · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Neural voice puppetry: Audio-driven facial reenactment

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:10.041306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:06.492254Z digest=sha256:a1564c1d6795a9b8e66f191437ac43ac68d5f2fa9225d9db8e130402ca0f643f

Observation 060bec6f-0892-4851-8333-5a8919ee5369 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:09.862468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:06.616378Z digest=sha256:6cba0c7b294985ae197e1412143be1a1433ec0beb6efd1631592b4e3abbd860b

Observation bd18af0b-fee2-41f1-a369-5dba1b63a086 · outbound

This paper cites End-to-end speech-driven realistic facial animation with temporal gans.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation End-to-end speech-driven realistic facial animation with temporal gans

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:09.698081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:06.789509Z digest=sha256:cfa63f2f5bf596278cbdbae891d7288dd32e881b70e47f0f412fad982626b2d6

Observation 1d9de22c-4ae4-47f7-9b79-9f859a04a1f4 · outbound

This paper cites One- shot talking face generation from single-speaker audio-visual correlation learning.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation One- shot talking face generation from single-speaker audio-visual correlation learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:09.545374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:06.924625Z digest=sha256:5fb59281ea177438a8293aa639276105e13a4caba81d7b9087c6aae1422e0495

Observation 44cbe3cd-bdcc-4584-b2da-c1c17694c446 · outbound

This paper cites Codetalker: Speech-driven 3d facial animation with discrete motion prior.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Codetalker: Speech-driven 3d facial animation with discrete motion prior

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:09.282431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:07.037851Z digest=sha256:f7ccc0af9b24d12313c97d1929d07938a4640ae9a8b76a0f3ac1b0f4ef655f91

Observation 37acc315-2d75-487b-ad54-50a0676394f1 · outbound

This paper cites Feng, Stacy Marsella, and Ari Shapiro.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Feng, Stacy Marsella, and Ari Shapiro

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:09.033295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:07.194342Z digest=sha256:6669830b9b02f60559f3abc8843d8e98b743c6d7d4054b9da262d80c5d58a577

Observation 3ff9156e-1d75-40b2-886e-aac50714b777 · outbound

This paper cites You only speak once to see.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation You only speak once to see

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:08.764484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:07.342653Z digest=sha256:20cc30f74725c0047a6941554ad6f2cdd0d0126fcff2a99e797057e6712d79bf

Observation a57b2918-b078-455e-96e5-8d9bd0cce1d0 · outbound

This paper cites Mining audio, text and visual information for talking face generation.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Mining audio, text and visual information for talking face generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:08.580082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:07.509187Z digest=sha256:3ef75e3ba46102e7780c82f4f6460977b0fe8fff1239eb15b7d81f6073de2163

Observation 250eb052-f068-48e3-ae08-fc2ee0989261 · outbound

This paper cites Motiondif- fuse: Text-driven human motion generation with diffusion model.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Motiondif- fuse: Text-driven human motion generation with diffusion model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:08.278339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:07.598810Z digest=sha256:9e8606f0e0c6d88914da016f4570a49afdbadf50566813285c4fa0ee201b6dbc

Observation 1d34f567-4d6b-45b9-b940-0482e0c8f529 · outbound

This paper cites Media2face: Co-speech facial animation gener- ation with multi-modality guidance.

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation Media2face: Co-speech facial animation gener- ation with multi-modality guidance

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:07.968350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:07.692944Z digest=sha256:bddcfac00623c61a07e7dddab477c895785bd9027952e80ce49345b30dcbd7d9

Pith citing papers

Observation e50333f4-7b0c-4c81-a95b-b4b2d291f5ec · inbound

RAM: Recover Any 3D Human Motion in-the-Wild cites this paper.

RAM: Recover Any 3D Human Motion in-the-Wild Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:45:19.107798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T08:44:49.087457Z digest=sha256:078247514b9c7eac595794d19303ef663fc49cfa8665acc682fc46997440fa7d