Pith. sign in

Paper Citation Record · LEDGER

TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2304.00334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.00334 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:23:42.243682Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T22:48:38.173879Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d572bcdd-c9d7-4801-899e-090da81f8b37 · inbound

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities cites this paper.

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-09T18:23:42.243682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:23:42.243682Z digest=sha256:acac7f904d0deba8be7596385824b6ad4b2a9d9b1b605da02f91fc6186469a3d

Observation 6c502d7a-fa97-41c4-8b13-25c3d36acaa5 · inbound

Exploring Timeline Control for Facial Motion Generation cites this paper.

Exploring Timeline Control for Facial Motion Generation TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:50.626113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:50.626113Z digest=sha256:08712e5cfcbdabb95a6d91c5311e8c2af7bbdd7cfc57fd83e9f260270d47544a

Observation 64cca734-c5d4-4cfc-b158-db51f42a451a · inbound

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding cites this paper.

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:56.996444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:56.996444Z digest=sha256:3a6acde28f61f58fef4a47ff8dcaade878e6190c3100a91493ea53c67b7a7c96

Observation 05d2aa18-834e-4629-82bc-df357b077bd1 · inbound

CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation cites this paper.

CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T19:32:44.266588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:32:44.266588Z digest=sha256:9c9943eea7f1ebee8dcda694347c74b9bd931c2fd86c5fca52302e1044f2325a

Observation d108b1ca-21c1-43c9-af87-5e2c40af0373 · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:40.503235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:40.503235Z digest=sha256:ba1b0fe324acbd8dac3912e92750758c434f034a93ce286db727cb898082985c

Observation aafaaea3-8f20-4361-a88c-38dd72fe881e · inbound

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation cites this paper.

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T11:47:50.527329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:47:50.527329Z digest=sha256:c5d6193d26f22990cb421a30cdbe277ae3a4560f20f0fec0c2b854a9273d3761

Observation 3a40e739-f33a-4a78-9729-aed7e8e10021 · inbound

KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes cites this paper.

KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:48:38.175829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T22:46:51.954117Z digest=sha256:b17e666512f02fd4b9581da327b16fe6bcbb1c9bd9398c30a54a326a4faf9dbe

Observation 0dd9bd23-9571-421e-a853-2f245cd8f8f2 · inbound

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence cites this paper.

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.245024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T08:27:41.839123Z digest=sha256:ef548eb1c36b41bc3e4ded4c516a130ef6be189698da49b7c72b2a4f5b9419d4

Observation 73213a83-b957-4825-bc89-a3b3c87e2f69 · inbound

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions cites this paper.

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T07:49:23.168452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:49:23.168452Z digest=sha256:7600bb9cbab84b016fecd89b1ae0e8fa494237a50db8a048d84a1f77499965b6