Pith. sign in

Paper Citation Record · LEDGER

Emotional Face-to-Speech

As of 10 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 2 inbound Pith citation observations for arXiv:2502.01046.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01046 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:50:43.962905Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T22:26:14.531002Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:03:13.782299Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact2
  • verified fuzzy51
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fc86a2a-63b2-486e-896e-48f12d1f1a68 · outbound

This paper cites write newline.

Emotional Face-to-Speech write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.604844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.604844Z digest=sha256:e4dd5aa4ba8c76fa81e2176e0e371d9df63b61ee9552af9a5d938ee76bd3d8d9

Observation 4a82a1ca-ec06-4862-b264-10040089832f · outbound

This paper cites write newline.

Emotional Face-to-Speech write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.611319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.611319Z digest=sha256:f657d9d8ebd12a7d9b044e050630342a09686e16fa495c80a233b9f6e6a8778e

Observation 8c784a1f-2db3-45d0-982b-b09334faca72 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Emotional Face-to-Speech LRS3-TED: a large-scale dataset for visual speech recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.616656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.616656Z digest=sha256:7b3985e406b4737e36ac41ebde449f5d4709317fe0f924ee1c9d320db87ed355

Observation 228fd9bc-ab5a-4b86-8cbf-37803986be1f · outbound

This paper cites SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing.

Emotional Face-to-Speech SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.141971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.625383Z digest=sha256:30e390e2e902db6910a6ede3e7d7f6022fc5779e2584a995840de2406724e0ff

Observation 1ac5b807-2b4e-45ab-ad6b-f84a0dbfcb44 · outbound

This paper cites D., Ho, J., Tarlow, D., and van den Berg, R.

Emotional Face-to-Speech D., Ho, J., Tarlow, D., and van den Berg, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.125362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.630497Z digest=sha256:a6ad18f8fbb94e7c7401bc0d546595858fdc388686b4e035112b9af518d88f10

Observation bba212cb-3264-4084-8f3b-1641ebdffaea · outbound

This paper cites W., Fidler, S., and Kreis, K.

Emotional Face-to-Speech W., Fidler, S., and Kreis, K

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.109829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.635301Z digest=sha256:5517fa44e9f612d0fa142508060cd075557c3a8f0916f416f80d3032b2679e40

Observation 4001e787-5070-46db-ab23-5ef9159ade57 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:45.095205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.640656Z digest=sha256:8b58f9db3d8de9b5ce1c92e27974f361794abadd84baab3a1434f2f00206f856

Observation 0ef881d3-5e1b-429f-9015-27998dd27e63 · outbound

This paper cites and Zisserman, A.

Emotional Face-to-Speech and Zisserman, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.079738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.645714Z digest=sha256:6aadf2f8c41b71d833974239f11e1869f86d6c481821db6a653fdaa98ce98c46

Observation 7f473141-0a36-416a-a612-0b4510cb8fe6 · outbound

This paper cites V2C: Visual voice cloning.

Emotional Face-to-Speech V2C: Visual voice cloning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.062125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.651033Z digest=sha256:9846df00eee43766956c3107b43ba29f428ffb6e86dc57c748f5153fb0c826b9

Observation 83a53979-b4a1-44f2-8a66-6c3259c198c0 · outbound

This paper cites S., Nagrani, A., and Zisserman, A.

Emotional Face-to-Speech S., Nagrani, A., and Zisserman, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.045828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.655804Z digest=sha256:bd865f0e365982dd5e7d83188c79d5449fb203a6e0a98dec689748119a60e9d4

Observation 88efc1a8-5fdc-4fe4-89f3-006ecdceea6c · outbound

This paper cites Learning to dub movies via hierarchical prosody models.

Emotional Face-to-Speech Learning to dub movies via hierarchical prosody models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.030003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.661008Z digest=sha256:b3b35f48a81e7fc29934ea137a03318901b52b012126db2bb3e6d4da87d29b54

Observation 9d01dd86-d44e-436e-b922-49c544605764 · outbound

This paper cites StyleDubber : Towards multi-scale style learning for movie dubbing.

Emotional Face-to-Speech StyleDubber : Towards multi-scale style learning for movie dubbing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.013211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.665906Z digest=sha256:04c254700925731444fba22738e153d3b89c6c74d71d0b7a1d5f8f4880c2a621

Observation 9b997226-46ca-46b9-b267-0d56038f4fa9 · outbound

This paper cites High fidelity neural audio compression.

Emotional Face-to-Speech High fidelity neural audio compression

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.994838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.670452Z digest=sha256:66b66ade1aa7370e27090b0ba07ac7d340de8fe07c704961efbc402775a6a5e6

Observation 90c26510-fbdd-45a1-ac15-96b6d63c93ec · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Emotional Face-to-Speech Arcface: Additive angular margin loss for deep face recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.978383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.675081Z digest=sha256:0bdcec98bb9ea4ebf9fdc784470f39e5932ebdfe6e48fde763e9a33eaaf4976b

Observation df752898-ec6a-4c72-8170-b51272e7ac6d · outbound

This paper cites and Shutov, V.

Emotional Face-to-Speech and Shutov, V

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.963515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.679918Z digest=sha256:8beb311ccc5f84e6c8b7409bf17588b538ec59955b3dc387ff65dcee1766f2ea

Observation 201828de-4436-4332-9442-1cd2f16623d7 · outbound

This paper cites Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature.

Emotional Face-to-Speech Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.948604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.684435Z digest=sha256:55f5d3167dad3e90a9d7dd26dcafa004f70bd0d8a405f8b76694a978b487a36b

Observation 43a1cc26-d1ee-4d6a-aba2-90c7f0720dbe · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation.

Emotional Face-to-Speech Efficient emotional adaptation for audio-driven talking-head generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.933128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.689001Z digest=sha256:fcfe1d67df37c833d50d28878e3dc2040d55c9ed9646881cf89dca2e8d966b24

Observation 6ce5e821-1e50-408b-9497-39e726890cc7 · outbound

This paper cites Improving adversarial energy-based model via diffusion process.

Emotional Face-to-Speech Improving adversarial energy-based model via diffusion process

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.918063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.693418Z digest=sha256:267d839a040c034b7aea38cc0bf7fb81664cceffe8c8d40f40b47e92651dd269

Observation a6996e8a-99a5-4913-8545-3280c0c30699 · outbound

This paper cites Face2Speech : Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.

Emotional Face-to-Speech Face2Speech : Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.901896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.697928Z digest=sha256:0af4725be0617ada3350a376a0f20e96fab8a39ad668c70daa52719c2b067742

Observation 0dcbacbf-7564-47ef-85b0-061428d49176 · outbound

This paper cites EGC: Image generation and classification via a diffusion energy-based model.

Emotional Face-to-Speech EGC: Image generation and classification via a diffusion energy-based model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.886498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.702510Z digest=sha256:7b2846d437499517d11e4df18f61d4feb7e162eb5c568dfe037d253621897ea4

Observation 7fc0aca9-365b-4dad-96ae-1f9d0521f517 · outbound

This paper cites Emodiff : Intensity controllable emotional text-to-speech with soft-label guidance.

Emotional Face-to-Speech Emodiff : Intensity controllable emotional text-to-speech with soft-label guidance

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.870210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.707083Z digest=sha256:98d1ea616b76753f276fcd353e7ae8ec185f0dc7a07f190f5db71444ceea7864

Observation 7870b00c-eed2-4471-bd49-c34b75d4b41c · outbound

This paper cites An investigation of multi-speaker training for wavenet vocoder.

Emotional Face-to-Speech An investigation of multi-speaker training for wavenet vocoder

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.852149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.711763Z digest=sha256:1c9e9dcec5a85ffd65fe25632b434d1ac9530625833a45ab21ab0c501d66fdfa

Observation bc283b0d-32bc-419e-82b7-98c5ae23e077 · outbound

This paper cites and Salimans, T.

Emotional Face-to-Speech and Salimans, T

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.837142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.716419Z digest=sha256:2825ec033a8c8ec6e1df58efb533fd486715b63f9f6438a42c7f504f1d92cb7e

Observation fe0ceb16-2d99-4d55-bc37-d4d864504a47 · outbound

This paper cites Denoising diffusion probabilistic models.

Emotional Face-to-Speech Denoising diffusion probabilistic models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.821971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.721209Z digest=sha256:63ddfd3e57c01801f866e01410b181a6eccfa869742eb11667f8ade1c1d19659

Observation b596290c-7592-4d5b-ba2c-cdf081b6ec5e · outbound

This paper cites and Johnson, L.

Emotional Face-to-Speech and Johnson, L

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.806622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.725912Z digest=sha256:d56667c6d9fcdf22795a6e34d8df3bce16220b5a84c38e37b66ee1fe53d227d5

Observation 0eaf4e25-4cfb-41a3-9591-b6a03d4ab89a · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.791581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.730525Z digest=sha256:90f0f0c83983cd722267907fee617ff4d67ea7e3eb0abe3a16bfc7329bceab41

Observation e6f77063-e505-4adc-bede-bbd5ddb52ac3 · outbound

This paper cites Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping.

Emotional Face-to-Speech Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.735052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.735052Z digest=sha256:a175156df3d204a2807efe3b38a6fa448bd656830ac811eea7f2222ec9239e1b

Observation c1567119-44ca-4a97-bd0a-c4ba830e4cfe · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.740224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.740224Z digest=sha256:5059e526abbed4dbf0e64b041a5f860a33b25289fd45f45b595fca30c05d060a

Observation 4b7a60f2-fa75-42f0-83c8-bb0b68734653 · outbound

This paper cites Speak, read and prompt: High -fidelity text-to-speech with minimal supervision.

Emotional Face-to-Speech Speak, read and prompt: High -fidelity text-to-speech with minimal supervision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.765114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.744826Z digest=sha256:6b9e93c3616e59d57b2a3cbf154b74323d87cb452be94771095a92ce0d37773b

Observation b425ebf1-8a0f-488a-a629-1bac92c2f7ba · outbound

This paper cites Deep Directed Generative Models with Energy-Based Probability Estimation.

Emotional Face-to-Speech Deep Directed Generative Models with Energy-Based Probability Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.749850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.749850Z digest=sha256:7075f3485b79efd2b323ac3215cdc6553c92fbe9848644ca2aed588bf8e6c263

Observation 67f83871-f69a-45dd-93dd-0b227ece1eae · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.749956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.755018Z digest=sha256:7a730fc4727c5d7ef5da5de90bfe20b8b9760b15269f633ee8afd556de426ff2

Observation 02329e41-fe90-41a0-b8e9-deeb3d81768c · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.734438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.759986Z digest=sha256:5acd13682825ca8b0c0aa8dafce22e664851fed6648b01044db8233994058d65

Observation ef8844f7-c50f-409d-abfc-fe157fb86835 · outbound

This paper cites S., and Chung, S.

Emotional Face-to-Speech S., and Chung, S

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.718803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.764727Z digest=sha256:22bf9f1798046c3b665585d26d2edc90aa4fe5ca5898e26b684e422011be0666

Observation 7badaced-9882-408d-af97-7de45d9d8ef8 · outbound

This paper cites Hear Your Face: Face-based voice conversion with F0 estimation.

Emotional Face-to-Speech Hear Your Face: Face-based voice conversion with F0 estimation

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:50:44.138143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.769527Z digest=sha256:620dd411f234de6d505ae07385d765054c7636fad2c1795cf5c376d1d726f611

Observation e147df56-4a2d-43a1-a161-0e72be10814a · outbound

This paper cites UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts.

Emotional Face-to-Speech UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.774408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.774408Z digest=sha256:d26a65df2d39e45d3d0715ab19766acfe6ee263ac5056d84c2e421148a3fa2fd

Observation 694e3347-2324-4734-a331-2f27283a99a4 · outbound

This paper cites A., Han, C., Raghavan, V.

Emotional Face-to-Speech A., Han, C., Raghavan, V

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.702628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.779447Z digest=sha256:d22aab0f87c45edd92e9f82f98eb84e7766c5c829847e843bbe4298e93b47180

Observation 75876484-b02c-44ce-a574-ebd534e4271b · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.686565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.783888Z digest=sha256:90056d110c6a58b183ff90236fb3c73c526adacf18286d47fa6eef5005125a73

Observation 4eb7e847-dc45-48ec-a169-667fce75cfed · outbound

This paper cites Towards a simultaneous and granular identity-expression control in personalized face generation.

Emotional Face-to-Speech Towards a simultaneous and granular identity-expression control in personalized face generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.668157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.788335Z digest=sha256:cb1d1e68506166682167ce5f37293d07d6c9d5aa1b31bef88d53426b200974e0

Observation b071e055-cb81-4e53-b369-b905175722f6 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.650702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.792876Z digest=sha256:518c5bbc51c7df1999b8bfab3792ad3c117592108dd0e25d44975c5c04700bfc

Observation 05cb4b77-1554-4915-a91f-04da4147c73a · outbound

This paper cites and Hutter, F.

Emotional Face-to-Speech and Hutter, F

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.635934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.797480Z digest=sha256:bcfe5f0aa84acf7a3cd232ff86f56ecf9f3fa06634dadae59e910e8968f090a0

Observation 8884668d-d95d-4799-9f20-f312e427d6c6 · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data distribution.

Emotional Face-to-Speech Discrete diffusion modeling by estimating the ratios of the data distribution

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.620993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.801954Z digest=sha256:c79ae93e911be08a3b6b69f50701e057df977fe8c8fa6b4766d4608e620b035c

Observation 487895ef-ea72-49bc-aca9-17699ef1d044 · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Emotional Face-to-Speech emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.605341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.806366Z digest=sha256:2271482885a68b2cbe53f7c0332c120888a49db8488d4e1017e4db21d33f3b5c

Observation 95de6b4f-48c9-4fcd-821b-bae05433aa0c · outbound

This paper cites POSTER++: A simpler and stronger facial expression recognition network.

Emotional Face-to-Speech POSTER++: A simpler and stronger facial expression recognition network

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.810875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.810875Z digest=sha256:fa45cc1bd19e487b15880109014833ff45c5db57c2a59357d9b0772aa7dba33c

Observation 7282e835-8cd3-4b1b-a943-84abdcd98820 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.589517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.816300Z digest=sha256:066847388d662d010534c746e1bea722e7e5be46b6f92c5af1213ff9cef62bfb

Observation 83977cf2-68c9-4aee-8aa0-e5e0e741473a · outbound

This paper cites Concrete score matching: Generalized score matching for discrete data.

Emotional Face-to-Speech Concrete score matching: Generalized score matching for discrete data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.574543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.820798Z digest=sha256:3ce9f94ac6ee235f43d298da9dd14229d60f34c22b1ef12f6ff463400e1e6d48

Observation 7e9dd3b9-8940-4494-a023-8fe45fcc6085 · outbound

This paper cites HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis.

Emotional Face-to-Speech HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.825414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.825414Z digest=sha256:4212f40f20ce877a06bf56c29a98025c2ad695cb5274babaed2099d5ed547128

Observation afdb7d01-8e67-4cb2-a693-ad523e366df3 · outbound

This paper cites Unlocking Guidance for Discrete State-Space Diffusion and Flow Models.

Emotional Face-to-Speech Unlocking Guidance for Discrete State-Space Diffusion and Flow Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.830118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.830118Z digest=sha256:b0ed7a0fc38deea3900e30007cecedc5e7bc1916c56d64c8ace6d275c9b9e303

Observation 21379b9e-a8fb-44e7-953f-b299025db48a · outbound

This paper cites Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data.

Emotional Face-to-Speech Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.834886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.834886Z digest=sha256:36c668a4d73671d7860e67e5cc7fee625de54df6b3c5e2d073e35eb61863d00e

Observation 4fec4e79-0764-4560-826d-869f4fbe7d28 · outbound

This paper cites Visual form predictions facilitate auditory processing at the n1.

Emotional Face-to-Speech Visual form predictions facilitate auditory processing at the n1

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.559350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.840930Z digest=sha256:ffc33aab2a44e10786e984b1ab3b232ca18afc3945cb89718cc42de07573366e

Observation d8b23048-440a-428e-8eeb-b845b9442490 · outbound

This paper cites and Xie, S.

Emotional Face-to-Speech and Xie, S

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.544780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.845641Z digest=sha256:3acdcc4602f257afb384221857548b8dced7683f6215d6599b5b8172a0924a7a

Observation 5ee9044b-52e0-4d3c-b73f-0b813e45db40 · outbound

This paper cites Hearing faces: Target speaker text-to-speech synthesis from a face.

Emotional Face-to-Speech Hearing faces: Target speaker text-to-speech synthesis from a face

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.530430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.850193Z digest=sha256:dcd4a8a79f058ae030969e0526e35020571169f1410a7305a849e0a2d2daca4d

Observation e6f466b0-cfd9-4b98-8ff9-fbdf8debc2f3 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

Emotional Face-to-Speech W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.515015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.854922Z digest=sha256:e4471860c73f6b703e9f5503e78116a43dca1b376aa6ae29d18c1445324d6e4f

Observation f319cb0e-356d-4538-b860-072a1b7ee870 · outbound

This paper cites A., Bengio, Y., and Courville, A.

Emotional Face-to-Speech A., Bengio, Y., and Courville, A

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.500660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.859270Z digest=sha256:546d9c7956ac939822cc11f4019282e4c5c6fdd504ede0baef6ec8d9ed4164a1

Observation d808a201-85eb-4347-9039-17a490fddc9e · outbound

This paper cites FastSpeech 2: Fast and high-quality end-to-end text to speech.

Emotional Face-to-Speech FastSpeech 2: Fast and high-quality end-to-end text to speech

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.486244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.863822Z digest=sha256:3175dd58ace1ff9791d611a904aa10228ad8b850c18796c6396a98bd8b768c7b

Observation b467fb4c-8a20-45dc-97aa-341d8420bd84 · outbound

This paper cites J., Jin, Q., and Guo, B.

Emotional Face-to-Speech J., Jin, Q., and Guo, B

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.471369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.868134Z digest=sha256:b2016853480fc58d8a8eda59bb013ecd825d6873ca9478db35d3f7e916e925fd

Observation 1e9a56be-cefb-4282-99bd-0850c636f88e · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering.

Emotional Face-to-Speech Facenet: A unified embedding for face recognition and clustering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.456550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.872724Z digest=sha256:d72001613035d4b16105ca58ee0f189aecc691d3d91b38837d1c2c07ac70ce62

Observation a14e57f9-f64d-4bb3-924b-27c0db7805b6 · outbound

This paper cites NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers.

Emotional Face-to-Speech NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.441381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.877177Z digest=sha256:4a9a593af4cc05b6c7f710bf09277b3784818edb8637eb10861089168f020a64

Observation 1e4269af-0220-4bbe-936b-d2ea22cbce0b · outbound

This paper cites Denoising diffusion implicit models.

Emotional Face-to-Speech Denoising diffusion implicit models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.426446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.881486Z digest=sha256:d2c911f7318e843f3fa29821aa5b303544cf424181a4e29ee4955833a9180141

Observation a4f43710-7e79-4477-b3fc-93cce2bdb6cb · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.411261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.886082Z digest=sha256:3fcfef0800b39f2edf7e3f9d59058a5882adadbcd07d7b4d57a0493fdaedd20b

Observation 46ef9fa6-b41a-465a-9401-7e7a96969847 · outbound

This paper cites Attention is all you need in speech separation.

Emotional Face-to-Speech Attention is all you need in speech separation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.396768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.890897Z digest=sha256:124ae760565adf62d83412685a7c37cd1417b562614533afcf22c775b0d5bc0f

Observation 40747989-37d4-44d9-9ee2-83348a6edb91 · outbound

This paper cites Score-based continuous-time discrete diffusion models.

Emotional Face-to-Speech Score-based continuous-time discrete diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.382407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.895228Z digest=sha256:b347bb1cf11e40222802a045b4479329d40c2cf0a15a2bf9a40eabc71d9b2cb1

Observation 67357277-30e9-4a4e-867e-96a6c819c725 · outbound

This paper cites and Fostick, L.

Emotional Face-to-Speech and Fostick, L

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.367580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.899637Z digest=sha256:1f7c2ed96648c54d1299bffa43c5c05f86940912b9278a912c4ab98708af0b5d

Observation 47f5461b-95aa-4381-a1bf-dba0f1d2c3ea · outbound

This paper cites and Hinton, G.

Emotional Face-to-Speech and Hinton, G

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.352130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.904043Z digest=sha256:58f519d938e8f392d5cb609b9a940f60e6be8477a3ea0e06a2bd369c2358807a

Observation 51fafd75-1f5b-4a6c-bde3-d68fa5123445 · outbound

This paper cites and Vanathi, P.

Emotional Face-to-Speech and Vanathi, P

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.336646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.908535Z digest=sha256:ace7724b0529bed02a7ec72a0459127897fc460917fb2974691a066813986351

Observation 47d97212-12bb-42ba-ba0d-5eb6f0572c2b · outbound

This paper cites N., Kaiser, L., and Polosukhin, I.

Emotional Face-to-Speech N., Kaiser, L., and Polosukhin, I

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.319986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.912912Z digest=sha256:cd7efde1845e0129b3b9c8b9f387ca52b1460d7afc269d885a50475b200b1f55

Observation 505ab147-c5c9-486f-80ae-edd969806099 · outbound

This paper cites Generalized end-to-end loss for speaker verification.

Emotional Face-to-Speech Generalized end-to-end loss for speaker verification

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.304950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.917352Z digest=sha256:97f40b7c77fcf41c34fe612d69bd9fb2d29cf91429fea467145d2c2c2cdfe39d

Observation 7215bd37-6571-45cc-9288-b4aa56e3d42f · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Emotional Face-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.921870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.921870Z digest=sha256:9d724b5ac8203b729879b41df3eee1484fafdfd657003eb24d5360452c0abb48

Observation f604beb7-c73f-410f-b68d-c28c69d65f01 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.290346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.926776Z digest=sha256:32c15d86af1638ca2e54837b93784f6080f0098d04150b3a0b4f2ede9099ec73

Observation 7424b5fb-18ff-4f5e-b46e-3ec3804fc604 · outbound

This paper cites J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R.

Emotional Face-to-Speech J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.274761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.931175Z digest=sha256:978c3725d0dd086f43ffe80fc2bf7fee42ac409bb68d0d123c8f46ec143fa6aa

Observation 637cb4a9-52c4-4d75-abe9-81f326f42521 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Emotional Face-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.935681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.935681Z digest=sha256:c863aecf0a640ee351216892e3ab9495fa69d733d0c859a142d1616acb41e892

Observation 0137943d-8457-47dd-bd86-8f5564f3e4b2 · outbound

This paper cites DCTTS: discrete diffusion model with contrastive learning for text-to-speech generation.

Emotional Face-to-Speech DCTTS: discrete diffusion model with contrastive learning for text-to-speech generation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.258792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.940309Z digest=sha256:d49b9d83b057b0e1def97626eff8b296d75f137f59b549d8230ade140dfb2018

Observation b1310808-0727-4144-83c7-132e9c8ab306 · outbound

This paper cites FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model.

Emotional Face-to-Speech FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:50:44.007474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.944733Z digest=sha256:e8c8a29664d6ac3a4e7fd741c3f462cbb9c3bc832b62bd277b58025ef610ed01

Observation a275ec1a-d076-4ea0-ba38-02192b2dc34d · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

Emotional Face-to-Speech Diffsound: Discrete diffusion model for text-to-sound generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.949596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.949596Z digest=sha256:6a6c9dee4dc1e9f38ecf17b9580c8e33ad75e2aea46226778383d7a60850684e

Observation 7af2d74f-5042-49d0-aaad-507e99131309 · outbound

This paper cites SoundStream : An end-to-end neural audio codec.

Emotional Face-to-Speech SoundStream : An end-to-end neural audio codec

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.232022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.953942Z digest=sha256:195f30eda62bbbc88cb551e597e38ae45dd27591f2b20d632cdbc2bc9cfce1c3

Observation 1f177877-1668-4d84-9463-adb26d58ab72 · outbound

This paper cites SpeechTokenizer : Unified speech tokenizer for speech language models.

Emotional Face-to-Speech SpeechTokenizer : Unified speech tokenizer for speech language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.216139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.958341Z digest=sha256:439ac8913ec818561b8432093af7f21a0b2a4164023f2109154d1a6671cce37e

Observation 1302bf96-acf4-4fef-90e8-34fb688272a8 · outbound

This paper cites Srcodec: Split -residual vector quantization for neural speech codec.

Emotional Face-to-Speech Srcodec: Split -residual vector quantization for neural speech codec

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.200703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.962905Z digest=sha256:0537405fb2f2ace8e0639342d1e2c5fe119c7ded49775d656bea4e2f198a6c88

Pith citing papers

Observation 107ff3ea-8160-4e01-853b-32bfeccd21b6 · inbound

Archon: A Unified Multimodal Model for Holistic Digital Human Generation cites this paper.

Archon: A Unified Multimodal Model for Holistic Digital Human Generation Emotional Face-to-Speech

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.784053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:03:04.294439Z digest=sha256:03a28ca23c00eb4af0ea8b4b77ef51793befd31ed4870e3ee3c7233a6413fd5f

Observation cc278dfb-24cf-454f-8d8b-6e0842a3be83 · inbound

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model cites this paper.

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model Emotional Face-to-Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T22:26:14.531002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:26:14.531002Z digest=sha256:833570eeb4e40d24c5e803418de8eaa40725481e6bec8a0ee725dd19a657f6d9