Pith. sign in

Paper Citation Record · LEDGER

Emotional Face-to-Speech

As of 10 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 2 inbound Pith citation observations for arXiv:2502.01046.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01046 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:50:43.962905Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T22:26:14.531002Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:03:13.782299Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact2
  • verified fuzzy51
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fc86a2a-63b2-486e-896e-48f12d1f1a68 · outbound

This paper cites write newline.

Emotional Face-to-Speech write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.604844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.604844Z digest=sha256:e4dd5aa4ba8c76fa81e2176e0e371d9df63b61ee9552af9a5d938ee76bd3d8d9

Observation 4a82a1ca-ec06-4862-b264-10040089832f · outbound

This paper cites write newline.

Emotional Face-to-Speech write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.611319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.611319Z digest=sha256:f657d9d8ebd12a7d9b044e050630342a09686e16fa495c80a233b9f6e6a8778e

Observation 8c784a1f-2db3-45d0-982b-b09334faca72 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Emotional Face-to-Speech LRS3-TED: a large-scale dataset for visual speech recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.616656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.616656Z digest=sha256:7b3985e406b4737e36ac41ebde449f5d4709317fe0f924ee1c9d320db87ed355

Observation 228fd9bc-ab5a-4b86-8cbf-37803986be1f · outbound

This paper cites SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing.

Emotional Face-to-Speech SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.141971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.625383Z digest=sha256:30302bc8f1065bb947c3b2479c22ce30b5d3d1f0237df0a43cd31d373da17988

Observation 1ac5b807-2b4e-45ab-ad6b-f84a0dbfcb44 · outbound

This paper cites D., Ho, J., Tarlow, D., and van den Berg, R.

Emotional Face-to-Speech D., Ho, J., Tarlow, D., and van den Berg, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.125362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.630497Z digest=sha256:b896e57451af1b1ece61666943b4fe608552741cf4cc720fb60a0c73d9ff5799

Observation bba212cb-3264-4084-8f3b-1641ebdffaea · outbound

This paper cites W., Fidler, S., and Kreis, K.

Emotional Face-to-Speech W., Fidler, S., and Kreis, K

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.109829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.635301Z digest=sha256:6a3ce4b62dd2ce4a61dd7f01d5323a0045ce95144e96efe0d079f56f8233c5c4

Observation 4001e787-5070-46db-ab23-5ef9159ade57 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:45.095205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.640656Z digest=sha256:ccc40509ad42ff712626c8f855e2c5a854a41f9cc9c4acb43c2bbb90bf69e09f

Observation 0ef881d3-5e1b-429f-9015-27998dd27e63 · outbound

This paper cites and Zisserman, A.

Emotional Face-to-Speech and Zisserman, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.079738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.645714Z digest=sha256:e32631b994d8d73de69a05e99a2f16fc1609d1fe49639ffe1aa556944446365a

Observation 7f473141-0a36-416a-a612-0b4510cb8fe6 · outbound

This paper cites V2C: Visual voice cloning.

Emotional Face-to-Speech V2C: Visual voice cloning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.062125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.651033Z digest=sha256:20a8c0d9f2d795d8f5652e8e37937ccbbb7c61cf2bbad76cc0d38df949cdc306

Observation 83a53979-b4a1-44f2-8a66-6c3259c198c0 · outbound

This paper cites S., Nagrani, A., and Zisserman, A.

Emotional Face-to-Speech S., Nagrani, A., and Zisserman, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.045828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.655804Z digest=sha256:91d1f10244ab2dda1067e656271abd767db280a24c307ef830aa8c099da7d1c0

Observation 88efc1a8-5fdc-4fe4-89f3-006ecdceea6c · outbound

This paper cites Learning to dub movies via hierarchical prosody models.

Emotional Face-to-Speech Learning to dub movies via hierarchical prosody models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.030003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.661008Z digest=sha256:598d40c9200b3a67b82a2edb73f62b427d244d48b3e5c55d4a5292da2f344cd5

Observation 9d01dd86-d44e-436e-b922-49c544605764 · outbound

This paper cites StyleDubber : Towards multi-scale style learning for movie dubbing.

Emotional Face-to-Speech StyleDubber : Towards multi-scale style learning for movie dubbing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.013211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.665906Z digest=sha256:5053fcdd33fcb3d3e2a2a28c50985bdf9da6f6d273336eb0e351a93d3933e747

Observation 9b997226-46ca-46b9-b267-0d56038f4fa9 · outbound

This paper cites High fidelity neural audio compression.

Emotional Face-to-Speech High fidelity neural audio compression

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.994838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.670452Z digest=sha256:14e7a320a5abfc3f65732f764425d1807efff03013fae457b41a9e57aab2fc66

Observation 90c26510-fbdd-45a1-ac15-96b6d63c93ec · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Emotional Face-to-Speech Arcface: Additive angular margin loss for deep face recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.978383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.675081Z digest=sha256:2702f1e773d0371a112a167aa6ef8f07566de4aa533d250ae391aadd3e9e182b

Observation df752898-ec6a-4c72-8170-b51272e7ac6d · outbound

This paper cites and Shutov, V.

Emotional Face-to-Speech and Shutov, V

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.963515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.679918Z digest=sha256:b28a0f69b4bd8d2cfd5162bfefd50d5dd643690573827bfe69f776c263a9b420

Observation 201828de-4436-4332-9442-1cd2f16623d7 · outbound

This paper cites Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature.

Emotional Face-to-Speech Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.948604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.684435Z digest=sha256:2941c780a0270128c0f6bceeb400d1c67f22f49382cbada23955df0aa13be985

Observation 43a1cc26-d1ee-4d6a-aba2-90c7f0720dbe · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation.

Emotional Face-to-Speech Efficient emotional adaptation for audio-driven talking-head generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.933128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.689001Z digest=sha256:07f2c793bfa3430fe287ad36465b7e6a7285e54e3b86288251c22049362dba28

Observation 6ce5e821-1e50-408b-9497-39e726890cc7 · outbound

This paper cites Improving adversarial energy-based model via diffusion process.

Emotional Face-to-Speech Improving adversarial energy-based model via diffusion process

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.918063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.693418Z digest=sha256:deb9015610f04b1264c05363060e4541b75c12ef6ff4c9e1b4001eda3e28833f

Observation a6996e8a-99a5-4913-8545-3280c0c30699 · outbound

This paper cites Face2Speech : Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.

Emotional Face-to-Speech Face2Speech : Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.901896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.697928Z digest=sha256:604b6365052784abab4a87270bf07d85287efbfd2228c0d454541d007dfc5644

Observation 0dcbacbf-7564-47ef-85b0-061428d49176 · outbound

This paper cites EGC: Image generation and classification via a diffusion energy-based model.

Emotional Face-to-Speech EGC: Image generation and classification via a diffusion energy-based model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.886498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.702510Z digest=sha256:6625bf1c0becdeb28974c758aaaadcaebbec0803f43c06fb555d4064cb414746

Observation 7fc0aca9-365b-4dad-96ae-1f9d0521f517 · outbound

This paper cites Emodiff : Intensity controllable emotional text-to-speech with soft-label guidance.

Emotional Face-to-Speech Emodiff : Intensity controllable emotional text-to-speech with soft-label guidance

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.870210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.707083Z digest=sha256:66c16900f62ddf9e92a4dfaa830cf580ce2f75a762bce3942fde6c9d1f3c5f52

Observation 7870b00c-eed2-4471-bd49-c34b75d4b41c · outbound

This paper cites An investigation of multi-speaker training for wavenet vocoder.

Emotional Face-to-Speech An investigation of multi-speaker training for wavenet vocoder

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.852149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.711763Z digest=sha256:8cbcbf5e52c5f244ee6f20f7650923cae7a90df2d1b2b19b5b3b46ea50a95f37

Observation bc283b0d-32bc-419e-82b7-98c5ae23e077 · outbound

This paper cites and Salimans, T.

Emotional Face-to-Speech and Salimans, T

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.837142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.716419Z digest=sha256:8a2a59e68e9565429cb0ae7401b8f6bb44a12c2887a4691baa582ad76d7879d1

Observation fe0ceb16-2d99-4d55-bc37-d4d864504a47 · outbound

This paper cites Denoising diffusion probabilistic models.

Emotional Face-to-Speech Denoising diffusion probabilistic models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.821971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.721209Z digest=sha256:1401bf3eb4c2922919cd54a657da836f3303b828ea12fb4c1e1e0815f1841bc2

Observation b596290c-7592-4d5b-ba2c-cdf081b6ec5e · outbound

This paper cites and Johnson, L.

Emotional Face-to-Speech and Johnson, L

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.806622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.725912Z digest=sha256:87d72c7c0de09392ed9af455aacf0d7cc23baac84b9874e19bee535454201715

Observation 0eaf4e25-4cfb-41a3-9591-b6a03d4ab89a · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.791581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.730525Z digest=sha256:088da7f643e3460c505d1def152c3801a0fcadeca82056fb085fcdcad3282e5f

Observation e6f77063-e505-4adc-bede-bbd5ddb52ac3 · outbound

This paper cites Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping.

Emotional Face-to-Speech Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.735052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.735052Z digest=sha256:a175156df3d204a2807efe3b38a6fa448bd656830ac811eea7f2222ec9239e1b

Observation c1567119-44ca-4a97-bd0a-c4ba830e4cfe · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.740224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.740224Z digest=sha256:5059e526abbed4dbf0e64b041a5f860a33b25289fd45f45b595fca30c05d060a

Observation 4b7a60f2-fa75-42f0-83c8-bb0b68734653 · outbound

This paper cites Speak, read and prompt: High -fidelity text-to-speech with minimal supervision.

Emotional Face-to-Speech Speak, read and prompt: High -fidelity text-to-speech with minimal supervision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.765114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.744826Z digest=sha256:f4910f3d4a61134803d4941abfcbaa62b8a8429b47e5db7e3ba9736003db9bcc

Observation b425ebf1-8a0f-488a-a629-1bac92c2f7ba · outbound

This paper cites Deep Directed Generative Models with Energy-Based Probability Estimation.

Emotional Face-to-Speech Deep Directed Generative Models with Energy-Based Probability Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.749850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.749850Z digest=sha256:7075f3485b79efd2b323ac3215cdc6553c92fbe9848644ca2aed588bf8e6c263

Observation 67f83871-f69a-45dd-93dd-0b227ece1eae · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.749956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.755018Z digest=sha256:763c9937db1ea27eddf1bf4a41766593961f31134ec3d108185863e7789f6559

Observation 02329e41-fe90-41a0-b8e9-deeb3d81768c · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.734438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.759986Z digest=sha256:22719eee836386da3ed65984a4d451d4a7ab7c33d4b6ce65808ec7f377b1c4c1

Observation ef8844f7-c50f-409d-abfc-fe157fb86835 · outbound

This paper cites S., and Chung, S.

Emotional Face-to-Speech S., and Chung, S

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.718803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.764727Z digest=sha256:b54c33a7e12791040e4b09c8823335fee2ee1c10dd31dd07cf31757e43285fb2

Observation 7badaced-9882-408d-af97-7de45d9d8ef8 · outbound

This paper cites Hear Your Face: Face-based voice conversion with F0 estimation.

Emotional Face-to-Speech Hear Your Face: Face-based voice conversion with F0 estimation

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:50:44.138143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.769527Z digest=sha256:65f77463a2708dc57b11a4a6a81ac639b80c999f76fd3e7c5d78ebefacf567ff

Observation e147df56-4a2d-43a1-a161-0e72be10814a · outbound

This paper cites UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts.

Emotional Face-to-Speech UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.774408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.774408Z digest=sha256:d26a65df2d39e45d3d0715ab19766acfe6ee263ac5056d84c2e421148a3fa2fd

Observation 694e3347-2324-4734-a331-2f27283a99a4 · outbound

This paper cites A., Han, C., Raghavan, V.

Emotional Face-to-Speech A., Han, C., Raghavan, V

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.702628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.779447Z digest=sha256:4ec3783866495441dffe9f46a6162958397334277ad232569a014af1aa2a47f6

Observation 75876484-b02c-44ce-a574-ebd534e4271b · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.686565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.783888Z digest=sha256:a59aa05b447a9d5495039f212b629d60ed5b88a0dd208ba516ab07e380868317

Observation 4eb7e847-dc45-48ec-a169-667fce75cfed · outbound

This paper cites Towards a simultaneous and granular identity-expression control in personalized face generation.

Emotional Face-to-Speech Towards a simultaneous and granular identity-expression control in personalized face generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.668157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.788335Z digest=sha256:c1e7d9e26bae96c96a373efd742da34777d49b7ba3a60cb7e50c92f94efa5112

Observation b071e055-cb81-4e53-b369-b905175722f6 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.650702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.792876Z digest=sha256:292450cd5c33b841a4f5c14775f5fb3579e4ef17849a9e0deb2d2054cc8d801d

Observation 05cb4b77-1554-4915-a91f-04da4147c73a · outbound

This paper cites and Hutter, F.

Emotional Face-to-Speech and Hutter, F

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.635934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.797480Z digest=sha256:747fea06c6e6ad4cc77bb6248ceacb135401600584dbfa50fb3e4a233c0f393b

Observation 8884668d-d95d-4799-9f20-f312e427d6c6 · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data distribution.

Emotional Face-to-Speech Discrete diffusion modeling by estimating the ratios of the data distribution

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.620993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.801954Z digest=sha256:57abf0683728eb33d874672e4746831beebbc890c58ff09e09ff5b24d7844d55

Observation 487895ef-ea72-49bc-aca9-17699ef1d044 · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Emotional Face-to-Speech emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.605341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.806366Z digest=sha256:d537f537f6c655337378295d57b43ff1124714960b736abe3af3c27c40af966a

Observation 95de6b4f-48c9-4fcd-821b-bae05433aa0c · outbound

This paper cites POSTER++: A simpler and stronger facial expression recognition network.

Emotional Face-to-Speech POSTER++: A simpler and stronger facial expression recognition network

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.810875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.810875Z digest=sha256:fa45cc1bd19e487b15880109014833ff45c5db57c2a59357d9b0772aa7dba33c

Observation 7282e835-8cd3-4b1b-a943-84abdcd98820 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.589517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.816300Z digest=sha256:271346f204e6c13ea2a65b9055bb49e731667ab8cfc8fd0b1174b5f70766edf9

Observation 83977cf2-68c9-4aee-8aa0-e5e0e741473a · outbound

This paper cites Concrete score matching: Generalized score matching for discrete data.

Emotional Face-to-Speech Concrete score matching: Generalized score matching for discrete data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.574543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.820798Z digest=sha256:76b4ed7efcc8b4867e92cc470954750110e8c0b5df45cf210565946f9ff418c9

Observation 7e9dd3b9-8940-4494-a023-8fe45fcc6085 · outbound

This paper cites HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis.

Emotional Face-to-Speech HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.825414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.825414Z digest=sha256:4212f40f20ce877a06bf56c29a98025c2ad695cb5274babaed2099d5ed547128

Observation afdb7d01-8e67-4cb2-a693-ad523e366df3 · outbound

This paper cites Unlocking Guidance for Discrete State-Space Diffusion and Flow Models.

Emotional Face-to-Speech Unlocking Guidance for Discrete State-Space Diffusion and Flow Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.830118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.830118Z digest=sha256:b0ed7a0fc38deea3900e30007cecedc5e7bc1916c56d64c8ace6d275c9b9e303

Observation 21379b9e-a8fb-44e7-953f-b299025db48a · outbound

This paper cites Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data.

Emotional Face-to-Speech Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.834886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.834886Z digest=sha256:36c668a4d73671d7860e67e5cc7fee625de54df6b3c5e2d073e35eb61863d00e

Observation 4fec4e79-0764-4560-826d-869f4fbe7d28 · outbound

This paper cites Visual form predictions facilitate auditory processing at the n1.

Emotional Face-to-Speech Visual form predictions facilitate auditory processing at the n1

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.559350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.840930Z digest=sha256:5db2baa9832eb9b4c7f2eca572efa0f2167d9c637265a91cb77d88a99bcef716

Observation d8b23048-440a-428e-8eeb-b845b9442490 · outbound

This paper cites and Xie, S.

Emotional Face-to-Speech and Xie, S

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.544780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.845641Z digest=sha256:fa1686f1691a331d6d37d35f0889e33a9371c641e3362c304c242dfdea9f8c71

Observation 5ee9044b-52e0-4d3c-b73f-0b813e45db40 · outbound

This paper cites Hearing faces: Target speaker text-to-speech synthesis from a face.

Emotional Face-to-Speech Hearing faces: Target speaker text-to-speech synthesis from a face

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.530430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.850193Z digest=sha256:19012b3b5ea7525c0daca839fc340ae47306e1a7d619596574f94817ad7ef4c5

Observation e6f466b0-cfd9-4b98-8ff9-fbdf8debc2f3 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

Emotional Face-to-Speech W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.515015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.854922Z digest=sha256:7cde0a6277a2d00d292a53c4a9bf32abb414ec247010aed53d82a3761ebc3324

Observation f319cb0e-356d-4538-b860-072a1b7ee870 · outbound

This paper cites A., Bengio, Y., and Courville, A.

Emotional Face-to-Speech A., Bengio, Y., and Courville, A

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.500660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.859270Z digest=sha256:9e545ab4036b56cb045fd4650ab70a18b1d6f388f774a94c094bc2b67a50e217

Observation d808a201-85eb-4347-9039-17a490fddc9e · outbound

This paper cites FastSpeech 2: Fast and high-quality end-to-end text to speech.

Emotional Face-to-Speech FastSpeech 2: Fast and high-quality end-to-end text to speech

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.486244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.863822Z digest=sha256:1ce14a3c7825306fdba04369876ca10dac27cba31c82b04be0fd17c3e65735da

Observation b467fb4c-8a20-45dc-97aa-341d8420bd84 · outbound

This paper cites J., Jin, Q., and Guo, B.

Emotional Face-to-Speech J., Jin, Q., and Guo, B

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.471369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.868134Z digest=sha256:f7e298bf80f1dea721dee9d61a012841d75d9a6b23cb87f1cdb10c8d48264d72

Observation 1e9a56be-cefb-4282-99bd-0850c636f88e · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering.

Emotional Face-to-Speech Facenet: A unified embedding for face recognition and clustering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.456550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.872724Z digest=sha256:2870947865c10d6e9065edde76553cacb17143ac849941a5cd3387ec262a85cc

Observation a14e57f9-f64d-4bb3-924b-27c0db7805b6 · outbound

This paper cites NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers.

Emotional Face-to-Speech NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.441381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.877177Z digest=sha256:968295aa7fc4dc44b2fee4be9376ddc30f69d9729aba4a1c2e5a6a5b524129db

Observation 1e4269af-0220-4bbe-936b-d2ea22cbce0b · outbound

This paper cites Denoising diffusion implicit models.

Emotional Face-to-Speech Denoising diffusion implicit models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.426446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.881486Z digest=sha256:0a6eff9dafcddcd0d037a39145ea2b66eeb84f725ce896317718801aa15ffb2a

Observation a4f43710-7e79-4477-b3fc-93cce2bdb6cb · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.411261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.886082Z digest=sha256:a485a8e34f02a1ea1997dc39dc6fd495e1cd579aee78e0be49fe772e25482e78

Observation 46ef9fa6-b41a-465a-9401-7e7a96969847 · outbound

This paper cites Attention is all you need in speech separation.

Emotional Face-to-Speech Attention is all you need in speech separation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.396768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.890897Z digest=sha256:03d4df66c09072374031d6244a1ef374037cc9e0e09aeec7840c4e4669f3cd25

Observation 40747989-37d4-44d9-9ee2-83348a6edb91 · outbound

This paper cites Score-based continuous-time discrete diffusion models.

Emotional Face-to-Speech Score-based continuous-time discrete diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.382407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.895228Z digest=sha256:d034cc71937ccefe712cea83d321b80cce6fac5e2406a5a9dad1f6f446b1c226

Observation 67357277-30e9-4a4e-867e-96a6c819c725 · outbound

This paper cites and Fostick, L.

Emotional Face-to-Speech and Fostick, L

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.367580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.899637Z digest=sha256:c1fbfe6fbe4fe8d256c7b1eaf7240c73ade049c58c47a93bc31dfc5146cef3ea

Observation 47f5461b-95aa-4381-a1bf-dba0f1d2c3ea · outbound

This paper cites and Hinton, G.

Emotional Face-to-Speech and Hinton, G

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.352130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.904043Z digest=sha256:171bd1cfb7285e36ce3f82018bdaf78ecfeb111bbd91f52bcac4b494d364b60e

Observation 51fafd75-1f5b-4a6c-bde3-d68fa5123445 · outbound

This paper cites and Vanathi, P.

Emotional Face-to-Speech and Vanathi, P

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.336646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.908535Z digest=sha256:d1172028be59b4909a82231d6eaa3dc0ed29647c1f7e4625f1c2894e9a6a04bf

Observation 47d97212-12bb-42ba-ba0d-5eb6f0572c2b · outbound

This paper cites N., Kaiser, L., and Polosukhin, I.

Emotional Face-to-Speech N., Kaiser, L., and Polosukhin, I

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.319986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.912912Z digest=sha256:4300c584c26b041e1d6e830aba57f50758e97987bb6226790cd84996e4c5e041

Observation 505ab147-c5c9-486f-80ae-edd969806099 · outbound

This paper cites Generalized end-to-end loss for speaker verification.

Emotional Face-to-Speech Generalized end-to-end loss for speaker verification

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.304950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.917352Z digest=sha256:3f156683ee04d355c19b996e4791b9ee11e2623646b0b93279608d3e370b0b90

Observation 7215bd37-6571-45cc-9288-b4aa56e3d42f · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Emotional Face-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.921870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.921870Z digest=sha256:9d724b5ac8203b729879b41df3eee1484fafdfd657003eb24d5360452c0abb48

Observation f604beb7-c73f-410f-b68d-c28c69d65f01 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.290346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.926776Z digest=sha256:b29553080b1a71f0d83b25232f135d6ab417bdf5b3a3cc5c22f900f52bc4aca8

Observation 7424b5fb-18ff-4f5e-b46e-3ec3804fc604 · outbound

This paper cites J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R.

Emotional Face-to-Speech J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.274761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.931175Z digest=sha256:cd8b9306488a8fd2443fbc304c0bbcb6f97897a6d12045f8d69cfc3c6547bafb

Observation 637cb4a9-52c4-4d75-abe9-81f326f42521 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Emotional Face-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.935681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.935681Z digest=sha256:c863aecf0a640ee351216892e3ab9495fa69d733d0c859a142d1616acb41e892

Observation 0137943d-8457-47dd-bd86-8f5564f3e4b2 · outbound

This paper cites DCTTS: discrete diffusion model with contrastive learning for text-to-speech generation.

Emotional Face-to-Speech DCTTS: discrete diffusion model with contrastive learning for text-to-speech generation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.258792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.940309Z digest=sha256:8dac67f263159d0c987ded310669a48b4c5be5eed0c3e23fcedf23b13f1e5ede

Observation b1310808-0727-4144-83c7-132e9c8ab306 · outbound

This paper cites FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model.

Emotional Face-to-Speech FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:50:44.007474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.944733Z digest=sha256:2a692e7eabecca7587f6bf15285f0b316cc06c683ba29c633800d39a55fe0efa

Observation a275ec1a-d076-4ea0-ba38-02192b2dc34d · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

Emotional Face-to-Speech Diffsound: Discrete diffusion model for text-to-sound generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.949596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.949596Z digest=sha256:6a6c9dee4dc1e9f38ecf17b9580c8e33ad75e2aea46226778383d7a60850684e

Observation 7af2d74f-5042-49d0-aaad-507e99131309 · outbound

This paper cites SoundStream : An end-to-end neural audio codec.

Emotional Face-to-Speech SoundStream : An end-to-end neural audio codec

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.232022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.953942Z digest=sha256:a6e7659efcebce9262949622b34b36ac60302882a72e1dee3eb2bdf5f5378006

Observation 1f177877-1668-4d84-9463-adb26d58ab72 · outbound

This paper cites SpeechTokenizer : Unified speech tokenizer for speech language models.

Emotional Face-to-Speech SpeechTokenizer : Unified speech tokenizer for speech language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.216139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.958341Z digest=sha256:dfa23d1e4396b157b36dac5cc67b600b1b9ebc634ca7a1a58badf58471f117f2

Observation 1302bf96-acf4-4fef-90e8-34fb688272a8 · outbound

This paper cites Srcodec: Split -residual vector quantization for neural speech codec.

Emotional Face-to-Speech Srcodec: Split -residual vector quantization for neural speech codec

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.200703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.962905Z digest=sha256:6617cbbf96aba67d3ab1809c5df53c1f7652afa2eb7275c7c9a2a3aafead72cc

Pith citing papers

Observation 107ff3ea-8160-4e01-853b-32bfeccd21b6 · inbound

Archon: A Unified Multimodal Model for Holistic Digital Human Generation cites this paper.

Archon: A Unified Multimodal Model for Holistic Digital Human Generation Emotional Face-to-Speech

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.784053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:03:04.294439Z digest=sha256:8b46b9f152df0ccff134026f94ae72d03ed759f603adca99c9ac0adaa8ba7361

Observation cc278dfb-24cf-454f-8d8b-6e0842a3be83 · inbound

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model cites this paper.

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model Emotional Face-to-Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T22:26:14.531002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:26:14.531002Z digest=sha256:833570eeb4e40d24c5e803418de8eaa40725481e6bec8a0ee725dd19a657f6d9