Pith. sign in

Paper Citation Record · LEDGER

Emotional Face-to-Speech

As of 17 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 4 inbound Pith citation observations for arXiv:2502.01046.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01046 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:50:43.962905Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:03.314503Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:03:13.782299Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact2
  • verified fuzzy51
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fc86a2a-63b2-486e-896e-48f12d1f1a68 · outbound

This paper cites write newline.

Emotional Face-to-Speech write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.604844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.604844Z digest=sha256:b780b26c98ebf8ae9939239afb2225b1ee92dc57616a0bef037e50158e7490c8

Observation 4a82a1ca-ec06-4862-b264-10040089832f · outbound

This paper cites write newline.

Emotional Face-to-Speech write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.611319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.611319Z digest=sha256:b3a070fd3bddb05bc669e61654795324904e1f68d913929c077dc9aab34b6635

Observation 8c784a1f-2db3-45d0-982b-b09334faca72 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Emotional Face-to-Speech LRS3-TED: a large-scale dataset for visual speech recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.616656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.616656Z digest=sha256:a12a9b60d4fcc3e0919d0a9eb59f9cdbebf525366fcfda32bba7ba1d534c9ecc

Observation 228fd9bc-ab5a-4b86-8cbf-37803986be1f · outbound

This paper cites SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing.

Emotional Face-to-Speech SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.141971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.625383Z digest=sha256:dac2e6b08f5f6b12874f63dc4959bbc7fb31d1af5fc1d196eec0a468fc4ebf1b

Observation 1ac5b807-2b4e-45ab-ad6b-f84a0dbfcb44 · outbound

This paper cites D., Ho, J., Tarlow, D., and van den Berg, R.

Emotional Face-to-Speech D., Ho, J., Tarlow, D., and van den Berg, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.125362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.630497Z digest=sha256:48b484ad19bcdbadfe827e9d5479f7062c89d13269eab99ecf27f91391f455cf

Observation bba212cb-3264-4084-8f3b-1641ebdffaea · outbound

This paper cites W., Fidler, S., and Kreis, K.

Emotional Face-to-Speech W., Fidler, S., and Kreis, K

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.109829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.635301Z digest=sha256:85c8bf5ca73c8b1787c39ae23d97b8623e0e2988358fca2fcd8a3a07484410d1

Observation 4001e787-5070-46db-ab23-5ef9159ade57 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:45.095205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.640656Z digest=sha256:34ab5ad2433cb530e006bc883566d49707849bbec50135c378859dc529a5ef49

Observation 0ef881d3-5e1b-429f-9015-27998dd27e63 · outbound

This paper cites and Zisserman, A.

Emotional Face-to-Speech and Zisserman, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.079738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.645714Z digest=sha256:01868242ee8a5f033b80170a086021271dcad714db37a53623a822f0f78cc4d3

Observation 7f473141-0a36-416a-a612-0b4510cb8fe6 · outbound

This paper cites V2C: Visual voice cloning.

Emotional Face-to-Speech V2C: Visual voice cloning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.062125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.651033Z digest=sha256:e496a1ec7a093bb56b2df8911e1855521b4324b81370b112863da3c22ac344a5

Observation 83a53979-b4a1-44f2-8a66-6c3259c198c0 · outbound

This paper cites S., Nagrani, A., and Zisserman, A.

Emotional Face-to-Speech S., Nagrani, A., and Zisserman, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.045828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.655804Z digest=sha256:323f5c8f60924b5f3d46950b10c213e665296df8df516cf880dd4861484e22d9

Observation 88efc1a8-5fdc-4fe4-89f3-006ecdceea6c · outbound

This paper cites Learning to dub movies via hierarchical prosody models.

Emotional Face-to-Speech Learning to dub movies via hierarchical prosody models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.030003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.661008Z digest=sha256:05af5903c2e4cbd3eb7442a620ebc8119a7b0c96824eaeb02e72a21da7dd3a36

Observation 9d01dd86-d44e-436e-b922-49c544605764 · outbound

This paper cites StyleDubber : Towards multi-scale style learning for movie dubbing.

Emotional Face-to-Speech StyleDubber : Towards multi-scale style learning for movie dubbing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.013211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.665906Z digest=sha256:cb9a0f8095f655731a1ca1cb3dbd8f740feddd22fff89b327605018425350d96

Observation 9b997226-46ca-46b9-b267-0d56038f4fa9 · outbound

This paper cites High fidelity neural audio compression.

Emotional Face-to-Speech High fidelity neural audio compression

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.994838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.670452Z digest=sha256:7f6f9710925e839f75030948c3846f39ca59bf12287ad942975487746cf80d6b

Observation 90c26510-fbdd-45a1-ac15-96b6d63c93ec · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Emotional Face-to-Speech Arcface: Additive angular margin loss for deep face recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.978383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.675081Z digest=sha256:08cb9daea2e480d7068a5dc780aa65987dd9de74d84f4a4e96caf10264707031

Observation df752898-ec6a-4c72-8170-b51272e7ac6d · outbound

This paper cites and Shutov, V.

Emotional Face-to-Speech and Shutov, V

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.963515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.679918Z digest=sha256:6350492da5f1f7bbf10ddf191ee1bfff41657b09e8e1cb941191dcdf42c879af

Observation 201828de-4436-4332-9442-1cd2f16623d7 · outbound

This paper cites Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature.

Emotional Face-to-Speech Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.948604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.684435Z digest=sha256:01589aa2e8df5539786fcf3db2a1ec4d7dd76c506255d18e0f872bafe39062e6

Observation 43a1cc26-d1ee-4d6a-aba2-90c7f0720dbe · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation.

Emotional Face-to-Speech Efficient emotional adaptation for audio-driven talking-head generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.933128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.689001Z digest=sha256:f0d10a2c0b5557c357fa1794b3031fca36cf52acd8e18e8a6c80916951f841a3

Observation 6ce5e821-1e50-408b-9497-39e726890cc7 · outbound

This paper cites Improving adversarial energy-based model via diffusion process.

Emotional Face-to-Speech Improving adversarial energy-based model via diffusion process

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.918063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.693418Z digest=sha256:cb2e3a252a2a506424e1637813efe3602490a62468d1031608b59313ac13074f

Observation a6996e8a-99a5-4913-8545-3280c0c30699 · outbound

This paper cites Face2Speech : Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.

Emotional Face-to-Speech Face2Speech : Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.901896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.697928Z digest=sha256:e0961c71a0839929b6f7e28c1aa96bc4c3b94b16ed8d44c3fb26e5189725641f

Observation 0dcbacbf-7564-47ef-85b0-061428d49176 · outbound

This paper cites EGC: Image generation and classification via a diffusion energy-based model.

Emotional Face-to-Speech EGC: Image generation and classification via a diffusion energy-based model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.886498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.702510Z digest=sha256:ad5c8f68e1c953c0aca0ecba8f7c40c2e9f34b7f71cc62e6a373a9467e5f970e

Observation 7fc0aca9-365b-4dad-96ae-1f9d0521f517 · outbound

This paper cites Emodiff : Intensity controllable emotional text-to-speech with soft-label guidance.

Emotional Face-to-Speech Emodiff : Intensity controllable emotional text-to-speech with soft-label guidance

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.870210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.707083Z digest=sha256:e6033887ea8ab59a082ecf4b9345ceac7b8cef3a52af065c688be3fc173e2ffb

Observation 7870b00c-eed2-4471-bd49-c34b75d4b41c · outbound

This paper cites An investigation of multi-speaker training for wavenet vocoder.

Emotional Face-to-Speech An investigation of multi-speaker training for wavenet vocoder

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.852149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.711763Z digest=sha256:a83fa9e4d27247c120d1ee5352847923dffe988354224b39e9e988ad46b425e5

Observation bc283b0d-32bc-419e-82b7-98c5ae23e077 · outbound

This paper cites and Salimans, T.

Emotional Face-to-Speech and Salimans, T

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.837142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.716419Z digest=sha256:1007a9b93667042890ace121015730611436a5adcbe177d5322f12e0f00be46d

Observation fe0ceb16-2d99-4d55-bc37-d4d864504a47 · outbound

This paper cites Denoising diffusion probabilistic models.

Emotional Face-to-Speech Denoising diffusion probabilistic models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.821971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.721209Z digest=sha256:e328b0ed922679c15126c4b96b6750cf2523732a1bd7a9c5b0589a098f390b1a

Observation b596290c-7592-4d5b-ba2c-cdf081b6ec5e · outbound

This paper cites and Johnson, L.

Emotional Face-to-Speech and Johnson, L

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.806622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.725912Z digest=sha256:bbd37728b5eb5f9185cc6010086c6a3d4317b7070f914600e5fa1c16627cc2f0

Observation 0eaf4e25-4cfb-41a3-9591-b6a03d4ab89a · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.791581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.730525Z digest=sha256:306c2b592e788b2bb1efa348412e62a82bcc3083acd7c75a9ede9b44d20dde5b

Observation e6f77063-e505-4adc-bede-bbd5ddb52ac3 · outbound

This paper cites Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping.

Emotional Face-to-Speech Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.735052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.735052Z digest=sha256:3eb590d80886cbdb80f7b2d7a2aaeb33893fc328a85cea07166c35b6e1e729c0

Observation c1567119-44ca-4a97-bd0a-c4ba830e4cfe · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.740224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.740224Z digest=sha256:b7c280da78c5b34fd15301e6e6c264aabbae30ba98816e4f4212d3119b9bd7ca

Observation 4b7a60f2-fa75-42f0-83c8-bb0b68734653 · outbound

This paper cites Speak, read and prompt: High -fidelity text-to-speech with minimal supervision.

Emotional Face-to-Speech Speak, read and prompt: High -fidelity text-to-speech with minimal supervision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.765114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.744826Z digest=sha256:2db60762b7c8229da749b7bf043955db5b32ae59ffb81421149dd1a65c84684d

Observation b425ebf1-8a0f-488a-a629-1bac92c2f7ba · outbound

This paper cites Deep Directed Generative Models with Energy-Based Probability Estimation.

Emotional Face-to-Speech Deep Directed Generative Models with Energy-Based Probability Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.749850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.749850Z digest=sha256:2f4487a8939ae9512d13a578e6f0d062da35f08c27e5124af2bd88aa37f68e44

Observation 67f83871-f69a-45dd-93dd-0b227ece1eae · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.749956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.755018Z digest=sha256:777b89f1d48796ffdb02987d6c939df83f8eb603cd4ae9d303b75b7cfe78bbb2

Observation 02329e41-fe90-41a0-b8e9-deeb3d81768c · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.734438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.759986Z digest=sha256:d6291c7a2513653e740163c8e95ee848dd18f6f3a5bc1e90a75eb58e4a5017c3

Observation ef8844f7-c50f-409d-abfc-fe157fb86835 · outbound

This paper cites S., and Chung, S.

Emotional Face-to-Speech S., and Chung, S

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.718803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.764727Z digest=sha256:6a179c5cb8bce20b0ec4b34d4170e1a0e659fb987199a5996cc9a7e0c462167d

Observation 7badaced-9882-408d-af97-7de45d9d8ef8 · outbound

This paper cites Hear Your Face: Face-based voice conversion with F0 estimation.

Emotional Face-to-Speech Hear Your Face: Face-based voice conversion with F0 estimation

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:50:44.138143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.769527Z digest=sha256:d14a120d5584311479d112097699ce5c6b4020bdf37dfbdae3c205da0fdafe58

Observation e147df56-4a2d-43a1-a161-0e72be10814a · outbound

This paper cites UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts.

Emotional Face-to-Speech UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.774408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.774408Z digest=sha256:a7572cc4d80235d01e1bca99f63ecd0500696f4291ddcd3e06bbc8f208ac1b89

Observation 694e3347-2324-4734-a331-2f27283a99a4 · outbound

This paper cites A., Han, C., Raghavan, V.

Emotional Face-to-Speech A., Han, C., Raghavan, V

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.702628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.779447Z digest=sha256:dbfecc2843463faeaa1d000287af14371d69ddc254658ac8816f57e75636e8ee

Observation 75876484-b02c-44ce-a574-ebd534e4271b · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.686565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.783888Z digest=sha256:5bf428b9701faa0c2b78a6f2c2b94bc560a0d61c12c7cb357076984d25d54867

Observation 4eb7e847-dc45-48ec-a169-667fce75cfed · outbound

This paper cites Towards a simultaneous and granular identity-expression control in personalized face generation.

Emotional Face-to-Speech Towards a simultaneous and granular identity-expression control in personalized face generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.668157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.788335Z digest=sha256:3b19b3c69566ac53392bf02987cc216f24b94ffd258dffc140a736d985103e8d

Observation b071e055-cb81-4e53-b369-b905175722f6 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.650702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.792876Z digest=sha256:e80c854ef792505195339a224fc97a65552211c9421592bbf4255e44a9fb848e

Observation 05cb4b77-1554-4915-a91f-04da4147c73a · outbound

This paper cites and Hutter, F.

Emotional Face-to-Speech and Hutter, F

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.635934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.797480Z digest=sha256:da6e7b72b4a52aa9226fa566a4d20d38ef503ed06d93800ee4b006d55d4bdc6b

Observation 8884668d-d95d-4799-9f20-f312e427d6c6 · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data distribution.

Emotional Face-to-Speech Discrete diffusion modeling by estimating the ratios of the data distribution

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.620993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.801954Z digest=sha256:ce0034dd36a1f5367ce93455b53cbd58456c5fdeb5ac37bb89aa8fbb2b18830a

Observation 487895ef-ea72-49bc-aca9-17699ef1d044 · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Emotional Face-to-Speech emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.605341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.806366Z digest=sha256:5d110865e9d7b3779dc0c6d5fb9875b532428fdfff12bedbcb44b87cfdd94124

Observation 95de6b4f-48c9-4fcd-821b-bae05433aa0c · outbound

This paper cites POSTER++: A simpler and stronger facial expression recognition network.

Emotional Face-to-Speech POSTER++: A simpler and stronger facial expression recognition network

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.810875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.810875Z digest=sha256:9df32b1a4d7a8355a74ce0ac2ca719e36db401eb7666ff7c0ffd1911dd82e586

Observation 7282e835-8cd3-4b1b-a943-84abdcd98820 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.589517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.816300Z digest=sha256:b6ca961bfb3fc9048845031897524a964bd1b28d4374281ad9e935664a3ee23c

Observation 83977cf2-68c9-4aee-8aa0-e5e0e741473a · outbound

This paper cites Concrete score matching: Generalized score matching for discrete data.

Emotional Face-to-Speech Concrete score matching: Generalized score matching for discrete data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.574543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.820798Z digest=sha256:4bcaa4ed5a870951ec0bdad0bac599d440be08b11b2cf7756f4e31793f57d34d

Observation 7e9dd3b9-8940-4494-a023-8fe45fcc6085 · outbound

This paper cites HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis.

Emotional Face-to-Speech HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.825414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.825414Z digest=sha256:433f3083aaf153fce7306d33c8b4021b05a29853b9487c819ecede170a0edf1f

Observation afdb7d01-8e67-4cb2-a693-ad523e366df3 · outbound

This paper cites Unlocking Guidance for Discrete State-Space Diffusion and Flow Models.

Emotional Face-to-Speech Unlocking Guidance for Discrete State-Space Diffusion and Flow Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.830118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.830118Z digest=sha256:f199ba76cec94ea2216a7a06258e9904fa0695d02f41067f3216846d24cd4d14

Observation 21379b9e-a8fb-44e7-953f-b299025db48a · outbound

This paper cites Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data.

Emotional Face-to-Speech Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.834886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.834886Z digest=sha256:dedaf916d289720b434cd520c83276caa393151a7427a09c0fbec49393ab62d8

Observation 4fec4e79-0764-4560-826d-869f4fbe7d28 · outbound

This paper cites Visual form predictions facilitate auditory processing at the n1.

Emotional Face-to-Speech Visual form predictions facilitate auditory processing at the n1

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.559350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.840930Z digest=sha256:4e78a9e0e5a74f02cda0f039193ffca51c2ae5e1803d45269797f39ac1665c0b

Observation d8b23048-440a-428e-8eeb-b845b9442490 · outbound

This paper cites and Xie, S.

Emotional Face-to-Speech and Xie, S

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.544780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.845641Z digest=sha256:fa693970cad45b63f29b8683fc75122b142743ca60cfcc60d5cd9c375e4786a8

Observation 5ee9044b-52e0-4d3c-b73f-0b813e45db40 · outbound

This paper cites Hearing faces: Target speaker text-to-speech synthesis from a face.

Emotional Face-to-Speech Hearing faces: Target speaker text-to-speech synthesis from a face

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.530430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.850193Z digest=sha256:c3eae6ba08c922a798c8227d1bfc6369a1940433dfb6cc3ae0c8a4c772b7cb02

Observation e6f466b0-cfd9-4b98-8ff9-fbdf8debc2f3 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

Emotional Face-to-Speech W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.515015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.854922Z digest=sha256:ae57481213cfdf43357080b22c804f3d1b83b96386b3ba63ec512da1fb9c186a

Observation f319cb0e-356d-4538-b860-072a1b7ee870 · outbound

This paper cites A., Bengio, Y., and Courville, A.

Emotional Face-to-Speech A., Bengio, Y., and Courville, A

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.500660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.859270Z digest=sha256:210759b861a55dc975e734bd5f706bfa3dff8ecc9e2a895b749655ebf3068fcd

Observation d808a201-85eb-4347-9039-17a490fddc9e · outbound

This paper cites FastSpeech 2: Fast and high-quality end-to-end text to speech.

Emotional Face-to-Speech FastSpeech 2: Fast and high-quality end-to-end text to speech

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.486244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.863822Z digest=sha256:5607ab4519c958b2e78e9c66484994ce00c9efa10364530e90b131204b88a723

Observation b467fb4c-8a20-45dc-97aa-341d8420bd84 · outbound

This paper cites J., Jin, Q., and Guo, B.

Emotional Face-to-Speech J., Jin, Q., and Guo, B

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.471369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.868134Z digest=sha256:e223fdce0deb317753c0da6592dc92894088a27d91a648208ed028be70f87cf2

Observation 1e9a56be-cefb-4282-99bd-0850c636f88e · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering.

Emotional Face-to-Speech Facenet: A unified embedding for face recognition and clustering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.456550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.872724Z digest=sha256:082cc67aa76c3cb0ff2e07fb3c213e71a983e51bdc115a5d7276984f3d7734e4

Observation a14e57f9-f64d-4bb3-924b-27c0db7805b6 · outbound

This paper cites NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers.

Emotional Face-to-Speech NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.441381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.877177Z digest=sha256:7d025d198b6902d5276131900b66cd0c167432afea1f21ec1dba6cd0138e6441

Observation 1e4269af-0220-4bbe-936b-d2ea22cbce0b · outbound

This paper cites Denoising diffusion implicit models.

Emotional Face-to-Speech Denoising diffusion implicit models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.426446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.881486Z digest=sha256:c691b1d8c560062b53ad6ab01442b8969d0741108e9ac7b0756d8b44777e7f1e

Observation a4f43710-7e79-4477-b3fc-93cce2bdb6cb · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.411261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.886082Z digest=sha256:fb50a42731ac52e6fa1c072bd090f3d4e8516f6848da3ed37ed5628ebd61faae

Observation 46ef9fa6-b41a-465a-9401-7e7a96969847 · outbound

This paper cites Attention is all you need in speech separation.

Emotional Face-to-Speech Attention is all you need in speech separation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.396768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.890897Z digest=sha256:ce2ecbd850aad2b43ca573734fd86459df9bca8716e2016a6b01d7a08f540598

Observation 40747989-37d4-44d9-9ee2-83348a6edb91 · outbound

This paper cites Score-based continuous-time discrete diffusion models.

Emotional Face-to-Speech Score-based continuous-time discrete diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.382407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.895228Z digest=sha256:9de806de2981054024bba2c4d2f9834f9a0583b8cfb85531e6090aa108550e83

Observation 67357277-30e9-4a4e-867e-96a6c819c725 · outbound

This paper cites and Fostick, L.

Emotional Face-to-Speech and Fostick, L

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.367580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.899637Z digest=sha256:2eefacf1a1230b0f39042db03108442c8b97f865d59bff1611e638f5b98c5d3c

Observation 47f5461b-95aa-4381-a1bf-dba0f1d2c3ea · outbound

This paper cites and Hinton, G.

Emotional Face-to-Speech and Hinton, G

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.352130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.904043Z digest=sha256:6be3b1f9363f6d38d82affa126e9e6672604c6063083c41cd9655cf445d6d39f

Observation 51fafd75-1f5b-4a6c-bde3-d68fa5123445 · outbound

This paper cites and Vanathi, P.

Emotional Face-to-Speech and Vanathi, P

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.336646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.908535Z digest=sha256:ad59405d2a01f8093ad42c51626e027558ffac8eddfea8185d0e693b3c470495

Observation 47d97212-12bb-42ba-ba0d-5eb6f0572c2b · outbound

This paper cites N., Kaiser, L., and Polosukhin, I.

Emotional Face-to-Speech N., Kaiser, L., and Polosukhin, I

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.319986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.912912Z digest=sha256:c67a49585d72cbba89c1ca027bef8b9ba03ebbe9d95812af781a780793e49641

Observation 505ab147-c5c9-486f-80ae-edd969806099 · outbound

This paper cites Generalized end-to-end loss for speaker verification.

Emotional Face-to-Speech Generalized end-to-end loss for speaker verification

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.304950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.917352Z digest=sha256:eb563c6eca42e64cc3322c8403631bec9ee8aeb93ba24ae774c7ca361ca861cf

Observation 7215bd37-6571-45cc-9288-b4aa56e3d42f · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Emotional Face-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.921870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.921870Z digest=sha256:d9e2cbe6e02ba423cb49edca7fe6277cc627e9cfcc1798ab81319d7b76d4c0bd

Observation f604beb7-c73f-410f-b68d-c28c69d65f01 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.290346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.926776Z digest=sha256:13a9cd358d6282d98affc2b4e6202615db81db97a3210cef4532f78b0c38e50c

Observation 7424b5fb-18ff-4f5e-b46e-3ec3804fc604 · outbound

This paper cites J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R.

Emotional Face-to-Speech J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.274761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.931175Z digest=sha256:e18e1b9eef1ee7b4db6f1ed4840d7d3968c9151b4a0e5b2e9daddb8baea0312f

Observation 637cb4a9-52c4-4d75-abe9-81f326f42521 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Emotional Face-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.935681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.935681Z digest=sha256:31e4f85ff5ea3039181d2691ddea29693155e94fce1df467c04dd15ac060d0cc

Observation 0137943d-8457-47dd-bd86-8f5564f3e4b2 · outbound

This paper cites DCTTS: discrete diffusion model with contrastive learning for text-to-speech generation.

Emotional Face-to-Speech DCTTS: discrete diffusion model with contrastive learning for text-to-speech generation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.258792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.940309Z digest=sha256:6421201cf8bb82d9b40be253c24cd160f7a3efec556342158a3c9cb1cbe08660

Observation b1310808-0727-4144-83c7-132e9c8ab306 · outbound

This paper cites FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model.

Emotional Face-to-Speech FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:50:44.007474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.944733Z digest=sha256:eb92324adcbfe856a58e8bdd96c2130f20ca1e20cb65ff3addf6e1c13cf8659f

Observation a275ec1a-d076-4ea0-ba38-02192b2dc34d · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

Emotional Face-to-Speech Diffsound: Discrete diffusion model for text-to-sound generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.949596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.949596Z digest=sha256:31533ae12378342c66e4e40f3ba61fd66451f2e8f2478b74a22fd350acea37ae

Observation 7af2d74f-5042-49d0-aaad-507e99131309 · outbound

This paper cites SoundStream : An end-to-end neural audio codec.

Emotional Face-to-Speech SoundStream : An end-to-end neural audio codec

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.232022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.953942Z digest=sha256:ee7874b60b057817ac22d04c8954f6839d943f71946db18c72cba5ec9807a210

Observation 1f177877-1668-4d84-9463-adb26d58ab72 · outbound

This paper cites SpeechTokenizer : Unified speech tokenizer for speech language models.

Emotional Face-to-Speech SpeechTokenizer : Unified speech tokenizer for speech language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.216139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.958341Z digest=sha256:5023a72a64c62ad36c97f271fd628effef984724230a51074b4d1b5aad04f5d7

Observation 1302bf96-acf4-4fef-90e8-34fb688272a8 · outbound

This paper cites Srcodec: Split -residual vector quantization for neural speech codec.

Emotional Face-to-Speech Srcodec: Split -residual vector quantization for neural speech codec

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.200703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.962905Z digest=sha256:f399348a4d71df129dc781395a476014baae5352fe5968ac8db73b594a865250

Pith citing papers

Observation 3c38875d-3e8f-4ddd-a3e0-d4bbc00889c8 · inbound

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing cites this paper.

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing Emotional Face-to-Speech

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T17:27:30.884916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:27:30.884916Z digest=sha256:97e100ae71147cbc84c04ee62807055c154c1e775d50e838d8ef322ba310482c

Observation 74f92805-58e2-422e-aece-09e6dec2d8ad · inbound

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing cites this paper.

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing Emotional Face-to-Speech

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:03.314503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:03.314503Z digest=sha256:c42c0683d79050e3c7898f4fc5d932ad30cb48415c0ffb78967facae3cc9fc49

Observation 107ff3ea-8160-4e01-853b-32bfeccd21b6 · inbound

Archon: A Unified Multimodal Model for Holistic Digital Human Generation cites this paper.

Archon: A Unified Multimodal Model for Holistic Digital Human Generation Emotional Face-to-Speech

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.784053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T08:03:04.294439Z digest=sha256:a9815c56948441340a71425fe941bdbea43431a5b7f4f9f2f04342f06fc277c2

Observation cc278dfb-24cf-454f-8d8b-6e0842a3be83 · inbound

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model cites this paper.

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model Emotional Face-to-Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T22:26:14.531002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:26:14.531002Z digest=sha256:08ef1b5aa4b7c8ec04714f217073e87cf8c22c835b13c9954348de888f882809