Pith. sign in

Paper Citation Record · LEDGER

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2507.04349.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04349 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:55:03.611900Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T15:21:28.820442Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:11:27.121428Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact4
  • verified fuzzy17
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5764825-324b-4fac-a39b-750df4887754 · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Adding Conditional Control to Text-to-Image Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.241628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.241628Z digest=sha256:57a93f5b48a740d43fc83cbb842b097a51e9f2762e3639e31db18ca29770f12b

Observation e546cec3-0c0f-4291-8e58-d3284f1384f9 · outbound

This paper cites An emotion speech synthesis method based on VITS.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet An emotion speech synthesis method based on VITS

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.445789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:02.318699Z digest=sha256:ba8f9f2a2d0923044202699b56774d76f0c146fe5b390f9e0a57a7319e1b466c

Observation 208baeef-84d1-4e7c-9ac4-f08ed652bcc1 · outbound

This paper cites V oiceBox: Text-guided multilingual universal speech generation at scale.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet V oiceBox: Text-guided multilingual universal speech generation at scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.309106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:02.438700Z digest=sha256:5d124bd56c3e635e4b80f6f66b80ee92fccbab7340cb88f7e4c4864456010c29

Observation c7f1d86e-2076-4a00-a0eb-69b5d39e2ef9 · outbound

This paper cites Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.574535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.574535Z digest=sha256:3b909739c65bb6f3f3cd2b1e38655dc1532cd9617c602852a71fe985f245bc53

Observation 44608682-5498-460d-a2ca-1803573f58dd · outbound

This paper cites E2 tts: Embarrass- ingly easy fully non-autoregressive zero-shot tts.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet E2 tts: Embarrass- ingly easy fully non-autoregressive zero-shot tts

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.144160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:02.684793Z digest=sha256:3509a723305e036d92afac9d0a67a25b08bce9c9302b02de8a6b66934abf9f0c

Observation 2c96ced3-fec9-4a79-bac3-b9a38325340c · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.966453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.966453Z digest=sha256:88a16566e134e6886731557e270b071fc1a802eb3bb5cace529a4bb0b1655607

Observation 29734753-f174-4dcb-b602-7a45227fba33 · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Tacotron: Towards End-to-End Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.102660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.102660Z digest=sha256:7a6446ba72bfeda2ca2d373a3033bc0b58b5211d144bc8cc81b8fadb840e90e6

Observation 9f6c56d8-d55d-4ea0-acde-d94a900578f6 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Fastspeech: Fast, robust and controllable text to speech

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.137493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.137493Z digest=sha256:99e0a768647c7ec63ac6dc4f2df776ab8aa04f76bc07fe101c3064315f7bf39a

Observation 300d42a3-b2bf-4d08-a03f-d3579d2f27cf · outbound

This paper cites ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.233086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.233086Z digest=sha256:8481dcb80a1cdaaae0483df97e0c915f2501b2017b0b3933e4b06f80aa8d3ae5

Observation b1df3dc2-cdee-42ec-a9e5-05ede80647d4 · outbound

This paper cites EmoDiff: Intensity controllable emotional text-to-speech with soft-label guidance.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet EmoDiff: Intensity controllable emotional text-to-speech with soft-label guidance

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.086938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.346441Z digest=sha256:72ce90daa2ffce2fd7d84ac9bc78874b13ecc1d2c4cd6df55610c37dfbd52c18

Observation aa2deb88-4ce8-4e88-be3e-a9407cfa5235 · outbound

This paper cites EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.419208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.419208Z digest=sha256:1d1cb9f82220153f5298ec3c79e307cd4cfc3d7710962e8f27f8a71f71b3fd54

Observation 366e1ecd-8c4f-4c60-a42c-bed69eb75c84 · outbound

This paper cites Speech synthesis with mixed emotions.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Speech synthesis with mixed emotions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.076900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.483681Z digest=sha256:dc93b5936cfef8e4844fbf2e62af1b79d8123e9f795be81aadeefcf432a68e20

Observation 26020b9b-e846-477d-b112-66bafcc54e32 · outbound

This paper cites QI-TTS: Questioning intonation control for emotional speech synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet QI-TTS: Questioning intonation control for emotional speech synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.064363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.518055Z digest=sha256:236631a33bf6e74bd7a948b851facad5b8be244354714b43d4ac643283c4c94a

Observation abe66048-9699-43e9-95d5-3f7a7d3ee0f5 · outbound

This paper cites MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.053177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.523695Z digest=sha256:9c8098b3720381649bf5c04623a89ec5f61b039f98230c36e59c123c3e750225

Observation ca31ec89-b39c-4d57-84fa-e2ae6f8c2ab1 · outbound

This paper cites Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:03.856194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.526791Z digest=sha256:dc8390b541e7aa8550dc6283b6217917b9ee010bf0a3771070a72170dc627dae

Observation 18355188-c91f-46cf-89a8-32016829c550 · outbound

This paper cites Emotional End-to-End Neural Speech Synthesizer.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emotional End-to-End Neural Speech Synthesizer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.529499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.529499Z digest=sha256:ad1f7e270dc0adf21838dd0fa99ded251c6267cfc631c8117785a6db5cec2d47

Observation 6a73c16a-4189-4908-8f33-4f0f697b58cc · outbound

This paper cites Controllable emotion transfer for end-to-end speech synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Controllable emotion transfer for end-to-end speech synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.043446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.532021Z digest=sha256:0ac3ae4d27a49e47444f8ee82b875d8fc263a202785d954865ecff149626b558

Observation 9b569e5c-3eaa-4f8d-be37-b62c236fe297 · outbound

This paper cites Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.032509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.534804Z digest=sha256:2cd57d8056b0d98552d6fba639cb9f3441b41ff302f26fe36493043f3d672878

Observation cd5dac71-3209-466d-bf22-a250b1ae6ad1 · outbound

This paper cites Emosphere- tts: Emotional style and intensity modeling via spherical emotion vector for controllable emotional text-to-speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emosphere- tts: Emotional style and intensity modeling via spherical emotion vector for controllable emotional text-to-speech

Reference 20

Resolution
verified exact
doi, observed 2026-08-06T19:55:03.658196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.537283Z digest=sha256:cb5a672f6675cda9a43380fa9dc2d9abab73f59128b1131039cf790e59311f61

Observation 24f69278-2478-400a-b62a-f7038cf8050b · outbound

This paper cites Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:03.834991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.540269Z digest=sha256:b86b3f4281e3835cab8e9e3c573b7c12255a1aebf2ce6da69aece8007f068b41

Observation c28b31f1-ae99-4660-a07b-f6aba23591c0 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.543011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.543011Z digest=sha256:5fac4c331ed311323b363d26f0d89c7b32e6f2822edf2079fd7151214b83fb2e

Observation ae120038-da12-4821-b592-095af89548a3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet High-resolution image synthesis with latent diffusion models, 2021

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.545404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.545404Z digest=sha256:5249c0b43ae85bdbd2356a7e10930f2ca32715a033179848666d4caa3777508b

Observation db240b9d-a81b-4cce-b5cb-3a8760823976 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.548339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.548339Z digest=sha256:39e0d69169519b0e4ae9fa604a0f83956f440cd737a01c437e5c7ec6757e5d39

Observation 74ab91bd-8e53-4437-8b1b-443cd01ffbdb · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet High-resolution image synthesis with latent diffusion models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.553864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.553864Z digest=sha256:e1d70c39819f66d147b1bce52b98c9029bac19d43f7bc2959252260d074627a1

Observation 3aba256c-3a5c-46b0-bc54-58c020af47e5 · outbound

This paper cites Scalable Diffusion Models with Transformers.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Scalable Diffusion Models with Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.556597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.556597Z digest=sha256:60a21074a9e406f8d392d8f0cbca781bfe10d6940a289db7353e44e5b2e14780

Observation dd543283-742f-407e-8f66-e0181482b3e6 · outbound

This paper cites an unresolved cited work.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.558680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.558680Z digest=sha256:44e7f79025b409c79cc70a64da491f838ad99987cf3f6e80fc55aa7c5edfd157

Observation 4722960f-6b47-4504-8095-59584de4ba11 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.561139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.561139Z digest=sha256:0b540dc131fbf6f781054763e7332acfed1f6b2264a83c8fed1872e1a6443de7

Observation b47965bb-61e1-41ba-9435-633e28f772fd · outbound

This paper cites Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.563909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.563909Z digest=sha256:676f5e910b64e2d9c5f399f7af624aab885a2deacfbf4348d1310d712e1e941f

Observation aca95d79-9e1d-4b04-a981-57869bda2470 · outbound

This paper cites Stable Flow: Vital Layers for Training-Free Image Editing.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Stable Flow: Vital Layers for Training-Free Image Editing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.566512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.566512Z digest=sha256:08e2e444d49a720d598aaec0d7295c3a819fc648a900f145df4faf573f92466f

Observation 1cfc32c3-f63a-4151-ae45-23d83dec5180 · outbound

This paper cites an unresolved cited work.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:55:04.003235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.569359Z digest=sha256:22ed52d270d5dab6195d5272ace5fc15c383405e6a67abe6d688c2337ffd0594

Observation 7c0a4d10-0711-41e7-9bcb-e6d557b03047 · outbound

This paper cites Jointly predicting arousal, valence and dominance with multi-task learning.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Jointly predicting arousal, valence and dominance with multi-task learning

Reference 33

Resolution
verified exact
doi, observed 2026-08-06T19:55:03.647994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.572713Z digest=sha256:7c7b6b1b0fee909fb7eaa2c36458c3bd385c9215ed5188d75aad18a27bb122b4

Observation c0bef902-9eaf-4a1a-93a1-9fae01b79ccf · outbound

This paper cites Automatic speech emotion recognition using recurrent neural networks with local attention.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Automatic speech emotion recognition using recurrent neural networks with local attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.575345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.575345Z digest=sha256:e1f0a3c0be885e4cbe789313b511be10b70c9b5cc9d7966c2fd131dc53141c35

Observation 57d46f65-0a4c-4e40-8a49-541b498ce322 · outbound

This paper cites DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.577522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.577522Z digest=sha256:56e0af387980f2685c7c652b553c5cc982c65dd48decb018ce964fffbfdcef35

Observation 1c380437-f813-455d-95bc-6a23c1569fa7 · outbound

This paper cites Dawn of the transformer era in speech emotion recognition: Closing the valence gap.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Dawn of the transformer era in speech emotion recognition: Closing the valence gap

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.991773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.580775Z digest=sha256:1c63d6fc19ce532a7d5f07edfad7556a85b886f72e4b1cffd58df6dac3a59b7d

Observation f0c04121-3ed6-4d39-b2ef-27e9337b95c5 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.583254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.583254Z digest=sha256:ca4b2c4a252e61bb4b9058c6a8081b4f74791926f082520c96e3ca0e16bd2acb

Observation 2a13b751-60f8-42e4-b698-360ce9de2ead · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.585947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.585947Z digest=sha256:a3da2632f1145cf3f1b1f88c95f7a41d6f3069d6e3ba0e89f18c6ae46415000e

Observation 9b0e4c01-87bd-4f7a-b487-cfea8791fdd4 · outbound

This paper cites Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.975384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.589132Z digest=sha256:11a166731075b7316b436f3e0bec688aa789c904740237980d78c9f21fd6d5d7

Observation 76e9d925-af9f-4cdc-a122-cb4e78110d5a · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.964072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.591411Z digest=sha256:466063920ec73586be13a451ba62e59dae8c0e5ed65971582930426632423487

Observation 04826e61-70a3-40b4-9541-418907ce9bcd · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Iemocap: Interactive emotional dyadic motion capture database

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.952788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.593862Z digest=sha256:66c2133431bed015494266f0f1edd37af42434a5136e67800b47ca2ffa1902b8

Observation 786b6ee7-f2df-484f-adf6-d87af02a949e · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.942258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.596052Z digest=sha256:c9494dc5e3753e7f34a4e036031c561d8d99f092feff1999b5abfffa235dc33e

Observation 50fd03a3-26bf-4b1e-a145-3c47d6137f59 · outbound

This paper cites Expresso: A benchmark and analysis of discrete expressive speech resynthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Expresso: A benchmark and analysis of discrete expressive speech resynthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.930958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.598786Z digest=sha256:40847ac38166700a55e28ef22a14f8b8c859dcccd248e34dee46a3fc610d3c4a

Observation d39eaf97-f758-4234-9303-b0fcedb68c25 · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet The ryerson audio-visual database of emotional speech and song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.918607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.603813Z digest=sha256:c90a82db2518bce07bf338a8949fbd448f249c55223e2b2304cce203a606ad60

Observation 57707053-24ee-4ccf-b3fa-4bc411d1448c · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.908017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:03.606423Z digest=sha256:9cd388833eca5147c3aab81891a6ed16e7f58fa9a98141a47dba569234e41787

Observation 39cf408f-2267-471a-813e-88f5f227bf17 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.608406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.608406Z digest=sha256:5554518402c85a3a66449dfa17da7ec13f10e67c469a1c6571612b0ddcd193c6

Observation ce67d789-761f-4502-8d41-2ab1bd58d570 · outbound

This paper cites supple_demo/index.html.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet supple_demo/index.html

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.611900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.611900Z digest=sha256:4653315b2bd69519bfce5336ec2b47c0fafafaca554d5a52367ff1024c50450f

Observation 94723fb0-b9a0-4101-abd0-d7076b1748b1 · outbound

This paper cites an unresolved cited work.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.600978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.600978Z digest=sha256:38bfe72318a69def7e879455ad2143a93b7c84e7dc40526ede180df470bcf0f4

Pith citing papers

Observation 440da532-b107-4cfc-b485-61e310bb88d9 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.124142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:0b5bf6f5afdf020f02370aeb3aff6b991a7b1e159cb3ab1135dd3731f91accbf

Observation a0b54a62-c850-489d-8430-cbcb74ef4ad6 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.820442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.820442Z digest=sha256:0b12c39c5518b9aa98bcea310ef0245d0840bcb768a4f9aacb4f3ae46648c22b