Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Control of Emotion Rendering in Speech Synthesis

As of 22 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2412.12498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12498 v3

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:06:08.843002Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:32:24.055265Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:21:12.229573Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy62
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08d3e8ac-5307-4189-8835-5b7965c3d74e · outbound

This paper cites An overview of affective speech synthesis and conversion in the deep learning era,.

Hierarchical Control of Emotion Rendering in Speech Synthesis An overview of affective speech synthesis and conversion in the deep learning era,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.529362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.529362Z digest=sha256:195e9b42522233cc23c8e3f309043bce587d401ecbc87425bbce2075578656f6

Observation 7ea42af2-61d1-49a6-a472-6ccc45c4f23b · outbound

This paper cites The age of artificial emotional intelligence,.

Hierarchical Control of Emotion Rendering in Speech Synthesis The age of artificial emotional intelligence,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.534138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.534138Z digest=sha256:4bc38fe907a8f1d89065b806c4efcaa3cde3ea1ad2ee35abdfbe144e2104d113

Observation f5a65bdb-402c-4b67-8377-f3a677b3e779 · outbound

This paper cites Pittermann, A.

Hierarchical Control of Emotion Rendering in Speech Synthesis Pittermann, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.537992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.537992Z digest=sha256:26267162e92259004b4039c8d5ca19a13ac592dad43f9baa367825e1ced65dc0

Observation de745932-fa9b-48d9-8ed3-98cb491f9fe0 · outbound

This paper cites Emotion modelling for speech generation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotion modelling for speech generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.542142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.542142Z digest=sha256:237ed7a0973495b43d6417995983f6f06cef8c5e9cf6aadcd85db4cdae17bcf4

Observation d8e9decb-2849-4676-b019-2d95944d4b8d · outbound

This paper cites An overview of affective speech synthesis and conversion in the deep learning era,.

Hierarchical Control of Emotion Rendering in Speech Synthesis An overview of affective speech synthesis and conversion in the deep learning era,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.545989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.545989Z digest=sha256:0a6f4c8bd907b8c38eed1ba419675ba13ec9c5c4ced2bedec8d644bc9bbe21f2

Observation bbb3fce4-6774-4428-86de-191eef892933 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.554067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.554067Z digest=sha256:694b226470cfd280a73d9c6f49331d4e23e4edc4767deb4ed463ce696ee513f2

Observation 8566ab21-8c13-4f2a-8c06-cf4144af8221 · outbound

This paper cites Phonetic enhanced language modeling for text-to-speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Phonetic enhanced language modeling for text-to-speech synthesis,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.740594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.557786Z digest=sha256:87d8318a41c43e0b6272720c11ad0b2e55c3ee41c00a61f6c7a6688b2f32905e

Observation a7099819-e79b-4cad-a1d3-31fbe54e0631 · outbound

This paper cites Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.729060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.561903Z digest=sha256:19bac35d7033dd23055a8ab376062cdf2f0bcbe474490f1ca8f15811b0ed3097

Observation 2bb81916-c203-42e1-9bbb-fdd053892d84 · outbound

This paper cites Emotional speech synthesis: A review,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotional speech synthesis: A review,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.718028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.565933Z digest=sha256:db1b3d239d7d1c25a236a2b6e85c8a81a44a6010da984a9de5980e648bb15952

Observation 1251d4f1-acd2-460b-87af-466245531d99 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Hierarchical Control of Emotion Rendering in Speech Synthesis A Survey on Neural Speech Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.569719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.569719Z digest=sha256:e74e7d02f8449d640d1df96d16f55acfbf26d5eddc6912569a1ee5c672a082b4

Observation d8e301cf-7f52-4092-9ec3-5cf01d6801e8 · outbound

This paper cites Pragmatics and intonation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Pragmatics and intonation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.705759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.574210Z digest=sha256:9378241a4d05ba4937ba0dcb5d3f0804cebf316c2acd12376db704dd76b13bf8

Observation 72648521-38b0-47c5-afc8-d4598095268d · outbound

This paper cites Perception of affective and linguistic prosody: an ale meta-analysis of neuroimaging studies,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Perception of affective and linguistic prosody: an ale meta-analysis of neuroimaging studies,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.693377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.577953Z digest=sha256:9e2e35a914b1a777b37ca9193e5a91cfd0faa1effba7386f74309d78fb081efe

Observation 0940779d-18ad-40c5-9d9c-a35dededc153 · outbound

This paper cites Laukka,Vocal Communication of Emotion.

Hierarchical Control of Emotion Rendering in Speech Synthesis Laukka,Vocal Communication of Emotion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.581669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.581669Z digest=sha256:8323465fa060c16cdd8eddca573720e0ff6ae66c646179396f1e982c32705056

Observation 553e88ce-f7e0-4a08-8928-aaea88dd6c48 · outbound

This paper cites Vaw-gan for disentanglement and recomposition of emotional elements in speech,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Vaw-gan for disentanglement and recomposition of emotional elements in speech,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.681139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.585131Z digest=sha256:f36afc8306b84a6c63e65ec2e7456c080add749893cd0b5a55484f39ed940ed9

Observation 86230675-45ef-4fb5-aef9-c97c60873202 · outbound

This paper cites Converting anyone’s emotion: Towards speaker-independent emotional voice conversion,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Converting anyone’s emotion: Towards speaker-independent emotional voice conversion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.669949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.588437Z digest=sha256:29e26be2d24aabc86cb0cabbe0074ceac80184c56bc100f13688fedc6348cdcb

Observation 5cf6d6f2-b8ca-4c5c-b2c8-ae8e739f3d18 · outbound

This paper cites opensmile – the munich versatile and fast open-source audio feature extractor,.

Hierarchical Control of Emotion Rendering in Speech Synthesis opensmile – the munich versatile and fast open-source audio feature extractor,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.658229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.591916Z digest=sha256:175e16cedf40780e2d8c96cb148e214d830dd1b562ec809661043bae2a0b5489

Observation af631296-ce47-423f-8982-1fb9d22ec81a · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.647402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.595316Z digest=sha256:5a8b9eb0d0c0a87ef6a9c5db0da7f0a2690192aa8e185975004fc1efe41c1500

Observation f375cf66-d0ae-4c25-bd10-e11b0e70e90c · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.636109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.598849Z digest=sha256:b8a88d1505ddfd682d796b0c910d1c590793b769f2fd11a9c72c75d286ec54ad

Observation 6b3de789-6f1f-46f8-9513-141b08bd1b1d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.624187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.602516Z digest=sha256:fabb645b9c44723c3603df89809032a449abf9cd9d5f21b2973d2b9cd4e1f0f9

Observation 51bc479e-69e0-4c27-9e2d-5146282bcce9 · outbound

This paper cites Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.611419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.606036Z digest=sha256:a5a0e08c3a4740c4407ccbaa1d0a763b6abea4848174412e636f46c648b7377b

Observation b4b56386-85af-455d-b678-41e89379a2d3 · outbound

This paper cites End-to-end emotional speech synthesis using style tokens and semi-supervised training,.

Hierarchical Control of Emotion Rendering in Speech Synthesis End-to-end emotional speech synthesis using style tokens and semi-supervised training,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.599773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.609465Z digest=sha256:ba59c8ba83202be4c67d548eaa0597f97c65d0c496163aafcdcfb9b1d0798454

Observation 9e75c06d-5f43-4436-9fad-9df94ae907c3 · outbound

This paper cites Controllable emotion transfer for end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Controllable emotion transfer for end-to-end speech synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.587506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.613031Z digest=sha256:c3bb675758b32dd378e7a73e4661fe816a59c4303b4fe55fcded73a8bef347ec

Observation fbf80c48-966e-4927-9fe7-771385373f4e · outbound

This paper cites Semi-supervised learning for contin- uous emotional intensity controllable speech synthesis with disentangled representations,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Semi-supervised learning for contin- uous emotional intensity controllable speech synthesis with disentangled representations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.576093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.616553Z digest=sha256:b4428a49a503449fe5adcd27281a8e676f8dffad291da262d8da411f5fe71767

Observation 58483e6a-fdcb-424b-8ea5-20ef5f380240 · outbound

This paper cites iemotts: Toward robust cross-speaker emotion transfer and control for speech synthesis based on disentanglement between prosody and timbre,.

Hierarchical Control of Emotion Rendering in Speech Synthesis iemotts: Toward robust cross-speaker emotion transfer and control for speech synthesis based on disentanglement between prosody and timbre,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.563496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.620007Z digest=sha256:4344bafc779b2eff3d17c0ff8b11c257af27af09c77461fca63e64cac24ad405

Observation df59d7bb-c414-461c-930c-afd84cc6d1fb · outbound

This paper cites Cross-speaker emotion disentangling and transfer for end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Cross-speaker emotion disentangling and transfer for end-to-end speech synthesis,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.551617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.623521Z digest=sha256:13e793e7fd795f08a4f5412b89e85f63e7e8f040db717b00f8ca227e1d5e5738

Observation 1bd9a0a4-7126-4814-98ee-cc77c2016b49 · outbound

This paper cites Relative attributes,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Relative attributes,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.540602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.626913Z digest=sha256:d95d086f66f37a9b8f9e5129c99d1ed4e468aa62bcd940727f42792fbd6228b6

Observation e53eaca7-7203-4ded-ae55-c3c69f3a4bfe · outbound

This paper cites Emotion inten- sity and its control for emotional voice conversion,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotion inten- sity and its control for emotional voice conversion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.529277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.630492Z digest=sha256:b19fea1e6bf0341d35372a07985311da552bf27bf88115699f5751efcbf76a1e

Observation 007f184e-c906-4cc2-8aea-4c71f99e0347 · outbound

This paper cites Controlling emotion strength with relative attribute for end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Controlling emotion strength with relative attribute for end-to-end speech synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.517716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.633927Z digest=sha256:b27b6fc8edc8921b099ac3c0c99b9574d24ec50f88f29dc3512e640a1a41ecca

Observation 62d61b50-b565-476d-9b3f-ee3540482272 · outbound

This paper cites Speech synthe- sis with mixed emotions,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Speech synthe- sis with mixed emotions,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.505413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.637550Z digest=sha256:82f0bf1d4772bed76aec1e2d20ca100c267fd2c7a6d5ed5ec14c697d9fa54596

Observation a7917f73-2356-4f6a-951d-d65bd4af3bd7 · outbound

This paper cites Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.494097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.641090Z digest=sha256:f3ba2f7414f5b48782eecee1db9a8613be161e22eaa8d44fafe6d434d2de8f2a

Observation bd28a430-d85f-4c44-a79f-783e1804141a · outbound

This paper cites Hierarchical emotion prediction and control in text-to-speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Hierarchical emotion prediction and control in text-to-speech synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.482373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.644651Z digest=sha256:01c96c4ff78d0b097df6ad0c520a6d1c435ef82f6e7ccb82ebe2d953ae5a004d

Observation 29abbc49-652d-41bf-8193-ac8a04c500ec · outbound

This paper cites Fine-grained quantitative emotion editing for speech generation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fine-grained quantitative emotion editing for speech generation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.470470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.648316Z digest=sha256:181670150741b152290b7240714d4dee0a5bdac47453d30a03178740f9726ead

Observation c6ec9f19-999d-4bc1-a840-16411c1d93cf · outbound

This paper cites an unresolved cited work.

Hierarchical Control of Emotion Rendering in Speech Synthesis Unresolved cited work

Reference 33

Resolution
verified exact
doi, observed 2026-08-11T14:06:08.897451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.656176Z digest=sha256:0032dec968dfb7fea2f09e2cfc52b499a10c9765f35adb74b1f26702450f13e5

Observation 49a28819-f679-415d-a143-4435751171f9 · outbound

This paper cites Transforming spectrum and prosody for emotional voice conversion with non-parallel training data,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Transforming spectrum and prosody for emotional voice conversion with non-parallel training data,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.458206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.660221Z digest=sha256:52e864987950225a2f34afc21614a193c24567ba13def78a30aee7bf07fe82da

Observation 4e76db84-c4fc-481c-bb83-acf416247de6 · outbound

This paper cites Intonation and emotion: Influence of pitch levels and contour type on creating emotions,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Intonation and emotion: Influence of pitch levels and contour type on creating emotions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.446239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.664011Z digest=sha256:2378bea5e445b3d5f54e7ce286dcc1f1ccd477ebc7de3828296ff73cdf678600

Observation d171ff62-628a-42d7-a24d-56491c56f374 · outbound

This paper cites Norms of valence, arousal, and dominance for 13,915 english lemmas,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Norms of valence, arousal, and dominance for 13,915 english lemmas,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.434523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.667486Z digest=sha256:fd802817e1483f01bb1fc07b2d130a7ec94d88adbd8c21bcef98edb7d6868035

Observation 7a113808-64e3-4cfd-94b5-99c826fbe3c7 · outbound

This paper cites The influence of pitch range, duration, amplitude and spectral features on the interpretation of the rise-fall-rise intonation contour in english,.

Hierarchical Control of Emotion Rendering in Speech Synthesis The influence of pitch range, duration, amplitude and spectral features on the interpretation of the rise-fall-rise intonation contour in english,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.422051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.671081Z digest=sha256:2afb0e0cbc8bd2af9359f04b3891e01e418a6a4150694230a3ff23e36514b61e

Observation 74cbc6ab-b04a-4be1-8957-d538deadb96b · outbound

This paper cites Using prosody to avoid ambiguity: Effects of speaker awareness and referential context,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Using prosody to avoid ambiguity: Effects of speaker awareness and referential context,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.408439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.674691Z digest=sha256:325ade062f48534e9c7d6934ab44a5c327edd6ba4286fb80f69bd7449d59e86a

Observation 846906b9-1916-4720-af85-33c619be4bbb · outbound

This paper cites Analysis of emotionally salient aspects of fundamental frequency for emotion detection,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Analysis of emotionally salient aspects of fundamental frequency for emotion detection,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.396219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.678149Z digest=sha256:287cb21e441d5969df78a68669af6038d7fa2cb74900a978c9c05a0274436da4

Observation 6f37876d-3586-4f80-821e-57263b1036d0 · outbound

This paper cites Survey on speech emotion recognition: Features, classification schemes, and databases,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Survey on speech emotion recognition: Features, classification schemes, and databases,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.384194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.681733Z digest=sha256:209f45771940ca801d18d0ace1bf253d3e99cdd9035f3dce36a8402847671703

Observation 507fc731-cad1-408c-bd43-699b8498c694 · outbound

This paper cites Expression of emotional–motivational connotations with a one-word utterance,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Expression of emotional–motivational connotations with a one-word utterance,

Reference 41

Resolution
verified exact
doi, observed 2026-08-11T14:06:08.885442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.685558Z digest=sha256:bd38f31329d70bba0f1028cc477f3b3962dbee6e84b44ebdc594147dc5adac1f

Observation 04ad7066-1700-451b-93f0-d389f3b42c8b · outbound

This paper cites Controllable accented text- to-speech synthesis with fine and coarse-grained intensity rendering,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Controllable accented text- to-speech synthesis with fine and coarse-grained intensity rendering,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.372004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.689285Z digest=sha256:3c384db6a1508aa5b8d539ecc3d92783d27af8da1c765fe922fc8d91137a4b00

Observation 3a71807e-85b6-4a64-acb3-087cbf41ba31 · outbound

This paper cites Connecting cross-modal rep- resentations for compact and robust multimodal sentiment analysis with sentiment word substitution error,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Connecting cross-modal rep- resentations for compact and robust multimodal sentiment analysis with sentiment word substitution error,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.359568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.693128Z digest=sha256:3cfd62c0f0b28a7b3375edd6831a9377dd511bb34f3f0ce2ef42189aa2cb20bf

Observation 7adf3645-7708-4a2d-9e70-c517a3ce8768 · outbound

This paper cites Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.697089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.697089Z digest=sha256:c88b80a7e801d2f72d61e5474d969ef44556e89a8002f1a105824ab9af41d91e

Observation f055f3e1-1a1f-4e76-8ad5-c13e09ec8813 · outbound

This paper cites Improve emotional speech synthesis quality by learning explicit and im- plicit representations with semi-supervised training.

Hierarchical Control of Emotion Rendering in Speech Synthesis Improve emotional speech synthesis quality by learning explicit and im- plicit representations with semi-supervised training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.339091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.700600Z digest=sha256:6605a89a4a3fbf63a54fce10e06b16c55877a1b81b0580b5252dfd5656274b99

Observation 18323115-3318-40cb-a572-c9f6c843ff6b · outbound

This paper cites Multi-speaker emotional speech synthesis with fine-grained prosody modeling,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Multi-speaker emotional speech synthesis with fine-grained prosody modeling,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.326666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.704111Z digest=sha256:3234b9fb930b74bb94c644d0b907f715abf004e2061474cb48eced576517bcad

Observation 3451055b-47ef-40eb-a121-e6c3f51d4ca1 · outbound

This paper cites Language model-based emotion prediction methods for emotional speech synthesis systems,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Language model-based emotion prediction methods for emotional speech synthesis systems,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.313000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.707638Z digest=sha256:aafb6be458859a7531aa04d10ae3d06e6b7eca220c842f13d39ddc5913429f05

Observation 12a6575f-239c-484a-8120-57e33d627da4 · outbound

This paper cites Fine-grained style modeling, transfer and prediction in text-to-speech synthesis via phone-level content-style disentanglement,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fine-grained style modeling, transfer and prediction in text-to-speech synthesis via phone-level content-style disentanglement,

Reference 48

Resolution
malformed identifier
no resolver link, observed 2026-08-11T14:06:08.711157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.711157Z digest=sha256:6082c5f9d96212087b33d92293e00f472c5d37462c935f37b5793c44f7bc8458

Observation ff31eb97-8b64-437d-9424-ac8e06022f16 · outbound

This paper cites Hierarchical multi-grained generative model for ex- pressive speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Hierarchical multi-grained generative model for ex- pressive speech synthesis,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.300941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.714934Z digest=sha256:ab6bbc9ac3ee79005d5fea78dbc3212abf411a159a7a60bf852a76403f4ef2be

Observation f030e8e2-866b-4215-9605-ade282db58e8 · outbound

This paper cites Towards multi-scale style control for expressive speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Towards multi-scale style control for expressive speech synthesis,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.289100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.718757Z digest=sha256:a56f9bb5ca49579174cd5ad9fb837986e846fbc1ccabc81ee89ac1e4822ffebf

Observation b9119c76-aa80-461b-b5eb-8da500f6d1df · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.276828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.722320Z digest=sha256:2a3a60875d503e2a84c8ef0a3930cc9d9b3d59f8fe24ce4dc8f22c1fd908d434

Observation 415ea64f-bcf3-4a3d-8b6d-6ca1ffc4c713 · outbound

This paper cites An emotion speech synthesis method based on vits,.

Hierarchical Control of Emotion Rendering in Speech Synthesis An emotion speech synthesis method based on vits,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.265049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.725748Z digest=sha256:e8016f0a7dc5861f0e4a33f004a25c0ef54115f9cab7766ac78b186b5d16c13c

Observation 0da5d858-9530-45f4-9673-0adb61c6a828 · outbound

This paper cites Generative adversarial networks,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Generative adversarial networks,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.729316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.729316Z digest=sha256:5b786b8a52c529a8272f1a8f91505cab0ec884548a41fa3a8e13c79390f36544

Observation 890f59db-3461-4d35-81dc-43fdaa98605c · outbound

This paper cites FastDiff 2: Revisiting and incorporating GANs and diffusion models in high-fidelity speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis FastDiff 2: Revisiting and incorporating GANs and diffusion models in high-fidelity speech synthesis,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.246376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.732675Z digest=sha256:b8e0a160f6abe624171c365a3c991212863aea0e5c86039580464bd61710eeb1

Observation 04b22edc-df59-45d9-9737-0b869f81c49e · outbound

This paper cites Bddm: Bilateral denoising diffusion models for fast and high-quality speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Bddm: Bilateral denoising diffusion models for fast and high-quality speech synthesis,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.234257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.736092Z digest=sha256:e8c179377cb80bcb5ce87c4aebcef4c9499157f670e44e47601b801722b10601

Observation 96ff01ce-bd93-44f2-8a9a-31ac9b5725e2 · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.221999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.739604Z digest=sha256:d29b026c7be8875fa10cd19e7d233d8ab6c791ab1166d85ff979518398bfb90d

Observation 91f90457-53fd-40e5-b56e-c505c5bab4e0 · outbound

This paper cites Flow matching for generative modeling,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Flow matching for generative modeling,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.743172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.743172Z digest=sha256:df9029c15b3bd29d27b34a28bf75d310b3dfcb52e9f7dc014536b0aa9827ad25

Observation e82fc72d-6b5d-42c9-b002-1ce0533b03ba · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.190533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.752366Z digest=sha256:be88cb41ba9c19a6795721422362e4c66631b3eebf6d5da4e5eaecd2ef4b17de

Observation e6c0991c-45fc-4f31-b163-f22bef0a4add · outbound

This paper cites Emotional voice conversion: Theory, databases and esd,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotional voice conversion: Theory, databases and esd,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.179120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.755963Z digest=sha256:29e88df7d3cc979793bbac93d99a48ceba03f4c289660f02345ff70b01681d05

Observation d33e773c-9774-419f-8da0-05ba17a9d5b9 · outbound

This paper cites Attention is all you need,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Attention is all you need,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.759549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.759549Z digest=sha256:82eb31681dc24f17255861c965c27e218985b81b510f1439e3892157f5f4afe2

Observation 813f499e-a0c9-4543-b437-e51b04cbfddc · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Glow-tts: A generative flow for text-to-speech via monotonic alignment search,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.160130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.763267Z digest=sha256:6e16847d9ba1bfe698429d7b6f265e4a8f8a8ce84dcd66457fe0786b29930f8b

Observation dba3efa7-c20e-480c-ba76-366bc30dde28 · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Generalized end-to-end loss for speaker verification,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.202501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.766815Z digest=sha256:0348e35f8041826487aa53c12290ee99d3627fe46f90b0b0563b2f922fa02737

Observation 132e3ca8-96e5-4561-a8a9-7fec55ef931e · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.148508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.770234Z digest=sha256:7597edc6c33efaf335262c92d055464f932628fa072dbe21cd3f099f5b14f00a

Observation f9939591-14e3-46db-9ec5-f21518a35f09 · outbound

This paper cites Adam: A method for stochastic optimization,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Adam: A method for stochastic optimization,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.773936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.773936Z digest=sha256:da9f901d22e28df0ac28159dc6f8fb683c0c88cd055a7f31e1bbfba06c950448

Observation 753e61bd-483a-4493-8e97-ca4741da447b · outbound

This paper cites Unsupervised domain adaptation by backpropagation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Unsupervised domain adaptation by backpropagation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.777748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.777748Z digest=sha256:7409afdfa8320ffacd096b75fa01990e10903066f07f11b1f86e7b74f4387c6e

Observation a3b2ab0a-2312-4a12-9903-23778bf20223 · outbound

This paper cites Adversarial domain generalized transformer for cross-corpus speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Adversarial domain generalized transformer for cross-corpus speech emotion recognition,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.123636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.781281Z digest=sha256:5d838e130cb6b349ec82241b613d5ad9edfdb002754c9c6ea0471eded5a81340

Observation f68cf9c0-3cba-4c10-b9e7-08500f99dc13 · outbound

This paper cites Mel-cepstral distance measure for objective speech qual- ity assessment,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Mel-cepstral distance measure for objective speech qual- ity assessment,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.111777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.785019Z digest=sha256:549b0ed62557a591ff1c93843b2d6be2023a80c23ad65084dcaadd17d80c5b60

Observation ea67377c-8dd6-4316-b004-1b36882f95a1 · outbound

This paper cites 3-d convolutional recurrent neural networks with attention model for speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis 3-d convolutional recurrent neural networks with attention model for speech emotion recognition,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.099517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.788524Z digest=sha256:0a97bbf8969357e376d434fe6a7d41e8a887b4021a5ef87daa919ede83431cb0

Observation c64804da-90f8-4de1-acfa-c3e2a175ff24 · outbound

This paper cites Multi-conditioning and data augmentation using generative noise model for speech emotion recognition in noisy conditions,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Multi-conditioning and data augmentation using generative noise model for speech emotion recognition in noisy conditions,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.792294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.792294Z digest=sha256:c6b55d27a60a463f6ca96f1783e91f3361d0fe12fea46a484fa6b9324d5efe3b

Observation 27af8343-f844-40ba-8bd9-e22be958a71a · outbound

This paper cites Speech emotion recognition in noisy and reverberant environments,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Speech emotion recognition in noisy and reverberant environments,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.080298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.795971Z digest=sha256:5a19220e429c76a2e14a08858e0b87bf296f15491515efda891e8449a8fe2e8b

Observation 736ee137-ad4e-43ef-9ad9-991c1a8b1e36 · outbound

This paper cites Deep learning techniques for speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Deep learning techniques for speech emotion recognition,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.068317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.799554Z digest=sha256:6432d2f7d6da3f7cee1c7cd5a2c799f1ef546dfafc1a987f1fcbdbf03930df7a

Observation a122c003-a410-4541-8099-139f18cbf217 · outbound

This paper cites Improved emotion recognition using gaussian mixture model and extreme learning machine in speech and glottal signals,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Improved emotion recognition using gaussian mixture model and extreme learning machine in speech and glottal signals,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.055938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.803083Z digest=sha256:7cd8c806b3dcd0e71377ee9b0190c3aba45d949579ced00649dc75a50b4109c3

Observation 82bef523-36d6-4aa2-824a-553e97ed47a7 · outbound

This paper cites Specaugment: A simple data augmentation method for automatic speech recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Specaugment: A simple data augmentation method for automatic speech recognition,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.806926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.806926Z digest=sha256:43a8cdc69f31fa974c0c0aaacc9f59f8215c77a4c980e0535b5bcab25a24587b

Observation 75740fd4-f651-4a51-b9ef-0a835da7b6e3 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.036690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.810492Z digest=sha256:68436e45046741ec0df2ae463b69ce53744e930a2fc943146b3413a26f567a42

Observation a6217b18-5345-455d-99fe-21903b1531c3 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Hierarchical Control of Emotion Rendering in Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.814123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.814123Z digest=sha256:bd86ddf3411a3e72c3de2d2bc559b2dd8fb7271c2988a0f81885bf476d491b5c

Observation 8fa8ec19-2b99-47ff-a705-448931d98a54 · outbound

This paper cites Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.024616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.818127Z digest=sha256:ad8add59045c783140d6310ee0fd064fdb07ea70d659ae6e77b96bdc5b619cb7

Observation bdd7e523-0143-4a73-8f5d-990a8eaa6e35 · outbound

This paper cites Measuring disentanglement: A review of metrics,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Measuring disentanglement: A review of metrics,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.013442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.821706Z digest=sha256:30d1db5fd01595a5947bbeef3b51080fd297c758a32f98c133b3b96ffe85d7bf

Observation e5f835b3-95da-444b-81f4-e42c7df6c66e · outbound

This paper cites Isolating sources of disentanglement in vaes,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Isolating sources of disentanglement in vaes,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.001787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.825361Z digest=sha256:8651b92c133a13d78a5e8c2a35f9eb69210b2ac8b73f640e5f73ef629008808b

Observation a51c0188-1dce-41ab-9916-5432d91f5b93 · outbound

This paper cites A framework for the quantitative evaluation of disentangled representations,.

Hierarchical Control of Emotion Rendering in Speech Synthesis A framework for the quantitative evaluation of disentangled representations,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:08.989499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.829068Z digest=sha256:ec39d83d93e5bd3cfcce7a79de6f4a6721062a511726a43deb6106618a9da9e2

Observation aae9dfeb-e6b1-4f37-8c91-5e959dccd5d8 · outbound

This paper cites Learning deep disentangled embeddings with the f-statistic loss,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Learning deep disentangled embeddings with the f-statistic loss,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:08.977717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.832368Z digest=sha256:0ae1c0a05d57b4bc7f709831ac57eaaf5abbb61d016064afa9f5b0fb5e13fd19

Observation 2007afa7-cd4a-4a23-8ff2-adc7d7eea17f · outbound

This paper cites Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.835917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.835917Z digest=sha256:2398771c7588f2c156eb9145ddd03dd850c9887cebe0a93676cfd232f883ae02

Observation 7058b55d-b386-4746-87bc-7f3790192bbd · outbound

This paper cites Dawn of the transformer era in speech emotion recognition: Closing the valence gap,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Dawn of the transformer era in speech emotion recognition: Closing the valence gap,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.839539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.839539Z digest=sha256:1b5eaa4402e401242077f27a907f6a2688d55c68642a5e7b23a5ba3e4423d8ee

Observation 77a008ac-5af3-470c-9ab4-4fc7747fd84a · outbound

This paper cites Contrastive learn- ing based modality-invariant feature acquisition for robust multimodal emotion recognition with missing modalities,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Contrastive learn- ing based modality-invariant feature acquisition for robust multimodal emotion recognition with missing modalities,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:08.951996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.843002Z digest=sha256:ea6215e97a802eebefc553a8349a239d0773d250d083d861de5fbca1d11dd4da

Observation ebcb22a9-6a66-4103-b963-4c4c2beeed13 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 252762216.

Hierarchical Control of Emotion Rendering in Speech Synthesis Available: https://api.semanticscholar.org/CorpusID: 252762216

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.759306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.550085Z digest=sha256:64669c267f0c4f4a4e28ff400eadef6eeeaac105379f41b3ffd962b79bda5f12

Observation de2b0f40-376a-4d8e-8abf-e3ef02beb990 · outbound

This paper cites Fine-Grained Quantitative Emotion Editing for Speech Generation.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fine-Grained Quantitative Emotion Editing for Speech Generation

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:06:08.929237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:06:08.652070Z digest=sha256:3260272e30a31c1e7e7b99e83b7b199a7ed66d8c484e1afa70efd5958887e8f5

Pith citing papers

Observation 99ae3806-5571-4a12-a529-2c9c05b018ec · inbound

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey cites this paper.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Hierarchical Control of Emotion Rendering in Speech Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.055265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.055265Z digest=sha256:b348eac28b41f2d1cdf2b3abbf8e9cb35ad8eb11087c730f44baefc54668300f

Observation 4473eeaa-47f2-4a02-bada-4098ec1aec00 · inbound

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions cites this paper.

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions Hierarchical Control of Emotion Rendering in Speech Synthesis

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:21:12.319458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:21:09.698405Z digest=sha256:7d1d69ad15dbf18cf738f21c5e37fa2ef2c51defb30dae888fc79812ce66d77a