REVIEW 3 major objections 5 minor 67 references
Emotional voice conversion can be driven by relative natural-language instructions that say how the source should change, not by fixed target emotions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 00:49 UTC pith:REMLAD4E
load-bearing objection Solid new relative-instruction EVC task and source-anchored flow; competitive numbers, with the main soft spot being LLM-instruction fidelity rather than a broken method. the 3 major comments →
TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that free-form natural-language instructions describing how a source utterance's affect should change can replace target emotion labels and reference speech. TRACE-EVC, via Emo-Compass, models each conversion as a source-anchored rectified flow in a continuous affective space and predicts the displacement induced by the instruction, enabling zero-shot relative control across categorical transitions, intensity modifications, and open-ended changes while remaining competitive with conventional EVC systems.
What carries the argument
Emo-Compass: a source-anchored rectified flow that starts from the source emotion embedding and predicts a velocity equal to the affective displacement the instruction describes, rather than generating a target embedding from noise.
Load-bearing premise
The whole pipeline assumes that language-model instructions built from source and target style descriptions plus measured affective differences truly name the emotional change, without leaking the target style or inventing artifacts.
What would settle it
If human instruction-following scores stay low when the same source is paired with different relative instructions, or if shuffling instructions across test pairs leaves emotion similarity and conversion quality unchanged, the claim that Emo-Compass follows relative instructions would be false.
If this is right
- Users can steer emotional speech with relative phrases instead of picking discrete emotion labels or supplying reference clips.
- Intensity and open-ended affective edits become usable without forcing the user to supply a target intensity level or VAD coordinate.
- Unseen speakers can be converted zero-shot because speaker identity is taken from the source and emotion is a continuous embedding.
- The same framework still supports ordinary categorical emotion conversion and remains competitive with label- and reference-based systems.
- TRACE-Instruct supplies paired relative instructions for training and evaluating inter-emotion, intensity, and open-ended transformations.
Where Pith is reading between the lines
- Relative-instruction control could transfer to other speech attributes—prosody, speaking rate, formality—where users naturally think in deltas rather than absolute targets.
- The same LLM-plus-difference recipe used for TRACE-Instruct could build relative-instruction data for other expressive modalities without hand-labeling every transition.
- If SER-derived affective differences or residual target leakage dominate the instructions, instruction-following ratings and objective emotion metrics will systematically diverge under open-ended prompts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces instruction-guided relative emotional voice conversion (EVC), in which free-form natural-language instructions specify source-conditioned affective changes (e.g., “slightly calmer,” “more confident”) rather than absolute target labels, references, or style descriptions. To support the task it constructs TRACE-Instruct by prompting an LLM with source/target style captions and SER-derived ΔVAD, then rule-filtering explicit target labels. The proposed TRACE-EVC system centers on Emo-Compass, which models the conversion as a source-anchored rectified flow whose velocity is the displacement v* = z_tgt − z_src in a concatenated emotion2vec–VAD–prosody space, followed by a DurFlex-style synthesis module for zero-shot conversion. Experiments on ESD (inter-emotion) and MEAD (intra-emotion intensity, including unseen speakers) report competitive or better EECS, ACC_cls, SECS, UTMOS/NISQA, and subjective instruction-following scores versus label- and reference-based baselines, with ablations supporting the flow module and multi-cue embedding.
Significance. If the relative-control claim holds, the work is a meaningful step beyond target-oriented EVC: relative instructions match many practical editing scenarios and enable continuous intensity and open-ended affective edits without requiring a target label or reference. The source-anchored rectified-flow formulation is a clean architectural fit for that setting, the TRACE-Instruct construction and planned release of data/code/demos are concrete community contributions, and the system remains competitive on standard categorical conversion while supporting zero-shot speakers on MEAD. The graded arousal-shift analysis (Fig. 3) is a useful falsifiable check of intensity control. These strengths make the paper of clear interest to the expressive-speech community, provided the relative (vs. absolute) nature of the supervision and evaluation is more rigorously established.
major comments (3)
- [Sec. III-A, Eq. (3)] Sec. III-A / Eq. (3): TRACE-Instruct is generated by prompting Qwen3 with both p_src and p_tgt plus SER-derived ΔVAD, then applying only rule-based filters that strip explicit emotion labels, intensity tags, numerical VAD, and metadata. Residual target-style phrasing can still leak into the instruction (e.g., absolute descriptors of the target delivery). Because the paper’s central novelty is relative rather than absolute control, this is load-bearing: without a leakage audit (e.g., human ratings of “relative vs. absolute,” or a held-out set of human-written relative instructions with no access to p_tgt) it is hard to know whether Emo-Compass is learning true source-conditioned displacements or largely recovering target-style cues that remain in the prompt. A small human-written or human-rewritten instruction test set, and/or an ablation that removes p_tgt from the LLM prompt, would subs
- [Tables I, II, IV; Sec. III-B] Tables I, II, IV and the definition of z: emotion2vec appears both inside the emotion embedding z = [e; a; r] that Emo-Compass predicts and as the primary objective emotion metric (EECS, emo acc). VAD L2 and intensity MSE likewise reuse the same continuous cues used in training. This creates mild metric–model dependence: high EECS/emo-acc can partly reflect consistency with the same representation family rather than independently verified affective change. The paper already reports an independent wav2vec2-XLSR ACC_cls on ESD and subjective IF; extending independent SER/VAD evaluators (or a frozen SER never used in training) to MEAD and to the Emo-Compass intrinsic table would make the emotion-conversion evidence more robust.
- [Table III; Sec. IV-A Subjective Evaluation] Table III and the subjective protocol: instruction-following is the paper’s primary new claim, yet IF is reported only as two scalar means (4.135 / 4.296) over 10 utterances × 15 raters, with raters always shown the instruction alongside source and converted audio. There is no comparison condition (e.g., absolute-style prompts, shuffled instructions, or a label-based system forced to follow the same relative request), no inter-rater reliability, and no open-ended / out-of-distribution instruction set. A modest expansion—more items, a shuffled-instruction control, and at least one absolute-prompt baseline—would better isolate whether the system is following relative change rather than producing a generally more expressive or SER-aligned output.
minor comments (5)
- [Table I] Table I includes VEVO trained on 101k h as a “large-scale reference.” Even with the caveat, the row invites unfair comparison on SECS/UTMOS; consider moving it to a separate reference row or appendix so the 11.9 h comparison remains the primary reading.
- [Eq. (6); Sec. IV-A Implementation Details] Eq. (6) and the implementation details leave the number of Euler steps K and the loss weights λ_cos, λ_emo, λ_int, λ_aff unspecified in the main text. Reporting chosen values (and a brief sensitivity note) would aid reproducibility.
- [Sec. II-A; Table V] The paper correctly notes that public PromptEVC-style systems are unavailable and describe absolute rather than relative targets. A short qualitative discussion of how TRACE-Instruct instructions differ from TextrolSpeech-style absolute captions (beyond Table V diversity stats) would still help readers place the contribution.
- [Fig. 2] Fig. 2 is dense; the training-only dashed paths for target audio are easy to miss. A one-sentence callout in the caption that inference never sees target audio would reduce ambiguity.
- [Title; Sec. II-A] Minor wording: “V oice” appears with a space in several headings (e.g., title and Sec. II-A); standardize to “Voice.”
Circularity Check
No circular derivation: supervised source-anchored flow matching with external baselines, independent classifiers, and human ratings.
full rationale
The paper is an empirical systems paper proposing a new task, a constructed instruction dataset, and a model (Emo-Compass as source-anchored rectified flow plus a DurFlex-style synthesizer). Training defines the constant target velocity by construction as v* = z_tgt - z_src (Eq. 5) and regresses the network to it (Eq. 7); this is ordinary supervised flow matching, not a claimed first-principles prediction that reduces to its inputs. At inference the target is unknown and the network integrates from the source under the free-form instruction alone. Evaluation reports audio-level EECS (emotion2vec on waveforms), an independent wav2vec2 ACC_cls, SECS, WER, UTMOS/NISQA, and human nMOS/IF scores against external label-based and reference-based baselines (Tables I–III). Intrinsic embedding metrics in Table IV simply confirm that the trained module recovers held-out targets better than a random permutation; they are not presented as surprising predictions. TRACE-Instruct generation uses an LLM conditioned on style descriptions and SER-derived ΔVAD, but this is data construction, not a circular claim about the model. No self-citation supplies a uniqueness theorem or ansatz that forces the central result, and no fitted scalar is renamed a prediction. The derivation chain is therefore self-contained and non-circular.
Axiom & Free-Parameter Ledger
free parameters (3)
- loss weights λ_cos, λ_emo, λ_int, λ_aff
- number of Euler steps K at inference
- learning rate 1e-4, epochs/steps, Transformer depth 6
axioms (4)
- domain assumption Linear interpolation between source and target emotion embeddings is a valid training trajectory for learning relative affective displacement (rectified-flow assumption).
- domain assumption emotion2vec + VAD + mean F0/energy form a sufficiently complete continuous emotion representation for both categorical and intensity control.
- ad hoc to paper LLM (Qwen3) prompted with p_src, p_tgt, ΔVAD produces faithful relative instructions after rule-based filtering of target labels.
- domain assumption Speaker embedding from EASE plus emotion projection yields style vectors that disentangle speaker from emotion sufficiently for zero-shot conversion.
invented entities (2)
-
Emo-Compass (source-anchored rectified-flow emotion predictor)
no independent evidence
-
TRACE-Instruct dataset
no independent evidence
read the original abstract
Traditional emotional voice conversion (EVC) conditions generation on explicit target emotions like labels or references, defining the target affective state but omitting the direction or nature of the transition. We introduce instruction-guided relative emotional voice conversion, a task where natural-language instructions specify source-conditioned affective transformations (e.g., "make the speech slightly calmer" or "sound noticeably more confident") instead of fixed targets. To support this task, we construct TRACE-Instruct, a dataset of relative emotion instructions covering categorical transitions, intensity modifications, and open-ended affective changes. We propose TRACE-EVC, a zero-shot framework built around Emo-Compass, a module that models each conversion as a source-anchored rectified flow. Rather than conditioning on an explicit target, it predicts the direction and degree of the affective change. Experiments demonstrate that TRACE-EVC accurately follows relative emotion instructions while preserving speaker identity, linguistic content, and speech quality, and remains competitive with conventional EVC systems on standard categorical emotion conversion.
Figures
Reference graph
Works this paper leans on
-
[1]
Emotional voice conversion: Theory, databases and esd,
K. Zhou, B. Sisman, R. Liu, and H. Li, “Emotional voice conversion: Theory, databases and esd,”Speech Communication, vol. 137, pp. 1–18, 2022
2022
-
[2]
An Overview & Analysis of Sequence-to-Sequence Emotional V oice Conversion,
Z. Yang, X. Jing, A. Triantafyllopoulos, M. Song, I. Aslan, and B. W. Schuller, “An Overview & Analysis of Sequence-to-Sequence Emotional V oice Conversion,” inInterspeech 2022, 2022, pp. 4915–4919
2022
-
[3]
Emotion inten- sity and its control for emotional voice conversion,
K. Zhou, B. Sisman, R. Rana, B. W. Schuller, and H. Li, “Emotion inten- sity and its control for emotional voice conversion,”IEEE Transactions on Affective Computing, vol. 14, no. 1, pp. 31–48, 2023
2023
-
[4]
Mixed-EVC: Mixed Emotion Synthesis and Control in V oice Conversion,
K. Zhou, B. Sisman, C. Busso, B. Ma, and H. Li, “Mixed-EVC: Mixed Emotion Synthesis and Control in V oice Conversion,” inThe Speaker and Language Recognition Workshop (Odyssey 2024), 2024, pp. 180– 186
2024
-
[5]
Vaw-gan for disentanglement and recomposition of emotional elements in speech,
K. Zhou, B. Sisman, and H. Li, “Vaw-gan for disentanglement and recomposition of emotional elements in speech,” in2021 IEEE Spoken Language Technology Workshop (SLT), 2021, pp. 415–422
2021
-
[6]
An Improved StarGAN for Emotional V oice Conversion: Enhancing V oice Quality and Data Augmentation,
X. He, J. Chen, G. Rizos, and B. W. Schuller, “An Improved StarGAN for Emotional V oice Conversion: Enhancing V oice Quality and Data Augmentation,” inInterspeech 2021, 2021, pp. 821–825
2021
-
[7]
Limited Data Emotional V oice Con- version Leveraging Text-to-Speech: Two-Stage Sequence-to-Sequence Training,
K. Zhou, B. Sisman, and H. Li, “Limited Data Emotional V oice Con- version Leveraging Text-to-Speech: Two-Stage Sequence-to-Sequence Training,” inInterspeech 2021, 2021, pp. 811–815
2021
-
[8]
Durflex-evc: Duration- flexible emotional voice conversion leveraging discrete representations without text alignment,
H.-S. Oh, S.-H. Lee, D.-H. Cho, and S.-W. Lee, “Durflex-evc: Duration- flexible emotional voice conversion leveraging discrete representations without text alignment,”IEEE Transactions on Affective Computing, vol. 16, no. 3, pp. 1660–1674, 2025
2025
-
[9]
Emotional V oice Conversion with Semi-Supervised Generative Modeling,
H. Zhu, H. Zhan, H. Cheng, and Y . Wu, “Emotional V oice Conversion with Semi-Supervised Generative Modeling,” inInterspeech 2023, 2023, pp. 2278–2282
2023
-
[10]
Zero shot audio to audio emotion trans- fer with speaker disentanglement,
S. Dutta and S. Ganapathy, “Zero shot audio to audio emotion trans- fer with speaker disentanglement,” inICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 10 371–10 375
2024
-
[11]
Textless speech emotion conversion using discrete & decomposed representations,
F. Kreuk, A. Polyak, J. Copet, E. Kharitonov, T. A. Nguyen, M. Rivi `ere, W.-N. Hsu, A. Mohamed, E. Dupoux, and Y . Adi, “Textless speech emotion conversion using discrete & decomposed representations,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y . Goldberg, Z. Kozareva, and Y . Zhang, Eds. Abu Dhabi, United...
2022
-
[12]
Speaking style conversion in the waveform domain using discrete self-supervised units,
G. Maimon and Y . Adi, “Speaking style conversion in the waveform domain using discrete self-supervised units,” inFindings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 8048–8061. [Online]. Available: https://aclanthology.org/2023.fi...
2023
-
[13]
Vevo: Control- lable zero-shot voice imitation with self-supervised disentanglement,
X. Zhang, X. Zhang, K. Peng, Z. Tang, V . Manohar, Y . Liu, J. Hwang, D. Li, Y . Wang, J. Chan, Y . Huang, Z. Wu, and M. Ma, “Vevo: Control- lable zero-shot voice imitation with self-supervised disentanglement,” in ICLR. OpenReview.net, 2025
2025
-
[14]
HybridVC: Efficient V oice Style Conversion with Text and Audio Prompts,
X. Niu, J. Zhang, and C. P. Martin, “HybridVC: Efficient V oice Style Conversion with Text and Audio Prompts,” inInterspeech 2024, 2024, pp. 4368–4372
2024
-
[15]
PromptEVC: Controllable Emotional V oice Conversion with Natural Language Prompts,
T. Qi, S. Wang, C. Lu, T. Song, H. Yang, Z. Wu, and W. Zheng, “PromptEVC: Controllable Emotional V oice Conversion with Natural Language Prompts,” inInterspeech 2025, 2025, pp. 4588–4592
2025
-
[16]
ClapFM-EVC: High-Fidelity and Flexible Emotional V oice Conversion with Dual Control from Natural Language and Speech,
Y . Pan, Y . Hu, Y . Yang, J. Yao, J. Ye, H. Zhou, L. Ma, and J. Zhao, “ClapFM-EVC: High-Fidelity and Flexible Emotional V oice Conversion with Dual Control from Natural Language and Speech,” inInterspeech 2025, 2025, pp. 4583–4587
2025
-
[17]
Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,
J. Yao, Y . Yang, Y . Lei, Z. Ning, Y . Hu, Y . Pan, J. Yin, H. Zhou, H. Lu, and L. Xie, “Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,” inICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 10 571–10 575
2024
-
[18]
Textrolspeech: A text style control speech corpus with codec language text-to-speech models,
S. Ji, J. Zuo, M. Fang, Z. Jiang, F. Chen, X. Duan, B. Huai, and Z. Zhao, “Textrolspeech: A text style control speech corpus with codec language text-to-speech models,” inICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 10 301–10 305
2024
-
[19]
Speechcraft: A fine-grained expressive speech dataset with natural language description,
Z. Jin, J. Jia, Q. Wang, K. Li, S. Zhou, S. Zhou, X. Qin, and Z. Wu, “Speechcraft: A fine-grained expressive speech dataset with natural language description,” inProceedings of the 32nd ACM International Conference on Multimedia, ser. MM ’24. New York, NY , USA: Association for Computing Machinery, 2024, pp. 1255–1264. [Online]. Available: https://doi.org...
-
[20]
Stylecap: Automatic speaking- style captioning from speech based on speech and language self- supervised learning models,
K. Yamauchi, Y . Ijima, and Y . Saito, “Stylecap: Automatic speaking- style captioning from speech based on speech and language self- supervised learning models,” inICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 11 261–11 265
2024
-
[21]
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive V oice Conversion,
S ¸. Akti, T.-N. Nguyen, and A. Waibel, “Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive V oice Conversion,” in Interspeech 2025, 2025, pp. 1358–1362
2025
-
[22]
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot V oice Conversion,
K. Wang, W. Guan, Z. Jiang, H. Huang, P. Chen, W. Wu, Q. Hong, and L. Li, “Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot V oice Conversion,” inInterspeech 2025, 2025, pp. 1383–1387
2025
-
[23]
Speaker Normalization and Content Restoration for Zero-Shot V oice Conversion with Attention- Enhanced Discriminator,
D. Hu, Y . Xiang, J. Lu, X. Hu, and X. Xu, “Speaker Normalization and Content Restoration for Zero-Shot V oice Conversion with Attention- Enhanced Discriminator,” inInterspeech 2025, 2025, pp. 1403–1407
2025
-
[24]
Emotional V oice Conversion Using Neural Networks with Different Temporal Scales of F0 based on Wavelet Transform,
Z. Luo, T. Takiguchi, and Y . Ariki, “Emotional V oice Conversion Using Neural Networks with Different Temporal Scales of F0 based on Wavelet Transform,” in9th ISCA Workshop on Speech Synthesis Workshop (SSW 9), 2016, pp. 140–145
2016
-
[25]
Deep Bidirectional LSTM Modeling of Timbre and Prosody for Emotional V oice Conversion,
H. Ming, D. Huang, L. Xie, J. Wu, M. Dong, and H. Li, “Deep Bidirectional LSTM Modeling of Timbre and Prosody for Emotional V oice Conversion,” inInterspeech 2016, 2016, pp. 2453–2457
2016
-
[26]
Transforming Spectrum and Prosody for Emotional V oice Conversion with Non-Parallel Training Data,
K. Zhou, B. Sisman, and H. Li, “Transforming Spectrum and Prosody for Emotional V oice Conversion with Non-Parallel Training Data,” inThe Speaker and Language Recognition Workshop (Odyssey 2020), 2020, pp. 230–237
2020
-
[27]
Nonparallel Emotional Speech Conversion Using V AE-GAN,
Y . Cao, Z. Liu, M. Chen, J. Ma, S. Wang, and J. Xiao, “Nonparallel Emotional Speech Conversion Using V AE-GAN,” inInterspeech 2020, 2020, pp. 3406–3410
2020
-
[28]
Disentanglement of Emotional Style and Speaker Identity for Expressive V oice Conversion,
Z. Du, B. Sisman, K. Zhou, and H. Li, “Disentanglement of Emotional Style and Speaker Identity for Expressive V oice Conversion,” inInter- speech 2022, 2022, pp. 2603–2607
2022
-
[29]
Converting Anyone’s Emotion: Towards Speaker-Independent Emotional V oice Conversion,
K. Zhou, B. Sisman, M. Zhang, and H. Li, “Converting Anyone’s Emotion: Towards Speaker-Independent Emotional V oice Conversion,” inInterspeech 2020, 2020, pp. 3416–3420
2020
-
[30]
DiffEmotionVC: A Dual-Granularity Disentangled Diffusion Framework for Any-to-Any Emotional V oice Conversion,
X. Su, B. Yang, X. Yi, and Y . Cao, “DiffEmotionVC: A Dual-Granularity Disentangled Diffusion Framework for Any-to-Any Emotional V oice Conversion,” inInterspeech 2025, 2025, pp. 4393–4397
2025
-
[31]
ZSDEVC: Zero-Shot Diffusion-based Emotional V oice Conversion with Disentan- gled Mechanism,
H.-H. Chou, Y .-S. Lin, C.-C. Sung, Y . Tsao, and C.-C. Lee, “ZSDEVC: Zero-Shot Diffusion-based Emotional V oice Conversion with Disentan- gled Mechanism,” inInterspeech 2025, 2025, pp. 4398–4402
2025
-
[32]
Prompttts: Controllable text-to-speech with text descriptions,
Z. Guo, Y . Leng, Y . Wu, S. Zhao, and X. Tan, “Prompttts: Controllable text-to-speech with text descriptions,” inICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5
2023
-
[33]
Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,
D. Yang, S. Liu, R. Huang, C. Weng, and H. Meng, “Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 2913–2925, 2024
2024
-
[34]
PromptStyle: Controllable Style Transfer for Text-to-Speech with Nat- ural Language Descriptions,
G. Liu, Y . Zhang, Y . Lei, Y . Chen, R. Wang, L. Xie, and Z. Li, “PromptStyle: Controllable Style Transfer for Text-to-Speech with Nat- ural Language Descriptions,” inInterspeech 2023, 2023, pp. 4888–4892
2023
-
[35]
Controlling Emotion in Text-to-Speech with Natural Language Prompts,
T. Bott, F. Lux, and N. T. Vu, “Controlling Emotion in Text-to-Speech with Natural Language Prompts,” inInterspeech 2024, 2024, pp. 1795– 1799
2024
-
[36]
PL-TTS: A Generalizable Prompt-based Diffusion TTS Augmented by Large Language Model,
S. Li, Q. Mao, and J. Shi, “PL-TTS: A Generalizable Prompt-based Diffusion TTS Augmented by Large Language Model,” inInterspeech 2024, 2024, pp. 4888–4892
2024
-
[37]
Exploring speech style spaces with language models: Emotional TTS without emotion labels,
S. S. Chandra, Z. Du, and B. Sisman, “Exploring speech style spaces with language models: Emotional TTS without emotion labels,” inThe Speaker and Language Recognition Workshop (Odyssey 2024), 2024, pp. 194–200
2024
-
[38]
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval,
H. Sun, J. Tian, J. Zhou, H. Wang, J. He, S. Zhao, X. Kong, D. Hu, X. Xu, X. Hu, and Y . Qin, “RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval,” inInterspeech 2025, 2025, pp. 2995–2999
2025
-
[39]
Cross-Cultural Comparison of Gradient Emotion Perception: Human vs. Alexa TTS V oices,
I. Gessinger, M. Cohn, G. Zellou, and B. M ¨obius, “Cross-Cultural Comparison of Gradient Emotion Perception: Human vs. Alexa TTS V oices,” inInterspeech 2022, 2022, pp. 4970–4974
2022
-
[40]
Cross- linguistic Emotion Perception in Human and TTS V oices,
I. Gessinger, M. Cohn, B. R. Cowan, G. Zellou, and B. M ¨obius, “Cross- linguistic Emotion Perception in Human and TTS V oices,” inInterspeech 2023, 2023, pp. 5222–5226
2023
-
[41]
Emotional dimension control in language model-based text-to-speech: Spanning a broad spectrum of human emotions,
K. Zhou, Y . Zhang, D. Ng, S. Zhao, H. Wang, and B. Ma, “Emotional dimension control in language model-based text-to-speech: Spanning a broad spectrum of human emotions,” inICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 17 257–17 261
2026
-
[42]
Emotion Arithmetic: Emotional Speech Synthesis via Weight Space Interpolation,
P. Kalyan, P. Rao, P. Jyothi, and P. Bhattacharyya, “Emotion Arithmetic: Emotional Speech Synthesis via Weight Space Interpolation,” inInter- speech 2024, 2024, pp. 1805–1809
2024
-
[43]
VECL- TTS: V oice identity and Emotional style controllable Cross-Lingual Text-to-Speech,
A. Gudmalwar, N. Shah, S. Akarsh, P. Wasnik, and R. R. Shah, “VECL- TTS: V oice identity and Emotional style controllable Cross-Lingual Text-to-Speech,” inInterspeech 2024, 2024, pp. 3000–3004
2024
-
[44]
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis,
H. Li, L. Qu, J. Hu, and T. Li, “EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis,” inInterspeech 2025, 2025, pp. 4368–4372
2025
-
[45]
DiEmo-TTS: Disentan- gled Emotion Representations via Self-Supervised Distillation for Cross- Speaker Emotion Transfer in Text-to-Speech,
D.-H. Cho, H.-S. Oh, S.-B. Kim, and S.-W. Lee, “DiEmo-TTS: Disentan- gled Emotion Representations via Self-Supervised Distillation for Cross- Speaker Emotion Transfer in Text-to-Speech,” inInterspeech 2025, 2025, pp. 4373–4377
2025
-
[46]
EATS-Speech: Emotion- Adaptive Transformation and Priority Synthesis for Zero-Shot Text-to- Speech,
J. Xing, Z. Li, S. Chen, X. Xing, and X. Xu, “EATS-Speech: Emotion- Adaptive Transformation and Priority Synthesis for Zero-Shot Text-to- Speech,” inInterspeech 2025, 2025, pp. 4358–4362
2025
-
[47]
RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning,
A. Sigurgeirsson and S. King, “RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning,” in13th edition of the Speech Synthesis Workshop, 2025, pp. 130–136
2025
-
[48]
Affective story teller: a TTS system for emotional expressivity,
M. A. M. Shaikh, A. R. F. Rebord ˜ao, and K. Hirose, “Affective story teller: a TTS system for emotional expressivity,” inInterspeech 2010, 2010, pp. 518–521
2010
-
[49]
EmoSphere- TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech,
D.-H. Cho, H.-S. Oh, S.-B. Kim, S.-H. Lee, and S.-W. Lee, “EmoSphere- TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech,” inInterspeech 2024, 2024, pp. 1810–1814
2024
-
[50]
Emosphere++: Emotion-controllable zero-shot text-to-speech via emotion-adaptive spherical vector,
D.-H. Cho, H.-S. Oh, S.-B. Kim, and S.-W. Lee, “Emosphere++: Emotion-controllable zero-shot text-to-speech via emotion-adaptive spherical vector,”IEEE Transactions on Affective Computing, vol. 16, no. 3, pp. 2365–2380, 2025
2025
-
[51]
Mead: A large-scale audio-visual dataset for emotional talking-face generation,
K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y . Qiao, and C. C. Loy, “Mead: A large-scale audio-visual dataset for emotional talking-face generation,” inComputer Vision – ECCV 2020, A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds. Cham: Springer Interna- tional Publishing, 2020, pp. 700–717
2020
-
[52]
Odyssey 2024 - Speech Emotion Recognition Challenge: Dataset, Baseline Framework, and Results,
L. Goncalves, A. N. Salman, A. R. Naini, L. Moro-Vel ´azquez, T. The- baud, P. Garcia, N. Dehak, B. Sisman, and C. Busso, “Odyssey 2024 - Speech Emotion Recognition Challenge: Dataset, Baseline Framework, and Results,” inThe Speaker and Language Recognition Workshop (Odyssey 2024), 2024, pp. 247–254
2024
-
[53]
A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lvet al., “Qwen3 technical report,”arXiv preprint arXiv:2505.09388, 2025
Pith/arXiv arXiv 2025
-
[54]
Text embeddings by weakly-supervised contrastive pre- training,
L. Wang, N. Yang, X. Huang, B. Jiao, L. Yang, D. Jiang, R. Majumder, and F. Wei, “Text embeddings by weakly-supervised contrastive pre- training,”arXiv preprint arXiv:2212.03533, 2022
Pith/arXiv arXiv 2022
-
[55]
emotion2vec: Self-supervised pre-training for speech emotion representation,
Z. Ma, Z. Zheng, J. Ye, J. Li, Z. Gao, S. Zhang, and X. Chen, “emotion2vec: Self-supervised pre-training for speech emotion representation,” inFindings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V . Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 15 747–15 760. [Online]...
2024
-
[56]
Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021
2021
-
[57]
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 17 022–17 033. [Online]. Available: https://proceedings.neurips.cc/p...
2020
-
[58]
Amphion: An open-source audio, music and speech generation toolkit,
X. Zhang, L. Xue, Y . Gu, Y . Wang, J. Li, H. He, C. Wang, T. Song, X. Chen, Z. Fang, H. Chen, J. Zhang, T. Y . Tang, L. Zou, M. Wang, J. Han, K. Chen, H. Li, and Z. Wu, “Amphion: An open-source audio, music and speech generation toolkit,” inIEEE Spoken Language Technology Workshop, SLT 2024, 2024
2024
-
[59]
Overview of the amphion toolkit (v0.2),
J. Li, X. Zhang, Y . Wang, H. He, C. Wang, L. Wang, H. Liao, J. Ao, Z. Xie, Y . Huang, J. Zhang, and Z. Wu, “Overview of the amphion toolkit (v0.2),”arXiv preprint arXiv:2501.15442, 2025
Pith/arXiv arXiv 2025
-
[60]
Robust speech recognition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. Mcleavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” inProceedings of the 40th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202. PMLR, 23–2...
2023
-
[61]
Generalized end-to-end loss for speaker verification,
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 4879– 4883
2018
-
[62]
UTMOS: UTokyo-SaruLab System for V oiceMOS Chal- lenge 2022,
T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-SaruLab System for V oiceMOS Chal- lenge 2022,” inInterspeech 2022, 2022, pp. 4521–4525
2022
-
[63]
NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Pre- diction with Crowdsourced Datasets,
G. Mittag, B. Naderi, A. Chehadi, and S. M ¨oller, “NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Pre- diction with Crowdsourced Datasets,” inInterspeech 2021, 2021, pp. 2127–2131
2021
-
[64]
Un- supervised Cross-Lingual Representation Learning for Speech Recogni- tion,
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Un- supervised Cross-Lingual Representation Learning for Speech Recogni- tion,” inInterspeech 2021, 2021, pp. 2426–2430
2021
-
[65]
LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,
H. Zen, V . Dang, R. Clark, Y . Zhang, R. J. Weiss, Y . Jia, Z. Chen, and Y . Wu, “LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,” inInterspeech 2019, 2019, pp. 1526–1530
2019
-
[66]
Mtld, vocd-d, and hd-d: A validation study of sophisticated approaches to lexical diversity assessment,
P. M. McCarthy and S. Jarvis, “Mtld, vocd-d, and hd-d: A validation study of sophisticated approaches to lexical diversity assessment,”Be- havior research methods, vol. 42, no. 2, pp. 381–392, 2010
2010
-
[67]
Texygen: A benchmarking platform for text generation models,
Y . Zhu, S. Lu, L. Zheng, J. Guo, W. Zhang, J. Wang, and Y . Yu, “Texygen: A benchmarking platform for text generation models,” inThe 41st international ACM SIGIR conference on research & development in information retrieval, 2018, pp. 1097–1100
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.