{"id":"51fcb464-b04e-4ffd-b6f9-ecaec8051653","arxiv_id":"2505.20678","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"PromptEVC uses natural language prompts, a diffusion-based prompt mapper, and prosody control to perform emotional voice conversion, reporting improved controllability over label- and reference-based methods.","lead":"PromptEVC is a system that lets a user control the emotion in converted speech by typing a natural language description, like 'a very happy tone with a hint of surprise.' It converts these descriptions into fine-grained emotion embeddings, then adjusts rhythm and prosody while preserving the speaker, and it reports better controllability scores than label- or reference-based baselines on TextrolSpeech.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Controllability superiority is asserted but never tested against baselines: Tables 2-3 report PromptEVC alone, so 'outperforms' in intensity/mixed-emotion/prosody is unsupported.","rationale":"The reader's verdict of CONDITIONAL is appropriate, but I locate the load-bearing gap differently. The reader emphasizes unvalidated ChatGPT-4 prompt rewriting; that is a real data-quality concern, yet it is indirectly checked by the attribute accuracy results in Tables 2-3, which use the original labels as ground truth. The more consequential gap is the absence of any baseline comparison for the specific controllability dimensions named in the central claim. Table 1 compares global quality only, so the claim that PromptEVC outperforms prior controllable EVC methods in intensity control, mixed emotion synthesis, and prosody manipulation is not directly supported by any reported experiment. This is a fixable but essential omission: the conclusion overreaches the measurements. My recommendation does not change the reader's verdict because the proposed condition--adding baseline controllability comparisons--is already consistent with a CONDITIONAL acceptance, and no evidence in the paper suggests the core architecture is invalid.","tokens_in":8359,"tokens_out":6520,"duration_ms":72958,"concrete_test":"Re-run the Section 3.4 protocols on all comparison models with identical test prompts and source utterances: use the same attribute classifiers and the same five-participant listening test, and report per-attribute accuracy for Ein, Emx, P, S, V along with a paired significance test (e.g., bootstrap or Wilcoxon) for PromptEVC versus the strongest baseline. If PromptEVC does not significantly exceed Emovox or Mixed-EVC on the relevant attributes, the superiority claim should be reduced to 'competitive on global quality' and the abstract's comparative wording should be adjusted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim is comparative: PromptEVC 'outperforms state-of-the-art controllable EVC methods in emotion conversion, intensity control, mixed emotion synthesis, and prosody manipulation' (Abstract). The only head-to-head comparisons are in Table 1, which measures MCD, CER, RMSE_F0, ACC_cls, MOS naturalness, and similarity. These are global quality metrics; they do not directly measure intensity control, mixed-emotion fidelity, or prosody manipulation. The controllability evidence in Section 3.4 (Tables 2 and 3) reports classification accuracy for Ein, Emx, P, S, V for PromptEVC only. No baseline (Emovox, Mixed-EVC, ZEST, Textless-EVC) is evaluated on these same attribute-classification or listener protocols. Therefore the data support 'PromptEVC produces recognizable attributes,' not 'PromptEVC outperforms SOTA on those attributes.' The comparative claim for three of four capabilities is an inference from quality metrics, not a measurement. Additionally, the attribute classifier used in Table 2 is not described in terms of architecture, training data, or whether it was trained on the same TextrolSpeech distribution, which makes the absolute accuracy numbers hard to interpret without this information.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes PromptEVC, an emotional voice conversion system that uses natural language prompts to control emotional expression. The system comprises a RoBERTa-based emotion descriptor, a diffusion-based prompt mapper that aligns text embeddings with emotion2vec reference embeddings, a prosody modeling and control module with a duration regulator and prosody predictor, and a speaker encoder with an F0 constraint. Experiments on the TextrolSpeech dataset compare PromptEVC against four baselines on global objective and subjective metrics (MCD, CER, RMSE_F0, ACC_cls, MOS naturalness, similarity) and include ablations. Controllability is assessed via classifier-based and listener-based accuracy on emotional intensity, mixed emotion, pitch, speaking speed, and volume.","tokens_in":8632,"tokens_out":5394,"duration_ms":53934,"significance":"If the claims hold, PromptEVC would be a valuable step toward flexible, text-driven control of emotional expression in voice conversion, removing the need for reference audio or numeric values at inference. The architecture is plausible and internally consistent: the components (VITS, HuBERT units, emotion2vec, diffusion conditioning) are integrated in a sensible way, and the ablations show that each proposed module contributes to the reported metrics. The main contribution is a new combination of existing techniques for the EVC setting, with a prompt-mapping mechanism that appears trainable with available paired data. However, the current experimental evidence does not support the comparative controllability claims, and the evaluation has statistical gaps that prevent the reported superiority from being established.","major_comments":[{"comment":"The abstract and conclusion claim that PromptEVC outperforms state-of-the-art controllable EVC methods in intensity control, mixed emotion synthesis, and prosody manipulation, but Tables 2 and 3 report attribute classification accuracies for PromptEVC only. No baseline (Textless-EVC, ZEST, Emovox, Mixed-EVC) is evaluated with the same pre-trained classifier or the same five-participant listening protocol on E_in, E_mx, P, S, or V. The superiority for these three capabilities is therefore inferred from the global quality metrics in Table 1 rather than measured on the controllability tasks themselves. Please run the same controllability protocols on at least the most relevant baselines (Emovox for intensity, Mixed-EVC for mixed emotion) or revise the claim to state that PromptEVC enables control of these attributes rather than outperforming existing methods on them.","section":"Section 3.4, Tables 2 and 3"},{"comment":"The ChatGPT-4 rewritten prompts are the sole training signal for the emotion descriptor and the prompt mapper, yet no validation is reported that the rewrites preserve the original five style factors (E_in, P, S, V, E_cg) or the emotional content of the paired audio. The rewrites also deliberately remove gender information, which is reasonable for the task, but the extent to which other factors are distorted is unknown. If the rewrites systematically alter intensity, speed, volume, or mixed-emotion annotations, the text-conditioned emotion embeddings will be trained on incorrect targets, and the controllability results in Tables 2 and 3 may not transfer to real user prompts. Please provide a human or automatic validation of the rewritten prompts (e.g., ratings against the original factor values, or a classifier-based consistency check) and report the agreement.","section":"Section 3.1, Data preparation"},{"comment":"The objective metrics MCD, CER, RMSE_F0, and ACC_cls are reported as point estimates with no variance, no test set size, and no significance tests, and only MOS has confidence intervals. The comparison is also performed on a single corpus (TextrolSpeech). Under these conditions, the claimed superiority of PromptEVC over Emovox and Mixed-EVC on these metrics is not statistically established. Please report the number of test utterances, per-utterance or per-condition standard deviations, and perform significance testing (e.g., paired tests with bootstrap or Wilcoxon) for at least the primary metrics. Also specify how many speakers and which emotion categories are included in the evaluation set.","section":"Table 1 and Section 3.2"},{"comment":"ACC_cls is attributed to a 'pre-trained speech emotion recognition (SER) model [31]', but reference [31] is the FunASR speech recognition toolkit, not a SER model. This is either a mis-citation or a serious mismatch between the stated metric and the actual implementation. Please specify the exact emotion classification model used, its training data, and whether it was independently trained on the same TextrolSpeech distribution. Likewise, the pre-trained classifier used for Table 2 is not described in terms of architecture, training data, or whether it was trained on TextrolSpeech; without this information, the absolute accuracy values (77.58%, 61.25%, etc.) are difficult to interpret, and the reader cannot judge whether the classifier is sensitive enough to support the controllability claims.","section":"Section 3.2, Table 1; Section 3.4, Table 2"}],"minor_comments":[{"comment":"The diffusion loss is written with a double vertical bar but no exponent; it should be the squared L2 norm. Please also define the dimensions of x_t and e_txt explicitly.","section":"Section 2.1, Eq. (3)"},{"comment":"The duration loss expression `log cosh(\\hat{y}_i - log(y_i+1))` is unconventional; the log-cosh loss is usually applied to the residual y - \\hat{y}, not to a residual after logarithmic scaling. Please clarify the exact functional form and why this particular transformation is used.","section":"Section 2.2, Eq. (4)"},{"comment":"The intensity cutoffs of 30% / 40% / 30% for high/medium/low are presented without justification. Please cite a source or provide a rationale (e.g., distributional analysis) for these thresholds.","section":"Section 3.1"},{"comment":"The Ground Truth row has no entry for E_mx. Please clarify whether mixed-emotion ground truth is unavailable by design, and how the 60.04% accuracy for E_mx should be interpreted in that case.","section":"Table 3"},{"comment":"Figure 1 is quite dense and hard to read in the preprint. Please enlarge the figure and provide a more detailed caption that explains the data flow, especially for the prosody modeling and control module shown in Figure 1(b).","section":"Figure 1"},{"comment":"The phrase 'outperforms state-of-the-art controllable EVC methods' and 'significantly improves' are stronger than what the current experimental design supports, given the missing baseline comparisons and statistical tests noted above. Please soften these statements until the evidence is added.","section":"Abstract and Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper is a plausible short-format EVC contribution, but the central comparative claims outrun the reported evidence. The mismatched citation of FunASR as an SER model is a concrete red flag that should be corrected. The lack of baseline controllability comparisons is the main technical gap; I believe it is fixable within the scope of a revision, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is bringing natural language prompts to emotional voice conversion. PromptVC and PromptTTS2 do this for style and TTS, but not for emotion specifically, and the combination of a RoBERTa descriptor, a diffusion-based prompt mapper that aligns text to emotion2vec reference embeddings, and a prosody control pipeline is a sensible way to get fine-grained control without labels, reference clips, or sliders. Credit where it's due: the architecture is coherent, the ablation study shows each removed component hurts in the direction you'd expect, and Table 1 consistently favors PromptEVC on MCD, CER, RMSE_F0, ACC_cls, and MOS. That is a real empirical result for a first system.\n\nThe soft spots are mostly about the gap between what the abstract claims and what the data actually show. The abstract says PromptEVC outperforms state-of-the-art methods in intensity control, mixed emotion synthesis, and prosody manipulation. But Tables 2 and 3 only report PromptEVC's own attribute classification accuracy; no baseline (Emovox, Mixed-EVC, ZEST, Textless-EVC) is run on those protocols. So the data support \"PromptEVC produces recognizable attributes,\" not \"PromptEVC beats the baselines on those attributes.\" That's a real overreach, and the stress-test note is right to flag it. The fix is either add baselines to the controllability evaluation or soften the language.\n\nThe next soft spot is the ChatGPT-4 prompt rewriting. The emotion descriptor and prompt mapper are trained on the rewritten text, and there's no human validation that the rewrites preserved intensity, speed, volume, or mixed-emotion annotations. That's a load-bearing assumption, and it's untreated. Also, the objective results in Table 1 are reported without variance or significance tests, and ACC_cls is attributed to FunASR, which is an ASR toolkit, not a speech emotion recognizer—likely a mis-citation. One corpus and no code or data artifacts limit reproducibility, though the demo page is a plus.\n\nUnderneath that, the central idea is credible and the system is not internally contradictory. The citation pattern is fine; self-cites to the authors' earlier EVC work are relevant and not hiding anything. The paper deserves a serious referee because it opens a useful direction and the core mapping is plausible, but it needs a revision that either delivers the missing baseline comparison for controllability or restates the claim, plus validation of the prompt rewriting.\n\nRecommendation: send to peer review. It's not ready as is, but it's a solid workshop-to-conference submission with a novel capability and a fixable evaluation gap.","headline":"Text prompts for emotional voice conversion is a genuinely new capability and the system looks plausible, but the paper claims a comparative win on controllability that its own tables don't actually test.","tokens_in":9155,"tokens_out":1693,"would_cite":false,"duration_ms":18237,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PromptEVC claims that free-text descriptions alone can control emotional voice conversion at fine granularity—emotion category, intensity, mixed emotions, and prosody—without labels, reference audio, or numeric values, and that it beats…","keywords":["emotional voice conversion","natural language prompts","fine-grained emotion control","diffusion prompt mapper","prosody modeling","mixed emotion synthesis","speaker identity preservation","text-conditioned speech synthesis"],"falsifier":"Synthesize one neutral utterance under prompts that differ only in the stated intensity ('slightly happy', 'happy', 'very happy') and only in prosody words ('faster' vs 'slower'); if a pretrained emotion model's intensity scores and listener ratings do not move monotonically in the stated direction, the claimed fine-grained control does not hold.","tokens_in":8160,"feed_emoji":"🎤","tokens_out":10880,"duration_ms":97713,"temperature":0.7,"pith_summary":"The paper sets out to show that a voice-conversion system can be controlled entirely by a free-text description of the desired emotion, replacing fixed labels, reference recordings, and hand-set numeric values. It proposes a text-to-emotion pathway in which a text encoder produces a coarse emotion embedding and a diffusion-based prompt mapper refines it against emotion embeddings extracted from real speech, so the final embedding tracks words like 'very', 'slightly', and 'with a hint of surprise'. A prosody pipeline then adjusts rhythm using the linguistic content plus emotional cues, and a speaker encoder with a fundamental-frequency constraint preserves identity. On the TextrolSpeech corpus the paper reports better emotion conversion, intensity control, mixed-emotion synthesis, and prosody manipulation than label-, reference-, and value-based EVC baselines, with ablations attributing a large share of the gain to the prompt mapper. If the claim holds, users could simply type how they want an utterance to sound and the system would deliver it.","feed_headline":"Text prompts steer emotion, intensity, and prosody in voice conversion","feed_subtitle":"No labels, reference audio, or sliders: a diffusion mapper turns plain wording into fine-grained emotional control.","key_machinery":"The load-bearing mechanism is the two-stage text-to-emotion embedding pathway: a coarse text embedding produced by a pretrained language model, refined by a diffusion model that is jointly trained to reproduce reference emotion embeddings from speech. The diffusion prompt mapper is what makes open-vocabulary wording usable—without it, directly predicting emotion embeddings from text raises pitch error and lowers emotion-classification accuracy and naturalness. Supporting mechanisms are the prosody pipeline (deduplicated speech units, a duration regulator, and a prosody predictor) and an F0-constrained speaker encoder, which together keep content intelligible and identity stable while rhythm and pitch are edited.","core_discovery":"The authors' central claim is that natural-language prompts can serve as the sole control signal for fine-grained emotional voice conversion, provided the gap between text and speech is bridged in two stages. The emotion descriptor turns a prompt into a coarse embedding $e_{txt}$; the prompt mapper, a diffusion model trained jointly with reference embeddings $e_{ref}$ from a self-supervised speech-emotion model, refines this into $e_{pm}$, and only this text-conditioned embedding is used at inference. Prosody is handled separately: quantized self-supervised speech units provide linguistic tokens, a duration regulator adjusts rhythm from content and emotion, and a prosody predictor synthesizes natural pitch and energy. A speaker encoder with an F0 constraint keeps identity fixed while intonation changes. The paper reports that this design beats label-, reference-, and value-based controllable EVC baselines on emotion conversion, intensity control, mixed-emotion synthesis, and prosody manipulation, and its ablations show that removing the prompt mapper is the most damaging single change.","pith_inferences":["One consequence the paper leaves implicit is that the reference encoder may be needed only for training: since inference uses only the text-conditioned embedding, the same design could run without any reference speech at deployment, which suits streaming or on-device use.","The controllability claim is bounded by the five style factors of the training corpus; prompts describing states outside that coverage, such as 'bored but urgent' or non-English descriptions, are a natural stress test that the paper does not report.","Because the training prompts were rewritten by a large language model with no reported human validation, the intensity-control numbers should be read as conditional on those rewrites preserving the original style-factor tags.","The same text-conditioned diffusion-mapper design could plausibly extend to other suprasegmental speech attributes—politeness, emphasis, confidence—whenever paired text descriptions and speech are available."],"forward_implications":["Users can specify the target emotion in free text, such as 'very happy with a hint of surprise,' and get converted speech without selecting a reference clip or typing numeric values.","Because the prompt mapper is trained against reference speech embeddings, the same text interface can express emotional intensity, mixed emotions, pitch, speed, and volume within a single prompt.","The prosody pipeline keeps content intelligible during emotional edits; the paper reports lower character error rate than the label- and reference-based baselines.","The F0-constrained speaker encoder means emotional expression can be altered without the speaker identity drifting, which is needed for dubbing and assistive applications."],"supporting_citations":[{"why":"It supplies the paired text-description and emotional speech corpus with the five style factors used to train and evaluate the system.","marker":"[29]"},{"why":"It defines the relative-attribute-ranking intensity measurement the paper uses to set high, medium, and low emotional intensity levels.","marker":"[13]"},{"why":"It provides the self-supervised speech emotion representation from which reference embeddings are extracted for joint training of the prompt mapper.","marker":"[23]"},{"why":"It provides the pretrained text encoder used as the backbone of the emotion descriptor.","marker":"[24]"},{"why":"It provides the self-supervised speech representations that are quantized into linguistic tokens for prosody modeling.","marker":"[25]"},{"why":"It provides the conditional variational autoencoder architecture on which the decoder and adversarial training are built.","marker":"[30]"},{"why":"It is the prior voice-conversion work that introduced natural-language-prompt control, which PromptEVC extends to emotion conversion.","marker":"[21]"},{"why":"It serves as the reference-audio-based controllable baseline the paper compares against.","marker":"[11]"},{"why":"It serves as the mixed-emotion synthesis baseline that uses manually defined attribute vectors.","marker":"[15]"},{"why":"It provides the speech emotion recognition classifier used to compute objective emotion-classification accuracy.","marker":"[31]"}],"fun_headline_variants":["Natural language prompts control emotional voice conversion","Say it with words: AI voice changes emotion via prompts","Prompt-to-emotion: text steers voice intensity and prosody","No sliders, just sentences: emotional voice control via prompts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the large-language-model rewrites of the training prompts preserve the original style-factor tags and the emotional content of the paired audio, because the emotion descriptor and prompt mapper are trained entirely on that alignment and no human validation of the rewrites is reported.","fun_headline_variants_meta":{"raw":{"variants":["Natural language prompts control emotional voice conversion","Say it with words: AI voice changes emotion via prompts","Prompt-to-emotion: text steers voice intensity and prosody","No sliders, just sentences: emotional voice control via prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1338,"prompt_tokens":913,"completion_tokens":425,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":359}},"tokens_in":529,"tokens_out":425,"duration_ms":4839,"temperature":1.0,"reasoning_tokens":359,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:48:54.826613+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Synthesize one neutral utterance under prompts that differ only in the stated intensity ('slightly happy', 'happy', 'very happy') and only in prosody words ('faster' vs 'slower'); if a pretrained emotion model's intensity scores and listener ratings do not move monotonically in the stated direction, the claimed fine-grained control does not hold.","supporting_citations":[{"cited_title":"emotion2vec: Self-supervised pre-training for speech em otion representation,","cited_arxiv_id":null,"evidence_quote":"It supplies the paired text-description and emotional speech corpus with the five style factors used to train and evaluate the system."},{"cited_title":"Towards realistic emotional voice conversion using controllable e motional intensity,","cited_arxiv_id":null,"evidence_quote":"It defines the relative-attribute-ranking intensity measurement the paper uses to set high, medium, and low emotional intensity levels."},{"cited_title":"De Beaugrande and W","cited_arxiv_id":null,"evidence_quote":"It provides the self-supervised speech emotion representation from which reference embeddings are extracted for joint training of the prompt mapper."},{"cited_title":"V ocal expression of affect,","cited_arxiv_id":null,"evidence_quote":"It provides the pretrained text encoder used as the backbone of the emotion descriptor."},{"cited_title":"Prompttts: Control- lable text-to-speech with text descriptions,","cited_arxiv_id":null,"evidence_quote":"It provides the self-supervised speech representations that are quantized into linguistic tokens for prosody modeling."},{"cited_title":"Mixed-evc : Mixed emotion synthesis and control in voice conversion,","cited_arxiv_id":null,"evidence_quote":"It is the prior voice-conversion work that introduced natural-language-prompt control, which PromptEVC extends to emotion conversion."},{"cited_title":"Textless speech emotion conversion using discrete & decom - posed representations,","cited_arxiv_id":null,"evidence_quote":"It serves as the reference-audio-based controllable baseline the paper compares against."},{"cited_title":"Seen and unseen emo- tional style transfer for voice conversion with a new emotio nal speech dataset,","cited_arxiv_id":null,"evidence_quote":"It serves as the mixed-emotion synthesis baseline that uses manually defined attribute vectors."},{"cited_title":"Hubert: Self-supervised speech repre sen- tation learning by masked prediction of hidden units,","cited_arxiv_id":null,"evidence_quote":"It provides the speech emotion recognition classifier used to compute objective emotion-classification accuracy."}],"review_version":1}