{"id":"6607e103-95a5-4029-a11e-7a2fa3c3e3c3","arxiv_id":"2501.08791","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A CCNF-based manipulation block inside YourTTS shifts synthesized voices along perceptual voice quality axes, with expert-confirmed changes for breathiness and high roughness, at a cost in speaker similarity.","lead":"This paper adds a voice-manipulation block to a text-to-speech system that shifts voices along perceptual qualities such as breathiness, roughness, resonance, and weight on a continuous scale. Phonetic experts confirmed the effect for breathiness and for strong roughness on both voices the system had heard and voices it had not.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that resonance and weight are manipulated rests on a circular objective metric: the same regressors that generated the training pseudo-labels are used for evaluation, with no expert validation for these two PVQs.","rationale":"The reader's weakest assumption identified the pseudo-labeling pipeline as the core risk, and I agree: the regressors trained on 296 atypical-voice files are the only source of training signal for all four PVQs, and the objective evaluation for resonance and weight reuses exactly those regressors. This makes Table II a self-consistency check for the two gendered PVQs, not a perceptual validation. The paper's own observation that predicted scores are compressed relative to expert ratings (e.g., breathiness low/med/high: 11/19/35 vs. 24/46/70) further weakens the regressors as an absolute measure. The subjective listening test for breathiness and roughness partially breaks the circularity and is the strongest evidence in the paper, showing that the approach can work for at least one quality. However, the conclusion generalizes to all four PVQs without expert evidence for resonance and weight. This does not require rejection: the approach is promising and the breathiness result is convincing. But the scope of the claim should be narrowed, or the missing perceptual validation supplied. Since the reader already set the verdict to CONDITIONAL, my read does not change the verdict; it sharpens the specific reason why conditionality is needed: the unvalidated generalization to resonance and weight. The concrete test I propose would settle whether that generalization holds by applying the same expert protocol that was already used successfully for the other two qualities.","tokens_in":9909,"tokens_out":5092,"duration_ms":56984,"concrete_test":"Run the same expert listening protocol used for roughness/breathiness (Section IV, Table III) on resonance and weight manipulations at low/medium/high conditioning for a small set of seen and unseen speakers, and test whether expert-rated severity increases significantly across levels using a paired Wilcoxon test. If expert ratings do not increase monotonically and significantly, then the resonance and weight results are not perceptually established and the central claim overreaches. This directly tests whether the pseudo-labeling pipeline produces perceptually valid manipulation for the unevaluated PVQs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that the system 'effectively manipulates PVQs' covers all four qualities, but for resonance and weight the only evaluation is Table II, which uses the same random-forest regressors that generated the pseudo-labels (Section III-B: 'These fitted regressors were used to pseudo-label the LibriTTS-R dataset'). These regressors were trained on only 296 PVQD+ files and then applied to the 585-hour LibriTTS-R corpus; the paper itself notes their outputs are compressed relative to expert ratings (e.g., breathiness predicted 11/19/35 vs. expert-rated 24/46/70 in Tables II–III). Thus Table II largely measures whether the CCNF can re-invert the regressor mapping, not whether listeners perceive the intended voice-quality change. The expert ratings for roughness and breathiness break this circularity for those two qualities, but no such independent evidence exists for resonance and weight. The conclusion nonetheless claims success for all four PVQs, so the weakest link is the unsupported generalization from two perceptually validated qualities to four.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper integrates a Conditional Continuous Normalizing Flow (CCNF) into YourTTS to manipulate perceptual voice qualities (PVQs) on a continuous scale, targeting roughness, breathiness, resonance, and weight. Speaker embeddings are transformed by conditioning on an attribute vector of PVQ severities (plus speaking rate), which is estimated by random forest regressors trained on the PVQD+ dataset and used to pseudo-label LibriTTS-R. The method is evaluated objectively with the same regressors and with EER speaker-similarity measures, and subjectively by phonetic experts for roughness and breathiness only. The authors report significant perceived changes in breathiness across all levels, significant roughness change only at the highest extrapolated setting, modest MOS degradation, and increasing EER with strong manipulation.","tokens_in":10044,"tokens_out":3500,"duration_ms":37495,"significance":"If the result holds, the system would be a useful tool for generating graded voice-quality examples for training speech pathologists and for data augmentation in atypical-voice processing, and it would extend TTS controllability beyond prosody and emotion to lower-level perceptual attributes without hand-crafted acoustic manipulations. The work has clear strengths: the CCNF approach avoids direct acoustic-correlate manipulation; the system is evaluated on both seen and unseen speakers; and the subjective expert listening test provides genuinely independent evidence for two of the four qualities. The main weakness is that the evidence for the other two qualities (resonance and weight) rests entirely on a circular objective metric, so the paper's broad conclusion that all four PVQs are 'effectively manipulated' is not yet supported.","major_comments":[{"comment":"The objective evaluation of resonance and weight is circular: the random forest regressors trained on PVQD+ are used both to generate the pseudo-labels for LibriTTS-R training data and to evaluate the manipulated speech in Table II. For resonance and weight, Table II is the only reported evidence, so the predicted severities largely measure whether the CCNF can invert the regressors' own mapping rather than whether listeners perceive the intended changes. The paper itself notes the compression of predicted scores relative to expert ratings (e.g., breathiness 11/19/35 vs. 24/46/70 in Tables II and III). Since no expert validation is provided for resonance or weight, the Section V conclusion that the system 'effectively manipulates PVQs' overreaches for these two qualities. Please provide independent perceptual evidence for resonance and weight, or explicitly restrict the perceptual claim to breathiness and roughness.","section":"Section III-B, Table II"},{"comment":"The roughness manipulation result does not support a continuous-control claim. Only the 'high' condition, which uses an extrapolated scale value of 200 beyond the CAPE-V maximum of 100, produces a statistically significant perceived change (p < 0.005); low and medium conditions are not significantly different from the original recording. The paper should either present the roughness claim as limited to strong extrapolated settings or provide additional evidence that intermediate levels are perceptually distinguishable. The choice of 0/100/200 is also an ad-hoc modification of the evaluation scale that deserves explicit justification.","section":"Section IV-A, Table III"},{"comment":"The pseudo-labeling pipeline is a load-bearing risk. Random forest regressors are trained on only 296 PVQD+ files (containing atypical voices) and then applied to 585 hours of typical LibriTTS-R speech. If the regressors' HuBERT-based features do not generalize from atypical to typical voices, the pseudo-labels themselves are misaligned with the perceptual qualities the system claims to control. The fact that predicted severities for original typical speech are strongly compressed relative to expert ratings (e.g., breathiness 11 vs. expert 24) suggests this risk is real. For breathiness and roughness, the expert listening test partly mitigates the concern, but for resonance and weight no external check exists. Please report the distribution of pseudo-labels on LibriTTS-R and, if possible, a validation of the regressors on typical-voice samples with expert ratings.","section":"Section III-B"}],"minor_comments":[{"comment":"The text refers to a 'paired Wilcoxon rank sum test'; this is a terminological inconsistency, as 'rank sum' (Mann-Whitney U) is for unpaired data while a paired test would be the Wilcoxon signed-rank test. Please correct the wording.","section":"Section IV-A"},{"comment":"The EER values for resonance and weight are strongly non-monotonic (e.g., seen resonance: 8.8, 1.4, 22.2; seen weight: 29.6, 1.7, 19.4). The explanation in Section IV-C about gender spoofing is plausible, but the table would benefit from confidence intervals or a statistical comparison to establish that the medium-condition drop is meaningful and not noise.","section":"Table II"},{"comment":"The definition of Δlogp(θ) appears to have a sign inconsistency: the initial condition sets log pZ1(z(t1)) - log pS(s|a; θ) = 0, yet this term is added to the trace integral. Please check the signs and clarify the relation between Δlogp and the standard instantaneous change-of-variables formula.","section":"Section II, Eq. (3)"},{"comment":"The listening-test design description is not fully consistent: 15 raters x 32 samples = 480 total, and with 4 speakers and conditions (roughness 3, breathiness 3, plus original) the per-speaker count appears to be 7, not 8. Please clarify the exact number of conditions per speaker.","section":"Section IV"},{"comment":"Reference [29] is cited for the choice of the 6th HuBERT layer, but [29] is the WavLM paper; please ensure the citation correctly supports the HuBERT feature claim, and align the text with the cited source.","section":"Section III-B"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within scope for a speech/audio conference and the core idea is interesting. My main concern is that the headline claim covers four PVQs while convincing evidence exists only for breathiness and, at the extreme setting, roughness. The circularity of the resonance/weight evaluation is the kind of issue that can be fixed by either adding a perceptual study for these qualities or carefully restricting the claims. The roughness issue is more subtle because the only significant effect uses an out-of-scale value; if the authors cannot demonstrate graded perception, they should reposition the contribution accordingly. I would not reject the paper, but the revision needs to either supply missing evidence or narrow the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a real step forward for perceptual voice quality control. The authors put a conditional continuous normalizing flow inside YourTTS and show zero-shot manipulation of four PVQs. Prior work either stuck to one speaker (GTR-Voice) or handled gendered qualities only (PerMod). Their breathiness results are credible: expert ratings rose significantly across low/medium/high, for both seen and unseen speakers. Roughness only worked at an aggressive extrapolated setting, which they own up to. The idea of learning manipulation from examples instead of tweaking jitter/shimmer is the right kind of approach.\n\nThe soft spots are real but not disqualifying. The objective table for resonance and weight uses the same random-forest regressors that generated the pseudo-labels used to train the flow. That's a fitted-value check, not an independent measure of perception. The stress-test note is correct: the paper claims 'effectively manipulates PVQs' for all four, but only two have expert ratings. I'd put that as the main referee ask—replace or supplement the circular metric with independent predictors or a listening test for resonance and weight. No code or demo audio is included, and there is no baseline run against PerMod or GTR-Voice, which would have made the novelty claim cleaner. The pseudo-labeling pipeline is risky—296 atypical-voice files to pseudo-label 585 hours of typical speech—but the authors acknowledge the resulting skew, and the monotonic trends are at least consistent.\n\nThe roughness scale extrapolated to 200 is arbitrary, but they justify it by data skew and the expert ratings show it was needed. I'd call that minor.\n\nOverall: the central mechanism is plausible, the expert data for two qualities is genuine evidence, and the limitations are stated rather than hidden. The paper deserves a serious referee; it should not be desk-rejected. My recommendation: send to review, request the artifacts, and push for an independent evaluation of resonance and weight before acceptance. It's a useful paper for speech-pathology training and data augmentation, and a good warning example for circular pseudo-label evaluation.","headline":"A solid, genuinely new voice-manipulation system whose breathiness and roughness claims are validated by experts, but whose resonance and weight claims rest on a circular metric.","tokens_in":10641,"tokens_out":1861,"would_cite":true,"duration_ms":19651,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A text-to-speech system can turn perceptual voice qualities — breathiness, roughness, resonance, and weight — up or down on a continuous scale by conditioning a normalizing flow on perceptual labels.","keywords":["perceptual voice qualities","text-to-speech","conditional continuous normalizing flows","voice modification","breathiness","roughness","resonance","speaker embeddings"],"falsifier":"Take a held-out set of ordinary read-speech utterances, have three professional voice-rating experts score breathiness and roughness on the 100-point clinical scale, and compare the ordering of their scores with the ordering of the pseudo-labels used in training. If the orderings disagree substantially, the conditioning values do not encode the claimed perceptual qualities, and the manipulation results would not generalize beyond the pseudo-label distribution.","tokens_in":9643,"feed_emoji":"🎚️","tokens_out":8171,"duration_ms":79557,"temperature":0.7,"pith_summary":"This paper aims to show that a text-to-speech system can manipulate perceptual voice qualities — breathiness, roughness, resonance, and weight — along a continuous severity scale, without directly editing acoustic correlates such as jitter or shimmer. The authors insert a conditional continuous normalizing flow between the speaker encoder and the decoder of an existing multi-speaker TTS model, and train the flow to move a speaker embedding according to an attribute vector. The attribute vectors come from statistical predictor models that pseudo-label a large read-speech corpus, so the model learns the manipulation from examples rather than from hand-designed acoustic rules. Phonetic experts rated the synthesized voices and found breathiness shifts convincing, while roughness required stronger settings; speaker similarity degrades as severity increases. If the approach holds, it gives clinicians and voice practitioners a label-driven way to generate graded voice-quality examples for training and data augmentation.","feed_headline":"One flow gives TTS a continuous dial for voice quality","feed_subtitle":"Phonetic experts confirmed gradable breathiness on speakers never heard in training.","key_machinery":"The load-bearing object is the Conditional Continuous Normalizing Flow (CCNF), a generative model that turns samples into a standard normal distribution by solving an ordinary differential equation, and can generate new samples by solving the ODE backwards under a changed conditioning vector. Here the conditioning vector has eight entries: seven perceptual voice qualities plus speaking rate. The flow is inserted between a fixed speaker-embedding extractor and the TTS decoder: it first encodes the original embedding under the original attribute vector, then decodes it under the manipulated vector, so the ODE trajectory itself carries the requested voice-quality change. This design is what lets the system manipulate qualities without ever touching acoustic correlates directly.","core_discovery":"The central claim is that a Conditional Continuous Normalizing Flow, trained to map speaker embeddings under an attribute vector, can act as a voice-quality dial inside a TTS system: at synthesis time the original speaker embedding is transformed to a latent Gaussian with the original attributes and then transformed back with the desired attributes, and the difference between those two paths is the modification. The paper demonstrates this on four perceptual voice qualities and reports that the manipulated severity, as predicted by the regressors, rises monotonically with the requested level, and that the shift is statistically significant for breathiness at all levels and for roughness only at strong modifications. The work also claims to generalize to unseen speakers, because the modifications operate on the speaker embedding rather than on the acoustic waveform, and to preserve audio quality except under extreme settings. Expert-rated breathiness reached roughly 70–79 on the 100-point clinical scale from a typical-voice baseline near 24, which the authors take as evidence that the label-driven manipulation is perceptually real.","pith_inferences":["A natural extension is that adding a new perceptual quality is mostly a matter of obtaining reliable predictor models and pseudo-labels; the flow-training pipeline itself need not be redesigned.","The observed covariation between roughness and breathiness could be exploited to study how perceptual voice qualities interact, since the flow lets one axis be moved while others are held fixed.","The same design pattern suggests a broader recipe: any attribute that can be pseudo-labeled on a large unlabeled speech corpus could become a controllable TTS dimension, as long as it leaves a trace in the speaker embedding."],"forward_implications":["Clinicians could generate graded breathy or rough voice examples for training speech pathologists, using ordinary read speech as the starting point.","Voice-quality editing can work for speakers never seen during training, because the manipulation lives in speaker-embedding space rather than in a per-speaker model.","Strong manipulations move the voice away from the original speaker, so the method can also prod how much of speaker identity is carried by voice quality.","The usable range differs per quality: breathiness moves convincingly at moderate settings, while roughness needs aggressive conditioning, so a deployment tool would need per-quality calibration."],"supporting_citations":[{"why":"It supplies the multi-speaker TTS backbone into which the manipulation block is inserted.","marker":"[19]"},{"why":"It supplies the perceptual voice quality database and the regression models used to pseudo-label the large training corpus and to evaluate severity.","marker":"[15]"},{"why":"It introduces the conditional continuous normalizing flow editing recipe that this paper adapts to voice-quality control.","marker":"[18]"},{"why":"It provides the attribute-conditioned flow block design and the forward-backward editing procedure used at inference.","marker":"[20]"},{"why":"It is the prior latent-diffusion system for perceptual voice quality modification that this work extends to non-gendered qualities.","marker":"[10]"},{"why":"It defines the 100-point clinical severity scales used by the expert raters in the listening test.","marker":"[11]"},{"why":"It supplies the normalizing flows foundation on which the conditional continuous flow is built.","marker":"[16]"}],"fun_headline_variants":["TTS dial for perceptual voice dimensions: rough to smooth","Continuous voice-quality control from examples, not acoustics","Flow-based TTS turns knobs on breathiness and roughness","Experts confirm gradable voice quality with unseen speakers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The labels used to train the voice-quality dial come from statistical models trained on only 296 atypical-voice recordings; if those models misjudge ordinary read speech, the dial is tuned to qualities that do not match what listeners hear.","fun_headline_variants_meta":{"raw":{"variants":["TTS dial for perceptual voice dimensions: rough to smooth","Continuous voice-quality control from examples, not acoustics","Flow-based TTS turns knobs on breathiness and roughness","Experts confirm gradable voice quality with unseen speakers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1377,"prompt_tokens":880,"completion_tokens":497,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":432}},"tokens_in":496,"tokens_out":497,"duration_ms":6023,"temperature":1.0,"reasoning_tokens":432,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:17:17.884614+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out set of ordinary read-speech utterances, have three professional voice-rating experts score breathiness and roughness on the 100-point clinical scale, and compare the ordering of their scores with the ordering of the pseudo-labels used in training. If the orderings disagree substantially, the conditioning values do not encode the claimed perceptual qualities, and the manipulation results would not generalize beyond the pseudo-label distribution.","supporting_citations":[{"cited_title":"YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero- Shot V oice Conversion for Everyone,","cited_arxiv_id":null,"evidence_quote":"It supplies the multi-speaker TTS backbone into which the manipulation block is inserted."},{"cited_title":"Towards an Interpretable Representation of Speaker Identity via Perceptual V oice Qualities,","cited_arxiv_id":null,"evidence_quote":"It supplies the perceptual voice quality database and the regression models used to pseudo-label the large training corpus and to evaluate severity."},{"cited_title":"StyleFlow: Attribute- conditioned Exploration of StyleGAN-Generated Images using Condi- tional Continuous Normalizing Flows,","cited_arxiv_id":null,"evidence_quote":"It provides the attribute-conditioned flow block design and the forward-backward editing procedure used at inference."},{"cited_title":"PerMod: Perceptually Grounded V oice Modification with Latent Diffusion Mod- els,","cited_arxiv_id":null,"evidence_quote":"It is the prior latent-diffusion system for perceptual voice quality modification that this work extends to non-gendered qualities."},{"cited_title":"Consensus Auditory-Perceptual Evaluation of V oice: Development of a Standardized Clinical Protocol,","cited_arxiv_id":null,"evidence_quote":"It defines the 100-point clinical severity scales used by the expert raters in the listening test."},{"cited_title":"Variational Inference with Normalizing Flows,","cited_arxiv_id":null,"evidence_quote":"It supplies the normalizing flows foundation on which the conditional continuous flow is built."}],"review_version":1}