{"id":"21e124c3-d96a-4bcb-aa50-978c34a44ec9","arxiv_id":"2507.00227","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Rectified flows with a tunable sampling temperature offer the best naturalness-diversity trade-off among stochastic prosody predictors for text-to-speech.","lead":"This paper compares three stochastic methods (normalizing flows, conditional flow matching, and rectified flows) for predicting pitch, energy, and duration in text-to-speech, against a deterministic system and cloned human speech. It finds that a rectified-flow predictor with a tunable sampling temperature gives the best practical balance between natural-sounding and varied prosody.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Diversity ratings in Section 3.4 are never significance-tested; the central claim that RF 0.8 offers the best naturalness–diversity trade-off rests on an untested ordering of boxplot medians.","rationale":"The reader's weakest_assumption is prosody-cloning fidelity and stimulus breadth. I agree those are limitations, but the single most load-bearing gap for the paper's headline is the absence of inferential statistics on the diversity dimension. The central claim is an ordinal ranking: RF 0.8 is the best on the diversity/naturalness frontier. Ordinal claims require evidence that the ordering is not noise. The authors collected 7680 ratings, a sufficiently powered dataset for pairwise tests; they used Kruskal-Wallis and Dunn for naturalness but omit the same analysis for diversity. This is not a subtle modeling assumption but a missing statistical inference directly supporting the conclusion.\n\nThe proposed test would settle the issue: run the same nonparametric procedure on the diversity ratings. If RF 0.8 is not significantly more diverse than CFM 0.8, the abstract's \"most favorable trade-off\" is overstated. If it is significant, the claim stands, and the remaining issue is the reader's cloning-fidelity concern, which is secondary. Either way, the paper benefits from the check.\n\nI keep the verdict CONDITIONAL (UNCHANGED) because the concern is addressable: the data already exist and the authors can supply the missing tests without new experiments. I do not see internal inconsistency in the method; the paper provides open-source code and objective evidence for temperature-controlled variance, which are real strengths. The issue is the statistical support for the subjective ranking.","tokens_in":8232,"tokens_out":5982,"duration_ms":56816,"concrete_test":"Re-analyze the raw diversity ratings from the subjective study (Section 3.4, Figure 6). Run a Kruskal-Wallis test across the eight conditions, then Dunn's posthoc with Bonferroni correction for all pairwise comparisons. Specifically compare RF 0.8 vs CFM 0.8 and RF 0.8 vs NF 0.8 on diversity, and RF 0.8 vs CFM 0.8 on naturalness. If the diversity differences are not significant at alpha=0.05, or if CFM 0.8 has statistically equivalent diversity with equivalent naturalness, the \"most favorable trade-off\" claim fails. The authors should also report the full pairwise matrix for both rating dimensions regardless of outcome.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion — that \"the RF based model offers the most favorable trade-off between the two desired properties\" (Section 4) — is derived from the subjective study in Section 3.4. The authors report Kruskal-Wallis and Dunn posthoc tests only for naturalness ratings (Figure 5), stating that RF at 0.4, CFM at 0.4, and the deterministic baseline \"tie for best performance\" and that RF 0.8 is \"on par with the human baseline.\" No equivalent significance test is reported for the diversity ratings (Figure 6). The claim that RF 0.8 \"produces the highest diversity after the human baseline\" is thus a descriptive statement about boxplot medians, not a statistically supported ranking. Without a pairwise comparison of diversity across conditions, the observed advantage of RF 0.8 over CFM 0.8 or NF 0.8 could be within sampling noise. Moreover, \"best trade-off\" is a joint claim: even if RF 0.8 were significantly more diverse, the trade-off is only favorable if its naturalness is not significantly worse than competing stochastic systems at comparable diversity. The manuscript does not report this comparison either. This concern is independent of the prosody-cloning fidelity issue: even if cloning were perfect, the headline ranking would remain unsupported.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares stochastic prosody predictors (Normalizing Flows, Conditional Flow Matching, Rectified Flows) against a deterministic baseline and human recordings in an explicit prosody modeling pipeline for TTS. The authors evaluate cascaded versus joint prediction of pitch, energy, and duration, the effect of predictor order, and the effect of sampling temperature on prosodic variance. Subjective ratings from 40 crowd-sourced raters measure naturalness and diversity of prosody across eight conditions. The main claims are that cascaded prediction benefits duration modeling, that sampling temperature steers prosodic diversity, that naturalness and diversity are inversely related, and that Rectified Flows at temperature 0.8 provide the most favorable naturalness–diversity trade-off. The paper provides open-source code and uses public datasets (LibriTTS, RAVDESS, ADEPT).","tokens_in":8486,"tokens_out":4749,"duration_ms":51986,"significance":"If the central claims hold, the paper provides a practically useful comparison for TTS systems that need explicit, controllable prosody: it identifies Rectified Flows as a strong stochastic predictor and demonstrates a temperature-based control mechanism for diversity. The subjective evaluation is thoughtfully designed, notably using prosody cloning to isolate prosody from voice and audio-quality differences, and the naturalness ratings are analyzed with non-parametric significance tests. The paper also contributes open-source code and a reproducible experimental pipeline, which strengthens its value for the community. However, the headline trade-off claim is currently supported only by a descriptive comparison of boxplot medians, because diversity ratings are not significance-tested, and the objective temperature-variance analysis lacks uncertainty quantification. These gaps mean the significance of the central conclusion remains partly unsubstantiated.","major_comments":[{"comment":"The central claim that the RF model at temperature 0.8 offers the most favorable naturalness–diversity trade-off is not statistically supported. The authors report Kruskal-Wallis and Dunn posthoc tests only for naturalness ratings (Figure 5), stating that RF 0.4, CFM 0.4, and the deterministic baseline tie for best naturalness and that RF 0.8 is on par with the human baseline. No equivalent significance test is reported for the diversity ratings in Figure 6. The assertion that RF 0.8 produces the highest diversity after the human baseline is therefore based on an untested ordering of boxplot medians, and the joint trade-off claim would require comparing RF 0.8 against other stochastic systems at comparable diversity while confirming its naturalness is not significantly worse. Please add significance testing for diversity ratings and a formal or at least clearly defined comparison for the trade-off claim.","section":"§3.4, Figures 5–6"},{"comment":"Table 1 reports Jensen–Shannon divergences without variance, confidence intervals, or any significance test, yet the text states that the difference between cascade and joint prediction is 'not significant' for pitch and energy and that ordering differences are 'never significant.' No statistical procedure is described for these conclusions. Since the cascade configuration is adopted for all subsequent experiments based on this finding, the claim that only duration prediction requires prior knowledge of other prosodic variables should be supported by an explicit test, for example a bootstrap over utterances or a repeated-generation experiment with reported intervals.","section":"§3.3, Table 1"},{"comment":"The claim that sampling temperature effectively controls prosodic variance rests on visual inspection of curves without error bars or uncertainty estimates. Figure 2 and Figure 3 show averages over 200 utterances, but no measure of variability across utterances or across random seeds is provided, so the reader cannot judge whether the monotonic trends are stable or whether differences between NF, CFM, and RF are meaningful. Please add confidence intervals or error bars, and ideally a quantitative measure such as the correlation between temperature and variance, to support the controllability claim.","section":"§3.3, Figures 2–3"},{"comment":"The abstract states that 'stochastic methods produce natural prosody on par with human speakers,' which is broader than the reported evidence: the significance tests identify only RF at 0.8 as statistically on par with the human baseline, and other stochastic conditions may be worse. Section 4 similarly concludes that 'Rectified Flows offer the best overall performance' without qualifying that this conclusion depends on the unsupported trade-off analysis. Please temper these statements to match the statistical findings, or add the missing analyses that would justify them.","section":"Abstract and §4"}],"minor_comments":[{"comment":"The notation in Equation (1) is inconsistent with standard flow-matching conventions: in the usual CFM formulation x1 denotes the data sample and x0 the noise, whereas the text says x1 is sampled from the noise distribution and x is an observation in data space. Please align the notation with the cited flow-matching literature to avoid confusion.","section":"§2.4, Eq. (1)"},{"comment":"The caption of Figure 4 says the orange distribution is from the CFM system, but the surrounding text says the synthetic distribution is generated with the RF-based model at temperature 0.8. This inconsistency should be corrected.","section":"§3.3, Figure 4"},{"comment":"The phrase 'exact prosody cloning' from [26] is asserted without any validation in this paper's setup. A sentence acknowledging this reliance on the prior method, or a brief perceptual check, would strengthen the interpretation of the human baseline.","section":"§3.4"},{"comment":"The statement 'We verified these results in informal listening tests' is vague; please describe the informal tests or remove the reference to them.","section":"§3.3"},{"comment":"The term 'Kruskall-Wallis' is misspelled; the correct spelling is 'Kruskal-Wallis.'","section":"Throughout"},{"comment":"The axis label 'Variance of Mean' is ambiguous; please clarify that it denotes the variance across utterances of the per-utterance mean pitch or duration.","section":"§3.3, Figures 2–3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid empirical comparison with a well-designed subjective protocol and useful open-source contributions. The main issue is that the headline trade-off claim lacks significance testing for diversity, which is fixable with additional analyses or a more cautious framing. The scope is appropriate for a speech/audio venue, and I do not see a novelty-disclosure concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look: a competent empirical comparison of normalizing flows, conditional flow matching, and rectified flows for explicit prosody prediction in TTS, with open code and a sensible human study design. The new content is real: nobody else has put RF into pitch/energy/duration prediction, and the temperature-as-diversity-knob plus cascade-versus-joint finding are genuinely useful for practitioners.\n\nThe subjective setup is the best part. Cloning human prosody into the synthetic voice to isolate prosody is a good trick, and they used 40 raters with 960 ratings per condition, which is reasonable for this subfield. Naturalness results are significance-tested with Kruskal-Wallis and Dunn, and the ordering (deterministic and low-temp stochastic tie; RF 0.8 on par with human) is supported.\n\nThe soft spot is the diversity ratings. They do not report any significance test for Figure 6. The claim that RF 0.8 offers the best naturalness-diversity trade-off rests on an ordering of boxplot medians, and the stress-test concern lands: without a pairwise comparison, the advantage of RF over CFM or NF at the same temperature could be noise. And 'best trade-off' is a joint claim; you need to show RF 0.8's naturalness is not worse than the alternatives at comparable diversity. They only show it is comparable to the human baseline, not to CFM 0.8 or NF 0.8. That is a load-bearing omission for the paper's central conclusion, but it is fixable.\n\nThe objective sections have smaller issues. Figures 2 and 3 lack error bars, and Table 1 reports JS divergences without variance or significance, so the 'difference is not significant' statements are not actually supported by the table. The abstract also overstates: 'natural prosody on par with human speakers' is really only RF 0.8 in one condition, and human samples did not even score highest on naturalness (the deterministic did). That should be softened.\n\nThe math and code look fine. The authors reuse ToucanTTS as infrastructure and cite their own prosody-cloning method, but the central comparisons are against external human ratings and public datasets, so circularity isn't a concern. This paper deserves a serious referee; the empirical work is useful and the flaws are addressable rather than fatal. I would send it out, with the expectation that the authors add proper statistics for diversity and error bars where they make quantitative claims.","headline":"Useful empirical comparison of stochastic prosody predictors, but the headline trade-off claim lacks significance tests on diversity ratings.","tokens_in":9024,"tokens_out":1974,"would_cite":true,"duration_ms":20166,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Rectified flows give the best prosody trade-off in TTS, claim authors","keywords":["prosody modeling","speech synthesis","rectified flow","flow matching","normalizing flow","sampling temperature","explicit prosody prediction","text-to-speech"],"falsifier":"Re-running the listening study with more speakers, more sentences, and a different rater pool could refute the ranking if Rectified Flow at temperature 0.8 is no longer judged as diverse as the CFM baseline at equal naturalness; specifically, if CFM at 0.8 is rated as diverse and natural as RF 0.8, the paper's central trade-off claim fails. A second, more direct check is to measure whether the cloned human prosody matches the original recordings' pitch, energy, and duration distributions; a mismatch there would weaken the human baseline itself.","tokens_in":8065,"feed_emoji":"🎚️","tokens_out":9158,"duration_ms":82721,"temperature":0.7,"pith_summary":"The paper claims that stochastic generative models can predict the prosody of an utterance — its pitch, energy, and duration contours — as naturally as human speakers while preserving the one-to-many variability that deterministic systems collapse. It compares Normalizing Flows, Conditional Flow Matching, and Rectified Flows against a deterministic baseline and against cloned human recordings, using both distributional distances and a 40-rater listening study. The central result is that a Rectified Flow predictor offers the most favorable trade-off between prosodic naturalness and diversity: at a sampling temperature of 0.8 it produces the widest variety after the human baseline while remaining statistically indistinguishable from human naturalness. The paper also establishes that a single scalar, the sampling temperature, lets the user steer this trade-off at inference time, turning a monotone most-likely reading into varied renditions of the same sentence.","feed_headline":"Rectified flows win the naturalness-diversity trade-off for prosody","feed_subtitle":"At temperature 0.8, rectified flow matches human naturalness with the widest variety of renditions.","key_machinery":"The central object is a probabilistic variance predictor that treats the mapping from an encoded phoneme sequence to a prosodic contour as a transport problem between a Gaussian noise distribution and the distribution of valid contours, rather than predicting a single averaged contour. Three transport mechanisms are compared: Normalizing Flows (invertible transformations applied in one step), Conditional Flow Matching (a time-dependent vector field trained with a flow-matching objective to approximate an optimal-transport path), and Rectified Flows (the same vector-field approach, but with a ReFlow post-training stage that replaces arbitrary noise–data couplings with deterministic ones, straightening the learned paths). During inference the noise is scaled by a sampling temperature before the flow is solved, which is the mechanism that makes the variance of generated contours controllable. A secondary mechanism is the cascaded predictor structure, energy → pitch → duration, where each predictor is conditioned on previously predicted prosodic features; this ordering outperforms joint prediction for duration, and the order of pitch and energy has negligible impact.","core_discovery":"On the paper's own terms, the discovery is that explicit prosody prediction need not sacrifice expressivity for controllability. A prosody predictor trained as a Rectified Flow — a flow that is post-trained with the ReFlow procedure to couple noise and data deterministically and straighten the transport path — produces pitch, energy, and duration contours whose naturalness is indistinguishable from human recordings while the diversity across repeated takes of the same sentence approaches the human range. The evidence is a combination of objective Jensen–Shannon divergence comparisons between predicted and human contour distributions and a rated listening study. The rated study also reveals that naturalness and diversity are inversely related even for human recordings, so the trade-off is inherent to the task rather than a model defect. The paper concludes that Rectified Flows offer the best overall performance for prosody modeling among the studied approaches, with sampling temperature as the effective control over the naturalness–diversity trade-off.","pith_inferences":["The same cascade-plus-RF recipe could plausibly be carried over to conversational and spontaneous speech, where the one-to-many prosody problem is more acute; the paper limits its experiments to read speech and does not test this.","The observed unimodality and narrower spread of the synthetic contour distributions suggest that scaling up the prosody predictor's capacity, or using a more structured noise prior, might close the remaining diversity gap to human bimodal distributions.","The prosody-cloning evaluation protocol could serve as a general harness for comparing any stochastic prosody generator, since it isolates prosody as the only variable while holding voice and audio quality fixed.","A calibrated perceptual diversity scale, built from the temperature–variance mapping the paper reports, could turn the raw sampling temperature into a user-friendly setting for content producers."],"forward_implications":["A TTS system with an RF prosody predictor can produce multiple takes of the same sentence that sound like a human speaker saying it differently, with a single temperature knob controlling how different the takes are.","Cascading the predictors (energy → pitch → duration) is preferable to predicting all three jointly, at least for duration; the order of pitch and energy can be chosen freely without hurting quality.","The inverse naturalness–diversity relationship appears even in human recordings, so no model is expected to maximize both simultaneously; the practical target is a controllable point on the trade-off curve.","Because temperature scaling produces a near-logarithmic variance response, the user-facing control can be made intuitive by exposing a log-temperature dial rather than a raw temperature."],"supporting_citations":[{"why":"Defines the conditional flow-matching objective that the CFM prosody predictor is trained with.","marker":"[16]"},{"why":"Introduces the ReFlow post-training procedure that turns a CFM model into a Rectified Flow with straightened transport paths.","marker":"[17]"},{"why":"Prior work showing stochastic duration models benefit TTS; the paper's CFM setup and objective evaluation follow its methodology.","marker":"[21]"},{"why":"Provides the prosody-cloning procedure used to convert human recordings to synthetic samples, isolating prosody as the only variable in the subjective study.","marker":"[26]"},{"why":"LibriTTS is the training corpus whose read-speech prosodic variation supports learning expressive prosody.","marker":"[30]"},{"why":"RAVDESS supplies multiple emotional realizations of the same sentences, used as the human distribution for objective JS-divergence comparisons.","marker":"[31]"},{"why":"ADEPT's topical-emphasis subset supplies the multiple diverse takes per sentence used in the human naturalness and diversity ratings.","marker":"[32]"}],"fun_headline_variants":["Rectified flows match human prosody with tunable diversity","Flow-based prosody hits human-level naturalness and variety","Stochastic prosody: human-like output, temperature-controlled variety","Rectified flow prosody: natural as human, tuneable diversity","Explicit prosody gets stochastic: human-level naturalness, wider variety"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison rests on the assumption that the prosody-cloning procedure used for the human baseline recreates the exact prosody of a human recording in a synthetic voice without introducing audible artifacts, so that ratings of the cloned samples reflect prosody alone.","fun_headline_variants_meta":{"raw":{"variants":["Rectified flows match human prosody with tunable diversity","Flow-based prosody hits human-level naturalness and variety","Stochastic prosody: human-like output, temperature-controlled variety","Rectified flow prosody: natural as human, tuneable diversity","Explicit prosody gets stochastic: human-level naturalness, wider variety"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1185,"prompt_tokens":860,"completion_tokens":325,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":238}},"tokens_in":476,"tokens_out":325,"duration_ms":3922,"temperature":1.0,"reasoning_tokens":238,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:21:01.764191+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the listening study with more speakers, more sentences, and a different rater pool could refute the ranking if Rectified Flow at temperature 0.8 is no longer judged as diverse as the CFM baseline at equal naturalness; specifically, if CFM at 0.8 is rated as diverse and natural as RF 0.8, the paper's central trade-off claim fails. A second, more direct check is to measure whether the cloned human prosody matches the original recordings' pitch, energy, and duration distributions; a mismatch there would weaken the human baseline itself.","supporting_citations":[{"cited_title":"Variational inference with normal- izing flows,","cited_arxiv_id":null,"evidence_quote":"Prior work showing stochastic duration models benefit TTS; the paper's CFM setup and objective evaluation follow its methodology."},{"cited_title":"Varianceflow: High-quality and controllable text-to-speech using variance information via nor- malizing flow,","cited_arxiv_id":null,"evidence_quote":"Provides the prosody-cloning procedure used to convert human recordings to synthetic samples, isolating prosody as the only variable in the subjective study."},{"cited_title":"This dataset is com- prised exclusively of read speech in English and features 2,456 speakers","cited_arxiv_id":null,"evidence_quote":"LibriTTS is the training corpus whose read-speech prosodic variation supports learning expressive prosody."},{"cited_title":"Con- former: Convolution-augmented Transformer for Speech Recog- nition,","cited_arxiv_id":null,"evidence_quote":"RAVDESS supplies multiple emotional realizations of the same sentences, used as the human distribution for objective JS-divergence comparisons."},{"cited_title":"Exact Prosody Cloning in Zero- Shot Multispeaker Text-to-Speech,","cited_arxiv_id":null,"evidence_quote":"ADEPT's topical-emphasis subset supplies the multiple diverse takes per sentence used in the human naturalness and diversity ratings."}],"review_version":1}