{"id":"0acbc459-6ed2-47f4-8b70-81a37530d773","arxiv_id":"2608.08362","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"CtrlSpeech adds phone-level pitch, loudness, and duration controls to a diffusion-transformer TTS backbone, improving local prosody control while preserving zero-shot speaker similarity.","lead":"CtrlSpeech is a text-to-speech system that lets users adjust pitch, loudness, and duration at the level of individual phones while keeping a target speaker's voice. It builds on an existing autoregressive diffusion model (DiTAR) and reports competitive zero-shot voice cloning plus improved fine-grained prosody control.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Oracle-conditioned RMSE/MAE in Tables 3–4 show the model can reconstruct ground-truth prosody, not that arbitrary user edits are realized; no perturbation or listening test supports the coarse-to-fine control claim.","rationale":"The reader identified the same experimental weakness, and I think it is the right one to stress. The paper's contribution is explicitly control; the objective results for zero-shot TTS are secondary (and themselves rest on a self-reproduced DiTAR baseline without error bars). Once the control evaluation is seen as reconstruction fidelity, the paper's novelty reduces to conditioning DiTAR on oracle prosody tokens, which is a modest extension. I would keep the verdict CONDITIONAL because the architecture and zero-shot numbers are plausible and the missing evidence is obtainable; a perturbation/edit test and a listening test would resolve the concern. No ad hominem or circularity accusation: the issue is an evaluation gap, not an indication of fraud. I do not see a deeper internal inconsistency in the method; the patch-wise AR+diffusion formulation is standard and the conditioning scheme is described in enough detail to reproduce. The main risk is that the system cannot synthesize natural speech under non-oracle controls, which would invalidate the claimed UI workflow. Therefore verdict unchanged (CONDITIONAL) and agreement partial: same broad concern, but I would sharpen it to a concrete perturbation/isolation test and de-emphasize the 'copying' language, since copying an oracle condition is itself evidence of conditioning strength; the real gap is arbitrary-edit realization.","tokens_in":9095,"tokens_out":6376,"duration_ms":63666,"concrete_test":"On LJSpeech test utterances, extract the standard control signals, then apply deliberately non-ground-truth edits: e.g., increase the pitch bin of a target phone by +30 Mel, multiply its duration by 1.5, and lower its loudness by 3 dB, while leaving other phones untouched. Synthesize with CtrlSpeech (0.6B) and measure the extracted pitch/loudness/duration of the edited phone, the change in neighboring phones, and speaker similarity (SIM-o); also run a small CMOS test comparing naturalness of the edited utterance against the unedited synthesis. If edited phones track the edited controls and neighbors stay fixed, controllability holds; if the model reverts to ground-truth prosody, the Tables 3–4 RMSE gains are reconstruction fidelity, not user control.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central controllability evidence (Tables 3–4) is an oracle-reconstruction test. The model is fed phone-aligned pitch, loudness, and duration values extracted from the target utterance, and then scored by RMSE/MAE against those same ground-truth values. This establishes that the model can exploit strong oracle conditions to reproduce reference prosody, which is necessary for control but not sufficient. The paper's claimed user-facing workflow ('adjust expressive attributes at a fine temporal granularity,' §5.4) requires that arbitrary edits be realized: if a user raises pitch on one phone, stretches a duration, or changes loudness, the output should change in the intended direction on that phone, leave neighboring phones and other attributes approximately unchanged, and remain natural. None of these properties is measured. There is no perturbation test, no isolation test, and no human listening test for edited (non-ground-truth) controls. The duration numbers are especially ambiguous: because duration is supplied as a frame count and the model has an auxiliary stop-prediction loss (§3.4), reduced duration MAE may largely reflect the model counting frames to the specified length, not producing naturally paced, editable speech. The authors' own Limitations section concedes dependence on pitch-extraction and forced-alignment quality; noisy oracle controls would make the RMSE test even less diagnostic. Therefore the load-bearing claim of fine-grained controllability is currently under-supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CtrlSpeech, a controllable expressive TTS system built on the DiTAR autoregressive-diffusion backbone. It augments phone embeddings with phone-aligned pitch, loudness, and duration signals, and combines these with a global speaker embedding for zero-shot voice cloning. The authors report competitive zero-shot TTS results on LibriSpeech-PC test-clean and Seed-TTS test-en, and claim fine-grained controllability, supported by reduced pitch/loudness RMSE and phoneme-level duration MAE when ground-truth control signals are provided. The paper also describes a coarse-to-fine interactive refinement interface.","tokens_in":9360,"tokens_out":4997,"duration_ms":47940,"significance":"If the controllability claims hold, the paper makes a practically useful contribution: it shows that phone-level prosodic control and zero-shot speaker conditioning can be combined in a continuous autoregressive diffusion TTS system, with code and model weights promised. The zero-shot results at 0.6B are respectable, and the speaker-conditioning ablation in Table 2 clearly shows the complementary value of speaker embeddings and prompt speech. However, the central fine-grained controllability claim currently rests on oracle-conditioned reconstruction metrics (feeding ground-truth prosody and measuring error to the same values), which is necessary but not sufficient evidence for user-facing control. The paper would be strengthened by perturbation-based controllability tests, listening tests on edited controls, and statistical significance assessment.","major_comments":[{"comment":"The controllability evaluation is an oracle-reconstruction test. Pitch, loudness, and duration are extracted from the target utterance and fed as inputs, and RMSE/MAE are measured against those same target values. This demonstrates that the model can exploit strong ground-truth conditions to reproduce reference prosody, but it does not show that arbitrary user edits are realized locally. The paper should add perturbation experiments in which non-ground-truth controls are applied (e.g., raising pitch on a specific phone, stretching a single duration, or changing loudness) and verify that the output changes in the intended direction on that phone, leaves neighboring phones and other attributes approximately unchanged, and remains natural. A listening test for edited, non-oracle controls would directly support the coarse-to-fine control claim in §5.3–5.4.","section":"§4.4, Tables 3-4"},{"comment":"The duration MAE reduction may largely reflect the oracle duration condition combined with the stop-prediction loss (§3.4), which classifies each patch as first, middle, or last and can act as a frame-counting mechanism. The paper does not isolate whether the model produces naturally paced, editable speech or simply counts frames to the supplied length. The authors should analyze responses to non-oracle duration edits, for example by feeding a uniformly stretched or locally modified duration vector and reporting whether the output follows the edit, and by checking naturalness of the resulting pacing.","section":"§5.3, Table 4"},{"comment":"All objective and subjective results are single point estimates without confidence intervals or significance tests. The key zero-shot comparisons are close (WER 2.46 vs 2.55 on LibriSpeech-PC; CMOS -0.16 vs -0.22), so it is not clear whether CtrlSpeech is actually competitive with the reproduced DiTAR baseline beyond random variation. Per-utterance or bootstrap confidence intervals for WER/SIM-o and appropriate statistical tests for CMOS/SMOS are needed to support the 'competitive' claim.","section":"§5.1, Tables 1-2"},{"comment":"The coarse-to-fine iterative refinement workflow is a core claimed contribution but is never evaluated. The paper should present at least one experiment or case study showing a user making a coarse selection, refining local pitch/loudness/duration, and the output changing accordingly across rounds while preserving speaker identity. Without this, the interface description in §5.4 is unsupported by evidence.","section":"§5.4, Figure 1"},{"comment":"The pitch controllability result is not compared to DrawSpeech's 'With Sketch' condition on equal footing. DrawSpeech with sketch achieves 27.78 Hz pitch RMSE, which is lower than CtrlSpeech's 38.39 Hz with explicit control signals, although CtrlSpeech is better on loudness RMSE (4.56 vs 13.15 dB). Since DrawSpeech is the most relevant controllability baseline, the authors should either discuss this discrepancy or provide a matched comparison, rather than only contrasting CtrlSpeech with its own text-only setting.","section":"Table 3"}],"minor_comments":[{"comment":"The text says 'L1 flow-matching objective' but Eq. (4) is a squared L2 norm. Please clarify which loss is actually used and fix the inconsistency.","section":"§3.4 vs Eq. (4)"},{"comment":"Please specify the forced-alignment tool and the exact data filtering criteria for the Emilia and GigaSpeech subsets, including how the 20,000-hour total is distributed between the two corpora.","section":"§4.1"},{"comment":"The bin assignment for f0 values above 650 Hz is described as 'clipped to bin 127', while voiced frames map to bins 1–126. Clarify the boundary between bin 126 and bin 127, and whether 650 Hz itself maps to bin 126 or 127.","section":"§3.3.2"},{"comment":"The VAE encoder produces a posterior from which latent tokens are sampled, but it is not stated whether sampling is also used at inference or whether the mean is used. This affects reconstruction fidelity and should be clarified.","section":"§3.1"},{"comment":"The classifier-free guidance description says 'the time and speaker embeddings are treated as a unified conditioning signal'; please specify whether the prosodic control signals are also dropped during training for CFG or whether they are always provided.","section":"§4.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a system paper with potentially useful engineering contributions, but the central controllability evidence is currently too weak to support the load-bearing claim of fine-grained, user-editable control. The missing perturbation experiments and listening tests are in scope for a major revision. I also note that the DiTAR baseline is a self-reproduction from non-open-source checkpoints; the authors should provide enough detail (or release their reproduction) for reviewers and readers to judge the comparison. If the code/model release is not actually complete, the reproducibility claims should be moderated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a clean, useful integration job rather than a new paradigm. The genuinely new bit is feeding phone-aligned pitch, loudness, and duration into the DiTAR backbone alongside a global speaker embedding, and showing that zero-shot TTS quality roughly holds (WER 2.46% on LibriSpeech-PC, 2.58% on Seed-TTS test-en at 0.6B) while explicit prosodic conditions cut pitch RMSE roughly in half and duration MAE by more than half. That is a practical combination that hasn't been reported before, and it deserves credit.\n\nWhat it does well: the architecture is described clearly, the coarse/fine conditioning split is sensible, and the ablations (speaker condition removal, text-only vs. controlled) show the conditions are actually being used. The authors also honestly list limitations: English-only, dependence on extraction/alignment quality, and text-only prosody prediction being hard. That is the right tone.\n\nSoft spots, in proportion. The biggest one is the controllability evaluation. Tables 3–4 feed ground-truth phone-level pitch, loudness, and duration into the model and measure RMSE/MAE against those same ground-truth values. That is an oracle-reconstruction test: it shows the model can exploit strong conditions to reproduce the reference prosody, but it does not show that arbitrary user edits are realized. If a user raises pitch on one phone or stretches a duration, does the output change only there, in the intended direction, without degrading neighboring phones or naturalness? No perturbation test, no isolation test, and no listening test for edited (non-ground-truth) controls address this. The duration numbers are especially ambiguous because duration is provided as a frame count and there's an auxiliary stop-loss; reduced MAE might partly reflect the model counting frames to hit the target length rather than producing editable pacing. This is a genuine gap, but not a fatal one—the text-only ablations show large degradation when the controls are removed, so the conditions clearly carry signal. The paper just needs to test edited controls directly.\n\nMinor issues: the DiTAR baseline is a self-reproduction (fine, since DiTAR isn't open-source, but it adds uncertainty), there are no error bars or significance tests, and the abstract/§5.4 claims an iterative coarse-to-fine UI that is described but never evaluated. The paper says demo/code/weights are available but gives no URL—that's a reproducibility problem given the emphasis on practical control.\n\nBottom line: this is a legitimate, potentially useful system paper that deserves a serious referee. I'd send it to review, but the referee should insist on a perturbation test and ideally a released demo before publication. The main claim of fine-grained control is directionally supported but not yet fully demonstrated.\n\nFor my own work, I'd cite it as evidence that local prosodic control can coexist with zero-shot cloning in an autoregressive diffusion framework. Worth a reading group slot if the demo ever shows up.","headline":"Solid integration of phone-level prosody control into DiTAR, but the controllability test is oracle-reconstruction and needs perturbation evidence before the main claim is fully supported.","tokens_in":9957,"tokens_out":1410,"would_cite":true,"duration_ms":14674,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CtrlSpeech adds phone-level pitch, loudness, and duration control to zero-shot TTS without sacrificing voice cloning.","keywords":["text-to-speech","speech synthesis","prosody control","zero-shot voice cloning","autoregressive diffusion","phone-aligned conditioning","pitch control","flow matching"],"falsifier":"Perturb the supplied control signals during inference—shift pitch up one bin, stretch all durations by 20%, or replace the ground-truth contours with user-drawn sketches—and measure pitch, loudness, and duration RMSE against the edited targets; if the output does not track the edits, or if the oracle-conditioned gains vanish once the signals are not the same values used in training, the controllability claim collapses.","tokens_in":8865,"feed_emoji":"🎚️","tokens_out":9520,"duration_ms":71846,"temperature":0.7,"pith_summary":"CtrlSpeech is a text-to-speech system that aims to show that fine-grained prosody control and zero-shot voice cloning can live in the same model. It augments the DiTAR autoregressive-diffusion backbone with phone-aligned pitch, loudness, and duration signals, while a global speaker embedding plus prompt audio carries timbre. At 0.6B scale the system reaches 2.46% WER on LibriSpeech-PC test-clean and 2.58% WER on Seed-TTS test-en, matching or slightly beating the authors' DiTAR reproduction. When the phone-aligned signals are supplied, pitch RMSE falls from 67.86 to 38.39 Hz, loudness RMSE from 6.35 to 4.56 dB, and phoneme duration MAE from 28.08 to 11.86 frames. The intended payoff is an editing workflow where a user first generates an utterance and then iteratively refines local prosodic events without losing the target voice.","feed_headline":"Phone-level prosody control keeps zero-shot TTS quality","feed_subtitle":"Phone-aligned prosody cues cut pitch RMSE ~44% and duration MAE by more than half.","key_machinery":"The mechanism is the patch-wise diffusion-autoregressive factorization of DiTAR, where a causal autoregressive transformer emits a context representation for each patch of four continuous speech tokens and a local diffusion transformer denoises the next patch conditioned on that context. CtrlSpeech inserts control by augmenting phone embeddings with phone-aligned pitch, loudness, and duration tokens, and by adding a speaker embedding and prompt-speech condition to the autoregressive stream. The model is optimized with a flow-matching velocity objective plus an auxiliary loss that classifies each acoustic patch as first, middle, or last. The phone-aligned control stream is what lets the decoder attend to local prosodic targets, while the autoregressive context preserves long-range coherence across patches.","core_discovery":"The paper's central claim is that local prosodic control does not have to come at the cost of zero-shot synthesis quality. CtrlSpeech concatenates each phone embedding with quantized control tokens—pitch in 128 Mel-scale bins over 65–650 Hz, loudness in 64 bins over -60 to 0 dB, and duration in acoustic frames per phone—and conditions the utterance on a global speaker embedding and optional prompt speech. On LibriSpeech-PC test-clean the 0.6B model obtains 2.46% WER and 0.65 SIM-o, and on Seed-TTS test-en 2.58% WER and 0.63 SIM-o, both slightly ahead of the reproduced DiTAR backbone. With ground-truth control signals provided, pitch RMSE on LJSpeech drops from 67.86 to 38.39 Hz and loudness RMSE from 6.35 to 4.56 dB, while phoneme duration MAE on LibriSpeech-PC drops from 28.08 to 11.86 frames. The authors interpret these numbers as evidence that aligned low-level prosodic conditions give users direct local control while the global condition preserves timbre.","pith_inferences":["An implication the paper leaves implicit is that the quantized control tokens could be transplanted from one utterance to another, turning the model into a prosody-transfer tool: extract pitch, loudness, and duration contours from a reference recording and apply them to new text with a different speaker embedding.","Because the reported controllability metrics are measured against the same ground-truth signals fed into the model, a stricter test would have human users draw or tweak contours and check whether the audio follows the edited curves rather than the original recording; the paper does not run that experiment.","The coarse-to-fine conditioning scheme suggests a natural extension to additional phone-aligned attributes such as emotion tags, voice quality, or articulation rate, provided those attributes can be extracted or annotated at phone granularity.","One testable robustness extension is to corrupt or jitter the input control signals during inference and measure how gracefully pitch and duration RMSE degrade; the paper only reports the oracle-conditioned case."],"forward_implications":["At 0.6B scale, CtrlSpeech matches or slightly beats a reproduced DiTAR on zero-shot WER and speaker similarity on both LibriSpeech-PC test-clean and Seed-TTS test-en.","Providing phone-aligned pitch and loudness reduces frame-level RMSE substantially on LJSpeech, from 67.86 to 38.39 Hz and from 6.35 to 4.56 dB.","Providing phone-level target durations cuts phoneme duration MAE on LibriSpeech-PC from 28.08 to 11.86 frames, showing effective pacing control.","Combining speaker embedding and prompt speech outperforms either alone for speaker similarity (SIM-o 0.63 vs 0.53 for prompt-only), so the two coarse conditions are complementary.","The system's design supports iterative coarse-to-fine editing: generate with global conditions, then refine local pitch, loudness, and duration."],"supporting_citations":[{"why":"Supplies the DiTAR patch-wise autoregressive diffusion backbone that CtrlSpeech extends with control signals.","marker":"[1]"},{"why":"Defines the LibriSpeech-PC test-clean zero-shot benchmark and the F5-TTS baseline lineage.","marker":"[5]"},{"why":"CosyVoice provides the voiceprint-model approach for extracting the global speaker embedding.","marker":"[20]"},{"why":"WORLD vocoder used to extract the raw F0 contour before Mel scaling and quantization.","marker":"[22]"},{"why":"DIO algorithm performs the raw period extraction for F0, later refined by StoneMask.","marker":"[23]"},{"why":"Supplies the Seed-TTS test-en evaluation set for zero-shot synthesis.","marker":"[28]"},{"why":"Provides the Semantic-VAE codec checkpoints that map waveforms to 40 Hz continuous latent tokens.","marker":"[30]"},{"why":"Whisper-large-v3 is used to compute word error rates for intelligibility.","marker":"[32]"},{"why":"WavLM speaker verification model is used for the SIM-o similarity metric.","marker":"[33]"},{"why":"DrawSpeech serves as the comparison system for pitch and loudness controllability on LJSpeech.","marker":"[34]"}],"fun_headline_variants":["Fine-grained prosody control without zero-shot quality loss","CtrlSpeech: coarse-to-fine prosody control for expressive speech","44% lower pitch error, 58% lower duration error in TTS","Zero-shot TTS with precise local prosody control","Local prosody control for expressive speech synthesis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that feeding the model ground-truth phone-level pitch, loudness, and duration and then measuring error against those same ground-truth values is a faithful test of how well a user can control prosody; if the extracted control signals are noisy or the model is simply copying strong oracle inputs, the reported control gains may not carry over to real manual editing.","fun_headline_variants_meta":{"raw":{"variants":["Fine-grained prosody control without zero-shot quality loss","CtrlSpeech: coarse-to-fine prosody control for expressive speech","44% lower pitch error, 58% lower duration error in TTS","Zero-shot TTS with precise local prosody control","Local prosody control for expressive speech synthesis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000352,"raw_usage":{"total_tokens":1904,"prompt_tokens":914,"completion_tokens":990,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":908}},"tokens_in":530,"tokens_out":990,"duration_ms":8514,"temperature":1.0,"reasoning_tokens":908,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:06:48.039199+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Perturb the supplied control signals during inference—shift pitch up one bin, stretch all durations by 20%, or replace the ground-truth contours with user-drawn sketches—and measure pitch, loudness, and duration RMSE against the edited targets; if the output does not track the edits, or if the oracle-conditioned gains vanish once the signals are not the same values used in training, the controllability claim collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DiTAR patch-wise autoregressive diffusion backbone that CtrlSpeech extends with control signals."},{"cited_title":"CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis","cited_arxiv_id":"2608.08362","evidence_quote":"Defines the LibriSpeech-PC test-clean zero-shot benchmark and the F5-TTS baseline lineage."},{"cited_title":"V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,","cited_arxiv_id":null,"evidence_quote":"DIO algorithm performs the raw period extraction for F0, later refined by StoneMask."},{"cited_title":"Mela-tts: Joint transformer-diffusion model with representation alignment for speech synthesis,","cited_arxiv_id":null,"evidence_quote":"Supplies the Seed-TTS test-en evaluation set for zero-shot synthesis."},{"cited_title":"World: a vocoder-based high-quality speech synthesis system for real-time applications,","cited_arxiv_id":null,"evidence_quote":"DrawSpeech serves as the comparison system for pitch and loudness controllability on LJSpeech."}],"review_version":1}