Pith. sign in

REVIEW 3 major objections 6 minor 40 references

Descriptor-conditioned speech generation frequently changes non-target acoustic and prosodic traits, even in outputs that clearly respond to the prompt; a training-free reranker, VoDER-Cal, reduces this collateral movement within a three-ca

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A paired audit of three TTS systems shows that descriptor-aligned voice changes come with off-target acoustic shifts, and a candidate selector reduces these shifts at inference time.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection Useful paired audit of TTS attribute control; the 'off-target' finding is contingent on author-defined feature sets, but the paper is transparent and the candidate selector is a reasonable contribution. the 3 major comments →

arxiv 2608.00545 v1 pith:LTTMRGW5 submitted 2026-08-01 cs.SD

Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation

classification cs.SD
keywords controllable speech generationvoice descriptorsprompt adherenceattribute preservationpaired auditinference-time rerankingTTS evaluationacoustic analysis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Text-to-speech models that accept natural-language voice descriptions (like 'deep' or 'bright') are usually evaluated on whether the output matches the prompt. This paper asks the complementary question: does the system change only the requested attribute, or does it also alter other vocal characteristics? Through a paired audit of 5,940 outputs from three reference-conditioned systems, the authors find that outputs moving in the expected target direction frequently also change non-target acoustic and prosodic features, and that this coupling persists even when the target response clearly exceeds baseline seed variation. The paper then introduces VoDER-Cal, a training-free candidate selector that keeps a strong target response while favoring candidates with smaller off-target deviation, reducing held-out off-target deviation from 0.344 to 0.276 within a three-candidate budget. The upshot is that prompt adherence and attribute preservation are distinct capabilities that evaluation should measure separately.

Core claim

On the paper's own terms, the central discovery is that natural-language voice descriptors produce coupled, not isolated, acoustic changes. Across CosyVoice3, VoxCPM2, and Fish-Speech-S2, the same descriptor—deep, bright, or rough—is realized through different combinations of pitch, timing, energy, and spectral movement, and the majority of target-responsive outputs exhibit at least one substantial off-target change: between 54.5% and 95.8% in the eight settings with enough data. This collateral movement survives conditioning on outputs whose target response exceeds the seed-to-seed noise floor, so it is not simply a failed attempt to express the descriptor. The paper further claims that a t

What carries the argument

The paired audit design: each descriptor-conditioned output is compared against a matched neutral baseline generated from the same system, reference speaker, text, and random seed, with changes measured across eight signal-level features. Descriptor-specific target sets (deep: lower F0 and lower 85% roll-off; bright: higher spectral centroid and higher roll-off; rough: higher spectral flatness and higher zero-crossing rate) define what counts as on-target, so every other measured change is off-target for the audit. VoDER-Cal builds on this by computing a direction-aligned target score (a normalized average of the target-feature changes), filtering candidates that pass content, speaker, and t

Load-bearing premise

The audit's central claim rests on the paper's own mapping from each descriptor to a fixed set of measurable signal features (deep: F0 and rolloff; bright: centroid and rolloff; rough: flatness and zero-crossing rate); if a word like 'deep' perceptually includes spectral darkening or 'rough' includes other qualities, then changes the paper labels off-target may actually be part of what listeners mean by the word.

What would settle it

Run the same paired audit on a system whose descriptor-to-feature mapping is first validated perceptually for each word, and measure off-target movement on listener-defined non-target dimensions. If, among outputs with an above-noise target response, fewer than ~10% show substantial movement on any non-target dimension, the paper's central coupling claim would collapse for that system. Conversely, if a perceptual preservation metric shows no advantage of VoDER-Cal over target-only selection, the method's benefit would be shown to be an artifact of its acoustic proxy.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Benchmarks for instruction-following TTS should combine prompt-adherence scores with matched-baseline preservation metrics; a high match score can coexist with large collateral acoustic changes.
  • The same descriptor maps to different physical realizations across systems, so cross-model comparisons of controllability need feature-level measurements, not only global preference or semantic agreement.
  • Inference-time reranking is an effective training-free lever: a three-candidate pool raises the joint success rate (target, content, speaker, preservation) from about 4.8% with a single sample to roughly 14% regardless of the ranking rule, and VoDER-Cal specifically improves preservation at the same budget.
  • Training losses that jointly reward target expression and penalize off-target change are a natural extension, and the paper's paired-audit protocol supplies the measurement template for such losses.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's descriptor-to-feature mappings (e.g., deep = F0 and rolloff) are signal-level operational choices; a perception-validated mapping for each word could strengthen or qualify the off-target verdicts for that word.
  • Because VoDER-Cal's selector and evaluator features come from the same correlated acoustic family, a preservation signal trained on listener judgments of 'same speaker, same pace, same loudness' might yield larger and more perceptually meaningful gains than the current held-out feature set.
  • Neutral and nonsense control prompts themselves perturb several features in some systems, suggesting that conditioning language alone can shift the voice; this underscores the paper's matched-baseline protocol as the appropriate standard for future controllability evaluation.
  • If the coupling pattern generalizes beyond these three systems, descriptor-based voice control may be too coarse for applications that need repeatable, localized edits (e.g., audiobook narration), making either candidate reranking or model-level disentanglement necessary.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a paired, preservation-sensitive audit of three reference-conditioned speech-generation systems (CosyVoice3, VoxCPM2, Fish-Speech-S2). For each descriptor (deep, bright, rough), the authors define narrow signal-level target sets (Eqs. 10–12) and measure, relative to matched neutral baselines, both target-aligned movement and changes in non-target acoustic/prosodic features. Across 5,940 outputs, they report that target responses are frequently accompanied by off-target changes, including after conditioning on above-noise target response (Sec. 4.2). They then introduce VoDER-Cal, a training-free candidate selector that retains strong target responses while minimizing selector-side off-target deviation, and report automatic and listener-based evaluations showing reduced held-out off-target deviation relative to target-only selection (Table 2, Fig. 3).

Significance. If the central claim holds, the paper makes a useful methodological contribution by distinguishing prompt adherence from attribute-preserving control and by demonstrating that current systems frequently couple requested changes with collateral acoustic/prosodic movement. The study design is careful in several respects: a factorial generation matrix, matched neutral baselines, request-level bootstrap inference, calibration/evaluation splits, pre-defined target sets, and a shared code repository. The VoDER-Cal method is training-free and uses held-out features for evaluation, which mitigates direct overfitting concerns. The listening study, despite its limited size, provides some perceptual grounding. However, the strongest interpretive claims—that changes are genuinely 'off-target' and that VoDER-Cal improves preservation—depend on load-bearing operational and statistical choices that need further support or more cautious framing.

major comments (3)
  1. [Sec. 3.2, Eqs. (10)–(12); Sec. 7.3] The central audit claim that systems produce changes 'outside the requested attribute' is entirely relative to the paper's narrow signal-level target sets. For example, 'deep' is defined only as F0↓ and rolloff↓; perceptually, a deep voice may naturally include spectral darkening, slower rate, or lower energy, which the paper counts as off-target. Similarly, 'rough' is acknowledged to be weakly captured by flatness and ZCR. The Limitations section transparently calls these 'operational evaluation choices,' but the abstract and conclusion use the stronger wording 'characteristics outside the intended change' and 'preserves less than prompt adherence suggests.' The H1 listening study does not resolve this: listeners were asked to judge the requested attribute and 'non-target' dimensions selected using the same mapping, not whether listeners consider those dimensions part of the requested a
  2. [Sec. 5.2, Table 2; Sec. 13.2, Table 7] The headline VoDER-Cal result—reduction in mean held-out off-target deviation from 0.344 (Target-only) to 0.276 (VoDER-Cal)—is reported without a confidence interval or significance test. The paper provides request-level bootstrap CIs for the binary joint-success difference (Table 7: +0.19 points, CI [−0.56, 0.93]), which is statistically indistinguishable from zero, but the continuous deviation comparison, which is the main quantitative evidence for the method's benefit, has no such interval. Given that the macro average is over only nine system–descriptor settings and the eligible-request count is modest, this could be sampling noise. Please provide request-level clustered bootstrap CIs for the paired difference in held-out deviation (and preferably per-system results), or explicitly relegate the deviation reduction to a descriptive finding if inference is not feasible.
  3. [Sec. 4.2, Fig. 2b; Sec. 6, Fig. 4] The conditional analysis defines a 'substantial off-target deviation' as an absolute normalized change reaching 0.5 of a fixed physical scale (Sec. 11.3). This threshold is arbitrary and, more importantly, the off-target features themselves are the same narrow set from Eqs. (10)–(12). The claim that '54.5%–95.8% of responsive outputs exhibit at least one substantial off-target deviation' is therefore a statement about the paper's chosen features and threshold, not about perceptually distinguishable unwanted changes. The H2 listening study partially addresses perceptibility, but it only compares VoDER-Cal vs. Target-only on four pre-selected dimensions; it does not validate the off-target label. I recommend either adding a sensitivity analysis over the off-target threshold and feature set, or softening the wording to 'measured off-target changes' throughout the results section.
minor comments (6)
  1. [Throughout] The typeset name 'V oDER-Cal' contains an odd spacing artifact (also in the abstract and figures); please unify it as 'VoDER-Cal'.
  2. [Sec. 3.2 / Table 4] The fixed normalization scales are presented without justification for their physical ranges (e.g., 3 semitones for F0, 0.8 words/s for rate). A brief note on how these were chosen (prior literature? pilot data?) would improve reproducibility.
  3. [Sec. 4.1 / Fig. 2a] The heatmap values are means over paired changes normalized by fixed scales; because several features are correlated (e.g., speaking rate and duration, centroid and rolloff), the 'Non-target CIs' count can overstate breadth. The paper acknowledges this in Sec. 3.2, but it would help to include a correlation matrix or to group correlated features in the displayed results.
  4. [Sec. 6 / Fig. 4a] H1's request-level estimates (51.1% for both target and non-target changes) have wide CIs (roughly 38–67%) based on only 45 request-tasks. The conclusion that target and off-target changes are 'perceptually noticeable' is directionally supported but quantitatively fragile; please explicitly note this weakness in the main text, not only in the supplementary material.
  5. [Sec. 13.4 / Table 9] The sensitivity table shows that ρ=0 reduces held-out deviation to 0.269 while increasing eligibility, but with substantially lower target retention (87.9%). The prose says the main ρ=0.75 'retains 99.1%' but does not discuss the trade-off visible in the table; a sentence interpreting this would help readers understand the operating point.
  6. [Sec. 10.2 / Equation (8)] The equation for total outputs (3×(6+4+1)×6×10×3 = 5,940) is correct, but the derivation is compressed; it would be clearer to write 3 systems × 11 conditions × 6 speakers × 10 texts × 3 seeds.

Circularity Check

0 steps flagged

No circularity: the audit and VoDER-Cal evaluation are empirical and do not reduce to their inputs.

full rationale

The paper's central claims are empirical measurements of paired neutral vs descriptor-conditioned outputs from three external systems. The descriptor-target sets (Eqs. 10-12) are explicit operational definitions, and the paper repeatedly labels them as signal-level proxies rather than complete perceptual decompositions (Sec. 3.2, 7.3, Supp. 15), so the off-target finding is a conditional empirical result, not a tautology. VoDER-Cal is evaluated against baselines using a non-overlapping held-out feature set, with thresholds and feature partitions set on a calibration split and applied unchanged (Sec. 5.1); the selection criterion is not identical to the evaluation metric. There are no load-bearing self-citations or imported uniqueness theorems; all cited systems and metrics are external. The acknowledged limitations—narrow target definitions, correlated features, limited listening studies—are assumption risks, not circular derivations.

Axiom & Free-Parameter Ledger

6 free parameters · 4 axioms · 0 invented entities

The central audit relies on hand-defined target sets, normalization scales, and thresholds. None of these are fitted to produce a pre-specified conclusion, but they are author-chosen and the conclusions about 'off-target' movement are contingent on them. VoDER-Cal introduces a new selection rule but no new underlying entity.

free parameters (6)
  • Fixed normalization scales (Table 4) = F0: 3 st, rate: 0.8 w/s, centroid: 220 Hz, rolloff: 450 Hz, flatness: 0.03, ZCR: 0.03, RMS: 0.04, duration: 1.0 s
    Hand-chosen scales define the normalized magnitude of every feature change; they determine which changes count as 'substantial' in the conditional analysis and the off-target deviation in VoDER-Cal.
  • Substantial off-target threshold = 0.5 of fixed scale
    An off-target deviation is called substantial when its absolute normalized magnitude reaches 0.5. This threshold is chosen by the authors and affects the conditional percentages in Fig. 2b.
  • Target-retention fraction rho = 0.75
    VoDER-Cal retains candidates that keep at least 75% of the strongest target score in the pool. Sensitivity analysis in Table 9 shows this choice affects eligibility and preservation.
  • Candidate budget B = 3
    The three-candidate pool is the main operating point; the paper reports only B=3 in detail.
  • Target-noise threshold tau_noise = estimated from baseline seed variation on calibration split
    The one-sided target-direction distribution of paired baseline seeds sets the noise threshold for target responsiveness and feasibility. This is a data-derived parameter.
  • Content and speaker validity thresholds = binary thresholds calibrated on the calibration split
    Binary content (ASR transcript match) and speaker (ECAPA-TDNN cosine) thresholds are set on calibration data and applied to evaluation. These gates determine feasibility.
axioms (4)
  • domain assumption The three evaluated systems (CosyVoice3, VoxCPM2, Fish-Speech-S2) are representative of current reference-conditioned controllable TTS.
    The empirical conclusion is limited to these systems and may not generalize to all systems (Sec. 7.3).
  • ad hoc to paper The signal-level target sets in Eqs. (10) to (12) correctly operationalize deep, bright, and rough for the audit.
    These sets are defined by the authors and the paper notes they are not exhaustive perceptual models; if a descriptor perceptually includes other features, the off-target classification changes.
  • standard math Request-level block bootstrap with 10,000 replicates yields valid confidence intervals for the clustered data.
    The paper resamples system-descriptor-speaker-text blocks to account for seed dependence (Sec. 11.3).
  • domain assumption ECAPA-TDNN cosine similarity is a valid proxy for speaker preservation.
    Used as a validity gate; x-vectors provide a cross-check but the primary gate relies on this embedding model.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation." pith.science (2026). https://pith.science/paper/LTTMRGW5

@misc{pith2026260800545,
  author       = {Pith},
  title        = {Pith review of: Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LTTMRGW5}},
  note         = {Machine review of arXiv:2608.00545}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Natural-language descriptions have become a flexible interface for controlling generated speech. Existing evaluations largely assess whether an output matches a prompt, but prompt matching alone does not reveal whether characteristics outside the intended change remain stable. We examine this distinction through a controlled paired audit of three speech-generation systems: CosyVoice3, VoxCPM2, and Fish-Speech-S2. The evaluation contains 5,940 outputs spanning six reference speakers, ten texts, three random seeds, and eleven conditions. Using acoustic, prosodic, content, and speaker measurements, we find that responses in the expected target direction are frequently accompanied by changes outside descriptor-specific signal-level target sets. This pattern remains among outputs whose target response exceeds baseline seed variation, and the accompanying changes differ substantially across systems. We further introduce VoDER-Cal, a training-free candidate selector that retains sufficiently strong target responses while favoring smaller off-target deviations. A three-candidate pool raises the joint success rate from 4.8% under single-sample direct generation to approximately 14% for all candidate-selection policies. Within the matched three-candidate budget, VoDER-Cal reduces held-out off-target deviation from 0.344 under target-only selection to 0.276 and improves listener-rated preservation. Preservation-sensitive evaluation therefore complements prompt-adherence evaluation, while candidate reranking offers a practical inference-time improvement. Code, configuration files, and analysis scripts are available at https://github.com/intelland/VoDER

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

40 extracted references · 32 canonical work pages · 4 internal anchors

  1. [1]

    Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation

    INTRODUCTION Natural language allows users to control generated speech without relying on predefined style labels or low-level acoustic parameters. A desired voice can be described using attributes such asdeep,bright, or rough, or through more detailed free-form instructions. PromptTTS, InstructTTS, PromptTTS 2, PromptStyle, and PromptTTS++ demon- strated...

  2. [2]

    Natural-language control of speech Controllable speech generation has traditionally relied on either global style representations or explicitly defined acoustic factors

    RELATED WORK 2.1. Natural-language control of speech Controllable speech generation has traditionally relied on either global style representations or explicitly defined acoustic factors. Reference-based prosody transfer and global style tokens encode speaking characteristics in learned latent representations [ 14, 15], while architectures such as FastSpe...

  3. [3]

    Systems and generation matrix We evaluate CosyV oice3, V oxCPM2, and Fish-Speech-S2, three in- dependently developed reference-conditioned speech-generation sys- tems

    PAIRED EV ALUATION OF ATTRIBUTE-LEVEL CONTROL 3.1. Systems and generation matrix We evaluate CosyV oice3, V oxCPM2, and Fish-Speech-S2, three in- dependently developed reference-conditioned speech-generation sys- tems. Each system is tested using the same factorial matrix of six reference speakers, ten English texts, three random seeds, and eleven conditi...

  4. [4]

    Target dir

    VOICE DESCRIPTORS PRODUCE COUPLED CHANGES 4.1. Models respond, but not through isolated changes Figure 2a shows that the three systems respond systematically to the tested descriptors, but their responses frequently extend beyond the corresponding target features. Fordeep, outputs move in the expected target direction in 84.4% of CosyV oice3 pairs, 77.2% ...

  5. [5]

    Eligible

    VODER-CAL: PRESERV ATION-A W ARE CANDIDATE SELECTION The paired audit reveals substantial variation among candidates gen- erated for the same request. Different candidates can achieve sim- ilar target responses while exhibiting markedly different off-target changes. V oDER-Cal exploits this variation to improve attribute preservation without retraining th...

  6. [6]

    H1 evaluates whether target and off-target changes are perceptible relative to a matched baseline, while H2 compares the candidates selected by V oDER-Cal and Target-only selection

    AUXILIARY PERCEPTUAL CHECK We conduct two blinded listening studies with ten adult listeners. H1 evaluates whether target and off-target changes are perceptible relative to a matched baseline, while H2 compares the candidates selected by V oDER-Cal and Target-only selection. Each study con- tains 45 samples balanced across the three systems and three prim...

  7. [7]

    Content preserved 2 3 Select cleaner output 1 3 Smaller non-target change (b) Joint success rate Direct Random T arget-only VoDER-Cal Oracle 0.0 2.5 5.0 7.5 10.0 12.5 15.0Requests satisfying all checks (%) 4.8 4.7 14.1 14.3 14.6 (c) Candidate ranking Direct Random T arget-only VoDER-Cal Oracle 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35Held-out off-target dev...

  8. [8]

    Prompt adherence and attribute preservation are comple- mentary Prompt adherence and attribute preservation capture different aspects of controllability

    DISCUSSION 7.1. Prompt adherence and attribute preservation are comple- mentary Prompt adherence and attribute preservation capture different aspects of controllability. Prompt-adherence evaluation asks whether gener- ated speech expresses the requested concept, whereas preservation- sensitive evaluation asks whether characteristics outside that concept r...

  9. [9]

    When attribute-specific editing is desired, evaluation must also test whether characteristics outside the requested target remain stable

    CONCLUSION Prompt adherence alone does not fully characterize natural-language voice control. When attribute-specific editing is desired, evaluation must also test whether characteristics outside the requested target remain stable. Across three speech-generation systems, our paired audit shows that descriptor-aligned responses frequently co-occur with off...

  10. [10]

    PromptTTS: Controllable text-to-speech with text descriptions,

    Zhifang Guo, Yichong Leng, Yihan Wu, Sheng Zhao, and Xu Tan, “PromptTTS: Controllable text-to-speech with text descriptions,”Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2023

  11. [11]

    InstructTTS: Modelling expressive tts in discrete latent space with natural language style prompt,

    Dongchao Yang, Songxiang Liu, Rongjie Huang, Chao Weng, and Helen Meng, “InstructTTS: Modelling expressive tts in discrete latent space with natural language style prompt,” IEEE/ACM Trans. Audio, Speech, Lang. Process., 2024

  12. [12]

    PromptTTS 2: Describing and generating voices with text prompt,

    Yichong Leng, Zhifang Guo, Kai Shen, Xu Tan, Zeqian Ju, Yanqing Liu, Yufei Liu, Dongchao Yang, Leying Zhang, Kaitao Song, Lei He, Xiang-Yang Li, Sheng Zhao, Tao Qin, and Jiang Bian, “PromptTTS 2: Describing and generating voices with text prompt,” inProc. Int. Conf. Learning Representations (ICLR), 2024

  13. [13]

    PromptStyle: Controllable style transfer for text-to-speech with natural language descriptions,

    Guanghou Liu, Yongmao Zhang, Yi Lei, Yunlin Chen, Rui Wang, Zhifei Li, and Lei Xie, “PromptStyle: Controllable style transfer for text-to-speech with natural language descriptions,” arXiv preprint arXiv:2305.19522, 2023

  14. [14]

    PromptTTS++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions,

    Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, and Kentaro Tachibana, “PromptTTS++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions,”Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2024

  15. [15]

    ControlSpeech: Towards simultane- ous and independent zero-shot speaker cloning and zero-shot language style control,

    Shengpeng Ji, Qian Chen, Wen Wang, Jialong Zuo, Minghui Fang, Ziyue Jiang, Hai Huang, Zehan Wang, Xize Cheng, Siqi Zheng, and Zhou Zhao, “ControlSpeech: Towards simultane- ous and independent zero-shot speaker cloning and zero-shot language style control,” inProc. 63rd Annu. Meeting Assoc. Comput. Linguistics (ACL), 2025, pp. 6966–6981

  16. [16]

    CosyV oice 3: Towards in-the-wild speech generation via scaling-up and post-training,

    Zhihao Du, Changfeng Gao, Yuxuan Wang, Fan Yu, Tianyu Zhao, Hao Wang, Xiang Lv, Hui Wang, et al., “CosyV oice 3: Towards in-the-wild speech generation via scaling-up and post-training,”arXiv preprint arXiv:2505.17589, 2025

  17. [17]

    V oxCPM2 technical report,

    Yixuan Zhou, Guoyang Zeng, Xin Liu, Xiang Li, Renjie Yu, Jiancheng Gui, Jiaheng Wu, et al., “V oxCPM2 technical report,” arXiv preprint arXiv:2606.06928, 2026

  18. [18]

    Fish Audio S2 technical report,

    Shijia Liao, Yuxuan Wang, Songting Liu, Yifan Cheng, Ruoyi Zhang, Tianyu Li, Shidong Li, et al., “Fish Audio S2 technical report,”arXiv preprint arXiv:2603.08823, 2026

  19. [19]

    OV-InstructTTS: To- wards open-vocabulary instruct text-to-speech,

    Yong Ren, Jiangyan Yi, Jianhua Tao, Haiyang Sun, Zhengqi Wen, Hao Gu, Le Xu, and Ye Bai, “OV-InstructTTS: To- wards open-vocabulary instruct text-to-speech,”arXiv preprint arXiv:2601.01459, 2026

  20. [20]

    FlexiV oice: Enabling flexible style control in zero-shot tts with natural language instructions,

    Dekun Chen, Xueyao Zhang, Yuancheng Wang, Kenan Dai, Li Ma, and Zhizheng Wu, “FlexiV oice: Enabling flexible style control in zero-shot tts with natural language instructions,” arXiv preprint arXiv:2601.04656, 2026

  21. [21]

    InstructTTSEval: Benchmarking complex natural-language in- struction following in text-to-speech systems,

    Kexin Huang, Qian Tu, Liwei Fan, Chenchen Yang, Dong Zhang, Shimin Li, Zhaoye Fei, Qinyuan Cheng, and Xipeng Qiu, “InstructTTSEval: Benchmarking complex natural-language in- struction following in text-to-speech systems,”arXiv preprint arXiv:2506.16381, 2025

  22. [22]

    SPAM: Style prompt adherence metric for prompt-based tts,

    Chanhee Cho, Nayeon Kim, and Bugeun Kim, “SPAM: Style prompt adherence metric for prompt-based tts,”arXiv preprint arXiv:2601.05554, 2026

  23. [23]

    Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,

    R. J. Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron J. Weiss, Rob Clark, and Rif A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” inProc. Int. Conf. Machine Learning (ICML), 2018, pp. 4693–4702

  24. [24]

    Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,

    Yuxuan Wang, Daisy Stanton, Yu Zhang, R. J. Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” inProc. Int. Conf. Machine Learning (ICML), 2018, pp. 5180–5189

  25. [25]

    FastSpeech 2: Fast and high-quality end-to- end text to speech,

    Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, “FastSpeech 2: Fast and high-quality end-to- end text to speech,” inProc. Int. Conf. Learning Representations (ICLR), 2021

  26. [26]

    FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations

    Yoonhyung Lee, Hyunsin Park, Jinhwan Park, and Jinkyu Lee, “FC-TTS: Style and timbre control in zero-shot text- to-speech with disentangled speech representations,”arXiv preprint arXiv:2605.24618, 2026

  27. [27]

    DisCo-Speech: Controllable zero-shot speech generation with a disentangled speech codec,

    Tao Li, Wengshuo Ge, Zhichao Wang, Zihao Cui, Yong Ma, Yingying Gao, Chao Deng, Shilei Zhang, and Junlan Feng, “DisCo-Speech: Controllable zero-shot speech generation with a disentangled speech codec,”arXiv preprint arXiv:2512.13251, 2025

  28. [28]

    Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment

    Taehyung Yu and Seongjae Kang, “Best-of- n tts evalua- tion is confounded by asr family alignment,”arXiv preprint arXiv:2607.08256, 2026

  29. [29]

    Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

    Ali Asaria, Tony Salomone, and Deep Gandhi, “Reliable neural-codec text-to-speech by asr self-verification and distilla- tion: Near-zero catastrophic failures across models and codecs,” arXiv preprint arXiv:2606.18323, 2026

  30. [30]

    librosa: Audio and music signal analysis in python,

    Brian McFee, Colin Raffel, Dawen Liang, Daniel P. W. Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto, “librosa: Audio and music signal analysis in python,”Proc. Python in Science Conf., pp. 18–25, 2015

  31. [31]

    pYIN: A fundamental fre- quency estimator using probabilistic threshold distributions,

    Matthias Mauch and Simon Dixon, “pYIN: A fundamental fre- quency estimator using probabilistic threshold distributions,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2014, pp. 659–663

  32. [32]

    Robust speech recognition via large-scale weak supervision,

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Chris- tine McLeavey, and Ilya Sutskever, “Robust speech recognition via large-scale weak supervision,” inProc. Int. Conf. Machine Learning (ICML), 2023, pp. 28492–28518

  33. [33]

    ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,

    Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,” inProc. Interspeech, 2020, pp. 3830–3834

  34. [34]

    X-vectors: Robust dnn em- beddings for speaker recognition,

    David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, “X-vectors: Robust dnn em- beddings for speaker recognition,” inProc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 5329–5333. Supplementary Material Beyond Prompt Adherence: Auditing Attribute-Level V oice Control in Speech Generation

  35. [35]

    Systems and model identifiers Table 3 lists the three systems, model identifiers, and control inter- faces used for generation

    REPRODUCIBILITY AND EV ALUATION SCOPE 10.1. Systems and model identifiers Table 3 lists the three systems, model identifiers, and control inter- faces used for generation. Environment specifications, inference settings, evaluation configurations, and analysis scripts are provided in the accompanying code repository. Table 3. Systems, model identifiers, an...

  36. [36]

    Acoustic and prosodic measurements The analysis uses F0, speaking rate, duration, RMS energy, spectral centroid, 85% spectral roll-off, spectral flatness, and zero-crossing rate

    MEASUREMENTS AND STATISTICAL DEFINITIONS 11.1. Acoustic and prosodic measurements The analysis uses F0, speaking rate, duration, RMS energy, spectral centroid, 85% spectral roll-off, spectral flatness, and zero-crossing rate. F0 is converted to semitone change for the primary paired audit: ∆fst = 12 log2 f(y d) f(y 0) .(9) Table 4 gives the fixed scales u...

  37. [37]

    Responsive

    DETAILED AUDIT RESULTS 12.1. Conditional target-response analysis Table 5 gives the complete conditional summary. In eight sufficiently populated settings, 54.5%–95.8% of target-responsive outputs exhibit at least one substantial off-target deviation. The Fish-Speech-S2 deep setting contains only two target-responsive outputs under the compos- ite thresho...

  38. [38]

    Policies and operating point The candidate budget counts descriptor-conditioned generations; the matched neutral baseline y0 is shared across all policies and is not included in B

    VODER-CAL DETAILS AND ROBUSTNESS 13.1. Policies and operating point The candidate budget counts descriptor-conditioned generations; the matched neutral baseline y0 is shared across all policies and is not included in B. Direct uses the descriptor-conditioned candidate generated with the fixed first seed. Random draws one seed and returns the corresponding...

  39. [39]

    no difference

    LISTENING-STUDY DETAILS 14.1. Design Ten adult listeners completed two blinded studies. H1 used 45 matched-baseline–descriptor pairs balanced across three systems, three descriptors, and five requests per system–descriptor cell. H2 used 45 matched-baseline–candidate triplets with Target-only and V oDER-Cal identities hidden and positions balanced. Each sa...

  40. [40]

    V oDER-Cal shows that preservation-aware ranking can select candidates with lower measured off-target devia- tion from a fixed pool

    EVIDENCE SCOPE The empirical conclusion is behavioral and limited to the tested sys- tems and signal-level target definitions: descriptor-aligned outputs can also change acoustic and prosodic measurements outside the corresponding target set. V oDER-Cal shows that preservation-aware ranking can select candidates with lower measured off-target devia- tion ...

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.