{"id":"949e7cf4-92ad-4e6b-af73-3a4b558a363a","arxiv_id":"2509.15626","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new public voice impression corpus plus a reference-free TTS conditioning method reduces numerical voice impression control error and suppresses reference-audio leakage.","lead":"This paper releases LibriTTS-VI, a public dataset of voice impression labels for text-to-speech, and proposes two ways to stop a reference voice from leaking its own impressions into the synthesized output. The authors report lower objective and subjective control error, with the best method dropping mean squared impression error from 0.61 to 0.41 and from 1.15 to 0.92.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Objective MSE gains may be an artifact of training and evaluating with the same Voice Impression Estimator, since the VIE is trained on 100 low-agreement annotations and its labels are propagated without validation; independent perceptual replication is needed.","rationale":"The reader's weakest assumption—that the VIE is a valid, reliable measure of human-perceived voice impression—is exactly the load-bearing concern. The paper's objective evaluation is entangled with the training signal through the shared VIE, and the VIE's limited training data and low inter-annotator agreement make this more than a theoretical risk. The proposed concrete test directly addresses this by measuring perceptual improvement with independent annotators, which would either validate the claim or reveal that the gains are an artifact of the evaluation metric. Since the reader already flagged this and assigned CONDITIONAL, my stress-test does not change the verdict; it strengthens the rationale by pinpointing the circularity in Sec. 2.2/4.2 and the missing validation of label propagation in Sec. 3. The paper has genuine strengths: a public corpus, clear method descriptions, reproducible code, and an honest reporting of low alpha values. These do not resolve the circularity, but they justify a conditional acceptance pending independent perceptual evaluation.","tokens_in":10206,"tokens_out":2504,"duration_ms":27422,"concrete_test":"Recruit a new pool of annotators who were not involved in creating LibriTTS-VI, and have them rate the exact synthesized stimuli from the RVI-MSE evaluation (Table 3) and the multiple-VI modulation evaluation (Table 5) on the 11 VI scales. Compute the MSE between their mean ratings and the specified target VI vectors for VIC-base and VIC-rfg. If VIC-rfg's improvement over VIC-base (e.g., RVI-MSE 0.61→0.41) does not replicate with these held-out human ratings, the central claim that VIC-rfg improves perceptual controllability is unsupported. This is a single, decisive check that bypasses the VIE circularity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims rest on the Voice Impression Estimator (VIE), which is used twice: it generates the target VI vectors during control-module fine-tuning (Sec. 2.2, Eq. 5) and it computes the objective metrics VI-MSE, RVI-MSE, and ΔV (Sec. 4.2, Eqs. 6–8). The reported headline improvement (RVI-MSE from 0.61 to 0.41) is therefore not an independent measure of perceptual controllability—it measures how well each system can make the VIE's output match the target, a metric the training directly optimizes. The VIE's own reliability is doubtful: it is trained on only 100 manually annotated utterances with a low average inter-annotator agreement (Krippendorff's α = 0.464, Table 1), and the labels are propagated to up to 100 acoustically similar utterances per seed without any validation that acoustic similarity preserves the perceived VI. The subjective evaluation uses the same four annotators who created the training labels and reports only one speaker (ID 8555), so it does not break the circularity. Thus the 0.61→0.41 objective gain and even the 1.15→0.92 subjective gain may reflect fitting to the VIE's idiosyncrasies and to the original annotators' biases rather than genuinely improved voice impression control.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LibriTTS-VI, a public corpus of 11-dimensional voice impression (VI) annotations built on LibriTTS-R, and proposes two methods to reduce impression leakage in zero-shot TTS: VIC-sep, which uses separate utterances for speaker and VI conditioning, and VIC-rfg, which removes the reference audio entirely and generates a speaker embedding from the target VI. The authors report objective and subjective improvements in controllability (RVI-MSE from 0.61 to 0.41; multiple-VI subjective MSE from 1.15 to 0.92) while maintaining synthesis quality.","tokens_in":10563,"tokens_out":5641,"duration_ms":64877,"significance":"If the evaluation is valid, the paper makes a useful contribution: LibriTTS-VI would be the first public VI corpus with explicit annotation standards, and the reference-free VIC-rfg architecture is a practical step toward privacy-preserving and leakage-free VI control. The authors also provide reproducible experimental details and release the corpus. However, the central quantitative claims currently rest on a partially self-referential evaluation loop and a small, non-independent subjective test, so the evidence for the headline improvements is not yet convincing.","major_comments":[{"comment":"The objective controllability metrics VI-MSE, RVI-MSE, and ΔV are computed with the same VIE that supplies target VI vectors during fine-tuning. The VIE is trained on 100 manually annotated utterances with average Krippendorff's α = 0.464 (Table 1) and labels propagated by acoustic similarity without validation. The reported 0.61→0.41 improvement therefore shows alignment with this particular estimator, not necessarily with human perception. An independent perceptual test, or at least a VIE validated on held-out human ratings, is needed to support the headline claim.","section":"Sec. 4.2, Eqs. (6)-(8); Sec. 2.2, Eq. (5); Table 3"},{"comment":"The subjective controllability evaluation uses the same four annotators who created the training labels and reports only speaker 8555. This does not break the circularity: the annotators may be rating consistently with their own earlier labels, and a single speaker is not representative. The MOS test with 30 independent raters is a strength, but the controllability claim needs more speakers and independent (non-author) raters.","section":"Sec. 4.3, Table 5"},{"comment":"For VIC-rfg, the reference audio r is not used at all by construction, so ΔV = RVI-MSE - VI-MSE does not measure leakage from a reference; it only compares the model's error for two different target vectors. The near-zero ΔV (0.05) is therefore largely an architectural consequence, not an empirical demonstration of leakage reduction. The paper should state this explicitly and not present ΔV for VIC-rfg as comparable to the baseline's leakage measure.","section":"Sec. 4.2, Eq. (8); Sec. 2.4; Table 3"},{"comment":"The abstract claims that a comparison with a prompt-based TTS reveals imprecise numerical control and VI-text entanglement that the proposed methods overcome. No such experiment or result appears in the manuscript. Either add the missing comparison or remove this claim from the abstract.","section":"Abstract, last sentence"}],"minor_comments":[{"comment":"K) Slow–Fast is listed with α = '-' because it is derived objectively, but the reported average α = 0.464 appears to be over the other 10 dimensions. Please clarify in the table caption or footnote.","section":"Table 1"},{"comment":"The target VI is modulated from -3 to +3, but the annotation scale is 1–7. Specify how the anchor vector and modulation range are mapped to the 1–7 scale.","section":"Sec. 4.2, modulation experiment"},{"comment":"The text says 'the same audio generated' from the modulation experiment (50 sentences per condition) is used, but Table 5 reports 10 sentences per condition. Clarify the selection or the number of sentences.","section":"Sec. 4.3"},{"comment":"VIC-rfg uses a sampled Gaussian noise vector, so synthesis is stochastic; Eq. (8) implicitly assumes determinism. State whether the metrics are averaged over noise samples.","section":"Eqs. (6)-(8)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a speech/audio venue and the corpus release is valuable. The main concern is the self-referential evaluation: the objective metrics use the same estimator that provides training targets, and the subjective test uses the same annotators and a single speaker. I believe this can be fixed with additional experiments or a careful re-framing of the claimed evidence, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid engineering paper with a genuinely useful public corpus. The two methods are simple variations on existing ideas, but the corpus alone justifies a look. That said, the headline numeric gains are computed with the same Voice Impression Estimator used to generate training targets, so they measure how well each system fits that estimator, not human perception. Don't quote the 0.61→0.41 without that caveat.\n\nWhat's new: LibriTTS-VI is the first public VI corpus, built on LibriTTS-R with 11 perceptual scales, manual annotations, and propagated labels. That's a real resource. The VIC-sep idea—using two different utterances from the same speaker to separate identity from impression—is a clean and easy-to-replicate trick. VIC-rfg borrows reference-free embedding generation from the speaker-privacy literature and applies it to VI control; not deeply novel, but the combination with VIC is new. The RVI-MSE / ΔV protocol for quantifying leakage is a useful evaluation idea.\n\nThe soft spots are exactly where the reader says. The VIE is trained on 100 utterances with Krippendorff's alpha 0.464, labels are propagated to similar utterances without checking that the impression actually transfers, and the same VIE supplies targets and computes objective metrics. So the objective results are circular to an important degree. The subjective evaluation uses the same four annotators who made the corpus, only one speaker is reported, and the prompt-based comparison promised in the abstract is nowhere in the body. Those are fixable but they matter. The MOS test with 30 independent raters is a plus, and the objective quality metrics (CER/WER/UTMOS) are fine.\n\nIs the central argument sound? I think the leakage problem is real and the reference-free approach does reduce it by construction. The direction of the effect is believable. But the magnitude—and whether it transfers to actual perceptual control—needs an independent VIE or at least more speakers and fresh annotators.\n\nBottom line: worth engaging with for the corpus and the evaluation protocol, but treat the numbers as provisional. I'd send it to peer review, with a request for an independent perceptual check and the missing prompt comparison.\n\nRecommendation: send to peer review, conditional on revision.","headline":"Useful corpus and a sensible leakage-reduction trick, but the central controllability numbers are mostly self-referential; the subjective check is too thin to fully break the loop.","tokens_in":11008,"tokens_out":2846,"would_cite":true,"duration_ms":27606,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that text-to-speech can take an 11-number voice-impression vector as the only control signal, without a reference recording, and follow it more faithfully than prior reference-based methods.","keywords":["voice impression control","text-to-speech","zero-shot TTS","reference-free synthesis","disentanglement","impression leakage","corpus","VITS"],"falsifier":"Have new, independent annotators rate outputs of VIC-rfg versus VIC-base on a fresh set of utterances and targets, with annotators blind to which system generated each sample and with the VIE entirely out of the loop. If the reference-free model's perceived accuracy is no better than baseline (or no better than chance), the claim that reference-free conditioning improves controllability would be refuted.","tokens_in":10133,"feed_emoji":"🎙️","tokens_out":6163,"duration_ms":63389,"temperature":0.7,"pith_summary":"Voice impression control lets users adjust qualities like brightness or calmness numerically in synthesized speech, but prior systems needed a reference audio clip whose own impression leaked into the output. This paper introduces the first public corpus for this task, LibriTTS-VI, and proposes two ways to reduce leakage: training with separate utterances for speaker identity versus impression, and a reference-free model that generates the speaker embedding from the target impression vector alone. The paper reports that the reference-free method lowers the mean-squared error of 11-dimensional impression vectors from 0.61 to 0.41 objectively and from 1.15 to 0.92 subjectively, while keeping audio quality about the same. If right, fine-grained and privacy-friendly voice control becomes feasible without any reference voice.","feed_headline":"Reference-free TTS follows numeric voice impressions more faithfully","feed_subtitle":"A reference-free model cuts 11-dimensional voice-impression error from 0.61 to 0.41 and nearly eliminates reference-bias leakage.","key_machinery":"The control module inserted into the VITS-based TTS backbone. VIC-base fuses the target VI vector with the reference audio's encoded vector, using dropout and a gradient reversal layer to suppress impression information. VIC-sep feeds the module a second utterance from the same speaker as the speaker-identity source while the original utterance supplies target VI and synthesis text. VIC-rfg removes the reference: the module gets only the VI vector plus a random Gaussian vector, and the stochastic duration predictor is fine-tuned jointly to keep speaking rate stable.","core_discovery":"Central claim: impression leakage—the unwanted bias of the synthesized voice toward the reference audio's impression—can be largely removed by architecture. VIC-rfg, the reference-free system, replaces the reference audio's encoded vector with Gaussian noise inside the speaker-encoder control module, leaving the 11-dimensional target VI vector as the only speaker-related conditioning. In zero-shot evaluation on 39 unseen speakers, this reduces the leakage gap ΔV from 0.22 to 0.05, improves the average control slope from 0.096 to 0.177, and lowers subjective multi-VI MSE from 1.15 to 0.92. The paper also releases LibriTTS-VI, a public corpus of manual and estimated 11-dimensional impression l","pith_inferences":["Inference: Since the same VIE supplies training labels and objective metrics and was built from only 100 annotated utterances with average inter-annotator agreement α=0.464, the reported MSE gains may partially reflect fitting to the estimator; a blind listening study with fresh raters would test whether perceived controllability improves as much as the numbers say.","Inference: The reference-free architecture opens an interface where users adjust a voice's impression via sliders or a text-to-vector mapping without ever providing a reference recording; the paper does not evaluate such an interface.","Inference: The RVI-MSE and ΔV protocol could serve as a general leakage test for any conditioning attribute in controllable TTS, not just voice impression; that broader use is not explored here."],"forward_implications":["VIC-rfg synthesizes speech from a target VI vector alone, so a user can request e.g. a brighter or calmer voice numerically without a reference recording.","LibriTTS-VI, with manual annotations and estimated VIs over LibriTTS-R, becomes a public benchmark for comparing VI control systems.","RVI-MSE and ΔV quantify leakage; the paper reports ΔV dropping from 0.22 (baseline) to 0.14 (VIC-sep) and 0.05 (VIC-rfg).","Control remains partial: the average modulation slope rises to 0.177 for VIC-rfg but is near zero for Powerful-Weak, so some dimensions still respond weakly.","Reference-free control trades speaker similarity: SECS falls from 0.82–0.84 to 0.76, a quality to weigh in applications needing voice preservation."],"fun_headline_variants":["Reference-free TTS cuts voice-impression error by 33%","New corpus plus reference-free method slashes impression leakage","Without reference audio, TTS follows numeric voice scores better","LibriTTS-VI enables precise voice control without reference bias"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central results rest on the Voice Impression Estimator (VIE) being a reliable proxy for human-perceived voice impression, because the same model creates the target labels during training and computes the objective error used to report success.","fun_headline_variants_meta":{"raw":{"variants":["Reference-free TTS cuts voice-impression error by 33%","New corpus plus reference-free method slashes impression leakage","Without reference audio, TTS follows numeric voice scores better","LibriTTS-VI enables precise voice control without reference bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000579,"raw_usage":{"total_tokens":2561,"prompt_tokens":737,"completion_tokens":1824,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":1753}},"tokens_in":481,"tokens_out":1824,"duration_ms":17801,"temperature":1.0,"reasoning_tokens":1753,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T16:11:19.960015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have new, independent annotators rate outputs of VIC-rfg versus VIC-base on a fresh set of utterances and targets, with annotators blind to which system generated each sample and with the VIE entirely out of the loop. If the reference-free model's perceived accuracy is no better than baseline (or no better than chance), the claim that reference-free conditioning improves controllability would be refuted.","supporting_citations":[],"review_version":1}