{"id":"c22744bc-0799-491a-9b4a-0dcd4a8e14da","arxiv_id":"2507.01356","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A voice conversion pipeline conditions FastSpeech 2 on predicted likability ratings, achieving partial subjective control but with degraded identity and content at strong settings.","lead":"This paper builds a system that makes speech sound more or less likable by converting the voice while keeping the speaker and words intact. It trains a likability scorer on human ratings, then uses that scorer to automatically label thousands of synthetic speech samples for training.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The likability-control result depends on a predictor that both creates the training labels and scores the outputs; its transfer to JVS/JTES is unvalidated, so the objective evidence may be circular.","rationale":"The reader's weakest assumption is exactly the load-bearing concern: the automatic likability predictor trained on CocoNut-Humoresque must produce reliable pseudo-labels for JVS/JTES, and the same predictor must be a valid objective measure for converted speech. I agree with this identification. The paper's strongest claim is that the method effectively controls voice likability while preserving speaker identity and linguistic content, but the objective demonstration of likability control is circular because the same model that generates the training labels also scores the outputs. The paper itself flags limitations that corroborate the concern: CER rises for female speakers at high targets (Figure 4), speaker cosine similarity drops at extremes (Figure 5), and the subjective preference test shows no 0-vs-1 effect and a reversal for speaker m49 (Section 5.3). These internal admissions weaken the abstract's unqualified claim. The proposed concrete test would settle the issue by providing independent human ratings on both the original corpora and the converted utterances across the full target range. If that test passes, the circularity concern is resolved; if it fails, the central claim needs substantial qualification. The verdict should remain conditional, with the condition being independent human validation and release of the predictor/VC code, so no change to the reader's verdict is needed.","tokens_in":8435,"tokens_out":4440,"duration_ms":55992,"concrete_test":"Run one human-rating study on a held-out subset of JVS/JTES: (a) original utterances, to test predictor transfer, and (b) the same utterances converted at target likability values -2, -1, 0, 1, 2, with each utterance rated by at least 20 listeners per demographic group. Compute LCC/SRCC between human mean ratings and (a) predictor outputs on the original utterances, and (b) target values on the converted utterances. If (a) is not significantly positive (e.g., below ~0.3) or (b) does not show a monotonic increase in human-rated likability with target, then the Section 5.2 objective evidence is an artifact of the circular predictor, and the control claim must be re-evaluated with human labels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 trains the voice conversion model on pseudo-labels produced by the Section 2 predictor, and Section 5.2 evaluates success with the same predictor. This loop is only valid if the predictor's CocoNut-trained mapping (Table 2: LCC=0.46) transfers to studio-recorded JVS/JTES speech, but the paper provides no held-out human ratings on JVS/JTES to test that transfer. The post-filter in Eq. (1) only aligns marginal mean and variance; it cannot correct per-utterance ranking errors. If the predictor systematically keys on corpus artifacts (e.g., recording conditions, Fo range, source-separation noise in CocoNut) rather than on perceived likability, the VC model will learn to manipulate those artifacts and the same predictor will report success. Figure 3's narrow predicted range (-0.51 to -0.23) for targets -2 to 2 may reflect such bias rather than a real perceptual ceiling. The only independent evidence is the four-speaker preference test in Section 5.3, which shows partial, speaker-dependent effects (no significant 0 vs. 1 difference; speaker m49 reverses), so it does not by itself establish the full claim. The speaker-identity and content-preservation parts use independent metrics, but those also degrade at extreme targets (Figures 4 and 5). Thus the central claim of effective, identity- and content-preserving likability control is not yet supported by non-circular objective evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a pipeline for controlling perceived voice likability in voice conversion. A neural predictor (TDNN with statistics pooling) is trained on the CocoNut-Humoresque corpus to output likability ratings for four listener groups, with a linear post-filter that rescales predictions to match human rating variance. This predictor is then used to pseudo-label the JVS and JTES speech corpora, and a FastSpeech 2-based voice conversion model conditioned on HuBERT discrete units, speaker embeddings, and target likability ratings is trained on these pseudo-labels. The authors report objective evaluations of likability control (using the same predictor), content preservation (CER), and speaker identity (embedding cosine similarity), plus a subjective pairwise preference test on four speakers. The paper claims effective likability control while preserving speaker identity and linguistic content.","tokens_in":8713,"tokens_out":3609,"duration_ms":42798,"significance":"If the central claim were fully supported, the work would be a valuable contribution to controllable speech synthesis, since it offers a scalable route from subjective ratings to a voice conversion model without per-sample manual annotation. The predictor itself shows a significant correlation with human ratings (LCC = 0.46 for the combined listener group) and a reasonably good liked/disliked classification accuracy of 74%. The subjective listening test uses a nontrivial number of participants and provides some evidence that the model can shift perceived likability for three of four speakers in one comparison direction. However, the principal evidence for likability control is circular, and the independent subjective evidence is only partial. The paper also shows clear degradation in content preservation at the extreme target value and speaker-dependent failures. The strengths are the integration of an existing likability corpus, the multi-group listener modeling, and the explicit trade-off parameter for identity preservation; the weaknesses are the unvalidated transfer of the predictor to new corpora and the overstatement of the experimental results in the abstract and conclusion.","major_comments":[{"comment":"The objective likability evaluation is circular. The predictor used to score the converted speech is the same model that generated the pseudo-labels on which the conversion model was trained in Section 3. Consequently, Figure 3 primarily verifies that the conversion model has learned to match this particular predictor, not that human-perceived likability is controlled. The post-filter in Eq. (1) only aligns the marginal mean and variance of predictions; it cannot correct per-utterance ranking errors, and no human ratings on the JVS/JTES corpora are provided to validate that the CocoNut-Humoresque-trained predictor transfers to studio-recorded speech. To support the central claim, the authors should either provide held-out human likability ratings for converted JVS/JTES speech or validate the predictor on those corpora independently.","section":"Section 5.2, Figure 3"},{"comment":"The subjective evaluation provides only partial support for the control claim. The paper reports no significant difference between target values 0 and 1, and speaker m49 even shows a significant decrease in perceived likability from target -1 to 1, contradicting the intended control. Only comparisons involving the -1 target, and only for three of four speakers, support the effect. The abstract's statement that the method 'effectively controls voice likability' is too strong relative to this evidence. The claims should be tempered, or additional subjective tests covering a wider target range and more speakers should be provided.","section":"Section 5.3, Figure 6"},{"comment":"The claim that linguistic content is preserved is inconsistent with the reported results. Figure 4 shows a marked increase in CER at target likability 2, especially for female speakers, and the text in Section 5.2 explicitly acknowledges that the conversion process resulted in a higher CER at that target. The abstract and conclusion nonetheless claim that linguistic content is preserved without qualification. The content-preservation claim should be restricted to the range for which the data provide support (approximately -2 to 1), and the conclusion should reflect the acknowledged limitation.","section":"Section 5.2, Figure 4 and Conclusion"},{"comment":"The speaker-identity preservation claim is also overgeneralized. The text states that speaker identity was preserved 'within the target ratings of -1 to 1,' but Figure 5 shows degradation at the extreme target values, with the mean cosine similarity approaching or falling below the EER threshold at target 2 for some speakers. The abstract's unqualified claim that speaker identity is preserved therefore goes beyond the presented evidence. Please qualify the identity-preservation guarantee and coordinate the abstract, Section 5.2, and the conclusion.","section":"Section 5.2, Figure 5 and Abstract"}],"minor_comments":[{"comment":"The figure contains a typo: 'Mel-spetrogram Decoder' should be 'Mel-spectrogram Decoder.'","section":"Figure 2"},{"comment":"The preference test does not state which statistical test was used to determine significance. Please specify the test and whether multiple-comparison corrections were applied.","section":"Section 5.3"},{"comment":"The 'All' row aggregates the four listener groups, but the aggregation method is not described. Please clarify whether the rating is averaged over groups or over listeners.","section":"Section 4.2, Table 2"},{"comment":"The paper reports the classification accuracy of the predictor on a liked/disliked dichotomy but does not define the threshold used to map the continuous rating to the binary class. Please provide this definition.","section":"Section 4.2"},{"comment":"The manuscript does not mention code or data availability. Since the experiments rely on several open-source components and custom training pipelines, a reproducibility statement would be valuable.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central methodological concern is not that the idea is wrong but that the main objective evidence is self-referential. I would encourage the editor to require either a human-rated validation of the predictor on JVS/JTES or a reformulated paper that explicitly frames the objective likability result as a model-predictor consistency check rather than evidence of perceptual control. The subjective results are promising but not strong enough to carry the full abstract claim, and the content-preservation failure at target 2 should be visible in the title-level claims. This seems fixable within the scope of a revision, hence major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is genuinely new: using an existing likability corpus to train a predictor, auto-annotating a large TTS corpus, and then conditioning a FastSpeech 2 + HuBERT voice converter on a continuous likability score. That pipeline is a sensible answer to the scalability problem in subjective speech control, and the paper is clearly written with enough detail to reproduce. I also credit the authors for reporting the narrow predicted range in Figure 3 and the CER/speaker degradation at extreme targets; they don't hide the blemishes entirely.\n\nThe main soft spot is the one the stress-test flags: the objective evaluation in Section 5.2 uses the exact predictor that generated the training labels. That loop would be fine for a sanity check, but it is presented as evidence of effective control. The only independent evidence is the subjective preference test, which is partial—significant for -1 vs 0 and -1 vs 1 overall, but not for 0 vs 1, and speaker m49 reverses. So the central claim of reliable, identity- and content-preserving likability control is not fully supported yet. The predictor's moderate correlation (LCC=0.46) also means the pseudo-labels are noisy, and no held-out human ratings on JVS/JTES confirm that the predictor transfers to studio-recorded speech.\n\nThat said, the circularity does not sink the paper. The subjective test gives some real signal, and the idea itself is worth building on. What the paper needs is (a) human ratings on a subset of the actual synthesized JVS/JTES speech, (b) a tempered abstract that acknowledges the partial control and the degradation at extreme targets, and (c) ideally release of code/data. The authors already admit some of these limitations in the body, which makes me think they would respond well to revision.\n\nThis is a solid workshop to good conference level paper, not a paradigm shift. Read it if you work on paralinguistic control or auto-labeling pipelines; it is a good discussion piece for a reading group on evaluation validity. I would not cite it in my own work until the evidence is stronger, but I would send it to reviewers rather than desk-reject: the novelty is real, the flaws are identifiable and fixable, and the subjective data point to something real.","headline":"A novel auto-annotated likability-conditioned VC system with a circular objective evaluation and an abstract that overreaches; worth a serious referee but needs major revision.","tokens_in":9253,"tokens_out":1782,"would_cite":false,"duration_ms":23143,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A voice conversion system that raises or lowers perceived likability of any speaker's voice while preserving who is speaking and what is said.","keywords":["voice conversion","voice likability","paralinguistic voice control","speech synthesis","likability prediction","automatic corpus annotation","FastSpeech 2","HuBERT discrete units"],"falsifier":"Run a fresh listener panel on the existing converted audio: collect human likability ratings for utterances spanning the full target range (−2 to 2) across several speakers, and compare those ratings against both the target values and the predictor's outputs. If human ratings do not move monotonically with the target — or diverge from the predictor's scores while the predictor still tracks the targets — the control effect is an artifact of the auto-annotator echoing its own labels, and the central claim fails. A cheaper partial check is to verify whether the predictor's LCC of about 0.46 holds on a human-rated holdout drawn from JVS/JTES alone.","tokens_in":8227,"feed_emoji":"🎙️","tokens_out":7632,"duration_ms":75921,"temperature":0.7,"pith_summary":"This paper sets out to establish that voice likability — the subjective likeability of a speaker's voice — can be treated as a controllable dimension of speech synthesis, alongside who is speaking and what is said. To make this practical, it trains an automatic likability predictor on an existing corpus of human ratings, then uses that predictor to label thousands of utterances in large Japanese speech corpora, removing the need for expensive manual ratings. On top of those pseudo-labels, a voice conversion model based on discrete speech units, speaker embeddings, and a target likability score synthesizes the same words in a more or less likable voice. Subjective listening tests with 100 participants confirm that listeners perceive the intended likability differences for three of the four test speakers, and objective checks show linguistic content and speaker identity are mostly preserved. If correct, the result is a scalable recipe for controlling a subjective voice quality without collecting new human ratings for every corpus.","feed_headline":"Dial a voice's likability up or down, speaker stays the same","feed_subtitle":"A predictor auto-rates big Japanese speech corpora, then a converter shifts perceived likability per target.","key_machinery":"The load-bearing object is the automatic likability predictor, a single network that maps a log-Mel spectrogram through three time-delay neural network layers and a statistics-pooling layer (the x-vector-style backbone) to four time-invariant ratings, one per listener group defined by gender and age. A post-filtering step — a linear rescaling of the raw predictions to match the mean and variance of human ratings on the validation set — corrects for the predictor's tendency to regress toward the center of the rating scale. This predictor does double duty: it produces the pseudo-labels used to train the voice conversion model on the JVS and JTES corpora, and it is also the objective metric used to verify likability control on converted speech. The voice conversion model itself is a FastSpeech 2 TTS backbone, fed with compressed HuBERT cluster indices (k-means with k = 1000), an ECAPA-TDNN speaker embedding, and the target likability scalar; a trade-off multiplier s (1 at training, 2.5 at inference) scales the likability conditioning to balance control strength against speaker preservation.","core_discovery":"The paper's central claim is that perceived voice likability is a controllable acoustic attribute that can be manipulated by conditioning a TTS-based voice converter on a single scalar likability target, while keeping the speaker and the words fixed. The discovery chain runs through two components that feed each other: a TDNN-based predictor, trained on the CocoNut-Humoresque corpus to output mean likability for four listener groups (gender × age), generalizes well enough (LCC = 0.46, p < 3×$10^{-17}$, 74% liked/disliked classification accuracy) to serve as an automatic annotator; and a FastSpeech 2 model conditioned on HuBERT discrete units, an ECAPA-TDNN speaker embedding, and the predicted likability rating learns to slide the output voice along the likability axis. The authors demonstrate the control objectively — predicted likability of converted speech tracks the target for all listener groups — and subjectively, with listeners significantly preferring the intended voice in paired comparisons for three of the four test speakers. They also report the expected trade-off: pushing likability far from the speaker's natural range degrades speaker identity and intelligibility, which they expose through a scalar multiplier that the user can tune at inference.","pith_inferences":["An immediate next test the authors do not run: check whether the likability predictor transfers across languages or recording conditions, since CocoNut-Humoresque, JVS, and JTES are all Japanese speech, leaving the transfer assumption untested outside this language and domain.","A practical application implied by the trade-off analysis: the scalar multiplier s could be set automatically by a speaker-verification-style threshold, so the system pushes likability as far as possible while keeping the converted voice inside a guaranteed identity-acceptance region.","Because the predictor outputs separate ratings for four listener groups, the same trained model could generate reference voices targeted at a specific demographic — e.g., maximally likable to women over 40 — a fine-grained control the paper demonstrates but does not pursue as an application."],"forward_implications":["A user can take any utterance from a covered speaker and re-synthesize it at a higher or lower likability level, with speaker identity and linguistic content largely intact within the −1 to 1 target range.","Because the predictor requires no human raters at conversion time, the same pipeline can in principle be moved to any moderately sized multi-speaker corpus: predict, annotate, train.","The trade-off knob s gives an explicit, tunable dial between how much likability changes and how much of the original speaker's identity survives, making the method usable in applications where one of the two matters more.","The predictor itself can serve as a cheap, continuous substitute for some listening-test functions in future synthesis work, since its outputs correlate significantly with human ratings and it reports per-listener-group predictions.","The observed failure mode — weak control between targets 0 and 1, and speaker m49's converted speech becoming less likable at higher targets — implies that control quality depends on the base converter's naturalness and identity preservation, not just on the likability signal."],"supporting_citations":[{"why":"Supplies the CocoNut-Humoresque corpus of human likability ratings with listener group labels that the predictor is trained and evaluated on.","marker":"[11]"},{"why":"Provides the x-vector TDNN-plus-statistics-pooling architecture that the likability predictor is built on.","marker":"[12]"},{"why":"One of the two baseline voice conversion methods using self-supervised discrete speech units that the conversion model extends.","marker":"[17]"},{"why":"The companion study on discrete versus soft speech units for voice conversion that motivates the HuBERT k-means tokenization.","marker":"[18]"},{"why":"Provides the ECAPA-TDNN model used to extract the speaker embeddings that condition synthesis and verify identity preservation.","marker":"[19]"},{"why":"Supplies HuBERT, the self-supervised model whose hidden units are clustered into the discrete tokens that carry linguistic content.","marker":"[20]"},{"why":"Provides FastSpeech 2, the TTS backbone whose variance adapter and decoder are conditioned on the likability ratings.","marker":"[21]"},{"why":"The JVS corpus is one of the two large Japanese multi-speaker corpora that are auto-annotated and used to train the conversion model.","marker":"[24]"},{"why":"The JTES corpus supplies the remaining training utterances and the four held-out evaluation speakers used in the listening tests.","marker":"[25]"}],"fun_headline_variants":["AI predictor auto-scores speech data to let you adjust voice likability","Voice conversion shifts likability, keeps speaker and words intact","Auto-rated corpora enable on-the-fly voice likability control","New voice converter manipulates perceived likability per listener","Control voice likability with automated corpus rating and TTS"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the predictor's modest agreement with human raters (LCC = 0.46) is good enough that the pseudo-labels it stamps on the JVS and JTES corpora teach the conversion model the right thing, and that the same predictor is then a fair objective gauge of the converted speech.","fun_headline_variants_meta":{"raw":{"variants":["AI predictor auto-scores speech data to let you adjust voice likability","Voice conversion shifts likability, keeps speaker and words intact","Auto-rated corpora enable on-the-fly voice likability control","New voice converter manipulates perceived likability per listener","Control voice likability with automated corpus rating and TTS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3255,"prompt_tokens":923,"completion_tokens":2332,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":2247}},"tokens_in":539,"tokens_out":2332,"duration_ms":19017,"temperature":1.0,"reasoning_tokens":2247,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:52:58.348339+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a fresh listener panel on the existing converted audio: collect human likability ratings for utterances spanning the full target range (−2 to 2) across several speakers, and compare those ratings against both the target values and the predictor's outputs. If human ratings do not move monotonically with the target — or diverge from the predictor's scores while the predictor still tracks the targets — the control effect is an artifact of the auto-annotator echoing its own labels, and the central claim fails. A cheaper partial check is to verify whether the predictor's LCC of about 0.46 holds on a human-rated holdout drawn from JVS/JTES alone.","supporting_citations":[{"cited_title":"V oice transformation: A survey,","cited_arxiv_id":null,"evidence_quote":"Supplies the CocoNut-Humoresque corpus of human likability ratings with listener group labels that the predictor is trained and evaluated on."},{"cited_title":"PromptTTS: Controllable text-to-speech with text descriptions,","cited_arxiv_id":null,"evidence_quote":"Provides the x-vector TDNN-plus-statistics-pooling architecture that the likability predictor is built on."},{"cited_title":"“Would you buy a car from me?","cited_arxiv_id":null,"evidence_quote":"One of the two baseline voice conversion methods using self-supervised discrete speech units that the conversion model extends."},{"cited_title":"Who finds this voice attractive? A large-scale experiment using in-the-wild data,","cited_arxiv_id":null,"evidence_quote":"The companion study on discrete versus soft speech units for voice conversion that motivates the HuBERT k-means tokenization."},{"cited_title":"X-vectors: Robust DNN embeddings for speaker recogni- tion,","cited_arxiv_id":null,"evidence_quote":"Provides the ECAPA-TDNN model used to extract the speaker embeddings that condition synthesis and verify identity preservation."},{"cited_title":"AutoMOS: Learning a non-intrusive as- sessor of naturalness-of-speech,","cited_arxiv_id":null,"evidence_quote":"Supplies HuBERT, the self-supervised model whose hidden units are clustered into the discrete tokens that carry linguistic content."},{"cited_title":"UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,","cited_arxiv_id":null,"evidence_quote":"Provides FastSpeech 2, the TTS backbone whose variance adapter and decoder are conditioned on the likability ratings."},{"cited_title":"Any-to-one sequence- to-sequence voice conversion using self-supervised discrete speech representations,","cited_arxiv_id":null,"evidence_quote":"The JVS corpus is one of the two large Japanese multi-speaker corpora that are auto-annotated and used to train the conversion model."},{"cited_title":"A comparison of discrete and soft speech units for improved voice conversion,","cited_arxiv_id":null,"evidence_quote":"The JTES corpus supplies the remaining training utterances and the four held-out evaluation speakers used in the listening tests."}],"review_version":1}