{"id":"40465d46-4906-447c-9752-3c1046cb43a3","arxiv_id":"2601.18094","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"OneVoice unifies zero-shot voice conversion for speech, expressive, and singing scenarios in one model via MoE routing and progressive training, matching specialized performance.","lead":"OneVoice is a single zero-shot voice conversion model that handles linguistic preservation, expressive speech, and singing using a continuous language model with Mixture-of-Experts and two-stage training. A smart generalist might read it to understand whether one AI system can replace multiple specialized models for voice cloning tasks.","discovery_kind":"unification","skeptic_critique":{"model":"grok-4.3","headline":"Two-stage LoRA training's ability to eliminate trade-offs from speech-singing data imbalance is the least-secured step in the unification claim.","rationale":"The reader's weakest_assumption is identical to the load-bearing concern identified above. Because the supplied text is still abstract-only and contains no quantitative ablations or data statistics, the UNVERDICTED / LOW verdict remains appropriate; the proposed concrete_test would directly test whether the assumption holds.","tokens_in":1730,"tokens_out":331,"duration_ms":36968,"concrete_test":"Re-run the singing-VC and speech-VC test sets after ablating the second-stage LoRA enhancement (i.e., using only the foundational checkpoint); if singing F0 accuracy or speaker similarity drops by more than 15 % relative while speech metrics stay within 3 %, the training strategy is load-bearing; if speech metrics also degrade, trade-offs exist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (matching or surpassing specialized models in linguistic-preserving, expressive, and singing VC) requires that the two-stage progressive training—foundational pre-training followed by LoRA-based domain-expert enhancement—successfully injects singing-specific expressivity while preserving speech performance. The abstract asserts this resolves the abundant-speech vs. scarce-singing imbalance via shared-expert isolation and scenario-aware routing, yet supplies no data-volume ratios, LoRA rank, or ablation numbers showing zero degradation on speech metrics when singing data is added. This assumption is therefore the single point whose failure would invalidate the “no trade-offs” part of the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents OneVoice, a unified zero-shot voice conversion framework that handles linguistic-preserving, expressive, and singing scenarios in a single model. It builds on a continuous language model trained with VAE-free next-patch diffusion, employs a Mixture-of-Experts architecture with dual-path routing (shared expert isolation plus scenario-aware domain expert assignment using global-local cues), fuses scenario-specific prosodic features via a gated mechanism, and uses two-stage progressive training (foundational pre-training followed by LoRA-based domain-expert enhancement) to address abundant-speech versus scarce-singing data imbalance. The central empirical claim is that OneVoice matches or surpasses specialized models across all three scenarios while enabling flexible scenario control and a fast 2-step decoding variant.","tokens_in":1872,"tokens_out":603,"duration_ms":38914,"significance":"If the performance claims and absence of trade-offs are substantiated, the work could consolidate three previously fragmented VC subfields into one deployable model, reducing the need for scenario-specific systems and offering practical gains in flexibility and inference speed. The explicit separation of shared conversion knowledge from scenario-specific expressivity via MoE and the progressive LoRA strategy constitute a clear methodological contribution, provided they are backed by ablations and reproducible metrics.","major_comments":[{"comment":"The unification claim without performance trade-offs rests on the two-stage progressive training successfully injecting singing-specific expressivity while preserving speech metrics. The manuscript provides no data-volume ratios between speech and singing corpora, no LoRA rank or adapter configuration details, and no ablation tables comparing speech-only metrics before versus after the LoRA singing-enhancement stage. Without these, it is impossible to verify that the “no trade-offs” assertion holds.","section":"Training Procedure and Experimental Results"},{"comment":"Table or figure reporting cross-scenario comparisons: the claim that OneVoice matches or surpasses specialized models requires explicit numerical results (e.g., MOS, WER, speaker similarity, F0 correlation) against named baselines for each of the three scenarios, together with statistical significance tests. The current presentation leaves the magnitude of any gains or equivalences unclear.","section":"Experiments"}],"minor_comments":[{"comment":"Notation for the dual-path routing and gated prosody fusion should be introduced with a single diagram or equation block early in the method section to improve readability.","section":"Method"},{"comment":"The fast-decoding (2-step) variant is mentioned only briefly; a short paragraph or table quantifying the quality-speed trade-off relative to the full model would strengthen the efficiency claim.","section":"Experiments"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of an audio/speech signal-processing venue. Citation of prior MoE and LoRA work in VC appears adequate, but the authors should confirm that no concurrent unified-VC submissions were overlooked."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed comments. We address each major comment below with clarifications and commit to revisions that will incorporate the requested details and tables to strengthen the manuscript.","responses":[{"response":"We agree that these specifics are necessary to substantiate the no-trade-offs claim. In the revised manuscript we will report the exact data-volume ratios between the speech and singing corpora. We will also document the LoRA rank and full adapter configuration in the training procedure section. In addition, we will add an ablation table that directly compares speech metrics (WER and speaker similarity) before versus after the LoRA singing-enhancement stage, thereby allowing readers to verify that speech performance is preserved.","revision_made":"yes","referee_comment":"[Training Procedure and Experimental Results] The unification claim without performance trade-offs rests on the two-stage progressive training successfully injecting singing-specific expressivity while preserving speech metrics. The manuscript provides no data-volume ratios between speech and singing corpora, no LoRA rank or adapter configuration details, and no ablation tables comparing speech-only metrics before versus after the LoRA singing-enhancement stage. Without these, it is impossible to verify that the “no trade-offs” assertion holds."},{"response":"We acknowledge that the current presentation summarizes results without the requested level of numerical detail or statistical tests. In the revision we will insert a new table (or expanded figure) that reports explicit values for MOS, WER, speaker similarity, and F0 correlation for OneVoice against the named specialized baselines in each of the three scenarios. We will also include the results of statistical significance tests (paired t-tests or Wilcoxon signed-rank tests) to quantify the observed equivalences or improvements.","revision_made":"yes","referee_comment":"[Experiments] Table or figure reporting cross-scenario comparisons: the claim that OneVoice matches or surpasses specialized models requires explicit numerical results (e.g., MOS, WER, speaker similarity, F0 correlation) against named baselines for each of the three scenarios, together with statistical significance tests. The current presentation leaves the magnitude of any gains or equivalences unclear."}],"tokens_in":1463,"tokens_out":457,"duration_ms":30745,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"OneVoice puts linguistic-preserving, expressive, and singing voice conversion into a single zero-shot model. The key question is how well the two-stage progressive training with LoRA-based experts handles the speech-singing data imbalance without creating trade-offs.","headline":"OneVoice unifies three VC scenarios in one model via MoE routing and staged LoRA training, but the no-trade-off claim on speech-singing data imbalance is the part that still needs tighter evidence.","tokens_in":2383,"tokens_out":134,"would_cite":false,"duration_ms":46090,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"Its core design for unification lies in a Mixture-of-Experts (MoE) designed to explicitly model shared conversion knowledge and scenario-specific expressivity... two-stage progressive training that includes foundational pre-training and scenario enhancement with LoRA-based domain experts."},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"Experiments show that OneVoice matches or surpasses specialized models across all three scenarios"}],"headline":"MoE routing + two-stage LoRA training for speech/singing VC has no structural overlap with RS cost or forcing chain","alignment":"orthogonal","rationale":"The paper's core machinery (dual-path MoE with shared/domain experts, gated prosody fusion, next-patch diffusion, and progressive pre-train + LoRA enhancement to handle speech/singing imbalance) operates entirely within standard ML engineering for audio generation. It invokes no recognition cost J, no φ-ladder identities, no 8-tick periodicity, and no parameter-free derivation of constants. RS modules such as Cost.FunctionalEquation (J-uniqueness) and Foundation.RealityFromDistinction therefore supply neither confirmation nor contradiction.","tokens_in":52670,"confidence":"high","tokens_out":330,"duration_ms":10354,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A single zero-shot model unifies linguistic-preserving, expressive, and singing voice conversion without trade-offs.","keywords":["voice conversion","zero-shot","mixture of experts","unified model","speech synthesis","singing voice conversion","prosody conditioning"],"falsifier":"Listening tests or objective metrics showing that OneVoice underperforms a dedicated singing voice conversion model on melody or pitch accuracy when the LoRA domain experts are removed would indicate the unification claim does not hold.","tokens_in":2635,"feed_emoji":"🎤","tokens_out":671,"duration_ms":29395,"temperature":0.7,"pith_summary":"The paper introduces OneVoice as a unified framework that performs zero-shot voice conversion across three scenarios—linguistic-preserving speech, expressive speech, and singing—using one model instead of specialized systems. It relies on a continuous language model trained via VAE-free next-patch diffusion and introduces a Mixture-of-Experts architecture with dual-path routing to separate shared conversion knowledge from scenario-specific expressivity. A two-stage progressive training process, including foundational pre-training followed by LoRA-based domain expert enhancement, addresses the imbalance between abundant speech data and scarce singing data. If the approach holds, practitioners could replace multiple dedicated models with a single flexible system that supports high-fidelity output and rapid decoding.","feed_headline":"One model covers speech, expressive, and singing voice conversion","feed_subtitle":"A unified zero-shot system matches specialized models across scenarios and supports decoding in two steps.","key_machinery":"Mixture-of-Experts with dual-path routing (shared expert isolation plus scenario-aware domain expert assignment using global-local cues) that explicitly separates shared conversion knowledge from scenario-specific expressivity, augmented by gated per-layer fusion of prosodic features.","core_discovery":"OneVoice achieves performance that matches or surpasses specialized models across linguistic-preserving, expressive, and singing voice conversion by combining a Mixture-of-Experts design with dual-path routing for shared and scenario-aware experts, gated fusion of scenario-specific prosodic features at every layer, and a two-stage training regime that uses LoRA-based domain experts to mitigate data imbalance while preserving a fast 2-step decoding option.","pith_inferences":["Deployment in resource-limited settings becomes simpler because one set of weights covers multiple voice conversion use cases.","The routing design may generalize to other audio domains where shared structure coexists with distinct stylistic requirements, such as music style transfer.","Further scaling the expert count or training data could reveal whether unification remains stable when singing data grows closer to speech volume."],"forward_implications":["The same model delivers competitive results in linguistic-preserving speech conversion, expressive speech conversion, and singing voice conversion.","Scenario control remains flexible through the routing and prosody mechanisms.","Decoding can be reduced to as few as two steps while retaining quality.","The architecture supports high-fidelity sequence modeling without a VAE."],"fun_headline_variants":["OneVoice unifies speech expressive and singing voice conversion","Single model covers all three zero-shot voice conversion scenarios","OneVoice matches specialists in three VC scenarios","Unified model for linguistic expressive singing VC zero-shot"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The two-stage progressive training with LoRA-based domain experts can sufficiently alleviate the data imbalance between abundant speech and scarce singing data to enable high performance in all three scenarios without trade-offs.","fun_headline_variants_meta":{"raw":{"variants":["OneVoice unifies speech expressive and singing voice conversion","Single model covers all three zero-shot voice conversion scenarios","OneVoice matches specialists in three VC scenarios","Unified model for linguistic expressive singing VC zero-shot"]},"model":"grok-4.3","cost_usd":0.007004,"raw_usage":{"total_tokens":3248,"prompt_tokens":678,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":70037000,"prompt_tokens_details":{"text_tokens":678,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2512,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":678,"tokens_out":58,"duration_ms":26772,"temperature":1.0,"reasoning_tokens":2512,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-22T11:32:07.701752+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Listening tests or objective metrics showing that OneVoice underperforms a dedicated singing voice conversion model on melody or pitch accuracy when the LoRA domain experts are removed would indicate the unification claim does not hold.","supporting_citations":[],"review_version":1}