{"id":"649c49a9-bf7c-4de4-8ee4-f8db4514c7e8","arxiv_id":"2412.01078","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A 60,000-hour bilingual synthetic speech dialogue dataset and a trained speech language model demonstrate strong speech interaction performance.","lead":"This paper introduces KE-Omni, a speech language model that can listen and respond in Chinese and English, trained on a new 60,000-hour synthetic speech dialogue dataset called Ke-SpeechChat. The work shows how large amounts of synthetic voice conversations can improve speech interaction quality and data efficiency.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Primary evidence for real-world speech interaction is in-distribution: both training and chat-test audio are CosyVoice TTS, so the headline advantage may reflect synthetic-speech overfitting rather than transfer to human speech.","rationale":"The reader's weakest assumption correctly identifies the synthetic-to-real transfer problem as the most load-bearing issue. The paper's headline results in Table 8 are obtained on chat-test, which is generated by the same CosyVoice TTS pipeline used to create the training data, so strong in-distribution performance does not by itself establish that the model can handle real human speech. The internal quality checks (DNSMOS/UTMOS, ASR on AISHELL-1/LibriSpeech, TTS on SeedTTS test sets) provide some evidence that Ke-SpeechChat is acoustically clean and intelligible, but they do not test the conversational interaction setting that the central claim targets. VoiceBench includes real spoken instructions, yet the paper reports only pooled scores, making it impossible to determine whether KE-Omni's competitiveness there comes from real or synthetic inputs. A secondary but related weakness is that the scaling study in Section 5.5 varies batch size across subsets to keep step counts consistent, so the observed improvements are not a clean function of data quantity; however, the external-validity concern is more fundamental because it affects the interpretation of every reported interaction result. For these reasons, the paper's core claim should remain conditional: it is plausible and partially supported, but it requires a dedicated real-speech interaction evaluation before it can be accepted. The reader's conditional verdict, with moderate confidence, is therefore appropriate, and no change to that verdict is needed.","tokens_in":18804,"tokens_out":6599,"duration_ms":64430,"concrete_test":"Re-evaluate KE-Omni-L and KE-Omni-XL on the real-speech instruction subsets of VoiceBench (and, if feasible, on a newly recorded set of human speakers asking the chat-test questions) using the same S2TIF and VoiceBench protocols, reporting results separately for real versus synthetic audio. If the KE-Omni advantage over LLaMA-Omni and Qwen2-Audio observed in Table 8 shrinks or disappears on real speech, the synthetic-data proxy assumption fails; if the margin persists on real speech, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Ke-SpeechChat is high quality for training a bilingual speech language model, and that KE-Omni outperforms baselines, rests on the assumption that CosyVoice-synthesized dialogues are a sufficient proxy for real human speech interaction. This premise is untested in the paper. All training audio is generated by CosyVoice from Qwen-rewritten text (§3.2), and the main interaction benchmark, chat-test, is also synthesized by CosyVoice with the same pipeline (§5.2, Table 10). The only external benchmark, VoiceBench, mixes real and synthetic instructions, but Table 9 reports only pooled scores, so real-speech transfer cannot be isolated. The limitation section (§6) admits the synthetic data are clean, noiseless, and single-turn, but it does not acknowledge that the central evaluation is in-distribution. If KE-Omni's large gains in Table 8 are partly due to matching TTS artifacts, its performance on real human speech may be much closer to—or worse than—the baselines, which would undermine the paper's claim to advance real-time speech interaction. This is a correctness risk, not merely a scope limitation, because the paper presents KE-Omni as a seamless large speech language model for practical use.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Ke-SpeechChat, a large-scale synthetic bilingual speech dialogue dataset containing roughly 6.97 million Chinese and English conversations (about 60,000 hours) generated by rewriting open-source instruction data with Qwen LLMs and synthesizing speech with CosyVoice using virtual speaker embeddings derived from WenetSpeech4TTS. It also presents KE-Omni, an end-to-end speech-to-speech model built on a Whisper encoder, a LLaMA-3.1-8B backbone, and a duration-predictor/unit-generator/vocoder decoder. The authors report that KE-Omni outperforms existing speech language models on their own chat-test set across instruction-following, modality alignment, and speech quality metrics, and achieves competitive results on VoiceBench. The dataset quality is further assessed via DNSMOS/UTMOS, ASR fine-tuning on Whisper, and zero-shot TTS based on CosyVoice.","tokens_in":19082,"tokens_out":7198,"duration_ms":60369,"significance":"If the reported results are robust, the paper would provide a substantial resource for bilingual speech language model research, especially for Chinese, and a detailed pipeline for constructing synthetic speech interaction data at scale. The authors deserve credit for the detailed documentation of the data construction prompts (Appendix A), explicit quality-control stages (CER/WER filtering, DNSMOS-based voice selection), and the multi-task evaluation spanning ASR, TTS, speech interaction, and safety benchmarks. The dataset statistics and the plan for a two-stage training recipe are clearly presented. However, the central evidence for the model's real-world interaction capability is weakened by the in-distribution nature of the main test set and by the confounded scaling analysis. The significance of the contribution would be substantially strengthened by addressing these evaluation issues.","major_comments":[{"comment":"The scaling study across the XS, S, M, L, and XL subsets confounds data volume with batch size. Section 5.1 states that 'in order to make the number of training steps as consistent as possible across all datasets, we adopted different batch sizes for different datasets in the second stage,' and Section 5.5 attributes the non-monotonic CER/WER (L better than XL) to the larger batch size in XL. As a result, the comparison does not isolate the effect of training data scale, and the claim that scaling data improves performance is not cleanly supported. Please either hold the batch size fixed across subsets, run a controlled ablation (e.g., XL with the L batch size), or explicitly frame the results as a joint function of data size and batch size. Reporting the batch sizes used for each subset would at least make the confound transparent.","section":"Section 5.1 and Section 5.5"},{"comment":"The primary interaction benchmark, chat-test, is generated with the same CosyVoice pipeline and the same style of virtual-speaker audio as the training data. Training and evaluation are therefore in-distribution with respect to the TTS system and voice library. The large margins over baselines in Table 8 may partly reflect adaptation to CosyVoice artifacts rather than generalizable speech interaction ability, especially since the baseline models were not trained on CosyVoice audio. The authors should evaluate KE-Omni on real human speech (e.g., recorded human instructions on the same tasks or a public spoken dialogue benchmark) and, for the VoiceBench results in Table 9, report the real and synthetic instruction subsets separately instead of pooled scores. Section 6 should explicitly acknowledge the in-distribution nature of the main test set.","section":"Section 5.2 and Table 8 / Table 10"},{"comment":"All reported results are single runs without error bars, confidence intervals, or significance tests. Several headline comparisons are close: e.g., L vs XL in Chinese CER (5.03 vs 5.16), L vs XL in UTMOS (3.39 vs 3.43), and KE-Omni-L vs KE-Omni-XL on AlpacaEval (3.74 vs 3.78). Without variance estimates, it is impossible to judge whether these differences are meaningful. Please provide results over at least three random seeds (or repeated evaluations with different holdout subsets) and apply appropriate statistical tests for the main comparisons in Table 8 and Table 9.","section":"Tables 6-9 and Section 5.5"},{"comment":"The claim that 'KE-Omni achieves significantly better performance than baseline systems when trained on datasets of comparable size' is not substantiated by a controlled comparison. The baseline models (LLaMA-Omni, SpeechGPT, Qwen2-Audio) were trained on their own datasets, not on the Ke-SpeechChat subsets, so the comparison confounds model architecture, training data, and training procedure. To support this claim, the authors would need to train the baseline architectures on the same Ke-SpeechChat subsets, or at least train them on the same amount of data drawn from the same distribution. Otherwise, the claim should be rephrased as 'KE-Omni outperforms off-the-shelf baselines on our test set,' which is a weaker but accurate statement.","section":"Section 5.5, Table 8"},{"comment":"The quality-assurance step in dataset construction uses Whisper-family ASR models to compute CER/WER and filter dialogues, and the subsequent modality-alignment evaluation of KE-Omni also uses Whisper-large-v3 for transcription. This creates a closed-loop risk: the training data is selected to be easily transcribable by Whisper, which may inflate the measured CER/WER of the final model. The authors should evaluate the generated speech with a different ASR system (e.g., an independently trained model) and discuss the potential bias. At minimum, they should acknowledge this circularity in Section 6.","section":"Section 3.2.3 and Section 5.3"}],"minor_comments":[{"comment":"The sentence '...totaling over 60,000 hours, This contributes significantly...' is a run-on with a missing period; 'This' should begin a new sentence and likely refer to 'this dataset' or 'this work'.","section":"Abstract"},{"comment":"The text says 'including 40,000 users and 2 agents,' but Table 1 lists 21,000 male users and 21,000 female users, totaling 42,000 users. Please correct the number to 42,000 and clarify whether the same user speaker set is used for both languages.","section":"Section 3.3.2"},{"comment":"The sentence 'KE-Omni use LLaMA-3.1-8B-Instruct(Fang et al., 2024)' cites Fang et al. (LLaMA-Omni) for the LLaMA-3.1 model; the correct reference is Dubey et al. (2024), the LLaMA 3 herd paper. Also, 'use' should be 'uses'.","section":"Section 5.1"},{"comment":"The phrase 'KE-Omni outperforms to other baseline systems' is grammatically incorrect; it should be 'outperforms other baseline systems.'","section":"Section 5.5"},{"comment":"The condition 'Max Count ≥ 5 × ⌊n/10⌋' appears in the figure caption but is not defined in the main text; please introduce the notation for ⌊n/10⌋ and explain the threshold clearly in Section 3.2.1.","section":"Figure 2 and Section 3.2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a substantial dataset and an interesting training recipe, but the evaluation currently does not support real-world interaction claims strongly enough for acceptance. The in-distribution test set and the batch-size confound are the key issues; both are addressable with additional experiments. I would encourage the editor to request a revision that adds real-speech evaluation, controls batch size, and reports variance. The manuscript also cites '7 million conversations' while Table 1 sums to 6.97 million; this is minor, but a consistent rounding or exact figure would be preferable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Ke-SpeechChat is the real contribution. The 60k-hour bilingual synthetic dialogue dataset, built by rewriting instruction data into spoken-style Q&A and synthesizing with CosyVoice across 42k virtual voices, is a valuable resource. The pipeline is sensible, the quality filters (ASR CER/WER, DNSMOS/UTMOS) are appropriate, and the downstream ASR/TTS checks give some evidence of quality. The virtual-speaker construction via weighted averaging of real embeddings is a nice privacy-preserving move.\n\nKE-Omni itself is a fairly standard LLaMA-Omni-style architecture: Whisper encoder, LLaMA backbone, HuBERT units plus HiFi-GAN decoder. The scaling story is the main empirical finding: more synthetic data clearly helps S2TIF and speech quality, and CER/WER drop sharply from XS to S. That is useful for the field.\n\nNow the soft spots. The strongest one: the chat-test used in Table 8 is also CosyVoice TTS, with the same speaker library design. So the big gains over LLaMA-Omni are measured in-distribution. The model may be learning TTS artifacts rather than general speech understanding. Without a benchmark on real human speech, the claim that KE-Omni advances practical real-time interaction is overstated. This doesn't invalidate the dataset, which is openly synthetic, but it should be labeled as a synthetic-speech evaluation, not a general speech interaction evaluation.\n\nSecond, the scaling experiment confounds data size with batch size. The authors themselves note that XL's higher CER/WER than L may be because the larger batch size in XL negatively affects performance. That's an uncontrolled variable, so the scaling conclusions are weaker than the headline suggests.\n\nThird, no error bars or significance tests appear anywhere. Differences of 0.05 in S2TIF style scores are probably noise. Fourth, Whisper is used both to filter the synthetic audio and to compute the modality-alignment metrics, so there is a closed loop; a different ASR might give different numbers.\n\nOn citations, they cover the relevant prior work adequately, including LLaMA-Omni, Moshi, VoiceBench, and CosyVoice. I don't see a citation problem.\n\nBottom line: this is a dataset paper, and a useful one. The model is a competent baseline. The overclaim is in framing the work as advancing speech interaction when the evidence is confined to synthetic speech. A serious referee should ask for an evaluation on a real-speech benchmark, or at least a clear statement that results are for synthetic speech; deconfounding batch size; and error bars. I'd accept it for review, but with the expectation of significant revision.","headline":"A large bilingual synthetic speech dialogue dataset is the real contribution; the model gains are real but measured only on synthetic speech, so the practical-interaction claim is overstated.","tokens_in":19576,"tokens_out":2647,"would_cite":true,"duration_ms":24266,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 60,000-hour synthetic speech dialogue dataset can train a bilingual speech language model that outperforms comparable baselines.","keywords":["speech language model","synthetic speech dialogue dataset","supervised fine-tuning","bilingual speech interaction","text-to-speech data augmentation","speech-text modality alignment","KE-Omni","Ke-SpeechChat"],"falsifier":"Evaluate KE-Omni on a held-out set of real, spontaneous human conversations (with background noise, disfluencies, and overlapping speech) and measure WER/CER and human-rated naturalness; if scores drop materially relative to the synthetic chat-test, the claim that synthetic speech is a sufficient training proxy is overturned.","tokens_in":18633,"feed_emoji":"🗣️","tokens_out":3790,"duration_ms":32771,"temperature":0.7,"pith_summary":"This paper argues that large-scale synthetic speech dialogue data can substitute for expensive real recordings when training a speech language model. It contributes a dataset construction pipeline—LLM rewriting of existing instruction text into spoken-style dialogues, TTS synthesis with privacy-protecting virtual voices, and CER/WER filtering—that produced 7 million Chinese and English conversations totalling over 60,000 hours. Trained on subsets of this data, the authors' KE-Omni model outperforms English-only baselines at comparable data scale on instruction-following, speech-text alignment, and speech quality. The practical stake is lowering the data cost and privacy barriers to building real-time speech assistants, especially for Chinese.","feed_headline":"60,000 hours of synthetic speech trains a better voice AI","feed_subtitle":"Rewritten text plus TTS voices beat English-only speech models in Chinese and English.","key_machinery":"The load-bearing mechanism is the synthetic data factory. Open-source instruction datasets are rewritten by LLMs (Qwen2.5-14B and 72B) through rewriting, filtering, and spoken-style post-processing to produce conversational text. A virtual voice library is built from premium WenetSpeech4TTS clips, using WavLM x-vectors to cluster clips by speaker, then averaging pairs of same-gender, same-rate voice embeddings to create privacy-protecting synthetic speakers that do not correspond to real individuals. CosyVoice synthesizes speech for these dialogues, AudioSeal watermarks the audio, and ASR-based CER/WER filtering drops low-quality clips. The KE-Omni model itself combines a frozen Whisper encoder with a trainable adapter compressing audio to 10 frames per second, a LLaMA-3.1-8B backbone, and a speech decoder with a duration predictor, a chunk-based autoregressive speech unit generator, and a HiFi-GAN vocoder.","core_discovery":"The paper's central claim is that Ke-SpeechChat, a fully synthetic speech dialogue dataset built from rewritten text instructions, is high quality, and that a speech language model trained on it scales effectively. KE-Omni, trained on this dataset, achieves significantly better speech-to-text instruction-following scores than LLaMA-Omni, SpeechGPT, and Qwen2-Audio when trained on comparable data. Modality alignment, measured by CER/WER of spoken output, improves sharply as training data grows from 1,500 to 6,000 hours and saturates at the largest subsets, while speech quality (UTMOS) keeps improving up to the full 60,000 hours. The authors also report that KE-Omni performs competitively on the VoiceBench benchmark, with the caveat that its refusal style and spoken-form paragraph cues lower scores on AdvBench and IFEval.","pith_inferences":["The same pipeline likely transfers to other languages or domains where instruction text exists but conversational audio is scarce.","Because both training and test sets are synthetic and clean, real-world robustness to noise, disfluency, and overlapping speech remains an open question the paper does not settle; mixing synthetic with real dialogue data is a natural next step.","The fact that the L subset beats XL on CER/WER hints at a saturating or non-monotonic benefit of synthetic scale that deserves isolation from batch-size effects."],"forward_implications":["Speech language models can be trained from text-only instruction corpora plus TTS, removing the need for recorded conversational audio.","Data scale matters: modality alignment and speech quality improve as synthetic data grows, with alignment improving sharply beyond 4,000 hours.","Privacy can be built into the data pipeline via composite virtual voices and watermarking, addressing voice-misuse concerns.","Bilingual speech interaction research is extended beyond English-only systems such as LLaMA-Omni and SpeechGPT."],"supporting_citations":[{"why":"CosyVoice: the TTS engine that converts rewritten dialogues into all training and test audio, and the baseline for the TTS evaluation.","marker":"[Du et al., 2024]"},{"why":"WenetSpeech4TTS: supplies the premium clips that form the real-speaker library later blended into virtual voices.","marker":"[Ma et al., 2024]"},{"why":"WavLM: provides x-vector extraction for grouping clips by speaker and building stable speaker embeddings.","marker":"[Chen et al., 2022]"},{"why":"Whisper: used for CER/WER filtering of synthetic audio, for transcribing outputs in evaluation, and as the frozen speech encoder of KE-Omni.","marker":"[Radford et al., 2023]"},{"why":"LLaMA-3.1-8B-Instruct: the LLM backbone that is fine-tuned on Ke-SpeechChat and generates the text responses.","marker":"[Dubey et al., 2024]"},{"why":"HuBERT: extracts continuous speech representations converted into discrete units for the speech decoder's unit generation.","marker":"[Hsu et al., 2021]"},{"why":"LLaMA-Omni: source of the architecture pattern for the speech decoder and the primary English-speaking baseline in the comparison.","marker":"[Fang et al., 2024]"},{"why":"DNSMOS: the quality metric used to select premium clips and to compare Ke-SpeechChat's audio quality with existing datasets.","marker":"[Reddy et al., 2022]"}],"fun_headline_variants":["60k hours of synthetic speech beat LLaMA-Omni and Qwen2-Audio","Synthetic dialogue at 60,000 hours improves speech LM scaling","Scaling voice AI with 60k hours of synthetic speech data","From 1,500 to 60k hours: synthetic speech scales voice LLMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Speech synthesized by CosyVoice from LLM-rewritten text is a good enough stand-in for real human conversation to train a speech assistant, and the synthetic test sets fairly measure real-world ability.","fun_headline_variants_meta":{"raw":{"variants":["60k hours of synthetic speech beat LLaMA-Omni and Qwen2-Audio","Synthetic dialogue at 60,000 hours improves speech LM scaling","Scaling voice AI with 60k hours of synthetic speech data","From 1,500 to 60k hours: synthetic speech scales voice LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000437,"raw_usage":{"total_tokens":2195,"prompt_tokens":891,"completion_tokens":1304,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":1219}},"tokens_in":507,"tokens_out":1304,"duration_ms":8822,"temperature":1.0,"reasoning_tokens":1219,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:42:06.068684+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate KE-Omni on a held-out set of real, spontaneous human conversations (with background noise, disfluencies, and overlapping speech) and measure WER/CER and human-rated naturalness; if scores drop materially relative to the synthetic chat-test, the claim that synthetic speech is a sufficient training proxy is overturned.","supporting_citations":[],"review_version":1}