{"id":"204b6c45-1bf7-4768-9605-cd9df530facd","arxiv_id":"2505.22029","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"The paper introduces LLM-Dys, a 12,790-hour synthetic dysfluent speech corpus generated by LLM plus TTS, and claims state-of-the-art dysfluency detection with a Whisper-based transcriber.","lead":"This paper creates a huge synthetic dataset of stuttered speech, 12,790 hours across 11 dysfluency types, using an LLM to write the disfluent text and a TTS model to speak it. It then shows a Whisper-based system trained on this data can detect dysfluencies in real speech, though the evaluation method is questionable.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed SOTA on real stuttered speech is unsupported because SEP-28k evaluation uses a self-annotated test set in the model's own label format, and baseline comparisons use different test sets and label schemes.","rationale":"I agree with the reader's assessment. The paper's main claim is that LLM-Dys enables SOTA performance on real stuttered speech. However, the evidence for this is compromised by a non-standard test set: Section 4.1 explicitly says the SEP-28k evaluation set was manually annotated by the authors in the model's output format, not the official SEP-28k schema. Tables 3 and 4 compare against baselines whose scores come from a different test set and different label scheme; the paper even states that block-level baseline scores are applied to both word- and phoneme-level rows. This makes the SOTA claim untestable from the paper's reported numbers. The UCLASS evaluation is also on a small self-segmented subset with binary 1:1 labels, not the standard UCLASS protocol, though the SEP-28k issue alone is sufficient. I give credit for releasing the dataset and code, and for the internal consistency of the synthetic benchmark (Table 2 shows the model learns the synthetic task), but the external validity claim fails. A single official-benchmark evaluation would settle it.","tokens_in":14199,"tokens_out":3179,"duration_ms":34771,"concrete_test":"Run the released model zero-shot on the official SEP-28k test set using the original Lea et al. annotations (same clips and label schema as Wagner et al. [11]) and report per-class F1. If the model does not exceed the [11] numbers on identical data and labels, the SOTA claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central SOTA claim is load-bearing and unsupported. In Section 4.1 the authors create their own SEP-28k test set by manually annotating 200 samples per level 'following our model's output format with dysfluency tokens added to clean text.' This is not the official SEP-28k label schema, so scores in Tables 3-4 are not comparable to Wagner et al. [11] or Yolo-Stutter [14], which were evaluated on official SEP-28k labels. Section 4.3.2 even states that the block-level scores from [11] are 'applied to both word-level and phoneme-level metrics' — i.e., the baseline numbers are not re-computed on the authors' test set. Thus the 'consistently achieving state-of-the-art performance' claim in the abstract and conclusion is not demonstrated. The self-annotated set is also tiny (200 samples per level) and annotated by the authors, creating risk of unconscious label bias and format overfitting. The dataset itself may be a useful resource, but the paper's central empirical claim requires independent validation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces LLM-Dys, a synthetic dysfluent speech corpus generated by prompting a large language model (Claude 3.5 Sonnet) to produce dysfluent text with token-level labels, then synthesizing speech with VITS and E2-TTS. Pauses and prolongations are inserted post-hoc at specified durations. The corpus spans 11 dysfluency types at word and phoneme levels and amounts to over 12,000 hours across 109 speakers. The authors train a Whisper-based transcription model on selected subsets and report results on (i) held-out synthetic data, (ii) a self-annotated subset of SEP-28k (200 samples per level), and (iii) a self-defined binary split of UCLASS. They claim state-of-the-art performance on real-world dysfluency detection and discuss the effect of training set size and fluent-to-disfluent ratio.","tokens_in":14489,"tokens_out":6288,"duration_ms":62108,"significance":"If validated, LLM-Dys would be a valuable contribution: it is one of the first efforts to use LLM-generated text diversity for dysfluency simulation, and the open-sourced dataset and code are commendable. The idea of using LLMs to generate more naturalistic dysfluent text is a plausible step beyond rule-based text insertion. However, the central empirical claim—that the resulting model achieves state-of-the-art performance on real stuttered speech—is not currently supported by the evaluation protocol: the SEP-28k test set is re-annotated in the authors' own format, baselines are cited from other papers without rerunning on the same test set, and no error bars are provided. The near-perfect synthetic test scores are expected given shared generation pipelines. The dataset itself may still be useful to the community, but the paper's headline claims require substantially stronger validation.","major_comments":[{"comment":"The SEP-28k evaluation is not conducted on the official SEP-28k labels. The authors manually annotate 200 samples per level 'following our model's output format with dysfluency tokens added to clean text,' and in Section 4.3.2 they state that block-level scores from [11] are 'applied to both word-level and phoneme-level metrics.' Consequently, the numbers in Tables 3-4 are not directly comparable to Wagner et al. [11] or Yolo-Stutter [14], which were evaluated on official SEP-28k labels. The claimed state-of-the-art performance on real stuttered speech is therefore not established. To support the claim, the authors should evaluate on the standard SEP-28k benchmark (or at least the same test clips and label schema as prior work), or rerun the baselines on their new test set, and report inter-annotator agreement for the new annotations.","section":"Section 4.1 and Table 3"},{"comment":"The near-perfect results on LLM-Dys are expected because the test set is drawn from the same LLM prompt/TTS pipeline used for training; the authors' own explanation ('consistent patterns in LLM-generated dysfluencies and our standardized TTS pipeline') confirms that this evaluation cannot measure generalization to real dysfluent speech. These results should be presented as a sanity check, not as evidence of system capability.","section":"Section 4.3.1 and Table 2"},{"comment":"The UCLASS evaluation uses a self-defined 200-sample split (80 fine-tuning, 120 testing) with binary labels, and the comparison to StutterNet is made against previously published results. It is unclear whether the splits and label protocols are matched; no variance or confidence intervals are reported anywhere. The authors should report results over multiple random splits or runs and, if possible, use the same evaluation protocol as the baseline.","section":"Section 4.1 and Table 5"},{"comment":"The claim that LLM-Dys has 'superior synthesis quality' and is 'comparable to real fluent speech' rests on the Meta Audiobox Aesthetics model, a generic audio quality predictor. There is no evidence that this model is sensitive to the naturalness of dysfluency patterns specifically, and it may simply reward pleasant prosody. A human perceptual study with dysfluent speakers or clinicians would be needed to support the naturalness claim.","section":"Section 2.4 and Fig. 2"}],"minor_comments":[{"comment":"The duration ranges for pauses (0.8-3.5s word-level, 0.3-1.5s phoneme-level) and prolongations (0.17-0.8s) are stated without justification; please provide at least a brief explanation of how they were chosen, or a sensitivity analysis.","section":"Section 2.2"},{"comment":"'Scaling Law' is an overstatement for a training-set-size ablation; consider renaming to 'Effect of Training Data Size'.","section":"Section 4.3.3"},{"comment":"The reported hours are computed by multiplying per-speaker durations by 109 speakers, which inflates the apparent size of the corpus. The number of unique LLM-generated utterances per type is the better measure of text diversity and should be highlighted.","section":"Table 1"},{"comment":"These figures are difficult to read in the current version; consider enlarging fonts, expanding abbreviations, and ensuring all axis labels are legible.","section":"Figures 2, 3, and 5"}],"recommendation":"major_revision","confidential_remarks":"The paper has a promising dataset and methodology, but the evaluation protocol for the central state-of-the-art claim is not defensible. I recommend requiring the authors to re-run experiments on official SEP-28k labels (or an equally standard protocol) and to compare with baselines on the identical test set. If this is not feasible, the claims of state-of-the-art performance should be withdrawn, and the paper should be refocused on dataset construction. The bar for acceptance should include independent validation of the test annotations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the dataset is the contribution, not the benchmark numbers. LLM-Dys is a 12,790-hour, 11-category synthetic dysfluency corpus generated with an LLM plus VITS/E2-TTS, with data and code released. That is a real resource and a step beyond rule-based simulation in VCTK-TTS, VCTK-Token, and Libri-Dys. The POS analysis and the scaling/ratio ablations are reasonable and useful. But the paper's central claim—state-of-the-art on real stuttered speech—does not hold up as written.\n\nThe SEP-28k evaluation is self-referential. Section 4.1 says the authors manually annotated 200 samples per level \"following our model's output format with dysfluency tokens added to clean text.\" That is not the official SEP-28k label schema, so the numbers in Tables 3–4 are not comparable to Wagner et al. or Yolo-Stutter, which used official labels. Section 4.3.2 explicitly says block-level scores from Wagner are applied to both word-level and phoneme-level metrics; the baselines were not recomputed on this test set. There are also no error bars. The synthetic test results are high but expected, since test and train share the same generation pipeline. The zero-shot framing is not a substitute for a common evaluation set.\n\nThe UCLASS experiment is more standard but small: 80 fine-tuning clips, 120 test clips, binary labels. It is suggestive, not a load-bearing claim. The audio quality evaluation uses Meta Audiobox Aesthetics, an automatic model, rather than human raters; that is fine as a first-pass check but weaker than the \"comparable to real fluent speech\" phrasing in the introduction.\n\nWho gets value: researchers building or using simulated dysfluency corpora for clinical speech assessment. The dataset and code are likely to be useful even if the benchmark claims need rework.\n\nRecommendation: send to peer review, but with major revision. The authors should either annotate the standard SEP-28k label schema (or use an existing official subset), recompute all baselines on the same test set, and report variance. If they do that, the dataset could stand on its own without the unsupported SOTA claim.","headline":"The dataset is a genuinely useful resource; the claimed SOTA on real speech is not supported by the evaluation.","tokens_in":15031,"tokens_out":2564,"would_cite":true,"duration_ms":28161,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that prompting LLMs to write dysfluent text and synthesizing it with TTS produces a 12,790-hour corpus, and that a detector trained on it beats prior results on real stuttered speech.","keywords":["speech dysfluency detection","synthetic data generation","large language models","text-to-speech","stuttering detection","dysfluency corpus","Whisper","SEP-28k"],"falsifier":"Have independent annotators label the same SEP-28k subset using the dataset's original annotation scheme, then rerun the zero-shot evaluation; if the LLM-Dys model's F1 falls below the reported baselines (for example, word insertion 0.77, word repetition 0.64, word pause 0.62 from Wagner et al.), the state-of-the-art claim on real stuttered speech collapses.","tokens_in":14027,"feed_emoji":"🗣️","tokens_out":7007,"duration_ms":64328,"temperature":0.7,"pith_summary":"Speech dysfluency detection is held back by scarce, inconsistently annotated real stutter data. The paper's central claim is that synthetic data can close that gap when the text simulations are written by a large language model rather than hand-built rules: it introduces LLM-Dys, a corpus of over 12,790 hours covering 11 dysfluency types at word and phoneme levels, which it reports sounds as natural as fluent speech and more natural than prior simulated corpora. Fine-tuning a token-based Whisper detector on this corpus yields state-of-the-art F1 scores on real stuttered speech benchmarks, SEP-28k and UCLASS, while the authors also document a scaling plateau and an optimal fluent-to-disfluent training ratio.","feed_headline":"12,790 hours of LLM-synthesized stutter speech sets detection highs","feed_subtitle":"LLM prompting plus TTS yields a huge naturalistic corpus; detection on real stuttered speech improves.","key_machinery":"The load-bearing object is the paired dysfluent-text-plus-label generator, where the LLM supplies authentic dysfluency patterns from its learned language model and its output tokens double as training labels, eliminating manual annotation. The TTS stage uses VITS's duration-prediction matrix to realize pauses and prolongations at precise timestamps, and E2-TTS for filler-word insertions, which is what keeps the synthesized audio natural rather than rule-pasted. On the detection side, Whisper-large-v3-turbo is fine-tuned to transcribe dysfluency tokens inserted into clean text, outputting dysfluency type and location jointly for eleven categories.","core_discovery":"LLM-Dys rests on a two-stage pipeline: first a large language model rewrites clean sentences into dysfluent text, generating word-level insertions, repetitions, pauses, deletions, and substitutions plus phoneme-level versions from CMU/IPA transcriptions, while emitting matching token labels at the same time. These texts are synthesized with the VITS text-to-speech model, with pauses inserted as silent gaps of 0.3-3.5 seconds and prolongations produced by stretching phoneme durations in VITS's duration matrix, and E2-TTS is used for word-level insertions because VITS handles filler words poorly. The authors report that this pipeline produces higher Audiobox Aesthetics scores than VCTK-Token, VCTK, LibriTTS, and LibriStutter, and that their Whisper-large-v3-turbo transcriber reaches F1 0.91 for word insertion, 0.67 for word repetition, and 0.79 for word pause on a self-annotated SEP-28k test subset, beating prior published numbers, plus 0.977 word-level accuracy on UCLASS after fine-tuning on only 80 clips.","pith_inferences":["Inference: if synthetic dysfluent speech keeps closing the quality gap with real recordings, the field can benchmark detectors on procedurally generated patients, dialects, and speaking styles, making evaluation fair for groups underrepresented in SEP-28k and UCLASS.","Inference: the SEP-28k evaluation relies on a self-annotated 200-sample subset labeled in the model's own token format; a natural next test is the original SEP-28k annotations used by the compared baseline, and if the advantage shrinks there, the state-of-the-art claim is benchmark-specific.","Inference: the same LLM-prompt-to-TTS recipe could generate dysfluency data for other languages or for non-stutter disorders such as apraxia and aphasia, provided a TTS with sufficient phonetic control exists.","Inference: the duration-matrix manipulation for pauses and prolongations suggests TTS acoustic models can be instrumented to emit event-level timing supervision, a technique that might transfer to other timestamp-annotated audio tasks."],"forward_implications":["If correct, end-to-end dysfluency detection can be trained almost entirely on synthetic speech, with real data reserved for a small test set.","The eleven-type, word-and-phoneme unified annotation scheme gives one detector format for tasks that previously used binary, event-level, and ASR-filler labels separately.","The scaling plateau at 3,000-4,000 samples per type implies that far less than the full 12,790 hours may suffice, allowing cheap subsampling for fast iteration.","The roughly 0.05 optimal fluent-to-disfluent ratio tells practitioners to keep fluent speech in the training mix, or synthetic-only models will over-predict dysfluencies.","The UCLASS result suggests that synthetic pretraining transfers to binary stuttering classification with minimal real fine-tuning."],"supporting_citations":[{"why":"Supplies the prior simulation pipeline and token-level label format that LLM-Dys extends from rule-based to LLM-generated text.","marker":"[23]"},{"why":"Provides the VITS text-to-speech model that synthesizes dysfluent text and whose duration matrix enables pause and prolongation insertion.","marker":"[18]"},{"why":"Gives the Whisper-large-v3-turbo pretrained model that is fine-tuned with token labels for end-to-end dysfluency detection.","marker":"[32]"},{"why":"Supplies SEP-28k, the real stuttered speech dataset from which the authors manually annotate their zero-shot test subset.","marker":"[16]"},{"why":"Supplies UCLASS, the real stuttered speech archive used for the binary classification evaluation with 80 fine-tuning and 120 test clips.","marker":"[15]"},{"why":"Provides the Wagner et al. baseline whose block-level F1 scores are compared against the LLM-Dys detector on SEP-28k.","marker":"[11]"},{"why":"Provides the Yolo-Stutter baseline and its recall figures for comparison on stuttered speech detection.","marker":"[14]"},{"why":"Supplies the Meta Audiobox Aesthetics evaluation tool whose CE, CU, and PQ metrics support the claim that LLM-Dys audio is natural.","marker":"[26]"},{"why":"Provides LibriTTS as a fluent comparison corpus and as fluent training data in the fluent-to-disfluent ratio experiments.","marker":"[21]"}],"fun_headline_variants":["LLM-TTS pipeline yields 11 dysfluency types, boosts detection","LLM-Dys: synthetic stutter corpus beats prior synthetic data","Two-stage LLM + TTS improves real stutter detection","LLM rewrites plus TTS synth data lift dysfluency F1"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The narrowest load-bearing premise is that a test set the authors built themselves, by hand-labeling 200 SEP-28k clips per level in the model's own token format, fairly measures real-world dysfluency detection.","fun_headline_variants_meta":{"raw":{"variants":["LLM-TTS pipeline yields 11 dysfluency types, boosts detection","LLM-Dys: synthetic stutter corpus beats prior synthetic data","Two-stage LLM + TTS improves real stutter detection","LLM rewrites plus TTS synth data lift dysfluency F1"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000172,"raw_usage":{"total_tokens":1259,"prompt_tokens":913,"completion_tokens":346,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":268}},"tokens_in":529,"tokens_out":346,"duration_ms":4209,"temperature":1.0,"reasoning_tokens":268,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:16:06.708706+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent annotators label the same SEP-28k subset using the dataset's original annotation scheme, then rerun the zero-shot evaluation; if the LLM-Dys model's F1 falls below the reported baselines (for example, word insertion 0.77, word repetition 0.64, word pause 0.62 from Wagner et al.), the state-of-the-art claim on real stuttered speech collapses.","supporting_citations":[{"cited_title":"Frame-level stut- ter detection,","cited_arxiv_id":null,"evidence_quote":"Provides the VITS text-to-speech model that synthesizes dysfluent text and whose duration matrix enables pause and prolongation insertion."},{"cited_title":"Detect- ing dysfluencies in stuttering therapy using wav2vec 2.0,","cited_arxiv_id":null,"evidence_quote":"Supplies UCLASS, the real stuttered speech archive used for the binary classification evaluation with 80 fine-tuning and 120 test clips."},{"cited_title":"Self- supervised speech models for word-level stuttered speech de- tection,","cited_arxiv_id":null,"evidence_quote":"Provides the Wagner et al. baseline whose block-level F1 scores are compared against the LLM-Dys detector on SEP-28k."},{"cited_title":"VCTK-TTS demonstrates significantly bet- ter intelligibility and naturalness than all previous datasets","cited_arxiv_id":null,"evidence_quote":"Provides the Yolo-Stutter baseline and its recall figures for comparison on stuttered speech detection."}],"review_version":1}