{"id":"fedd59b1-5285-48ea-b293-14536e8b5efc","arxiv_id":"2501.15877","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new open multilingual stuttered speech dataset for Indian languages, with read and spontaneous speech, word-level annotations, and experiential questionnaire data.","lead":"Project Boli is a new multilingual dataset of stuttered speech from 28 Indian speakers, with word-level annotations of five stutter types and questionnaire responses about living with stuttering. It aims to support speech technology and research for people who stutter, especially in Indian languages.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Boli dataset's central claim is unverifiable as written: no download link or repository is provided, and the word-level annotations come with no protocol or inter-annotator agreement, so neither open access nor label reliability can be checked.","rationale":"Reading in good faith, the paper's contribution is a resource, not primarily a claim about stuttering itself, so the decisive tests are data availability and label quality. The descriptive statistics are internally consistent: the five stutter-type counts sum to 280, matching Table IV's total events, and demographic counts sum to 28. The cross-dataset classification is framed as validation rather than a state-of-the-art result, so modest F1-scores or the unusual balancing rule in Eq. (1) are not, by themselves, fatal. The two candidate weaknesses are the missing download mechanism and the unvalidated annotation process; the reader's weakest assumption identified exactly this pair, and I agree. Accessibility is the more immediate gate because without a download no independent re-annotation is possible, but reliability is equally load-bearing for the word-level claim. The proposed concrete test—release the data and run a blinded two-annotator reliability study on a sample—directly settles whether either condition actually holds. The reader's CONDITIONAL verdict is appropriate: the preprint describes a plausible and needed resource, but the central claim is conditional on materials and evidence that are not yet provided. Hence no verdict adjustment is needed.","tokens_in":5603,"tokens_out":8514,"duration_ms":84164,"concrete_test":"Open the project website and follow any links to verify that a downloadable dataset archive exists with a license, README, and annotation guidelines. If it does, take a random 20% sample of the 280 word-level stutter events and have two independent annotators, blind to the original labels, re-annotate the clips using the published guidelines; report per-class Cohen's kappa or Krippendorff's alpha. If no download link exists, or if per-class agreement is below 0.6 for any stutter type, the open-access and word-level annotation claims are not supported; if the materials are released and agreement is at least 0.6 for all types, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Boli is an open-access, word-level annotated stuttered speech dataset for Indian languages. For this claim to hold, two conditions are necessary: (i) the data are actually released in a usable form, and (ii) the manual word-level labels are trustworthy. Section II describes the collection website (https://project-boli.vercel.app/) but gives no repository URL, DOI, license, checksum, or download procedure, so a reader cannot verify that the dataset exists or access it. If no data can be downloaded, the contribution collapses regardless of label quality. Conditional on access, the paper also reports no annotation guidelines, no number of annotators, no adjudication procedure, and no inter-annotator agreement measure. Every downstream quantity—Table IV stuttering rates, the class distribution (SR=140, B=70, PR=41, WR=21, IN=8), Table V classification labels, and the ASR verification in Table VI—is derived from these unvalidated word-level labels. Section III further omits how continuous recordings were segmented into the '≈5 s' clips used in classification, so the technical validation cannot be reconstructed from the paper alone. This does not mean the labels are wrong; it means the central claim currently rests on an unverified manual annotation process.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Boli, a crowd-sourced multilingual dataset of read and spontaneous speech from people who stutter, together with demographic metadata, responses to a 25-item questionnaire, and word-level annotations of five stutter types. The authors report descriptive statistics on stuttering rates, a cross-dataset stutter-type classification experiment trained on Sep-28k and tested on English Boli clips, ASR word-error-rate comparisons, and qualitative summaries of the questionnaire. The stated goal is to provide an open-access resource for stuttered speech research, particularly for Indian languages.","tokens_in":5863,"tokens_out":6095,"duration_ms":52079,"significance":"If the dataset is made available as described, it would fill a genuine gap: there is currently no Indian-language stuttered speech dataset with word-level annotations and linked experiential questionnaire data. The independent cross-dataset evaluation design is a strength, since the classifiers are trained on external data and tested on Boli, avoiding circularity. The manual collection of both read and spontaneous speech from 28 speakers with stutter is a useful contribution, as are the reported patterns of stuttering rates across speech modes. However, the paper as written does not allow a reader to access the dataset, assess label reliability, or interpret the quantitative validation; these gaps make the contribution contingent.","major_comments":[{"comment":"The central claim that Boli is an open-access dataset is not verifiable. The manuscript names only the data-collection website (https://project-boli.vercel.app/) and provides no repository URL, DOI, license, checksum, or download procedure. Without a persistent data availability statement, a reader cannot confirm that the dataset exists or use it. Please add a stable repository link or DOI and a clear data-access section, and state the license under which the audio and annotations are released.","section":"Section II, data collection"},{"comment":"No annotation protocol is reported: the paper does not state the number of annotators, their training or screening, the annotation unit (word vs. syllable), the annotation tool, or the adjudication procedure. No inter-annotator agreement measure is provided. This matters because Table IV's stuttering rates, the class distribution (SR=140, B=70, PR=41, WR=21, IN=8), Table V's classification labels, and Table VI's ASR verification all depend on the unvalidated manual labels. Please report annotation guidelines and inter-annotator agreement on a representative subset.","section":"Section II, manual annotation"},{"comment":"The stutter-type classification evaluation is under-specified. The paper does not state the number of Boli test clips per class, how continuous recordings were segmented into the approximately 5-second clips, or how the English-only subset was extracted from the 28 speakers. F1 scores are reported without confidence intervals; with class counts as small as IN=8, values such as 0.99 for interjection are not interpretable. Please report test-set size, per-class support, and bootstrap or equivalent confidence intervals, along with an explicit segmentation rule.","section":"Section III, Table V"},{"comment":"The hybrid-sampling procedure is not described precisely enough to reproduce. Equation (1) defines only an average of N1 and N2, and the surrounding text does not specify how many samples are drawn from the minority and majority classes, whether replacement is used, or how the balanced dataset is constructed from the Sep-28k training set. Because the paper's main conclusion that random forest is the best classifier rests on the balanced-training comparison, please specify the exact resampling algorithm.","section":"Section III, Eq. (1)"},{"comment":"The questionnaire analysis is presented as qualitative summary statements, but the instrument itself is not included, the response distributions are not tabulated, and no statistical test or analysis method is described. For example, the claim that 'language plays a crucial role in stuttering' is not connected to any reported quantitative result. Please add the full questionnaire or a link to it, summary statistics for each item, and a description of how the qualitative findings in Figure 3 were derived.","section":"Section IV, questionnaire analysis"}],"minor_comments":[{"comment":"The cross-reference 'Table reftable:demographics' appears to be a LaTeX error; it should be Table III.","section":"Section II, Table III"},{"comment":"Table II lists the Boli duration as 2.5 hours, while the text states approximately 2.8 hours of audio; please reconcile these numbers.","section":"Table II vs. Section II"},{"comment":"The label 'PRBIN SR WR' in Figure 2 is unclear; it should list the five stutter types as PR, B, IN, SR, WR.","section":"Figure 2"},{"comment":"Table VI reports WER but does not specify the ASR model versions, whether reference transcripts are the manual word-level annotations, how punctuation and casing were handled, or how the concatenated-speech setting was aligned with the manual annotations.","section":"Section III, ASR evaluation"},{"comment":"The paper does not report ethics approval, informed consent, or anonymization procedures for the human participants; a dataset paper of this type should include an ethics and consent statement.","section":"General"},{"comment":"The abstract promises 'severity assessment of stuttering events,' and the index terms include 'Intelligibility assessment,' but the manuscript does not actually model severity or intelligibility; it only reports self-reported severity in Table III. Please align the claims with the content.","section":"Abstract and index terms"}],"recommendation":"major_revision","confidential_remarks":"The main risk is accessibility: if the repository or DOI cannot be provided, the paper's contribution cannot be evaluated. In that case rejection would be appropriate. Otherwise, with annotation guidelines and a proper validation appendix, the resource could be a useful contribution to the stuttered-speech community. I would also encourage the authors to deposit the annotation protocol and a data statement with the final version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The Boli dataset is a real gap-filler: the first resource combining multiple Indian languages, read and spontaneous speech, word-level annotations of five stutter types, and questionnaire data from people who stutter. The scope is modest (28 speakers, ~2.8 hours), but the combination is new, and the authors are honest that they are not aiming for state-of-the-art classification. The descriptive statistics about read versus spontaneous stuttering rates and the questionnaire findings are useful for the community.\n\nThe paper's strengths: the collection procedure is explained, the annotation types follow the Sep-28k conventions, and the cross-dataset evaluation (training on Sep-28k, testing on Boli) is a sensible way to get some signal despite the small size. The authors also include the questionnaire summaries, which is unusual and valuable.\n\nThe soft spots are real and they matter. First, the paper claims the dataset is open access, but no repository, DOI, license, or download procedure is provided anywhere. The project website is mentioned, but it does not obviously host the data, so as written a reader cannot check whether the dataset exists at all. Second, the word-level annotations have no associated protocol: no annotation guidelines, no number of annotators, no adjudication process, and no inter-annotator agreement. Every downstream result, including Table IV stuttering rates and the classification labels in Table V, rests on labels that cannot be verified. Third, the segmentation into ~5-second clips used for classification is not described, so the technical validation cannot be reconstructed. The F1-scores are also reported on a tiny test set without confidence intervals, so they are illustrative at best.\n\nThese issues are fixable, not fatal. If the authors supply a direct data link with a clear license, release annotation guidelines, report inter-annotator agreement, and describe the segmentation, the contribution becomes solid. As is, the paper's central claims are conditional on materials that are absent.\n\nThis paper deserves peer review because the dataset fills a genuine gap and the authors have put in real effort. I would send it out with a strong request to make the data accessible and to document label reliability. The reader's conditional verdict is fair, and the stress-test concern about verifiability is accurate and should be the primary request for revision.\n\nFor my own work, I would not cite this yet because I cannot access the data, but I would revisit it once the dataset is released. I would bring it to a reading group if the discussion is about dataset construction pitfalls in speech pathology resources.","headline":"A genuinely new stuttered-speech resource for Indian languages, but the paper currently makes it impossible to verify the dataset exists or that the labels are trustworthy.","tokens_in":6333,"tokens_out":1562,"would_cite":false,"duration_ms":16815,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Boli dataset provides word-level, timestamped labels of five stutter types in read and spontaneous Indian-language speech, along with questionnaire responses from people who stutter.","keywords":["stuttered speech dataset","Indian languages","word-level annotation","stutter type classification","read speech","spontaneous speech","questionnaire","stuttering event detection"],"falsifier":"Try to download the audio from the project site given in the paper and have two independent annotators relabel a random sample of the same files; if the files are not downloadable or the annotators disagree substantially, the dataset's central promise of open access and reliable word-level labels is not yet met.","tokens_in":5440,"feed_emoji":"🗣️","tokens_out":7426,"duration_ms":65604,"temperature":0.7,"pith_summary":"This paper introduces Boli, a dataset of stuttered speech from people who stutter in India, designed to support both scientific study and speech technology. The central claim is that Boli fills a gap left by existing stutter datasets, which are mostly English, file-level, or non-Indian, by offering an open-access multilingual Indian-language resource. The dataset pairs read speech (Rainbow Passage) and spontaneous speech (image descriptions) with word-level annotations of five stutter types — blocks, prolongations, interjections, sound repetitions, and word repetitions — and adds questionnaire data on how stuttering affects daily life. If the data are made available as claimed, Boli would give researchers a benchmark for stutter detection and automatic speech recognition evaluation in Indian languages that did not previously exist.","feed_headline":"Boli dataset brings word-level stutter labels to Indian languages","feed_subtitle":"28 speakers, five stutter types, read and spontaneous audio, plus self-reported experience data.","key_machinery":"The load-bearing object is the Boli dataset itself: paired read and spontaneous recordings, word-level stutter-type labels with timestamps, and anonymized questionnaire responses. The word-level annotation scheme is what lets the dataset support event detection rather than merely file-level classification, and the questionnaire connects acoustic phenomena to lived experience. For validation, the paper uses a standard pipeline — mel-frequency cepstral coefficients as audio features, hybrid class balancing, and a five-way classifier trained on an existing English dataset and tested on Boli — to show that the audio admits stutter-type classification.","core_discovery":"The paper's discovery is the Boli dataset: audio and metadata from 28 people who stutter, selected after screening 67 volunteers, totaling about 2.5 to 2.8 hours of speech. Each participant read part of the Rainbow Passage and described an image, in English and in their mother tongue, spanning Hindi, Telugu, Bengali, Marathi, and Assamese. The recordings carry word-level, time-stamped annotations of five stutter types: blocks, prolongations, interjections, sound repetitions, and word repetitions. The paper also reports that stuttering was less frequent in spontaneous speech (2.02 events per minute) than in read speech (6.76 events per minute), and it validates the audio through a five-class stutter-type classification experiment in which a balanced random forest trained on an existing English dataset reaches an average F1 of 0.87 on Boli's English utterances. It further reports word-error-rate baselines from two automatic speech recognition systems on stuttered Indian English, with Whisper outperforming Wav2Vec2.0, and it summarizes the 25-question responses from 67 participants on self-reported triggers, speech-therapy experience, and social effects of stuttering.","pith_inferences":["A natural next step is to release annotation guidelines and conduct an inter-annotator agreement study; without those, the word-level labels cannot be independently verified.","The current class imbalance (140 sound repetitions vs 8 interjections) means stutter-type models need explicit balancing, and expansion should oversample rare types and severe cases.","The read/spontaneous paired design transfers naturally to many under-resourced languages and could be paired with a standardized severity instrument to make experiential data comparable across studies.","A first practical check is to verify that the stated project website actually serves the audio, since the paper gives no download link."],"forward_implications":["Word-level labels with timestamps create a training target for stutter-event detection in Indian languages, not just file-level classification.","The paired read/spontaneous design shows stuttering rate drops in spontaneous speech (2.02 vs 6.76 events per minute), so Boli can support studies of speaking-mode effects.","The questionnaire responses tie acoustic events to self-reported triggers and coping strategies, enabling experience-aware speech analysis.","Cross-dataset classification results (average F1 0.87 with a balanced random forest) suggest models trained on English stutter data transfer to Indian English, so Boli could serve as a target benchmark.","ASR word-error-rate baselines for stuttered Indian English give a concrete starting point for improving speech recognition for people who stutter."],"supporting_citations":[{"why":"Provides the longstanding English clinical stutter dataset whose format and limitations motivate Boli's design.","marker":"[6]"},{"why":"Supplies the training data for the cross-dataset stutter-type classification experiments used to validate Boli's audio.","marker":"[10]"},{"why":"Offers a comparable interview-based stutter dataset with similar annotations, used in the comparison table.","marker":"[11]"},{"why":"Represents recent non-English stutter data collection, showing the gap Boli addresses for Indian languages.","marker":"[12]"},{"why":"Closest prior Indian stutter dataset, limited to syllable-level Kannada; Boli extends to multiple languages with word-level labels.","marker":"[13]"},{"why":"One of the two ASR systems evaluated on Boli to measure word error rate on stuttered Indian English.","marker":"[14]"},{"why":"The other ASR system evaluated on Boli, reporting lower word error rate.","marker":"[15]"}],"fun_headline_variants":["Boli: word-level stutter labels for 5 Indian languages","Boli dataset: 28 speakers, 5 stutter types, open access","Boli: stutter events drop in spontaneous vs read speech","Boli: Indian stutter dataset beats ASR baselines","Boli: 2.8 hours of labeled stuttered speech from India"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the audio is actually available for download and that the manual word-level stutter labels are consistent enough to be reused, because the paper provides no repository link, annotation guideline, or inter-annotator agreement figures.","fun_headline_variants_meta":{"raw":{"variants":["Boli: word-level stutter labels for 5 Indian languages","Boli dataset: 28 speakers, 5 stutter types, open access","Boli: stutter events drop in spontaneous vs read speech","Boli: Indian stutter dataset beats ASR baselines","Boli: 2.8 hours of labeled stuttered speech from India"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1273,"prompt_tokens":955,"completion_tokens":318,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":224}},"tokens_in":571,"tokens_out":318,"duration_ms":3358,"temperature":1.0,"reasoning_tokens":224,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:48:59.874823+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Try to download the audio from the project site given in the paper and have two independent annotators relabel a random sample of the same files; if the files are not downloadable or the annotators disagree substantially, the dataset's central promise of open access and reliable word-level labels is not yet met.","supporting_citations":[{"cited_title":"”The university college london archive of stuttered speech (uclass)”","cited_arxiv_id":null,"evidence_quote":"Provides the longstanding English clinical stutter dataset whose format and limitations motivate Boli's design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the training data for the cross-dataset stutter-type classification experiments used to validate Boli's audio."},{"cited_title":"B., & MacWhinney, B","cited_arxiv_id":null,"evidence_quote":"Offers a comparable interview-based stutter dataset with similar annotations, used in the comparison table."},{"cited_title":"S., Mahesh, S., Barche, P., Mirishkar, S","cited_arxiv_id":null,"evidence_quote":"Closest prior Indian stutter dataset, limited to syllable-level Kannada; Boli extends to multiple languages with word-level labels."},{"cited_title":"”wav2vec 2.0: A framework for self-supervised learning of speech representations”","cited_arxiv_id":null,"evidence_quote":"One of the two ASR systems evaluated on Boli to measure word error rate on stuttered Indian English."},{"cited_title":"W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I","cited_arxiv_id":null,"evidence_quote":"The other ASR system evaluated on Boli, reporting lower word error rate."}],"review_version":1}