{"id":"381a385e-066b-43ba-a38b-2cca8c2fd305","arxiv_id":"2411.17149","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A new Indian English stutter corpus and an augmented cepstral feature set are reported to classify typical vs atypical disfluencies with 85% F1, though the evaluation is hampered by missing split details and cherry-picked results.","lead":"Researchers built a new speech dataset of stuttered Indian English and a classifier that distinguishes ordinary disfluencies, like hesitations and repetitions, from stuttering-related ones. The work targets voice assistants that cut off people who stutter, and early stutter detection in children.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported 85.01% F1 appears to be a per-class best-configuration selection, not a single-model result, and no held-out split is described.","rationale":"The reader's weakest assumption correctly identifies the missing train/validation/test split and speaker-disjoint protocol as a fundamental issue. My stress-test adds a more specific and demonstrable problem: the reported 85.01% average F1 does not correspond to any single configuration in Table IV; it is consistent with a per-class best selection across different SDC parameter settings. This is an internal inconsistency that strengthens the reader's rejection, but it is a different angle from the missing-split concern. The reader's rationale did mention 'best-per-class selections from different hyperparameter configurations,' so there is partial overlap. I credit the paper for the dataset collection and annotation effort, but the central comparative claim is not supported without a valid, single-model evaluation. The verdict remains REJECT; no adjustment is needed because the concern reinforces the existing rejection rather than moving it.","tokens_in":8403,"tokens_out":3262,"duration_ms":28072,"concrete_test":"Re-run the classification with a single SDC configuration (e.g., 13-2-3-6) selected only on a validation set, using a speaker-disjoint 70/15/15 split, and report macro-F1 for the three disfluency classes. If the single-configuration macro-F1 falls below 85.01% or below the per-class-best average, the headline claim is an artifact of selection. As a secondary check, compute the macro-F1 of the per-class best configuration values from Table IV and confirm whether it matches the stated 85.01%.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that PE-ZTWCC+SDC achieves 85.01% average F1 and 'outperforms traditional features' depends entirely on how the evaluation was conducted. Two concrete problems undermine it. First, Table IV shows that the best F1 for each class comes from different SDC configurations: 13-2-3-6 gives 86.98% for repetitions, 13-2-3-5 gives 85.45% for filled pauses, and 13-2-3-7 gives 82.78% for prolongations. The average of these three per-class best values is 85.07%, essentially the abstract's 85.01%, yet no single configuration in Table IV averages above 84% (e.g., 13-2-3-6 averages 84.00%). The headline result is therefore not attributable to one fixed model; it is a best-of-configurations composite chosen using the test labels, which is an optimistic ceiling rather than a generalization estimate. Second, Section V reports results without specifying any train/validation/test split or speaker-disjoint protocol. Since clips are derived from only 30 PWS speakers plus IIITH-IED-E speakers, it is unknown whether the same speaker appears in both training and test partitions or whether hyperparameter selection used validation or test labels. Without a single fixed configuration evaluated on a held-out speaker-disjoint test set, the comparative claim against MFCC, PLP, and SFCC baselines is not supported as stated. The dataset contribution may be useful, but the performance claim needs a rigorous evaluation protocol before it can be accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces IIITH-TISA, described as the first Indian English stammer corpus (10 hours, 30 PWS speakers, 3,251 three-second clips), and an extension of the IIITH-IED dataset (IIITH-IED-E) with additional typical disfluency annotations. The authors propose Perceptually Enhanced Zero-Time Windowed Cepstral Coefficients (PE-ZTWCC) combined with Shifted Delta Cepstra (SDC) as input to a shallow Time Delay Neural Network (TDNN) for binary classification of typical versus atypical disfluencies in three categories (repetition, filled pause, prolongation). The central claim is an average F1 score of 85.01%, outperforming traditional features such as MFCC, PLP, and SFCC. The paper also details the corpus creation, annotation procedure, feature extraction equations, and a small hyperparameter study over SDC N-d-p-K parameters.","tokens_in":1427,"tokens_out":1877,"duration_ms":57309,"significance":"If the empirical claim were validly supported, the paper would make two useful contributions: a new resource for stuttering research in Indian English, and a feature representation that captures both spectral and temporal context in a low-data regime. The dataset collection and annotation process, including SEP-28k-aligned labeling and involvement of a speech-language pathologist, is a genuine asset for the community. The feature extraction is grounded in established signal processing and is described with enough detail to be reproduced. However, the performance claim is currently not supported by the reported experiments, so the significance of the method contribution depends on a rigorous re-evaluation.","major_comments":[{"comment":"The reported average F1 of 85.01% is not achieved by any single system. In Table IV, the best per-class scores come from different SDC configurations: 86.98% for repetitions with 13-2-3-6, 85.45% for filled pauses with 13-2-3-5, and 82.78% for prolongations with 13-2-3-7. The mean of these three per-class bests is 85.07%, essentially the abstract's 85.01%. No single configuration in Table IV has an average above 84% (e.g., 13-2-3-6 averages 84.00%). Thus the headline number is a best-of-configurations composite selected using the test labels, not a generalization estimate of one trained model. The paper must report a fixed SDC configuration and a single model trained under a proper protocol, with its average F1.","section":"Abstract, Table IV in Section V"},{"comment":"No train/validation/test split or speaker-disjoint protocol is described anywhere in the paper. The evaluation thus cannot rule out speaker overlap between training and test partitions, nor can it establish whether the SDC parameters (Section IV.A) and TDNN architecture (Section IV.B) were tuned on the test set. Because the paper says the parameters were varied \"in steps and top results are tabulated,\" the reported F1 values are at risk of being optimistic ceiling estimates from test-set peeking. The authors must specify the split, number of speakers in each partition, and how hyperparameter selection was performed (e.g., on a validation set), and they should report results with error bars or confidence intervals.","section":"Section V (Results and Discussions)"},{"comment":"The comparison against baseline features is presented without any statistical significance testing or repeated-run variance. For example, the gap between PE-ZTWCC+SDC and PLP+SDC in Table III is about 1-3 percentage points depending on the class, which may not be significant given the limited number of speakers (30 PWS plus typical speakers). Without significance tests or confidence intervals, the claim that PE-ZTWCC+SDC \"outperforms traditional features\" is not established.","section":"Tables II and III, Section V"},{"comment":"The paper does not state how many clips from each speaker are used, whether the model is trained per disfluency category (binary typical vs atypical) or jointly, or how the datasets are combined (e.g., are IIITH-TISA and IIITH-IED-E simply pooled?). If a single binary classifier is trained across all categories, the label space and loss should be described; if separate classifiers are trained per category, that should be stated. This ambiguity affects the interpretation of all reported F1 scores.","section":"Sections II.A and V"}],"minor_comments":[{"comment":"The text contains a typo: \"eN-d-p-K\" should be \"N-d-p-K\".","section":"Section IV.C"},{"comment":"Equation (1) is missing parentheses around the sine term and uses \"f or\" for \"for\"; Equation (2) is missing a closing parenthesis. Please correct the formatting.","section":"Equations (1) and (2), Section III.A"},{"comment":"The notation in Eq. (5) is ambiguous: the exponent 1/5 should be applied to the whole expression X_WE[n,k], e.g., X_WEP[n,k] = (X_WE[n,k])^(1/5), and the sentence describing the inverse transform should clarify that an inverse DFT, not a continuous inverse Fourier transform, is used.","section":"Section III.A, Eq. (5)"},{"comment":"The table title says \"across three datasets,\" but IIITH-IED-E is an extension of IIITH-IED and includes the IIITH-IED counts; please clarify whether the counts are cumulative or disjoint.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is real and potentially valuable, but the evaluation protocol is sufficiently flawed that the central performance claim cannot be accepted. If the authors are able to disclose a proper speaker-disjoint split, report results for a single fixed configuration, and add significance testing, the paper could become suitable for publication. If not, the manuscript may need to be reframed as a purely descriptive dataset paper. The paper's scope and the journal's standards should guide this decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The IIITH-TISA corpus is the real contribution here. A ten-hour Indian English stammer dataset, aligned with SEP28k conventions, is something the field can use, and the extension of IIITH-IED adds useful typical-disfluency data. The annotation work looks careful, and the decision to publish the corpus fills an actual language gap. Credit where due: that alone makes the paper worth a look.\n\nThe feature work is more incremental. PE-ZTWCC is ZTWCC plus standard perceptual processing (Mel warping, equal-loudness, power-law compression), and SDC is well-trodden. The TDNN is a standard shallow architecture. None of that is a problem, but it is not a breakthrough.\n\nThe abstract says the method achieves 85.01% average F1 and outperforms traditional features. That claim does not hold up against the paper's own tables. Table IV lists F1 scores for PE-ZTWCC+SDC under different N-d-p-K settings. The repetition best (86.98%) comes from 13-2-3-6, the filled-pause best (85.45%) from 13-2-3-5, and the prolongation best (82.78%) from 13-2-3-7. The average of those three per-class bests is 85.07%, essentially the abstract's 85.01%. But no single configuration in Table IV averages above 84% — the best average is 13-2-3-6 at 84.00%. So the headline number is a composite of different models, each selected per class on what amounts to the test labels. That is an optimistic ceiling, not a generalization estimate.\n\nJust as important, the paper never describes a train/validation/test split or any speaker-disjoint protocol. Clips come from 30 PWS speakers plus the IIITH-IED-E speakers. Without knowing whether the same speaker appears in both training and test partitions, or how hyperparameters were selected, the comparison against MFCC, PLP, and SFCC baselines is not interpretable. There are also no error bars or significance tests. The evaluation section reads like a sketch, not a measurement.\n\nThere is a smaller overstatement: the introduction says no published work has used ML for this classification problem. Stutter-event detection exists (SEP28k, FluentNet), and the authors cite those. The specific typical-vs-atypical framing may be underexplored, but that is a narrower claim than the one they actually make.\n\nWho gets value from this paper? Someone building or benchmarking Indian English disfluency corpora, and anyone who wants a cautionary example of how per-class best selection creates illusory gains. The corpus is worth engaging with; the performance claim is not, in its current form.\n\nRecommendation: I would send this to peer review rather than desk-reject, because the dataset contribution deserves scrutiny and the evaluation flaw is fixable. The authors need to specify the split, report a single fixed configuration with error bars, and stop reporting composites of per-class bests. With that, the paper could be a solid resource.","headline":"The IIITH-TISA corpus is a genuine new resource, but the 85% F1 headline is a per-class best-of-configurations average, not the result of any single system.","tokens_in":9246,"tokens_out":1735,"would_cite":false,"duration_ms":17268,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new feature representation combining perceptually enhanced zero-time windowed cepstral coefficients with shifted delta cepstra classifies typical versus atypical speech disfluencies in Indian English with an average F1 of 85.01%.","keywords":["speech disfluency","stuttering","zero-time windowing","cepstral features","shifted delta cepstra","time-delay neural network","Indian English","voice assistants"],"falsifier":"A strict speaker-disjoint evaluation (e.g., leave-speakers-out cross-validation on the IIITH-TISA and extended IIITH-IED corpora) that reports the same 85.01% average F1 would confirm the claim; a drop below the strong baselines (such as PLP+SDC or ZTWCC+SDC) would refute it. Alternatively, collecting a new cohort of 30 PWS and 30 controls in Indian English and running the pretrained classifier would show whether the features generalize beyond the original recordings.","tokens_in":37,"feed_emoji":"🗣️","tokens_out":7367,"duration_ms":112440,"temperature":0.7,"pith_summary":"The paper sets out to show that a compact set of handcrafted features can tell typical speech disfluencies (pauses, repetitions, and prolongations that occur in fluent speech) apart from atypical ones that mark stuttering. To do this it introduces a new Indian English stammer corpus and extends an existing typical-disfluency corpus, then combines perceptually enhanced zero-time windowed cepstral coefficients with shifted delta cepstra as input to a shallow time-delay neural network. The reported result is an average F1 score of 85.01% across repetition, filled-pause, and prolongation classes, which the paper claims outperforms traditional features such as MFCCs, PLP, and SFCC. If this holds, a lightweight classifier could be inserted after disfluency detection in voice assistants to avoid premature cutoffs for people who stutter, and could support early screening for stuttering in children.","feed_headline":"New audio features sort stutter disfluencies at 85% F1.","feed_subtitle":"A compact cepstral feature set and shallow network tell typical from stuttering disfluencies in Indian English.","key_machinery":"The load-bearing object is the PE-ZTWCC+SDC feature vector: zero-time windowing estimates a spectrum from only a few samples with a heavily decaying window, and the numerator of group delay plus double differencing and Hilbert envelope yields the ZTW spectrum. Perceptual enhancement (Mel warping, equal-loudness contour, power-law compression) shapes that spectrum like human hearing, and SDC computes blockwise delta cepstra over N-d-p-K parameters to encode local and longer-range temporal movement. The shallow TDNN with two dilated convolutional layers then maps these 104-dimensional frame features to a binary typical-versus-atypical decision per disfluency event.","core_discovery":"On its own terms, the paper claims that the specific combination of Perceptually Enhanced Zero-Time Windowed Cepstral Coefficients (PE-ZTWCC) with Shifted $\\Delta$ Cepstra (SDC), fed through a shallow two-layer dilated time-delay neural network, is the first feature representation to reliably separate typical from atypical disfluencies in Indian English. The PE-ZTWCC features exploit zero-time windowing's high temporal resolution and group-delay spectral estimation to capture laryngeal tension cues, then apply Mel warping, equal-loudness pre-emphasis, and a one-fifth power-law nonlinearity to mimic human auditory perception. Adding SDC extends the cepstra across multiple frames, and the shallow TDNN with tuned dilation rates captures wider temporal context without overfitting the small dataset. The paper reports per-class F1 scores of 86.98% for repetitions, 85.45% for filled pauses, and 82.78% for prolongations, with an average of 85.01%, and states that these numbers beat traditional feature sets in both plain and SDC-augmented forms.","pith_inferences":["The same PE-ZTWCC+SDC pipeline could be tested on other atypical-disfluency types (e.g., blocks, interjections) and on child speech, since the features target laryngeal tension cues that are not language-specific.","A strict speaker-disjoint evaluation protocol would likely change the reported numbers; the paper's missing split description makes the 85.01% F1 an upper-bound estimate rather than a guaranteed generalization.","The N-d-p-K parameter search in Table IV suggests per-class optimal contexts differ (e.g., prolongations need longer context), so a multi-expert or adaptive-context system might yield further gains beyond a single global feature configuration.","If this feature set works at short clip lengths, it could be integrated into streaming endpointers with low latency, but the paper does not test streaming or noisy conditions."],"forward_implications":["Voice assistants could use such a classifier to set separate endpointing thresholds for speakers who stutter, reducing premature cutoffs without slowing responses for others.","Early stutter screening in children becomes feasible from short 3-second clips, if the typical/atypical distinction transfers to developmental speech.","The introduced corpus and the extended typical-disfluency corpus give researchers a public benchmark for Indian English disfluency classification.","The strong performance of handcrafted temporal context features suggests that shallow, parametric models can rival deeper networks on small pathological-speech datasets.","Feature engineering choices (base cepstrum, perceptual warping, SDC parameters) have a measurable effect, so future systems can tune these rather than relying solely on model capacity."],"supporting_citations":[{"why":"Supplies the annotation protocol and 3-second clip standardization used to curate the new stammer corpus.","marker":"[9]"},{"why":"Provides the base Indian English typical-disfluency dataset that this paper extends with additional speakers and annotations.","marker":"[12]"},{"why":"Establishes zero-time windowing cepstral coefficients, the spectral representation the proposed features build on.","marker":"[16]"},{"why":"Shows ZTWCC outperforms SFCC in stutter detection, the prior result that motivates choosing ZTWCC as the base feature.","marker":"[17]"},{"why":"Introduces the time-delay neural network formulation for modeling long temporal contexts that the shallow classifier uses.","marker":"[23]"},{"why":"Defines shifted delta cepstra and the N-d-p-K parameterization that extends the static features with temporal context.","marker":"[24]"}],"fun_headline_variants":["New Indian English corpus tells stutter disfluencies from typical at 85% F1","PE-ZTWCC+SDC feature combo sorts speech disfluencies at 85% F1","First stammer corpus for Indian English classifies disfluency at 85% F1","Temporal context features hit 85% F1 separating typical vs atypical disfluency"],"cache_read_input_tokens":11392,"weakest_assumption_plain":"The reported F1 scores assume the models were evaluated on held-out clips from speakers not used to tune the SDC parameters and network architecture, but the paper never specifies a train/validation/test split or speaker-disjoint protocol.","fun_headline_variants_meta":{"raw":{"variants":["New Indian English corpus tells stutter disfluencies from typical at 85% F1","PE-ZTWCC+SDC feature combo sorts speech disfluencies at 85% F1","First stammer corpus for Indian English classifies disfluency at 85% F1","Temporal context features hit 85% F1 separating typical vs atypical disfluency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000245,"raw_usage":{"total_tokens":1556,"prompt_tokens":988,"completion_tokens":568,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":472}},"tokens_in":604,"tokens_out":568,"duration_ms":5233,"temperature":1.0,"reasoning_tokens":472,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:27:33.146288+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A strict speaker-disjoint evaluation (e.g., leave-speakers-out cross-validation on the IIITH-TISA and extended IIITH-IED corpora) that reports the same 85.01% average F1 would confirm the claim; a drop below the strong baselines (such as PLP+SDC or ZTWCC+SDC) would refute it. Alternatively, collecting a new cohort of 30 PWS and 30 controls in Indian English and running the pretrained classifier would show whether the features generalize beyond the original recordings.","supporting_citations":[{"cited_title":"Sep-28k: A dataset for stuttering event detection from podcasts with people who stutter,","cited_arxiv_id":null,"evidence_quote":"Supplies the annotation protocol and 3-second clip standardization used to curate the new stammer corpus."},{"cited_title":"Towards a database for detection of multiple speech disfluencies in indian english,","cited_arxiv_id":null,"evidence_quote":"Provides the base Indian English typical-disfluency dataset that this paper extends with additional speakers and annotations."},{"cited_title":"Breathy to tense voice discrim- ination using zero-time windowing cepstral coefficients (ztwccs)","cited_arxiv_id":null,"evidence_quote":"Establishes zero-time windowing cepstral coefficients, the spectral representation the proposed features build on."},{"cited_title":"Enhancing stutter detection in speech using zero time windowing cepstral coefficients and phase information,","cited_arxiv_id":null,"evidence_quote":"Shows ZTWCC outperforms SFCC in stutter detection, the prior result that motivates choosing ZTWCC as the base feature."},{"cited_title":"A time delay neural net- work architecture for efficient modeling of long temporal contexts","cited_arxiv_id":null,"evidence_quote":"Introduces the time-delay neural network formulation for modeling long temporal contexts that the shallow classifier uses."},{"cited_title":"Approaches to language identification using gaussian mixture models and shifted delta cepstral features","cited_arxiv_id":null,"evidence_quote":"Defines shifted delta cepstra and the N-d-p-K parameterization that extends the static features with temporal context."}],"review_version":1}