{"id":"3a250bf6-2b43-40ea-9439-57dc95ba6a31","arxiv_id":"2506.00644","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new multimodal expert-annotated stuttering dataset with disfluency types, secondary behaviors, tension scores, and a consensus gold standard test set.","lead":"This paper adds expert clinician annotations to FluencyBank videos, labeling stuttering types, secondary behaviors, and tension levels. It also releases a consensus-based test set and audio, video, and multimodal baselines for automatic stuttering severity assessment.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gold standard labels are produced by the same annotators they evaluate, and the test set is deliberately selected from highest-disagreement files; without independent validation, the annotator and model F1 scores in Tables 3–4 do not establish clinically reliable performance.","rationale":"The reader's weakest assumption correctly identifies the circularity of the consensus gold standard and the unrepresentative selection of the test set. This is the most load-bearing concern because every quantitative result that supports the dataset's value—annotator F1 (Table 3), aggregation method comparisons (Table 3), and baseline model F1 (Table 4)—is evaluated against this self-derived, non-independent reference on a deliberately hard subset. If the consensus is biased, all benchmark numbers lose external meaning; if the test set is unrepresentative, the difficulty estimates are misleading for real-world deployment. The paper is otherwise transparent: it documents the annotation process, reports low tension agreement, and acknowledges challenges. The concern is not about dishonesty but about the absence of independent validation, which is fixable. A conditional acceptance with a requirement for external validation (or at least public release with clear caveats) remains the appropriate verdict; no new concern changes that. I agree with the reader's assessment and would keep the verdict unchanged. The proposed concrete test—independent SLP annotation of the test set—directly settles whether the gold labels are clinically valid and whether the reported F1 scores should be trusted.","tokens_in":8340,"tokens_out":3403,"duration_ms":36693,"concrete_test":"Recruit two or three independent Speech-Language Pathologists with comparable clinical experience who were not involved in the original annotation. Have them annotate the full 8-file test set (4 reading + 4 interview) using the same annotation manual and ELAN setup, without access to the consensus gold labels. Compute per-class F1 and Krippendorff's alpha between the independent annotators' (majority) labels and the published gold labels. If agreement is high for primary disfluency types (e.g., F1 > 0.8) and secondary behaviors, the gold standard is externally valid; if agreement is markedly lower than the original annotator F1 scores in Table 3, the gold labels are idiosyncratic and the benchmark is not clinically validated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the authors provide a clinically aligned, consensus-based gold standard for stuttering severity assessment. Section 4.3 describes how the gold labels were created: all three annotators discussed and resolved disagreements until consensus. Table 3 then reports each annotator's F1 against these gold labels. This is not an external evaluation: every annotator co-produced the reference they are scored against. Shared training, shared experience, or group dynamics during the consensus meetings could systematically bias the gold labels, and the reported F1 would not reveal it. The problem is compounded by test-set selection: the eight files were chosen because the two most experienced annotators had the most disagreements in classifying disfluency types. The test set is therefore a hard, non-representative subset, so the F1 scores in Tables 3 and 4 cannot be read as typical accuracy on FluencyBank or in clinical practice. Additionally, the paper reports very low agreement on tension (Krippendorff's alpha = 0.18, Section 4.2), yet tension is one of the three annotation dimensions claimed to align with clinical practice; no baseline or validation shows this dimension is usable. The assertion that the scheme 'aligns with clinical practice' is supported only by the annotators' qualifications and use of known taxonomies, not by any external criterion. If the consensus process is biased or the test set is unrepresentative, the benchmark numbers are inflated in unknown directions and the resource cannot serve as a reliable evaluation set.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multi-dimensional annotation scheme for stuttering severity assessment on FluencyBank, in which three expert speech-language pathologists annotate stuttering moments with disfluency type (LBDL), secondary behaviors (SSI-4), and a tension scale. The authors report inter-annotator agreement, construct a consensus-based 'gold standard' test set from files with the most disagreements, evaluate individual annotators and aggregation methods against that gold, and provide audio, video, and multimodal baselines. The stated goal is to make the annotations publicly available to enable clinically aligned automatic stuttering assessment.","tokens_in":8642,"tokens_out":6713,"duration_ms":62018,"significance":"If the dataset and the consensus test set are valid and released, this would be a valuable resource. Existing stuttering annotation efforts either lack visual dimensions, rely on non-expert annotators, or provide only primary disfluency types; a clinically informed multi-dimensional annotation effort involving expert SLPs would be a useful contribution. The paper is transparent about its methodology, uses established taxonomies, and reports detailed annotation processes and challenges. However, the central validity of the gold standard is questionable because it is derived from the same annotators it is used to evaluate, and the tension dimension shows very low agreement despite being a core part of the claimed multi-dimensional scheme. These issues need to be addressed before the resource can be used as a reliable benchmark.","major_comments":[{"comment":"The gold standard used to evaluate annotators in Table 3 is produced by the same three annotators whose individual labels are scored against it, and the test set is deliberately selected from files where the two most experienced annotators had the most disagreements. For any instance where all three annotators initially agreed, an annotator's label is identical to the gold label, and for disagreements the gold is the outcome of discussion among those same annotators. The reported F1 scores therefore measure consistency with the group's discussion process rather than independent clinical accuracy, and they are likely inflated by shared training, group dynamics, and selection effects. These scores cannot be interpreted as typical annotator performance on FluencyBank or in clinical practice. The authors should validate the gold standard against an external reference (e.g., a fourth expert not involved in the original annotation), report agreement with that reference, and discuss how the deliberate choice of high-disagreement files affects the test set's representativeness for model evaluation.","section":"4.3, Table 3"},{"comment":"Tension is presented as one of the three core annotation dimensions (Section 3.1, Table 1) and is part of the claim that the scheme aligns with clinical practice, yet the inter-annotator agreement for tension is very low (Krippendorff's alpha = 0.18, KS = 0.38, sigma = 0.34) in Section 4.2. No analysis, baseline, or validation is provided for the tension dimension, and it is absent from the test-set evaluation in Tables 3 and 4. Because a central contribution is a comprehensive multi-dimensional scheme, the unusably low agreement for this dimension substantially weakens the claim. The authors should either provide evidence that tension annotations are reliable after the consensus process (e.g., agreement on the gold test set for tension), refine the tension annotation protocol, or clearly delimit the contribution to the other two dimensions.","section":"3.1, 4.2"},{"comment":"The counts in Table 2 are internally inconsistent. The sum of the primary-type rows (SR 190 + ISR 143 + MUR 94 + P 93 + B 265 + None 25) is 810, not the reported total of 732. Additionally, the secondary-behavior columns do not sum to the row totals; for example, the SR row lists 23 + 114 + 38 + 1 + 53 = 229 secondary-behavior occurrences for 190 SR events. The table should clarify whether multiple secondary behaviors can be labeled per event, and the totals must be corrected, since this table is the quantitative description of the test set.","section":"Table 2"},{"comment":"The central artifact of the paper, the annotations themselves, is not actually available: the abstract states the annotations 'will be made publicly available' and the conclusion says they 'will be released,' but no data link is provided. For a dataset contribution, reviewers and readers need access to the annotations and the annotation manual to verify the scheme and reproduce the analysis. Please provide an anonymous download link (or a clear availability statement with the actual repository) as part of the manuscript or supplementary materials.","section":"1, 6"}],"minor_comments":[{"comment":"The baseline F1 scores in Table 4 are reported without standard deviations or statistical significance tests. Given the small differences between models (e.g., the Any-class F1 values ranging from 0.90 to 0.95) and the use of overlapping 5-second segments, the claim that the multi-modal approach is best for the Any class and generally better for secondary behaviors is not statistically supported. Please report means and variances over multiple runs or seeds.","section":"5, Table 4"},{"comment":"The text states that 'we see higher IAA scores in the interview section which was annotated after the annotation and discussion of the reading section,' but per-section agreement scores are not reported anywhere. Please provide the relevant IAA values for the reading and interview sections separately to substantiate this observation.","section":"4.2"},{"comment":"The row label 'None' for primary disfluency type is not defined in Section 3.1. Please clarify what a 'None' primary type denotes (e.g., a secondary behavior observed without a stuttering moment) and how such cases are handled in the span-based annotation scheme.","section":"4.3, Table 2"},{"comment":"The annotation manual is described in detail but is not provided. Including the manual as supplementary material would greatly improve reproducibility and allow other clinicians to apply the same scheme.","section":"3.1"},{"comment":"'Mutli-dimensional' is a typo for 'Multi-dimensional'.","section":"Table 1"},{"comment":"The term 'test set' is used for the consensus evaluation set of Section 4.3 and later for model evaluation in Section 5. Please clarify whether the baseline evaluations in Table 4 were performed on the consensus gold test set only, or on the full annotated dataset, and specify the train/dev/test split used.","section":"4.3, 5"}],"recommendation":"major_revision","confidential_remarks":"The core dataset contribution is promising, but the paper currently functions more as a protocol description because the annotations are not released. The low tension agreement and the self-derived gold standard are substantial validity concerns that the authors must address, and the Table 2 arithmetic errors suggest the quantitative description needs careful checking. If the authors provide the data, an external validation of the gold standard, and a usable tension protocol (or a justified de-emphasis of the tension dimension), the paper could become a solid resource for the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this one. First, it is a genuinely new resource: expert SLPs annotated FluencyBank videos for disfluency type, secondary behaviors, and tension, using LBDL, SSI-4, and the Boey scale. No prior stuttering dataset combines those dimensions, and none uses clinician annotators on audiovisual data. That part is real. Second, the numbers in Tables 3 and 4 should not be read as performance benchmarks. The gold labels were produced by the same three annotators who are then scored against them, and the test files were deliberately selected from the highest-disagreement cases. So the F1 scores measure group-conformity on a hard subset, not accuracy against an independent standard.\n\nWhat the paper does well: the annotation methodology is detailed and transparent, including the manual, ELAN setup, and boundary rules. The authors report the low tension agreement (alpha=0.18) honestly rather than hiding it. The 'reported challenges' section is genuinely useful for anyone designing a similar study. The baselines are basic but sufficient to show the task is hard.\n\nThe soft spots are the usual dataset-paper ones, and they are disclosed. The consensus process is standard practice in clinical annotation, but using that consensus as the gold standard to score the annotators themselves is circular. The test set selection from high-disagreement files makes it a deliberately hard set, so the numbers are not representative of typical performance. Tension is the weakest dimension; at alpha=0.18 it is close to unusable, and the claim that the scheme 'aligns with clinical practice' is not independently validated. Baselines lack error bars, and the data is not yet released, so the numbers cannot be verified. None of these are fatal if the paper is framed as a resource description rather than a benchmark paper.\n\nWho this is for: researchers working on automatic stuttering detection or clinical speech assessment. They will get value from the annotation scheme and, once the data is out, a much-needed multimodal benchmark. A serious referee should engage with it, but should push for an independent validation set or at least a clear statement that the reported F1 values are agreement-based, not ground-truth-based. I'd accept it for review and recommend major revisions that reframe the evaluation and release the data.","headline":"A genuinely useful multimodal stuttering annotation resource, but the self-referential gold standard and the deliberately hard test set mean the reported F1 scores are group-conformity numbers, not clinical ground truth.","tokens_in":9198,"tokens_out":2751,"would_cite":true,"duration_ms":26663,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a multi-dimensional, clinically aligned annotation scheme for stuttering severity — covering disfluency type, secondary behavior, and tension — and applies it to audiovisual recordings of adults who stutter, yielding a…","keywords":["stuttering severity assessment","disfluency annotation","clinical speech-language pathology","multi-modal annotation","inter-annotator agreement","consensus gold standard","secondary behaviors","tension scoring"],"falsifier":"If a fresh panel of independently trained clinicians, who did not participate in the original consensus, annotates the same test files and their majority labels disagree substantially with the published gold labels (for example, macro F1 well below the original annotators' scores), the claim that the test set is a reliable consensus gold standard would be refuted.","tokens_in":8172,"feed_emoji":"🗣️","tokens_out":7072,"duration_ms":64018,"temperature":0.7,"pith_summary":"This paper argues that clinical stuttering severity assessment cannot be supported by existing audio-only, non-expert annotations, and that a multi-dimensional scheme is needed that captures disfluency type, secondary behaviors, and tension together. It proposes such a scheme, grounded in established clinical taxonomies, and applies it to 66 audiovisual recordings of adults who stutter, annotated independently by three speech-language pathologists. It also constructs a consensus-based test set in which disagreed-upon stuttering events were re-reviewed and resolved, and uses it to score individual annotators and aggregation methods. The reported inter-annotator agreement and baseline model results quantify how hard the task is, and the released labels are intended to let automatic stuttering assessment be trained and evaluated against clinically meaningful targets.","feed_headline":"Clinician-written stuttering labels ground AI severity scoring","feed_subtitle":"Three expert pathologists label disfluency type, secondary behavior, and tension in 66 videos.","key_machinery":"The carrier of the argument is the annotation scheme itself: each stuttering span receives a primary disfluency type from a standard behavioral taxonomy of stuttering-like disfluencies, a secondary-behavior category (verbal, facial grimace, head movement, extremity movement), and a tension level on a 0-3 scale; annotators mark span boundaries using the acoustic waveform alongside video and transcript in a unified multi-modal annotation tool. The consensus process is the second load-bearing mechanism: a combined file with a disagreement tier focuses discussion on disputed events, and the resolved labels form the gold test set. Aggregation baselines (majority vote and distance-based selection methods) and segment-level evaluation metrics supply the quantitative frame that lets annotator quality and model performance be compared against the gold labels.","core_discovery":"On its own terms, the paper's central discovery is that a clinically valid stuttering annotation can be operationalized as a triple of labels per stuttering moment — primary disfluency type from a behavioral taxonomy, secondary behavior category, and a 0-3 tension score — and applied across a public corpus of adult stuttered speech at scale, yielding 1,654 reading spans and 4,037 interview spans. The authors show that expert clinicians disagree substantially, especially on tension, and that a structured consensus process can produce gold-standard labels for a test set. Against those labels, individual clinicians achieve macro F1 scores between about 0.67 and 0.79, simple aggregation methods do not consistently beat the best annotator, and baseline machine-learning models reach 0.95 F1 for detecting any stuttering event but much lower scores for specific types, with audio alone best for primary disfluencies and video or multi-modal input best for secondary behaviors. The claim being established is that this resource, with its multiple dimensions and consensus labels, is a necessary step toward automatic systems that assess severity rather than merely detect disfluency.","pith_inferences":["If the released labels are used as training targets, the very low inter-annotator agreement on tension means models trained on tension scores will inherit noisy supervision; an ordinal or anchor-based relabeling of tension may be needed before it can be predicted reliably.","Because the gold labels were produced by the same three annotators whose individual labels are scored, and the test files were deliberately selected from high-disagreement cases, the published annotator F1 numbers likely overstate how well a new, independent clinician would match the gold; an external validation panel would settle this.","A natural extension the authors do not build is to combine the annotated spans into clinical severity indices — percentage of stuttered syllables, average tension, and secondary-behavior counts — which would directly test whether the scheme predicts therapist severity ratings."],"forward_implications":["Automatic stuttering assessment can move from binary disfluency detection toward clinically meaningful outputs: type, secondary behavior, and tension labels for each stuttering moment.","The consensus test set gives a reproducible target for comparing future annotators and models, so reported F1 scores across studies become comparable.","The multimodal baseline results indicate that fusing audio and video improves detection of any stuttering event and of secondary behaviors, while audio-only remains stronger for primary disfluency typing; future systems should decide per-label modality.","Because agreement improved from the reading section to the later-annotated interview section, iterative discussion rounds on small batches may be an effective protocol for maintaining expert label reliability in larger annotation efforts."],"supporting_citations":[{"why":"Supplies the clinical assessment instrument and manual that the annotation guidelines are based on.","marker":"[12]"},{"why":"Supplies the audiovisual recordings and orthographic transcripts that are annotated.","marker":"[15]"},{"why":"Provides the podcast-derived stuttering dataset with non-expert labels that the paper contrasts as clinically insufficient.","marker":"[16]"},{"why":"Provides an existing speech-language-pathologist-annotated audio corpus that lacks the visual dimensions the scheme adds.","marker":"[18]"},{"why":"Supplies the disfluency taxonomy used for primary stuttering types.","marker":"[23]"},{"why":"Supplies the secondary-behavior classification and the reading passage used in the reading samples.","marker":"[13]"},{"why":"Supplies the 0-3 tension scale used for severity scoring.","marker":"[26]"},{"why":"Supplies the span-level agreement calculation method used for inter-annotator agreement.","marker":"[27]"},{"why":"Supplies the annotation aggregation methods compared in the test set evaluation.","marker":"[28]"},{"why":"Supplies the segment-based evaluation metrics used for annotator and model F1 scores.","marker":"[29]"}],"fun_headline_variants":["Clinician-crafted stuttering labels with consensus gold test set","Stuttering severity dataset: expert labels on 66 videos","Audio best for disfluency, video for secondary behaviors","Expert annotators disagree: consensus needed for stuttering labels","Multi-modal expert annotations target stuttering severity AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The gold-standard test labels come from consensus among the same three clinicians whose individual annotations are being evaluated, and the test files were chosen because they contained the most disagreements between two of those clinicians; the assumption is that this internal consensus is a valid external ground truth for measuring annotation quality and model performance.","fun_headline_variants_meta":{"raw":{"variants":["Clinician-crafted stuttering labels with consensus gold test set","Stuttering severity dataset: expert labels on 66 videos","Audio best for disfluency, video for secondary behaviors","Expert annotators disagree: consensus needed for stuttering labels","Multi-modal expert annotations target stuttering severity AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1554,"prompt_tokens":883,"completion_tokens":671,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":591}},"tokens_in":499,"tokens_out":671,"duration_ms":7218,"temperature":1.0,"reasoning_tokens":591,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:00:46.313075+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a fresh panel of independently trained clinicians, who did not participate in the original consensus, annotates the same test files and their majority labels disagree substantially with the published gold labels (for example, macro F1 well below the original annotators' scores), the claim that the test set is a reliable consensus gold standard would be refuted.","supporting_citations":[{"cited_title":"Guitar,Stuttering: An integrated approach to its nature and treatment","cited_arxiv_id":null,"evidence_quote":"Supplies the clinical assessment instrument and manual that the annotation guidelines are based on."},{"cited_title":"The impact of stuttering on the quality of life in adults who stutter,","cited_arxiv_id":null,"evidence_quote":"Supplies the audiovisual recordings and orthographic transcripts that are annotated."},{"cited_title":"Stuttering and the international classification of functioning, disability, and health (icf): An up- date,","cited_arxiv_id":null,"evidence_quote":"Provides the podcast-derived stuttering dataset with non-expert labels that the paper contrasts as clinically insufficient."},{"cited_title":"Yairi and C","cited_arxiv_id":null,"evidence_quote":"Provides an existing speech-language-pathologist-annotated audio corpus that lacks the visual dimensions the scheme adds."},{"cited_title":"Stuttering moments","cited_arxiv_id":null,"evidence_quote":"Supplies the disfluency taxonomy used for primary stuttering types."},{"cited_title":"Long-term consequences of child- hood bullying in adults who stutter: Social anxiety, fear of nega- tive evaluation, self-esteem, and satisfaction with life,","cited_arxiv_id":null,"evidence_quote":"Supplies the secondary-behavior classification and the reading passage used in the reading samples."},{"cited_title":"The uclass archive of stut- tered speech","cited_arxiv_id":null,"evidence_quote":"Supplies the 0-3 tension scale used for severity scoring."},{"cited_title":"Fluentnet: End-to- end detection of stuttered speech disfluencies with deep learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the span-level agreement calculation method used for inter-annotator agreement."},{"cited_title":"Ksof: The kassel state of fluency dataset – a therapy centered dataset of stuttering,","cited_arxiv_id":null,"evidence_quote":"Supplies the annotation aggregation methods compared in the test set evaluation."},{"cited_title":"A longitudinal study of stuttering in children: A preliminary report,","cited_arxiv_id":null,"evidence_quote":"Supplies the segment-based evaluation metrics used for annotator and model F1 scores."}],"review_version":1}