{"id":"77dcd847-258f-4330-b80f-1922ba81e559","arxiv_id":"2508.02255","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"StutterCut uses uncertainty-guided normalised cut on speech embeddings to segment stuttering dysfluencies at frame level, outperforming existing methods on real and synthetic benchmarks.","lead":"The paper introduces StutterCut, a semi-supervised framework that turns dysfluency segmentation into a graph partitioning problem using speech embeddings from overlapping windows. It also extends the FluencyBank dataset with frame-level dysfluency boundaries, and reports higher F1 scores and more precise stuttering onset detection on real and synthetic datasets.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The supplied full text is an unrelated economics paper (arXiv:2508.02252), not the StutterCut manuscript (arXiv:2508.02255), so the central empirical claim is unsupported and cannot be stress-tested.","rationale":"I read the abstract and the supplied full text. The central claim is an empirical performance claim: StutterCut outperforms existing methods on dysfluency segmentation. To validate it, one needs the method description, dataset construction, evaluation protocol, baselines, and results. None of these are present in the submitted full text, which is an unrelated economics paper. This is a stronger, more fundamental concern than the reader's weakest_assumption about pseudo-oracle labels, because it blocks any technical assessment of the method. The reader's report correctly identified the full-text mismatch as a red flag and reached UNVERDICTED; my concern does not change that verdict. I did not find an internal inconsistency in the abstract's description itself, but the absence of the actual manuscript makes the correctness_risk unknowable. The concrete test is simply to obtain the real paper; if the mismatch persists, the claim should remain unverified. I am not alleging misconduct; I am stating that the evidence supplied for this review does not support the abstract's assertions. Agreement_with_reader is 'disagree' because the reader's stated weakest_assumption is about the pseudo-oracle mechanism, not the full-text mismatch, even though the reader's rationale flags the mismatch separately.","tokens_in":2220,"tokens_out":4595,"duration_ms":58673,"concrete_test":"Fetch the actual full text of arXiv:2508.02255. Verify that it is the StutterCut paper and that it contains a methods section describing the graph-partitioning pipeline, the pseudo-oracle/uncertainty mechanism, the FluencyBank extension, and an experiments section reporting F1 and onset metrics against baselines on real and synthetic datasets. If the full text is unavailable or still mismatched, the central claim must remain unverified and the UNVERDICTED verdict stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing condition for the abstract's central claim is that a full technical paper exists whose methods, experiments, and comparisons match the abstract. That condition fails: the manuscript body is a completely different paper, 'FX-constrained growth: Fundamentalists, chartists and the dynamic trade-multiplier' (arXiv:2508.02252v1 [econ.GN]), while the abstract is from a cs.SD submission. There is no description of StutterCut's graph construction, no specification of the pseudo-oracle classifier or the Monte Carlo dropout gating, no dataset annotation protocol for the FluencyBank extension, and no results or baseline comparisons. Thus the claim of higher F1 and more precise stuttering onset detection is entirely unverified. The reader's identified weakest assumption about the pseudo-oracle is a plausible technical risk, but it is secondary: even if that mechanism were internally valid, the absence of any evaluative content leaves the central claim unsupported. The correct posture is that the paper is unverdictable from the provided materials.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript under review, identified as arXiv:2508.02255 (cs.SD), presents an abstract for a paper titled 'StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation.' The abstract claims a semi-supervised framework that formulates dysfluency segmentation as a graph partitioning problem, uses a pseudo-oracle classifier trained on weak utterance-level labels, controls the classifier's influence via Monte Carlo dropout uncertainty, and reports improved F1 scores and more precise stuttering onset detection on real and synthetic datasets. The abstract also describes an extension of the FluencyBank dataset with frame-level dysfluency boundaries for four dysfluency types. However, the full text supplied with the submission is not this paper; it is a completely unrelated economics manuscript titled 'FX-constrained growth: Fundamentalists, chartists and the dynamic trade-multiplier' by Marwil J. Dávila-Fernández and Serena Sordi (arXiv:2508.02252, econ.GN). The submitted body contains no description of StutterCut, no methods, no experimental setup, no results, and no baseline comparisons. Thus, the central claims of the abstract are unsupported by the manuscript text.","tokens_in":2401,"tokens_out":2380,"duration_ms":27668,"significance":"If the StutterCut method and its reported results were fully described and validated, the contribution could be significant: it would offer a novel formulation of dysfluency segmentation as a graph-partitioning problem with weak supervision, and the claimed frame-level extension of FluencyBank would be a useful community resource. The use of Monte Carlo dropout uncertainty to gate a pseudo-oracle is an interesting mechanism, and higher F1 with more precise onset detection would be a meaningful practical advance for speech therapy tools. However, because the supplied full text is an unrelated economics paper, none of these contributions are actually present or assessable in the manuscript. The significance of the reported claims cannot be evaluated, and the submission as it stands does not constitute a coherent scientific paper.","major_comments":[{"comment":"The submitted manuscript body is not the paper described in the abstract. The abstract describes StutterCut, a dysfluency segmentation method with graph partitioning, weak-label pseudo-oracle refinement, and Monte Carlo dropout uncertainty gating, along with experiments on FluencyBank and synthetic datasets. The full text, however, is an economics paper titled 'FX-constrained growth: Fundamentalists, chartists and the dynamic trade-multiplier' (arXiv:2508.02252), which contains none of these elements. There is no description of StutterCut's graph construction, no specification of the pseudo-oracle classifier, no details of the Monte Carlo dropout gating, no annotation protocol for the FluencyBank extension, and no experimental results or baseline comparisons. Consequently, the central empirical claim of the abstract—that StutterCut outperforms existing methods with higher F1 scores and more precise stuttering onset detection—is entirely unsupported by the manuscript. This is a load-bearing error that cannot be fixed without replacing the entire manuscript text.","section":"Full Text (entire)"},{"comment":"Even taken on its own, the abstract reports quantitative improvements ('higher F1 scores and more precise stuttering onset detection') without any supporting experimental details: no error bars, no statistical significance tests, no dataset splits, no baseline names, and no metrics for the claimed frame-level onset precision. The abstract also asserts that the weakly labelled FluencyBank dataset is extended with frame-level dysfluency boundaries for four dysfluency types, but provides no information about the annotation protocol, inter-annotator agreement, or how these labels were used in evaluation. These omissions are material because the reader cannot verify the strength or validity of the reported advantages. This comment is secondary to the full-text mismatch but would need to be addressed even if the correct manuscript were provided.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase 'normalised cut' is used without indicating the specific graph-partitioning objective or how the graph nodes are connected, which would be necessary for reproducibility.","section":"Abstract"},{"comment":"The term 'pseudo-oracle' is undefined in the abstract; it would be helpful to clarify whether this is a classifier trained on noisy labels or a specifically constructed heuristic.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"To the editor: this submission appears to have a serious integrity problem. The abstract and the full text are different papers from different arXiv categories. This is not a case of a minor formatting error; the entire technical content of the claimed StutterCut paper is absent. Even under a generous reading that treats the abstract as the sole content, the claims are unverifiable without any methods or results. I recommend rejection, and I would also flag this to the editorial office for possible submission-integrity review. I cannot recommend transfer to another venue, because the economics paper is outside this journal's scope and the StutterCut paper is not present."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The abstract for 2508.02255 describes a sensible idea: treat dysfluency segmentation as graph partitioning over speech embedding windows, refine edges with a pseudo-oracle trained on utterance-level weak labels, gate its influence by Monte Carlo dropout uncertainty, and benchmark on a frame-level extension of FluencyBank. That is a real gap, and the FluencyBank extension alone would be a useful community resource. I cannot assess more, because the supplied full text is a completely different paper—an FX/growth model in econ.GN (2508.02252). This is not a missing appendix or a formatting glitch; the manuscript body has nothing to do with StutterCut. There is no graph construction, no pseudo-oracle details, no annotation protocol, no results, no baselines. The central empirical claim in the abstract—\"outperforms existing methods, achieving higher F1 scores and more precise stuttering onset detection\"—is unsupported by anything I can read.\n\nThe reader's weakest assumption about the pseudo-oracle is secondary. Even if that mechanism were internally valid, the absence of the actual paper leaves the claim unverifiable. I would add that the abstract alone, even if it were backed by the right manuscript, would need error bars, statistical tests, and dataset split details before the F1 comparison could be trusted. But that is a minor concern relative to the mismatch.\n\nWho is this for? Nobody can use this submission as-is. If the authors intended to submit StutterCut, they uploaded the wrong file. As it stands, it should be desk rejected, not sent to referees. If a correct version exists, the idea is worth a look, but I cannot base that judgment on this submission.\n\nMy recommendation: return to authors as a broken submission. Do not send to peer review. If and when the actual StutterCut paper appears, I would be willing to take a serious look.","headline":"The abstract describes a plausible dysfluency segmentation method, but the submitted full text is an unrelated economics paper, so there is nothing to review.","tokens_in":2900,"tokens_out":2044,"would_cite":false,"duration_ms":22784,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"StutterCut: graph partitioning converts weak utterance labels into frame-level stutter boundaries, outperforming existing methods on real and synthetic speech.","keywords":["dysfluency segmentation","stuttering","graph partitioning","normalised cut","semi-supervised learning","Monte Carlo dropout","weak labels","FluencyBank"],"falsifier":"A direct evaluation on FluencyBank comparing StutterCut's frame-level boundaries against expert annotations for each of the four dysfluency types, reporting per-type F1 and onset timing error; if the method's precision on onsets does not exceed a simple frame-level classifier trained on a small amount of frame-level labels, the central claim would be undermined.","tokens_in":2041,"feed_emoji":"🗣️","tokens_out":2773,"duration_ms":29755,"temperature":0.7,"pith_summary":"This paper introduces StutterCut, a semi-supervised method that turns dysfluency segmentation into a graph partitioning problem. Speech embeddings from overlapping windows become graph nodes, and a pseudo-oracle classifier trained on utterance-level weak labels refines the connections, with Monte Carlo dropout uncertainty controlling how much that refinement affects the final cut. The authors also extend FluencyBank by adding frame-level dysfluency boundaries for four dysfluency types, creating a more realistic benchmark than synthetic datasets. If the claims hold, StutterCut provides frame-level stutter detection that outperforms existing methods, achieving higher F1 scores and more precise stuttering onset detection.","feed_headline":"Graph-cut method sharpens stutter onset detection","feed_subtitle":"Semi-supervised StutterCut finds frame-level stutter boundaries more precisely on real and synthetic speech.","key_machinery":"The central object is the graph partitioning formulation: overlapping speech windows become graph nodes, and a normalised cut partitions the graph into fluent and dysfluent segments. The key refinement mechanism is the pseudo-oracle classifier, trained on utterance-level weak labels, which adjusts the connections between nodes; its influence is gated by Monte Carlo dropout uncertainty so that uncertain refinements do not distort the cut. The extended FluencyBank dataset with frame-level boundaries for four dysfluency types serves as the evaluation benchmark.","core_discovery":"The central claim is that frame-level dysfluency segmentation can be achieved from only utterance-level weak labels by combining normalised cut graph partitioning with an uncertainty-controlled pseudo-oracle. On both real and synthetic datasets, StutterCut reports higher F1 scores and more precise stuttering onset detection than existing methods, suggesting that graph partitioning can resolve boundaries that utterance-level classifiers miss.","pith_inferences":["The uncertainty-gated pseudo-oracle is the load-bearing component; a testable extension would be to compare Monte Carlo dropout with other uncertainty estimators, such as deep ensembles, to see whether the gating mechanism improves across the board.","The method's sensitivity to graph construction choices (window size, embedding type, similarity metric) is not visible in the abstract; if the method is highly sensitive to these hyperparameters, transfer to new languages or recording conditions may be limited.","Since the extended dataset annotates four dysfluency types, a natural next step is per-type segmentation evaluation, which could reveal whether the graph cut handles all types uniformly or favours particular types."],"forward_implications":["StutterCut enables frame-level stutter boundaries without frame-level training labels, substantially reducing annotation cost for dysfluency research.","The graph-partitioning-plus-pseudo-oracle recipe may transfer to other weakly supervised segmentation tasks in speech or audio, such as disfluency detection in spontaneous dialogue.","The extended FluencyBank dataset gives the community a realistic benchmark for comparing frame-level dysfluency segmentation methods.","More precise onset detection could make real-time feedback systems for speech therapy react to stutters earlier and more reliably."],"supporting_citations":[],"fun_headline_variants":["Graph-cut method locates stutter boundaries precisely","Semi-supervised graph partition finds frame-level stutter onset","Uncertainty-guided normalised cut segments dysfluencies","StutterCut: from weak labels to precise stutter boundaries","Graph partitioning improves stutter segmentation F1"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The refinement of the graph connections by a pseudo-oracle trained on utterance-level labels must be both accurate and correctly controlled by the Monte Carlo dropout uncertainty, so that the partition boundaries reflect true dysfluency rather than classifier error.","fun_headline_variants_meta":{"raw":{"variants":["Graph-cut method locates stutter boundaries precisely","Semi-supervised graph partition finds frame-level stutter onset","Uncertainty-guided normalised cut segments dysfluencies","StutterCut: from weak labels to precise stutter boundaries","Graph partitioning improves stutter segmentation F1"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000661,"raw_usage":{"total_tokens":2926,"prompt_tokens":753,"completion_tokens":2173,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":369,"completion_tokens_details":{"reasoning_tokens":2095}},"tokens_in":369,"tokens_out":2173,"duration_ms":17564,"temperature":1.0,"reasoning_tokens":2095,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:01:27.916483+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct evaluation on FluencyBank comparing StutterCut's frame-level boundaries against expert annotations for each of the four dysfluency types, reporting per-type F1 and onset timing error; if the method's precision on onsets does not exceed a simple frame-level classifier trained on a small amount of frame-level labels, the central claim would be undermined.","supporting_citations":[],"review_version":1}