{"id":"3cfdd8b1-6b59-4397-ba8a-f86bd0a29be6","arxiv_id":"1908.08668","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Combining CWT peak evidence with STM phone-boundary correction improves VOP localization accuracy at 10 ms deviation on TIMIT and Bengali read/conversation speech.","lead":"A two-stage detector combines wavelet peaks with sound-boundary corrections to locate vowel onset points more precisely in speech. It reports 82% identification within 10 ms on the TIMIT corpus, up from about 62% for an earlier method.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The two-stage correction rule in Section 3.3 assumes CWT VOP errors are predominantly rightward and that the nearest left STM boundary is the true VOP; the paper only supports this with an unquantified 'mostly' on a subset, so a non-trivial left-error or spurious-boundary rate could break the…","rationale":"I considered the alternative concern that the CWT and STM thresholds are tuned on a subset of TIMIT that may overlap the 120-utterance test set, which would bias the comparison. That is a legitimate methodological issue, but the paper's own text flags the correction rule's support as anecdotal ('reported for a subset... mostly'), and the reviewing rule directs us to weigh such self-asserted limitations explicitly. The correction rule is the component that converts a 52% within-10 ms CWT detector into an 82% system; if its directional assumption fails in a non-negligible fraction of cases, the headline result changes materially. The proposed concrete test directly measures the assumption on the same corpus, so it would settle whether the concern lands. The reader's weakest_assumption identifies the same mechanism, and I agree with the conditional verdict: the paper is plausible but not yet fully supported without a quantitative failure analysis of the correction rule.","tokens_in":14360,"tokens_out":8854,"duration_ms":92345,"concrete_test":"On the same 120 TIMIT utterances used for Table 3, compute for every true VOP: (a) the signed deviation of the CWT-detected peak, and (b) the phone identity of the nearest STM boundary to the left of that peak (vowel onset vs. other boundary). Then re-run the evaluation with two alternative correction rules: nearest STM boundary on the right, and global nearest STM boundary. If the 10 ms identification rate changes by more than a few points, or if more than 10% of corrections land on boundaries that are not vowel onsets, the reported advantage is not robust to the left-nearest assumption and the central claim should be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Table 3: 82% identification rate within 10 ms) is produced by the correction rule in Section 3.3, steps 3–4: for every CWT-detected VOP, the algorithm locates the nearest STM phone boundary to its left and redefines that boundary as the VOP. This rule is valid only if (i) CWT peaks systematically lie to the right of the true VOP and (ii) the nearest left STM boundary is the consonant–vowel boundary, not a spurious STM peak. The paper supports (i) with a single sentence, 'detected VOPs are mostly occurring right side of the actual VOPs,' described as 'reported for a subset of TIMIT' (Section 3.3), with no quantitative distribution. Condition (ii) is not examined, even though Table 2 shows that treating all STM boundaries as VOPs yields a 70% spurious rate, so many STM boundaries are not vowel onsets. If a CWT false peak falls before a VOP, or if the nearest left boundary is a consonant–consonant transition or an intra-vowel spurious peak, the corrected time is wrong. The reported 82% at 10 ms and 7 ms average deviation depend directly on these unverified conditions; without a failure-case analysis, the improvement over COMB-ESM and SE-GCI could be an artifact of the heuristic rather than of the proposed CWT+STM combination.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage method for vowel onset point (VOP) detection. In the first stage, the continuous wavelet transform (CWT) coefficients of the speech signal are averaged over scale to form a mean-signal, and VOP candidates are detected as peaks of the smoothed average absolute magnitude of this mean-signal. In the second stage, each detected VOP is corrected by relocating it to the nearest spectral-transition-measure (STM) phone boundary to its left. The method is evaluated on TIMIT and on Bengali read and conversation speech, and compared with two signal-processing baselines, COMB-ESM and SE-GCI. The central quantitative claim is an identification rate of 82% within 10 ms and an average deviation of about 7 ms on TIMIT (Table 3), compared with 62% and 51% for the baselines at the same tolerance. The paper also reports results for read and conversation modes of Bengali, where the proposed method shows smaller degradation in conversation mode than the baselines.","tokens_in":14701,"tokens_out":7887,"duration_ms":76370,"significance":"If the results are robust, the proposed method would be a simple, fully signal-processing-based approach that achieves substantially better VOP timing accuracy at low tolerances than two standard methods, and it addresses an under-studied speaking mode (conversation speech). The use of an unsupervised STM phone-boundary detector is a strength, as is the explicit comparison with two published baselines and the evaluation on two corpora. However, the significance is conditional on several unresolved points: the final VOP positions are, by construction, STM phone boundaries, so the reported accuracy is largely inherited from the STM stage rather than from the CWT analysis; the correction rule rests on an unquantified rightward-bias assumption; the mother wavelet and scales are unspecified; and the evaluation lacks independence between threshold tuning and testing, as well as any statistical significance testing. With these addressed, the contribution could be a meaningful practical improvement for applications needing VOPs within 10–20 ms.","major_comments":[{"comment":"The correction rule defines every final VOP as the nearest STM phone boundary to the left of a CWT-detected VOP. This is valid only if CWT-detected VOPs are systematically to the right of true VOPs and if the nearest left STM boundary is the consonant–vowel boundary rather than a spurious or unrelated boundary. The paper's support for the rightward bias is a single unquantified statement, \"detected VOPs are mostly occurring right side of the actual VOPs,\" based on a subset of TIMIT, and it does not examine the possibility that the nearest left STM boundary is a consonant–consonant transition or an intra-vowel spurious peak, even though Table 2 shows that 70% of STM boundaries are not VOPs. The Table 3 results (82% within 10 ms, 7 ms average deviation) are produced directly by this rule. Please provide a quantitative distribution of CWT VOP errors by direction and magnitude, and a breakdown of correction outcomes by phonetic context (CV, semivowel–vowel, diphthong, CC boundary) to demonstrate that the heuristic, not the CWT evidence, is not the source of the reported improvement.","section":"Section 3.3, steps 3–4"},{"comment":"The mother wavelet φ(t) and the set of scales qs are never specified. The CWT stage is the core of the proposed first stage, and its behavior—including the claims about suppressing nasals, fricatives, and unvoiced regions in Figure 1—depends entirely on these choices. Without specifying the wavelet family, the scale range, and the number of scales, the experiments cannot be reproduced, and the particular CWT implementation cannot be compared with other wavelet-based approaches. Please specify these parameters and justify them empirically or theoretically.","section":"Section 3.1, Eqs. (1)–(2)"},{"comment":"The thresholds th1=15% and th2=12% are selected by experiments on a subset of TIMIT, and for Bengali the text states that \"in a similar manner, experiments are performed with Bengali read and conversation speech signals,\" implying that thresholds are also tuned on the evaluation corpus. The paper does not state that the 120 TIMIT evaluation utterances (Section 4.1) are disjoint from the threshold-selection subset. If tuning and evaluation use the same corpus, the reported identification rates are not independent estimates of performance on unseen data. Please specify the exact tuning/evaluation split, or fix thresholds a priori, or use cross-validation, and report results accordingly.","section":"Section 3.1 (Table 1) and Section 3.2"},{"comment":"The claim that the proposed method outperforms COMB-ESM and SE-GCI is based on single point estimates with no error bars, confidence intervals, or significance tests. With 120 TIMIT utterances and roughly 100 Bengali utterances per mode, per-utterance metrics are feasible. Please add a paired significance test (e.g., Wilcoxon signed-rank) for identification rate at each tolerance and for average deviation, or provide confidence intervals, so that the 20–30 percentage-point gaps at 10 ms can be assessed statistically rather than impressionistically.","section":"Section 4.1 (Table 3)"},{"comment":"The derivation of ground-truth VOPs is not described. TIMIT provides phone segment labels, not vowel onset labels; the paper should state explicitly how the reference VOPs were computed (for example, as the boundary of the first vowel phone in each syllable) and what criterion was used for matching detected VOPs to reference VOPs when computing missed and spurious rates. The same information is needed for the Bengali corpus, including whether manual annotation or automatic segmentation was used. Without this, the evaluation is not reproducible.","section":"Section 4.1"}],"minor_comments":[{"comment":"The heading \"Performance Evaluaion\" contains a typo; it should read \"Performance Evaluation.\"","section":"Section 4 heading"},{"comment":"The captions of Figures 5 and 6 state that \"Actual phone boundaries\" and \"detected phone boundaries\" are marked, but the panels display VOPs, not phone boundaries; the captions should say \"VOPs\" to match panels (a) and (d).","section":"Figures 5 and 6 captions"},{"comment":"The dataset description is confusing: the text says \"Altogether, 20 utterances are collected from 5 distinct speakers\" and then \"About 100 utterances from each mode are selected for evaluating\"; please clarify whether the 20 utterances are a subset used for threshold tuning and what the relation is to the 100 utterances per mode used in the evaluation.","section":"Section 4.2"},{"comment":"The text says that \"the well-known difficulties in VOP detection are finding false VOPs in case of semi-vowels, nasals, and fricatives as they are periodic in nature,\" but fricatives are typically aperiodic; later in the same paragraph the paper states that CWT coefficients are \"almost zero in fricatives and unvoiced speech regions,\" which is consistent with aperiodicity. Please reconcile this wording.","section":"Section 3.1, paragraph 4"},{"comment":"In the extracted text, the integral symbol in Eq. (1) appears garbled as \"/uni222B.dsp\"; please ensure the equation renders correctly in the final manuscript.","section":"Eq. (1)"},{"comment":"The claim that \"there is no work related to CWT for VOP detection\" is a strong negative claim; please soften it to \"to the best of the authors' knowledge\" or provide a brief literature survey to support it.","section":"Section 3.1, paragraph 2"}],"recommendation":"major_revision","confidential_remarks":"The paper's main methodological risk is that the final VOP positions are literally the STM phone boundaries selected by the CWT candidates; the reported 82%-within-10 ms result is therefore essentially the STM phone-boundary accuracy, with the CWT stage acting as a spurious-boundary filter. I recommend asking the authors for an ablation study (e.g., using all STM boundaries with the same spurious-removal rule, or substituting randomly chosen left-boundary candidates) to demonstrate the marginal contribution of the CWT stage beyond the selection. The paper is not a reject; the two-stage idea is legitimate and the conversation-mode evaluation is a useful contribution, but the load-bearing assumptions in Section 3.3 and the missing experimental details need to be addressed. I would also encourage the editor to ensure the paper is reviewed by someone who can assess whether the unspecified wavelet/scale choices are conventional in the speech-processing community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper has a plausible new idea—CWT for VOP detection plus STM phone-boundary correction—but the numbers that matter come from an unvalidated correction heuristic. I'd send it to review, but the authors need to be pushed on that.\n\nWhat's genuinely new: nobody seems to have used CWT this way for VOPs, and the two-stage design is a sensible way to combine coarse VOP evidence with a finer boundary estimator. The evaluation on Bengali read and conversation speech is a plus, and the paper is clearly written with useful figures. The reported 82% within 10 ms on TIMIT is a big jump over the two baselines, and the spurious-rate drop is in the right direction.\n\nThe central weakness is Section 3.3. The algorithm takes each CWT-detected VOP and replaces it with the nearest STM phone boundary to the left. That works only if CWT errors are mostly rightward and if the nearest left boundary is actually the consonant-vowel boundary. The paper's evidence for the first point is a single sentence saying detected VOPs are 'mostly' to the right, based on a subset, with no distribution. The second point is not examined at all, even though Table 2 shows that treating all STM boundaries as VOPs gives a 70% spurious rate; many STM boundaries are not vowel onsets. The paper claims the spurious STM boundaries don't affect the two-stage method, but that claim is presented without supporting experiments. If a CWT peak falls left of the true VOP, or if the nearest left boundary is a consonant-consonant transition, the correction moves the VOP to the wrong place. A few failure-case examples would settle this.\n\nOther soft spots are more minor: the mother wavelet and scale set are never specified, thresholds are tuned on a subset of the same corpus used for evaluation, there are no error bars or significance tests, and the baselines are from 2009 and 2012, not the recent methods the paper itself cites. These are all fixable.\n\nThe method is not circular in a damaging sense: the CWT stage does real filtering, and the combination is a reasonable engineering heuristic. The claims are just stronger than the current evidence. I'd want a revised version with a quantitative analysis of the correction rule's assumptions and the missing experimental details. For a researcher working on VOP detection or speech segmentation, this is worth reading as a point of departure.","headline":"CWT+STM two-stage VOP detector has a plausible idea and good results, but the correction step that produces the accuracy numbers is under-validated.","tokens_in":15204,"tokens_out":4444,"would_cite":false,"duration_ms":42108,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two-stage CWT plus phone-boundary correction detects vowel onsets within about 7 ms on average, nearly doubling the 10 ms identification rate of prior signal-processing methods.","keywords":["vowel onset point detection","continuous wavelet transform","spectral transition measure","phone boundary","TIMIT corpus","Bengali speech","read speech","conversation speech"],"falsifier":"Restrict the evaluation to a subset of TIMIT and Bengali utterances containing semivowel-vowel and diphthong transitions, and compare the corrected VOP positions against manual labels; if the nearest left STM boundary fails to match the marked onset for a substantial share, the correction rule is shifting those VOPs to the wrong place.","tokens_in":14163,"feed_emoji":"🗣️","tokens_out":5065,"duration_ms":46877,"temperature":0.7,"pith_summary":"The paper argues that vowel onset points (VOPs), the instants vowels begin in speech, can be located far more precisely than prior signal-processing methods allow by combining two kinds of evidence: peaks in the smoothed average absolute magnitude of continuous wavelet transform (CWT) coefficients, and phone boundaries found by the spectral transition measure (STM). On TIMIT, the proposed two-stage method identifies 82 percent of VOPs within 10 ms deviation, with an average deviation of about 7 ms, compared with 62 percent for SE-GCI and 51 percent for COMB-ESM at the same tolerance. The same method is applied to Bengali read and conversation speech, where it degrades less in conversation mode than the baselines do. The central claim is that CWT evidence localizes candidate onsets, and STM phone boundaries correct their position.","feed_headline":"Vowel onsets pinned to 7 ms by wavelet plus phone boundaries","feed_subtitle":"Two-stage method lifts 10 ms VOP accuracy to 82 percent on TIMIT while cutting average deviation to 7 ms.","key_machinery":"The load-bearing machinery is the two-stage correction rule. The CWT stage produces a mean-signal $\\frac{1}{N}\\sum_{q} C_x(p,q)$ from wavelet coefficients at several scales; its frame-wise average absolute magnitude, smoothed and thresholded, yields VOP hypotheses. The STM stage computes $S_g = \\frac{1}{D}\\sum_i r_i^2(g)$, a mean-squared spectral-transition measure per frame, whose peaks mark phone boundaries. The correction step then sets each detected VOP to the nearest phone boundary on its left. This rule is what turns a coarse 40 ms-level detector into one whose reported average deviation is about 7 ms.","core_discovery":"The central discovery is that VOP candidates detected from CWT coefficients are systematically shifted to the right of the true onsets, and that this bias can be corrected without any training by snapping each candidate to the nearest detected phone boundary on its left. The paper's recipe is: compute CWT coefficients of the speech signal, average them over scales to form a mean-signal, take the average absolute magnitude (AAM) over 20 ms frames, smooth over 40 ms, retain local peaks above 15 percent of the smoothed maximum, and remove smaller peaks within a 50 ms window; separately compute STM from 39-dimensional MFCC features and keep phone-boundary peaks above 12 percent of the STM contour maximum; then replace each CWT VOP with the closest phone boundary to its left. The authors report that this yields an identification rate of 82 percent within 10 ms and an average deviation of about 7 ms on TIMIT, and that the advantage persists on Bengali conversation speech, where the 10 ms identification rate is 72 percent versus 45 percent for SE-GCI and 40 percent for COMB-ESM.","pith_inferences":["Editorial extension: the same snap-to-nearest-left-phone-boundary correction could be applied to other VOP detectors, such as COMB-ESM or SE-GCI, and would likely reduce their deviation as well, since the bias correction is independent of the CWT front end.","Editorial extension: the method's success depends on STM boundaries being accurate; on semivowel-vowel and diphthong transitions, where spectral changes are gradual, the nearest left boundary may be a spurious STM peak, so a targeted evaluation on those phone classes is the natural next test.","Editorial extension: the approach suggests a symmetric counterpart for vowel end point detection by snapping to the nearest right phone boundary, an extension the authors mention as future work."],"forward_implications":["If the claim holds, speech segmentation and consonant-vowel unit spotting can use VOPs accurate to roughly 7 ms instead of accepting 40 ms error bars.","For Indian-language speech applications built on CV units, the method provides accurate onsets without training a vowel detector.","Because the method is unsupervised and signal-based, it transfers to new languages without retraining, as the Bengali read and conversation results illustrate.","In conversation speech, where vowels shorten and paralinguistic events create spurious peaks, the correction stage still reduces spurious rates to 7 percent versus 13 percent for SE-GCI and 15 percent for COMB-ESM in the Bengali data."],"supporting_citations":[{"why":"Supplies the STM method for phone-boundary detection that provides the correcting evidence in the second stage.","marker":"[25]"},{"why":"Defines the mean-squared spectral transition measure used to build the STM contour.","marker":"[26]"},{"why":"Supplies the continuous wavelet transform framework used in the first stage.","marker":"[23]"},{"why":"Is the COMB-ESM baseline whose identification rate, deviation, and spurious rate are compared against the proposed method.","marker":"[13]"},{"why":"Is the SE-GCI baseline and the source of the 50 ms single-VOP window heuristic used for peak pruning.","marker":"[14]"},{"why":"Is the TIMIT corpus used for the main identification-rate and deviation evaluation.","marker":"[28]"},{"why":"Is the Bengali read and conversation corpus used for the mode comparison.","marker":"[29]"},{"why":"Motivates the importance of VOPs for consonant-vowel unit speech recognition in Indian languages.","marker":"[8]"}],"fun_headline_variants":["CWT plus phone boundaries cuts VOP error to 7 ms","Two-stage VOP detection hits 82% at 10 ms on TIMIT","Wavelet VOPs corrected by phone boundaries to 7 ms","Accurate vowel onset detection via CWT and phone boundaries","VOP detection improves to 7 ms with phone-boundary snapping"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that for every CWT-detected VOP the nearest STM phone boundary to its left is the true vowel onset; that is plausible for simple consonant-to-vowel transitions but is not guaranteed for semivowel-vowel transitions, diphthongs, or cases where the CWT peak sits left of the true VOP.","fun_headline_variants_meta":{"raw":{"variants":["CWT plus phone boundaries cuts VOP error to 7 ms","Two-stage VOP detection hits 82% at 10 ms on TIMIT","Wavelet VOPs corrected by phone boundaries to 7 ms","Accurate vowel onset detection via CWT and phone boundaries","VOP detection improves to 7 ms with phone-boundary snapping"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000501,"raw_usage":{"total_tokens":2453,"prompt_tokens":949,"completion_tokens":1504,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":1422}},"tokens_in":565,"tokens_out":1504,"duration_ms":9036,"temperature":1.0,"reasoning_tokens":1422,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:32:27.085166+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Restrict the evaluation to a subset of TIMIT and Bengali utterances containing semivowel-vowel and diphthong transitions, and compare the corrected VOP positions against manual labels; if the nearest left STM boundary fails to match the marked onset for a substantial share, the correction rule is shifting those VOPs to the wrong place.","supporting_citations":[{"cited_title":"Dusan, L","cited_arxiv_id":null,"evidence_quote":"Supplies the STM method for phone-boundary detection that provides the correcting evidence in the second stage."},{"cited_title":"Furui, On the role of spectral transition for speech pe rception, The Journal of the Acoustical Society of America 80 (4) (1986) 1016–1025","cited_arxiv_id":null,"evidence_quote":"Defines the mean-squared spectral transition measure used to build the STM contour."},{"cited_title":"Stephane, A wavelet tour of signal processing, The Spa rse W ay","cited_arxiv_id":null,"evidence_quote":"Supplies the continuous wavelet transform framework used in the first stage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the COMB-ESM baseline whose identification rate, deviation, and spurious rate are compared against the proposed method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the SE-GCI baseline and the source of the 50 ms single-VOP window heuristic used for peak pruning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the TIMIT corpus used for the main identification-rate and deviation evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the Bengali read and conversation corpus used for the mode comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the importance of VOPs for consonant-vowel unit speech recognition in Indian languages."}],"review_version":1}