{"id":"93a7fb1b-1e34-4fbc-9229-212f5d5d50eb","arxiv_id":"2501.09159","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":21,"one_line_summary":"Fully convolutional networks trained on synthesized subharmonic phonation classify subharmonic period M=1-4 with >98% synthetic accuracy, with qualitative only evidence on real voices.","lead":"Two convolutional neural networks, trained on thousands of synthesized voice signals, label short audio clips as normal or as subharmonic with period two, three, or four, reaching above 98% accuracy in simulation. The networks also respond plausibly to real disordered-voice recordings, but the paper has not yet proven how well they work on real patients.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 98.9% accuracy is measured in-distribution on the paper's own synthesis model; because Table I excludes biphonation, intermittency, and fo variation, and Section V's tremor case shows FCN-785 failing to hold M=1 on clean segments, the synthetic corpus has not been shown representative of real…","rationale":"The reader's weakest assumption is exactly the load-bearing one: the synthetic training/evaluation distribution is not shown to be representative of real pathological subharmonic voicing. I found no critical internal error in the architecture, training, or held-out synthetic evaluation; the confusion matrices, SHR/fo analyses, and honest statement of limitations are strengths. But the central practical claim, detecting pathological subharmonic voicing in real recordings, depends on distributional realism, and the paper's own Section V and Conclusions undercut that assumption with concrete evidence: fixed-fo training causes both FCNs to misclassify clean harmonic segments under tremor, and biphonation/intermittency are not in the training set. The 98.9% synthetic accuracy is therefore best interpreted as an upper bound on a narrow, well-controlled distribution, not as clinical detection performance. This supports a CONDITIONAL disposition rather than ACCEPT or REJECT: the methodology is sound and honestly reported, but external validation with broader synthetic coverage or real annotated recordings is needed. The proposed test would directly determine whether the missing synthesis modes are responsible for the observed real-voice failures.","tokens_in":12009,"tokens_out":6223,"duration_ms":68051,"concrete_test":"Synthesize an extended test set with the same kinematic model but add (a) sinusoidal fo contours with 35-Hz peak-to-peak tremor as in Fig. 8, (b) intermittent subharmonic locking with unlocked intervals, and (c) left-right asymmetry to create entrained biphonation; evaluate the already-trained FCN-785 frame-by-frame. If clean M=1 tremor segments are labeled M>1 at a high rate, or accuracy on the extended set falls materially below 98%, the synthetic-corpus realism assumption is the limiting factor and external validation is required before the detection claim can be accepted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (98.9% for FCN-785) is an in-distribution result: the held-out test signals are drawn from the same 21-parameter generator and the same uniform ranges used for training (Section III, Table I). The generator only produces symmetric-fold subharmonic modulation with a fixed fo per signal; it does not include entrained biphonation, intermittent locking/unlocking, or fo tremor. These are not rare edge cases for the intended application: the paper's own Section V case studies encounter unlocked modulation (Fig. 6), suspected biphonation (Fig. 7), and severe vocal tremor (Fig. 8), and the Conclusions explicitly list biphonation, intermittency, and fo variation as missing training ingredients. The tremor case is the sharpest evidence: FCN-785 never outputs M=1 even in the clean harmonic segments at t=0.05-0.13 s and 0.77-0.84 s, which the authors attribute to training with fixed fo. Thus the 98.9% number cannot be read as evidence that the FCN detects pathological subharmonic voicing in general; it is evidence only for the narrow synthetic distribution. The paper is transparent about this gap, but the abstract's 'encouraging outcomes' rests on three qualitative real-voice cases without ground truth, so the practical claim remains conditional.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two fully convolutional neural networks, FCN-401 and FCN-785, to classify the subharmonic period M (with M in {1,2,3,4}) of voice signals, where M=1 denotes normal phonation. The networks are trained exclusively on synthetic signals generated from a kinematic vocal fold model with subharmonic amplitude and frequency modulation, using random draws over the 21 synthesis parameters in Table I. Evaluation on held-out synthetic signals reports 98.1% and 98.9% overall accuracy for FCN-401 and FCN-785, respectively, with an analysis of performance versus subharmonic-to-harmonic ratio and fundamental frequency. The paper then presents three qualitative case studies on sustained /a/ recordings from a disordered voice database, showing encouraging but mixed results, including the failure of both networks on a voice with severe vocal tremor.","tokens_in":1523,"tokens_out":2506,"duration_ms":50581,"significance":"If the synthetic results were to transfer to clinical recordings, the proposed approach would fill a real gap in acoustic voice analysis: automatic and reliable detection of subharmonic phonation, which is known to degrade fundamental-frequency estimators and voice parameter measurements. The paper has notable strengths: the use of a physiologically motivated kinematic vocal fold model rather than simple AM/FM-cycle manipulation, a clearly described and reproducible Monte Carlo training procedure, and a credible analysis of the FCN-401 M=2 weakness in terms of SHR imbalance and window size. The central claim, however, is conditional: the 98% figure is an in-distribution result on the same generative model and parameter ranges used for training, and the real-voice evidence is qualitative and limited. The paper is transparent about its limitations, but the abstract and title do not carry that conditionality.","major_comments":[{"comment":"The reported classification accuracies (98.1% for FCN-401 and 98.9% for FCN-785) are measured on held-out signals drawn from the same synthesis model and the same uniform parameter ranges used for training. This is an in-distribution evaluation: it demonstrates that the networks learn the specific generator distribution, not that they detect pathological subharmonic voicing in general. The abstract's over 98% classification accuracy is presented without this caveat. The paper should either reframe the central claim as applying only to the synthetic distribution, or add out-of-distribution tests (e.g., parameter ranges outside Table I, biphonation, fo variability) that would support broader generalization.","section":"IV and III (Table I), Abstract"},{"comment":"The third case study (severe vocal tremor) directly contradicts the practical claim of encouraging outcomes at the level of segment classification: FCN-785 never outputs M=1 even in clean harmonic segments (t=0.05-0.13 s and 0.77-0.84 s), and FCN-401 also fails to hold M=1 for the entirety of these segments. The authors attribute this to training with fixed fo, which is a reasonable hypothesis, but the failure is not quantified and no ground truth is available for the real recordings. These qualitative cases cannot by themselves establish clinical utility. The paper should either provide quantitative evaluation on a larger real dataset with annotation, or explicitly limit the practical claims to a feasibility demonstration.","section":"V, Fig. 8"},{"comment":"The training dataset contains only symmetric-fold subharmonic modulation with fixed fo per signal, uniform independent parameter draws, and steady-state sustained vowels. The paper's own conclusions list biphonation, intermittency, and fo variation as necessary additions. This is not merely a future-work item: the current synthetic evaluation omits exactly the phenomena that the three case studies encounter (unlocked modulation, suspected biphonation, fo tremor). Consequently, the synthetic accuracy figure cannot be interpreted as a measure of performance on the clinically relevant population of subharmonic voices, and the paper's central claim needs to be scoped accordingly.","section":"III and VI"}],"minor_comments":[{"comment":"There are numerous typographical errors, for example JANURAY in the header, wholistic in the abstract, trainig in Section IV, Sythetic in the Fig. 3 caption, acuracy in Section IV, tranditional in the text before Eq. (23), radition in the lip radiation description, simultor in Section III, constititute in Section V, and entirity in Section V. These should be corrected before submission.","section":"Throughout"},{"comment":"The description of the network variants is difficult to parse: different allocations of its coefficients for the fourth convolution layer does not clearly explain how the same architecture yields different window sizes. A short explanation of how the receptive field is controlled would improve reproducibility.","section":"II, Fig. 1"},{"comment":"The notation r_M(t; phi) is used with subscript M, but the definition of r(t; phi) in Eq. (8) is clear only after reading both equations; consider writing r_M(t; phi) explicitly in the modulation definition and clarifying that M is the subharmonic period of the reference vibration.","section":"III, Eq. (6)"},{"comment":"The SHR analysis is informative, but the paper does not report the SHR computation parameters (e.g., whether the Hamming window is applied to the whole 1-s signal or to analysis frames, and how K_s and K_h are defined precisely). Adding these details would strengthen the reproducibility of the analysis.","section":"IV, Fig. 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid engineering contribution with a transparent methodology, but the central quantitative claim is in-distribution and the real-voice evaluation is qualitative. The authors' own conclusions acknowledge the missing training ingredients. I recommend major revision: the paper should either add out-of-distribution or real-data quantitative evaluation, or carefully scope the claims to the synthetic distribution and reposition the real cases as an illustrative pilot. The technical work on the FCN architecture and the analysis of SHR and fo effects is sound and worth publishing after such revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid proof-of-concept for FCN-based subharmonic-period classification, with an honest synthetic evaluation. The 98.9% number is real but measures learning of the paper's own generator; the real-voice support is qualitative. That is a limitation, but the paper says so itself. I'd send it to peer review.\n\nWhat's actually new: the authors are the first to frame subharmonic period detection as a four-class FCN classification problem (M=1..4), and they embed AM/FM subharmonic modulation into a kinematic vocal fold model rather than applying cycle-wise rescaling post-hoc. That is a meaningful step beyond prior period-doubling-only synthesis. The synthetic evaluation is well done: 4,000 held-out signals, separate confusion matrices, and a credible analysis of why the shorter-window FCN struggles on M=2 (SHR imbalance and spectral resolution). The two-architecture comparison is useful.\n\nSoft spots, in order of size. First, the central accuracy figure is in-distribution: training and test draw from the same 21-parameter generator and the same uniform ranges. The generator excludes biphonation, intermittent locking, and fo variation. Those are not edge cases for the intended use; the paper's own case studies hit all three. The tremor case is the sharpest evidence: FCN-785 never outputs M=1 even in clean harmonic segments, which the authors attribute to fixed-fo training. So the 98.9% does not transfer as-is to real voices. The paper is transparent about this in the conclusions, but the abstract's 'encouraging outcomes' rests on three qualitative recordings with no ground truth. Second, there is no code or data release, which makes the result harder to build on. The synthesis model is described in detail, so reimplementation is possible, but artifacts would help. Third, the biphonation case is interesting but speculative; the networks were not trained for it, and the 'suspected' locking is interpreted generously.\n\nNone of this is fatal. The claims are appropriately scoped, the synthesis is more realistic than prior work, and the failure analysis is honest. This is a conditional accept: it needs external validation on real pathological voices or a clearly stated in-distribution claim, and ideally artifacts.\n\nWho should read it: anyone working on acoustic analysis of type-2 voices, or on DL for pathological speech. I'd bring it to our reading group, and I'd cite it as the current FCN baseline for subharmonic-period classification. Yes to peer review; the correct outcome is likely a revision with more real-voice data or a narrower title.","headline":"A genuinely new FCN approach to subharmonic-period classification with an honest in-distribution synthetic evaluation; real-voice evidence is qualitative, but the paper deserves a serious referee.","tokens_in":12879,"tokens_out":2350,"would_cite":true,"duration_ms":23852,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that fully convolutional networks trained on a synthetic corpus of subharmonic voices can assign each 2-ms snapshot a subharmonic period in {1,2,3,4} with over 98% accuracy on held-out synthetic signals, with case studies…","keywords":["subharmonic phonation","voice disorders","fully convolutional network","subharmonic period classification","synthetic voice corpus","kinematic vocal fold model","sustained vowel analysis","acoustic voice analysis"],"falsifier":"If a laryngoscopy-verified clinical corpus showed that FCN-785 mislabels clean harmonic segments under fo tremor at rates no better than chance while synthetic accuracy stays near 99%, the transfer assumption would be refuted; the paper's own tremor case already points in this direction but with only a single recording.","tokens_in":11815,"feed_emoji":"🎙️","tokens_out":7071,"duration_ms":61323,"temperature":0.7,"pith_summary":"The paper seeks a reliable way to detect subharmonic phonation in voice recordings, where the vocal folds repeat a vibration pattern every two, three, or four glottal cycles instead of every cycle. Existing pitch and roughness measures either miss subharmonics or only handle period-doubling, and subjective spectrogram inspection remains the standard for more complex cases. The proposal is to train fully convolutional neural networks on a large Monte Carlo corpus of synthesized subharmonic vowels so that each network labels short snapshots with the subharmonic period M in {1,2,3,4}. On held-out synthetic signals the longer-window network FCN-785 reaches 98.9% overall accuracy, and every one of the four classes is above 98.5% conditional accuracy. On three real pathological /a/ recordings the networks track period-tripling modulation, partially detect locked biphonation, and mislabel clean segments under severe fo tremor, which the paper reads as evidence that the approach works but the synthetic training distribution needs more realism.","feed_headline":"Neural net labels vocal-fold subharmonic periods 98.9% of the time","feed_subtitle":"Trained only on simulated voices, the network separates normal cycles from subharmonic periods 2-4; case studies expose remaining gaps.","key_machinery":"The load-bearing object is the synthesis-driven FCN pair. On the synthesis side, the kinematic vocal fold model defines fold displacement from a reference vibration r(t; phi), and subharmonic behavior is realized by amplitude and frequency modulation of that reference, so the modulation passes through vocal-fold collision, nonlinear glottal flow, and vocal-tract acoustics rather than being superimposed on the output waveform. On the classifier side, a five-layer fully convolutional network with four max-pooling layers reduces the output rate by a factor of 8 to emit one subharmonic-period estimate per 2-ms snapshot, using sigmoid outputs instead of softmax because the label set is not exhaustive of pathological voices. The training corpus randomizes 21 synthesis parameters, including fo, modulation extents and phases, glottal geometry, tract lengths and areas, lung pressure, and noise level, and the two networks share the same coefficient count while differing in filter length and count at the fourth convolutional layer.","core_discovery":"The central claim is that subharmonic period estimation can be cast as a per-snapshot classification problem solved by a fully convolutional network operating on raw, mean-variance-normalized 8 kHz audio, and that a network trained exclusively on synthesized modulation-type subharmonics can classify M=1 through M=4 with over 98% accuracy on held-out synthetic signals and can partially transfer to real clinical recordings. The paper establishes this by generating training data with a kinematic vocal-fold model coupled to a wave-propagation vocal tract, introducing subharmonics through amplitude and frequency modulation of the fold reference vibration, and training two FCN variants: FCN-401 with a 50.1-ms window and FCN-785 with a 98.1-ms window. The longer window performs better on low-fo and low-SHR signals, while the shorter window tracks transitions between subharmonic states slightly better; both ignore weak unlocked modulation, but both mislabel clean harmonic segments in a severe vocal-tremor case because the training corpus fixes fo per signal.","pith_inferences":["A testable extension that the paper does not run is to add slow sinusoidal fo drift to the synthetic corpus; if tremor-case mislabeling disappears while synthetic accuracy stays above 98%, fixed-fo training would be identified as the cause rather than the FCN architecture.","Because the output layer uses sigmoids rather than softmax, the per-class probabilities can be read as confidence values, so a clinical system could flag low-confidence snapshots for human review, a use the paper does not develop.","The same FCN pipeline is input-duration agnostic, so a straightforward extrapolation is to evaluate it on connected speech once the training corpus includes fo variation; the paper only demonstrates sustained vowels, so this is an extension beyond its reported results.","If a larger clinical dataset with simultaneous high-speed videoendoscopy ground truth were available, one could quantify the synthetic-to-real transfer gap in terms of conditional per-M accuracy instead of relying on three case studies."],"forward_implications":["If the synthetic accuracy transfers to clinical use, subharmonic detection can be automated with a single feed-forward network, giving clinicians a numerical stream of M labels over time instead of requiring subjective spectrogram reading.","The M labels can feed downstream measures such as NSH and SHR and can steer fundamental-frequency estimators away from the subharmonic period, reducing the underestimation of pathology severity that the introduction identifies.","The 98.1% versus 98.9% comparison indicates that the 98-ms window should be preferred for sustained vowels, while the 50-ms window is preferable when transition timing matters, so the two networks cover complementary use cases.","Training on fixed-fo sustained vowels is a known source of failure, as the tremor case shows clean harmonic stretches being mislabeled, so extending the corpus with fo variation, intermittency, and biphonation is the paper's own predicted path to a dependable detector.","Weak unlocked modulation is mostly ignored by both networks, which is desirable if the goal is to report only strongly locked subharmonic periods rather than every cycle-to-cycle irregularity."],"supporting_citations":[{"why":"Provide the kinematic vocal fold model that generates the glottal area and displacement waveforms used to synthesize subharmonic phonation.","marker":"[30, 31]"},{"why":"Supply the two-port wave-propagation vocal tract model that turns the glottal flow into the radiated acoustic signal used for training and evaluation.","marker":"[32, 33]"},{"why":"Demonstrates the fully convolutional architecture for fundamental frequency estimation that motivates the FCN design and the 8 kHz sampling rate.","marker":"[23]"},{"why":"Provides the KayPENTAX Disordered Voice Database recordings used in the three sustained-vowel case studies.","marker":"[48]"},{"why":"Introduces the subharmonic-to-harmonic ratio, the prior numerical approach whose limitation to period-doubling the paper aims to overcome.","marker":"[17]"},{"why":"Strengthens Sun's SHR method with improved initial fo estimation but remains restricted to period-doubling, motivating classification of M>2.","marker":"[18]"},{"why":"Establishes the type-2 signal classification and the recommendation for subjective inspection that the proposed automated detector is meant to replace.","marker":"[14]"},{"why":"Shows that machine-learning pitch estimators are robust to subharmonics, which motivates training a network to find subharmonics rather than ignore them.","marker":"[12]"}],"fun_headline_variants":["FCN trained on synthetic voices labels subharmonic periods 98%","Deep FCN identifies subharmonic voicing from synthetic training","Convolutional net trained on simulated voices detects subharmonics","FCN labels subharmonic periods in voice with 98% accuracy","Synthetic-trained FCN finds real subharmonic periods in voice"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 21-parameter Monte Carlo corpus of symmetric-fold, modulation-type subharmonic sustained vowels, each with a fixed fo and a steady state, is representative enough of real pathological subharmonic phonation that the high synthetic accuracy transfers to clinical recordings.","fun_headline_variants_meta":{"raw":{"variants":["FCN trained on synthetic voices labels subharmonic periods 98%","Deep FCN identifies subharmonic voicing from synthetic training","Convolutional net trained on simulated voices detects subharmonics","FCN labels subharmonic periods in voice with 98% accuracy","Synthetic-trained FCN finds real subharmonic periods in voice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000611,"raw_usage":{"total_tokens":2810,"prompt_tokens":881,"completion_tokens":1929,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":1842}},"tokens_in":497,"tokens_out":1929,"duration_ms":12286,"temperature":1.0,"reasoning_tokens":1842,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:10:17.790392+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a laryngoscopy-verified clinical corpus showed that FCN-785 mislabels clean harmonic segments under fo tremor at rates no better than chance while synthetic accuracy stays near 99%, the transfer assumption would be refuted; the paper's own tremor case already points in this direction but with only a single recording.","supporting_citations":[{"cited_title":"Fully-convolutional net- work for pitch estimation of speech signals,","cited_arxiv_id":null,"evidence_quote":"Demonstrates the fully convolutional architecture for fundamental frequency estimation that motivates the FCN design and the 8 kHz sampling rate."},{"cited_title":"Disordered V oice Database and Program [Model 4337],","cited_arxiv_id":null,"evidence_quote":"Provides the KayPENTAX Disordered Voice Database recordings used in the three sustained-vowel case studies."},{"cited_title":"A pitch determination algorithm based on subharmonic-to-harmonic ratio,","cited_arxiv_id":null,"evidence_quote":"Introduces the subharmonic-to-harmonic ratio, the prior numerical approach whose limitation to period-doubling the paper aims to overcome."},{"cited_title":"Acoustic tracking of pitch, modal, and subhar- monic vibrations of vocal folds in Parkinson’s Disease and Parkinsonism,","cited_arxiv_id":null,"evidence_quote":"Strengthens Sun's SHR method with improved initial fo estimation but remains restricted to period-doubling, motivating classification of M>2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the type-2 signal classification and the recommendation for subjective inspection that the proposed automated detector is meant to replace."},{"cited_title":"Evaluation of machine-learning pitch estimation algorithms,","cited_arxiv_id":null,"evidence_quote":"Shows that machine-learning pitch estimators are robust to subharmonics, which motivates training a network to find subharmonics rather than ignore them."}],"review_version":1}