{"id":"2935c2ba-2ebf-4072-9970-c99bbb8162b4","arxiv_id":"2506.19253","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A stimulus-aware harmonic filterbank with prominence peak picking estimates the F0 contour of frequency-following responses more accurately than autocorrelation.","lead":"The paper introduces a new way to track the pitch of brain responses to speech, using a filterbank that sums harmonic energy near the known stimulus pitch. In tests on 16 people, it reduced tracking errors by 9 to 47 percent compared to the standard method.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post hoc selection of harmonic count K on the test data inflates HAS-PR's reported RMSE gains over ACF; the central claim needs validation with a fixed or cross-validated K.","rationale":"The reader's weakest assumption (50 Hz search window) is a real limitation for generalizability, but it is not the most load-bearing issue for the central claim, because the stimulus F0 prior is an inherent and defensible part of the FFR setting. The more serious problem is the post hoc selection of K. Since K is chosen by minimizing the very RMSE metric that is then reported as evidence, the 8.8–47.4% improvements are optimistically biased. This is a classical in-sample selection effect: any method given a free parameter tuned on the test set will appear better than a fixed baseline. The paper's own statistical tests are circular in this respect. The proposed concrete test—using a fixed K=2 or a cross-validated K, and matching the ACF baseline to receive the same prior—directly addresses whether the method's advantage is real. If the advantage holds, the conditional acceptance is justified; if not, the claim should be weakened. The 50 Hz window should also be stressed by testing stimuli with larger F0 deviations, but that is secondary. We therefore agree with the reader's overall conditional verdict, but our primary concern differs from their stated weakest assumption.","tokens_in":9083,"tokens_out":6008,"duration_ms":57883,"concrete_test":"Fix K=2 for all four conditions (or select K via leave-one-subject-out cross-validation on the 16 subjects), rerun Table 1 using an ACF baseline restricted to the same 50 Hz window around the stimulus F0, and recompute the paired t-tests in §3.2. If the mean RMSE reductions over ACF on Female Happy drop below the reported 47.4% by a large margin or become non-significant, the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim (Abstract) is an 8.8–47.4% mean RMSE reduction over ACF on the four stimuli. This claim rests on Table 1, but Section 3.2 reveals that the number of harmonics K in Eq. (3) was chosen per stimulus by taking the value that 'resulted in the minimum RMSE in Table 1'—K=4 for Male Sad, K=2 for the others. Selecting K on the same test data used for evaluation makes the reported RMSE values in-sample optima rather than unbiased performance estimates. The paired t-tests in §3.2 that compare HAS-PR with the optimal K against ACF or K=1 are likewise post hoc and do not estimate how the method would perform with a fixed parameter choice. The comparison is further confounded by the under-specified ACF baseline: it is not stated whether ACF was allowed the same 50 Hz stimulus-F0 search window that HAS-PR uses (§2.4). If not, part of the advantage could derive from the prior, not the harmonic summation or prominence peak-picking. Therefore the quantitative claim that HAS-PR 'outperformed' ACF is not yet established; it needs a re-analysis with K fixed or cross-validated and the ACF baseline matched. This is a load-bearing concern because the entire reported improvement could shrink or vanish when these confounds are removed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HAS-PR, a frequency-domain pitch-tracking method for the Frequency Following Response (FFR). The method builds a filterbank of harmonic templates (Eq. 3), aggregates harmonic amplitudes (Eq. 7), and selects the most prominent peak within a ±50 Hz window centered on the known stimulus F0 (Section 2.4). The authors evaluate HAS-PR against ACF and several speech-pitch algorithms on FFRs from 16 participants to four natural speech stimuli, and report that HAS-PR reduces mean RMSE relative to ACF by 8.8% to 47.4% depending on stimulus. They also report Gross Pitch Error and RMSE20 metrics, and analyze performance versus number of averaged sweeps.","tokens_in":9300,"tokens_out":2401,"duration_ms":27133,"significance":"If the reported gains are unbiased, the method could be a practically useful advance for FFR F0 tracking, which is currently dominated by ACF. The paper is one of the first to explicitly adapt harmonic-structure-based pitch tracking to FFRs, and the use of prominence-based peak picking to counter spectral tilt is a sensible idea. The inclusion of GPE and RMSE20 metrics, along with the sweep-count analysis, is a strength. However, the significance hinges on whether the comparison to ACF is fair and whether the reported improvements are in-sample artifacts of parameter selection.","major_comments":[{"comment":"The number of harmonics K was selected post hoc on the same test data used for the performance evaluation. The text in Section 3.2 states that the optimal K 'resulted in the minimum RMSE in Table 1' (K=4 for Male Sad, K=2 for the others). This makes the reported RMSE values in-sample optima rather than unbiased performance estimates. The paired t-tests in Section 3.2 comparing HAS-PR with the optimal K against ACF or K=1 are therefore also post hoc. The central claim that HAS-PR 'outperformed' ACF is not established by this analysis. Please re-run the evaluation with a fixed K chosen a priori, or with K selected by cross-validation on a development subset, and report the corresponding RMSE values and test statistics.","section":"§3.2, Table 1 and Eq. (3)"},{"comment":"The ACF baseline is described only as 'the Autocorrelation Function' and the implementation details are not given. In particular, it is not stated whether ACF was given the same ±50 Hz prior centered on the stimulus F0 that HAS-PR uses, nor whether the same frame length, step size, and band-pass filtering were applied. If ACF was not given the stimulus-F0 prior, part of the reported RMSE reduction could derive from the prior rather than from harmonic summation or prominence peak picking. Please specify the ACF implementation completely, and also report an ACF variant that uses the same F0 search window so that the comparison isolates the algorithmic contribution.","section":"§2.4 and Table 1"},{"comment":"The evaluation may be partly circular because the error metric is the difference between the estimated response F0 and the known stimulus F0, while the algorithm restricts its search to a ±50 Hz interval centered on that same stimulus F0. This design guarantees that any estimate is within 50 Hz of the reference, and it mixes the effect of a strong prior with the effect of the harmonic-sum filterbank. The paper should quantify how much of the error reduction is due to the search window itself. A concrete test: compare HAS-PR with a version of ACF that is given the same 50 Hz window, and also report results when the window is widened (e.g., ±100 Hz or a full 80-500 Hz search) to show that the algorithm still tracks accurately without the tight prior.","section":"§2.4 and §3.1"},{"comment":"The statistical analysis is underpowered for the strength of the claims. Paired t-tests are performed on the same data used to select K, and multiple comparisons (three sets of comparisons across four stimuli) are made without correction. The p-value of 0.078 for the FS condition in the K vs. K=1 comparison is reported as non-significant, but the text does not discuss whether the pattern of results would survive a multiple-comparison correction. Please provide a corrected analysis or clearly state the number of comparisons and the correction procedure.","section":"§3.2"}],"minor_comments":[{"comment":"The definition of 'prominence' is not given explicitly; the text only cites Kirmse & de Ferranti (2017). Please specify the algorithm used for prominence computation, as the exact definition affects reproducibility.","section":"§2.4"},{"comment":"The matrix H is defined via the DFT of each filter, but the size of H and the normalization of the DFT are not stated. Please clarify how the filters are normalized and whether the summation in Eq. (7) is over the positive-frequency bins only.","section":"§2.4, Eq. (5)-(7)"},{"comment":"The stimulus F0 extraction uses the median of several MATLAB pitch methods with manual correction of octave errors. This is a reasonable approach, but the manual correction step is subjective; please state how many frames were corrected and whether the correction was blinded to the subsequent FFR analysis.","section":"§2.2"},{"comment":"The table reports means across participants but not standard deviations, even though the text says 'mean and standard deviation' are presented. Please add the SD values or remove the claim.","section":"Table 1"},{"comment":"In Figure 3, the performance is shown versus the number of averaged sweeps, but it is unclear whether the number of sweeps was varied per participant and whether the results are averaged over participants or shown as medians. Please clarify the aggregation procedure.","section":"§3.1"},{"comment":"The phrase 'within each response and stimulus F0 contour pair' in the abstract is awkward; consider rephrasing to 'across pairs of response and stimulus F0 contours.'","section":"Abstract and §4"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a real need in FFR analysis, but the evaluation design has two entangled problems: in-sample selection of K and an under-specified ACF baseline. These are fixable with a re-analysis, so I recommend major revision rather than rejection. The novelty is modest—the method is a fairly direct application of harmonic summation with a known-F0 prior—but it is a reasonable engineering contribution if the comparison is made fair. I would also note that the use of the stimulus F0 as a prior is not inherently disqualifying, but the paper should be explicit that the method estimates the response F0 under the assumption that it remains near the stimulus F0, which is physiologically plausible but should be tested or acknowledged as a limitation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read the paper. The core idea is a genuine adaptation: they take harmonic amplitude summation, a known speech-pitch method, and tailor it to FFRs by exploiting the fact that the stimulus F0 is known. That lets them center the search on the stimulus F0 and use prominence peak-picking instead of the highest peak. The motivation is clear, and the evaluation is more thorough than most in this subfield: 16 subjects, four natural speech stimuli spanning a wide F0 range, and comparisons against ACF, Praat, PEFAC, HarmoF0, and a Bayesian tracker. They also report both RMSE and gross pitch error, which is good practice.\n\nThe paper is genuinely new as the first FFR-specific harmonic-structure F0 estimator. It also seems to work better than ACF on the female-talker stimuli where ACF struggles, which is a clinically relevant result if it holds up.\n\nThe problem is the main quantitative claim. Section 3.2 reveals that the number of harmonics K was selected per stimulus by taking the value that produced the minimum RMSE in the same Table 1 used to report the final results. That makes the reported RMSE values in-sample optima. The paired t-tests then compare this optimally tuned HAS-PR against ACF and other methods. That is not a fair test of the algorithm, and the improvement could shrink or disappear with a fixed or cross-validated K. The ACF baseline is also under-specified: it is not stated whether ACF received the same 50 Hz stimulus-F0 search window that HAS-PR uses. If not, part of the advantage might come from the prior rather than the harmonic summation or prominence picking. On top of that, the 50 Hz window assumption is untested—if the response F0 deviates by more than that from the stimulus F0, the method would fail, and the paper does not characterize that failure mode. Finally, the improvement over ACF is not statistically significant for the Male Sad condition, so the benefit is not uniform.\n\nThese are not fatal objections. The method is plausible and the paper does a lot right. But the central claim of an 8.8–47.4% RMSE reduction is not yet supported. The authors need to re-run the analysis with K chosen in a parameter-free way (fixed across stimuli, or cross-validated on held-out subjects or frames), match the ACF baseline by giving it the same prior, and report the contribution of the stimulus-F0 prior separately. I would send this to peer review, but my recommendation would be major revision. The paper is worth engaging with; I would not cite it yet until the parameter-selection issue is resolved.","headline":"A useful adaptation of harmonic summation to FFR pitch tracking, but the headline RMSE gains over ACF are inflated by post hoc selection of the harmonic count K and an un-matched ACF baseline.","tokens_in":9892,"tokens_out":1643,"would_cite":false,"duration_ms":17171,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A stimulus-aware harmonic filterbank estimates F0 contours in the Frequency Following Response more accurately than autocorrelation, cutting average RMSE by 8.8–47.4%.","keywords":["Frequency Following Response","fundamental frequency tracking","harmonic amplitude summation","pitch tracking","auditory brainstem response","prominence peak picking","autocorrelation function"],"falsifier":"Simulate a synthetic FFR frame exactly as in Eq. (1) with a true F0 set 60 Hz above the stimulus F0 and K=2 harmonics added, then run HAS-PR and check whether the selected peak is near the true F0 or falls back to an alias inside the ±50 Hz window; a correct estimate outside the window would refute the method's search constraint as load-bearing.","tokens_in":1822,"feed_emoji":"🧠","tokens_out":2091,"duration_ms":84481,"temperature":0.7,"pith_summary":"The paper claims that pitch tracking in the Frequency Following Response (FFR) — the brain's scalp-recorded response to sound — is more accurate when the algorithm knows the stimulus pitch in advance and searches only near it. It introduces the Harmonic Amplitude Summation (HAS) filterbank, which sums amplitudes at candidate F0 and its first few harmonics while suppressing non-harmonic frequencies, and combines it with prominence-based peak picking (HAS-PR). On FFRs from 16 normal-hearing listeners responding to four natural-speech \"balloon\" stimuli with F0 ranging from 89 Hz to 452 Hz, the paper reports that HAS-PR lowers root-mean-square error between stimulus and response F0 contours by 8.8% to 47.4% relative to the standard autocorrelation method, with larger gains for higher-pitched female voices. If correct, this gives auditory researchers a more reliable way to trace how the brainstem tracks pitch in natural speech, with potential value for clinical screening.","feed_headline":"Harmonic filterbank tracks pitch in brain responses more accurately","feed_subtitle":"Speech-evoked FFR pitch error drops 8.8% to 47.4% versus autocorrelation in 16 listeners.","key_machinery":"The central object is the Harmonic Amplitude Summation (HAS) filterbank: a bank of synthetic periodic filters, one for each F0 candidate at 1 Hz resolution between 80 Hz and 500 Hz, each containing a fundamental plus K harmonics (optimally K=4 for the low-F0 male speech and K=2 for the high-F0 female speech). Each filter's magnitude DFT is multiplied pointwise with the response frame's magnitude DFT and summed, so a candidate F0 scores high only when energy lines up at its fundamental and its harmonics. The second mechanism is prominence-based peak picking: instead of taking the tallest peak in the score curve, the algorithm takes the peak that stands out most against its local surroundings, which makes the estimate insensitive to the overall spectral downward slope produced by noise. The F0 search window is centered on the known stimulus F0 of the same aligned frame, with a ±50 Hz threshold, which prevents octave errors and lets the method keep frequency resolution.","core_discovery":"The central claim is that harmonic structure, not just periodicity, carries reliable pitch information in FFRs, and that a filterbank can exploit it once the search is constrained by the known stimulus F0. The authors argue that previous speech and music pitch trackers fail on FFRs because they search too wide a range, weight harmonics logarithmically, include too many harmonics, and pick the highest rather than the most prominent peak. HAS-PR addresses each issue: it uses a DFT on a linear scale, aggregates only the first 2–4 harmonics, restricts candidate F0s to a 50 Hz window around the stimulus F0 of the time-aligned frame, and selects the most prominent peak in the summed spectrum. In their recordings, this reduces RMSE relative to ACF by 8.8% for Male Sad, 31.2% for Male Happy, 37.8% for Female Sad, and 47.4% for Female Happy stimuli, and also lowers gross pitch error in all conditions. The method is also reported to have lower RMSE than Praat, PEFAC, HarmoF0, and a Bayesian pitch tracker in all four stimulus conditions.","pith_inferences":["If the response F0 ever deviates by more than 50 Hz from the stimulus F0, HAS-PR will be blind to it; testing this boundary by comparing HAS-PR with a broad-range pitch tracker on recordings where the pitch percept is shifted (for example, via a missing fundamental or masking) would directly probe the method's main assumption.","The prominence-based peak selector could transfer to speech pitch trackers as a detrending-free way to handle spectral tilt, an extension the paper only hints at when discussing why the highest peak is unreliable for FFRs.","The condition-dependent optimal harmonic count (K=4 for low-F0 stimuli, K=2 for high-F0 stimuli) suggests an adaptive harmonic-count rule might improve accuracy across speakers and pitches, since the paper fixes K per condition rather than adapting it frame by frame.","Because the method needs only the stimulus F0 contour and a DFT, it could be implemented in near-real-time for neurofeedback or closed-loop auditory experiments, although the paper does not address real-time operation."],"forward_implications":["If HAS-PR works as reported, researchers can extract F0 contours from FFRs evoked by natural speech with lower error than ACF, especially for high-pitched female talkers where the response harmonic structure is weaker.","Because the method uses only the first few harmonics and a constrained search, it is designed to avoid octave errors, a failure mode that the paper reports affects several speech-based baseline methods.","The method's reliance on the known stimulus F0 means it can be applied to any FFR paradigm in which the stimulus F0 contour is available, not just the four emotional speech stimuli tested here.","The paper's sweep-count comparisons indicate that the advantage of HAS-PR over ACF persists as the number of averaged response sweeps changes, suggesting the method could support shorter or more flexible recording protocols.","Since HAS-PR operates on a per-frame DFT, it can be combined with existing FFR preprocessing such as band-pass filtering and coherent averaging, making it a drop-in replacement for autocorrelation in current pipelines."],"supporting_citations":[{"why":"Supplies the autocorrelation function method that the paper uses as the standard FFR pitch-tracking baseline and compares against.","marker":"Rabiner (1977)"},{"why":"PEFAC, a harmonic-summation pitch estimator whose log-scale normalization the paper adapts and replaces with a linear-scale DFT filterbank.","marker":"Gonzalez & Brookes (2014)"},{"why":"HarmoF0, a neural harmonic-summation estimator that the paper includes as a speech-pitch baseline and as a source of the logarithmic harmonic summation idea.","marker":"Wei et al. (2022)"},{"why":"Bayesian pitch tracker, included as a robust speech-pitch baseline in the comparison, and noted for performing reasonably on one female stimulus but not others.","marker":"Shi et al. (2019)"},{"why":"Supplies the prominence-based peak-picking algorithm that the paper uses instead of highest-peak selection.","marker":"Kirmse & de Ferranti (2017)"},{"why":"Provides the 16-subject FFR dataset, the four emotional speech stimuli, and the recording protocol used to evaluate the method.","marker":"Karimi Boroujeni et al. (2024)"},{"why":"Source of the Emotional Speech Database from which the \"balloon\" stimuli were derived.","marker":"Zhou et al. (2022)"},{"why":"Earlier FFR work that constrains autocorrelation analysis to a neighborhood around the stimulus pitch period, the prior-knowledge concept that HAS-PR extends.","marker":"Jeng et al. (2011)"}],"fun_headline_variants":["Harmonic filterbank cuts FFR pitch error up to 47%","Brain pitch tracking: harmonic filter outperforms autocorrelation","FFR pitch: harmonic amplitude summation beats ACF by up to 47%","New filterbank improves pitch tracking in brain responses","Pitch from brain waves: harmonic filter reduces error by 47%"],"cache_read_input_tokens":12032,"weakest_assumption_plain":"The method assumes the neural F0 of the response never wanders more than 50 Hz away from the stimulus F0; the reported accuracy only covers frames where that window contains the true response pitch, and the paper does not test what happens when this assumption is violated.","fun_headline_variants_meta":{"raw":{"variants":["Harmonic filterbank cuts FFR pitch error up to 47%","Brain pitch tracking: harmonic filter outperforms autocorrelation","FFR pitch: harmonic amplitude summation beats ACF by up to 47%","New filterbank improves pitch tracking in brain responses","Pitch from brain waves: harmonic filter reduces error by 47%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000717,"raw_usage":{"total_tokens":3301,"prompt_tokens":1102,"completion_tokens":2199,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":718,"completion_tokens_details":{"reasoning_tokens":2110}},"tokens_in":718,"tokens_out":2199,"duration_ms":15140,"temperature":1.0,"reasoning_tokens":2110,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:06:51.752154+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a synthetic FFR frame exactly as in Eq. (1) with a true F0 set 60 Hz above the stimulus F0 and K=2 harmonics added, then run HAS-PR and check whether the selected peak is near the true F0 or falls back to an alias inside the ±50 Hz window; a correct estimate outside the window would refute the method's search constraint as load-bearing.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the autocorrelation function method that the paper uses as the standard FFR pitch-tracking baseline and compares against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"PEFAC, a harmonic-summation pitch estimator whose log-scale normalization the paper adapts and replaces with a linear-scale DFT filterbank."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the prominence-based peak-picking algorithm that the paper uses instead of highest-peak selection."}],"review_version":1}