{"id":"79424425-c79a-48a3-bc81-6f8cf4754ac8","arxiv_id":"2412.08762","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A Bayesian curve registration method for functional mixed membership models recovers distinct EEG spectral features and links peak alpha frequency to age and ASD status.","lead":"This paper introduces a statistical method that aligns brain-wave (EEG) curves from different children while allowing each child's curve to be a mix of two underlying patterns. It uses this to measure how the alpha brain rhythm peak shifts with age and autism diagnosis.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The distinct-recovery claim rests on fixing rho=0 after the sampler found a competing mode where the 1/f shape acquires a peak; the simulation evidence also discards non-converged runs and rescales feature 2, so the central claim is not yet supported.","rationale":"The reader's weakest assumption is that a single shared warp shape with scalar rescaling is sufficient for both features, with rho fixed at 0 in the case study. My strongest concern is closely related but sharper and more direct: the paper itself reports that estimating rho leads to a bimodal posterior and that the rho=1 mode causes the 1/f feature to display a peak, which is exactly the regime that would undermine the distinct-recovery claim. The decision to fix rho=0 is therefore not a harmless modeling simplification but a response to an identifiability failure that is load-bearing for the abstract's central claim. I also note that the simulation study, which is supposed to validate estimation of rho and the feature shapes, excludes datasets that converged to low-likelihood modes and standardizes feature 2 before computing error, so the reported accuracy is not representative of the model's behavior on all simulated datasets. These issues justify keeping the paper conditional rather than accepting its claims as established. The requested concrete test would settle whether the distinct-recovery and PAF regression results are robust when rho is not fixed, or whether they are artifacts of the enforced constraint.","tokens_in":16498,"tokens_out":4006,"duration_ms":44146,"concrete_test":"Re-run the case-study model with rho estimated rather than fixed at 0, using multiple random initializations and a parallel-tempering or other multimodal-sampling scheme to explore both posterior modes; report the posterior distribution of rho and the resulting feature shapes under each mode. If the high-mass mode remains near rho=1 and feature 1 then contains a peak, the claim that the 1/f noise is recovered distinctly from the alpha peak is not supported; if the rho near 0 mode dominates and feature 1 remains peak-only, the concern is resolved. Additionally, recompute the PAF regression summaries under each mode to see whether the age and clinical effects are stable.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the model recovers the 1/f pink noise feature 'distinctly' from the alpha peak and quantifies age/clinical effects on PAF. The paper's own Section 3.2 admits that the posterior for rho (Eq. 4) is bimodal, with modes near 0 and near 1; when the sampler converges to the rho=1 mode, the 1/f shape itself acquires a small peak that can be warped, creating identifiability problems. The authors then fix rho=0 in the case study rather than estimating it. Thus the distinctness of the two recovered features is not a free finding but is enforced by setting the second feature's warping to zero. If the true warps differ in shape between the two features, or if the rho=1 mode reflects genuine data structure, the recovered feature shapes and any PAF regression are biased. The simulation study does not resolve this: Section 3.1 discards datasets that converge to low-likelihood modes (e.g., only 26 of 40 datasets for N=50), and feature 2 accuracy is reported only after standardizing both the true and estimated functions via Eqs. 25-26, which removes amplitude mismatches. Therefore, the evidence that the model can estimate two distinct features and a warp scale is conditional on favorable runs and a post-hoc rescaling, leaving the central claim insufficiently supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian hierarchical curve registration model for functional mixed membership data with two features. Individual curves are modeled as mixtures of two population-level shapes, each subject has a time-transformation function modeled through B-splines with the Jupp transformation for monotonicity, and covariates enter through a regression model on the unconstrained warping coefficients. The method is applied to resting-state EEG spectral data from children with ASD and typically developing children. The central claims are that the model recovers the 1/f pink-noise feature distinctly from the alpha-band peak and that the regression component quantifies the effects of age and clinical designation on peak alpha frequency location.","tokens_in":16824,"tokens_out":3779,"duration_ms":41454,"significance":"If the central claims were fully supported, the paper would make a useful contribution by extending Bayesian curve registration to mixed membership models and by providing a principled way to regress warping functions on covariates. The Jupp-transformation framework for constrained warping and the semi-supervised identifiability strategy are appealing ideas. However, the current empirical evidence is substantially conditional: the simulation discards a non-negligible fraction of datasets that converge to a low-likelihood mode, the case study fixes the key warp-scaling parameter rho at 0 after observing bimodality, and the reported errors for feature 2 are computed after a post-hoc standardization that removes amplitude information. Because these decisions directly affect the two headline claims, the paper needs additional work before the conclusions can be accepted as stated.","major_comments":[{"comment":"The simulation study reports results only for datasets that converged to the high-likelihood mode, discarding 14 of 40 datasets for N=50, 11 of 40 for N=100, and 10 of 40 for N=150. The paper states that the remaining runs converged to biased values of rho and inaccurate shapes, but it does not report the corresponding error metrics over all datasets or provide diagnostics showing that the discarded runs are merely convergence failures rather than evidence of posterior multimodality. Because the simulation is the primary quantitative support for the method's ability to recover the true features and rho, reporting only favorable runs makes the MSE and R-MISE results conditional on the mode selected. The authors should report results including all datasets, or remedy the multimodality with better initialization, tempering, or additional identifiability constraints, and show that the conclusions are robust to this choice.","section":"Section 3.2, Eq. (4)"},{"comment":"The central claim that the model 'recovers the 1/f pink noise feature distinctly from the peak in the alpha band' is weakened by the decision to fix rho=0 in the case study. The paper itself reports that the posterior for rho is bimodal with modes near 0 and near 1, and that at the rho=1 mode the 1/f shape acquires a small peak that can be warped. Setting rho=0 means the second feature is not warped at all, so the distinctness of the recovered 1/f shape is enforced by modeling choice rather than estimated from the data. This is load-bearing because the scientific interpretation of the recovered features and the subsequent peak-alpha-frequency regression depends on the warping specification. The authors should provide a sensitivity analysis over plausible values of rho, or develop an identifiability strategy that allows rho to be estimated reliably, and discuss how the recovered shapes and PAF regression change.","section":"Section 3.1, Eqs. (25)-(26)"},{"comment":"The reported R-MISE for feature 2 is computed after standardizing both the true and estimated functions using Eqs. (25)-(26), which removes amplitude and level differences. The paper acknowledges a mismatch in amplitude for feature 2 and introduces the rescaling as a post-processing step. This means the simulation does not assess whether the model recovers the amplitude of the second feature, which is scientifically meaningful in the EEG application (e.g., peak prominence and group differences in loadings). The authors should report errors without the standardization, or justify the standardization as a principled identification condition rather than a correction applied after seeing the results.","section":"Section 3.1, Eq. (4)"},{"comment":"The simulation study fixes the generative value rho=0.4, while the case study fixes rho=0. No simulation is conducted for the rho=0 regime actually used in the data analysis, and no simulation is conducted for the bimodal scenario documented in the case study. Consequently, the simulation provides no direct evidence that the model works in the setting where the paper claims its main empirical result. The authors should include a simulation scenario with rho=0 and ideally a scenario with a bimodal posterior to examine whether the proposed workflow (including the decision to fix rho) yields unbiased estimates of the features and the PAF regression.","section":"Section 3.2"},{"comment":"The semi-supervised labeling of subjects in the case study is based on the data: subjects with the most prominent peaks are assigned to feature 2 and subjects with the poorest linear fit are assigned to feature 1. Because these same labeled subjects are then used to identify the features, this procedure can create a circularity that partly explains why the recovered feature 1 looks like 1/f noise and feature 2 looks like an alpha peak. The paper should report how sensitive the recovered shapes and the PAF regression are to the choice and number of labeled subjects, or use a labeling criterion that does not depend on the outcome being modeled.","section":"Section 3.2"}],"minor_comments":[{"comment":"The abstract contains the placeholder 'Keywords and phrases: First keyword, second keyword.' and should be replaced with actual keywords.","section":"Section 3"},{"comment":"The text says the code is available in the package crMM 'available on Github (put link)'; the link is missing and must be provided for reproducibility.","section":"Section 2.5.2"},{"comment":"The posterior predictive fit in Eq. (18) applies the same warping function to both features and omits the rho rescaling used in the sampling model in Eq. (5). This is inconsistent unless rho=1 is assumed; the authors should clarify which quantity is being plotted and why.","section":"Section 2.5.2"},{"comment":"The sentence 'A simulation of twenty-five curves from the generative model described is illustrated in Figure 1' is ambiguous because the simulation study uses datasets of size N=50, 100, and 150; the twenty-five curves appear to refer only to the illustrative figure.","section":"Section 3.1"},{"comment":"The paper notes that the model is 'sensitive to prior choices' but does not report the tuning procedure or a sensitivity analysis. Since the case study hyperparameters are chosen after tuning for convergence, some summary of the sensitivity of the main conclusions to these choices would be helpful.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the methodological idea is potentially valuable, but the current validation strategy is too heavily shaped by post-hoc choices to support the headline claims. The revisions requested in the major comments are feasible within the manuscript's scope: report full simulation results, address the rho identifiability issue, and provide sensitivity analyses for the labeling and rescaling decisions. I recommend major revision rather than rejection because the underlying model and Bayesian machinery are coherent and the limitations are explicitly acknowledged."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper combines Bayesian curve registration (Telesca and Inoue 2008) with functional mixed membership models (Marco et al. 2024a) and adds covariate-dependent warping. That synthesis is genuinely new, and the model is thoughtfully built: the Jupp transformation makes unconstrained warping regression straightforward, and the semi-supervised labeling is a sensible way to handle identifiability. The case study's PAF findings align with the neuroimaging literature, which is encouraging. The authors also deserve credit for being transparent about the rho multimodality and the limitations of their approach.\n\nThe soft spots are real, though. The simulation reports only on retained datasets (26/40 for N=50, 29/40 for N=100, 30/40 for N=150). That is selective reporting: the method fails to converge on a quarter to a third of runs, and the error metrics are computed only on the successful ones. The post-hoc rescaling of feature 2 (Eqs. 25-26) removes amplitude mismatch, so the R-MISE numbers do not reflect the actual estimation error. In the case study, rho is fixed to 0 because the sampler found a competing mode where the 1/f feature gains a peak. That means the central claim about 'recovering the 1/f pink noise feature distinctly from the alpha peak' is not a free finding—it is enforced by setting the second feature's warping to zero. The authors acknowledge this, but the abstract and introduction present it as a success rather than a constraint.\n\nReproducibility is also weaker than it should be: the code link is a placeholder ('put link') and the supplements are 'available upon request.' That matters for a methods paper that depends on MCMC tuning.\n\nOverall, the model is a legitimate contribution and the paper deserves a serious referee, but it needs revision before publication. The simulation should report all runs or justify why the discarded ones are invalid, the rescaling should be justified or replaced with an identifiability constraint that is part of the model, and the rho=0 decision should be presented as a limitation, not a validated choice. If the authors can address these, the method will be a solid addition to the functional data analysis toolkit.","headline":"A useful synthesis of Bayesian registration and functional mixed membership, but the central EEG claim rests on fixing rho=0 and selective simulation reporting.","tokens_in":17338,"tokens_out":2045,"would_cite":false,"duration_ms":21866,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62R10","62F15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A Bayesian curve-registration model for functional mixed memberships separates the 1/f pink-noise feature from the alpha-band peak in EEG spectra and quantifies how age and clinical status shift the peak alpha frequency.","keywords":["functional data analysis","curve registration","mixed membership models","Bayesian hierarchical models","EEG spectral analysis","peak alpha frequency","time warping","B-spline basis"],"falsifier":"Run the same EEG case study with two freely estimated warp functions per subject, one per feature, instead of a single warp scaled by $\\rho$; if the recovered 1/f shape, the posterior PAF distribution, or the age-by-diagnosis regression coefficients change materially, the paper's shared-warp assumption is the weak link. A simulation with two genuinely different warp shapes would show the same bias directly.","tokens_in":16341,"feed_emoji":"🧠","tokens_out":11492,"duration_ms":102818,"temperature":0.7,"pith_summary":"Functional data like EEG spectra are often misaligned: the same physiological feature appears at different frequencies in different people. The paper proposes a Bayesian model in which each person's curve is not a single shared shape but a mixture of two population-level feature shapes, each delayed or stretched by a person-specific time warp. Fitting this model to resting-state EEG from children with and without autism separates the 1/f pink-noise background from the alpha-band peak and, through a regression on the warp functions, estimates how peak alpha frequency changes with age and diagnosis. The central payoff is that phase variation and amplitude/feature variation are estimated together, so the location of the alpha peak and the difference between groups can be quantified without pre-aligning curves by hand.","feed_headline":"Age shifts the EEG alpha peak only in typical children","feed_subtitle":"A Bayesian warping model separates pink noise from the alpha peak and measures how age and autism move it.","key_machinery":"The central object is the mixed-membership registration equation $Y_i(t)=c_i+\\pi_i f_1(h_i(t))+(1-\\pi_i)f_2(\\rho(h_i(t)-t)+t)+\\epsilon_i(t)$, where $f_1$ and $f_2$ are B-spline feature shapes, $\\pi_i$ is the membership weight, $h_i$ is a subject-specific monotone time warp, and $\\rho$ rescales the warp for the second feature. The warp functions are B-splines constrained to be monotone and boundary-fixing; monotonicity is imposed through the Jupp transformation, which maps the constrained spline coefficients to an unconstrained space where a linear regression $E[\\tilde{\\eta}_i]=\\tilde{\\Upsilon}+B'X_i$ can be placed on warp coefficients. That regression is what lets age, diagnosis, and their interaction shift the peak $\\alpha$ frequency, while the semi-supervised labeling of a few subjects anchors the two features to identifiable scientific meanings.","core_discovery":"The paper claims that curve registration and mixed membership can be combined in a single Bayesian hierarchical model in which each observed curve is $Y_i(t) = c_i + \\pi_i f_1(h_i(t)) + (1-\\pi_i) f_2(\\rho(h_i(t)-t)+t) + \\epsilon_i(t)$, with $f_1$ and $f_2$ B-spline feature shapes, $\\pi_i$ the membership weight for feature 1, $h_i$ a monotone time-transformation function, and $\\rho$ a warp rescaling for the second feature. Identifiability comes from anchoring warps at the identity, centering intercepts and feature levels, and assigning a small number of subjects to full membership in one feature. In the EEG case study, $\\rho$ is fixed at zero, so the 1/f feature is effectively unwarped while the $\\alpha$-peak feature is warped. The model recovers the 1/f pink-noise component separately from the $\\alpha$ peak, produces a posterior distribution for peak $\\alpha$ frequency concentrated near 9.4–9.5 Hz, and shows through the warp regression that typically developing (TD) children's PAF increases with age while ASD children's does not, with TD children loading more heavily on the peak feature.","pith_inferences":["A stress test the paper does not report: simulate from a generative model where the two features have genuinely different warp shapes rather than a common shape rescaled by $\\rho$, then check whether the recovered 1/f shape and PAF-age regression become biased; this would directly probe the shared-warp assumption.","The Jupp-space warp regression could be extended from age and diagnosis to electrode-level or multivariate models, which the discussion gestures at; borrowing information across electrodes might sharpen the ASD age-interaction estimate in smaller samples.","The two-feature mixed-membership setup with registration is generic enough to transfer to other misaligned functional data with two interpretable subpopulations, such as gene-expression time courses, provided a few subjects can be labelled as pure types."],"forward_implications":["If the model is right, EEG spectral analyses can estimate aligned feature shapes and subject-specific warps jointly, so the peak alpha frequency and its uncertainty are available directly from the posterior without a separate pre-registration step.","Separating the 1/f noise from the alpha-peak shape removes the aperiodic background as a confound, making group comparisons of the peak cleaner.","The warp regression supplies a quantitative age-by-diagnosis picture: the alpha peak moves to higher frequencies with age in typically developing children but stays nearly fixed in children with ASD, matching prior clinical findings.","The semi-supervised design shows that labelling only a few subjects (5% in the case study) is enough to identify the two features, so the method does not require full cluster labels."],"supporting_citations":[{"why":"Introduces the Bayesian hierarchical curve-registration framework the paper extends to mixed membership.","marker":"Telesca and Inoue (2008)"},{"why":"Supplies the unconstrained B-spline time-transformation model and Jupp-transformation regression used for warps.","marker":"Brumback and Lindstrom (2004)"},{"why":"Defines the two-feature functional mixed membership model and first applies it to this EEG dataset.","marker":"Marco et al. (2024a)"},{"why":"Provides the separability identifiability condition for mixed membership models that motivates semi-supervised labeling.","marker":"Chen et al. (2023)"},{"why":"Contributes the first-order random-walk penalized B-spline prior for the feature shapes.","marker":"Lang and Brezger (2004)"},{"why":"Defines the Jupp transformation used to enforce monotone warp functions in an unconstrained parameter space.","marker":"Jupp (1978)"},{"why":"Introduces the EEG cohort and the age-related PAF finding for TD versus ASD children that the case study reproduces.","marker":"Dickinson et al. (2018)"},{"why":"Gives the independent result that ASD children lack the normal age–PAF relationship, used to validate the warp regression.","marker":"Edgar et al. (2015)"},{"why":"Supplies normative TD peak-alpha-frequency values against which the posterior PAF distribution is compared.","marker":"Miskovic et al. (2015)"},{"why":"Identifies the T8 electrode as the strongest contributor to ASD diagnosis, justifying the data focus.","marker":"Scheffler et al. (2019)"}],"fun_headline_variants":["Age shifts EEG alpha peak only in typical kids","Warped model shows age effect on alpha peak only in typical children","Bayesian warping separates alpha peak from pink noise; age effect only in typical children","EEG warping model: age moves alpha peak only in typical children"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model assumes a single shared time-warp shape per subject across both spectral features, with only a scalar rescaling $\\rho$ telling the features apart, and in the EEG case study that rescaling is fixed at zero so the 1/f feature is not warped at all; if the true warps differ in shape between features, the recovered feature shapes and peak-$\\alpha$-frequency estimates could be biased.","fun_headline_variants_meta":{"raw":{"variants":["Age shifts EEG alpha peak only in typical kids","Warped model shows age effect on alpha peak only in typical children","Bayesian warping separates alpha peak from pink noise; age effect only in typical children","EEG warping model: age moves alpha peak only in typical children"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1448,"prompt_tokens":961,"completion_tokens":487,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":410}},"tokens_in":577,"tokens_out":487,"duration_ms":5241,"temperature":1.0,"reasoning_tokens":410,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:35:42.804853+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same EEG case study with two freely estimated warp functions per subject, one per feature, instead of a single warp scaled by $\\rho$; if the recovered 1/f shape, the posterior PAF distribution, or the age-by-diagnosis regression coefficients change materially, the paper's shared-warp assumption is the weak link. A simulation with two genuinely different warp shapes would show the same bias directly.","supporting_citations":[{"cited_title":"and Inoue, L","cited_arxiv_id":null,"evidence_quote":"Introduces the Bayesian hierarchical curve-registration framework the paper extends to mixed membership."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the unconstrained B-spline time-transformation model and Jupp-transformation regression used for warps."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the separability identifiability condition for mixed membership models that motivates semi-supervised labeling."},{"cited_title":"and Brezger, A","cited_arxiv_id":null,"evidence_quote":"Contributes the first-order random-walk penalized B-spline prior for the feature shapes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Jupp transformation used to enforce monotone warp functions in an unconstrained parameter space."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the EEG cohort and the age-related PAF finding for TD versus ASD children that the case study reproduces."},{"cited_title":"C., Heiken, K., Chen, Y","cited_arxiv_id":null,"evidence_quote":"Gives the independent result that ASD children lack the normal age–PAF relationship, used to validate the warp regression."},{"cited_title":"A., Fan, M., Owens, M., Sayama, H., and Gibb, B","cited_arxiv_id":null,"evidence_quote":"Supplies normative TD peak-alpha-frequency values against which the posterior PAF distribution is compared."},{"cited_title":"W., Telesca, D., Sugar, C","cited_arxiv_id":null,"evidence_quote":"Identifies the T8 electrode as the strongest contributor to ASD diagnosis, justifying the data focus."}],"review_version":1}