{"id":"7bf3cc8a-43ce-4be2-828f-a33af73dd537","arxiv_id":"2607.06797","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MAPLE estimates conditional class probabilities by local averaging over Mapper-graph neighborhoods with data-driven cover selection and proves consistency under regularity conditions.","lead":"MAPLE turns the Mapper graph from topological data analysis into a nonparametric classifier that averages labels inside geometry-aware neighborhoods, with a bias-variance rule for choosing the cover. It can matter for biomedical prediction when disease stages or tumor grades sit on branching manifolds that global models miss.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Filter dependence is the real soft spot, but the paper already flags it and the theory is standard local averaging; no stronger internal break found.","rationale":"The Reader correctly isolates the filter as the weakest link and rates the paper CONDITIONAL with medium correctness risk. After re-reading the construction (§2.2–2.4), the bias–variance cover lemma, and Theorem 1 under (A1)–(A5), I find no deeper internal inconsistency: once neighborhoods are well-behaved the consistency proof is textbook local averaging, and the cover scaling is a transparent extension of classical bandwidth arguments. The simulations are well-matched to the heterogeneity story and the real-data graphs are interpretable. The only material soft spot remains filter quality, which the authors already list as a limitation and partially probe. Therefore no verdict change is warranted; the Reader’s CONDITIONAL assessment already reflects the right degree of caution pending fuller filter-sensitivity and supplement-proof checks.","tokens_in":19326,"tokens_out":572,"duration_ms":6434,"concrete_test":"Re-run the Scenario-1 simulation (n=500, σ=0.2) and the PPMI 10-fold CV using three alternative filters (first principal component, unsupervised MDS, and a deliberately poor random linear projection) while keeping the same cover rule and clustering. If O-MAPLE’s κ_w advantage over RF/OLR disappears or reverses under the non-OASDA filters, the headline gains are filter-driven rather than Mapper-driven; if the advantage persists, the concern is overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Theorem 1) is ordinary nonparametric consistency for a local average once neighborhoods satisfy |U_n(x)|→∞ and diam(U_n)→0 in probability (A4–A5), with the filter only required to be Lipschitz (A2). That argument is standard and does not appear broken. The practically load-bearing premise is the one the Reader already named: that the one-dimensional OASDA filter f preserves the response-relevant geometry so that Mapper pullbacks induce neighborhoods that are actually local for η_r. If f collapses branches or mixes heterogeneous regimes (possible under high-dimensional noise or when OASDA’s ordinal sparse projection is misspecified), the graph neighborhoods become misspecified even while formal consistency still holds for the projected process. The paper acknowledges filter dependence (§5) and supplies limited sensitivity (Supp. Table S4), but the main simulations and both applications use the same OASDA filter, so the empirical gains cannot be cleanly separated from filter quality. This is a genuine limitation on the strength of the “geometry-adaptive” claim, not a contradiction of Theorem 1.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes MAPLE, a nonparametric estimator of conditional class probabilities that uses Mapper-graph neighborhoods for localized averaging in high-dimensional predictor space. A one-dimensional filter (OASDA) induces overlapping intervals whose cover parameters (l, S, q) are chosen by a bias–variance criterion L(S,l)=ρS^{-1}+(1−ρ)l^{-4}, with optimal scaling given in Lemma 1; predictions use inverse-squared-distance weighted averages of labels in primary Mapper neighborhoods, with ordinal and multinomial decision rules. Theorem 1 claims pointwise consistency of ˆη_r(t) and Bayes-risk consistency of the plug-in classifier under (A1)–(A5). Simulations under designed branching heterogeneity show gains over multinomial/ordinal logistic regression and competitive or better performance versus random forest; applications to PPMI Hoehn–Yahr staging and TCGA glioma grade report competitive accuracy and topology-aware visualizations.","tokens_in":19631,"tokens_out":1176,"duration_ms":11494,"significance":"If the claims hold, MAPLE is a useful bridge between Mapper-style TDA and supervised nonparametric classification: it supplies an explicit local-average estimator, a data-driven cover rule with asymptotic scaling, and standard consistency guarantees, plus a permutation importance measure and code. The simulation design (branch-specific vs global latent scores) and two biomedical applications make the geometry-adaptive claim falsifiable and practically relevant for heterogeneous biomedical data where global parametric models are misspecified. Strengths include the explicit bias–variance cover criterion, reported SEs, and public code; the main limitation is that empirical gains are not cleanly separated from the quality of the OASDA filter.","major_comments":[{"comment":"§2.2 and Theorem 1 / (A2): The load-bearing premise for the “geometry-adaptive” claim is that the one-dimensional OASDA filter preserves response-relevant structure so that Mapper pullbacks induce neighborhoods that are local for η_r in the original space. Formal consistency only requires Lipschitz f and |U_n|→∞, diam→0; if f collapses branches or mixes regimes, neighborhoods can be misspecified for prediction while the projected process remains consistent. Main simulations and both applications use the same OASDA filter; Supp. Table S4 is only limited sensitivity. Please expand filter-sensitivity experiments (PCA, random projections, unsupervised filters) in the main text and state more carefully that gains are conditional on a response-relevant filter.","section":null},{"comment":"§2.3 Lemma 1 and Eq. (1): The optimal scaling l*∼n^{1/5}, S*∼2n/l*, q*/S*=1/2 is derived under classical bias–variance rates (var∼S^{-1}, bias²∼l^{-4}) for a one-dimensional smoother. It is not shown that Mapper clustering and graph connectivity preserve these rates once neighborhoods are irregular graph-induced sets rather than fixed-width intervals. Either sketch why the rates still hold under (A4)–(A5), or reframe Lemma 1 as a heuristic cover selector motivated by classical theory rather than a Mapper-specific optimality result.","section":null},{"comment":"§3.1–3.2 and Table 1: Scenario 1 is constructed so that branch-specific latent scores and thresholds make a single global surface misspecified; MAPLE’s largest gains appear precisely there. That is informative but risks overstating general superiority over RF/OLR. Please add at least one heterogeneous setting that is not explicitly branch-engineered (e.g., smooth manifold with local coefficient variation, or mixture of Gaussians without designed branches) so that the advantage is not tied only to the simulation’s generative geometry.","section":null}],"minor_comments":[{"comment":"§2.4: Inverse-squared-distance weights are one of several schemes (Supp. Table S2); a one-sentence justification for preferring ∥·∥^{-2} over inverse distance or uniform weights would help.","section":null},{"comment":"Figure 3 and Figure 4: Node pie charts and colors are informative but hard to parse in grayscale; consider a legend that maps stages to hatch patterns or a clearer colorblind-safe palette.","section":null},{"comment":"§4.1–4.2: Variable screening (permutation Δ for PPMI; RF MDA for TCGA) uses different pipelines and thresholds (5% vs MDA>3). Clarify why the procedures differ and whether MAPLE’s own importance was infeasible for the gene set only for cost reasons (as Table S6 suggests).","section":null},{"comment":"Notation: t_new vs tnew, and ˆη_r(t) vs ˆη_r(t_new) are used interchangeably; unify indexing of the estimator.","section":null},{"comment":"References: Carrière–Oudot (2018) is cited for Mapper stability; a brief pointer to how (A4)–(A5) relate to their structural results would strengthen the theory section.","section":null}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid methods contribution for a statistics/biostatistics audience; the filter-dependence issue is real but already partially acknowledged and fixable with expanded experiments rather than a redesign. Fit for a methods journal is good if the authors temper the geometry-adaptive claim and strengthen the cover-rate justification. No novelty or citation concerns stood out."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: this is a usable packaging of Mapper into an explicit conditional-probability estimator with a data-driven cover selector, not just another exploratory TDA plot. That combination is new enough to matter for people who already care about geometry-aware local averaging.\n\nWhat they actually did well is clear. They define neighborhoods from Mapper connectivity, average class indicators with inverse-squared-distance weights, give ordinal and multinomial decision rules, and pick the cover (l, S, q) from a bias-variance criterion that yields the usual 1/5 scaling. Theorem 1 is ordinary nonparametric consistency once neighborhood size goes to infinity and diameter goes to zero; the argument is standard and does not look broken. Simulations match the story: under designed branching heterogeneity MAPLE beats multinomial/ordinal logistic by a wide margin and is competitive with or better than random forest, especially in smaller n and higher noise; under a global smooth signal the gap collapses, which is honest. PPMI and TCGA applications give competitive QWK/C-index and graphs that are more interpretable than PCA/t-SNE. Code is linked; methods are re-implementable. Limitations (filter choice, cost, limited gains when the signal is globally smooth) are stated in the discussion rather than hidden.\n\nThe soft spot is real but proportionate. Everything rides on the one-dimensional filter (here OASDA) preserving the response-relevant geometry. If the filter collapses branches or mixes regimes, the graph neighborhoods are misspecified even while formal consistency for the projected process still holds. The paper flags this and shows limited sensitivity in the supplement, but main simulations and both applications use the same filter, so the “geometry-adaptive” gains cannot be cleanly separated from filter quality. Cover discretization, silhouette clustering, and the primary-neighborhood definition are free parameters; none of that is fatal. Theory proofs live in a supplement we do not have fully in hand, so a referee will want to check the diameter and size arguments carefully.\n\nThis is for people working on nonparametric classification, TDA-for-prediction, or heterogeneous biomedical data who want a concrete estimator rather than a visualization. It is not a field-reorganizing result. I would send it to peer review; a serious referee can pressure the filter-sensitivity and the supplement proofs without the paper collapsing. Worth reading if that is your area; I would cite it when I need a supervised Mapper baseline.","headline":"Solid supervised Mapper estimator with a real bias-variance cover rule and standard consistency; filter dependence is the main practical limit, not a broken theorem.","tokens_in":20226,"tokens_out":572,"would_cite":true,"duration_ms":52973,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G08","62H30","55N31"],"pacs":[],"model":"grok-4.5","headline":"MAPLE turns Mapper graphs into local probability estimators with data-driven cover scaling, and proves they are consistent for classification on heterogeneous high-dimensional data.","keywords":["Topological data analysis","Mapper algorithm","localized prediction","nonparametric classification","high-dimensional data","Parkinson’s disease","glioma RNA sequencing","bias-variance cover selection"],"falsifier":"On synthetic data whose response depends on a branching manifold that a deliberately poor filter (for example pure PCA of noise coordinates) completely flattens, MAPLE should lose its advantage over ordinary k-NN or logistic regression and its estimated probabilities should fail to approach the true branch-specific probabilities as n grows.","tokens_in":20204,"feed_emoji":"🕸️","tokens_out":765,"duration_ms":9591,"temperature":0.7,"pith_summary":"High-dimensional biomedical data often live on curved, branched manifolds where a single global model misreads the relationship between predictors and outcomes. This paper proposes MAPLE: it builds a Mapper graph from a one-dimensional filter, then estimates class probabilities by averaging only inside the graph-connected neighborhoods that contain a new point. Cover size and overlap are chosen by an explicit bias–variance criterion that yields the classical n^{1/5} scaling for the number of intervals. Under standard regularity conditions the resulting estimator is pointwise consistent for the true conditional probabilities and the plug-in classifier is Bayes-risk consistent. Simulations and two clinical applications (Parkinson stage, glioma grade) show the largest gains precisely when the data are heterogeneous and noisy, while the same graph supplies readable topological summaries of patient subgroups.","feed_headline":"Mapper graphs become consistent local classifiers","feed_subtitle":"Data-driven cover scaling and neighborhood averaging beat global models on branched biomedical data","key_machinery":"MAPLE estimator: inverse-distance-weighted average of class indicators inside the union of Mapper clusters that contain a new observation, with interval number and overlap selected by minimizing ρ/S + (1-ρ)/l^4 subject to the cover constraint.","core_discovery":"The authors show that conditional class probabilities can be estimated by local averaging over data-adaptive neighborhoods induced by a Mapper graph whose cover parameters are chosen to balance bias and variance; under Lipschitz filter, twice-differentiable probabilities, and vanishing neighborhood diameter the estimator is pointwise consistent and the plug-in rule is Bayes-risk consistent.","pith_inferences":["The same construction could be lifted to survival outcomes by replacing class indicators with local Nelson–Aalen or Cox scores inside Mapper nodes.","If the filter itself is learned jointly with the cover (instead of fixed OASDA), the method may recover informative projections that pure unsupervised lenses miss.","Computational cost of repeated Mapper builds for variable importance suggests a natural next step: sparsity or screening inside the Mapper construction rather than as a pre-filter."],"forward_implications":["When predictor–response relations vary across latent branches, Mapper neighborhoods recover local structure that global multinomial or ordinal models miss.","Cover parameters can be set by a single scalar bias–variance trade-off instead of ad-hoc interval counts, giving asymptotic guidance for Mapper resolution.","The same graph that produces the predictions supplies topology-aware patient subgroups with distinct stage or grade compositions and clinical summaries.","Permutation importance computed from Mapper neighborhoods quantifies which covariates both predict and shape the topology.","Primary graph neighborhoods already balance locality and stability; deeper secondary neighborhoods need not improve accuracy."],"fun_headline_variants":["Mapper graphs yield consistent local class probabilities","Data-driven covers make Mapper neighborhoods Bayes-optimal","MAPLE turns adaptive Mapper graphs into local classifiers","Bias-variance Mapper covers beat globals on branched data","Topology-aware local averaging ensures pointwise consistency"],"cache_read_input_tokens":3328,"weakest_assumption_plain":"The one-dimensional filter must keep the geometry that actually matters for the response; if it collapses or distorts the relevant structure, the Mapper neighborhoods become the wrong sets and consistency no longer yields useful predictions.","fun_headline_variants_meta":{"raw":{"variants":["Mapper graphs yield consistent local class probabilities","Data-driven covers make Mapper neighborhoods Bayes-optimal","MAPLE turns adaptive Mapper graphs into local classifiers","Bias-variance Mapper covers beat globals on branched data","Topology-aware local averaging ensures pointwise consistency"]},"model":"grok-4.5","effort":"low","cost_usd":0.003484,"raw_usage":{"total_tokens":1150,"prompt_tokens":758,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":34840000,"prompt_tokens_details":{"text_tokens":758,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":338,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":758,"tokens_out":54,"duration_ms":4857,"temperature":1.0,"reasoning_tokens":338,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T21:06:48.601684+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On synthetic data whose response depends on a branching manifold that a deliberately poor filter (for example pure PCA of noise coordinates) completely flattens, MAPLE should lose its advantage over ordinary k-NN or logistic regression and its estimated probabilities should fail to approach the true branch-specific probabilities as n grows.","supporting_citations":[],"review_version":1}