{"id":"3ebb60e9-dc22-4cc5-a0c6-449276ddef80","arxiv_id":"2508.18058","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"BigFirst uses genetics and brain scans to split MCI patients into low-risk and high-risk subtypes, claiming faster progression and distinct biomarker profiles in the high-risk group.","lead":"A new machine learning pipeline called BigFirst sorts people with mild cognitive impairment into two groups, one that looks more like healthy aging and one that looks more like Alzheimer's disease, based on brain scans and genetic data. The authors report that the high-risk group declines faster and has different biomarkers, which could help select patients for clinical trials.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim of high conversion consistency is asserted without reported statistics; if the risk–conversion association is not significant, the subtyping is not validated by outcome.","rationale":"The reader's weakest assumption—alignment of the HC/AD boundary with MCI risk—is real, but the paper's own proposed evidence for that alignment is the sMCI/pMCI concordance. That evidence is asserted, not quantified. The strongest claim includes 'high consistency with conversion to AD over two years,' which is a testable empirical proposition. If the concordance is not statistically significant, the assumption fails empirically and the two subtypes are not validated as risk groups. If it is significant, the subtyping has predictive validity even if the exact biological meaning of the boundary is debated. Thus the missing statistics are the load-bearing gap. The GWAS circularity and multiple-testing issues are secondary because they concern biomarker interpretation, not the existence of risk stratification. I partially agree with the reader: the boundary assumption is the underlying vulnerability, but the precise, decisive flaw is the lack of a quantitative outcome test. A conditional verdict remains appropriate: the paper should be required to report the conversion statistics, ideally with an external cohort, before the central claim can be accepted. Therefore verdict_should_be = UNCHANGED (reader's CONDITIONAL is appropriate).","tokens_in":26464,"tokens_out":9129,"duration_ms":102635,"concrete_test":"Re-analyze the ADNI data used for Fig. 7: construct the 2×2 contingency table of BigFirst risk label (low/high) vs 24-month conversion status (stable/progressed), and compute Fisher's exact test (or χ²) with odds ratio and 95% CI. If the association is not statistically significant at p<0.05, the 'high consistency' claim fails; if significant, the outcome validation supports the subtyping.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that BigFirst subtypes are risk-stratified and show high consistency with two-year conversion to AD—rests on Section 2.6, which states 'substantial concordance' with sMCI/pMCI but reports no contingency table, risk ratio, p-value, or CI. Likewise, Section 2.4 and Fig. 5 present longitudinal trajectory differences with no statistical testing. The risk labels are produced by GMM clustering on HC/AD-trained representations (Section 4.2.3), but the paper never specifies how the two clusters are assigned 'low' vs 'high' risk; the only outcome-based validation is the unquantified sMCI/pMCI concordance. If that concordance is weak or non-significant, the HC/AD similarity axis does not correspond to MCI risk, and the 'low-risk'/'high-risk' labels are arbitrary clusters, invalidating the central claim regardless of the secondary biomarker and genetic analyses. This is more fundamental than the GWAS circularity: the conversion consistency is the only direct, non-circular evidence that the method captures risk.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes BigFirst, a three-module imaging-genetics fusion framework (RGE, RG2PG, CMPF) trained on HC vs. AD data and then deployed on 417 MCI subjects to learn a risk representation. Applying GMM to the learned multimodal representation yields two MCI subtypes, 86 'low-risk' and 331 'high-risk.' The authors report that the two subtypes differ in brain imaging QTs, CSF Aβ, cognitive domains, clinical symptoms, and covariates at baseline, in longitudinal trajectories of fluid biomarkers and cognition, and in GWAS/PheWAS genetic associations. They further claim high concordance with sMCI/pMCI conversion over two years. The main methodological components are a Mamba-based SNP encoder, a GAN-based pseudo-brain generator, and supervised contrastive fusion with GMM clustering.","tokens_in":26764,"tokens_out":5038,"duration_ms":64029,"significance":"If the central claim holds, BigFirst would be a potentially useful tool for cross-sectional MCI risk stratification using multimodal imaging and genetic data, with practical relevance for clinical trial enrichment. The paper has notable strengths: the architecture explicitly combines whole-genome SNP information with three imaging modalities, the authors report ablation results (Supplementary A.4.1), clustering quality metrics (Table 2), parameter sensitivity (Fig. A1), and external genetic annotation via PheWAS. The downstream biomarker comparisons are extensive. However, the two pillars of the central claim—that the subtypes are truly 'risk' subtypes and that the model captures MCI-to-AD conversion risk—are currently unsupported by quantitative outcome statistics. The low/high-risk labels are not defined operationally, and the only direct outcome validation (sMCI/pMCI concordance) is asserted without a contingency table or statistical test. These are correctable but load-bearing omissions.","major_comments":[{"comment":"The central validation of the risk subtypes is the claimed 'substantial concordance' with sMCI/pMCI conversion, but no quantitative result is reported. The text gives no contingency table, odds ratio, risk ratio, p-value, confidence interval, or sensitivity/specificity for two-year conversion. Without these statistics, the statement that BigFirst has 'a strong ability in AD conversion prediction' is unsupported. Please report the full 2x2 table for low/high-risk vs. stable/progressive MCI, an effect size with CI, and a test of association; if appropriate, also adjust for age, sex, and APOE.","section":"Section 2.6 / Fig. 7"},{"comment":"The paper never specifies how the two GMM clusters are assigned the labels 'low-risk' and 'high-risk.' Section 4.2.3 says only that GMM is applied to the learned neuroimaging representations. Since all downstream analyses depend on this labeling, the assignment rule must be explicit. For example, are clusters assigned by distance of cluster centroids to HC vs. AD representations, by the sign of a learned discriminant score, or by some monotonic ordering of cluster means? If the assignment is arbitrary or data-driven post hoc, the labels are not biologically grounded and the reported subtype differences could be artifacts of the ordering.","section":"Section 4.2.3 / Section 2.1"},{"comment":"The GWAS analysis has two serious problems. First, the significance threshold was lowered from 1e-8 to 1e-5 after no SNP passed the initial threshold; with 359,997 SNPs tested, approximately 3.6 hits are expected by chance at p < 1e-5, so the 20 reported loci (Table A3) are not adequately controlled for multiple testing. Second, and more fundamentally, the subtype labels are generated by a model whose RGE module consumes the same SNP data. A GWAS on those labels may therefore rediscover SNPs that the model used to make its representations, rather than independent biological associations. The claims about CACNA1C and ABCA13 need either replication in an independent MCI cohort, a permutation-based null that breaks the SNP-label relationship, or a clear statement that these are model-derived candidates and not genome-wide significant associations.","section":"Section 2.5.1 / Table A3"},{"comment":"The longitudinal findings are presented only as 'trajectories fitted from the mean values' with no statistical testing. Statements such as 'the progression of NfL was slower in the low-risk group' and 'the high-risk group exhibiting faster progression in later stages' are not supported by mean-curve inspection. The authors should fit a mixed-effects model (e.g., linear mixed model) to the individual longitudinal measurements, with a time-by-subtype interaction term, and report the interaction p-values and effect sizes. This applies to Fig. 5(a), (b), and (c).","section":"Section 2.4 / Fig. 5"},{"comment":"The paper reports three of 28 clinical symptoms as significant (headache p=6.50e-3, urinary frequency p=2.08e-2, falling p=1.51e-2) with no multiple-testing correction. With 28 χ2 tests, the Bonferroni threshold is 0.05/28 = 1.79e-3, and none of the reported p-values reaches this threshold. The claim that these symptoms 'reaching the significance level' is therefore incorrect under the stated correction policy. Either use an FDR/Bonferroni correction and report adjusted p-values, or clearly label these as uncorrected exploratory findings.","section":"Section 2.3.4"}],"minor_comments":[{"comment":"The generator loss contains a typo: 'Dm(Y′m − 1)' should presumably be '(Dm(Y′m) − 1)'. Please correct the equation.","section":"Section 4.2.2, Eq. (3)"},{"comment":"The text states 'three discontinuous covariates ... with t-test and two continuous covariates ... with χ2-test,' but the correct assignment is the reverse: discrete covariates require χ2 and continuous covariates require t-test. The subsequent p-values suggest the intended tests were used, but the wording is wrong.","section":"Section 2.3.5"},{"comment":"The interpretation step for genetic weights is underspecified: 'We performed a linear regression between the input (SNPs) and output of Mamba.' Please state the regression target explicitly (e.g., Mamba output for each tissue group), the standardization of SNPs, and whether coefficients were averaged across cross-validation folds.","section":"Section 2.2"},{"comment":"The GMM application is described as predicting 'disease status' when deployed on MCI. Clarify whether GMM is clustering the fused representations or performing a soft classification relative to HC/AD class centers, and define the input dimensionality and number of components selected.","section":"Section 4.2.3"},{"comment":"For the CH index, higher values are better, yet the complete model ('with RGE') has a slightly lower CH (419.80) than the reduced model ('without RGE', 420.38). The text says the proposed framework obtained 'higher CH values' than other methods; this is consistent for k-means/GMM but not for the ablation comparison. Please clarify or report the variance/standard error of these indices.","section":"Table 2"},{"comment":"There is no code or data availability statement for the method. Since the paper introduces a new algorithmic framework, sharing code (or at least a detailed pseudo-code for the GMM labeling step) would substantially improve reproducibility.","section":"General"},{"comment":"Numerous typographical errors remain, including 'data avaliability' (Section 5), 'MCIs with different level risks' (Abstract/introduction), 'were first accessed' (Section 2.1), 'comment genetic basis' (Section 2.5.2), and 'ore suitable' (Discussion). A careful proofread is needed.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"I have no editor-only concerns beyond those in the report. The paper's current form is more of an exploratory multi-modal clustering analysis than a validated risk-stratification tool; the authors should be encouraged to focus revision on the quantitative outcome validation and the explicit definition of the risk axis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague, quick take: the method is a genuine new assembly — Mamba SNP encoder, GAN-generated pseudo-brains, contrastive fusion — and the HC/AD classification accuracy (0.93) with RGE is decent. The paper also does a thorough job characterizing the two clusters across many modalities, and the direction of differences (more atrophy, worse cognition, lower Aβ42 in the high-risk group) is biologically plausible. Credit where due: they state the key assumption plainly in the Discussion, report parameter sensitivity, and the clustering metrics beat the comparison methods.\n\nThe load-bearing validation is missing, though. Section 2.6 claims 'substantial concordance' with sMCI/pMCI over two years but reports no contingency table, risk ratio, p-value, or CI. That concordance is the only direct, non-circular evidence that the HC/AD similarity axis is a risk axis. Without it, 'low-risk' vs 'high-risk' could just be proximity to the training endpoints. The stress-test note is right on this.\n\nThe GWAS also has a circularity problem: the subtype labels come from a model whose RGE module consumed the same SNPs, so the CACNA1C/ABCA13 hits could be shadows of the model rather than independent biology. Lowering the threshold to 1e-5 after nothing passed 1e-8 makes it worse. The 28 clinical symptom χ2 tests are not corrected — with Bonferroni, none of the three nominal p-values would survive — and the longitudinal trajectories in Fig. 5 are mean curves with no statistical testing, so 'slower progression' is an observation, not a result.\n\nThe clustering procedure itself is under-specified: GMM is mentioned in Section 4.2.3, but there is no description of how the two clusters get assigned 'low' vs 'high' risk labels. That is a reproducibility gap, along with no code release.\n\nNone of this kills the underlying idea — the subtyping concept is plausible and the biomarker patterns align with prior literature. But this is a methods paper with a promising architecture and provisional validation, not an established finding. Send it to peer review, yes, with major revision expected: proper statistics for conversion, a non-circular genetic analysis (e.g., leave-out-SNP), multiple-testing controls, and a described clustering/assignment procedure. I would not cite it as evidence yet, but I would follow it once these are addressed.","headline":"Novel architecture for MCI risk subtyping, but the key outcome validation is unquantified and the GWAS is circular — worth a serious referee, not yet a citable finding.","tokens_in":27234,"tokens_out":2861,"would_cite":false,"duration_ms":33862,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes that MCI patients can be split into low- and high-risk subtypes for Alzheimer's progression using only baseline brain scans and genetic data, and presents BigFirst as the method that does it.","keywords":["MCI subtyping","Alzheimer's disease","risk stratification","imaging genetics","multimodal fusion","brain imaging genetics","MCI heterogeneity","conversion prediction"],"falsifier":"If an independent cohort with long follow-up showed that MCI patients labeled low-risk convert to Alzheimer's at the same rate as those labeled high-risk, or that the two predicted subtypes have identical longitudinal atrophy, tau, and neurofilament-light slopes, the risk-axis claim would collapse. A direct test would be to take the model's decision scores on an external set and compare conversion-free survival of the two predicted subtypes; a hazard ratio near one would mean the stratification conveys no risk information.","tokens_in":26336,"feed_emoji":"🧠","tokens_out":6204,"duration_ms":80597,"temperature":0.7,"pith_summary":"This paper proposes a method, BigFirst, that sorts people with mild cognitive impairment (MCI) into low-risk and high-risk subtypes for progression to Alzheimer's disease, using only baseline brain imaging and genetic data. The core claim is that a model trained on the difference between healthy controls and Alzheimer's patients can serve as a risk ruler for the intermediate MCI population, exposing two biologically distinct subgroups. If the claim holds, clinical trials could pre-screen MCI participants by risk instead of waiting two years to see who converts, reducing the heterogeneity that often weakens trial results. The authors report that the high-risk group has greater brain atrophy, altered amyloid markers, worse memory and executive function, faster longitudinal decline in tau and neurofilament light, and two-year conversion rates largely consistent with the standard 'progressive MCI' label.","feed_headline":"Two MCI risk subtypes emerge from baseline brain scans and genes","feed_subtitle":"Trained only on healthy and Alzheimer brains, it labels MCI by risk and tracks two-year conversion.","key_machinery":"BigFirst is a three-module pipeline. The risk genetic information extraction module uses a selective state-space sequence model to compress roughly 360,000 brain-tissue-related genetic loci into disease-relevant genetic representations. The risk genetic guided pseudo-brain generation module sends those representations through adversarial generators to create per-modality attention masks, called pseudo-brains, that gate real imaging data from three modalities: amyloid PET, glucose PET, and structural MRI. The cross-modality pseudo-brain fusion module projects the gated modalities into a shared space with supervised contrastive learning and uses a Gaussian mixture model to separate healthy fro","core_discovery":"The central discovery is that deploying an imaging-genetics model trained only on healthy controls versus Alzheimer's patients onto an MCI population yields a stable two-cluster partition: 86 low-risk MCI and 331 high-risk MCI. The two clusters are not arbitrary; they differ on many independent measures, including brain atrophy patterns, cerebrospinal fluid amyloid levels, cognitive performance, clinical symptoms, and longitudinal trajectories of plasma tau and neurofilament light. The high-risk group behaves physiologically more like Alzheimer's patients and the low-risk group more like healthy controls, which is exactly the axis the model was trained on. The subtypes also show substantial","pith_inferences":["A direct extension would be to give each MCI patient a calibrated conversion probability from the model's fused representation, using follow-up conversion as supervision; the paper states its current method only assigns subtype, not risk probability.","The healthy-versus-Alzheimer's boundary is what defines 'risk', so an independent external cohort with long follow-up is the natural test of whether the boundary picks out the same high- and low-risk people; internal clustering metrics alone cannot establish biological validity.","Because the method learns a single axis from the two endpoint groups, it may collapse genuinely distinct MCI etiologies that share endpoints; applying the same fusion pipeline without those endpoint priors could reveal whether more than two MCI subtypes exist on other axes.","The genetic filtering step depends on brain-tissue expression and splicing annotations; alternative tissue-specific or functional annotation schemes could change the genetic representation and therefore the subtypes, so the result should be tested with alternative annotation choices."],"forward_implications":["MCI risk subtypes can be assigned from a single cross-sectional visit, so trials can enrich for high-risk converters earlier than follow-up-based labels allow.","Because the two subtypes differ on structural atrophy but not on amyloid PET, the results suggest brain atrophy is a more sensitive discriminator at the MCI stage and may precede metabolic and amyloid changes.","The differing longitudinal trajectories of tau and neurofilament light imply the high-risk subtype is the natural target for early intervention trials, while low-risk individuals may be spared aggressive therapy or used as controls.","Genetic loci in CACNA1C and ABCA13, plus lifestyle traits including physical activity, social engagement, and diet, associate with subtype membership and offer candidate risk markers.","The subtype labels overlap strongly with two-year stable versus progressive MCI labels but are not identical, so the method predicts conversion at the group level, not certainty for individuals."],"supporting_citations":[{"why":"Supplies the state-space sequence model used to condense the large SNP set into risk genetic representations.","marker":"[35]"},{"why":"Provides the adversarial generative design that the pseudo-brain generator adapts to build gene-guided imaging masks.","marker":"[64]"},{"why":"Supplies the supervised contrastive learning objective used to fuse the multiple imaging modalities.","marker":"[68]"},{"why":"Earlier biological-heterogeneity subtyping of MCI that motivates biomarker-driven stratification and serves as a comparator.","marker":"[12]"},{"why":"Identifies atrophy-based MCI subtypes from longitudinal MRI; the paper contrasts its risk axis with these progression patterns.","marker":"[14]"},{"why":"Data-driven classification of MCI subtypes that predicts progression, used as a comparison for validated subtyping.","marker":"[17]"},{"why":"Contrastive clustering baseline used in the clustering performance comparison.","marker":"[34]"},{"why":"Provides brain-tissue expression and splicing quantitative trait annotations used to filter whole-genome SNPs to roughly 360,000 loci.","marker":"[62]"},{"why":"Genome-wide association tool used to map subtype differences to specific genetic loci.","marker":"[32]"},{"why":"Gene-set enrichment analysis used to connect those loci to Alzheimer's-related biological pathways.","marker":"[33]"}],"fun_headline_variants":["Imaging-genetics splits MCI into two risk subtypes","Brain scans and genes separate MCI risk groups","Two MCI risk subtypes emerge from fused imaging-genetics","Imaging-genetics model sorts MCI by Alzheimer risk","MCI risk subtypes identified via imaging-genetics"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The whole scheme assumes that the biological axis separating healthy controls from Alzheimer's patients is the same axis that separates low-risk from high-risk MCI, so training only on healthy versus Alzheimer's brains gives a valid risk ruler for MCI.","fun_headline_variants_meta":{"raw":{"variants":["Imaging-genetics splits MCI into two risk subtypes","Brain scans and genes separate MCI risk groups","Two MCI risk subtypes emerge from fused imaging-genetics","Imaging-genetics model sorts MCI by Alzheimer risk","MCI risk subtypes identified via imaging-genetics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1113,"prompt_tokens":742,"completion_tokens":371,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":294}},"tokens_in":486,"tokens_out":371,"duration_ms":4559,"temperature":1.0,"reasoning_tokens":294,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:35:48.464466+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If an independent cohort with long follow-up showed that MCI patients labeled low-risk convert to Alzheimer's at the same rate as those labeled high-risk, or that the two predicted subtypes have identical longitudinal atrophy, tau, and neurofilament-light slopes, the risk-axis claim would collapse. A direct test would be to take the model's decision scores on an external set and compare conversion-free survival of the two predicted subtypes; a hazard ratio near one would mean the stratification conveys no risk information.","supporting_citations":[{"cited_title":"IEEE Transactio ns on Medical Imaging 41(9), 2348–2359 (2022)","cited_arxiv_id":null,"evidence_quote":"Provides the adversarial generative design that the pseudo-brain generator adapts to build gene-guided imaging masks."},{"cited_title":"Advances in neural information processing systems 33, 18661–18673 (2020)","cited_arxiv_id":null,"evidence_quote":"Supplies the supervised contrastive learning objective used to fuse the multiple imaging modalities."},{"cited_title":"Alzheimer’s & Dementia 10(5), 511– 521 (2014)","cited_arxiv_id":null,"evidence_quote":"Earlier biological-heterogeneity subtyping of MCI that motivates biomarker-driven stratification and serves as a comparator."},{"cited_title":"Brain 140(3), 735–747 (2017)","cited_arxiv_id":null,"evidence_quote":"Identifies atrophy-based MCI subtypes from longitudinal MRI; the paper contrasts its risk axis with these progression patterns."},{"cited_title":"Alzheimer’s & Dementia 20(5), 3442–3454 (2024)","cited_arxiv_id":null,"evidence_quote":"Data-driven classification of MCI subtypes that predicts progression, used as a comparison for validated subtyping."},{"cited_title":"In: Proceedings of the AAAI Conference on Artiﬁcial Intelligence, vol","cited_arxiv_id":null,"evidence_quote":"Contrastive clustering baseline used in the clustering performance comparison."},{"cited_title":"Science 369(6509), 1318–1330 (2020)","cited_arxiv_id":null,"evidence_quote":"Provides brain-tissue expression and splicing quantitative trait annotations used to filter whole-genome SNPs to roughly 360,000 loci."},{"cited_title":"Gigascience 4(1), 13742–015 (2015)","cited_arxiv_id":null,"evidence_quote":"Genome-wide association tool used to map subtype differences to specific genetic loci."},{"cited_title":"PLoS computational biology 11(4), 1004219 (2015)","cited_arxiv_id":null,"evidence_quote":"Gene-set enrichment analysis used to connect those loci to Alzheimer's-related biological pathways."}],"review_version":1}