{"id":"5c3b9b79-d0f7-4b4c-b78c-9272e1be014c","arxiv_id":"2508.13199","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A computational ranking of candidate 'plasticity marker' genes in pediatric AML from longitudinal scRNA-seq, using LSTM, Transformer, and BDM perturbation, without reported predictive performance.","lead":"This paper uses deep learning and algorithmic complexity methods on gene expression snapshots from pediatric leukemia patients to rank genes that may control whether leukemia cells differentiate or stay cancerous. The authors propose these genes as predictive biomarkers and therapeutic targets, but no prediction accuracy is reported, leaving the main claim unverified.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of 'predictive signatures' is unsupported: the paper never reports any predictive performance (accuracy, AUC, held-out validation) for the deep learning models from which the signatures are derived.","rationale":"The reader's verdict rejects the paper because the central claim of predictive signatures is unsupported by any predictive performance metric, and because the analysis is in-sample with arbitrary thresholds and no external validation. My stress-test identifies the same core issue and sharpens it: the deep learning models are never evaluated for their actual predictive ability, making the feature importance rankings (and hence the claimed 'predictive signatures') potentially artifacts of overfitting. The weakest assumption about three snapshots and pseudobulk aggregation is a contributing factor, but the absence of predictive validation is the decisive load-bearing flaw. I therefore agree with the reader's rejection and recommend no change in verdict; the paper would need to add held-out predictive performance and stability analyses to support even a conditional acceptance.","tokens_in":27738,"tokens_out":1994,"duration_ms":26744,"concrete_test":"Re-run the LSTM, BiLSTM, and Transformer pipelines with a rigorous held-out evaluation: leave-one-patient-out cross-validation on all 28 patients (or at least multiple fixed-seed random splits of the 14) and report the classification accuracy and AUC for distinguishing DX, REL, and REM states on the held-out patients. Compare against a label-permutation baseline to confirm performance is above chance. Also compute the Jaccard overlap of the top-100 feature-importance genes across multiple random patient subsets; if accuracy is near chance or the overlap is low, the 'predictive signatures' are not stable or predictive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and conclusions assert that the study 'identifies key plasticity markers as predictive signatures regulating developmental trajectories' and that the algorithms 'forecast' AML cell fate transitions. However, the Methods describe training LSTM/BiLSTM/Transformer models on pseudobulk gene expression from only 14 patients with 3 timepoints each, and the Results never report any measure of predictive accuracy, sensitivity, specificity, or held-out generalization for these models. The models are used solely to extract feature-importance weights, which are then interpreted as 'predictive signatures.' But without demonstrating that the models can classify or forecast held-out states (e.g., DX vs REL vs REM) better than chance, the 'predictive' descriptor is unearned. The feature rankings could simply reflect in-sample overfitting to the specific 14-patient subset (chosen randomly without a fixed seed) and to arbitrary thresholds (BDM binarization at 0.1 and 0.5; DEG threshold |log2FC|≥1, FDR<0.05). Moreover, pseudobulk aggregation across cells within each patient-timepoint discards single-cell resolution, and the 14-patient subset is not shown to be representative. Without predictive validation, the subsequent causal and therapeutic claims (e.g., 'causal drivers,' 'differentiation therapy targets') are unsupported. This is load-bearing because the paper's novelty and translational value rest entirely on the identification of *predictive* biomarkers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a 'neurosymbolic' framework that combines Algorithmic Information Dynamics (BDM perturbation), NNMF clustering, and deep learning models (LSTM, BiLSTM, Transformer, Bayesian Ridge Regression) applied to public bulk and longitudinal single-cell AML transcriptomic data. From a random subset of 14 of 28 patients with three clinical timepoints (diagnosis, remission, relapse), the authors derive DEGs, BDM perturbation scores, and feature-importance rankings, and interpret convergence across methods as the discovery of predictive plasticity biomarkers and causal regulators of AML cell fate transitions. The abstract and conclusions further claim that these signatures reveal ectoderm–mesoderm crosstalk, a brain–immune–hematopoietic axis, and actionable differentiation-therapy targets. The central scientific claim is that the models 'predict' or 'forecast' AML cell fate transitions and identify causal biomarkers.","tokens_in":28164,"tokens_out":3568,"duration_ms":48504,"significance":"If the central claim were established, a validated longitudinal deep-learning framework for predicting AML relapse from single-cell transcriptomics would be a meaningful contribution to computational oncology because true predictive biomarkers for paediatric AML relapse are an unmet clinical need. The manuscript has strengths: it uses publicly available datasets, provides a GitHub code repository, and attempts to combine several complementary methods rather than a single black-box model. However, the evidence presented does not support the central claim: no predictive performance is reported, no held-out validation is described, and the causal language greatly exceeds what the correlation- and perturbation-based analyses can establish. The contribution is therefore best viewed as an exploratory signature-discovery pipeline whose translational and predictive claims remain unsubstantiated.","major_comments":[{"comment":"The paper never reports any measure of predictive performance for the deep learning models. The Methods describe a 70/15/15 train/validation/test split for classification across DX/REL/REM states, but the Results contain no accuracy, sensitivity, specificity, AUC, or any held-out test prediction. Feature-importance weights extracted from trained model parameters are not equivalent to predictive accuracy; a model can assign high importance to features while still failing to generalize. Because the abstract's central claim is that the work 'identifies key plasticity markers as predictive signatures' and that the algorithms 'forecast' cell fate transitions, the absence of any predictive validation is load-bearing and unsupported.","section":"Methods (LSTM, BiLSTM, Transformer, BRR) and Results (AI-Driven Longitudinal Biomarker Discovery; Figure 3)"},{"comment":"There is a circularity in the biomarker discovery. The top-100 DEGs are selected from the same 14-patient pseudobulk cohort that is then used to train the deep learning models and to construct the BDM correlation networks. The feature rankings extracted from those models are subsequently presented as convergent evidence for the same genes. No held-out patients or independent cohort are used. The 14 patients are chosen 'randomly' from 28 without a reported seed, and the manuscript does not demonstrate that this subset is representative. Consequently, overlap between LSTM, Transformer, BDM, and NNMF gene lists may reflect shared input features and thresholds rather than biological convergence.","section":"Datasets and Preprocessing; BDM Analysis; AI-Driven Longitudinal Biomarker Discovery"},{"comment":"The causal claims are not supported by the methodology. BDM perturbation is applied to adjacency matrices constructed from Spearman correlations and binarized at arbitrary thresholds (0.5 or 0.1). Deleting a node and measuring the change in algorithmic complexity of a correlation graph does not establish that the gene is a causal regulator of cell fate, nor does it identify a 'critical tipping point' or 'causal control point.' Bayesian Ridge regression coefficients are also associative. The manuscript repeatedly uses 'causal drivers,' 'causal biomarkers,' and 'critical transition genes' in the Discussion and Conclusions, and Table 3 proposes CRISPR and small-molecule therapeutic combinations based on these rankings. Without perturbation experiments, causal or therapeutic claims exceed the evidence.","section":"BDM Analysis; Results (Single-Cell BDM Perturbation Signatures); Discussion"},{"comment":"The trajectory-inference and forecasting claims are undermined by the data structure. Expression is pseudobulk-aggregated from three clinical snapshots per patient, so the models observe three coarse timepoints rather than continuous cell fate transitions. The Limitations section correctly concedes that 'sparse clinical timepoints' constrain trajectory resolution and that multi-omic validation is a future priority, but the Conclusions nevertheless state that the algorithms 'forecast' cell fate trajectories and identify causal biomarkers. The manuscript does not control for batch or technical effects beyond log-normalization, and the 14-patient subset is not shown to preserve subtype diversity or temporal completeness in any quantitative way.","section":"Datasets and Preprocessing; Limitations and Mechanistic Interpretations"}],"minor_comments":[{"comment":"Typo: 'Dabatase' should be 'Database.' Also, Figure 2 panels E-G are described as scatter-like plots but the axes are not fully defined; please clarify what 'count frequency' and 'average BDM perturbation' represent.","section":"Results (first paragraph)"},{"comment":"No random seed, hyperparameter tuning, early stopping, or regularization strategy beyond dropout is reported for the deep learning models. With only 14 patients and 50 epochs, overfitting is a serious concern; please report seeds, tuning procedure, and any regularization used.","section":"Methods (LSTM, Transformer)"},{"comment":"Table 3 lists numerous therapeutic combinations (e.g., KDM5B inhibition, WNT activation, TPT1 deletion) without any experimental validation. These should be clearly labeled as speculative hypotheses, not 'precision targets' as implied in the text.","section":"Conclusions; Table 3"},{"comment":"References [61] and [62] appear to be the same STGRNS entry and should be consolidated.","section":"References"},{"comment":"Typo: 'linage' should be 'lineage'; in the Conclusions, 'PCHD1/2' should likely be 'PCDHA1/2.'","section":"Introduction; Conclusions"}],"recommendation":"reject","confidential_remarks":"The central predictive claim is unsupported as written: no test-set performance or held-out validation appears anywhere in the Results. The causal and therapeutic statements in the abstract, Discussion, and Table 3 go far beyond the evidence. I would be open to a substantially revised manuscript that reframes the work as exploratory signature discovery, adds a proper held-out evaluation with uncertainty quantification, and replaces causal language with associative/computational language. However, as it stands, the paper's main claims are not defensible within its current scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline: this paper's central claim—identifying 'predictive signatures' for AML cell fate—is not supported by what's actually in the paper. The LSTM, BiLSTM, Transformer, and Bayesian Ridge models are described in detail, but no accuracy, sensitivity, specificity, AUC, or held-out test results appear anywhere. The 70/15/15 split is mentioned, yet the test set is never used. The models serve only as feature-extraction devices, and the resulting gene rankings are then called predictive. That's a logical gap, not a minor omission.\n\nWhat the paper does well: it assembles a genuinely new combination—off-the-shelf deep time-series models plus BDM network perturbation—applied to the public Lambo pediatric AML longitudinal scRNA-seq cohort, and it ships public code on GitHub. The authors also write an unusually candid limitation section, acknowledging sparse timepoints and the need for multi-omic validation. The specific gene lists (IGLL1, AQP9, H3F3A/B, etc.) may be of heuristic interest to people hunting plasticity markers.\n\nThe soft spots, in order of severity. First, no prediction is demonstrated, so 'predictive signatures' is unearned. Second, the analysis is entirely in-sample: 14 of 28 patients selected randomly with no fixed seed; pseudobulk aggregation; arbitrary binarization thresholds (0.1, 0.5); DEG cutoffs. Feature importance could simply reflect overfitting to a small, unseeded subset. Third, causal language ('causal drivers,' 'forecast') far outstrips what BDM perturbation scores can claim. Fourth, the 'ectoderm-mesoderm crosstalk' and 'brain-immune-hematopoietic axis' are speculative readings of gene-set enrichment, not findings. None of this is fatal if you treat the paper as exploratory biomarker discovery, but the abstract and conclusions do not present it that way.\n\nWho should read it: method developers or bioinformaticians who want candidate genes to test in independent pediatric AML cohorts. It's not clinical evidence.\n\nI'd send this to peer review—the question is timely and the data/code are public—but the editor should demand a version that actually reports held-out predictive performance, uses a seeded and justified patient split, and tempers the causal interpretations. If the authors can add one external validation, even a simple classifier on an independent cohort, the paper becomes much more interesting.","headline":"Claims predictive biomarkers for pediatric AML but never evaluates whether any model predicts—feature importance is not forecast.","tokens_in":28596,"tokens_out":3971,"would_cite":false,"duration_ms":49036,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that combining deep temporal models with algorithmic complexity measurements on longitudinal single-cell transcriptomes identifies plasticity markers that predict paediatric AML cell-fate transitions and expose stalled-dif","keywords":["paediatric acute myeloid leukemia","single-cell RNA sequencing","cell fate plasticity","algorithmic information dynamics","recurrent neural networks","trajectory inference","attractor landscape","differentiation therapy"],"falsifier":"Collect a dense time series (or lineage tracing) across relapse in a validation cohort and test whether the top BDM-shifted genes change expression state before relapse; then perturb them with CRISPR knockdown or activation and measure whether differentiation trajectories in AML cells shift accordingly. Negative results on either leg would refute the causal-plasticity claim.","tokens_in":27691,"feed_emoji":"🧬","tokens_out":5726,"duration_ms":64632,"temperature":0.7,"pith_summary":"The paper sets out to show that the trajectory of paediatric acute myeloid leukaemia—from diagnosis to remission to relapse—can be predicted from single-cell gene expression by combining recurrent and attention-based neural networks with algorithmic information dynamics, a complexity measure that scores how much each gene or interaction contributes to the structure of the gene regulatory network. Applied to pseudobulk longitudinal scRNA-seq from 14 patients, the framework identifies a small set of \"plasticity markers\" whose perturbation scores and neural-network importance ranks are consistently high, and interprets them as bifurcation signatures that steer cell fate decisions. The authors use these signatures to argue that AML cells are not fixed leukemic stem-cell clones but occupy stalled, reprogrammable attractor states: developmental arrest blocks terminal differentiation, and the inferred state space mixes haematopoietic, lymphoid-like, erythroid, and neurodevelopmental programs. If true, the same markers become predictive biomarkers for relapse and concrete targets for differentiation therapy—WNT modulation, histone demethylase inhibition, or CRISPR-based reprogramming.","feed_headline":"RNNs plus algorithmic complexity decode AML cell fate","feed_subtitle":"Deep learning and perturbation complexity scores converge on stalled-differentiation genes as therapy targets.","key_machinery":"The load-bearing object is the Block Decomposition Method (BDM) perturbation analysis, a graph-complexity estimator that approximates Kolmogorov complexity by decomposing gene correlation networks into small blocks and summing their algorithmic probabilities. Node- and link-based perturbations measure how much network complexity drops when a gene or interaction is removed, yielding a causal-influence score beyond what correlation, entropy, or centrality capture. This BDM signal is combined with feature-importance rankings from LSTM, BiLSTM, and Transformer models trained on the three clinical timepoints; the union of these rankings defines the plasticity signatures that the paper treats as r","core_discovery":"The central claim is that causal drivers of AML plasticity can be discovered by integrating deep learning with algorithmic information dynamics, without needing multi-omic input. Using longitudinal single-cell transcriptomes at diagnosis, remission, and relapse, the authors identify a convergent set of transition genes—H3F3A/B, KDM5B, TPT1, TLE1, S100A9, KCNE5, CD14, and others—as plasticity markers regulating cell fate bifurcations. They interpret these signatures as evidence that AML cells are arrested near a common myeloid progenitor-like attractor with biased megakaryocyte–erythroid and granulocyte–monocyte lineage branches, and that disrupted bioelectric, immune, and epigenetic signalin","pith_inferences":["If the three-state representation is faithful, the diagnosis-to-remission changes alone should already forecast relapse; this is a holdout prediction the paper does not report, and a direct way to test the framework's causal claim.","The neurodevelopmental/ectoderm–mesoderm reading suggests a concrete experiment: ask whether neural-crest transcription factors (SOX10, DLX6, POU3F2) drive the hybrid identity in AML cell lines, e.g., via CRISPRa of these factors followed by scRNA-seq trajectory inference.","The BDM perturbation scores could be benchmarked against simple network centralities and against random gene deletions; if BDM adds no predictive power, the specific algorithmic-complexity contribution would be weakened.","Pseudobulk aggregation erases cell-level heterogeneity, so a natural extension is to apply the same pipeline to single-cell pseudo-time ordering to see whether the same plasticity markers appear within individual clones.",""],"forward_implications":["The 20–30 consistently top-ranked genes are proposed as causal regulators of AML plasticity, not merely statistical correlates.","Targeting these markers—via WNT activation with GSK3β inhibition, KDM5B or LSD1 inhibition, H3F3A/B knockdown, TPT1 deletion, or KLF1 activation—should push leukemic cells out of arrested states toward terminal differentiation.","The same signatures can serve as relapse-risk biomarkers, since they shift between diagnosis, remission, and relapse and could be monitored in liquid biopsies for early recurrence detection.","The overlap between AML plasticity markers and paediatric high-grade glioma signatures implies a shared developmental program, making cross-cancer differentiation-therapy strategies plausible.","AML cell-fate decisions appear biased near a common myeloid progenitor branch point with megakaryocyte–erythroid and granulocyte–monocyte lineages, rather than being purely stochastic.",""],"supporting_citations":[{"why":"Supplies the primary longitudinal single-cell RNA-seq dataset (GSE235063) from paediatric AML patients at diagnosis, remission, and relapse, on which all trajectory inference is performed.","marker":"[29]"},{"why":"Provides the algorithmic information calculus and perturbation methodology used to score causal network changes (AID/BDM).","marker":"[68]"},{"why":"Defines the Block Decomposition Method that estimates algorithmic complexity of the adjacency matrices used for perturbation scoring.","marker":"[67]"},{"why":"Introduces causal deconvolution by algorithmic generative models, the basis for treating BDM shifts as causal, not merely correlational, signatures.","marker":"[66]"},{"why":"Shows that mutant H3 histones drive pre-leukemic expansion and AML aggressiveness, anchoring the paper's claim that H3F3A/B are plasticity-driving epigenetic regulators.","marker":"[5]"},{"why":"Establishes that paediatric brain tumours arise from stalled developmental programs, which the paper extends to AML to interpret neurodevelopmental signatures as attractor states.","marker":"[25]"},{"why":"Prior algorithmic reconstruction of glioblastoma network complexity that supplies the TPT1/KDM5B plasticity signatures the paper finds again in AML.","marker":"[55]"},{"why":"Frames cancer dynamics as attractor landscapes and justifies using complexity measures instead of static pseudotime methods to infer cell-fate transitions.","marker":"[56]"}],"fun_headline_variants":["RNNs + algorithmic info dynamics decode AML fate switches","AI finds plasticity markers that drive AML cell fate","Deep learning maps AML differentiation arrest via attractors","Symbolic AI flags genes that tip AML cells toward relapse","RNNs and complexity scores reveal AML stemness drivers"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that three clinical snapshots (diagnosis, remission, relapse) from 14 patients, after pseudobulk aggregation, are enough to reconstruct AML's dynamic state space; if sparse sampling or batch effects dominate, the inferred transitions and attractors are artifacts.","fun_headline_variants_meta":{"raw":{"variants":["RNNs + algorithmic info dynamics decode AML fate switches","AI finds plasticity markers that drive AML cell fate","Deep learning maps AML differentiation arrest via attractors","Symbolic AI flags genes that tip AML cells toward relapse","RNNs and complexity scores reveal AML stemness drivers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000746,"raw_usage":{"total_tokens":3182,"prompt_tokens":784,"completion_tokens":2398,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":2333}},"tokens_in":528,"tokens_out":2398,"duration_ms":18903,"temperature":1.0,"reasoning_tokens":2333,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:41:58.929936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a dense time series (or lineage tracing) across relapse in a validation cohort and test whether the top BDM-shifted genes change expression state before relapse; then perturb them with CRISPR knockdown or activation and measure whether differentiation trajectories in AML cells shift accordingly. Negative results on either leg would refute the causal-plasticity claim.","supporting_citations":[{"cited_title":"A longitudinal single-cell atlas of treatment response in pediatric AML","cited_arxiv_id":null,"evidence_quote":"Supplies the primary longitudinal single-cell RNA-seq dataset (GSE235063) from paediatric AML patients at diagnosis, remission, and relapse, on which all trajectory inference is performed."},{"cited_title":"An Algorithmic Information Calculus for Causal Dis- covery and Reprogramming Systems","cited_arxiv_id":null,"evidence_quote":"Provides the algorithmic information calculus and perturbation methodology used to score causal network changes (AID/BDM)."},{"cited_title":"A Decomposition Method for Global Evaluation of Shannon Entropy and Local Estimations of Algorithmic Complexity","cited_arxiv_id":null,"evidence_quote":"Defines the Block Decomposition Method that estimates algorithmic complexity of the adjacency matrices used for perturbation scoring."},{"cited_title":"Causal deconvolution by algorithmic generative models","cited_arxiv_id":null,"evidence_quote":"Introduces causal deconvolution by algorithmic generative models, the basis for treating BDM shifts as causal, not merely correlational, signatures."},{"cited_title":"Mutant H3 histones drive human pre-leukemic hematopoi- etic stem cell expansion and promote leukemic aggressiveness","cited_arxiv_id":null,"evidence_quote":"Shows that mutant H3 histones drive pre-leukemic expansion and AML aggressiveness, anchoring the paper's claim that H3F3A/B are plasticity-driving epigenetic regulators."},{"cited_title":"Algorithmic reconstruction of glioblas- toma network complexity","cited_arxiv_id":null,"evidence_quote":"Prior algorithmic reconstruction of glioblastoma network complexity that supplies the TPT1/KDM5B plasticity signatures the paper finds again in AML."},{"cited_title":"A Review of Mathemat- ical and Computational Methods in Cancer Dynamics","cited_arxiv_id":null,"evidence_quote":"Frames cancer dynamics as attractor landscapes and justifies using complexity measures instead of static pseudotime methods to infer cell-fate transitions."}],"review_version":1}