{"id":"5f05165e-5c24-463e-b7fa-7f81979908af","arxiv_id":"2509.02639","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A graph-based embedding method (DAE) combines random-forest leaf proximity with a k-nearest-neighbor expression graph and LINE node embedding to improve scRNA-seq cell representations.","lead":"This paper introduces a new way to turn single-cell gene expression data into a compact map of cells, combining the usual expression levels with patterns of how genes relate to each other as learned by random forests. The authors claim this dual view improves detection of rare cell types and downstream clustering across six benchmark datasets.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing ablation: no test that CLG leaf co-assignment contributes beyond KNNG/expression, so the 'gene-gene interaction' benefit is unsubstantiated.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the CLG's leaf-proximity signal may be merely a re-encoding of expression levels, and if so the claimed benefit of integrating gene-gene interactions collapses. I agree with that diagnosis. The paper is not internally inconsistent: it openly states that the regulatory network is not extracted and that CLG uses leaf assignments, but the absence of any ablation means the central claim is not currently supported by the evidence presented. The comparison to expression-only baselines is insufficient because those baselines differ in many respects from DAE (different embedding algorithms, graph construction, hyperparameters), so the observed improvements cannot be attributed specifically to the CLG. The sigma tuning issue flagged by the reader is related but secondary: even if sigma were fixed honestly, the missing CLG ablation would still leave the mechanism unverified. I do not see grounds for rejection, because the method is plausible and the evaluation is extensive; a focused ablation could resolve the concern. Therefore the existing CONDITIONAL verdict is appropriate and unchanged.","tokens_in":13816,"tokens_out":3810,"duration_ms":45986,"concrete_test":"Run a controlled ablation on Cortex and PBMC (ideally all six datasets): (A) DAE-full as reported; (B) DAE without CLG (LINE embedding on the KNNG only); (C) DAE with a degree-preserving random permutation of CLG cell-to-leaf edges, so graph size and degree are unchanged but any biological leaf-co-assignment signal is destroyed; (D) DAE with CLG built from shuffled expression rows or from PCA scores, to test whether leaf co-assignment carries signal beyond expression. Compare NNE (Tables 2 and 3), ARI/NMI (Figures 5 and 6), and a quantitative rare-cell separation metric on Cortex (e.g., F1 for the three rare types after clustering, or silhouette separability of microglia/ependymal/mural). Also compute the overlap between KNNG neighborhoods and CLG common-leaf co-assignment: if the Jaccard overlap is high or if variant (B) or (C) matches (A) within a small margin, the gene-gene interactio","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that integrating gene-gene interactions with expression improves cell embeddings, and this improvement is attributed to the Cell-Leaf Graph (CLG) built from random-forest leaf co-assignment (Section 2.1). The load-bearing assumption is that the CLG provides information distinct from the expression matrix. But this is never demonstrated. Each random forest predicts one gene from all other genes, so leaf membership is a deterministic function of the expression values themselves; the paper even states: 'the final feature importance ranking and regulatory network extraction is not performed in this study' (Section 2.1). Thus the CLG is not an extracted regulatory network, and cell-to-leaf co-assignment may simply re-encode expression-level similarities already captured by the KNNG. No experiment isolates the CLG's contribution: DAE is compared against expression-only methods such as PCA, SIMLR, scVI, and RAFSIL, but never against DAE without the CLG, or with a randomized/permuted CLG, or with the KNNG alone embedded by LINE. Without such an ablation, the reported NNE and ARI improvements, and the rare-cell separation in Figures 3 and 4, could be due entirely to LINE on the KNNG, the test-set-tuned sigma in Equation (1), or other implementation choices. If the CLG signal is redundant with expression similarity, the method's novelty and biological interpretation collapse; the claimed integration of 'gene-gene interactions' would be a relabeling of expression neighborhoods, not a new data modality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes Dual Aspect Embedding (DAE), a method for scRNA-seq cell embedding that combines a K-nearest neighbor graph (KNNG) built from gene expression profiles with a Cell-Leaf Graph (CLG) constructed from random-forest leaf co-assignment, then embeds the fused Enriched Cell-Leaf Graph (ECLG) using the LINE algorithm. DAE is evaluated on six datasets against PCA, SVD, kPCA, t-SNE, SIMLR, scVI, RAFSIL, Seurat, and Monocle, using nearest-neighbor error (NNE), clustering ARI/NMI, visualization, sensitivity analyses, and runtime. The paper claims improved rare-cell identification, clustering, and trajectory inference, with the main technical novelty attributed to the CLG's capture of gene-gene interactions.","tokens_in":14200,"tokens_out":4535,"duration_ms":51335,"significance":"If the CLG genuinely adds information beyond expression-based similarities, the idea of integrating tree-based leaf co-assignment with a KNNG for cell embedding is a useful and potentially generalizable contribution. The paper includes a reasonably broad six-dataset benchmark and a runtime analysis on a 76,899-cell dataset, which are strengths. However, the current evidence does not establish the core claim: the reported NNE improvements rely on tuning the RBF bandwidth sigma on the same test datasets, no ablation isolates the CLG's contribution, and the abstract's trajectory-inference claim is not supported by any trajectory experiment. The central idea is plausible, but the evaluation needs substantial additional work.","major_comments":[{"comment":"The RBF bandwidth sigma is set to 0.3 by \"minimizing the Nearest Neighbor Error (NNE) across our experimental datasets,\" and Tables 2 and 3 then report NNE on those same datasets. This is test-set tuning: the reported NNE values are not independent estimates and are optimistically biased relative to the baselines, whose parameters are not tuned in the same way. Please use a separate calibration set or nested cross-validation for sigma selection, report the selected sigma per dataset, and include uncertainty estimates.","section":"Section 2.2, Eq. (1)"},{"comment":"The central claim is that the CLG contributes gene-gene interaction information beyond expression profiles, but no experiment isolates the CLG's contribution. Random-forest leaf co-assignment is a deterministic function of the same expression matrix used for the KNNG, and the paper itself states in Section 2.1 that \"the final feature importance ranking and regulatory network extraction is not performed in this study.\" Without ablations comparing DAE to (i) KNNG alone embedded with LINE, (ii) CLG alone, and (iii) ECLG with permuted leaf assignments or shuffled interaction scores, the improved NNE/ARI and rare-cell separation in Tables 2-3 and Figures 3-5 could be due entirely to LINE on the KNNG or other implementation choices. This ablation is necessary to substantiate the biological interpretation that gene-gene interactions drive the improvement.","section":"Sections 2.1 and 2.3"},{"comment":"The abstract and conclusion state that DAE improves \"downstream analyses such as visualization, clustering, and trajectory inference,\" but Section 3 contains no trajectory inference experiments. The evaluation covers similarity learning (Section 3.1), visualization (Section 3.2), and clustering (Section 3.3), with no trajectory reconstruction or ordering metrics. Either add trajectory-reconstruction experiments on suitable datasets or remove the trajectory-inference claim from the abstract and conclusion.","section":"Abstract and Section 3"},{"comment":"The reported rare-cell percentages for the Cortex dataset (microglia 0.03%, ependymal 0.008%, mural 0.02%) are inconsistent with the dataset size of 3,005 cells; 0.008% would correspond to about 0.24 cells. These percentages appear to be off by one or two orders of magnitude. Please provide exact cell counts and correct the percentages, since the quantitative rare-cell-detection claim and the interpretation of Figure 3 depend on them.","section":"Section 3.2"}],"minor_comments":[{"comment":"Seurat and Monocle are evaluated with UMAP while the other methods are evaluated with t-SNE applied to their embeddings. This conflates the embedding method with the visualization algorithm. Please use the same visualization algorithm for all methods, or clearly justify why UMAP is only used for these two baselines.","section":"Section 3.2 and Table 3"},{"comment":"No error bars, standard deviations, or confidence intervals are reported for the main NNE values. Given the stochasticity in random forests and LINE, please report repeated-run variability or at least summarize the seed-to-seed variation mentioned in Section 2.5.","section":"Tables 2 and 3"},{"comment":"Kmeans++ is given the correct number of clusters, while SC3, Phenograph, Seurat, and Monocle are unsupervised with respect to cluster count. Comparisons of ARI/NMI across clustering methods with different amounts of oracle information should be interpreted cautiously and this limitation should be stated.","section":"Section 3.3"},{"comment":"Figure 4 reports \"gene-gene interaction scores\" for rare cell types, but Section 2.1 states that feature-importance ranking and regulatory-network extraction are not performed in this study. Please clarify how these interaction scores were computed and whether they come from the same CLG used in the embedding or from a separate post-hoc analysis.","section":"Figure 4 and Section 2.1"},{"comment":"Several presentation issues need correction: \"Monoloce\" in the Figure 6 caption, \"Campbel\" in Table 1, and duplicated/malformed text at the end of reference list entry [41]. Also, the paper does not provide code or a data-availability statement; please add one.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be an Author Accepted Manuscript already published in Computers in Biology and Medicine, which is not itself a problem for this review. The main concerns are the test-set tuning of sigma, the missing ablation for the CLG, and the unsupported trajectory-inference claim; these are load-bearing and require additional experiments, not merely copyediting. I would not recommend acceptance without the ablation and evaluation corrections."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on the DAE paper. The method is a plausible re-packaging of existing ingredients, but the paper's central biological claim—that it integrates gene-gene interactions—is not actually demonstrated, and the main evaluation has a circular hyperparameter choice.\n\nWhat's actually new is the specific combination: random forest leaf co-assignment (from RAFSIL) to build a cell-leaf graph, fused with a KNNG, then embedded with LINE. That exact pipeline hasn't been published before, and the large-scale version using gene-cluster features is a sensible engineering contribution. The benchmarking is extensive: six datasets, several similarity and clustering metrics, and a runtime analysis. The results are consistently in DAE's favor, albeit with modest gains over RAFSIL and scVI in some cases.\n\nThe soft spots are real. The RBF kernel bandwidth sigma is set by minimizing NNE on the same six datasets, then NNE is reported on those datasets. That's fitting, not prediction, and it weakens all the comparative statements. There are no error bars in the main tables. More importantly, there is no ablation: the paper never tests DAE without the CLG, or with a permuted CLG. Since the CLG is built from random forests on the same expression data, it may simply re-encode the expression similarities already in the KNNG. The authors even state that no regulatory network is extracted, so calling it \"gene-gene interaction\" is a stretch. The abstract also promises trajectory inference, but there are no trajectory experiments in the paper.\n\nIf the goal is a practical embedding tool, the method is usable and the code seems straightforward to reproduce. If the goal is a claim about capturing regulatory information, the evidence is missing. A serious referee should ask for an ablation, a proper validation scheme for hyperparameters, and a correction of the trajectory claim.\n\nI'd send it to review, but as a conditional accept with major revisions. It's a useful incremental contribution to a crowded field. I wouldn't cite it in my own work in the next year, but I'd bring it to a reading group as an example of evaluation pitfalls in embedding papers.","headline":"A plausible embedding pipeline with an overstated biological story and a circular hyperparameter choice; worth refereeing, but the central claim needs an ablation.","tokens_in":14673,"tokens_out":2345,"would_cite":false,"duration_ms":25180,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes Dual Aspect Embedding (DAE), which fuses gene expression similarity with random-forest-derived gene-gene interaction information to produce cell embeddings that better identify rare cell types and improve clustering.","keywords":["single-cell RNA-seq","cell embedding","gene-gene interaction","random forest similarity","cell-leaf graph","graph embedding","rare cell type detection","similarity learning"],"falsifier":"Repeat the Cortex experiment after shuffling cell-to-leaf assignments in the CLG while keeping the graph's degree sequence and the KNNG unchanged; if the Nearest Neighbor Error and microglia/ependymal/mural separation stay as low as the real DAE's, then the leaf-proximity channel carries no unique biological signal.","tokens_in":13751,"feed_emoji":"🧬","tokens_out":5040,"duration_ms":50961,"temperature":0.7,"pith_summary":"The paper is trying to establish that single-cell RNA-seq embeddings work better when they encode not only how much each gene is expressed but also how genes regulate one another. It derives the regulatory view from the expression data itself: random forests trained to predict each gene's expression place cells in leaf nodes, and cells sharing leaves are treated as having similar gene-gene interaction profiles. These leaf relationships are fused with an ordinary nearest-neighbor graph and embedded with LINE. If the claim holds, researchers get a purely data-driven way to sharpen clustering, visualization, and rare-cell detection without needing external interaction databases.","feed_headline":"Gene-gene links in cell embeddings reveal rare cell types","feed_subtitle":"Random-forest leaf co-assignment plus expression similarity out-embeds expression-only methods on six scRNA-seq datasets.","key_machinery":"Enriched Cell-Leaf Graph (ECLG): a fused graph in which each cell is linked, with unit weight, to the random-forest leaf nodes it falls into (the Cell-Leaf Graph, a bipartite graph capturing regulatory similarity), then augmented with expression-based K-nearest-neighbor edges weighted by a Gaussian RBF kernel. LINE embedding on this graph produces cell vectors that preserve both first-order and second-order proximities.","core_discovery":"The paper claims that a cell embedding can be improved by fusing two views of the same scRNA-seq data: the usual expression-similarity view and a regulatory view inferred by random forests. For each gene, a random forest is trained to predict its expression from all other genes; cells that land in the same or nearby leaf nodes across the forest are taken to share gene-gene interaction context. The two views are combined into a single graph and embedded with LINE. Across six datasets, this Dual Aspect Embedding (DAE) reports lower nearest-neighbor error than PCA, SVD, kPCA, t-SNE, SIMLR, scVI, and RAFSIL, and it keeps rare cell types (microglia, ependymal, mural) visually separated in the mou","pith_inferences":["A testable consequence the paper does not report: removing the KNNG edges should degrade performance mostly for abundant cell types, while removing the CLG edges should degrade performance mostly for rare types; that asymmetry would confirm the two channels carry different information.","The same leaf-co-assignment construction could be plugged into other graph-based single-cell tools, for example as an additional similarity kernel, without retraining the rest of the pipeline.","Because the method learns interactions from the data, it could be extended to multi-omics settings by building leaf graphs on each modality; the paper lists this only as a future direction.","The reported rare-cell separation in Cortex is demonstrated visually rather than by a quantitative recall or precision metric, so a quantitative rare-cell detection benchmark would be a natural next check."],"forward_implications":["On six benchmark datasets, DAE achieves lower nearest-neighbor error than expression-only and deep-learning baselines, in both direct embedding and t-SNE visualization settings.","Rare cell populations — microglia, ependymal, and mural in mouse cortex — are separated more cleanly in DAE's embedding, which the paper attributes to preserved regulatory interactions.","Clustering algorithms applied on top of DAE embeddings give higher adjusted Rand index and normalized mutual information than on PCA/SVD/kPCA embeddings for most clustering methods tested.","The method requires no external gene-interaction database: the interaction layer is inferred from the same expression matrix, and performance plateaus with about 200 trees per forest.","The feature-cluster variant makes the approach runnable on a standard personal computer, with random-forest construction being parallelizable."],"supporting_citations":[{"why":"Supplies the random-forest similarity-learning approach and the gene-cluster feature-engineering module that DAE adapts.","marker":"[13]"},{"why":"Provides the tree-based gene regulatory network inference method whose leaf-assignment structure DAE repurposes.","marker":"[16]"},{"why":"Benchmarks gene regulatory network inference from single-cell transcriptomic data, supporting the use of the tree-based approach.","marker":"[17]"},{"why":"Defines random forests, the model class used to construct the Cell-Leaf Graph.","marker":"[22]"},{"why":"Defines K-nearest-neighbor graphs, the structure used for the expression-similarity view.","marker":"[23]"},{"why":"Guides the KNN graph construction via PCA preprocessing and supplies a standard single-cell clustering baseline.","marker":"[25]"},{"why":"Supplies the LINE graph embedding algorithm used to compute cell embeddings on the Enriched Cell-Leaf Graph.","marker":"[28]"},{"why":"Source of the mouse cortex and hippocampus dataset with the rare cell types used for the rare-cell claim.","marker":"[30]"},{"why":"Baseline kernel-based similarity learning method that DAE must outperform in the comparison.","marker":"[8]"},{"why":"Baseline deep generative embedding method that DAE is compared against.","marker":"[9]"}],"fun_headline_variants":["Dual-view graph boosts rare cell detection in scRNA-seq","Random forests link genes to sharpen cell embeddings","Cell-leaf graphs reveal rare cell types in single-cell data","Integrating gene-gene networks improves cell embedding","Gene-expression plus interactions: better cell embedding"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"That leaf co-assignment in random forests trained to predict each gene from the others captures gene-gene interaction structure that is genuinely distinct from the expression profiles themselves, rather than a redundant re-encoding of the same expression similarities.","fun_headline_variants_meta":{"raw":{"variants":["Dual-view graph boosts rare cell detection in scRNA-seq","Random forests link genes to sharpen cell embeddings","Cell-leaf graphs reveal rare cell types in single-cell data","Integrating gene-gene networks improves cell embedding","Gene-expression plus interactions: better cell embedding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000142,"raw_usage":{"total_tokens":1021,"prompt_tokens":775,"completion_tokens":246,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":172}},"tokens_in":519,"tokens_out":246,"duration_ms":3495,"temperature":1.0,"reasoning_tokens":172,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:10:22.222540+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the Cortex experiment after shuffling cell-to-leaf assignments in the CLG while keeping the graph's degree sequence and the KNNG unchanged; if the Nearest Neighbor Error and microglia/ependymal/mural separation stay as low as the real DAE's, then the leaf-proximity channel carries no unique biological signal.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the random-forest similarity-learning approach and the gene-cluster feature-engineering module that DAE adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the tree-based gene regulatory network inference method whose leaf-assignment structure DAE repurposes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Benchmarks gene regulatory network inference from single-cell transcriptomic data, supporting the use of the tree-based approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines K-nearest-neighbor graphs, the structure used for the expression-similarity view."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Guides the KNN graph construction via PCA preprocessing and supplies a standard single-cell clustering baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the LINE graph embedding algorithm used to compute cell embeddings on the Enriched Cell-Leaf Graph."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the mouse cortex and hippocampus dataset with the rare cell types used for the rare-cell claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Baseline kernel-based similarity learning method that DAE must outperform in the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Baseline deep generative embedding method that DAE is compared against."}],"review_version":1}