{"id":"ded21564-35f8-4eed-9953-d44fc3071b77","arxiv_id":"2507.22212","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"On a single mouse ileum MERFISH sample, VAE-derived stain embeddings modestly improved cluster homogeneity when combined with gene counts via multiplex Leiden clustering, while CellProfiler features did not.","lead":"This study tests whether simple variational autoencoders can pull useful biological signal from cell images (nuclear and membrane stains) in a small spatial transcriptomics dataset, and whether adding those learned features to gene counts improves cell-type clustering. A smart generalist should read it because it gives practical evidence on when cheap image features help or hurt in spatial biology, and suggests classical image features may lag behind learned ones.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The multiplex-clustering gain may be an artifact of the modality-weight optimization: the Figure 5 caption calls the reported weight the 'TC modality weight' while the text calls it the morphological-feature weight; if the caption is correct, the clustering is LS2-dominated, not TC augmented by LS2.","rationale":"I read the paper as a careful proof-of-concept, and the authors honestly caveat the homogeneity trade-off (lower AMI, lower label recovery) in the discussion. The strongest claim, however, is that multiplex clustering of TC with LS2 improves cluster homogeneity, and the gene-subset results are a key part of that claim. The manuscript contains a direct internal inconsistency about what the optimized weight means: the text says the 3%/11%/5% values are 'small contributions from morphological features,' while the Figure 5 caption labels the same quantity as the 'TC modality weight.' If the caption is literally correct, the multiplex graph is dominated by LS2, and the result does not show TC being augmented by a small morphological supplement; it shows LS2 dominating the graph after TC has been reduced to 19–25 genes. This is not an external-consensus disagreement but an internal consistency issue that determines whether the headline claim is supported. The reader's weakest assumption about ground-truth label quality is a real secondary limitation, but it is less immediately decisive because label noise would more likely attenuate than fabricate a stain-derived signal. The code is available, which makes the proposed test straightforward. I therefore do not see reason to change the conditional verdict; the concern should be resolved before the abstract's general phrasing is accepted.","tokens_in":27846,"tokens_out":11077,"duration_ms":139435,"concrete_test":"Inspect the released GitHub code to determine the exact definition of the reported weight: is it the weight assigned to the TC layer, to the LS2 layer, or to the morphological feature layer? Then, using the PCA-based 25-gene subset, rerun multiplex Leiden clustering with the TC weight fixed at 0.5, 0.7, and 0.9 (with LS2 weight = 1 - TC weight), and recompute homogeneity, AMI, and label-matched Jaccard scores against the ground-truth partition. If the homogeneity gain over TC-only disappears or reverses at high TC weight, the central claim should be revised; if it persists, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central integration claim depends on showing that adding a small amount of stain-derived information to transcript counts improves cluster homogeneity. In the gene-subset experiments (Results, 'Joint clustering of transcript counts and latent spaces enhanced cluster homogeneity'), the authors report optimized weights of 3%, 11%, and 5% for the GT-based, PCA-based, and random subsets, respectively, and interpret these as 'small contributions from morphological features.' However, the Figure 5 caption in the manuscript labels the same displayed quantity as the 'TC modality weight.' These two statements are mutually inconsistent. If the caption is correct, then in the multiplex graph the TC layer carries only 3–11% weight and LS2 carries 89–97%, meaning the clustering is effectively LS2-only with a tiny TC perturbation. In that case, the reported homogeneity improvement over TC-only is not evidence that 'augmenting only subsets of TC with LS2' helps; it is evidence that replacing most of the (degraded) TC signal with LS2 changes the partition. The same ambiguity could affect Figure 4 for the full gene set, where the TC modality weight is also displayed. Because the abstract and lay description generalize this result, the internal inconsistency is load-bearing.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript benchmarks strategies for integrating phenotypic imaging features with spatial transcriptomics using a public MERFISH mouse ileum dataset (5,192 cells, 241 genes, 19 annotated cell types). The authors train five VAEs (transcript-only, stain-only, and three multimodal combinations) with a binary multi-plane spot representation, extract latent spaces (LSs), and evaluate them via PERMANOVA/PERMDISP, random forest classification, unsupervised Leiden clustering, multiplex Leiden clustering with transcript counts (TC), and ridge regression for predicting gene expression from a stain-only LS. They report that TC generally outperforms all feature spaces, that LS2 captures some cell-type and gene-expression signal, that CellProfiler features underperform, and that multiplex clustering of TC with LSs yields higher homogeneity scores than TC alone, including when TC is restricted to gene subsets. The paper concludes that minimally tuned VAEs can extract biologically meaningful morphological signals under real-world constraints.","tokens_in":28117,"tokens_out":5957,"duration_ms":67295,"significance":"If the findings hold, this is a useful, reproducible benchmark (code and data are made available) for a practical question: can routine nuclear/membrane stains improve cell-type clustering in imaging-based spatial transcriptomics? The study's strengths include the use of multiple independent evaluation modalities, the explicit treatment of dispersion effects via PERMDISP, and the honest discussion of class-imbalance effects on the homogeneity metric. The main claims are plausible and appropriately restrained in the Discussion, though the abstract and lay summary are more assertive than the evidence warrants given the unresolved weight interpretation and the absence of repeated-run variability.","major_comments":[{"comment":"The manuscript contains a direct contradiction about what the optimized multiplex-clustering weights represent. The Results text reports optimized weights of 3%, 11%, and 5% for the GT-based, PCA-based, and random gene subsets and interprets them as 'small contributions from morphological features,' i.e., the weight of LS2 in the multiplex graph. The Figure 5 caption (and the Figure 4 caption) labels the same displayed quantity as the 'TC modality weight.' These cannot both be correct. If the captions are correct, then TC carries 3–11% of the total graph weight and LS2 carries 89–97%, so the clustering is effectively LS2-dominated with a minor TC perturbation; the homogeneity improvement over TC-only would not support the abstract's claim that 'augmenting only subsets of TC with the stain-derived LS2' produces gains. If the text is correct, the captions must be changed and the actual TC weights should be reported. Because the central integrative claim depends on this quantity, the authors must clarify the definition and report both weights.","section":"Results, 'Joint clustering of transcript counts and latent spaces enhanced cluster homogeneity'; Figure 5 caption"},{"comment":"The central claim of 'consistent gains' in homogeneity is supported only by point estimates from single clustering runs. The VAE is trained once on the full dataset (transductively), and each Leiden clustering and metric is reported without any repeated initialization, bootstrap interval, or alternative graph-construction settings. Since the homogeneity improvements in Figures 4 and 5 are modest in magnitude and the number of clusters changes from 16 to 6–8, the authors should provide variability across runs (e.g., different random seeds for the VAE sampling and Leiden initialization) or otherwise demonstrate that the gain is robust rather than a single realization. This is needed specifically because the Discussion attributes the homogeneity boost to improved recovery of the dominant mid-villus enterocyte population.","section":"Results, 'Unsupervised Leiden clustering...' and Figures 4–5"},{"comment":"For the multiplex clustering with LS3–5, the latent spaces were trained on the same transcript counts that are subsequently clustered together with TC. Any homogeneity gain in these combinations could partly reflect duplicated information rather than complementary signal. The paper does not address this circularity. The LS2 result is independent and remains the cleanest evidence, but the abstract generalizes the integration claim beyond LS2. Please state this limitation explicitly and, if possible, provide a control where the VAE is trained on a held-out subset of cells or on shuffled TC labels to show that the multiplex gain is not an artifact of information duplication.","section":"Methods, 'Models' (transductive training) and Results, multiplex clustering with LS3–5"}],"minor_comments":[{"comment":"The text states that filtering genes with fewer than 100 transcripts results in 196 genes for VAE3–5, but Table 1 lists 182 genes for all three multimodal models. Please correct the inconsistency.","section":"Table 1 and Methods, 'Models'"},{"comment":"The Focal loss formula uses α_t, γ, and p_t without defining these symbols in the text or in the equation; please add definitions.","section":"Equation 2"},{"comment":"The phrase 'TC modality weight' in the caption conflicts with the main text's description of the same quantity as the morphological-feature weight; this is already raised as a major point, but the caption should be corrected in any case.","section":"Figure 5 caption"},{"comment":"In the description of PERMANOVA, the combination of scikit-learn, scikit-bio, and SciPy is mentioned; please clarify which package implements the permutation test and how the distance matrix is standardized.","section":"Methods, 'Statistical analysis and modeling'"},{"comment":"The statement that the approach 'uncovered structure that was not captured by TC alone' refers to a stromal-cell cluster in TC+LS2, but the preceding sentences note a decline in overall label recovery; consider rephrasing to avoid overstatement.","section":"Discussion, paragraph on multiplex clustering"}],"recommendation":"major_revision","confidential_remarks":"The weight/caption ambiguity is the key issue; if the captions are correct, the manuscript's headline claim about augmenting TC with small amounts of LS2 is not supported by the presented analysis. The authors should also consider whether the term 'deconvolution' in the keywords and abstract accurately describes the unsupervised clustering tasks performed. The single-sample design is a limitation the authors acknowledge, and the paper's framing as an 'evaluation' is appropriate, but the abstract should be moderated accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new piece here is the head-to-head comparison of minimally tuned VAE latent spaces against CellProfiler features for integrating nuclear/membrane stains with MERFISH transcript counts under realistic constraints: 5k cells, 19 imbalanced cell types, and bounding-box crops instead of precise segmentation. The main positive result—stain-only LS2 improves multiplex cluster homogeneity and predicts a handful of genes—is backed by several independent analyses (PERMANOVA, random forest, ridge regression, clustering), and the authors are candid about trade-offs like lower AMI and worse label recovery. Code is public. That is a solid proof-of-concept.\n\nThe soft spots are real but mostly minor. One sample, transductive VAE training with no repeated runs or error bars, and ground-truth labels taken from the original MERFISH study without independent validation. These limit confidence but do not sink the paper; the authors acknowledge them.\n\nThe load-bearing issue is the weight ambiguity in the multiplex clustering. In the gene-subset section the text reports optimized weights of 3%, 11%, and 5% and interprets them as 'small contributions from morphological features.' The Figure 5 caption (and Figure 4) labels the same quantity as the 'TC modality weight.' If the caption is correct, the TC layer carries only 3–11% of the weight and LS2 carries 89–97%, meaning the clustering is LS2-dominated, not TC augmented with a dash of morphology. That turns the headline 'augmenting subsets of TC with LS2' into 'replacing most of the degraded TC signal with LS2.' The authors need to clarify which weight they optimized and how it enters the multiplex graph. If the small-morphology-weight reading is intended, the captions are wrong; if the captions are right, the text and abstract overstate the case. Either way, it must be fixed before the paper is usable. I do not think this is fatal—there is independent evidence that LS2 carries biological signal—but it is exactly the kind of inconsistency that should be caught in review.\n\nWho is this for? Anyone working on integrating image morphology with spatial transcriptomics, especially in the small-data regime. It will not settle the open question, but it is a careful, reproducible data point. I would send it to review: the code and multi-metric evaluation deserve a serious referee, and the weight issue is fixable. I would want to see the corrected version before relying on the headline claim.","headline":"Useful proof-of-concept on VAE-based morphology integration in spatial transcriptomics, but a weight-labeling inconsistency in the multiplex clustering undermines the headline claim until clarified.","tokens_in":28663,"tokens_out":2478,"would_cite":false,"duration_ms":28850,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under constrained real-world conditions, minimally tuned VAEs extract biological signal from stain images, and multiplex Leiden clustering with these latent spaces consistently improves cluster homogeneity.","keywords":["spatial transcriptomics","multi-modal integration","variational autoencoder","cell type deconvolution","morphological features","MERFISH","Leiden clustering","latent space"],"falsifier":"Re-run the identical VAE training and multiplex Leiden clustering on a second independently annotated spatial transcriptomics sample (or on this sample re-annotated with a different marker list). If the homogeneity gain from TC+LS2 over TC-only does not reproduce, or if the gain persists when cell-type labels are randomly shuffled, the central claim is falsified.","tokens_in":27602,"feed_emoji":"🔬","tokens_out":8019,"duration_ms":88716,"temperature":0.7,"pith_summary":"Spatial transcriptomics measures gene expression in tissue while preserving spatial context, but no standard method exists for adding the paired nuclear and membrane stain images that routine workflows already capture. This paper asks whether a minimally tuned variational autoencoder (VAE) can turn such images into usable cell representations under realistic constraints: one small dataset, heavily imbalanced cell types, and bounding-box crops instead of precise cell boundaries. Using a public MERFISH mouse ileum sample with 5,192 cells, 241 genes, and 19 annotated cell types, the authors report that VAE-derived latent spaces capture meaningful biological variation, with a stain-only latent space even predicting a small set of genes. The central result is that multiplex Leiden clustering, which combines transcript counts with these latent spaces, consistently improves cluster homogeneity compared with transcript counts alone, and classical image-profiling features fail to match the learned representations. If this holds, even simple deep learning models can make practical use of routine stains to sharpen cell-type identification in spatial transcriptomics.","feed_headline":"Stain-derived latent spaces consistently improve cell cluster purity","feed_subtitle":"Adding stain-derived features to gene expression improves cell-type cluster purity in spatial transcriptomics.","key_machinery":"The load-bearing mechanism is a convolutional variational autoencoder whose inputs are multi-plane binary spot maps (each gene as a plane, with spots expanded into local patches) and/or paired nuclear and membrane stain crops, mapped through a single encoder to a shared 50-dimensional latent space. Reconstruction is trained with a Dice-plus-Focal loss with dynamic gene weighting to counter extreme sparsity, and the KL term is omitted so the latent space prioritizes reconstruction fidelity over regularization. The integration step uses multiplex Leiden clustering, which builds layered k-nearest-neighbor graphs with a shared vertex set and distinct edge sets for each modality and optimizes a modality weight; that is the machinery that produces the reported homogeneity gains.","core_discovery":"On the paper's own terms, the discovery is that a standard convolutional VAE with a 50-dimensional latent space, trained without systematic hyperparameter optimization and evaluated through sampling rather than mean embeddings, can extract biologically informative features from cell crops in imaging-based spatial transcriptomics. Transcript counts remain the strongest feature space, but the latent spaces recover specific cell types—smooth muscle cells, mid-villus enterocytes, stromal cells, and CD8+ T cells—and a latent space trained only on nuclear and membrane stains can predict a handful of genes, including Acta2 and Neat1. The decisive finding is integrative: multiplex Leiden clustering of transcript counts with the VAE latent spaces (LS2–LS5) raises the homogeneity score of the resulting clusters even when the transcript input is reduced to small gene subsets, with optimized morphology weights of roughly 3–11%. In contrast, classical image-profiling features underperform across all evaluations, indicating that learned representations, not hand-crafted ones, carry the usable morphological signal in this constrained setting.","pith_inferences":["An extension the paper leaves implicit is applying the same VAE-plus-multiplex-Leiden recipe to multi-sample or disease-versus-health comparisons, where the homogeneity gain should be tested for whether it helps detect disease-associated cell states rather than just the dominant healthy enterocyte population.","Because the homogeneity improvement tracks the abundant mid-villus enterocyte population, the method may preferentially aid major cell types; rare-type recovery might require class-balanced training or a different clustering objective.","The near-zero optimized morphology weights hint that the latent morphology signal acts as a consistent prior or regularizer rather than a competing data source, so weight selection could potentially be fixed heuristically instead of grid-searched in future applications.","A cheap falsifiable extension is to shuffle the ground-truth labels within the same dataset and verify that the multiplex homogeneity gain disappears; if it persists under shuffled labels, the claimed signal would be an artifact of clustering bias rather than biology."],"forward_implications":["If the central claim holds, imaging-based spatial transcriptomics can be augmented with routine nuclear and membrane stains without requiring precise cell segmentation, since bounding-box crops sufficed for the reported gains.","Learned morphological embeddings can replace hand-crafted image features as the default complement to transcript counts, since the classical features collapsed when multiplexed with transcript data.","A stain-only latent representation can serve as a weak gene-expression predictor for a few morphologically or positionally anchored transcripts, which may help flag candidate genes for targeted follow-up in samples without a full transcript panel.","Small modality weights (around 3–11%) are enough for morphology to improve cluster purity, suggesting the integration benefit is real but modest and should not be expected to dominate the transcriptional signal.","The approach transfers to reduced gene panels (PCA-based, ground-truth-based, or random subsets), so it could be relevant to targeted or cost-limited spatial assays."],"supporting_citations":[{"why":"Supplies the mouse ileum MERFISH dataset and the segmentation derived with membrane-stain priors that all experiments build on.","marker":"[24]"},{"why":"Provides the public MERFISH measurements of the mouse ileum as the data source.","marker":"[25]"},{"why":"Defines the variational autoencoder formulation underlying the latent-space models.","marker":"[34]"},{"why":"Introduces the Dice loss used to handle the sparse multi-plane spot reconstruction.","marker":"[37]"},{"why":"Introduces the Focal loss combined with Dice to counter extreme class imbalance in spot masks.","marker":"[38]"},{"why":"Defines the Leiden algorithm used for both individual and multiplex clustering.","marker":"[47]"},{"why":"Implements the classical image-feature profiling baseline that the VAE latent spaces are compared against.","marker":"[26]"},{"why":"Provides the standard single-cell preprocessing, k-NN graph construction, and UMAP embedding infrastructure.","marker":"[31]"}],"fun_headline_variants":["Stain latents lift cluster purity in spatial transcriptomics","Morphology VAE sharpens cell-type clusters","Learned morphology features beat hand-crafted in ST","Multiplex clustering gains from stain-derived latents","VAE morphology enhances transcriptomic cluster homogeneity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation treats the published 19-cell-type labels of this single mouse ileum sample as ground truth for every metric; if those labels are systematically wrong, the claim that the latent spaces carry biological signal and improve cluster purity does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Stain latents lift cluster purity in spatial transcriptomics","Morphology VAE sharpens cell-type clusters","Learned morphology features beat hand-crafted in ST","Multiplex clustering gains from stain-derived latents","VAE morphology enhances transcriptomic cluster homogeneity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000621,"raw_usage":{"total_tokens":2928,"prompt_tokens":1046,"completion_tokens":1882,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":1807}},"tokens_in":662,"tokens_out":1882,"duration_ms":17767,"temperature":1.0,"reasoning_tokens":1807,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:57:36.661856+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the identical VAE training and multiplex Leiden clustering on a second independently annotated spatial transcriptomics sample (or on this sample re-annotated with a different marker list). If the homogeneity gain from TC+LS2 over TC-only does not reproduce, or if the gain persists when cell-type labels are randomly shuffled, the central claim is falsified.","supporting_citations":[],"review_version":1}