{"id":"2feedff0-8075-449a-97c4-9ad7be1f638c","arxiv_id":"2412.01561","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Existing spatial statistics methods (point pattern and lattice analysis) are adapted and demonstrated for spatial omics data, supported by a new tutorial resource called pasta.","lead":"A data-analysis guide shows how classic spatial statistics tools can be applied to modern spatial omics data, with worked examples in breast cancer and pancreas tissue. It also releases pasta, a free R and Python tutorial collection that walks analysts through point pattern and lattice methods step by step.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Triple-positive receptor map in the Xenium breast cancer example rests on an unsupported k=6 neighbourhood choice, and the paper itself shows such choices shift local Moran's I; a sensitivity check is needed.","rationale":"I read the paper as a perspective and reproducible resource whose central assertion is that classical spatial statistics, split into point pattern and lattice streams, can quantify biological phenomena in spatial omics data. The two main worked demonstrations are the triple-positive receptor analysis on Xenium breast cancer data and the DCIS/invasive tumour co-localization analysis on the same dataset. The reader's verdict identified the neighbourhood choice for local Moran's I as the weakest assumption, and I agree that this is the most load-bearing concern. The paper explicitly tells the reader that weight-matrix construction is critical and shows, on CosMx data, that different neighbourhood rules change local Moran's I values; yet the Xenium triple-positive classification uses k=6 with no sensitivity check or discussion of why six is appropriate. Since the triple-positive map is produced by thresholding local statistics, it is plausible that other reasonable neighbourhood definitions would reclassify cells near boundaries, changing the visual and quantitative claim of recapitulating Janesick et al. This is a genuine weakness, but not a rejection: the paper is a perspective, the analyses are illustrative, the vignette is shipped, and the broader conceptual points about point patterns versus lattice data do not depend on this one result. The reader's CONDITIONAL verdict is therefore appropriate. I do not see an additional concern that is more load-bearing than this neighbourhood sensitivity issue. The point pattern analysis of DCIS invasion does choose a data-derived restricted window, which is another potential sensitivity, but the paper compares homogeneous, inhomogeneous, and restricted-window L-functions in Supplementary Figure S3, so that choice is at least partially explored. The strength of the paper lies in its clear mapping of technologies to data modalities and its reproducible tutorial resource, which should be credited. The missing sensitivity analysis on the Xenium neighbourhood definition is best addressed by the concrete recomputation described above, or by an explicit statement that the triple-positive map is an illustration of one defensible but arbitrary choice.","tokens_in":23325,"tokens_out":4036,"duration_ms":40451,"concrete_test":"Recompute the Xenium breast cancer analysis (Figure 2B) from the pasta vignette using identical log-normalized counts and multiple weight matrices: k=6 as published, k=10, k=20, a fixed-radius graph at roughly 30 µm, and a contiguity/Delaunay graph. For each, recompute local Moran's I with the same permutation-based BH-adjusted p-values and classify cells as triple-positive only if all three receptors are significant high-high in the Moran scatter plot. Report the Jaccard overlap of the triple-positive cell set relative to k=6 and the fraction of cells in the arrow-indicated DCIS region that remain triple-positive under each choice. If overlap stays above roughly 90% and the DCIS region remains predominantly triple-positive, the finding is robust; if overlap drops substantially, the recapitulation claim needs a sensitivity analysis or explicit caveat.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most concrete biological claim is that spatial statistics recapitulate the Xenium breast cancer findings of Janesick et al., specifically the identification of single-, double-, and triple-positive receptor regions (Figure 2B). This classification is built from univariate local Moran's I values with the neighbourhood defined as the six nearest neighbours of each cell, followed by Moran scatter plot high-high classification and BH-adjusted permutation p-values. No sensitivity analysis for this neighbourhood choice is provided for the Xenium data. This matters because the paper itself demonstrates in the CosMx data (Figure 5D-F, Supplementary Figure S4C-D) that local Moran's I values change when the weight matrix is constructed from contiguous neighbours, 10 nearest neighbours, or a 1000-pixel distance band, and that some cells even have no contiguous neighbours. In the Xenium breast cancer sample, cell density varies across tissue compartments, so a fixed k=6 mixes very different physical scales: six nearest neighbours in a dense tumour region may cover tens of micrometres, whereas in a sparse stromal region they may span a much larger distance. Since the triple-positive map is the union of three univariate high-high classifications, small changes in local Moran's I near the significance or quadrant thresholds could shift the set of cells labelled triple-positive, especially at region boundaries. The central claim of recapitulation is therefore conditional on a weight-matrix choice that is not justified or tested for this dataset. The issue is not fatal to the paper's broader perspective-and-resource message, but it does undercut the strength of one of the two main illustrative findings unless robustness is shown or the limitation is acknowledged.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that classical spatial statistics, partitioned into point pattern analysis and lattice data analysis, provides a powerful and underused toolkit for spatially resolved omics data. The authors illustrate this claim through several public datasets: Xenium breast cancer data (receptor expression and tumour cell co-localization), IMC pancreas islets (scale and homogeneity effects), Visium mouse brain (local Moran's I), and CosMx lung cancer (weight matrix sensitivity). They also introduce pasta, a collection of R and Python vignettes that demonstrates the methods on real data, and discuss technical caveats such as window sampling, homogeneity assumptions, confounding between intensity and interaction, and weight matrix construction.","tokens_in":23616,"tokens_out":5478,"duration_ms":48485,"significance":"If the illustrative analyses are robust, the paper makes a valuable educational contribution by connecting two mature statistical literatures (point processes and lattice data) to the rapidly growing spatial omics field. The analyses are transparent, use standard estimators, and are accompanied by reproducible code, versioned source code, and a public vignette; the paper also candidly discusses scale, homogeneity, and confounding. The central risk is that the main biological demonstration, the triple-positive receptor map in breast cancer, depends on an analyst-chosen neighbourhood size that is not sensitivity-tested, and the recapitulation claim is not quantitatively anchored to the original publication. These issues are fixable and do not undermine the overall educational value of the resource.","major_comments":[{"comment":"The triple-positive receptor map in Figure 2B is built from univariate local Moran's I values with the neighbourhood fixed to the six nearest neighbours (k=6) of each cell, but the paper provides no sensitivity analysis for this choice on the Xenium data. This is load-bearing because the paper itself demonstrates in Figure 5D-F and Supplementary Figure S4C-D that local Moran's I values and even the presence of neighbours change when the weight matrix is constructed differently, and in the 'Definition of the neighbourhood' section it acknowledges that 'it remains to be investigated how much the construction of the weight matrix influences downstream analyses in spatial omics data.' Given that cell density varies across tissue compartments, a fixed k=6 mixes widely different physical scales, and small perturbations of the local Moran's I near the significance or high-high quadrant thresholds could shift the cells labelled triple-positive, especially at region boundaries. Please add a sensitivity analysis (e.g., k=4, 8, 10; distance-based and contiguity-based weights) and report the stability of the triple-positive regions and whether the qualitative recapitulation of Janesick et al. persists.","section":"Biological applications, 'Spatial autocorrelation reveals triple positive regions in breast tumours'"},{"comment":"The statement that spatial statistics 'was able to recapitulate the original findings in Janesick et al. [44]' is not accompanied by any quantitative comparison to the original annotations (e.g., overlap of the inferred triple-positive cell set with the pathologist-annotated regions), and the cross-reference to '(Figure 5)' is evidently incorrect because Figure 5 displays the Visium and CosMx analyses rather than the breast cancer receptor map. Please either provide a quantitative evaluation of the agreement with the original findings or temper the claim to a qualitative demonstration, and correct the figure citation.","section":"Biological applications, 'Spatial autocorrelation reveals triple positive regions in breast tumours'"}],"minor_comments":[{"comment":"The legend of Figure 3 contains the typo 'DICS 2' for 'DCIS 2', and Supplementary Figure S3 contains 'ductual carcinoma in situ' for 'ductal carcinoma in situ'; these should be corrected.","section":"Figure 3 legend and Supplementary Figure S3"},{"comment":"The sentence beginning 'Figure 5C-D show the difference' refers to panels D, E, and F for the contiguity, 10-nearest-neighbour, and 1000-pixel-distance weight matrices; the panel reference should read 'Figure 5D-F' to match the figure.","section":"Definition of the neighbourhood is critical in lattice data analysis"},{"comment":"The paper does not report the exact intensity threshold used to restrict the observation window in Figure 3; please provide this value in the Methods or figure caption, since the comparison in Supplementary Figure S3 is appreciated but the specific threshold remains an analyst choice.","section":"Methods and Figure 3"},{"comment":"In Figure 4, panels B and C have different y-axis ranges, which can make the visual comparison of clustering strength between the homogeneous and inhomogeneous K-functions misleading; consider using a common y-axis range.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"This is more of a perspective/educational resource than a methods paper, so the main criterion for publication should be the accuracy and robustness of the illustrative analyses. The requested sensitivity analysis for the Xenium neighbourhood choice is, in my view, necessary before the central recapitulation claim can be accepted. The authors' use of their own companion packages (sosta, spatialFDA) is disclosed in the Methods, which is good practice; the editorial team may wish to ensure the vignette is framed as an educational resource rather than as a benchmark for these packages."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read on the pasta paper. It's a perspective and resource, not a new method or discovery, and it is honest about that. The genuinely new piece is the vignette—curated, reproducible, with code in R and Python and a DOI. The framing of spatial omics data as either point patterns or lattice data, and the mapping of technologies onto those, is clear and well grounded in the spatial statistics literature. It also does a good job of discussing technical pitfalls—scale, homogeneity, edge effects, confounding—and of pointing to existing tools like Voyager, spatstat, and squidpy rather than pretending to compete with them.\n\nThe worked examples are mostly convincing. The point pattern analysis of DCIS and invasive tumour cells, with the restricted window and comparison of homogeneous vs inhomogeneous L-functions, is nicely done. The IMC islet example illustrates the scale issue well.\n\nThe soft spot is the one the stress-test flags: the triple-positive receptor classification in the breast cancer Xenium data uses a fixed k=6 nearest-neighbour weight matrix, with no sensitivity analysis for that dataset. That choice does matter. The paper itself shows on the CosMx data that contiguous, k=10, and distance-based neighbourhoods give different local Moran's I values, including cells with no contiguous neighbours. In the Xenium sample, cell density varies across tissue, so k=6 means different physical scales in different regions. Since the triple-positive map is a union of univariate high-high classifications, changes near thresholds could shift the region boundaries. So the 'recapitulation of Janesick et al.' is conditional on an untested choice. I would not call this fatal—the paper's main message is about methods and guidance, not about this specific tumour—but the authors should either run a small sensitivity analysis (k=6 vs k=10 vs distance-based) on the Xenium data or add an explicit caveat.\n\nMinor point: the window restriction for the point pattern analysis relies on an intensity threshold from their sosta package; that's also worth a sentence of justification. Self-citation of sosta and spatialFDA is not a problem here—they are their packages, and the resources are available.\n\nAll in all, this is a competent, useful resource paper. It deserves peer review and will be a handy reference for people entering spatial omics analysis. I'd recommend accepting it with a request for the sensitivity check or caveat on the weight matrix choice.","headline":"Solid, honest resource paper; the k=6 neighbourhood choice needs a sensitivity check but doesn't sink the paper.","tokens_in":24206,"tokens_out":2466,"would_cite":true,"duration_ms":21462,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that classical spatial statistics, matched to whether spatial omics data form point patterns or lattices, can quantify biological structures, recapitulating known breast cancer receptor zones and invasion patterns.","keywords":["spatial omics","spatial statistics","point patterns","lattice data","spatial autocorrelation","Moran's I","Besag's L","co-localization"],"falsifier":"Recompute the breast-cancer receptor analysis with $k=4$, $k=8$, a distance-based weight matrix, and a contiguity-based matrix; if the union of significant high-high clusters no longer delineates the same triple-positive DCIS region, then the identification depends on the chosen neighbourhood rather than on the biology.","tokens_in":23145,"feed_emoji":"🔬","tokens_out":8549,"duration_ms":70596,"temperature":0.7,"pith_summary":"Spatial omics data come in two statistical shapes: point patterns, where each cell or transcript is a stochastic point, and lattice data, where measurements sit on fixed spots or segmented cells. The paper argues that the choice of statistical lens should follow this data representation rather than the instrument, and that classical spatial statistics can then quantify biology from local gene expression to tissue organization. Re-analyzing a breast cancer imaging dataset, local Moran's I and the Moran scatter plot identify single-, double- and triple-positive zones for three hormone-receptor genes, and Besag's L function quantifies how differently two ductal carcinoma in situ subtypes associate with invasive tumour cells. Both results recapitulate the original publication's findings and add quantitative, significance-aware detail. The argument is accompanied by pasta, a vignette that demonstrates the analyses in R and Python.","feed_headline":"Spatial statistics pinpoints triple-positive breast cancer zones","feed_subtitle":"Classic point-pattern and lattice tools recapitulate known receptor regions and tumor invasion patterns from spatial omics.","key_machinery":"The machinery is the neighbourhood relation. In lattice analysis it is the weight matrix — a matrix of connection strengths between cells or spots — and the local autocorrelation statistics computed on it, especially local Moran's I with the Moran scatter plot for classifying 'high-high', 'low-low' and heterogeneous zones. In point pattern analysis it is the r-neighbourhood, a disk of radius r around each point, with complete spatial randomness (a uniform, independent scatter of points) as the reference; Besag's L compares observed neighbour counts at each radius to that random baseline. Observation windows, homogeneity assumptions, and edge corrections determine what the functions see, and the paper shows that changing the window or the neighbourhood definition can alter interpretations.","core_discovery":"The central claim is that the two streams of spatial statistics — point pattern analysis, which models the stochastic process generating point locations, and lattice analysis, which treats locations as fixed and models dependence among features through neighbourhood weights — are directly applicable to spatial omics and should be selected by data modality and mark type. The paper demonstrates that imaging-based data can be represented either way: cell centroids as a point pattern, or segmented cells as an irregular lattice of expression measurements. Using a six-nearest-neighbour weight matrix, local Moran's I plus the Moran scatter plot recover triple-positive ERBB2/ESR1/PGR regions in breast cancer, and bivariate Lee's L and multivariate Geary's c add pairwise and multivariate views of spatial correlation. On the point-pattern side, Besag's L in a restricted observation window shows DCIS1 cells are spaced apart from invasive tumour cells while DCIS2 cells mix with them at chance levels, recapitulating the original findings.","pith_inferences":["The sensitivity of local Moran's I to neighbourhood choice shown on the lung cancer example implies that the triple-positive breast-cancer regions should be checked for the same sensitivity; a sweep over $k$ or a distance-based matrix would tell whether those regions are an artifact of $k=6$.","Because imaging data can be represented as both a point pattern and a lattice, the same biological question could be cross-checked across the two streams, for example comparing cell-type co-localization from Besag's L with spatial autocorrelation of categorical cell-type marks.","Cells with no contiguous neighbours produce zero local Moran's I values, which suggests spatial statistics could double as a segmentation-quality diagnostic in imaging-based assays.","The paper's scale-dependent results point to a testable extension: normalization choices upstream, spatially aware versus global, may change which cells are classified as high-high, so combining normalization and neighbourhood choice in one sensitivity analysis would clarify how much of the biological readout is preprocessing-driven."],"forward_implications":["Analysts can pick point-pattern or lattice methods from how the data are represented, not from the brand of technology, so imaging data can be analyzed with both streams and spot-based data can be segmented into lattices or points.","Local spatial autocorrelation with a Moran scatter plot turns diffuse gene-expression maps into discrete regions, such as hormone-receptor-positive tumour zones, with per-cell significance attached.","Besag's cross L gives a scale-resolved, quantitative readout of whether cell types cluster, mix, or repel, replacing visual inspection of co-localization.","The choice of weight matrix, whether contiguous, k nearest neighbours, or distance-based, changes local Moran's I values for some cells, so reporting the choice and its rationale is part of the analysis.","Observation scale controls biological interpretation: islet cells that look clustered in a whole field of view look randomly distributed within an islet, so scale must follow the question."],"supporting_citations":[{"why":"Supplies the breast cancer imaging dataset whose receptor and DCIS findings the paper's spatial analyses recapitulate.","marker":"[44]"},{"why":"Provides the point-pattern theory used throughout: intensity, homogeneity, CSR, edge effects, and K and L functions.","marker":"[11]"},{"why":"Introduces local indicators of spatial association, the basis for the local Moran's I used to find receptor-positive regions.","marker":"[6]"},{"why":"Defines the Moran scatter plot, which separates local autocorrelation into high-high, low-low, high-low, and low-high clusters.","marker":"[7]"},{"why":"Supplies bivariate Lee's L, used to measure pairwise spatial association of the receptor genes.","marker":"[48]"},{"why":"Supplies multivariate Geary's c, used to flag locally heterogeneous multi-gene expression regions.","marker":"[4]"},{"why":"Provide the cross L (Besag's L) summary used to quantify co-localization and spacing of DCIS and invasive tumour cells.","marker":"[14, 73]"},{"why":"Introduces Moran's I, the global spatial autocorrelation statistic from which the local version is derived.","marker":"[57]"},{"why":"Supplies the lung cancer imaging dataset used to show how contiguity-, nearest-neighbour-, and distance-based weight matrices change local Moran's I results.","marker":"[43]"},{"why":"Supplies the spot-based mouse brain dataset used to illustrate local Moran's I and Moran scatter plot interpretation on a regular lattice.","marker":"[1]"}],"fun_headline_variants":["Spatial stats maps breast cancer receptor zones","Point patterns and lattices find breast cancer hotspots","Spatial statistics reveal triple-positive tumor zones","Classic spatial tools meet spatial omics in tumor mapping","Spatial stats decode breast cancer tissue maps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the six-nearest-neighbour neighbourhood defines the right scale for identifying receptor-positive regions; a different $k$ or a distance-based weight matrix could shift the triple-positive classification.","fun_headline_variants_meta":{"raw":{"variants":["Spatial stats maps breast cancer receptor zones","Point patterns and lattices find breast cancer hotspots","Spatial statistics reveal triple-positive tumor zones","Classic spatial tools meet spatial omics in tumor mapping","Spatial stats decode breast cancer tissue maps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000911,"raw_usage":{"total_tokens":3864,"prompt_tokens":841,"completion_tokens":3023,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":2950}},"tokens_in":457,"tokens_out":3023,"duration_ms":18808,"temperature":1.0,"reasoning_tokens":2950,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:17:04.814749+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the breast-cancer receptor analysis with $k=4$, $k=8$, a distance-based weight matrix, and a contiguity-based matrix; if the union of significant high-high clusters no longer delineates the same triple-positive DCIS region, then the identification depends on the chosen neighbourhood rather than on the biology.","supporting_citations":[{"cited_title":"Developing a Bivariate Spatial Association Measure: An Integration of Pearson’s r and Moran’s I","cited_arxiv_id":null,"evidence_quote":"Supplies bivariate Lee's L, used to measure pairwise spatial association of the receptor genes."}],"review_version":1}