{"id":"43a0f986-c845-434f-9756-e476962a0150","arxiv_id":"2412.13591","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of single-cell spatial omics technologies and data analysis methods, organized around data representation, visualization, and machine learning.","lead":"This survey reviews technologies and data analysis methods for single-cell spatial omics, which measure molecular features at cell resolution in space. It organizes the field around data structures, dimensionality reduction, and machine learning approaches, making it a useful entry point for researchers entering this area.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed breadth across all major omics modalities is not delivered: genomics and epigenomics receive only cursory mentions, undermining the central 'comprehensive overview' claim.","rationale":"The reader's weakest assumption concerns the common-coordinate-system requirement for the tensor representation Xs in Section 4.2. This is a legitimate caveat, but the authors themselves explicitly acknowledge it as 'a strong requirement' and cite alignment algorithms as a cornerstone for making it useful. More importantly, the survey's central value does not depend on Xs being universally applicable: Section 4.2 and Section 6 present graph-based and omics-per-cell representations as alternatives. The reader's CONDITIONAL verdict is therefore safe on that ground. The more load-bearing concern is the gap between the paper's stated scope and its actual coverage. The abstract and introduction promise coverage of all major omics modalities, yet genomics and epigenomics receive only brief mentions in the technology sections and no substantive treatment in the computational analysis sections. A reader picking up this review to learn about scs genomics or epigenomics data analysis would be misled. This is not a question of consensus or taste; it is an internal inconsistency between the claims and the content. The minor citation issues noted by the reader (FISSEQ efficiency, EEL-FISH citation) are real but secondary; they do not affect the central argument as much as the missing modality coverage. The CONDITIONAL verdict remains appropriate, but the condition should be expanded to include either adding material on scs genomics/epigenomics analysis or revising the scope statement to reflect the true focus on transcriptomics and MSI-based metabolomics/proteomics.","tokens_in":1045,"tokens_out":790,"duration_ms":57677,"concrete_test":"Count the number of paragraphs and cited computational tools devoted to scs genomics and epigenomics in Sections 3–6. If the combined total is below a small threshold (e.g., fewer than 5 paragraphs and no dedicated data-analysis methods for these modalities), the 'cover all major omics modalities' claim is unsupported. Additionally, search the reference list for any papers that specifically address computational analysis of scs genomics or epigenomics beyond technology descriptions; if none exist, the review's scope statement should be revised or those sections added.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and introduction state the goal is to 'cover all major omics modalities, including transcriptomics, genomics, epigenomics, proteomics and metabolomics' (Abstract, Section 1). The actual content does not deliver this. Section 3.1.1 gives two sentences to imaging-based genome/epigenome profiling; Section 3.1.2 mentions microfluidic barcoding for ATAC&RNA-seq and spatial CUT&Tag-RNA, plus DNA-GPS, but provides no detailed treatment. Sections 4–6, the computational core of the review, are devoted almost exclusively to transcriptomics and MSI-based metabolomics/proteomics: the pipeline discussion in Section 4.1 contrasts in situ transcriptomics with MSI, the data-structure discussion in Section 4.2 and the dimensionality-reduction survey in Section 5 use transcriptomics/MSI examples, and the machine-learning survey in Section 6 centers on SVGs, deconvolution, cell–cell communication, and MSI texture analysis. There is no equivalent discussion of computational challenges specific to scs genomics (e.g., spatial CNAs, somatic mutations) or epigenomics (e.g., spatial peak calling, chromatin domains). Thus the review falls short of its stated scope; a reader seeking analytical methods for scs genomics/epigenomics would not find them. This is an internal mismatch between the claimed comprehensiveness and the contents, not a disagreement with scientific consensus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a review of single-cell spatial (scs) omics from a data-analysis perspective. It introduces a glossary distinguishing spatial resolution from spatial localization, surveys acquisition technologies for transcriptomics/genomics and for proteomics/metabolomics, describes computational pipelines and a tensor data representation for spatially resolved data, reviews dimensionality reduction methods under the spatially agnostic/informed dichotomy, and surveys machine-learning approaches for modeling spatial information. It closes with future challenges centered on spatio-temporal modeling and experimental design integration.","tokens_in":27765,"tokens_out":8557,"duration_ms":74606,"significance":"If delivered at its stated breadth, the review would be a useful cross-modal entry point for researchers moving from single-cell transcriptomics into spatial methods, and it would help disseminate ideas from multivariate image analysis and chemometrics into the spatial-omics community. The paper is strongest where it synthesizes concepts: the resolution/localization distinction, the explicit treatment of spatially agnostic versus spatially informed dimensionality reduction, the discussion of tensor unfolding for omics-per-pixel data, the connection to MIA/MCR/spectral-analysis literature, and the acknowledgment that the tensor representation requires a common coordinate system across samples. It also provides a broad and current reference list. The main weakness is that the stated scope, covering all major omics modalities, is not matched by the actual content, particularly for genomics and epigenomics, so the review's central claim of comprehensive coverage needs either substantial expansion or a more modest framing.","major_comments":[{"comment":"The survey's stated goal of covering all major modalities, including genomics and epigenomics, is not met by the actual content. Section 3.1.1 gives only two sentences to imaging-based genome and epigenome profiling, and Section 3.1.2 mentions microfluidic barcoding for ATAC&RNA-seq, spatial CUT&Tag-RNA, and DNA-GPS, but Sections 4–6, the computational core that distinguishes this review, discuss almost exclusively transcriptomics and MSI-based metabolomics/proteomics. There is no treatment of computational challenges specific to scs genomics (e.g., spatial copy-number alterations, somatic mutation inference) or scs epigenomics (e.g., spatial peak calling, chromatin domains). The comprehensive-overview claim in the abstract is therefore not supported by the body of the paper. Please either add dedicated computational material for these modalities or revise the stated scope to match the content.","section":"Abstract; Section 1; Sections 3.1.1–3.1.2; Sections 4–6"},{"comment":"The tensor representation Xs (I × J × S1 × S2) is presented as the general data structure for spatial omics, but it depends on the assumption that samples from different individuals can be transformed to a common coordinate system. The authors acknowledge this as 'a strong requirement' and cite alignment algorithms [56]–[59], but the subsequent discussion in Sections 5 and 6 largely treats Xs as the default setting without explaining how often this assumption holds or what an analyst should do when alignment fails. The alternative neighborhood-network or coarse-grid representations are mentioned only briefly at the end of Section 4.2; the relationship between these alternatives and the tensor formalism, and the consequences for each downstream step, deserve more development given the review's stated focus on challenges in downstream analysis.","section":"Section 4.2"}],"minor_comments":[{"comment":"Reference [15], cited for sequencing preprocessing, is a deep-learning attention-mechanism paper and does not support the stated claim; please replace it with a relevant preprocessing reference or remove it.","section":"Section 1"},{"comment":"The sentence 'The detection efficiency of FISSEQ is lower than 0.005%' is an uncited quantitative claim; it should cite the primary FISSEQ work or be removed.","section":"Section 3.1.2"},{"comment":"The claim about a deep masked auto-encoder is attributed to 'Prasad et al. [113]', but reference [113] is BayesSpace by Zhao et al.; this citation appears to be mismatched and should be corrected.","section":"Section 6.2.1"},{"comment":"The technique name 'DBIiT-seq' is a typo and should read 'DBiT-seq'.","section":"Section 3.1.2"},{"comment":"In the sentence listing seqFISH+, MERFISH, and EEL-FISH with citations [29], [30], reference [30] is the DNA-GPS theoretical framework; please verify whether [30] supports the EEL-FISH sentence or replace it with [32].","section":"Section 3.1.1"},{"comment":"The bibliometric counts in Figure 1 are described only as extracted from the Web of Knowledge; the exact query strings, search dates, and inclusion criteria should be reported to make the figure reproducible.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a journal that publishes surveys of computational methods. I would ask the authors to reconcile the title and abstract with the actual coverage. If they prefer to narrow the scope to transcriptomics plus MSI-based proteomics/metabolomics, the abstract and introduction must be revised; if they wish to retain the claim of comprehensive coverage, they need to add a substantive computational section for scs genomics and epigenomics rather than a few sentences in Section 3."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Useful review, not original research. What you get is a clear organizational map of single-cell spatial omics data analysis: a glossary that actually helps, a sensible split between spatially agnostic and spatially informed methods, and a practical account of how spatial data can be arranged as tensors. The ML section is the strongest part; the taxonomy of encoding spatial information in the observation matrix versus in the learning algorithm is clean, and the bridge to multivariate image analysis and MCR is a real plus for readers coming from bioinformatics. The authors also flag the main limitation of their tensor representation, common coordinates across samples, so the central assumption is stated rather than hidden. I agree with the reader that there is no circularity problem; self-citations are used as supporting technical references and are fine.\n\nThe biggest soft spot is scope. The abstract promises coverage of all major omics modalities, but genomics and epigenomics are treated only in passing. Imaging-based genome and epigenome profiling gets a couple of sentences, and the computational core, including pipelines, tensor representation, dimensionality reduction, and ML, is almost entirely transcriptomics plus MSI proteomics and metabolomics. A reader looking for spatial CNA calling or spatial peak-calling methods would not find them. This is not a disagreement with the field; it is an internal mismatch between stated and delivered scope, and it should be fixed in revision.\n\nThere are also two accuracy issues worth correcting: the FISSEQ detection efficiency number appears without a citation, and the EEL-FISH citation looks mismatched. Both are minor and fixable.\n\nOverall, this is a competent survey. It will be most useful to newcomers who need a map of the field and to researchers writing an introduction to spatial omics analysis. It does not resolve an open question or contribute new methods, but it does not pretend to. I would send it to peer review; a firm referee can require the scope correction and the citation fixes. I would also cite the ML taxonomy in future work.","headline":"Competent scs omics survey with a solid ML taxonomy, but it overpromises coverage of genomics and epigenomics and needs citation fixes.","tokens_in":28179,"tokens_out":3380,"would_cite":true,"duration_ms":31340,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims that diverse single-cell spatial omics data, across transcriptomics, genomics, epigenomics, proteomics, and metabolomics, can be organized by one tensor structure, and that modeling spatial dependence, heterogeneity…","keywords":["single-cell spatial omics","spatial transcriptomics","spatial metabolomics","mass spectrometry imaging","tensor data representation","dimensionality reduction","spatially informed machine learning","spatial dependence"],"falsifier":"Take scs datasets of the same tissue produced by different laboratories or platforms and attempt to align their spatial domains to a common coordinate system; if cell-type distributions and spatial domains cannot be matched beyond chance, the tensor representation's core assumption fails for that tissue and the paper's framework cannot support cross-sample spatio-temporal inference.","tokens_in":27219,"feed_emoji":"🧬","tokens_out":7144,"duration_ms":62076,"temperature":0.7,"pith_summary":"Single-cell spatial (scs) omics data arrive from very different technologies, from imaging-based transcriptomics to mass spectrometry imaging for proteins and metabolites, but the paper argues they share a common mathematical skeleton: a tensor with one mode for individuals, one for omics features, and two or three spatial coordinates. By placing this tensor at the center, the review clarifies what each computational step does: preprocessing turns raw measurements into omics-per-pixel tensors or omics-per-cell matrices, dimensionality reduction methods either ignore or exploit the spatial indices, and machine-learning models encode spatial information either in the input matrix or in the learning algorithm. The authors' thesis is that scs data should be treated as spatial stochastic processes with dependence, heterogeneity, and scale, and that explicitly modeling these properties will determine whether scs omics delivers on its promise for studying disease across time and space. The paper also flags the load-bearing requirement that samples from different individuals be aligned to a common coordinate system, which it calls a strong assumption.","feed_headline":"One tensor organizes every single-cell spatial omics modality","feed_subtitle":"The review's tensor view makes clear which analysis methods use space and which ignore it.","key_machinery":"The central object is the tensor $\\mathbf{\\underline{X}}_s$ with dimensions $I \\times J \\times S_1 \\times S_2$ (and its three-dimensional extension), which represents omics-per-pixel or omics-per-grid-cell data for $I$ individuals, $J$ features, and $S_1 \\times S_2$ spatial locations. This tensor does the work of making spatial information explicit before cell segmentation; after segmentation, spatial information is carried instead by a coarse grid or a neighborhood graph. Around this object the paper organizes the whole survey: the tensor explains why spatially agnostic methods can be trustworthy for spatial patterns (they were not looking for them), why spatially informed projections need extra validation, and why alignment of samples to a common coordinate system is the prerequisite for spatio-temporal analysis.","core_discovery":"The paper's central contribution is a unified view of scs omics as a data-analysis problem. It defines spatial resolution and spatial localization as separate properties, surveys acquisition technologies across the five major omics modalities, and proposes the four-way tensor $\\mathbf{\\underline{X}}_s$ with dimensions $I \\times J \\times S_1 \\times S_2$ for spatially resolved measurements on a two-dimensional grid. Using this structure, it classifies dimensionality-reduction methods as spatially agnostic (PCA, t-SNE, UMAP, and relatives) or spatially informed (variants of PCA and matrix factorization that build spatial correlation into the model), and it organizes machine-learning approaches by whether spatial properties enter through the observation matrix or through the learning algorithm itself. The authors' main thesis is that the spatial properties of dependence, heterogeneity, and scale are information, not noise, and that future progress depends on modeling them explicitly while borrowing ideas from multivariate image analysis, multivariate curve resolution, and spectral analysis.","pith_inferences":["One testable extension: benchmark spatially agnostic against spatially informed projections on the same tissue, measuring how often spatial domains reproduce in held-out samples; this would quantify the paper's warning that informed methods can find spatial patterns even where none exist.","If cross-sample alignment becomes routine, the same tensor framework could enable cohort-level spatio-temporal atlases, comparing disease progression across individuals at matched spatial coordinates, a consequence the paper points toward but does not develop.","The paper's taxonomy suggests that modern attention-based and generative models, mentioned only briefly, could be grafted onto the tensor structure for super-resolution and missing-data imputation in scs data.","A practical extension of the review's logic is that method developers should report whether their pipeline is spatially agnostic or informed, since the trust one can place in discovered spatial patterns depends on that choice."],"forward_implications":["If the tensor representation is adopted, a single data organization can serve transcriptomics, genomics, epigenomics, proteomics, and metabolomics, easing cross-modal method transfer.","Because spatially agnostic visualizations can be trusted for spatial patterns while spatially informed ones require extra validation, the paper's framework implies that scs studies should report both kinds of projection to distinguish genuine tissue organization from method-induced smoothness.","Extending experimental-design-aware tools such as ASCA and PERMANOVA to spatial, massive, sparse data would let scs analyses incorporate randomization, replication, and blocking, increasing statistical power.","Alignment algorithms are a cornerstone for spatio-temporal studies, and their improvement determines whether the tensor structure can be used across individuals and time points.","Ideas from multivariate image analysis, multivariate curve resolution, and spectral analysis are ready-made sources of spatially informed methods for scs omics."],"supporting_citations":[{"why":"Supplies a widely used integration method and anchors for combining single-cell datasets, which the paper treats as a baseline for multimodal work.","marker":"[3]"},{"why":"Supplies the computational pipeline steps for spatial transcriptomics, including the in situ versus NGS distinction the paper generalizes.","marker":"[9]"},{"why":"Defines principles and challenges for modeling temporal and spatial omics data, providing the frontier the paper extends.","marker":"[22]"},{"why":"The closest prior survey of single-cell and spatial multi-omics, against which the paper positions its data-analysis-oriented focus.","marker":"[28]"},{"why":"Supplies a toolbox for spatial expression data and the neighborhood-graph representation used for omics-per-cell analysis.","marker":"[54]"},{"why":"The dominant visualization method that the paper analyzes for its spatially agnostic character and computational efficiency.","marker":"[61]"},{"why":"Provides the three structural properties of spatial data, dependence, heterogeneity, and scale, that organize the machine-learning section.","marker":"[117]"}],"fun_headline_variants":["Tensor framework unifies single-cell spatial omics analysis","One tensor organizes methods across five spatial omics","Spatial info is data, not noise: tensor view of scs omics","Five omics, one tensor: a map for spatial data analysis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that spatial samples from different individuals can be transformed to a common coordinate system so that the tensor $\\mathbf{\\underline{X}}_s$ is meaningful across a study; the paper itself calls this a strong requirement, and if alignment fails, cross-sample spatio-temporal analysis of scs data loses its foundation.","fun_headline_variants_meta":{"raw":{"variants":["Tensor framework unifies single-cell spatial omics analysis","One tensor organizes methods across five spatial omics","Spatial info is data, not noise: tensor view of scs omics","Five omics, one tensor: a map for spatial data analysis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000356,"raw_usage":{"total_tokens":1884,"prompt_tokens":849,"completion_tokens":1035,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":963}},"tokens_in":465,"tokens_out":1035,"duration_ms":7656,"temperature":1.0,"reasoning_tokens":963,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:57:28.923426+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take scs datasets of the same tissue produced by different laboratories or platforms and attempt to align their spatial domains to a common coordinate system; if cell-type distributions and spatial domains cannot be matched beyond chance, the tensor representation's core assumption fails for that tissue and the paper's framework cannot support cross-sample spatio-temporal inference.","supporting_citations":[{"cited_title":"Vandereyken, A","cited_arxiv_id":null,"evidence_quote":"The closest prior survey of single-cell and spatial multi-omics, against which the paper positions its data-analysis-oriented focus."},{"cited_title":"Dries et al., ‘Giotto: a toolbox for integrative analysis and visualization of spatial expression data’, Genome Biol., vol","cited_arxiv_id":null,"evidence_quote":"Supplies a toolbox for spatial expression data and the neighborhood-graph representation used for omics-per-cell analysis."},{"cited_title":"Nikparvar and J.-C","cited_arxiv_id":null,"evidence_quote":"Provides the three structural properties of spatial data, dependence, heterogeneity, and scale, that organize the machine-learning section."}],"review_version":1}