{"id":"0d09b62a-185b-4a6c-aa7e-6b90046198f7","arxiv_id":"2607.14163","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Converting transcriptomes into optimal-transport gene-layout images and pretraining a vision transformer on 72 million cells produces frozen embeddings that lead zero-shot cell-type annotation on all six held-out atlases tested and match the best token-based model on integration.","lead":"A new model, scVision, draws each single cell as a grayscale image by placing similar genes side by side, then learns to read such images from 72 million human cells. Tested on tissues it never trained on, its frozen representations label cell types and find gene programs that match or beat today's leading single-cell AI models — often needing far fewer labeled examples.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reference-bank/pretraining overlap is unexamined; if support cells appeared in the 72M-cell pretraining set, the zero-shot 'most accurate' ranking could reflect dataset familiarity rather than spatial representation.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the reference-bank/pretraining overlap is unspecified, and the headline 'most accurate on every atlas' cannot be trusted until this is resolved. I agree that this is the single most important threat to the central claim. The paper is otherwise strong: it uses frozen encoders, matched probes, present-class metrics, honest supplement tables reporting losses, permutation nulls for gene-program recurrence, and explicit limitations. None of those address the possibility that scVision was pretrained on the same studies that later serve as labelled support cells. If the overlap exists, the k-shot results (one label beating fifty labels of other models) and the six-atlas ranking could be partly explained by memorization of the reference cells rather than by the spatial gene layout. The proposed audit is feasible because all data are public through CELLxGENE Census and the paper already stores evaluation records; it would settle the concern without requiring new biological experiments. Since the reader already conditioned acceptance on this and related transparency issues, my stress-test does not move the verdict: it remains CONDITIONAL until the overlap is documented and, if necessary, the benchmarks are re-run with a clean reference bank.","tokens_in":47835,"tokens_out":3011,"duration_ms":38600,"concrete_test":"Audit the study-level and donor-level identifiers of three sets: the pretraining split, the ~25 held-out test studies, and every study contributing to the 87-type reference bank. If any reference-bank study or donor overlaps the pretraining split, re-run the six annotation benchmarks (Figs. 2–7) with the reference bank restricted to studies entirely absent from pretraining and report balanced accuracy and present-class macro-F1 under the same 20-NN probe. If scVision's margins persist under the overlap-free reference bank, the zero-shot ranking is supported; if they shrink or vanish, the headline claim is compromised by dataset leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central zero-shot claim rests on a labelled reference bank of 87 cell types 'from other studies' against which a frozen 20-NN probe is scored. The Methods ('Training data and quality control') state that roughly 25 studies are held out as the external test pool and that the remaining studies are split into training and validation, with no donor crossing the split. But the manuscript never states that the studies contributing to the reference bank were also excluded from the 72M-cell pretraining set. Both the reference bank and the pretraining data are drawn from the CZ CELLxGENE Census, so the same studies — indeed the same labelled support cells — could appear on both sides. If so, scVision's embedding could have been shaped by those exact cells during masked-image pretraining, giving it an advantage over classical baselines that never saw those studies and over token foundation models trained on different corpora. This would not invalidate the layout-ablation result, but it would undermine the stronger claim that scVision is 'the most accurate cell-type annotator' in a zero-shot sense. The supplement reports losses on cystic fibrosis and obstructive nephropathy, which shows honesty but does not resolve the overlap question. Because the reference-bank composition is never specified, the leakage path cannot be ruled out from the paper alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces scVision, a vision foundation model for single-cell transcriptomics. Each cell is rendered as a 104×104 image by placing 10,816 highly variable genes on an optimal-transport layout so that co-expressed genes are spatially adjacent; a ViT-B is pretrained with masked image modeling on 72 million human cells from CZ CELLxGENE Census data. The frozen encoder is then evaluated without fine-tuning on six held-out atlases using a shared 20-nearest-neighbour probe, and the authors claim it is the most accurate cell-type annotator, the most label-efficient, and that the spatial layout—not the encoder—carries most of the signal. The paper also reports attention-based gene-program discovery, cross-tissue program recurrence, multi-study integration comparable to scGPT/scVI on scIB total score, and higher deployment throughput than token-based foundation models.","tokens_in":48120,"tokens_out":5747,"duration_ms":65413,"significance":"If the central claims hold, scVision would be a meaningful conceptual advance: it replaces gene-token sequences with a continuous spatial image, preserves expression magnitude, and connects single-cell representation learning to mature computer-vision methods. The evaluation is above average in care: frozen encoders, a shared 20-NN probe for foundation models, present-class metrics that are appropriate for rare cell types, permutation nulls, robustness stress tests, and an honest supplement that reports losses on cystic fibrosis and obstructive nephropathy. The throughput and memory advantages are also concrete strengths. However, the headline 'most accurate zero-shot annotator' claim depends on an unexamined overlap between the pretraining corpus and the labelled reference bank, and the paper's own supplementary tables contradict a broad reading of 'most accurate on every atlas.' These issues are directly load-bearing and need to be resolved before the principal claim can be accepted.","major_comments":[{"comment":"The zero-shot claim requires that the studies used to build the 87-type reference bank were excluded from the 72M-cell pretraining split. The manuscript states that about 25 studies form the external test pool and that remaining studies are split by donor, but it never defines the reference bank's composition, nor states whether the reference-bank studies—or the six evaluation atlases—are disjoint from the pretraining studies. Since both the reference bank and the pretraining data are drawn from CZ CELLxGENE Census, the same studies, and even the same labelled support cells, could appear on both sides. If so, scVision's embedding could have been shaped by exactly those cells during masked-image pretraining, giving it a transfer advantage over baselines that never saw those studies. This would not invalidate the layout-ablation result, but it would undermine the central 'zero-shot most ac","section":"Methods: 'Training data and quality control'; Data availability"},{"comment":"The main text and Discussion state that scVision is 'the most accurate annotator on every held-out atlas,' and the Abstract says it is 'the most accurate cell-type annotator' relative to existing foundation models and classical baselines. The supplement itself contradicts a broad reading: Table S1, obstructive-nephropathy row under the 20-NN probe, gives scVision 0.482 balanced accuracy versus 0.503 for logistic regression (Δ = −0.021), and Table S4 reports that scFoundation wins cystic fibrosis. The claim may be intended to cover only the six named main-text atlases and only the matched-probe comparison, but as written it is overbroad. Please qualify the central claim explicitly to the six atlases and metrics where it holds, and either remove or prominently contextualize the supplementary losses. This is not a cosmetic issue: readers will otherwise draw a conclusion the paper's own data","section":"Table S1; Table S4; Discussion, final paragraph"},{"comment":"The mechanistic claim that 'the learned spatial arrangement of genes carries the larger share of the signal' is based on ablating the layout only at inference time: the encoder was pretrained on the OT layout, and a permuted layout is therefore an out-of-distribution input. This conflates 'biologically meaningful layout' with 'the layout the model was pretrained on.' A random but fixed layout, used consistently during pretraining under the same masked-image objective, could plausibly yield comparable zero-shot transfer for a sufficiently flexible ViT; the current controls do not rule this out. To support the biological-meaningfulness mechanism, please add a control in which a shuffled or otherwise non-biological but fixed layout is used during pretraining (same data, same objective) and evaluated on the same six atlases, or explicitly soften the claim to 'the learned spatial arrangement'","section":"Fig. 5e; Fig. 7e; Results: 'The advantage is in the spatial representation'"}],"minor_comments":[{"comment":"The 87-type reference bank is never described: no table of cell types, per-type cell counts, label hierarchy, or selection criterion. Since the entire annotation evaluation is scored against this bank, a data card is needed for reproducibility and for understanding apparent failures such as the coarse leukocyte label in Fig. 3g.","section":"Methods: 'Zero-shot evaluation and baselines'"},{"comment":"The paper repeatedly dismisses ROC-AUC as class-dominated. This is reasonable in principle, but no quantitative support is given (e.g., per-class AUC vs. abundance). A short supplement table showing that AUC ordering reverses or flattens under class imbalance would make the metric choice more persuasive.","section":"Results: Fig. 2; Fig. 3; 'ROC-AUC dismissed'"},{"comment":"In Table S1 the Δ column is defined as scVision minus the best of six baselines, but in the obstructive-nephropathy row the value −0.021 is listed while the table summary says 'scVision best (Δ>0): 4/5.' This is internally consistent, but the main text should cite these exceptions when summarizing 'best on every atlas' and should make clear whether 'best' means among foundation models or among all baselines.","section":"Table S1 vs. main text"},{"comment":"The statement 'The code for scVision ... is reserved for now and will be released publicly, with a persistent identifier, upon publication' prevents independent verification of the central claims, including the overlap analysis requested above. At minimum, release the scImage layout, pretrained weights, and evaluation scripts alongside the revised manuscript.","section":"Code availability"},{"comment":"The cross-tissue recurrence null is described as a 'label-permuted null' with 1000 permutations. For such high-dimensional Jaccard statistics, please also report the number of cell-type pairs that have exactly n=1 contributing atlas pair separately (Table S6 already marks n=1 entries as 'listed for completeness only'), and state whether the pooled z=41.3 is computed over all pairs or only cross-tissue pairs.","section":"Methods: 'Gene-program inference and cross-tissue recurrence'"},{"comment":"The phrase 'one of the largest pretrained models for single-cell analysis' is ambiguous: scVision has 85M parameters and was trained on 72M cells, while several token-based models report larger corpora. Please state the comparison explicitly (parameters, cells, studies) or soften the claim.","section":"Abstract and Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper's core idea is interesting and the evaluation is unusually thorough in many places. However, the central zero-shot claim cannot be assessed without knowing whether the reference-bank studies were part of the pretraining corpus. I recommend asking the authors to provide an explicit overlap audit and to rerun the benchmark with any overlapping studies removed. If the audit shows no overlap, the main claim is likely to survive in qualified form; if it shows overlap, the title-level claim will need substantial revision. The discrepancy between the main text and Tables S1/S4 should also be fixed in the same revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. This is a substantial empirical paper: they render each cell as a 104x104 image by placing ~10k genes on an OT-fixed pan-tissue layout, pretrain a ViT by masked image modeling on 72M cells, and evaluate the frozen encoder on held-out studies. The core construct is not entirely new — the same senior author's earlier Nat Commun paper (ref 46) already did image-based single-cell embeddings via genomic cartography — but the specific combination of a Gromov-Wasserstein layout, 72M-cell MAE pretraining, and systematic study-holdout zero-shot evaluation is new and produces results worth scrutinizing.\n\nWhat the paper does well: the evaluation is unusually careful for this literature. Frozen encoders, a shared 20-NN probe for all foundation models, present-class metrics justified against class-dominated ROC-AUC, and steelmanned classical baselines given their own native classifiers. The supplement reports its own losses — scFoundation wins cystic fibrosis, scVision trails logistic regression on obstructive nephropathy — and the robustness stress tests (gene dropout, count downsampling) are reported honestly, with scVision's depth sensitivity acknowledged rather than hidden. The layout ablation is the key control: shuffling the gene layout costs more than removing the encoder, which directly supports the central claim that spatial position carries signal. The cross-tissue attention-program recurrence with permutation nulls is also a strong, falsifiable result.\n\nThe soft spots are real but addressable. Most important: the reference bank of 87 cell types is never described (composition, per-type counts, label granularity), and the paper never states that the studies contributing to that bank were excluded from the 72M-cell pretraining split. Both come from the same CZ CELLxGENE Census, so a leakage path exists: if the same labelled support cells were in the pretraining corpus, scVision could be advantaged over classical baselines and token models with different training corpora. That would not invalidate the layout ablation, but it would weaken the \"most accurate zero-shot annotator\" headline. Second, code, weights, layout, and evaluation records are all withheld until publication, so none of the headline numbers can currently be checked. Third, the abstract and Figure 1's \"most accurate\" framing outruns the supplement's own qualified results, and the advertised perturbation-response and disease-axis capabilities have no reported results in the main text — Fig 1c promises them but they never appear.\n\nThe central argument holds up better than the abstract does. I read no internal contradiction, and the authors' willingness to report losses and null controls suggests clear thinking. But the paper cannot be accepted as-is: the overlap question must be answered with a concrete description of the reference bank and a statement of whether its studies appear in pretraining, and the code/weights need to be released for independent verification.\n\nBottom line: this deserves a serious referee. Send it to peer review with a request for the overlap analysis and code release. If the leakage path is closed, this is a significant contribution; if not, the ranking claims need substantial revision. I'd bring it to reading group either way, since the evaluation design is a useful template and the layout-ablation strategy is worth discussing.","headline":"A serious, unusually honest single-cell foundation model paper whose central representation claim is plausible, but the headline 'most accurate' ranking needs the withheld code and a reference-bank/pretraining overlap check before it can be trusted.","tokens_in":48749,"tokens_out":1056,"would_cite":true,"duration_ms":13445,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A cell rendered as an image, with genes placed by their co-expression, transfers to held-out tissues better than token-sequence representations.","keywords":["single-cell transcriptomics","foundation model","vision transformer","optimal transport","gene co-expression layout","masked image modeling","zero-shot cell-type annotation","gene program discovery"],"falsifier":"Check whether the studies that supply the labeled reference bank are present in the pretraining split, then retrain scVision excluding those studies and re-run the six-atlas balanced-accuracy evaluation. If the accuracy margin over token models collapses to parity, the spatial-layout claim is falsified.","tokens_in":47606,"feed_emoji":"🧬","tokens_out":2474,"duration_ms":31173,"temperature":0.7,"pith_summary":"scVision claims that how a transcriptome is arranged matters as much as the model reading it. Instead of treating genes as an unordered list of tokens, it uses optimal transport to place co-expressed genes at fixed neighboring pixels on a shared 104×104 lattice, painting each cell as a continuous image. A vision transformer pretrained by masked image modeling on 72 million human cells learns a frozen encoder that, with no fine-tuning, annotates cell types in six held-out studies more accurately than existing token-based foundation models and classical baselines. The paper argues the spatial layout itself carries most of the signal: permuting the gene layout drops accuracy more than removing the encoder. If true, this reframes single-cell representation learning as a vision problem and opens a direct route for computer-vision methods in biology.","feed_headline":"Arranging genes spatially makes cell-type annotation far more accurate","feed_subtitle":"A vision transformer trained on 72 million cells-as-images out-annotates token-sequence models in zero-shot transfer.","key_machinery":"The scImage: a 104×104 single-channel image with exactly one gene per pixel, built by Gromov–Wasserstein optimal transport that aligns a gene–gene co-expression distance matrix to a pixel–pixel Euclidean distance matrix. This fixed pan-tissue layout, computed once, makes co-expressed genes spatial neighbors, so a cell's continuous expression values become image intensities and gene programs become local texture. A ViT-base encoder is pretrained by masked image modeling on 72 million cells, and the frozen encoder's mean-pooled patch tokens serve as the zero-shot cell embedding.","core_discovery":"scVision establishes that a transcriptome can be represented as a fixed-layout image in which gene programs appear as local texture, and that this representation, when learned by a masked autoencoder on 72 million human cells, transfers across studies better than gene-token sequence representations. The central claim is that the learned spatial arrangement of genes—not the vision transformer alone—carries the larger share of the signal: randomly permuting the gene-to-pixel layout drops balanced accuracy from 0.52 to 0.29, while removing the encoder drops it only to 0.46. On six held-out atlases, the frozen scVision embedding yields the highest balanced accuracy and present-class macro-F1 amo","pith_inferences":["If the spatial layout is the main carrier of signal, then improving the layout itself—for example, learning tissue-specific maps or refining the optimal-transport cost—could yield further accuracy gains without changing the encoder.","The label-efficiency result suggests that representation quality, not classifier choice, is the bottleneck in zero-shot cell annotation, which could simplify future annotation pipelines in practice.","The image formulation naturally extends to spatial transcriptomics and paired modalities such as surface protein or chromatin accessibility, where the same layout logic might apply.","A direct test of the framework's generality would be to pretrain on multiple species and evaluate whether the learned layout transfers across species boundaries, since the current model is human-only."],"forward_implications":["Frozen scVision embeddings annotate held-out cell types more accurately than token-sequence foundation models on every tested atlas, with one labeled cell per type often beating other models given fifty labels.","Attention read from the last transformer block recovers cell-type-specific gene programs without pathway supervision, including a shared four-gene myeloid program found in microglia, macrophages, and dendritic cells across organs.","On multi-study integration, scVision matches the strongest token-based model on the combined score while preserving more biological structure, and does so without ever seeing a batch label.","The spatial formulation makes the representation robust to gene dropout—it remains above its clean-atlas score even with 70% of genes masked—while degrading gracefully with sequencing depth.","Spatial masking of a gene neighborhood provides a way to perturb a co-regulated module in silico, an operation with no direct counterpart in token-based models."],"fun_headline_variants":["Spatial gene layout, not the network, powers scVision's accuracy","Transcriptomes as images: spatial genes beat token sequences for cell typing","scVision: gene arrangement, not the model, drives cell typing accuracy","Cell-as-image model out-annotates token models in zero-shot","Gene positions, not the transformer, are the real signal in scVision"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The zero-shot evaluation assumes the reference bank used for scoring is truly independent of the 72-million-cell pretraining set; if the reference studies overlap with pretraining, the claimed transfer advantage could partly reflect memorization rather than generalization.","fun_headline_variants_meta":{"raw":{"variants":["Spatial gene layout, not the network, powers scVision's accuracy","Transcriptomes as images: spatial genes beat token sequences for cell typing","scVision: gene arrangement, not the model, drives cell typing accuracy","Cell-as-image model out-annotates token models in zero-shot","Gene positions, not the transformer, are the real signal in scVision"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001937,"raw_usage":{"total_tokens":7423,"prompt_tokens":758,"completion_tokens":6665,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":6570}},"tokens_in":502,"tokens_out":6665,"duration_ms":69965,"temperature":1.0,"reasoning_tokens":6570,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T05:22:48.578257+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether the studies that supply the labeled reference bank are present in the pretraining split, then retrain scVision excluding those studies and re-run the six-atlas balanced-accuracy evaluation. If the accuracy margin over token models collapses to parity, the spatial-layout claim is falsified.","supporting_citations":[],"review_version":1}