Pith. sign in

REVIEW 3 major objections 5 minor 86 references

Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims spatial gene-expression profiles give pathology image encoders a task-agnostic molecular awareness, improving gene-expression prediction, tissue classification, and whole-slide mutation-state prediction once the…

desk verdict A genuinely new large-scale recipe for aligning pathology images with spatial transcriptomics, but the paper's strongest evidence for its headline claim is compromised by same-patient leakage in the slice-level leave-one-out evaluation. read the letter →

arxiv 2412.00651 v1 pith:X2S34HW3 submitted 2024-12-01 cs.CV q-bio.GN

classification cs.CVq-bio.GN
keywords spatialtranscriptomicscomputationalpathologymultimodalrepresentationlearningcontrastivegeneexpressionpredictionwholeslideimagesViSTomics-4Mfoundationmodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that gene-expression profiles from spatial transcriptomics can supply the training signal that pathology image representations currently get from image-text pairs. The proposed framework, UMPIRE, first pre-trains a BERT-style gene encoder on roughly four million spatial transcriptomics spots, then aligns it with pre-trained pathology vision encoders by symmetric contrastive learning over 697,000 paired image-expression spots, so that the image embeddings acquire a specifically molecular awareness. On the paper's evaluations, this molecular awareness improves three families of downstream tasks, namely gene-expression prediction from H&E images, spot-level tissue classification, and WSI-level mutation-state prediction, and the gains transfer to a sequencing platform not seen in pre-training. If the claim holds, a single image encoder could answer molecular questions about cancer tissue at image-only cost, and the alignment recipe would be reusable for adding other molecular modalities to pathology foundation models.

What carries the argument

The load-bearing mechanism is a two-stage alignment whose first stage is Visiumformer, a 12-layer Transformer pre-trained with BERT-style masked-token prediction on tokenized spatial transcriptomics: each spot's gene-expression vector is normalised against per-gene means, sorted by expression level, and truncated to the top 1,500 gene indices, making the input order-agnostic. The second stage aligns a pre-trained pathology vision encoder (Phikon, ViT-B/16, or UNI, ViT-L/16) with the gene encoder using a symmetric contrastive loss in a shared 512-dimensional space over 696,636 pathology-image-expression pairs, pulling paired image and gene embeddings together and pushing unpaired ones apart. For gene-expression prediction at inference, the image is embedded as a query vector, compared by cosine similarity against a reference database of gene embeddings, and the top-K neighbours' expression profiles are weighted-aggregated into the prediction. The contrastive objective, rather than regression or reconstruction, is what the paper identifies as preserving the vision encoder's visual semantics while adding molecular structure; ablations replacing it with MSE or L1 loss degrade classification F1 by roughly 16 points.

What would settle it

Re-run gene-expression prediction with patient-level held-out splits, removing every slice of the test patient from both the fine-tuning set and the reference database; if the Pearson-correlation advantage over BLEEP collapses toward zero, the reported molecular awareness would be shown to be transductive memory rather than a general representation. A complementary check is zero-shot transfer to a tissue type absent from both pre-training corpora, where retention of the gains would confirm that the signal is task-agnostic.

Watch

Extended reading notes

Core claim

The central claim is that the molecular perspective is a robust, task-agnostic training signal for pathology image embeddings, in the sense that aligning image embeddings with gene-expression embeddings produces representations that are better at molecular-related tasks while preserving, and in some cases improving, purely visual performance. Concretely, the paper reports that its aligned encoders outperform the leading contrastive baseline BLEEP by an average of +42.9% in Pearson correlation for gene-expression prediction when fully fine-tuned, improve balanced accuracy in DLPFC cortical-layer classification by up to +42.3% relative to the base vision encoder, and improve WSI mutation-state AUC in three of four genes, while also outperforming the bulk-RNA method TANGLE in three of four mutation sub-tasks. The authors attribute the gains to the two-stage design: large-scale unimodal pre-training of the gene encoder, followed by contrastive alignment that teaches the image encoder molecular structure without destroying its visual semantics, which they support with ablations showing that replacing the symmetric contrastive loss with regression losses costs about 16 points of weighted F1.

Load-bearing premise

The central claim stands on the assumption that the leave-one-out slice evaluation measures learned molecular understanding rather than patient memory, since HLT's four sections come from one person, HPC's five from two, and the prediction reference database is built from remaining slices that can belong to the test patient.

Editorial extensions

If this is right

  • If the molecular signal is task-agnostic, a single aligned vision encoder can serve gene-expression prediction, spot-level tissue classification, and WSI-level mutation-state prediction, replacing task-specific models that must be retrained per task.
  • The reported +83.8% average PCC gain over BLEEP on the HER2+ dataset, sequenced with an older platform never seen in pre-training, implies the molecular alignment transfers across sequencing technologies, not just across tissue types.
  • Because the contrastive alignment preserves the vision encoder's original semantics, as evidenced by 10X Breast classification not degrading, the recipe could extend to other molecular modalities such as protein expression or methylation without sacrificing visual performance.
  • The ViSTomics-4M pre-training corpus, at 3.94 million spots across 1,363 slides and 30 tissues, is claimed to be the largest Visium-based dataset assembled to date, and the code and pre-trained weights are released, so the molecular encoder is reusable for downstream spatial-omics tasks.
  • Adapters using 0.3-0.8% of the fine-tuning parameters come within 2.8% of full fine-tuning, suggesting the molecular alignment can be attached to large vision encoders at low computational cost in resource-limited settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The leave-one-out slice protocol leaves a patient-level re-split as the natural next check; if the gains persist with all same-patient slices removed from the reference set, the task-agnostic claim would be considerably strengthened.
  • The tokenization scheme, sorting a spot's 20,310 genes by normalised expression and keeping the top 1,500, discards the spatial relationships among spots, so adding positional or neighbourhood context to the gene encoder is an implicit extension that could raise the ceiling on expression prediction.
  • The same two-stage recipe could in principle be applied to other tile-based histology readouts that are expensive to obtain, such as microsatellite-instability scoring or tumour purity, making molecular-aware image embeddings a cheap surrogate for assays that currently require sequencing.
  • Because the aligned embeddings place images and expression in one space, retrieval-based analysis becomes possible at inference without sequencing: a new H&E slide could be matched to reference expression programs, yielding an interpretable, patient-specific molecular sketch, a use the paper's query-reference design already anticipates but does not develop.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. UMPIRE proposes a two-stage framework for aligning pathology image patches with spatial transcriptomics gene expression. In the first stage, the authors pre-train a BERT-like gene encoder (Visiumformer) on roughly 3.94 million Visium spots from the new ViSTomics-4M collection; in the second stage, they use symmetric contrastive learning (Eq. 6) to align Phikon or UNI image encoders with Visiumformer over 697K paired pathology-image and gene-expression samples from the HEST dataset. The resulting representations are evaluated on three molecular-related tasks: query-reference gene expression prediction (Section 4.3), linear-probing spot/patch classification (Section 4.4), and MIL-based WSI mutation state prediction (Section 4.5). The paper reports consistent improvements over several baselines and releases code and pretrained weights.

Significance. The paper addresses a timely and important goal: injecting molecular information into pathology image representations. The scale of pre-training data, the systematic comparison of multiple vision encoders (Phikon and UNI), the use of two fine-tuning strategies (adapter and full fine-tuning), and the public release of code and weights are all strengths. If the reported results survive a leakage-free evaluation, UMPIRE would be a solid contribution to computational pathology and multimodal representation learning. However, the current evaluation does not yet establish the paper's central claim of a 'robust, task-agnostic training signal', because the main evidence is affected by a patient/donor leakage issue.

major comments (3)
  1. [Section 4.3, Eq. (8)-(9), Appendix B.2, D.3] The gene expression prediction evaluation uses leave-one-out cross-validation at the slice level. As the paper states in Appendix D.3, HLT has four sections from one individual, HPC has five sections from two patients, and HER2+ has 32 slides from seven patients. For each held-out slice, the reference database in Eq. (8)-(9) is built from the remaining slices, which therefore include adjacent or same-patient sections. Because prediction is a top-K retrieval of real expression profiles, the model can achieve high PCC by recognizing patient-specific expression programs or batch artifacts rather than by learning a general image-to-expression mapping. Section 4.3 says the downstream datasets were excluded from pre-training to eliminate data leakage, but that statement does not address patient-level leakage within the leave-one-out folds. I request a patient-level (or donor-level) split, or at minimum an analysis showing that the reported PCC gains over BLEEP are not driven by same-patient references (e.g., compare retrieval from same-patient versus cross-patient reference sets). Without this, the +39.0% and +42.9% average improvements over BLEEP in Table 1 may reflect transductive memory.
  2. [Section 4.4, Table 2, Appendix B.2] The DLPFC linear-probing evaluation in Table 2 also uses slice-level leave-one-out cross-validation, but DLPFC consists of 12 sections from only three healthy donors. The reported improvements in balanced accuracy (+28% to +42% depending on the encoder) could therefore be donor-specific: the classifier may learn donor identity from the 11 training slices and apply it to the held-out slice of the same donor. Since this task is a major piece of evidence for the claim that the molecular perspective improves image embeddings, the authors should report donor-level cross-validation or per-donor results, and should discuss whether the gains persist when the training set contains no sections from the test donor.
  3. [Section 4.5, Figure 5] The WSI mutation task uses patient-level five-fold cross-validation, which is the appropriate protocol, but the results are modest and inconsistent: UMPIRE outperforms the original vision encoder in only three of four subtasks, and it lags TANGLE on EGFR by about 3.65% in AUC. Consequently, this experiment alone cannot carry the abstract's claim that the molecular perspective provides a robust, task-agnostic training signal. The paper would be strengthened by a clearer statement of which evidence is meant to support task-agnosticism once the leakage-prone experiments are re-evaluated.
minor comments (5)
  1. [Title and Abstract] The word 'REpresentationn' is misspelled in the title and abstract; the correct spelling is 'Representation'.
  2. [Figure 2] The label 'Get Embedding durning Infernece' contains typographical errors; it should read 'Get Embedding during Inference'.
  3. [Section 3.2] In the description of the pre-training setup, 'per-training' appears where 'pre-training' is intended.
  4. [Section 4.3] The paper's use of 'task-agnostic' could be more precise, since all downstream tasks are molecular-related (gene expression prediction, spot classification of molecularly defined layers, and mutation prediction). Clarifying that the claim concerns transfer across molecular-related tasks would be helpful.
  5. [Section F, Limitations] The Limitations section does not mention the potential for patient/donor leakage in the slice-level leave-one-out evaluations; this should be acknowledged and addressed in the revision.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: UMPIRE's contrastive alignment is evaluated on external benchmarks; the slice-level leave-one-out leakage is a benchmark-validity caveat, not a reduction of the prediction to its fitted inputs.

full rationale

The paper's derivation chain is self-contained. Visiumformer is pre-trained with a masked-language-modeling loss on 3.94 million unpaired spatial transcriptomics profiles (Eq. 5), the vision encoder is an external pathology foundation model (Phikon or UNI), and the alignment stage optimizes a symmetric contrastive loss (Eq. 6) on HEST pairs from which the downstream evaluation datasets are explicitly excluded: 'the datasets used in this section were not included in the pre-training phase.' Downstream gene-expression prediction does not fit the held-out query slice's expression values: Eqs. 8-9 retrieve and weight reference profiles by cosine similarity, and the reported Pearson correlation is computed against a query slice that is not in the training fold. No parameter is fitted directly to the target values being predicted. The main residual concern is benchmark design rather than circularity: leave-one-out is performed at the slice level for HLT (four sections from one individual), HPC (five sections from two patients), and DLPFC (twelve sections from three donors), so the reference database and fine-tuning folds can contain adjacent sections of the same patient; this could inflate PCC through transductive memory of patient-specific expression programs. That is a leakage and external-validity limitation, not an equation-level reduction of the claimed molecular signal to its own inputs. The only self-citation (reference [54], a prior work by co-author Linhao Qu) is a non-load-bearing literature citation in the introduction. No load-bearing claim reduces by construction to a fitted parameter, a self-citation chain, or a renamed known result.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new physical or biological entities. The free parameters listed are hand-set hyperparameters that affect the central results; the axioms are the domain assumptions about image-gene correspondence and the soundness of the evaluation split.

free parameters (4)
  • context length N (gene tokens) = 1500
    Chosen in Section 3.2; the gene encoder sequence length is a hand-set hyperparameter that defines how much expression information is retained.
  • masking probability = 0.15
    Standard BERT-style masking, set in Table S3; not tuned.
  • temperature tau in contrastive loss = not reported
    Equation (6) includes temperature tau, but its value is never reported in the paper, leaving an unresolved hyperparameter that affects alignment.
  • top-K references for gene expression prediction = not specified
    Equation (8) uses the top K references, but K is never stated; the reported PCC depends on this choice.
assumptions (5)
  • domain assumption Gene expression anomalies correspond to discernible morphological patterns in pathology images
    Stated in Introduction (Section 1) as motivation; the entire alignment presupposes that paired image-gene similarities are learnable.
  • domain assumption Visium 55-micron spots align with 224x224 patches extracted at 50-100 um
    Section 3.1 and D.2 assume the patch covers the spot, so paired image-gene data are spatially matched.
  • domain assumption HEST filtering and exclusion of downstream datasets prevents data leakage
    Section D.2 claims certain HEST data used downstream are excluded, but the full list of downstream datasets and their provenance is not auditable from the paper.
  • ad hoc to paper Masked language modeling on sorted gene tokens is a valid self-supervised objective for gene expression
    The tokenization (Equation 1) discards absolute expression magnitudes and uses ranks; this is an ad hoc design choice borrowed from BERT but not independently validated for this domain.
  • standard math Transformer self-attention and symmetric contrastive loss are used as in prior work
    Sections 3.2 and 3.3 rely on standard architectures and loss functions from BERT and CLIP.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics." pith.science (2026). https://pith.science/paper/X2S34HW3

@misc{pith2026241200651,
  author       = {Pith},
  title        = {Pith review of: Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/X2S34HW3}},
  note         = {Machine review of arXiv:2412.00651}
}
read the original abstract

Recent advancements in multimodal pre-training models have significantly advanced computational pathology. However, current approaches predominantly rely on visual-language models, which may impose limitations from a molecular perspective and lead to performance bottlenecks. Here, we introduce a Unified Molecule-enhanced Pathology Image REpresentationn Learning framework (UMPIRE). UMPIRE aims to leverage complementary information from gene expression profiles to guide the multimodal pre-training, enhancing the molecular awareness of pathology image representation learning. We demonstrate that this molecular perspective provides a robust, task-agnostic training signal for learning pathology image embeddings. Due to the scarcity of paired data, approximately 4 million entries of spatial transcriptomics gene expression were collected to train the gene encoder. By leveraging powerful pre-trained encoders, UMPIRE aligns the encoders across over 697K pathology image-gene expression pairs. The performance of UMPIRE is demonstrated across various molecular-related downstream tasks, including gene expression prediction, spot classification, and mutation state prediction in whole slide images. Our findings highlight the effectiveness of multimodal data integration and open new avenues for exploring computational pathology enhanced by molecular perspectives. The code and pre-trained weights are available at https://github.com/Hanminghao/UMPIRE.

Figures

Figures reproduced from arXiv: 2412.00651 by the authors.

Figure 1
Figure 1. Overview of UMPIRE. First, approximately 4 million unlabeled spatial transcriptomics (ST) gene expression data were used to pre-train the Visiumformer for gene encoding. Next, a pre-trained pathologic Vision Transformer was adopted as the vision encoder. The symmetric contrastive loss LSCL is applied to align embeddings from both modalities. dataset, leading to suboptimal model performance due to limited training da… view at source ↗
Figure 2
Figure 2. Evaluation of Downstream Tasks. UMPIRE and baselines are assessed on: a. Bimodal gene expression prediction; b. Unimodal patch/spot classification; c. Vision-based WSI mutation state prediction. sample i. In this study, we set n to 1500, meaning that the context length for the gene encoder is 1500 tokens. Visiumformer Pre-training. Given a tokenized ST gene expression Ti ∈ R N , Visiumformer first applies an embed￾d… view at source ↗
Figure 3
Figure 3. Visualization of Bimodal Gene Expression Prediction. Ground truth and predicted spatially resolved expression levels for PIBF1 overlaying the whole slide image of sample patient-1-H2-5, visualized with a fixed (top) and a variable (bottom) color scale. the Spatial Transcriptomics platform [64], was then used to assess the transfer learning capabilities of UMPIRE across different technologies and platforms. Specifica… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Visualization of Linear Probing. a. Whole Slide Image and Ground Truth; b. Predicted spot/patch types for sample 151673, visualized before (top) and after (bottom) multimodal pre-training with contrastive loss; c. with reconstruction loss. this capability, UMPIRE-ADAPT…
Figure 5
Figure 5. Figure 5: MIL-based WSI Classification. Comparison of UMPIRE and baselines for WSI-level gene mutation state classification using MIL. a. Based on Phikon. b. Based on UNI. UMPIRE vs. UMPIRE-REC: In contrast to the improved consistency of UMPIRE across all datasets, the vision en…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

86 extracted references · 65 canonical work pages

  1. [1]

    Spatial deconvolution of her2-positive breast cancer delineates tumor- associated cell type interactions

    Alma Andersson, Ludvig Larsson, Linnea Stenbeck, Fredrik Salm´en, Anna Ehinger, Sunny Z Wu, Ghamdan Al-Eryani, Daniel Roden, Alex Swarbrick, ˚Ake Borg, et al. Spatial deconvolution of her2-positive breast cancer delineates tumor- associated cell type interactions. Nature communications, 12 (1):6012, 2021. 5, 20

  2. [2]

    Single-cell, single-nucleus, and spatial transcriptomics characterization of the immunological landscape in the healthy and psc human liver

    Tallulah S Andrews, Diana Nakib, Catia T Perciani, Xue Zhong Ma, Lewis Liu, Erin Winter, Damra Camat, Sai W Chung, Patricia Lumanto, Justin Manuel, et al. Single-cell, single-nucleus, and spatial transcriptomics characterization of the immunological landscape in the healthy and psc human liver. Journal of Hepatology, 80(5):730–743, 2024. 5, 19

  3. [3]

    Robust and data- efficient generalization of self-supervised machine learning for diagnostic imaging

    Shekoofeh Azizi, Laura Culp, Jan Freyberg, Basil Mustafa, Sebastien Baur, Simon Kornblith, Ting Chen, Nenad Tomasev, Jovana Mitrovi´c, Patricia Strachan, et al. Robust and data- efficient generalization of self-supervised machine learning for diagnostic imaging. Nature Biomedical Engineering, 7(6): 756–779, 2023. 4

  4. [4]

    Language models are few-shot learners

    Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 2

  5. [5]

    Emerg- ing properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 7

  6. [6]

    Stimage-1k4m: A histopathology image- gene expression dataset for spatial transcriptomics

    Jiawen Chen, Muqing Zhou, Wenrong Wu, Jinwei Zhang, Yun Li, and Didong Li. Stimage-1k4m: A histopathology image- gene expression dataset for spatial transcriptomics. arXiv preprint arXiv:2406.06393, 2024. 3, 18

  7. [7]

    Spatially resolved, highly mul- tiplexed rna profiling in single cells

    Kok Hao Chen, Alistair N Boettiger, Jeffrey R Moffitt, Siyuan Wang, and Xiaowei Zhuang. Spatially resolved, highly mul- tiplexed rna profiling in single cells. Science, 348(6233): aaa6090, 2015. 2

  8. [8]

    Towards a general-purpose foundation model for computational pathol- ogy

    Richard J Chen, Tong Ding, Ming Y Lu, Drew FK Williamson, Guillaume Jaume, Andrew H Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, et al. Towards a general-purpose foundation model for computational pathol- ogy. Nature Medicine, 30(3):850–862, 2024. 1, 2, 4, 5, 7, 9, 14

Show all 86 references
  1. [9]

    A simple framework for contrastive learning of visual representations

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020. 2

  2. [10]

    An empirical study of training self-supervised vision transformers

    Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9640–9649, 2021. 2

  3. [11]

    Pearson correlation coefficient

    Israel Cohen, Yiteng Huang, Jingdong Chen, Jacob Benesty, Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. Pearson correlation coefficient. Noise reduction in speech processing, pages 1–4, 2009. 6

  4. [12]

    Classification and mutation prediction from non–small cell lung cancer histopathology images using deep learning

    Nicolas Coudray, Paolo Santiago Ocampo, Theodore Sakel- laropoulos, Navneet Narula, Matija Snuderl, David Feny ¨o, Andre L Moreira, Narges Razavian, and Aristotelis Tsirigos. Classification and mutation prediction from non–small cell lung cancer histopathology images using dee...

  5. [13]

    scgpt: toward building a foundation model for single-cell multi-omics using generative ai

    Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengn- ing Luo, Nan Duan, and Bo Wang. scgpt: toward building a foundation model for single-cell multi-omics using generative ai. Nature Methods, pages 1–11, 2024. 2, 3

  6. [14]

    Vision transformers need registers, 2023

    Timoth´ee Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transformers need registers, 2023. 7 9

  7. [15]

    A cluster separation measure

    David L Davies and Donald W Bouldin. A cluster separation measure. IEEE transactions on pattern analysis and machine intelligence, (2):224–227, 1979. 8, 13, 15

  8. [16]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 2, 4, 5

  9. [17]

    Pathology-and-genomics multimodal transformer for survival outcome prediction

    Kexin Ding, Mu Zhou, Dimitris N Metaxas, and Shaoting Zhang. Pathology-and-genomics multimodal transformer for survival outcome prediction. In MICCAI, pages 622–631. Springer, 2023. 1, 2

  10. [18]

    Spatial pro- filing technologies illuminate the tumor microenvironment

    Ofer Elhanani, Raz Ben-Uri, and Leeat Keren. Spatial pro- filing technologies illuminate the tumor microenvironment. Cancer cell, 41(3):404–420, 2023. 2

  11. [19]

    Spatially resolved clonal copy number alter- ations in benign and malignant tissue

    Andrew Erickson, Mengxiao He, Emelie Berglund, Maja Marklund, Reza Mirzazadeh, Niklas Schultz, Linda Kvas- tad, Alma Andersson, Ludvig Bergenstr˚ahle, Joseph Bergen- str˚ahle, et al. Spatially resolved clonal copy number alter- ations in benign and malignant tissue. Nature, 60...

  12. [20]

    Scaling self-supervised learning for histopathology with masked image modeling

    Alexandre Filiot, Ridouane Ghermi, Antoine Olivier, Paul Jacob, Lucas Fidon, Alice Mac Kain, Charlie Saillard, and Jean-Baptiste Schiratti. Scaling self-supervised learning for histopathology with masked image modeling. medRxiv, pages 2023–07, 2023. 1, 2, 4, 5, 7, 13, 14

  13. [21]

    Light-sheet microscopy for slide-free non-destructive pathology of large clinical specimens

    Adam K Glaser, Nicholas P Reder, Ye Chen, Erin F McCarty, Chengbo Yin, Linpeng Wei, Yu Wang, Lawrence D True, and Jonathan TC Liu. Light-sheet microscopy for slide-free non-destructive pathology of large clinical specimens. Nature biomedical engineering, 1(7):0084, 2017. 1

  14. [22]

    Revealing dynamics of gene expression vari- ability in cell state space

    Dominic Gr¨un. Revealing dynamics of gene expression vari- ability in cell state space. Nature methods, 17(1):45–49, 2020. 13

  15. [23]

    Compre- hensive analysis of ubiquitously expressed genes in humans from a data-driven perspective

    Jianlei Gu, Jiawei Dai, Hui Lu, and Hongyu Zhao. Compre- hensive analysis of ubiquitously expressed genes in humans from a data-driven perspective. Genomics, Proteomics & Bioinformatics, 21(1):164–176, 2023. 13

  16. [24]

    Dimensional- ity reduction by learning an invariant mapping

    Raia Hadsell, Sumit Chopra, and Yann LeCun. Dimensional- ity reduction by learning an invariant mapping. In 2006 IEEE computer society conference on computer vision and pattern recognition (CVPR’06), pages 1735–1742. IEEE, 2006. 13

  17. [25]

    Integrating spatial gene expres- sion and breast tumour morphology via deep learning

    Bryan He, Ludvig Bergenstr˚ahle, Linnea Stenbeck, Abubakar Abid, Alma Andersson, ˚Ake Borg, Jonas Maaskola, Joakim Lundeberg, and James Zou. Integrating spatial gene expres- sion and breast tumour morphology via deep learning. Nature biomedical engineering, 4(8):827–834, 2020....

  18. [26]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 14

  19. [27]

    Momentum contrast for unsupervised visual rep- resentation learning

    Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual rep- resentation learning. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 9729–9738, 2020. 8

  20. [28]

    Masked autoencoders are scalable vision learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 16000– 16009, 2022. 2

  21. [29]

    Densely connected convolutional networks

    Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kil- ian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017. 14, 17

  22. [30]

    A visual–language foundation model for pathology image analysis using medical twitter

    Zhi Huang, Federico Bianchi, Mert Yuksekgonul, Thomas J Montine, and James Zou. A visual–language foundation model for pathology image analysis using medical twitter. Nature Medicine, pages 1–10, 2023. 1, 2, 9

  23. [31]

    Attention- based deep multiple instance learning

    Maximilian Ilse, Jakub Tomczak, and Max Welling. Attention- based deep multiple instance learning. In International con- ference on machine learning, pages 2127–2136. PMLR, 2018. 8, 17

  24. [32]

    Spatial transcriptomics in health and disease

    Sanjay Jain and Michael T Eadon. Spatial transcriptomics in health and disease. Nature Reviews Nephrology, pages 1–13,

  25. [33]

    High resolution mapping of the tumor microenvironment using integrated single-cell, spatial and in situ analysis

    Amanda Janesick, Robert Shelansky, Andrew D Gottscho, Florian Wagner, Stephen R Williams, Morgane Rouault, Ghezal Beliakoff, Carolyn A Morrison, Michelli F Oliveira, Jordan T Sicherman, et al. High resolution mapping of the tumor microenvironment using integrated single-cell, ...

  26. [34]

    Song, Ming Y

    Guillaume Jaume, Paul Doucet, Andrew H. Song, Ming Y . Lu, Cristina Almagro-Perez, Sophia J. Wagner, Anurag J. Vaidya, Richard J. Chen, Drew F. K. Williamson, Ahrong Kim, and Faisal Mahmood. HEST-1k: A Dataset for Spatial Transcriptomics and Histology Image Analysis. arXiv, 20...

  27. [35]

    Chen, Drew FK Williamson, Thomas Peeters, An- drew H

    Guillaume Jaume, Lukas Oldenburg, Anurag Jayant Vaidya, Richard J. Chen, Drew FK Williamson, Thomas Peeters, An- drew H. Song, and Faisal Mahmood. Transcriptomics-guided slide representation learning in computational pathology. In CVPR, 2024. 1, 2, 9

  28. [36]

    Modeling dense multimodal interactions between biological pathways and histology for survival prediction

    Guillaume Jaume, Anurag Vaidya, Richard Chen, Drew Williamson, Paul Liang, and Faisal Mahmood. Modeling dense multimodal interactions between biological pathways and histology for survival prediction. CVPR, 2024. 1

  29. [37]

    Scaling up visual and vision-language representation learning with noisy text supervision

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, pages 4904–

  30. [38]

    Thitogene: a deep learning method for predicting spatial transcriptomics from histological images

    Yuran Jia, Junliang Liu, Li Chen, Tianyi Zhao, and Yadong Wang. Thitogene: a deep learning method for predicting spatial transcriptomics from histological images. Briefings in Bioinformatics, 25(1):bbad464, 2024. 5, 14, 17, 18

  31. [39]

    Inferring histology-associated gene expression gradients in spatial transcriptomic studies

    Jan Kueckelhaus, Simon Frerich, Jasim Kada-Benotmane, Christina Koupourtidou, Jovica Ninkovic, Martin Dichgans, Juergen Beck, Oliver Schnell, and Dieter Henrik Heiland. Inferring histology-associated gene expression gradients in spatial transcriptomic studies. Nature Communica...

  32. [40]

    Dual-stream multiple instance learning network for whole slide image classification 10 with self-supervised contrastive learning

    Bin Li, Yin Li, and Kevin W Eliceiri. Dual-stream multiple instance learning network for whole slide image classification 10 with self-supervised contrastive learning. In CVPR, pages 14318–14328, 2021. 1

  33. [41]

    Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation. In ICML, 2022. 1

  34. [42]

    From bulk, single-cell to spatial rna sequencing

    Xinmin Li and Cun-Yu Wang. From bulk, single-cell to spatial rna sequencing. International journal of oral science, 13(1): 36, 2021. 2

  35. [43]

    Pibf1 regulates multiple gene expression via impeding long-range chromatin interaction to drive the ma- lignant transformation of hpv16 integration epithelial cells

    Xiaomin Li, Ci Ren, Anni Huang, Yue Zhao, Liming Wang, Hui Shen, Chun Gao, Bingxin Chen, Tong Zhu, Jinfeng Xiong, et al. Pibf1 regulates multiple gene expression via impeding long-range chromatin interaction to drive the ma- lignant transformation of hpv16 integration epitheli...

  36. [44]

    Exploiting geometric features via hierarchi- cal graph pyramid transformer for cancer diagnosis using histopathological images

    Mingxin Liu, Yunzan Liu, Pengbo Xu, Hui Cui, Jing Ke, and Jiquan Ma. Exploiting geometric features via hierarchi- cal graph pyramid transformer for cancer diagnosis using histopathological images. IEEE Transactions on Medical Imaging, 2024. 1

  37. [45]

    Data-efficient and weakly supervised computational pathology on whole- slide images

    Ming Y Lu, Drew FK Williamson, Tiffany Y Chen, Richard J Chen, Matteo Barbieri, and Faisal Mahmood. Data-efficient and weakly supervised computational pathology on whole- slide images. Nature biomedical engineering, 5(6):555–570,

  38. [46]

    A visual-language foun- dation model for computational pathology

    Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Ivy Liang, Tong Ding, Guillaume Jaume, Igor Odintsov, Long Phi Le, Georg Gerber, et al. A visual-language foun- dation model for computational pathology. Nature Medicine, 30:863–874, 2024. 1, 2, 3, 9, 20

  39. [47]

    Biomarkers in cancer staging, prognosis and treatment selection

    Joseph A Ludwig and John N Weinstein. Biomarkers in cancer staging, prognosis and treatment selection. Nature Reviews Cancer, 5(11):845–856, 2005. 1

  40. [48]

    Transcriptome-scale spatial gene expres- sion in the human dorsolateral prefrontal cortex

    Kristen R Maynard, Leonardo Collado-Torres, Lukas M We- ber, Cedric Uytingco, Brianna K Barry, Stephen R Williams, Joseph L Catallini, Matthew N Tran, Zachary Besich, Mad- havi Tippani, et al. Transcriptome-scale spatial gene expres- sion in the human dorsolateral prefrontal c...

  41. [49]

    Multimodal contrastive learning for spatial gene expression prediction using histology images

    Wenwen Min, Zhiceng Shi, Jun Zhang, Jun Wan, and Chang- miao Wang. Multimodal contrastive learning for spatial gene expression prediction using histology images. arXiv preprint arXiv:2407.08216, 2024. 2, 4, 5, 14, 17, 18

  42. [50]

    Caclust: linking genotype to transcriptional heterogene- ity of follicular lymphoma using bcr and exomic variants

    Kazimierz Oksza-Orzechowski, Edwin Quinten, Shadi Darvish Shafighi, Szymon M Kiełbasa, Hugo van Kessel, Ruben AL de Groen, Joost SP Vermaat, Julieta H Sepl´uveda-Y´a˜nez, Marcelo A Navarrete, Hendrik Veelken, et al. Caclust: linking genotype to transcriptional heterogene- ity ...

  43. [51]

    Repre- sentation learning with contrastive predictive coding

    Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. 13

  44. [52]

    Maxime Oquab, Timoth´ee Darcet, Theo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicola...

  45. [53]

    Leveraging infor- mation in spatial transcriptomics to predict super-resolution gene expression from histology images in tumors

    Minxing Pang, Kenong Su, and Mingyao Li. Leveraging infor- mation in spatial transcriptomics to predict super-resolution gene expression from histology images in tumors. BioRxiv, pages 2021–11, 2021. 2, 5, 14, 17, 18

  46. [54]

    Rethinking multiple instance learning for whole slide image classification: A good instance classifier is all you need

    Linhao Qu, Yingfan Ma, Xiaoyuan Luo, Qinhao Guo, Man- ning Wang, and Zhijian Song. Rethinking multiple instance learning for whole slide image classification: A good instance classifier is all you need. IEEE Transactions on Circuits and Systems for Video Technology, 2024. 2

  47. [55]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PMLR, 2021. 1, 2, 3, 6, 8, 20

  48. [56]

    Silhouettes: a graphical aid to the in- terpretation and validation of cluster analysis

    Peter J Rousseeuw. Silhouettes: a graphical aid to the in- terpretation and validation of cluster analysis. Journal of computational and applied mathematics, 20:53–65, 1987. 8, 13, 15, 20

  49. [57]

    H-optimus-0, 2024

    Charlie Saillard, Rodolphe Jenatton, Felipe Llinares-L´opez, Zelda Mariet, David Cahan´e, Eric Durand, and Jean-Philippe Vert. H-optimus-0, 2024. 1, 2

  50. [58]

    Nicheformer: a foundation model for single-cell and spatial omics

    Anna Christina Schaar, Alejandro Tejada-Lapuerta, Giovanni Palla, Robert Gutgesell, Lennard Halle, Mariia Minaeva, Larsen V ornholz, Leander Dony, Francesca Drummer, Mo- jtaba Bahrami, et al. Nicheformer: a foundation model for single-cell and spatial omics. bioRxiv, pages 202...

  51. [59]

    Transmil: Transformer based correlated multiple instance learning for whole slide image classification

    Zhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang, Jian Zhang, Xiangyang Ji, et al. Transmil: Transformer based correlated multiple instance learning for whole slide image classification. NIPS, 34:2136–2147, 2021. 1

  52. [60]

    A structure-aware hierarchical graph-based multiple instance learning framework for pt staging in histopathological image

    Jiangbo Shi, Lufei Tang, Yang Li, Xianli Zhang, Zeyu Gao, Yefeng Zheng, Chunbao Wang, Tieliang Gong, and Chen Li. A structure-aware hierarchical graph-based multiple instance learning framework for pt staging in histopathological image. IEEE Transactions on Medical Imaging, 42...

  53. [61]

    Vila-mil: Dual-scale vision-language multiple instance learning for whole slide image classification

    Jiangbo Shi, Chen Li, Tieliang Gong, Yefeng Zheng, and Huazhu Fu. Vila-mil: Dual-scale vision-language multiple instance learning for whole slide image classification. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11248–11258, 2024. 6

  54. [62]

    Multimodal prototyping for cancer survival prediction

    Andrew H Song, Richard J Chen, Guillaume Jaume, Anurag Jayant Vaidya, Alexander Baras, and Faisal Mah- mood. Multimodal prototyping for cancer survival prediction. In ICML, 2024. 1

  55. [63]

    Flex- ible experimental designs for valid single-cell rna-sequencing experiments allowing batch effects correction

    Fangda Song, Ga Ming Angus Chan, and Yingying Wei. Flex- ible experimental designs for valid single-cell rna-sequencing experiments allowing batch effects correction. Nature com- munications, 11(1):3274, 2020. 15

  56. [64]

    Visualization and analysis of gene expression in tis- sue sections by spatial transcriptomics

    Patrik L St˚ahl, Fredrik Salm´en, Sanja Vickovic, Anna Lund- mark, Jos ´e Fern ´andez Navarro, Jens Magnusson, Stefania Giacomello, Michaela Asp, Jakub O Westholm, Mikael Huss, 11 et al. Visualization and analysis of gene expression in tis- sue sections by spatial transcriptom...

  57. [65]

    Visualization and analysis of gene expression in tis- sue sections by spatial transcriptomics

    Patrik L St˚ahl, Fredrik Salm´en, Sanja Vickovic, Anna Lund- mark, Jos ´e Fern ´andez Navarro, Jens Magnusson, Stefania Giacomello, Michaela Asp, Jakub O Westholm, Mikael Huss, et al. Visualization and analysis of gene expression in tis- sue sections by spatial transcriptomics...

  58. [66]

    Trans- fer learning enables predictions in network biology

    Christina V Theodoris, Ling Xiao, Anant Chopra, Mark D Chaffin, Zeina R Al Sayed, Matthew C Hill, Helene Mantineo, Elizabeth M Brydon, Zexian Zeng, X Shirley Liu, et al. Trans- fer learning enables predictions in network biology. Nature, 618(7965):616–624, 2023. 2, 3

  59. [67]

    The expanding vistas of spatial transcriptomics

    Luyi Tian, Fei Chen, and Evan Z Macosko. The expanding vistas of spatial transcriptomics. Nature Biotechnology, 41 (6):773–782, 2023. 2

  60. [68]

    Visualizing data using t-sne

    Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11),

  61. [69]

    Vir- chow: a million-slide digital pathology foundation model

    Eugene V orontsov, Alican Bozkurt, Adam Casson, George Shaikovski, Michal Zelechowski, Siqi Liu, Kristen Severson, Eric Zimmermann, James Hall, Neil Tenenholtz, et al. Vir- chow: a million-slide digital pathology foundation model. arXiv preprint arXiv:2309.07778, 2023. 4

  62. [70]

    Mgiml: Cancer grading with incomplete radiology-pathology data via memory learning and gradient homogenization

    Pengyu Wang, Huaqi Zhang, Meilu Zhu, Xi Jiang, Jing Qin, and Yixuan Yuan. Mgiml: Cancer grading with incomplete radiology-pathology data via memory learning and gradient homogenization. IEEE Transactions on Medical Imaging ,

  63. [71]

    Transformer-based unsupervised contrastive learning for histopathological image classification

    Xiyue Wang, Sen Yang, Jun Zhang, Minghui Wang, Jing Zhang, Wei Yang, Junzhou Huang, and Xiao Han. Transformer-based unsupervised contrastive learning for histopathological image classification. Medical image analy- sis, 81:102559, 2022. 4

  64. [72]

    A pathology foundation model for cancer diagnosis and prognosis prediction

    Xiyue Wang, Junhan Zhao, Eliana Marostica, Wei Yuan, Ji- etian Jin, Jiayu Zhang, Ruijiang Li, Hongping Tang, Kanran Wang, Yu Li, et al. A pathology foundation model for cancer diagnosis and prognosis prediction. Nature, pages 1–9, 2024. 1

  65. [73]

    Scanpy: large-scale single-cell gene expression data analysis

    F Alexander Wolf, Philipp Angerer, and Fabian J Theis. Scanpy: large-scale single-cell gene expression data analysis. Genome biology, 19:1–5, 2018. 16

  66. [74]

    Cathepsin c promotes breast cancer lung metastasis by modulating neutrophil infiltration and neutrophil extracellular trap formation

    Yansen Xiao, Min Cong, Jiatao Li, Dasa He, Qiuyao Wu, Pu Tian, Yuan Wang, Shuaixi Yang, Chenxi Liang, Yajun Liang, et al. Cathepsin c promotes breast cancer lung metastasis by modulating neutrophil infiltration and neutrophil extracellular trap formation. Cancer cell, 39(3):42...

  67. [75]

    Spatially resolved gene expression prediction from histology images via bi- modal contrastive learning

    Ronald Xie, Kuan Pang, Sai Chung, Catia Perciani, Sonya MacParland, Bo Wang, and Gary Bader. Spatially resolved gene expression prediction from histology images via bi- modal contrastive learning. NIPS, 36, 2024. 2, 4, 5, 14, 16, 17, 18

  68. [76]

    Unsupervised spa- tially embedded deep representation of spatial transcriptomics

    Hang Xu, Huazhu Fu, Yahui Long, Kok Siong Ang, Raman Sethi, Kelvin Chong, Mengwei Li, Rom Uddamvathanak, Hong Kai Lee, Jingjing Ling, et al. Unsupervised spa- tially embedded deep representation of spatial transcriptomics. Genome Medicine, 16(1):12, 2024. 5, 20

  69. [77]

    A whole-slide foundation model for digital pathology from real-world data

    Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier Gonz ´alez, Yu Gu, et al. A whole-slide foundation model for digital pathology from real-world data. Nature, pages 1–8, 2024. 1, 4

  70. [78]

    A multimodal knowledge-enhanced whole-slide pathology foundation model

    Yingxue Xu, Yihui Wang, Fengtao Zhou, Jiabo Ma, Shu Yang, Huangjing Lin, Xin Wang, Jiguang Wang, Li Liang, Anjia Han, et al. A multimodal knowledge-enhanced whole-slide pathology foundation model. arXiv preprint arXiv:2407.15362, 2024. 1, 2

  71. [79]

    Stomicsdb: a comprehensive database for spatial transcriptomics data sharing, analysis and visualization

    Zhicheng Xu, Weiwen Wang, Tao Yang, Ling Li, Xizheng Ma, Jing Chen, Jieyu Wang, Yan Huang, Joshua Gould, Huifang Lu, et al. Stomicsdb: a comprehensive database for spatial transcriptomics data sharing, analysis and visualization. Nu- cleic acids research, 52(D1):D1053–D1061, 2...

  72. [80]

    Coca: Contrastive captioners are image-text foundation models

    Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mo- jtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022. 1, 2

  73. [81]

    Batch alignment of single-cell transcriptomics data using deep metric learning

    Xiaokang Yu, Xinyi Xu, Jingxiao Zhang, and Xiangjie Li. Batch alignment of single-cell transcriptomics data using deep metric learning. Nature communications, 14(1):960,

  74. [82]

    Heartsvg: a fast and accurate method for identifying spatially variable genes in large-scale spatial transcriptomics

    Xin Yuan, Yanran Ma, Ruitian Gao, Shuya Cui, Yifan Wang, Botao Fa, Shiyang Ma, Ting Wei, Shuangge Ma, and Zhang- sheng Yu. Heartsvg: a fast and accurate method for identifying spatially variable genes in large-scale spatial transcriptomics. Nature Communications, 15(1):5700, 2024. 13

  75. [83]

    Sodb facilitates comprehensive exploration of spatial omics data

    Zhiyuan Yuan, Wentao Pan, Xuan Zhao, Fangyuan Zhao, Zhimeng Xu, Xiu Li, Yi Zhao, Michael Q Zhang, and Jianhua Yao. Sodb facilitates comprehensive exploration of spatial omics data. Nature Methods, 20(3):387–399, 2023. 3, 18

  76. [84]

    Spatial transcriptomics prediction from histology jointly through transformer and graph neural networks

    Yuansong Zeng, Zhuoyi Wei, Weijiang Yu, Rui Yin, Yuchen Yuan, Bingling Li, Zhonghui Tang, Yutong Lu, and Yue- dong Yang. Spatial transcriptomics prediction from histology jointly through transformer and graph neural networks. Brief- ings in Bioinformatics , 23(5):bbac297, 2022...

  77. [85]

    Metabolic regulation of cell growth and proliferation

    Jiajun Zhu and Craig B Thompson. Metabolic regulation of cell growth and proliferation. Nature reviews Molecular cell biology, 20(7):436–450, 2019. 2 12 A. More Experimental Results A.1. Impact of Unimodal Pre-training The pre-training process of UMPIRE is divided into two sta...

  78. [86]

    Widespread Social Impact: Technological advancements should benefit a broader population

    developing multimodal models capable of robust general- ization across different platforms and technologies. Widespread Social Impact: Technological advancements should benefit a broader population. While molecular-level analyses of cancer significantly enhance diagnostic accu...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.