REVIEW 3 major objections 5 minor 86 references
Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims spatial gene-expression profiles give pathology image encoders a task-agnostic molecular awareness, improving gene-expression prediction, tissue classification, and whole-slide mutation-state prediction once the…
desk verdict A genuinely new large-scale recipe for aligning pathology images with spatial transcriptomics, but the paper's strongest evidence for its headline claim is compromised by same-patient leakage in the slice-level leave-one-out evaluation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a two-stage alignment whose first stage is Visiumformer, a 12-layer Transformer pre-trained with BERT-style masked-token prediction on tokenized spatial transcriptomics: each spot's gene-expression vector is normalised against per-gene means, sorted by expression level, and truncated to the top 1,500 gene indices, making the input order-agnostic. The second stage aligns a pre-trained pathology vision encoder (Phikon, ViT-B/16, or UNI, ViT-L/16) with the gene encoder using a symmetric contrastive loss in a shared 512-dimensional space over 696,636 pathology-image-expression pairs, pulling paired image and gene embeddings together and pushing unpaired ones apart. For gene-expression prediction at inference, the image is embedded as a query vector, compared by cosine similarity against a reference database of gene embeddings, and the top-K neighbours' expression profiles are weighted-aggregated into the prediction. The contrastive objective, rather than regression or reconstruction, is what the paper identifies as preserving the vision encoder's visual semantics while adding molecular structure; ablations replacing it with MSE or L1 loss degrade classification F1 by roughly 16 points.
What would settle it
Re-run gene-expression prediction with patient-level held-out splits, removing every slice of the test patient from both the fine-tuning set and the reference database; if the Pearson-correlation advantage over BLEEP collapses toward zero, the reported molecular awareness would be shown to be transductive memory rather than a general representation. A complementary check is zero-shot transfer to a tissue type absent from both pre-training corpora, where retention of the gains would confirm that the signal is task-agnostic.
Extended reading notes
Core claim
The central claim is that the molecular perspective is a robust, task-agnostic training signal for pathology image embeddings, in the sense that aligning image embeddings with gene-expression embeddings produces representations that are better at molecular-related tasks while preserving, and in some cases improving, purely visual performance. Concretely, the paper reports that its aligned encoders outperform the leading contrastive baseline BLEEP by an average of +42.9% in Pearson correlation for gene-expression prediction when fully fine-tuned, improve balanced accuracy in DLPFC cortical-layer classification by up to +42.3% relative to the base vision encoder, and improve WSI mutation-state AUC in three of four genes, while also outperforming the bulk-RNA method TANGLE in three of four mutation sub-tasks. The authors attribute the gains to the two-stage design: large-scale unimodal pre-training of the gene encoder, followed by contrastive alignment that teaches the image encoder molecular structure without destroying its visual semantics, which they support with ablations showing that replacing the symmetric contrastive loss with regression losses costs about 16 points of weighted F1.
Load-bearing premise
The central claim stands on the assumption that the leave-one-out slice evaluation measures learned molecular understanding rather than patient memory, since HLT's four sections come from one person, HPC's five from two, and the prediction reference database is built from remaining slices that can belong to the test patient.
Editorial extensions
If this is right
- If the molecular signal is task-agnostic, a single aligned vision encoder can serve gene-expression prediction, spot-level tissue classification, and WSI-level mutation-state prediction, replacing task-specific models that must be retrained per task.
- The reported +83.8% average PCC gain over BLEEP on the HER2+ dataset, sequenced with an older platform never seen in pre-training, implies the molecular alignment transfers across sequencing technologies, not just across tissue types.
- Because the contrastive alignment preserves the vision encoder's original semantics, as evidenced by 10X Breast classification not degrading, the recipe could extend to other molecular modalities such as protein expression or methylation without sacrificing visual performance.
- The ViSTomics-4M pre-training corpus, at 3.94 million spots across 1,363 slides and 30 tissues, is claimed to be the largest Visium-based dataset assembled to date, and the code and pre-trained weights are released, so the molecular encoder is reusable for downstream spatial-omics tasks.
- Adapters using 0.3-0.8% of the fine-tuning parameters come within 2.8% of full fine-tuning, suggesting the molecular alignment can be attached to large vision encoders at low computational cost in resource-limited settings.
Reading between the lines
- The leave-one-out slice protocol leaves a patient-level re-split as the natural next check; if the gains persist with all same-patient slices removed from the reference set, the task-agnostic claim would be considerably strengthened.
- The tokenization scheme, sorting a spot's 20,310 genes by normalised expression and keeping the top 1,500, discards the spatial relationships among spots, so adding positional or neighbourhood context to the gene encoder is an implicit extension that could raise the ceiling on expression prediction.
- The same two-stage recipe could in principle be applied to other tile-based histology readouts that are expensive to obtain, such as microsatellite-instability scoring or tumour purity, making molecular-aware image embeddings a cheap surrogate for assays that currently require sequencing.
- Because the aligned embeddings place images and expression in one space, retrieval-based analysis becomes possible at inference without sequencing: a new H&E slide could be matched to reference expression programs, yielding an interpretable, patient-specific molecular sketch, a use the paper's query-reference design already anticipates but does not develop.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. UMPIRE proposes a two-stage framework for aligning pathology image patches with spatial transcriptomics gene expression. In the first stage, the authors pre-train a BERT-like gene encoder (Visiumformer) on roughly 3.94 million Visium spots from the new ViSTomics-4M collection; in the second stage, they use symmetric contrastive learning (Eq. 6) to align Phikon or UNI image encoders with Visiumformer over 697K paired pathology-image and gene-expression samples from the HEST dataset. The resulting representations are evaluated on three molecular-related tasks: query-reference gene expression prediction (Section 4.3), linear-probing spot/patch classification (Section 4.4), and MIL-based WSI mutation state prediction (Section 4.5). The paper reports consistent improvements over several baselines and releases code and pretrained weights.
Significance. The paper addresses a timely and important goal: injecting molecular information into pathology image representations. The scale of pre-training data, the systematic comparison of multiple vision encoders (Phikon and UNI), the use of two fine-tuning strategies (adapter and full fine-tuning), and the public release of code and weights are all strengths. If the reported results survive a leakage-free evaluation, UMPIRE would be a solid contribution to computational pathology and multimodal representation learning. However, the current evaluation does not yet establish the paper's central claim of a 'robust, task-agnostic training signal', because the main evidence is affected by a patient/donor leakage issue.
major comments (3)
- [Section 4.3, Eq. (8)-(9), Appendix B.2, D.3] The gene expression prediction evaluation uses leave-one-out cross-validation at the slice level. As the paper states in Appendix D.3, HLT has four sections from one individual, HPC has five sections from two patients, and HER2+ has 32 slides from seven patients. For each held-out slice, the reference database in Eq. (8)-(9) is built from the remaining slices, which therefore include adjacent or same-patient sections. Because prediction is a top-K retrieval of real expression profiles, the model can achieve high PCC by recognizing patient-specific expression programs or batch artifacts rather than by learning a general image-to-expression mapping. Section 4.3 says the downstream datasets were excluded from pre-training to eliminate data leakage, but that statement does not address patient-level leakage within the leave-one-out folds. I request a patient-level (or donor-level) split, or at minimum an analysis showing that the reported PCC gains over BLEEP are not driven by same-patient references (e.g., compare retrieval from same-patient versus cross-patient reference sets). Without this, the +39.0% and +42.9% average improvements over BLEEP in Table 1 may reflect transductive memory.
- [Section 4.4, Table 2, Appendix B.2] The DLPFC linear-probing evaluation in Table 2 also uses slice-level leave-one-out cross-validation, but DLPFC consists of 12 sections from only three healthy donors. The reported improvements in balanced accuracy (+28% to +42% depending on the encoder) could therefore be donor-specific: the classifier may learn donor identity from the 11 training slices and apply it to the held-out slice of the same donor. Since this task is a major piece of evidence for the claim that the molecular perspective improves image embeddings, the authors should report donor-level cross-validation or per-donor results, and should discuss whether the gains persist when the training set contains no sections from the test donor.
- [Section 4.5, Figure 5] The WSI mutation task uses patient-level five-fold cross-validation, which is the appropriate protocol, but the results are modest and inconsistent: UMPIRE outperforms the original vision encoder in only three of four subtasks, and it lags TANGLE on EGFR by about 3.65% in AUC. Consequently, this experiment alone cannot carry the abstract's claim that the molecular perspective provides a robust, task-agnostic training signal. The paper would be strengthened by a clearer statement of which evidence is meant to support task-agnosticism once the leakage-prone experiments are re-evaluated.
minor comments (5)
- [Title and Abstract] The word 'REpresentationn' is misspelled in the title and abstract; the correct spelling is 'Representation'.
- [Figure 2] The label 'Get Embedding durning Infernece' contains typographical errors; it should read 'Get Embedding during Inference'.
- [Section 3.2] In the description of the pre-training setup, 'per-training' appears where 'pre-training' is intended.
- [Section 4.3] The paper's use of 'task-agnostic' could be more precise, since all downstream tasks are molecular-related (gene expression prediction, spot classification of molecularly defined layers, and mutation prediction). Clarifying that the claim concerns transfer across molecular-related tasks would be helpful.
- [Section F, Limitations] The Limitations section does not mention the potential for patient/donor leakage in the slice-level leave-one-out evaluations; this should be acknowledged and addressed in the revision.
Circularity Check
No circular derivation: UMPIRE's contrastive alignment is evaluated on external benchmarks; the slice-level leave-one-out leakage is a benchmark-validity caveat, not a reduction of the prediction to its fitted inputs.
full rationale
The paper's derivation chain is self-contained. Visiumformer is pre-trained with a masked-language-modeling loss on 3.94 million unpaired spatial transcriptomics profiles (Eq. 5), the vision encoder is an external pathology foundation model (Phikon or UNI), and the alignment stage optimizes a symmetric contrastive loss (Eq. 6) on HEST pairs from which the downstream evaluation datasets are explicitly excluded: 'the datasets used in this section were not included in the pre-training phase.' Downstream gene-expression prediction does not fit the held-out query slice's expression values: Eqs. 8-9 retrieve and weight reference profiles by cosine similarity, and the reported Pearson correlation is computed against a query slice that is not in the training fold. No parameter is fitted directly to the target values being predicted. The main residual concern is benchmark design rather than circularity: leave-one-out is performed at the slice level for HLT (four sections from one individual), HPC (five sections from two patients), and DLPFC (twelve sections from three donors), so the reference database and fine-tuning folds can contain adjacent sections of the same patient; this could inflate PCC through transductive memory of patient-specific expression programs. That is a leakage and external-validity limitation, not an equation-level reduction of the claimed molecular signal to its own inputs. The only self-citation (reference [54], a prior work by co-author Linhao Qu) is a non-load-bearing literature citation in the introduction. No load-bearing claim reduces by construction to a fitted parameter, a self-citation chain, or a renamed known result.
Assumptions & free parameters
free parameters (4)
- context length N (gene tokens) =
1500
- masking probability =
0.15
- temperature tau in contrastive loss =
not reported
- top-K references for gene expression prediction =
not specified
assumptions (5)
- domain assumption Gene expression anomalies correspond to discernible morphological patterns in pathology images
- domain assumption Visium 55-micron spots align with 224x224 patches extracted at 50-100 um
- domain assumption HEST filtering and exclusion of downstream datasets prevents data leakage
- ad hoc to paper Masked language modeling on sorted gene tokens is a valid self-supervised objective for gene expression
- standard math Transformer self-attention and symmetric contrastive loss are used as in prior work
Cite this review
Pith. "Pith review of Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics." pith.science (2026). https://pith.science/paper/X2S34HW3
@misc{pith2026241200651,
author = {Pith},
title = {Pith review of: Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics},
year = {2026},
howpublished = {\url{https://pith.science/paper/X2S34HW3}},
note = {Machine review of arXiv:2412.00651}
}
read the original abstract
Recent advancements in multimodal pre-training models have significantly advanced computational pathology. However, current approaches predominantly rely on visual-language models, which may impose limitations from a molecular perspective and lead to performance bottlenecks. Here, we introduce a Unified Molecule-enhanced Pathology Image REpresentationn Learning framework (UMPIRE). UMPIRE aims to leverage complementary information from gene expression profiles to guide the multimodal pre-training, enhancing the molecular awareness of pathology image representation learning. We demonstrate that this molecular perspective provides a robust, task-agnostic training signal for learning pathology image embeddings. Due to the scarcity of paired data, approximately 4 million entries of spatial transcriptomics gene expression were collected to train the gene encoder. By leveraging powerful pre-trained encoders, UMPIRE aligns the encoders across over 697K pathology image-gene expression pairs. The performance of UMPIRE is demonstrated across various molecular-related downstream tasks, including gene expression prediction, spot classification, and mutation state prediction in whole slide images. Our findings highlight the effectiveness of multimodal data integration and open new avenues for exploring computational pathology enhanced by molecular perspectives. The code and pre-trained weights are available at https://github.com/Hanminghao/UMPIRE.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Spatial deconvolution of her2-positive breast cancer delineates tumor- associated cell type interactions
Alma Andersson, Ludvig Larsson, Linnea Stenbeck, Fredrik Salm´en, Anna Ehinger, Sunny Z Wu, Ghamdan Al-Eryani, Daniel Roden, Alex Swarbrick, ˚Ake Borg, et al. Spatial deconvolution of her2-positive breast cancer delineates tumor- associated cell type interactions. Nature communications, 12 (1):6012, 2021. 5, 20
2021
-
[2]
Single-cell, single-nucleus, and spatial transcriptomics characterization of the immunological landscape in the healthy and psc human liver
Tallulah S Andrews, Diana Nakib, Catia T Perciani, Xue Zhong Ma, Lewis Liu, Erin Winter, Damra Camat, Sai W Chung, Patricia Lumanto, Justin Manuel, et al. Single-cell, single-nucleus, and spatial transcriptomics characterization of the immunological landscape in the healthy and psc human liver. Journal of Hepatology, 80(5):730–743, 2024. 5, 19
2024
-
[3]
Robust and data- efficient generalization of self-supervised machine learning for diagnostic imaging
Shekoofeh Azizi, Laura Culp, Jan Freyberg, Basil Mustafa, Sebastien Baur, Simon Kornblith, Ting Chen, Nenad Tomasev, Jovana Mitrovi´c, Patricia Strachan, et al. Robust and data- efficient generalization of self-supervised machine learning for diagnostic imaging. Nature Biomedical Engineering, 7(6): 756–779, 2023. 4
2023
-
[4]
Language models are few-shot learners
Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 2
arXiv 2005
-
[5]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 7
2021
-
[6]
Stimage-1k4m: A histopathology image- gene expression dataset for spatial transcriptomics
Jiawen Chen, Muqing Zhou, Wenrong Wu, Jinwei Zhang, Yun Li, and Didong Li. Stimage-1k4m: A histopathology image- gene expression dataset for spatial transcriptomics. arXiv preprint arXiv:2406.06393, 2024. 3, 18
arXiv 2024
-
[7]
Spatially resolved, highly mul- tiplexed rna profiling in single cells
Kok Hao Chen, Alistair N Boettiger, Jeffrey R Moffitt, Siyuan Wang, and Xiaowei Zhuang. Spatially resolved, highly mul- tiplexed rna profiling in single cells. Science, 348(6233): aaa6090, 2015. 2
2015
-
[8]
Towards a general-purpose foundation model for computational pathol- ogy
Richard J Chen, Tong Ding, Ming Y Lu, Drew FK Williamson, Guillaume Jaume, Andrew H Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, et al. Towards a general-purpose foundation model for computational pathol- ogy. Nature Medicine, 30(3):850–862, 2024. 1, 2, 4, 5, 7, 9, 14
2024
Show all 86 references
-
[9]
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020. 2
2020
-
[10]
An empirical study of training self-supervised vision transformers
Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9640–9649, 2021. 2
2021
-
[11]
Pearson correlation coefficient
Israel Cohen, Yiteng Huang, Jingdong Chen, Jacob Benesty, Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. Pearson correlation coefficient. Noise reduction in speech processing, pages 1–4, 2009. 6
2009
-
[12]
Classification and mutation prediction from non–small cell lung cancer histopathology images using deep learning
Nicolas Coudray, Paolo Santiago Ocampo, Theodore Sakel- laropoulos, Navneet Narula, Matija Snuderl, David Feny ¨o, Andre L Moreira, Narges Razavian, and Aristotelis Tsirigos. Classification and mutation prediction from non–small cell lung cancer histopathology images using dee...
2018
-
[13]
scgpt: toward building a foundation model for single-cell multi-omics using generative ai
Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengn- ing Luo, Nan Duan, and Bo Wang. scgpt: toward building a foundation model for single-cell multi-omics using generative ai. Nature Methods, pages 1–11, 2024. 2, 3
2024
-
[14]
Vision transformers need registers, 2023
Timoth´ee Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transformers need registers, 2023. 7 9
2023
-
[15]
A cluster separation measure
David L Davies and Donald W Bouldin. A cluster separation measure. IEEE transactions on pattern analysis and machine intelligence, (2):224–227, 1979. 8, 13, 15
1979
-
[16]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 2, 4, 5
2018 arXiv
-
[17]
Pathology-and-genomics multimodal transformer for survival outcome prediction
Kexin Ding, Mu Zhou, Dimitris N Metaxas, and Shaoting Zhang. Pathology-and-genomics multimodal transformer for survival outcome prediction. In MICCAI, pages 622–631. Springer, 2023. 1, 2
2023
-
[18]
Spatial pro- filing technologies illuminate the tumor microenvironment
Ofer Elhanani, Raz Ben-Uri, and Leeat Keren. Spatial pro- filing technologies illuminate the tumor microenvironment. Cancer cell, 41(3):404–420, 2023. 2
2023
-
[19]
Spatially resolved clonal copy number alter- ations in benign and malignant tissue
Andrew Erickson, Mengxiao He, Emelie Berglund, Maja Marklund, Reza Mirzazadeh, Niklas Schultz, Linda Kvas- tad, Alma Andersson, Ludvig Bergenstr˚ahle, Joseph Bergen- str˚ahle, et al. Spatially resolved clonal copy number alter- ations in benign and malignant tissue. Nature, 60...
2022
-
[20]
Scaling self-supervised learning for histopathology with masked image modeling
Alexandre Filiot, Ridouane Ghermi, Antoine Olivier, Paul Jacob, Lucas Fidon, Alice Mac Kain, Charlie Saillard, and Jean-Baptiste Schiratti. Scaling self-supervised learning for histopathology with masked image modeling. medRxiv, pages 2023–07, 2023. 1, 2, 4, 5, 7, 13, 14
2023
-
[21]
Light-sheet microscopy for slide-free non-destructive pathology of large clinical specimens
Adam K Glaser, Nicholas P Reder, Ye Chen, Erin F McCarty, Chengbo Yin, Linpeng Wei, Yu Wang, Lawrence D True, and Jonathan TC Liu. Light-sheet microscopy for slide-free non-destructive pathology of large clinical specimens. Nature biomedical engineering, 1(7):0084, 2017. 1
2017
-
[22]
Revealing dynamics of gene expression vari- ability in cell state space
Dominic Gr¨un. Revealing dynamics of gene expression vari- ability in cell state space. Nature methods, 17(1):45–49, 2020. 13
2020
-
[23]
Compre- hensive analysis of ubiquitously expressed genes in humans from a data-driven perspective
Jianlei Gu, Jiawei Dai, Hui Lu, and Hongyu Zhao. Compre- hensive analysis of ubiquitously expressed genes in humans from a data-driven perspective. Genomics, Proteomics & Bioinformatics, 21(1):164–176, 2023. 13
2023
-
[24]
Dimensional- ity reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun. Dimensional- ity reduction by learning an invariant mapping. In 2006 IEEE computer society conference on computer vision and pattern recognition (CVPR’06), pages 1735–1742. IEEE, 2006. 13
2006
-
[25]
Integrating spatial gene expres- sion and breast tumour morphology via deep learning
Bryan He, Ludvig Bergenstr˚ahle, Linnea Stenbeck, Abubakar Abid, Alma Andersson, ˚Ake Borg, Jonas Maaskola, Joakim Lundeberg, and James Zou. Integrating spatial gene expres- sion and breast tumour morphology via deep learning. Nature biomedical engineering, 4(8):827–834, 2020....
2020
-
[26]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 14
2016
-
[27]
Momentum contrast for unsupervised visual rep- resentation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual rep- resentation learning. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 9729–9738, 2020. 8
2020
-
[28]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 16000– 16009, 2022. 2
2022
-
[29]
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kil- ian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017. 14, 17
2017
-
[30]
A visual–language foundation model for pathology image analysis using medical twitter
Zhi Huang, Federico Bianchi, Mert Yuksekgonul, Thomas J Montine, and James Zou. A visual–language foundation model for pathology image analysis using medical twitter. Nature Medicine, pages 1–10, 2023. 1, 2, 9
2023
-
[31]
Attention- based deep multiple instance learning
Maximilian Ilse, Jakub Tomczak, and Max Welling. Attention- based deep multiple instance learning. In International con- ference on machine learning, pages 2127–2136. PMLR, 2018. 8, 17
2018
-
[32]
Spatial transcriptomics in health and disease
Sanjay Jain and Michael T Eadon. Spatial transcriptomics in health and disease. Nature Reviews Nephrology, pages 1–13,
-
[33]
High resolution mapping of the tumor microenvironment using integrated single-cell, spatial and in situ analysis
Amanda Janesick, Robert Shelansky, Andrew D Gottscho, Florian Wagner, Stephen R Williams, Morgane Rouault, Ghezal Beliakoff, Carolyn A Morrison, Michelli F Oliveira, Jordan T Sicherman, et al. High resolution mapping of the tumor microenvironment using integrated single-cell, ...
-
[34]
Song, Ming Y
Guillaume Jaume, Paul Doucet, Andrew H. Song, Ming Y . Lu, Cristina Almagro-Perez, Sophia J. Wagner, Anurag J. Vaidya, Richard J. Chen, Drew F. K. Williamson, Ahrong Kim, and Faisal Mahmood. HEST-1k: A Dataset for Spatial Transcriptomics and Histology Image Analysis. arXiv, 20...
2024
-
[35]
Chen, Drew FK Williamson, Thomas Peeters, An- drew H
Guillaume Jaume, Lukas Oldenburg, Anurag Jayant Vaidya, Richard J. Chen, Drew FK Williamson, Thomas Peeters, An- drew H. Song, and Faisal Mahmood. Transcriptomics-guided slide representation learning in computational pathology. In CVPR, 2024. 1, 2, 9
2024
-
[36]
Modeling dense multimodal interactions between biological pathways and histology for survival prediction
Guillaume Jaume, Anurag Vaidya, Richard Chen, Drew Williamson, Paul Liang, and Faisal Mahmood. Modeling dense multimodal interactions between biological pathways and histology for survival prediction. CVPR, 2024. 1
2024
-
[37]
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, pages 4904–
-
[38]
Thitogene: a deep learning method for predicting spatial transcriptomics from histological images
Yuran Jia, Junliang Liu, Li Chen, Tianyi Zhao, and Yadong Wang. Thitogene: a deep learning method for predicting spatial transcriptomics from histological images. Briefings in Bioinformatics, 25(1):bbad464, 2024. 5, 14, 17, 18
2024
-
[39]
Inferring histology-associated gene expression gradients in spatial transcriptomic studies
Jan Kueckelhaus, Simon Frerich, Jasim Kada-Benotmane, Christina Koupourtidou, Jovica Ninkovic, Martin Dichgans, Juergen Beck, Oliver Schnell, and Dieter Henrik Heiland. Inferring histology-associated gene expression gradients in spatial transcriptomic studies. Nature Communica...
2024
-
[40]
Dual-stream multiple instance learning network for whole slide image classification 10 with self-supervised contrastive learning
Bin Li, Yin Li, and Kevin W Eliceiri. Dual-stream multiple instance learning network for whole slide image classification 10 with self-supervised contrastive learning. In CVPR, pages 14318–14328, 2021. 1
2021
-
[41]
Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation. In ICML, 2022. 1
2022
-
[42]
From bulk, single-cell to spatial rna sequencing
Xinmin Li and Cun-Yu Wang. From bulk, single-cell to spatial rna sequencing. International journal of oral science, 13(1): 36, 2021. 2
2021
-
[43]
Pibf1 regulates multiple gene expression via impeding long-range chromatin interaction to drive the ma- lignant transformation of hpv16 integration epithelial cells
Xiaomin Li, Ci Ren, Anni Huang, Yue Zhao, Liming Wang, Hui Shen, Chun Gao, Bingxin Chen, Tong Zhu, Jinfeng Xiong, et al. Pibf1 regulates multiple gene expression via impeding long-range chromatin interaction to drive the ma- lignant transformation of hpv16 integration epitheli...
2024
-
[44]
Exploiting geometric features via hierarchi- cal graph pyramid transformer for cancer diagnosis using histopathological images
Mingxin Liu, Yunzan Liu, Pengbo Xu, Hui Cui, Jing Ke, and Jiquan Ma. Exploiting geometric features via hierarchi- cal graph pyramid transformer for cancer diagnosis using histopathological images. IEEE Transactions on Medical Imaging, 2024. 1
2024
-
[45]
Data-efficient and weakly supervised computational pathology on whole- slide images
Ming Y Lu, Drew FK Williamson, Tiffany Y Chen, Richard J Chen, Matteo Barbieri, and Faisal Mahmood. Data-efficient and weakly supervised computational pathology on whole- slide images. Nature biomedical engineering, 5(6):555–570,
-
[46]
A visual-language foun- dation model for computational pathology
Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Ivy Liang, Tong Ding, Guillaume Jaume, Igor Odintsov, Long Phi Le, Georg Gerber, et al. A visual-language foun- dation model for computational pathology. Nature Medicine, 30:863–874, 2024. 1, 2, 3, 9, 20
2024
-
[47]
Biomarkers in cancer staging, prognosis and treatment selection
Joseph A Ludwig and John N Weinstein. Biomarkers in cancer staging, prognosis and treatment selection. Nature Reviews Cancer, 5(11):845–856, 2005. 1
2005
-
[48]
Transcriptome-scale spatial gene expres- sion in the human dorsolateral prefrontal cortex
Kristen R Maynard, Leonardo Collado-Torres, Lukas M We- ber, Cedric Uytingco, Brianna K Barry, Stephen R Williams, Joseph L Catallini, Matthew N Tran, Zachary Besich, Mad- havi Tippani, et al. Transcriptome-scale spatial gene expres- sion in the human dorsolateral prefrontal c...
2021
-
[49]
Multimodal contrastive learning for spatial gene expression prediction using histology images
Wenwen Min, Zhiceng Shi, Jun Zhang, Jun Wan, and Chang- miao Wang. Multimodal contrastive learning for spatial gene expression prediction using histology images. arXiv preprint arXiv:2407.08216, 2024. 2, 4, 5, 14, 17, 18
2024 arXiv
-
[50]
Caclust: linking genotype to transcriptional heterogene- ity of follicular lymphoma using bcr and exomic variants
Kazimierz Oksza-Orzechowski, Edwin Quinten, Shadi Darvish Shafighi, Szymon M Kiełbasa, Hugo van Kessel, Ruben AL de Groen, Joost SP Vermaat, Julieta H Sepl´uveda-Y´a˜nez, Marcelo A Navarrete, Hendrik Veelken, et al. Caclust: linking genotype to transcriptional heterogene- ity ...
2024
-
[51]
Repre- sentation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. 13
2018 arXiv
-
[52]
Maxime Oquab, Timoth´ee Darcet, Theo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicola...
-
[53]
Leveraging infor- mation in spatial transcriptomics to predict super-resolution gene expression from histology images in tumors
Minxing Pang, Kenong Su, and Mingyao Li. Leveraging infor- mation in spatial transcriptomics to predict super-resolution gene expression from histology images in tumors. BioRxiv, pages 2021–11, 2021. 2, 5, 14, 17, 18
2021
-
[54]
Rethinking multiple instance learning for whole slide image classification: A good instance classifier is all you need
Linhao Qu, Yingfan Ma, Xiaoyuan Luo, Qinhao Guo, Man- ning Wang, and Zhijian Song. Rethinking multiple instance learning for whole slide image classification: A good instance classifier is all you need. IEEE Transactions on Circuits and Systems for Video Technology, 2024. 2
2024
-
[55]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PMLR, 2021. 1, 2, 3, 6, 8, 20
2021
-
[56]
Silhouettes: a graphical aid to the in- terpretation and validation of cluster analysis
Peter J Rousseeuw. Silhouettes: a graphical aid to the in- terpretation and validation of cluster analysis. Journal of computational and applied mathematics, 20:53–65, 1987. 8, 13, 15, 20
1987
-
[57]
H-optimus-0, 2024
Charlie Saillard, Rodolphe Jenatton, Felipe Llinares-L´opez, Zelda Mariet, David Cahan´e, Eric Durand, and Jean-Philippe Vert. H-optimus-0, 2024. 1, 2
2024
-
[58]
Nicheformer: a foundation model for single-cell and spatial omics
Anna Christina Schaar, Alejandro Tejada-Lapuerta, Giovanni Palla, Robert Gutgesell, Lennard Halle, Mariia Minaeva, Larsen V ornholz, Leander Dony, Francesca Drummer, Mo- jtaba Bahrami, et al. Nicheformer: a foundation model for single-cell and spatial omics. bioRxiv, pages 202...
2024
-
[59]
Transmil: Transformer based correlated multiple instance learning for whole slide image classification
Zhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang, Jian Zhang, Xiangyang Ji, et al. Transmil: Transformer based correlated multiple instance learning for whole slide image classification. NIPS, 34:2136–2147, 2021. 1
2021
-
[60]
A structure-aware hierarchical graph-based multiple instance learning framework for pt staging in histopathological image
Jiangbo Shi, Lufei Tang, Yang Li, Xianli Zhang, Zeyu Gao, Yefeng Zheng, Chunbao Wang, Tieliang Gong, and Chen Li. A structure-aware hierarchical graph-based multiple instance learning framework for pt staging in histopathological image. IEEE Transactions on Medical Imaging, 42...
-
[61]
Vila-mil: Dual-scale vision-language multiple instance learning for whole slide image classification
Jiangbo Shi, Chen Li, Tieliang Gong, Yefeng Zheng, and Huazhu Fu. Vila-mil: Dual-scale vision-language multiple instance learning for whole slide image classification. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11248–11258, 2024. 6
2024
-
[62]
Multimodal prototyping for cancer survival prediction
Andrew H Song, Richard J Chen, Guillaume Jaume, Anurag Jayant Vaidya, Alexander Baras, and Faisal Mah- mood. Multimodal prototyping for cancer survival prediction. In ICML, 2024. 1
2024
-
[63]
Flex- ible experimental designs for valid single-cell rna-sequencing experiments allowing batch effects correction
Fangda Song, Ga Ming Angus Chan, and Yingying Wei. Flex- ible experimental designs for valid single-cell rna-sequencing experiments allowing batch effects correction. Nature com- munications, 11(1):3274, 2020. 15
2020
-
[64]
Visualization and analysis of gene expression in tis- sue sections by spatial transcriptomics
Patrik L St˚ahl, Fredrik Salm´en, Sanja Vickovic, Anna Lund- mark, Jos ´e Fern ´andez Navarro, Jens Magnusson, Stefania Giacomello, Michaela Asp, Jakub O Westholm, Mikael Huss, 11 et al. Visualization and analysis of gene expression in tis- sue sections by spatial transcriptom...
2016
-
[65]
Visualization and analysis of gene expression in tis- sue sections by spatial transcriptomics
Patrik L St˚ahl, Fredrik Salm´en, Sanja Vickovic, Anna Lund- mark, Jos ´e Fern ´andez Navarro, Jens Magnusson, Stefania Giacomello, Michaela Asp, Jakub O Westholm, Mikael Huss, et al. Visualization and analysis of gene expression in tis- sue sections by spatial transcriptomics...
2016
-
[66]
Trans- fer learning enables predictions in network biology
Christina V Theodoris, Ling Xiao, Anant Chopra, Mark D Chaffin, Zeina R Al Sayed, Matthew C Hill, Helene Mantineo, Elizabeth M Brydon, Zexian Zeng, X Shirley Liu, et al. Trans- fer learning enables predictions in network biology. Nature, 618(7965):616–624, 2023. 2, 3
2023
-
[67]
The expanding vistas of spatial transcriptomics
Luyi Tian, Fei Chen, and Evan Z Macosko. The expanding vistas of spatial transcriptomics. Nature Biotechnology, 41 (6):773–782, 2023. 2
2023
-
[68]
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11),
-
[69]
Vir- chow: a million-slide digital pathology foundation model
Eugene V orontsov, Alican Bozkurt, Adam Casson, George Shaikovski, Michal Zelechowski, Siqi Liu, Kristen Severson, Eric Zimmermann, James Hall, Neil Tenenholtz, et al. Vir- chow: a million-slide digital pathology foundation model. arXiv preprint arXiv:2309.07778, 2023. 4
2023 arXiv
-
[70]
Mgiml: Cancer grading with incomplete radiology-pathology data via memory learning and gradient homogenization
Pengyu Wang, Huaqi Zhang, Meilu Zhu, Xi Jiang, Jing Qin, and Yixuan Yuan. Mgiml: Cancer grading with incomplete radiology-pathology data via memory learning and gradient homogenization. IEEE Transactions on Medical Imaging ,
-
[71]
Transformer-based unsupervised contrastive learning for histopathological image classification
Xiyue Wang, Sen Yang, Jun Zhang, Minghui Wang, Jing Zhang, Wei Yang, Junzhou Huang, and Xiao Han. Transformer-based unsupervised contrastive learning for histopathological image classification. Medical image analy- sis, 81:102559, 2022. 4
2022
-
[72]
A pathology foundation model for cancer diagnosis and prognosis prediction
Xiyue Wang, Junhan Zhao, Eliana Marostica, Wei Yuan, Ji- etian Jin, Jiayu Zhang, Ruijiang Li, Hongping Tang, Kanran Wang, Yu Li, et al. A pathology foundation model for cancer diagnosis and prognosis prediction. Nature, pages 1–9, 2024. 1
2024
-
[73]
Scanpy: large-scale single-cell gene expression data analysis
F Alexander Wolf, Philipp Angerer, and Fabian J Theis. Scanpy: large-scale single-cell gene expression data analysis. Genome biology, 19:1–5, 2018. 16
2018
-
[74]
Cathepsin c promotes breast cancer lung metastasis by modulating neutrophil infiltration and neutrophil extracellular trap formation
Yansen Xiao, Min Cong, Jiatao Li, Dasa He, Qiuyao Wu, Pu Tian, Yuan Wang, Shuaixi Yang, Chenxi Liang, Yajun Liang, et al. Cathepsin c promotes breast cancer lung metastasis by modulating neutrophil infiltration and neutrophil extracellular trap formation. Cancer cell, 39(3):42...
2021
-
[75]
Spatially resolved gene expression prediction from histology images via bi- modal contrastive learning
Ronald Xie, Kuan Pang, Sai Chung, Catia Perciani, Sonya MacParland, Bo Wang, and Gary Bader. Spatially resolved gene expression prediction from histology images via bi- modal contrastive learning. NIPS, 36, 2024. 2, 4, 5, 14, 16, 17, 18
2024
-
[76]
Unsupervised spa- tially embedded deep representation of spatial transcriptomics
Hang Xu, Huazhu Fu, Yahui Long, Kok Siong Ang, Raman Sethi, Kelvin Chong, Mengwei Li, Rom Uddamvathanak, Hong Kai Lee, Jingjing Ling, et al. Unsupervised spa- tially embedded deep representation of spatial transcriptomics. Genome Medicine, 16(1):12, 2024. 5, 20
2024
-
[77]
A whole-slide foundation model for digital pathology from real-world data
Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier Gonz ´alez, Yu Gu, et al. A whole-slide foundation model for digital pathology from real-world data. Nature, pages 1–8, 2024. 1, 4
2024
-
[78]
A multimodal knowledge-enhanced whole-slide pathology foundation model
Yingxue Xu, Yihui Wang, Fengtao Zhou, Jiabo Ma, Shu Yang, Huangjing Lin, Xin Wang, Jiguang Wang, Li Liang, Anjia Han, et al. A multimodal knowledge-enhanced whole-slide pathology foundation model. arXiv preprint arXiv:2407.15362, 2024. 1, 2
2024 arXiv
-
[79]
Stomicsdb: a comprehensive database for spatial transcriptomics data sharing, analysis and visualization
Zhicheng Xu, Weiwen Wang, Tao Yang, Ling Li, Xizheng Ma, Jing Chen, Jieyu Wang, Yan Huang, Joshua Gould, Huifang Lu, et al. Stomicsdb: a comprehensive database for spatial transcriptomics data sharing, analysis and visualization. Nu- cleic acids research, 52(D1):D1053–D1061, 2...
2024
-
[80]
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mo- jtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022. 1, 2
2022 arXiv
-
[81]
Batch alignment of single-cell transcriptomics data using deep metric learning
Xiaokang Yu, Xinyi Xu, Jingxiao Zhang, and Xiangjie Li. Batch alignment of single-cell transcriptomics data using deep metric learning. Nature communications, 14(1):960,
-
[82]
Heartsvg: a fast and accurate method for identifying spatially variable genes in large-scale spatial transcriptomics
Xin Yuan, Yanran Ma, Ruitian Gao, Shuya Cui, Yifan Wang, Botao Fa, Shiyang Ma, Ting Wei, Shuangge Ma, and Zhang- sheng Yu. Heartsvg: a fast and accurate method for identifying spatially variable genes in large-scale spatial transcriptomics. Nature Communications, 15(1):5700, 2024. 13
2024
-
[83]
Sodb facilitates comprehensive exploration of spatial omics data
Zhiyuan Yuan, Wentao Pan, Xuan Zhao, Fangyuan Zhao, Zhimeng Xu, Xiu Li, Yi Zhao, Michael Q Zhang, and Jianhua Yao. Sodb facilitates comprehensive exploration of spatial omics data. Nature Methods, 20(3):387–399, 2023. 3, 18
2023
-
[84]
Spatial transcriptomics prediction from histology jointly through transformer and graph neural networks
Yuansong Zeng, Zhuoyi Wei, Weijiang Yu, Rui Yin, Yuchen Yuan, Bingling Li, Zhonghui Tang, Yutong Lu, and Yue- dong Yang. Spatial transcriptomics prediction from histology jointly through transformer and graph neural networks. Brief- ings in Bioinformatics , 23(5):bbac297, 2022...
2022
-
[85]
Metabolic regulation of cell growth and proliferation
Jiajun Zhu and Craig B Thompson. Metabolic regulation of cell growth and proliferation. Nature reviews Molecular cell biology, 20(7):436–450, 2019. 2 12 A. More Experimental Results A.1. Impact of Unimodal Pre-training The pre-training process of UMPIRE is divided into two sta...
2019
-
[86]
Widespread Social Impact: Technological advancements should benefit a broader population
developing multimodal models capable of robust general- ization across different platforms and technologies. Widespread Social Impact: Technological advancements should benefit a broader population. While molecular-level analyses of cancer significantly enhance diagnostic accu...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.