REVIEW 19 cited by
PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology
read the original abstract
Foundation models in computational pathology promise to unlock the development of new clinical decision support systems and models for precision medicine. However, there is a mismatch between most clinical analysis, which is defined at the level of one or more whole slide images, and foundation models to date, which process the thousands of image tiles contained in a whole slide image separately. The requirement to train a network to aggregate information across a large number of tiles in multiple whole slide images limits these models' impact. In this work, we present a slide-level foundation model for H&E-stained histopathology, PRISM, that builds on Virchow tile embeddings and leverages clinical report text for pre-training. Using the tile embeddings, PRISM produces slide-level embeddings with the ability to generate clinical reports, resulting in several modes of use. Using text prompts, PRISM achieves zero-shot cancer detection and sub-typing performance approaching and surpassing that of a supervised aggregator model. Using the slide embeddings with linear classifiers, PRISM surpasses supervised aggregator models. Furthermore, we demonstrate that fine-tuning of the PRISM slide encoder yields label-efficient training for biomarker prediction, a task that typically suffers from low availability of training data; an aggregator initialized with PRISM and trained on as little as 10% of the training data can outperform a supervised baseline that uses all of the data.
Forward citations
Cited by 19 Pith papers
-
Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling
A new open pipeline and dataset enable training of a vision-language model for whole-slide pathology VQA that outperforms MedGemma on tissue identification, neoplasm detection, and differential diagnosis.
-
Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models
CRoMa scores each pathology image embedding by the margin between cross-site biological matches and same-site biological distractors, revealing distributional lower tails that pooled robustness scores hide.
-
Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models
Distilling TITAN and CARE slide embeddings into MIL aggregators gives reusable pretrained weights that beat from-scratch training on most of 15 pathology tasks, with the largest gains in few-shot and linear-probing settings.
-
How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology
Systematic factorial analysis shows optimized LLM input configurations for pathology WSIs raise GPT-5 performance from 15.1% to 39.5% on TCGA cancer classification and 38.1% to 62.9% on GTEx organ classification, with...
-
CRISP -- Clustering-Based Redundancy-Reduced Instance Sampling for Pathology Case Representation and Retrieval
CRISP is a clustering-based sampling framework that builds case-level representations from multiple whole-slide images for improved pathology retrieval, matching or exceeding single-slide selection on two breast cance...
-
Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning
PathCTM adaptively zooms from low to high magnification with attention-guided pruning and confidence-based early stopping, cutting WSI inference time by ~95% while maintaining or improving AUC.
-
MOOZY: A Patient-First Foundation Model for Computational Pathology
Patient-level pretraining with a case transformer and multi-task public supervision yields transferable WSI embeddings that beat larger slide-centric models on held-out pathology tasks.
-
GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
Distilled, Apache-2.0-licensed GigaPath-Flash and GigaTIME-Flash models deliver most of the original models' accuracy at a fraction of the compute and memory.
-
Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning
DICE ensembles frozen pathology foundation models, aligns them with deep mutual learning to make disagreement a reliable uncertainty proxy, and shows consensus-based localization on WSI tasks.
-
AGE-MIL: Anchor-Guided Evidence Learning for Patient-Level Prediction
AGE-MIL is a new MIL framework that uses patient-level anchors to direct patch selection and evidence accumulation for stable patient-level predictions on six pathology tasks, outperforming eight prior MIL methods.
-
Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning
PathCTM uses adaptive continuous reasoning across scales to reduce patch processing in whole slide images by over 95% while preserving diagnostic AUC.
-
Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning
PathCTM reduces processed patches in WSI analysis by 95.95% and inference time by 95.62% via adaptive scale-space reasoning while preserving AUC.
-
Atlas 2 -- Foundation models for clinical deployment
Atlas 2 and its distilled variants set new average state-of-the-art results across 80 pathology benchmarks, with larger robustness margins over prior models.
-
Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis
Foundation model representations from images and transcriptomics carry complementary signals for cancer classification; multimodal fusion improves results mainly when no modality dominates, and conformal prediction re...
-
Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation
A token-efficient VLM with frozen encoder, two-layer MLP aligner, and LLM decoder generates case-level synoptic pathology reports from multi-WSI inputs using 5x magnification patches and two-stage supervised training.
-
Genetically Aligned Patient Representations Improve Hematological Diagnosis
A two-stage training method using self-supervised pretraining on cell images followed by contrastive alignment with genetic data creates improved patient encoders for hematological diagnosis.
-
Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data
Benchmarking on TCGA shows TITAN foundation model edges out others for whole-slide retrieval but with only ~68% average accuracy, high organ-to-organ variation, and no consistent winner over patch-level baselines.
-
From Classical Machine Learning to Emerging Foundation Models: Review on Multimodal Data Integration for Cancer Research
A review mapping the transition from classical machine learning to foundation models for multimodal data integration in cancer research.
-
Data-Centric Foundation Models in Computational Healthcare: A Survey
The paper surveys data-centric strategies for foundation models in computational healthcare and supplies a curated list of related models and datasets.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.