Pith. sign in

REVIEW 19 cited by

PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.10254 v2 pith:XOEPVTU4 submitted 2024-05-16 eess.IV cs.CVcs.LG

PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology

classification eess.IV cs.CVcs.LG
keywords prismmodelsslideclinicalembeddingsfoundationaggregatordata
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Foundation models in computational pathology promise to unlock the development of new clinical decision support systems and models for precision medicine. However, there is a mismatch between most clinical analysis, which is defined at the level of one or more whole slide images, and foundation models to date, which process the thousands of image tiles contained in a whole slide image separately. The requirement to train a network to aggregate information across a large number of tiles in multiple whole slide images limits these models' impact. In this work, we present a slide-level foundation model for H&E-stained histopathology, PRISM, that builds on Virchow tile embeddings and leverages clinical report text for pre-training. Using the tile embeddings, PRISM produces slide-level embeddings with the ability to generate clinical reports, resulting in several modes of use. Using text prompts, PRISM achieves zero-shot cancer detection and sub-typing performance approaching and surpassing that of a supervised aggregator model. Using the slide embeddings with linear classifiers, PRISM surpasses supervised aggregator models. Furthermore, we demonstrate that fine-tuning of the PRISM slide encoder yields label-efficient training for biomarker prediction, a task that typically suffers from low availability of training data; an aggregator initialized with PRISM and trained on as little as 10% of the training data can outperform a supervised baseline that uses all of the data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling

    cs.CV 2025-12 conditional novelty 7.0

    A new open pipeline and dataset enable training of a vision-language model for whole-slide pathology VQA that outperforms MedGemma on tissue identification, neoplasm detection, and differential diagnosis.

  2. Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models

    cs.CV 2026-07 conditional novelty 6.0

    CRoMa scores each pathology image embedding by the margin between cross-site biological matches and same-site biological distractors, revealing distributional lower tails that pooled robustness scores hide.

  3. Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models

    cs.CV 2026-07 conditional novelty 6.0

    Distilling TITAN and CARE slide embeddings into MIL aggregators gives reusable pretrained weights that beat from-scratch training on most of 15 pathology tasks, with the largest gains in few-shot and linear-probing settings.

  4. How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology

    cs.CV 2026-06 unverdicted novelty 6.0

    Systematic factorial analysis shows optimized LLM input configurations for pathology WSIs raise GPT-5 performance from 15.1% to 39.5% on TCGA cancer classification and 38.1% to 62.9% on GTEx organ classification, with...

  5. CRISP -- Clustering-Based Redundancy-Reduced Instance Sampling for Pathology Case Representation and Retrieval

    cs.CV 2026-05 unverdicted novelty 6.0

    CRISP is a clustering-based sampling framework that builds case-level representations from multiple whole-slide images for improved pathology retrieval, matching or exceeding single-slide selection on two breast cance...

  6. Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning

    cs.CV 2026-05 conditional novelty 6.0

    PathCTM adaptively zooms from low to high magnification with attention-guided pruning and confidence-based early stopping, cutting WSI inference time by ~95% while maintaining or improving AUC.

  7. MOOZY: A Patient-First Foundation Model for Computational Pathology

    cs.CV 2026-03 conditional novelty 6.0

    Patient-level pretraining with a case transformer and multi-task public supervision yields transferable WSI embeddings that beat larger slide-centric models on held-out pathology tasks.

  8. GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

    cs.CV 2026-07 conditional novelty 5.0

    Distilled, Apache-2.0-licensed GigaPath-Flash and GigaTIME-Flash models deliver most of the original models' accuracy at a fraction of the compute and memory.

  9. Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning

    cs.CV 2026-06 unverdicted novelty 5.0

    DICE ensembles frozen pathology foundation models, aligns them with deep mutual learning to make disagreement a reliable uncertainty proxy, and shows consensus-based localization on WSI tasks.

  10. AGE-MIL: Anchor-Guided Evidence Learning for Patient-Level Prediction

    cs.CV 2026-06 unverdicted novelty 5.0

    AGE-MIL is a new MIL framework that uses patient-level anchors to direct patch selection and evidence accumulation for stable patient-level predictions on six pathology tasks, outperforming eight prior MIL methods.

  11. Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning

    cs.CV 2026-05 unverdicted novelty 5.0

    PathCTM uses adaptive continuous reasoning across scales to reduce patch processing in whole slide images by over 95% while preserving diagnostic AUC.

  12. Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning

    cs.CV 2026-05 unverdicted novelty 5.0

    PathCTM reduces processed patches in WSI analysis by 95.95% and inference time by 95.62% via adaptive scale-space reasoning while preserving AUC.

  13. Atlas 2 -- Foundation models for clinical deployment

    cs.CV 2026-01 conditional novelty 5.0

    Atlas 2 and its distilled variants set new average state-of-the-art results across 80 pathology benchmarks, with larger robustness margins over prior models.

  14. Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

    cs.LG 2026-06 unverdicted novelty 4.0

    Foundation model representations from images and transcriptomics carry complementary signals for cancer classification; multimodal fusion improves results mainly when no modality dominates, and conformal prediction re...

  15. Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation

    cs.CV 2026-05 unverdicted novelty 4.0

    A token-efficient VLM with frozen encoder, two-layer MLP aligner, and LLM decoder generates case-level synoptic pathology reports from multi-WSI inputs using 5x magnification patches and two-stage supervised training.

  16. Genetically Aligned Patient Representations Improve Hematological Diagnosis

    cs.CV 2026-05 unverdicted novelty 4.0

    A two-stage training method using self-supervised pretraining on cell images followed by contrastive alignment with genetic data creates improved patient encoders for hematological diagnosis.

  17. Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data

    cs.CV 2026-04 unverdicted novelty 4.0

    Benchmarking on TCGA shows TITAN foundation model edges out others for whole-slide retrieval but with only ~68% average accuracy, high organ-to-organ variation, and no consistent winner over patch-level baselines.

  18. From Classical Machine Learning to Emerging Foundation Models: Review on Multimodal Data Integration for Cancer Research

    q-bio.QM 2025-07 unverdicted novelty 3.0

    A review mapping the transition from classical machine learning to foundation models for multimodal data integration in cancer research.

  19. Data-Centric Foundation Models in Computational Healthcare: A Survey

    cs.LG 2024-01 unverdicted novelty 3.0

    The paper surveys data-centric strategies for foundation models in computational healthcare and supplies a curated list of related models and datasets.