Pith. sign in

REVIEW 8 cited by

A Multimodal Knowledge-enhanced Whole-slide Pathology Foundation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15362 v3 pith:AXKHLEAI submitted 2024-07-22 cs.CV cs.AI

classification cs.CVcs.AI
keywords pathologywhole-slidecontextpatchpretrainingdatafirstfoundation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Remarkable strides in computational pathology have been made in the task-agnostic foundation model that advances the performance of a wide array of downstream clinical tasks. Despite the promising performance, there are still several challenges. First, prior works have resorted to either vision-only or image-caption data, disregarding pathology reports with more clinically authentic information from pathologists and gene expression profiles which respectively offer distinct knowledge for versatile clinical applications. Second, the current progress in pathology FMs predominantly concentrates on the patch level, where the restricted context of patch-level pretraining fails to capture whole-slide patterns. Even recent slide-level FMs still struggle to provide whole-slide context for patch representation. In this study, for the first time, we develop a pathology foundation model incorporating three levels of modalities: pathology slides, pathology reports, and gene expression data, which resulted in 26,169 slide-level modality pairs from 10,275 patients across 32 cancer types, amounting to over 116 million pathological patch images. To leverage these data for CPath, we propose a novel whole-slide pretraining paradigm that injects the multimodal whole-slide context into the patch representation, called Multimodal Self-TAught PRetraining (mSTAR). The proposed paradigm revolutionizes the pretraining workflow for CPath, enabling the pathology FM to acquire the whole-slide context. To the best of our knowledge, this is the first attempt to incorporate three modalities at the whole-slide context for enhancing pathology FMs. To systematically evaluate the capabilities of mSTAR, we built the largest spectrum of oncological benchmark, spanning 7 categories of oncological applications in 15 types of 97 practical oncological tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Multi-modal Agentic Co-pilot for Evidence Grounded Computational Pathology

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    PathPocket constructs a 4.55M-entity pathology hypergraph from 110k graded documents and deploys a multi-agent framework that outperforms prior systems on 200k cases while raising pathologist accuracy in user studies.

  2. A Unified Low-level Foundation Model for Enhancing Pathology Image Quality

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A prompt-guided diffusion model pretrained on 190 million pathology patches outperforms task-specific models across most restoration and virtual staining benchmarks.

  3. PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis

    cs.CV 2025-08 conditional novelty 6.0 of 10

    PathMR applies pixel-level visual reasoning to pathology, jointly generating diagnostic text and cell-type segmentation masks, and introduces the GADVR gastric adenocarcinoma benchmark.

  4. Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping

    cs.CV 2025-08 conditional novelty 6.0 of 10

    PathPT improves few-shot rare cancer subtyping by using zero-shot vision-language models to create tile-level pseudo-labels and learning prompt tokens with spatial context, outperforming standard MIL baselines when th...

  5. Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Modality bias, an imbalanced attention to text or image during hallucinated outputs, is shown to be mitigated by a training-free attention intervention plus contrastive decoding.

  6. A Versatile Pathology Co-pilot via Reasoning Enhanced Multimodal Large Language Model

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A single pathology vision-language model trained with supervised and reinforcement fine-tuning beats four prior MLLMs across 72 ROI-level and slide-level tasks.

  7. Mitigating Batch Effects in Histopathology via Language-Mediated Robust Embedding Generation

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    GLMP generates robust pathology embeddings by routing histology images through an intermediate textual representation produced by general-purpose MLLMs to mitigate batch effects.

  8. Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    FG-PAN improves zero-shot brain tumor subtype classification by aligning refined visual patch features with LLM-generated fine-grained text prototypes.

Pith tools