REVIEW 9 cited by
PathBench: A comprehensive comparison benchmark for pathology foundation models towards precision oncology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The emergence of pathology foundation models has revolutionized computational histopathology, enabling highly accurate, generalized whole-slide image analysis for improved cancer diagnosis, and prognosis assessment. While these models show remarkable potential across cancer diagnostics and prognostics, their clinical translation faces critical challenges including variability in optimal model across cancer types, potential data leakage in evaluation, and lack of standardized benchmarks. Without rigorous, unbiased evaluation, even the most advanced PFMs risk remaining confined to research settings, delaying their life-saving applications. Existing benchmarking efforts remain limited by narrow cancer-type focus, potential pretraining data overlaps, or incomplete task coverage. We present PathBench, the first comprehensive benchmark addressing these gaps through: multi-center in-hourse datasets spanning common cancers with rigorous leakage prevention, evaluation across the full clinical spectrum from diagnosis to prognosis, and an automated leaderboard system for continuous model assessment. Our framework incorporates large-scale data, enabling objective comparison of PFMs while reflecting real-world clinical complexity. All evaluation data comes from private medical providers, with strict exclusion of any pretraining usage to avoid data leakage risks. We have collected 15,888 WSIs from 8,549 patients across 10 hospitals, encompassing over 64 diagnosis and prognosis tasks. Currently, our evaluation of 19 PFMs shows that Virchow2 and H-Optimus-1 are the most effective models overall. This work provides researchers with a robust platform for model development and offers clinicians actionable insights into PFM performance across diverse clinical scenarios, ultimately accelerating the translation of these transformative technologies into routine pathology practice.
Forward citations
Cited by 9 Pith papers
-
A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation
A lung-specific pathology foundation model built on Virchow2 achieved broad AUC gains across 32 tasks, 92.3% average AUC in a prospective study, and improved pathologist accuracy by 7.9 percentage points in a crossover RCT.
-
Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models
LogitProd fuses logits from heterogeneous pathology foundation models via sample-adaptive weights, ranking first on 20 of 22 benchmarks with a 3% average gain over the best single model and 12x lower training cost tha...
-
Democratizing and accelerating AI-driven pathology research through agentic intelligence
PathLab is an agentic framework that translates natural-language objectives into validated computational pathology workflows, achieving non-inferior performance on 12 datasets across four task families while enabling ...
-
A Pathology Foundation Model for Gastric Cancer with Real-World Validation
GRACE, a gastric-specific pathology foundation model trained on multicenter HE-stained slides, outperforms pancancer models on 28 tasks and improves pathologist accuracy, speed, and agreement in a reader study while e...
-
Spatial Transcriptomics-Guided Alignment Enhances Molecular Profiling in Pathology Foundation Model
STAMP uses a curated 1.8M-pair spatial transcriptomics atlas and pathway-informed alignment to augment pathology foundation models for molecular phenotype inference from H&E WSIs.
-
A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation
PulmoFoundation achieves 92.3% average AUC on 32 lung pathology tasks in prospective validation and raises pathologist accuracy from 83.8% to 91.7% in a crossover RCT.
-
A Unified Low-level Foundation Model for Enhancing Pathology Image Quality
A prompt-guided diffusion model pretrained on 190 million pathology patches outperforms task-specific models across most restoration and virtual staining benchmarks.
-
BRIGHT: A Collaborative Generalist-Specialist Foundation Model for Breast Pathology
Fine-tuning a generalist pathology model on 51,000 breast WSIs and concatenating its features with the original model yields top-1 performance on 21 of 24 internal breast-pathology tasks, but only 5 of 10 external tasks.
-
Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction
H-optimus-1 achieves the strongest externally validated survival prediction from histopathology images, with second-generation PFMs outperforming first-generation counterparts and a compact distilled model offering ef...
Discussion (0). Sign in to comment.