REVIEW 5 cited by
Training state-of-the-art pathology foundation models with orders of magnitude less data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Training state-of-the-art pathology foundation models with orders of magnitude less data
read the original abstract
The field of computational pathology has recently seen rapid advances driven by the development of modern vision foundation models (FMs), typically trained on vast collections of pathology images. Recent studies demonstrate that increasing the training data set and model size and integrating domain-specific image processing techniques can significantly enhance the model's performance on downstream tasks. Building on these insights, our work incorporates several recent modifications to the standard DINOv2 framework from the literature to optimize the training of pathology FMs. We also apply a post-training procedure for fine-tuning models on higher-resolution images to further enrich the information encoded in the embeddings. We present three novel pathology FMs trained on up to two orders of magnitude fewer WSIs than those used to train other state-of-the-art FMs while demonstrating a comparable or superior performance on downstream tasks. Even the model trained on TCGA alone (12k WSIs) outperforms most existing FMs and, on average, matches Virchow2, the second-best FM published to date. This suggests that there still remains a significant potential for further improving the models and algorithms used to train pathology FMs to take full advantage of the vast data collections.
Forward citations
Cited by 5 Pith papers
-
Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models
CRoMa scores each pathology image embedding by the margin between cross-site biological matches and same-site biological distractors, revealing distributional lower tails that pooled robustness scores hide.
-
When Are Multimodal Predictions Biologically Supported? A Diagnostic Evaluation Framework
DECAT classifies multimodal representations into four diagnostic scenarios using null-referenced metrics and a rule-based procedure to detect shared biology versus confounders without knowing the confounder identity.
-
PC-MIL: Decoupling Feature Resolution from Supervision Scale in Whole-Slide Learning
PC-MIL shows that anchoring supervision at a 2 mm scale and progressively mixing slide- and region-level labels improves cross-context accuracy in WSI cancer detection without reducing global performance.
-
MOOZY: A Patient-First Foundation Model for Computational Pathology
Patient-level pretraining with a case transformer and multi-task public supervision yields transferable WSI embeddings that beat larger slide-centric models on held-out pathology tasks.
-
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
A masked-diffusion pretrained convolutional model outperforms ViT pathology foundation models on cell-level dense prediction tasks in histology.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.