Pith. sign in

REVIEW 21 cited by

Phikon-v2, A large and public feature extractor for biomarker prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.09173 v1 pith:PSSPTBC7 submitted 2024-09-13 eess.IV cs.AIcs.CV

Phikon-v2, A large and public feature extractor for biomarker prediction

classification eess.IV cs.AIcs.CV
keywords modelphikon-v2datahistologymodelspredictionpubliclytasks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Gathering histopathology slides from over 100 publicly available cohorts, we compile a diverse dataset of 460 million pathology tiles covering more than 30 cancer sites. Using this dataset, we train a large self-supervised vision transformer using DINOv2 and publicly release one iteration of this model for further experimentation, coined Phikon-v2. While trained on publicly available histology slides, Phikon-v2 surpasses our previously released model (Phikon) and performs on par with other histopathology foundation models (FM) trained on proprietary data. Our benchmarks include eight slide-level tasks with results reported on external validation cohorts avoiding any data contamination between pre-training and evaluation datasets. Our downstream training procedure follows a simple yet robust ensembling strategy yielding a +1.75 AUC increase across tasks and models compared to one-shot retraining (p<0.001). We compare Phikon (ViT-B) and Phikon-v2 (ViT-L) against 14 different histology feature extractors, making our evaluation the most comprehensive to date. Our result support evidences that DINOv2 handles joint model and data scaling better than iBOT. Also, we show that recent scaling efforts are overall beneficial to downstream performance in the context of biomarker prediction with GigaPath and H-Optimus-0 (two ViT-g with 1.1B parameters each) standing out. However, the statistical margins between the latest top-performing FMs remain mostly non-significant; some even underperform on specific indications or tasks such as MSI prediction - deposed by a 13x smaller model developed internally. While latest foundation models may exhibit limitations for clinical deployment, they nonetheless offer excellent grounds for the development of more specialized and cost-efficient histology encoders fueling AI-guided diagnostic tools.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Self-supervision drives representational convergence in medical foundation models more than clinical supervision

    cs.CV 2026-07 conditional novelty 7.0

    Representational convergence among medical image encoders is modest, driven mainly by the self-supervised pretraining objective rather than clinical supervision or scale, yet still sufficient for cross-encoder and cro...

  2. Benchmarking Pathology Foundation Models for Spatial Domain Understanding

    cs.CV 2026-05 unverdicted novelty 7.0

    SpaPath-Bench evaluates spatial representation in 19 pathology foundation models via spatial domain identification on 42 paired WSI-ST slides using three agreement criteria across 83K runs.

  3. Topology-Driven Transferability Estimation for 3D Medical Vision Foundation Models

    cs.CV 2026-07 conditional novelty 6.5

    MST-based local boundary leakage and global topology divergence, fused by task complexity, rank SSL 3D medical encoders for segmentation without fine-tuning, beating prior TE metrics by 0.36 weighted Kendall τ at 56× speed.

  4. From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

    cs.CV 2026-08 conditional novelty 6.0

    MRPT, a multi-resolution hierarchical transformer pre-trained on 36K whole-slide images, is reported to outperform prior pathology foundation models on 34 classification, captioning, and VQA datasets.

  5. Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models

    cs.CV 2026-07 conditional novelty 6.0

    CRoMa scores each pathology image embedding by the margin between cross-site biological matches and same-site biological distractors, revealing distributional lower tails that pooled robustness scores hide.

  6. CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders

    cs.LG 2026-07 conditional novelty 6.0

    A chance-calibrated discordance measure shows frozen encoders encode fine-grained clinical findings weakly rather than blindly, with collapse rates ranging from 4.5% (bird species) to 50.0% (glaucoma).

  7. LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models

    cs.CV 2026-07 conditional novelty 6.0

    Language-guided adaptive multi-teacher distillation yields an 87M pathology encoder that matches or exceeds GigaPath and UNI on WSI captioning, VQA, and MIL tasks.

  8. Towards Autonomous and Auditable Medical Imaging Model Development

    cs.CV 2026-07 conditional novelty 6.0

    AMID, a verification-guided multi-agent MLE system for medical imaging, outperforms general MLE agents on 20 ReX-MLE challenges and approaches human challenge solutions on several tasks.

  9. DaX: Learning General Pathology Representations Across Scales

    eess.IV 2026-06 unverdicted novelty 6.0

    DaX is a pathology vision foundation model that extends DINOv3 with continuous magnification training and cross-scale consistency, achieving top average performance on a benchmark of 161 tasks from 44 datasets coverin...

  10. MOOZY: A Patient-First Foundation Model for Computational Pathology

    cs.CV 2026-03 conditional novelty 6.0

    Patient-level pretraining with a case transformer and multi-task public supervision yields transferable WSI embeddings that beat larger slide-centric models on held-out pathology tasks.

  11. Enabling clinical use of foundation models for computational pathology

    cs.CV 2026-02 conditional novelty 6.0

    Novel robustness losses added during downstream training on foundation-model features from pathology slides improve both robustness to technical variation and classification accuracy.

  12. HistoFID- Calibrating Frechet-distance evaluation across pathology foundation models

    eess.IV 2026-07 conditional novelty 5.0

    Raw Fréchet distances in pathology vary ~30-fold across encoders; dividing by each encoder's own within-cohort floor cuts cross-encoder variation by ~89% within and ~58% across cohorts.

  13. APRIL-MedSeg: A Modular Medical Image Segmentation Toolbox Embracing Modern Paradigms

    cs.CV 2026-06 unverdicted novelty 5.0

    APRIL-MedSeg is a new open-source modular toolbox that uses YAML configuration and component registries to unify multiple advanced paradigms for medical image segmentation.

  14. Mitigating Batch Effects in Histopathology via Language-Mediated Robust Embedding Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    GLMP generates robust pathology embeddings by routing histology images through an intermediate textual representation produced by general-purpose MLLMs to mitigate batch effects.

  15. Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction

    cs.CV 2026-05 unverdicted novelty 5.0

    A masked-diffusion pretrained convolutional model outperforms ViT pathology foundation models on cell-level dense prediction tasks in histology.

  16. Federated Distillation for Whole Slide Image via Gaussian-Mixture Feature Alignment and Curriculum Integration

    cs.CV 2026-05 unverdicted novelty 5.0

    FedHD is a federated learning framework for whole slide images that distills one-to-one synthetic features aligned via Gaussian mixtures and progressively integrates cross-site features through curriculum learning to ...

  17. Federated Distillation for Whole Slide Image via Gaussian-Mixture Feature Alignment and Curriculum Integration

    cs.CV 2026-05 unverdicted novelty 5.0

    FedHD performs federated distillation for whole slide images by generating one synthetic feature set per real slide via Gaussian-mixture alignment and adding them via curriculum integration, outperforming prior federa...

  18. Atlas 2 -- Foundation models for clinical deployment

    cs.CV 2026-01 conditional novelty 5.0

    Atlas 2 and its distilled variants set new average state-of-the-art results across 80 pathology benchmarks, with larger robustness margins over prior models.

  19. CellPrior-Net: Prior-Guided Nuclei Detection and Classification for H&E Whole-Slide Images

    cs.MM 2026-07 unverdicted novelty 4.0

    CellPrior-Net integrates hematoxylin channel prior into a lightweight CNN for nuclei detection and classification in H&E WSIs, claiming comparable accuracy to SOTA with significantly reduced inference time across 10.4...

  20. APRIL-MedSeg: A Modular Medical Image Segmentation Toolbox Embracing Modern Paradigms

    cs.CV 2026-06 unverdicted novelty 4.0

    Presents APRIL-MedSeg, a modular YAML-configurable toolbox for 2D medical image segmentation integrating semi-supervised, domain adaptation, distillation, weakly supervised, text-guided, and foundation model paradigms...

  21. Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions

    cs.AI 2026-07 conditional novelty 3.5

    AIIE is framed as an AI-dominated innovation ecosystem whose participant mix, coopetition, and goals shift by lifecycle stage, illustrated with Owkin and four open challenges.