REVIEW 10 cited by
HISTAI: An Open-Source, Large-Scale Whole Slide Image Dataset for Computational Pathology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advancements in Digital Pathology (DP), particularly through artificial intelligence and Foundation Models, have underscored the importance of large-scale, diverse, and richly annotated datasets. Despite their critical role, publicly available Whole Slide Image (WSI) datasets often lack sufficient scale, tissue diversity, and comprehensive clinical metadata, limiting the robustness and generalizability of AI models. In response, we introduce the HISTAI dataset, a large, multimodal, open-access WSI collection comprising over 60,000 slides from various tissue types. Each case in the HISTAI dataset is accompanied by extensive clinical metadata, including diagnosis, demographic information, detailed pathological annotations, and standardized diagnostic coding. The dataset aims to fill gaps identified in existing resources, promoting innovation, reproducibility, and the development of clinically relevant computational pathology solutions. The dataset can be accessed at https://github.com/HistAI/HISTAI.
Forward citations
Cited by 10 Pith papers
-
A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation
A lung-specific pathology foundation model built on Virchow2 achieved broad AUC gains across 32 tasks, 92.3% average AUC in a prospective study, and improved pathologist accuracy by 7.9 percentage points in a crossover RCT.
-
A Generative Foundation Model for Multimodal Histopathology
MuPD is a pretrained generative foundation model using a diffusion transformer with cross-modal attention that synthesizes histopathology images from text or RNA data and outperforms task-specific models on generation...
-
Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling
A new open pipeline and dataset enable training of a vision-language model for whole-slide pathology VQA that outperforms MedGemma on tissue identification, neoplasm detection, and differential diagnosis.
-
ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts
Multi-stage agglomerative distillation consolidates eight vision, vision-language, and slide-level pathology teachers into one backbone that ranks first on average across 96 downstream tasks.
-
DaX: Learning General Pathology Representations Across Scales
DaX is a pathology vision foundation model that extends DINOv3 with continuous magnification training and cross-scale consistency, achieving top average performance on a benchmark of 161 tasks from 44 datasets coverin...
-
A Pathology Foundation Model for Gastric Cancer with Real-World Validation
GRACE, a gastric-specific pathology foundation model trained on multicenter HE-stained slides, outperforms pancancer models on 28 tasks and improves pathologist accuracy, speed, and agreement in a reader study while e...
-
A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation
PulmoFoundation achieves 92.3% average AUC on 32 lung pathology tasks in prospective validation and raises pathologist accuracy from 83.8% to 91.7% in a crossover RCT.
-
A Breast Vision Pathology Foundation Model for Real-world Clinical Utility
BRAVE foundation model excludes 70-77% of negative breast biopsy and frozen-section cases with NPV above 0.95 while improving balanced accuracy from 88.5% to 95.1% in reader studies and predicting survival outcomes.
-
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
A masked-diffusion pretrained convolutional model outperforms ViT pathology foundation models on cell-level dense prediction tasks in histology.
-
Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation
A token-efficient VLM with frozen encoder, two-layer MLP aligner, and LLM decoder generates case-level synoptic pathology reports from multi-WSI inputs using 5x magnification patches and two-stage supervised training.
Discussion (0). Sign in to comment.