Pith. sign in

REVIEW 10 cited by

HISTAI: An Open-Source, Large-Scale Whole Slide Image Dataset for Computational Pathology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.12120 v1 pith:63XPPR2L submitted 2025-05-17 eess.IV cs.CV

classification eess.IVcs.CV
keywords datasethistaipathologyclinicalcomputationaldatasetsimagelarge-scale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in Digital Pathology (DP), particularly through artificial intelligence and Foundation Models, have underscored the importance of large-scale, diverse, and richly annotated datasets. Despite their critical role, publicly available Whole Slide Image (WSI) datasets often lack sufficient scale, tissue diversity, and comprehensive clinical metadata, limiting the robustness and generalizability of AI models. In response, we introduce the HISTAI dataset, a large, multimodal, open-access WSI collection comprising over 60,000 slides from various tissue types. Each case in the HISTAI dataset is accompanied by extensive clinical metadata, including diagnosis, demographic information, detailed pathological annotations, and standardized diagnostic coding. The dataset aims to fill gaps identified in existing resources, promoting innovation, reproducibility, and the development of clinically relevant computational pathology solutions. The dataset can be accessed at https://github.com/HistAI/HISTAI.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation

    eess.IV 2026-05 conditional novelty 7.0 of 10

    A lung-specific pathology foundation model built on Virchow2 achieved broad AUC gains across 32 tasks, 92.3% average AUC in a prospective study, and improved pathologist accuracy by 7.9 percentage points in a crossover RCT.

  2. A Generative Foundation Model for Multimodal Histopathology

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    MuPD is a pretrained generative foundation model using a diffusion transformer with cross-modal attention that synthesizes histopathology images from text or RNA data and outperforms task-specific models on generation...

  3. Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling

    cs.CV 2025-12 conditional novelty 7.0 of 10

    A new open pipeline and dataset enable training of a vision-language model for whole-slide pathology VQA that outperforms MedGemma on tissue identification, neoplasm detection, and differential diagnosis.

  4. ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts

    cs.CV 2026-07 accept novelty 6.0 of 10

    Multi-stage agglomerative distillation consolidates eight vision, vision-language, and slide-level pathology teachers into one backbone that ranks first on average across 96 downstream tasks.

  5. DaX: Learning General Pathology Representations Across Scales

    eess.IV 2026-06 unverdicted novelty 6.0 of 10

    DaX is a pathology vision foundation model that extends DINOv3 with continuous magnification training and cross-scale consistency, achieving top average performance on a benchmark of 161 tasks from 44 datasets coverin...

  6. A Pathology Foundation Model for Gastric Cancer with Real-World Validation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    GRACE, a gastric-specific pathology foundation model trained on multicenter HE-stained slides, outperforms pancancer models on 28 tasks and improves pathologist accuracy, speed, and agreement in a reader study while e...

  7. A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation

    eess.IV 2026-05 unverdicted novelty 6.0 of 10

    PulmoFoundation achieves 92.3% average AUC on 32 lung pathology tasks in prospective validation and raises pathologist accuracy from 83.8% to 91.7% in a crossover RCT.

  8. A Breast Vision Pathology Foundation Model for Real-world Clinical Utility

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    BRAVE foundation model excludes 70-77% of negative breast biopsy and frozen-section cases with NPV above 0.95 while improving balanced accuracy from 88.5% to 95.1% in reader studies and predicting survival outcomes.

  9. Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A masked-diffusion pretrained convolutional model outperforms ViT pathology foundation models on cell-level dense prediction tasks in histology.

  10. Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    A token-efficient VLM with frozen encoder, two-layer MLP aligner, and LLM decoder generates case-level synoptic pathology reports from multi-WSI inputs using 5x magnification patches and two-stage supervised training.

Pith tools