Pith. sign in

REVIEW 2 cited by

Quality Sentinel: Estimating Label Quality and Errors in Medical Segmentation Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.00327 v1 pith:VKXSZNOT submitted 2024-06-01 cs.CV

classification cs.CV
keywords qualitylabeldatasetssentinelannotationsmanualmedicalmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An increasing number of public datasets have shown a transformative impact on automated medical segmentation. However, these datasets are often with varying label quality, ranging from manual expert annotations to AI-generated pseudo-annotations. There is no systematic, reliable, and automatic quality control (QC). To fill in this bridge, we introduce a regression model, Quality Sentinel, to estimate label quality compared with manual annotations in medical segmentation datasets. This regression model was trained on over 4 million image-label pairs created by us. Each pair presents a varying but quantified label quality based on manual annotations, which enable us to predict the label quality of any image-label pairs in the inference. Our Quality Sentinel can predict the label quality of 142 body structures. The predicted label quality quantified by Dice Similarity Coefficient (DSC) shares a strong correlation with ground truth quality, with a positive correlation coefficient (r=0.902). Quality Sentinel has found multiple impactful use cases. (I) We evaluated label quality in publicly available datasets, where quality highly varies across different datasets. Our analysis also uncovers that male and younger subjects exhibit significantly higher quality. (II) We identified and corrected poorly annotated labels, achieving 1/3 reduction in annotation costs with optimal budgeting on TotalSegmentator. (III) We enhanced AI training efficiency and performance by focusing on high-quality pseudo labels, resulting in a 33%--88% performance boost over entropy-based methods, with a cost of 31% time and 4.5% memory. The data and model are released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HyperSORT: Self-Organising Robust Training with hyper-networks

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A hyper-network that predicts segmentation UNet weights from per-sample learned latent vectors yields a structured map of annotation styles and a way to flag erroneous labels.

  2. Beyond Pixel Agreement: Large Language Models as Clinical Guardrails for Reliable Medical Image Segmentation

    eess.IV 2025-06 conditional novelty 6.0 of 10

    A structured multi-stage LLM prompt can classify medical segmentation quality zero-shot, with accuracy comparable to trained vision models on a small test set.

Pith tools