Pith. sign in

REVIEW 2 cited by

Label Critic: Design Data Before Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.02753 v1 pith:D4VJCBEI submitted 2024-11-05 cs.CV

classification cs.CV
keywords labellabelsannotationsradiologistswhenbest-aibodycomparison
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As medical datasets rapidly expand, creating detailed annotations of different body structures becomes increasingly expensive and time-consuming. We consider that requesting radiologists to create detailed annotations is unnecessarily burdensome and that pre-existing AI models can largely automate this process. Following the spirit don't use a sledgehammer on a nut, we find that, rather than creating annotations from scratch, radiologists only have to review and edit errors if the Best-AI Labels have mistakes. To obtain the Best-AI Labels among multiple AI Labels, we developed an automatic tool, called Label Critic, that can assess label quality through tireless pairwise comparisons. Extensive experiments demonstrate that, when incorporated with our developed Image-Prompt pairs, pre-existing Large Vision-Language Models (LVLM), trained on natural images and texts, achieve 96.5% accuracy when choosing the best label in a pair-wise comparison, without extra fine-tuning. By transforming the manual annotation task (30-60 min/scan) into an automatic comparison task (15 sec/scan), we effectively reduce the manual efforts required from radiologists by an order of magnitude. When the Best-AI Labels are sufficiently accurate (81% depending on body structures), they will be directly adopted as the gold-standard annotations for the dataset, with lower-quality AI Labels automatically discarded. Label Critic can also check the label quality of a single AI Label with 71.8% accuracy when no alternatives are available for comparison, prompting radiologists to review and edit if the estimated quality is low (19% depending on body structures).

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PanTS: The Pancreatic Tumor Segmentation Dataset

    eess.IV 2025-07 conditional novelty 6.0 of 10

    PanTS is a new large CT dataset with expert-drawn pancreatic tumor and anatomy labels, and models trained on it beat prior public benchmarks.

  2. ShapeKit

    eess.IV 2025-06 reject novelty 5.0 of 10

    ShapeKit, a rule-based post-processing toolkit, reports Dice score improvements of up to 8.8 percentage points on two CT datasets without retraining the segmentation model.

Pith tools