Pith. sign in

REVIEW 2 cited by

Revisiting 3D Medical Scribble Supervision: Benchmarking Beyond Cardiac Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12834 v2 pith:DOFD5N5E submitted 2024-03-19 cs.CV

classification cs.CV
keywords medicalscribblesegmentationsupervisionacrosscardiacmethodsperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scribble supervision has emerged as a promising approach for reducing annotation costs in medical 3D segmentation by leveraging sparse annotations instead of voxel-wise labels. While existing methods report strong performance, a closer analysis reveals that the majority of research is confined to the cardiac domain, predominantly using ACDC and MSCMR datasets. This over-specialization has resulted in severe overfitting, misleading claims of performance improvements, and a lack of generalization across broader segmentation tasks. In this work, we formulate a set of key requirements for practical scribble supervision and introduce ScribbleBench, a comprehensive benchmark spanning over seven diverse medical imaging datasets, to systematically evaluate the fulfillment of these requirements. Consequently, we uncover a general failure of methods to generalize across tasks and that many widely used novelties degrade performance outside of the cardiac domain, whereas simpler overlooked approaches achieve superior generalization. Finally, we raise awareness for a strong yet overlooked baseline, nnU-Net coupled with a partial loss, which consistently outperforms specialized methods across a diverse range of tasks. By identifying fundamental limitations in existing research and establishing a new benchmark-driven evaluation standard, this work aims to steer scribble supervision toward more practical, robust, and generalizable methodologies for medical image segmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiverSeg: Scalable Interactive Segmentation of Biomedical Imaging Datasets with In-Context Guidance

    cs.CV 2024-12 conditional novelty 7.0 of 10

    MultiverSeg combines interactive prompting with a growing set of previously segmented image pairs to reduce the number of user interactions needed to segment a new biomedical dataset.

  2. ContextLoss: Context Information for Topology-Preserving Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A context-aware critical pixel mask loss improves topology preservation and gap closing for 2D and 3D segmentation.

Pith tools