{"id":"5e34fb6b-2c45-4294-8afc-c771a8a4c2fd","arxiv_id":"1908.10454","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of deep learning segmentation techniques for scarce and weak annotations, with cost-gain recommendations.","lead":"This paper is a survey of deep learning methods for medical image segmentation when training labels are scarce or weak, organized into a taxonomy with recommendations. It is a reference map for practitioners choosing among augmentation, semi-supervision, regularization, and weak-supervision techniques.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 6 recommendations are built on unreported synthesis criteria and at least one single-study generalization, so the practical cost-gain ranking is underdetermined.","rationale":"The paper's taxonomy of scarce versus weak annotations is coherent, and the survey is a useful reference; I found no internal inconsistency in the classification structure. The stress-test lands on the advisory layer. Table 7 does not actually tabulate performance gains, and Section 6's 'cost-gain trade-off' is a qualitative synthesis with no stated inclusion criteria, no normalization across heterogeneous experiments, and no explicit treatment of annotation cost or implementation difficulty. The single-study basis for some recommendations—Feng et al. (2017) for image-level labels and Mirikharaji et al. (2019) for noisy labels—makes the generalization to all medical segmentation scenarios particularly fragile. This matches and sharpens the reader's weakest assumption about comparability of reported gains. Therefore the CONDITIONAL verdict should stand: the paper should either make its evidence synthesis explicit or temper the groundedness of its recommendations.","tokens_in":46784,"tokens_out":3673,"duration_ms":39929,"concrete_test":"Construct a normalized evidence table for Section 6's recommendations: for each recommended method family, list all surveyed papers that report a quantitative segmentation gain, recording dataset, organ, backbone, baseline, and gain. Restrict the table to comparisons sharing a common baseline (e.g., U-Net without augmentation) or to public benchmarks (BraTS, ISIC, LiTS) and re-rank the families by median gain per annotation cost. If the Section 6 ordering changes, the recommendation logic is not robust to evidence synthesis; if it survives, the concern is resolved. As a targeted check, enumerate all papers in Section 5.1 that report Dice relative to full supervision: if Feng et al. (2017) is the only such report for CAM-based methods, the Section 6 recommendation is a single-study generalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central advisory claim—that Section 6 and Table 7 let a practitioner choose a method family by cost-gain trade-off—depends on the surveyed evidence being comparable and representative. That condition is not met. The paper states no search or inclusion protocol, uses no common benchmark, and does not normalize reported Dice gains across datasets, organs, backbones, baselines, or annotation regimes. Table 7 encodes only data requirements via color; it does not tabulate measured gain or implementation cost, so the 'cost-gain' ordering in Section 6 is a narrative judgment rather than a derived comparison. More concretely, at least one recommendation rests on a single primary study: the Section 6 recommendation of modified CAM-based methods for image-level labels is supported mainly by Feng et al. (2017), a pulmonary-nodule-specific result, and is then generalized to all medical image segmentation. The noisy-label recommendation similarly leans on Mirikharaji et al. (2019) for skin lesions. If those reported gains are dataset-specific, the recommended priorities could reorder under a more systematic evidence synthesis. This is not an internal inconsistency in the scarce/weak taxonomy, but it makes the survey's central practical output underdetermined by its evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews deep learning methods for medical image segmentation when training datasets are imperfect. It organizes the field into two broad problems—scarce annotations, where only limited labeled data exist, and weak annotations, where labels are sparse, noisy, or image-level only—and groups solution strategies accordingly: data augmentation, external labeled data, cost-effective annotation, unlabeled data, regularized training, and post-segmentation refinement for scarce annotations; and CAM/MIL, selective loss, and robust loss for weak annotations. For each strategy the paper summarizes representative works and reported results, compares methods in Table 7 according to required data resources, and offers practical recommendations in Section 6 based on cost-gain trade-offs. The survey is careful in places, explicitly noting where performance gains are regime-dependent and where comparisons are missing, but the advisory component rests on qualitative synthesis rather than a formal evidence table.","tokens_in":46975,"tokens_out":4997,"duration_ms":50503,"significance":"If its recommendations are accepted, the survey would provide a useful practical map for practitioners choosing among method families for imperfect medical segmentation datasets. Its main contribution is the scarce/weak taxonomy and the structured summaries in Tables 3–6, which organize a large and heterogeneous literature. The authors deserve credit for flagging unresolved issues, including that self-supervised and semi-supervised gains often narrow as labeled data grow (§4.4.1, §4.4.3), that one-shot synthetic augmentation results may not transfer to larger regimes (§4.1.3), and that CRF-based refinement shows mixed results in 3D (§4.6). The central weakness is that the Section 6 recommendations are not backed by a systematic, transparent synthesis of the reported Dice gains, so the practical cost-gain ordering is underdetermined by the evidence presented. This is a limitation of the advisory component, not an internal inconsistency in the taxonomy.","major_comments":[{"comment":"The paper's central practical claim—that Section 6 and Table 7 allow a practitioner to choose a method family by cost-gain trade-off—is underdetermined by the presented evidence. The survey states no search or inclusion protocol, uses no common benchmark, and does not normalize reported Dice gains across datasets, organs, architectures, baselines, or annotation regimes. Table 7 encodes only data requirements through color and provides method descriptions; it does not tabulate the measured performance gain or implementation cost behind the recommendations. As a result, the cost-gain ordering in Section 6 is a narrative judgment rather than a derived comparison. Please either add a systematic evidence table with reported gains and data regime per study, conduct a meta-analytic comparison, or explicitly reframe the recommendations as qualitative expert opinion.","section":"Section 6, Table 7"},{"comment":"Several Section 6 recommendations generalize from a single primary study. The recommendation of modified CAM-based approaches for image-level labels is supported chiefly by Feng et al. (2017), a pulmonary-nodule study on LIDC-IDRI, and is extended without further evidence to all medical image segmentation, including its use in the combined-scenario example in Section 6. Likewise, the noisy-label recommendation leans on Mirikharaji et al. (2019), a skin-lesion study with simulated polygon noise. If these gains are dataset- or noise-model-specific, the recommended priorities could reorder. Please add corroborating studies, or explicitly state the single-study basis and its scope when issuing a recommendation.","section":"Section 6, §5.1.1, §5.3.1"},{"comment":"The recommendations in Section 6 do not condition on the data regime despite the survey's own caveats. The manuscript repeatedly flags that gains are regime-dependent: Zhao et al. (2019a) is tested in a one-shot setting with unclear gains for larger training sets (§4.1.3); Models Genesis gains may change in the presence of data augmentation (§4.4.1); and semi-supervised methods without pseudo labels 'are not as effective when the training set grows' (§4.4.3). Section 6 nevertheless issues unconditional recommendations, e.g., for shape regularization, same-domain synthesis, and semi-supervised learning, without specifying for which labeled-set sizes or annotation budgets the supporting evidence was obtained. Please tie each recommendation to the data regime in which the evidence was gathered, or explicitly acknowledge this axis of uncertainty.","section":"Section 4.1.3, §4.4.1, §4.4.3, Section 6"}],"minor_comments":[{"comment":"There are several typographical errors: 'foregin' should be 'foreign' (§4.4.1), 'peior' should be 'prior' (§4.5.3), and 'to to combine' should be 'to combine' (§5.1.1). A careful copyedit is needed.","section":"§4.4.1, §4.5.3, §5.1.1"},{"comment":"Figure 1 and Table 7 rely on color shading to distinguish categories and data requirements; the distinctions are difficult to decode in grayscale. Please add symbols or hatching and increase the font sizes in Figure 1 for readability.","section":"Figure 1, Table 7"},{"comment":"Several summary tables list methods but omit the reported quantitative gains. Adding a column with the reported metric change and the data regime (labeled-set size, annotation type) would materially support the Section 6 recommendations and reduce the reader's need to consult the primary literature.","section":"Tables 1, 2, 4, 5, 6"},{"comment":"The survey does not state the literature search period, databases, or inclusion criteria. A short paragraph describing the selection scope would help readers judge coverage and would make the Section 6 synthesis more reproducible.","section":"Section 3, Section 6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a medical image analysis venue and the taxonomy is useful, but the Section 6 recommendations currently outrun the evidence. A revision that adds a transparent evidence table or scales back the advisory language would make the contribution acceptable. I also recommend adding a competing-interests statement, given the authors' industrial affiliation; I do not see an obvious citation-bias problem, as the self-citations are used as examples rather than as load-bearing evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a well-organized survey that will serve as a handy reference for practitioners facing scarce or weak annotations in medical image segmentation. The scarce/weak taxonomy is a real organizing contribution, and the paper is honest about many open issues. The soft spot is Section 6: the recommendations are built on unreported synthesis criteria and at least one single-study generalization. I don't think the taxonomy is in trouble, but the advisory claims are weaker than they look.\n\nWhat's new: the paper groups a large literature into two high-level problems (scarce annotations, weak annotations) and then into concrete strategies (augmentation, transfer/domain adaptation, active learning, self/semi-supervised, regularization, CRF post-processing; CAM/MIL, sparse-label losses, noisy-label losses). The tree in Figure 1 and the summary tables (especially Table 7) make it easy to locate methods by data limitation and data resources. That makes it a genuinely useful starting point for a newcomer and a quick reference for a practitioner. The authors also do something surveys often avoid: they flag when gains vanish as training set grows (e.g., self-supervised pre-training, semi-supervised without pseudo labels), when comparisons are missing, and when a result is dataset-specific. That is genuine scholarship.\n\nSoft spots: There is no stated search or inclusion protocol, so the coverage is selective rather than systematic. That's fine for a narrative review, but it matters because the recommendations in Section 6 are presented as a practical decision guide. The evidence base is heterogeneous—different datasets, organs, backbones, baselines, and metrics—and the paper does not normalize or report measured gains in Table 7. The cost-gain ordering is therefore a narrative judgment. The stress-test concern is on point: the modified-CAM recommendation leans on Feng et al. (2017), a pulmonary-nodule study, and then generalizes to all medical image segmentation; the noisy-label recommendation similarly leans on Mirikharaji et al. (2019) for skin lesions. Those may be the best available evidence, but a reader should be told how much weight to put on them.\n\nRecommendation: This deserves peer review at a journal that publishes surveys. The taxonomy alone is a useful contribution. The revision should either soften the recommendations to be explicitly provisional, or add a short meta-analytic summary table that tabulates the underlying studies, their datasets, and their reported gains. That would turn a suggestive discussion into something a practitioner can actually act on. Who should read it: anyone entering the label-efficient segmentation space, and researchers who need a structured map of the field.","headline":"A genuinely useful taxonomy of label-efficient medical image segmentation methods, but the Section 6 cost-gain recommendations are qualitative judgments over heterogeneous evidence, not a derived comparison.","tokens_in":47525,"tokens_out":4817,"would_cite":true,"duration_ms":41044,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Imperfect medical image segmentation datasets are not a dead end: the paper argues that scarce and weak annotations each have a matching toolbox of deep-learning methods, and that the right choice follows a cost-gain trade-off a…","keywords":["medical image segmentation","scarce annotations","weak annotations","sparse annotations","noisy labels","image-level labels","semi-supervised learning","active learning"],"falsifier":"Take a single segmentation backbone and a single medical dataset, then create controlled training conditions covering clean labels at 1%, 10%, and 50%, plus sparse, noisy, and image-level labels. Apply one representative method from each tier of Table 7 under identical data budgets and measure Dice on a fixed test set: if the green-tier methods (shape regularization, mixing, same-domain synthesis) are not among the best at low budgets, or if the overall ranking contradicts the cost-gain ordering, the survey's central recommendation logic would be falsified.","tokens_in":46589,"feed_emoji":"🩺","tokens_out":5879,"duration_ms":56355,"temperature":0.7,"pith_summary":"This survey establishes a working distinction between two failure modes of medical image segmentation datasets: scarce annotations, where too few images are densely labeled, and weak annotations, where labels are sparse, noisy, or image-level only. Against this taxonomy, it organizes the deep-learning literature into method families and argues that these families can be compared by cost-gain trade-offs, that is, how much performance they buy per unit of additional annotation or data effort. The paper's recommendations follow from this comparison: use no-extra-data methods first, add unlabeled or similar-domain data next, and call in experts only when richer annotation is genuinely needed. A sympathetic reader who accepts the survey's reading can choose a method for a given data budget without reading hundreds of primary papers, and can combine strategies when a dataset suffers from both scarcities and weaknesses.","feed_headline":"Scarce or weak labels: a guide to medical segmentation fixes","feed_subtitle":"Ranks method families by data cost so practitioners can pick one from a color-coded guide.","key_machinery":"The carrying object is a two-level taxonomy in which dataset limitations are first split into scarce and weak annotations, and then each is sub-divided by general strategy. Its practical engine is Table 7, which color-codes every methodology by required data resources: green for methods needing only the original limited annotated dataset, orange for methods needing additional unlabeled or similar-domain labeled data, and red for methods requiring experts in the loop. That color-coding turns a literature review into a decision procedure, because Section 6's recommendations are derived directly from it: the paper advises applying green-tier methods wherever possible, using orange-tier methods according to the availability of auxiliary data, and reserving red-tier methods for situations where more annotation is genuinely needed.","core_discovery":"The paper's central claim is that the practical bottleneck in medical image segmentation is not network architecture but the imperfection of training labels, and that this bottleneck splits into two tractable problems. Scarce annotations are addressed by increasing effective data through augmentation, synthesis, external datasets, unlabeled data, cost-effective annotation, regularization, or CRF post-processing; weak annotations are addressed per annotation type, with class activation maps and multiple instance learning for image-level labels, selective loss with or without mask completion for sparse labels, and robust loss with or without iterative mask refinement for noisy labels. The authors further claim that these methods differ more in the data resources they require than in the accuracy they achieve, and they rank them accordingly. The result is a set of concrete recommendations: shape regularization, mixing augmentation, and same-domain synthesis are suggested wherever possible; self-supervised pre-training is singled out as the most promising medium-resource approach; active learning and interactive segmentation are reserved for when expert annotation is unavoidable; and modified CAM-based approaches are preferred for image-level labels.","pith_inferences":["If the resource-based ranking is right, it implies that future segmentation papers should report annotation cost and data requirements alongside Dice, otherwise the cost-gain table cannot be updated as new methods appear.","The taxonomy suggests a modular benchmark: fix one backbone and one organ, then vary annotation quantity and quality across the three weak-annotation types; the predicted ordering of method families would be directly testable.","The survey's 'use wherever possible' tier is a strong empirical prediction: shape regularization, mixing, and same-domain synthesis should dominate more data-hungry methods at very small annotation budgets, but the ordering may reverse when budgets grow."],"forward_implications":["A team with a small but cleanly labeled dataset should first exhaust traditional augmentation, shape regularization, mixing augmentation, and same-domain synthesis before seeking more data.","A team with extra unlabeled images can expect self-supervised pre-training to give gains even at full training-set size, unlike many semi-supervised methods whose advantage shrinks as labeled data grows.","When labels are image-level, modified class activation map approaches are the recommended family, reportedly landing within a few Dice points of full supervision.","Datasets that combine scarce and weak annotations can be handled by composing solutions, for example a semi-supervised method plus a CAM-based weakly supervised branch."],"supporting_citations":[{"why":"Defines the U-Net architecture that anchors nearly every method family reviewed in the survey.","marker":"Ronneberger et al. (2015)"},{"why":"Models Genesis provides the key evidence for self-supervised pre-training, which Section 6 calls the most promising medium-resource approach.","marker":"Zhou et al. (2019b)"},{"why":"Shows a noise-resilient loss can match clean-annotation performance, underpinning the survey's noisy-label recommendations.","marker":"Mirikharaji et al. (2019)"},{"why":"Reports semi-supervised gains that persist even with the full labeled set, cited as rare evidence for semi-supervised learning without pseudo labels.","marker":"Chen et al. (2019c)"},{"why":"Introduces annotation cost into active learning, the basis for the survey's cost-sensitive annotation recommendations.","marker":"Kuo et al. (2018)"},{"why":"Demonstrates a two-stage CAM approach for image-level labels, the basis for the survey's preferred weakly supervised method family.","marker":"Feng et al. (2017)"},{"why":"Shows sparse dot-grid annotations can approximate dense-label performance, supporting the sparse-annotation guidance.","marker":"Silvestri and Antiga (2018)"}],"fun_headline_variants":["Imperfect labels: the real bottleneck in medical segmentation","Scarce or weak annotations? A practical review of fixes","Beyond architecture: solving the imperfect label problem","A field guide to handling imperfect medical datasets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that Dice-score gains reported in different papers on different datasets, organs, architectures, and baselines are roughly comparable, so the method rankings in Table 7 would survive re-measurement on one common benchmark.","fun_headline_variants_meta":{"raw":{"variants":["Imperfect labels: the real bottleneck in medical segmentation","Scarce or weak annotations? A practical review of fixes","Beyond architecture: solving the imperfect label problem","A field guide to handling imperfect medical datasets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000623,"raw_usage":{"total_tokens":2872,"prompt_tokens":916,"completion_tokens":1956,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":1895}},"tokens_in":532,"tokens_out":1956,"duration_ms":17802,"temperature":1.0,"reasoning_tokens":1895,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:42:16.243779+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a single segmentation backbone and a single medical dataset, then create controlled training conditions covering clean labels at 1%, 10%, and 50%, plus sparse, noisy, and image-level labels. Apply one representative method from each tier of Table 7 under identical data budgets and measure Dice on a fixed test set: if the green-tier methods (shape regularization, mixing, same-domain synthesis) are not among the best at low budgets, or if the overall ranking contradicts the cost-gain ordering, the survey's central recommendation logic would be falsified.","supporting_citations":[],"review_version":1}