{"id":"012f02c9-6055-4e28-a9f1-105fb6ea9ef3","arxiv_id":"1908.02994","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using expert-derived convexity and simplicity thresholds, the larger U-Net is shown to produce more anatomically implausible cardiac segmentations despite better Dice and distance scores.","lead":"The paper adds two geometric shape metrics to judge whether deep learning heart ultrasound segmentations look anatomically plausible. Applied to the public CAMUS dataset, these metrics re-rank two U-Net models: the larger network scored better on standard overlap measures but produced far more anatomically implausible shapes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 3× anatomical-outlier ratio that re-ranks U-Net 2 over U-Net 1 depends on thresholds from a single expert’s empirical minima; without inter-expert validation or sensitivity analysis, the re-ranking is unsupported.","rationale":"The reader’s weakest assumption is exactly the threshold validity, and I agree it is the central vulnerability. The paper’s novelty is the anatomical assessment, but the only evidence for that assessment is Table 1, where the outlier thresholds are the empirical minima of a single annotator. Because the conclusion is a ranking reversal, it is especially sensitive to the threshold location: U-Net 1 has 95 ana outliers (~5%) while U-Net 2 has 318 (~16%), so even a modest threshold shift could change the difference. The lack of statistical testing compounds this, but no test can help until threshold validity is established. I do not see an internal inconsistency or a lack of good faith; the concern is missing support. Since the required validation is feasible with existing CAMUS inter-expert data, the appropriate outcome is the same conditional verdict: accept only if threshold robustness is demonstrated. Thus I recommend no change from the reader’s CONDITIONAL verdict.","tokens_in":3535,"tokens_out":8547,"duration_ms":96942,"concrete_test":"Use the 50-patient inter-expert fold that already exists in the CAMUS study (Leclerc et al., in press): compute Cx/Sp thresholds separately from Expert A and Expert B, and also via bootstrap percentiles (e.g., 1st, 5th, 10th) of the pooled expert contours. Recompute the ana-outlier counts for U-Net 1 and U-Net 2 under each threshold set and report the ratio and ranking. If the U-Net 2-to-U-Net 1 ratio is not robustly around 3 across threshold choices, or if Expert B's own contours are frequently labeled as outliers under Expert A's thresholds, the central re-ranking claim is not supported. Also clarify and correct the Sp formula before recomputation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The re-ranking claim in §3.2 ('U-Net 2 ... produces three times less anatomically plausible shapes') is entirely carried by the ana-outlier column of Table 1. A prediction is called an anatomical outlier exactly when its Cx or Sp falls below the minimum values of the single expert’s annotations, as stated in §3.2. This threshold construction is the load-bearing link from raw scores to the conclusion that U-Net 2 is anatomically worse despite better Dice, MD and HD. The link is not secured: (i) no independent expert is used to show that the thresholds separate 'anatomically impossible' from plausible contours — the minima of one expert are a high-variance estimate of the lower bound of expert-plausible shapes; (ii) no sensitivity analysis is reported, so we do not know whether moving the thresholds by a small amount (or using another expert) preserves the 95 vs 318 gap or the ranking; (iii) the thresholds shown in Table 1 (e.g., Cx_endo > 0.741 vs expert mean±SD 0.975±0.022) are many SDs away from the reported expert distribution, so the derivation of 'minimum values' is not transparent. If a second expert’s own contours are frequently below these thresholds, the labels are not anatomical outliers but artifacts of threshold choice. The Sp formula is also ambiguously typeset, which would affect all threshold and outlier numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript extends the authors' prior evaluation of deep-learning segmentation on the CAMUS echocardiography dataset by adding two geometric shape metrics, convexity (Cx) and simplicity (Sp). Thresholds for separating acceptable from 'anatomical outlier' contours are set to the minimum values of these metrics over a single expert's annotations. Table 1 compares two U-Net variants and reports that U-Net 2, despite better Dice, mean distance, and Hausdorff distance, yields 318 anatomical outliers versus 95 for U-Net 1, leading the authors to conclude that traditional geometric metrics are insufficient for ranking segmentation algorithms. The paper also illustrates the criterion on two cases and argues that the proposed metrics detect local deformities that standard metrics miss.","tokens_in":4019,"tokens_out":5756,"duration_ms":58789,"significance":"The proposal to augment standard geometric metrics (Dice, MD, HD) with cheap, interpretable shape-based scores is timely, and the use of the public CAMUS dataset is a strength. The observation that a model with better Dice/MD/HD can produce more shape outliers is a valuable cautionary result for the cardiac-segmentation community. If the threshold validation were supplied, the criterion could become a useful quality-control tool. The paper is an extended abstract, so the scope is appropriately limited; however, the current evidence for the central re-ranking claim is insufficient because it rests on thresholds from a single expert and on an unvalidated identification of geometric thresholds with anatomical validity.","major_comments":[{"comment":"The anatomical-outlier label is assigned whenever a predicted Cx or Sp falls below the minimum expert value (Cx_endo > 0.741, Sp_endo > 0.529, Cx_epi > 0.960, Sp_epi > 0.694). These thresholds are minima from a single expert and are the sole classifier producing the 95 vs 318 gap that drives the re-ranking of U-Net 1 over U-Net 2. The paper offers no independent-expert validation, no sensitivity analysis, and no statistical test for the difference. Since the thresholds lie several standard deviations from the reported expert means (e.g., the Cx_endo threshold is about 10.6 standard deviations below the mean), the derivation of the minima is not transparent and the re-ranking conclusion is unsupported. Please add: (i) the distribution of expert minima, per contour and per patient, and the corresponding minima from a second expert if available; (ii) a sensitivity analysis showing outlier counts and model ranking for perturbed thresholds; (iii) a paired statistical test, such as McNemar's test on the per-case binary outlier labels.","section":"Section 3.2 and Table 1"},{"comment":"The paper equates violations of the Cx/Sp thresholds with 'anatomically impossible shapes' and calls the flagged contours 'anatomical outliers.' Cx and Sp are generic geometric shape measures; they do not encode any cardiac anatomical prior beyond convexity and compactness. The claim that these criteria are 'anatomical' is an unvalidated assumption rather than a demonstrated property. This matters because the paper's central message is about anatomical validity. Please either provide evidence, for example a blinded cardiologist review of a sample of flagged contours showing that they are indeed anatomically impossible, or rename the criterion to something like 'geometric shape outliers' and qualify the anatomical interpretation accordingly.","section":"Sections 3.1 and 3.2"},{"comment":"The typeset formula 'Sp(S) = sqrt(4π*Area(S)) / Perimeter(S)^2' is ambiguous: it is unclear whether the denominator is Perimeter squared or Perimeter times 2. Every Sp value, threshold, and outlier count in Table 1 depends on this formula. Please state the formula unambiguously, for example Sp = 4π Area / Perimeter^2, and confirm that all reported numbers use this definition.","section":"Section 3.1, definition of Sp"},{"comment":"The comparison of ana-outlier rates between U-Net 1 (21% ± 5%) and U-Net 2 (26% ± 16%) is reported without a statistical test, and the reported error bars appear to overlap. Because the classification is per-case and derived from a 10-fold cross-validation, a paired test (e.g., McNemar's test on the binary outlier labels) or a bootstrapped confidence interval for the rate difference should be provided to support the statement that U-Net 2 produces three times as many anatomical outliers.","section":"Section 3.2 and Table 1"}],"minor_comments":[{"comment":"The phrase 'it produces three times less anatomically plausible shapes' should read 'three times more anatomical outliers' or 'three times fewer anatomically plausible shapes.'","section":"Section 3.2, paragraph 2"},{"comment":"Please specify what the ± values represent (standard deviation across folds or across patients) and what the percentages in parentheses indicate (confidence intervals or standard deviations).","section":"Table 1"},{"comment":"In the caption, indicate which structure (LV-endo or LV-epi) the Cx and Sp values refer to, and reference panels (b) and (c) in the text in the order they are discussed.","section":"Figure 1"},{"comment":"The reference to 'Leclerc et al., in press' is incomplete in the bibliography; please add the journal, volume, and page numbers once the TMI article is published.","section":"References"},{"comment":"The abstract says 'The completed study sheds a new light on the ranking of models,' but the paper is an extended abstract; please clarify the relationship to the full in-press TMI study and avoid implying that the full study is contained here.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a short extended abstract that relies heavily on the authors' in-press TMI paper for context. The proposed shape-validity criterion is interesting and potentially useful for the echocardiography segmentation community, but the current evidence is not yet sufficient for a published standard: the thresholds are single-expert minima, the anatomical interpretation is asserted, and the key comparison lacks a statistical test. I recommend major revision rather than rejection because the required analyses (inter-expert validation, sensitivity analysis, a paired statistical test) are feasible within the scope of a short paper or could be reasonably deferred to a full-length follow-up."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely useful thing here is the demonstration that standard metrics can miss systematic shape failures: U-Net 2 has better Dice/MD/HD but produces roughly three times more anatomical outliers (318 vs 95 in Table 1). That is a concrete, falsifiable observation, and it should make anyone evaluating cardiac segmentation models pause before trusting a single aggregate metric. The paper is short and honest about being an extension of the authors' prior CAMUS study, and the geometric criteria (convexity, simplicity) are clearly borrowed with attribution from Zhu et al. on natural images. For what it is, the writing is clean and the direction is sensible.\n\nThe soft spots are real but not fatal to the overall idea. First, the thresholds are taken from the minima of one expert's annotations, and that single point of support carries the entire re-ranking. The stress-test note is right: we have no independent expert to show those thresholds separate anatomically impossible from plausible, and no sensitivity analysis. If a second expert's contours land below those thresholds, the 'anatomical outlier' label becomes an artifact. Second, the outlier-rate comparison (95 vs 318) is presented without a statistical test; even a paired McNemar-style test would help. Third, the simplicity formula in Section 3.1 is typeset ambiguously (the square root over the whole fraction vs just the numerator), which affects all derived numbers. Fourth, the claim that these criteria are 'anatomical' is asserted rather than established; the metrics are geometric and correlate with shape plausibility, but calling them anatomical goes a step beyond. Also, the thresholds in Table 1 look far from the expert mean±SD (e.g., Cx_endo >0.741 vs 0.975±0.022); that may indicate the minima come from some subset or a different computation, but it is not transparent as written.\n\nThat said, the central observation—that a larger, higher-Dice model can produce more implausible contours—is plausible and worth taking seriously. The paper would benefit from a more rigorous threshold validation, sensitivity analysis, and statistical testing, but it does not need to be invented from scratch. I would bring it to a reading group because it raises a practical evaluation question, not because the solution is final. I would not cite it until the threshold issue is addressed, but I would send it to peer review: it is a within-field methodology proposal with a concrete demonstration, and the flaws are fixable. The authors are clearly engaging with their own prior work and the literature; no signs of incoherence or fitting.\n\nRecommendation: send it to referees, but expect heavy revision. The idea is solid; the evidence currently overreaches its support.","headline":"The re-ranking result is real but brittle: the 3x outlier gap lives entirely on thresholds from one expert's minima, with no sensitivity analysis or statistical test, though the underlying idea is worth a serious referee.","tokens_in":4390,"tokens_out":673,"would_cite":false,"duration_ms":8847,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that standard segmentation metrics can look good while a model silently produces anatomically impossible heart shapes, and that two simple shape ratios catch these failures.","keywords":["echocardiography segmentation","CAMUS dataset","left ventricle","myocardium","convexity","simplicity","anatomical outliers","deep learning segmentation"],"falsifier":"Take the same CAMUS segmentations and recompute outlier rates using thresholds from a second expert's contours on the same 500 patients; if U-Net 2's outlier rate drops toward U-Net 1's or the two models switch rank, the reported re-ranking is an artifact of the single expert's minima.","tokens_in":3401,"feed_emoji":"🫀","tokens_out":6004,"duration_ms":56760,"temperature":0.7,"pith_summary":"This paper argues that the standard geometric metrics used to rank cardiac segmentation models (Dice score, mean distance, Hausdorff distance) can all look good while a model is silently producing anatomically impossible heart shapes. It adds two shape-validity criteria, convexity and simplicity, and sets thresholds at the lowest values found in an expert's manual contours. On the CAMUS dataset this separates \"anatomical outliers\" from acceptable segmentations and reveals a ranking change: U-Net 2, the model with better conventional scores, produces roughly three times more anatomical outliers than U-Net 1. The point of the paper is that evaluation of segmentation in echocardiography should include shape validity, not just overlap or distance errors.","feed_headline":"Better Dice score hides 3x more anatomical errors","feed_subtitle":"Convexity and simplicity thresholds from expert contours catch impossible heart shapes that standard metrics miss.","key_machinery":"The carrying mechanism is a two-criteria shape filter derived from expert contours. Convexity, $C_x(S)=\\mathrm{Area}(S)/\\mathrm{Area}(\\mathrm{ConvHull}(S))$, is near 1 for the oval left ventricle and bridge-like myocardium; simplicity, $S_p(S)=\\sqrt{4\\pi\\,\\mathrm{Area}(S)}/\\mathrm{Perimeter}(S)$, is a compactness measure. The expert's minimum values (LV-endo $C_x>0.741$, $S_p>0.529$; LV-epi $C_x>0.960$, $S_p>0.694$) define the \"anatomically possible\" boundary, and any predicted contour falling below a threshold is flagged as an anatomical outlier. This filter is what lets the authors count outliers per model and re-rank U-Net 1 against U-Net 2.","core_discovery":"The central claim is that a segmentation can pass Dice and distance thresholds yet still be anatomically invalid, and that two simple geometric ratios catch these failures. Convexity is the area of the structure divided by the area of its convex hull; simplicity is a circularity-like ratio of area to perimeter. Using the minimum convexity and simplicity values from one expert's annotations on the CAMUS dataset as thresholds, the authors label predictions below any threshold as anatomical outliers. Applying these labels to two U-Net models reverses the impression left by the conventional metrics: U-Net 2, with 18M trainable parameters and better Dice, mean absolute distance and Hausdorff distance, produces 318 anatomical outliers (16%), while U-Net 1 produces 95 (5%). The paper concludes that traditional metrics are insufficient to rank segmentation algorithms.","pith_inferences":["The expert-minimum thresholds are a single point of sensitivity: recomputing the outliers with a second expert's contours could change the absolute rates, and the paper does not test whether the 5-percent-versus-16-percent gap is stable.","The same shape filter could plausibly be used during training as a differentiable penalty or as a post-processing rejection rule, not only as an evaluation metric.","If applied to other cardiac datasets, the thresholds would need recalibration, since image-plane geometry and expert contouring style shift the natural range of convexity and simplicity."],"forward_implications":["Model ranking in cardiac segmentation depends on the evaluation metric: U-Net 2's lead on Dice and distance metrics disappears or reverses when anatomical outlier counts are used.","Traditional Dice- or distance-based evaluation can pass clinically unusable predictions, so published comparisons should report shape-validity scores alongside geometric scores.","The anatomical outlier label identifies cases that a clinician might need to review or correct, because the predicted shape is locally deformed even when global scores look acceptable.","The same convexity and simplicity criteria extend naturally to other structures and views in the CAMUS data, since the myocardium and left atrium also have characteristic shapes."],"supporting_citations":[{"why":"Provides the CAMUS dataset, the two U-Net models, and the Dice/MD/HD scores that the anatomical metrics re-rank.","marker":"(Leclerc et al., in press)"},{"why":"Supplies the convexity and simplicity shape criteria adapted for cardiac structures.","marker":"(Zhu et al., 2017)"},{"why":"Defines the U-Net architecture used by both compared models.","marker":"(Ronneberger et al., 2015)"}],"fun_headline_variants":["Dice score masks 3x more anatomical errors","Anatomical shape check flips U-Net ranking","Convexity test catches hidden shape errors","Better Dice, worse anatomy: 3x more outliers","Shape validity: the missing metric for echo segmentation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire outlier count depends on treating the lowest convexity and simplicity values present in one expert's manual outlines as a universal boundary between anatomically possible and impossible shapes, with no independent-expert validation or sensitivity analysis.","fun_headline_variants_meta":{"raw":{"variants":["Dice score masks 3x more anatomical errors","Anatomical shape check flips U-Net ranking","Convexity test catches hidden shape errors","Better Dice, worse anatomy: 3x more outliers","Shape validity: the missing metric for echo segmentation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000588,"raw_usage":{"total_tokens":2669,"prompt_tokens":763,"completion_tokens":1906,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":379,"completion_tokens_details":{"reasoning_tokens":1832}},"tokens_in":379,"tokens_out":1906,"duration_ms":15771,"temperature":1.0,"reasoning_tokens":1832,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:26:40.871951+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same CAMUS segmentations and recompute outlier rates using thresholds from a second expert's contours on the same 500 patients; if U-Net 2's outlier rate drops toward U-Net 1's or the two models switch rank, the reported re-ranking is an artifact of the single expert's minima.","supporting_citations":[{"cited_title":"Semantic Amodal Segmentation","cited_arxiv_id":null,"evidence_quote":"Supplies the convexity and simplicity shape criteria adapted for cardiac structures."},{"cited_title":"U-Net: Convolutional Networks for Biomedical Image Segmentation","cited_arxiv_id":null,"evidence_quote":"Defines the U-Net architecture used by both compared models."}],"review_version":1}