{"id":"614c1e4d-df2b-41e5-afb5-b4e631e3ba61","arxiv_id":"2508.14133","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An nnU-Net trained on 90 hepatobiliary-phase MRI scans segments liver anatomy with high Dice scores for most structures, while tumor segmentation remains variable.","lead":"Researchers trained an nnU-Net model to automatically outline liver tissue, tumors, and major vessels in contrast-enhanced MRI scans used for surgical planning. On a held-out set of 18 patients it matched manual outlines well for most structures, and the authors report that clinical use shortened the segmentation time from hours to about 15 minutes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manual segmentations by two technical physicians are treated as ground truth for all five structures, but no inter-observer variability is reported; without a human-human DSC baseline, the reported accuracy (e.g., portal vein DSC 0.74±0.06, tumor DSC 0.77±0.17) cannot be separated from label noise.","rationale":"The reader's weakest_assumption identifies the unmeasured reliability of the manual reference labels, and I agree this is the most load-bearing gap. Every accuracy number in the paper is a similarity between model output and that reference, so if the reference is noisy or systematically biased, the headline DSCs lose their meaning. The issue is particularly acute for portal vein, biliary tree, and tumors, where typical DSC values (0.74–0.80) are already moderate and where annotation errors are known to be largest. I also note two adjacent weaknesses that reinforce the conditional verdict: the assessment-set comparison is between the model and a human correction of that model rather than an independent reference, and the claimed reduction in segmentation time from several hours to ~15 minutes is not supported by any reported timing measurement. Neither of these invalidates the study, but they mean the central claim is not yet established beyond a conditional acceptance. The proposed test—an independent re-contouring of a test-set subset and comparison of model-human DSC to human-human DSC—would directly settle the label-noise concern. Because the reader already assigned CONDITIONAL and this concern supports that judgment rather than overturning it, I recommend UNCHANGED.","tokens_in":8847,"tokens_out":8667,"duration_ms":92201,"concrete_test":"Take a random subset of at least 10 patients from the 18-patient test set. Have a second independent observer—ideally a different technical physician or radiologist, blinded to model output—re-segment all five structures from the same hepatobiliary-phase MRI using the same 3D Slicer protocol. Compute pairwise human-human DSC (and, if feasible, average symmetric surface distance) between the original reference and the new contours, per structure. Then compare each model-vs-original-reference DSC to the corresponding human-human DSC distribution. If model DSCs fall within the human-human variability envelope (e.g., not significantly below median human-human DSC), the reported accuracy can be interpreted as human-level and the central claim holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All quantitative claims—test-set DSCs (Results, Figure 2), assessment-set DSCs (Results, Figure 3), and the comparison to Oh et al. in the Discussion—are computed against manual delineations made by two technical physicians and confirmed by a hepatobiliary surgeon (Methods, Manual segmentations). The manuscript reports no inter-observer variability, no repeat-segmentation study, and no independent second contour for any structure. This is load-bearing because the weakest structure scores are exactly the ones where annotation noise is largest: portal vein 0.74±0.06, biliary tree 0.79±0.07, hepatic vein 0.80±0.04, tumors 0.77±0.17. For thin vessels and small tumors, a one-voxel boundary shift or a missed small lesion changes DSC substantially. If the reference contains systematic over/under-segmentation, a model trained on those labels (n=72) will match them well while disagreeing with true anatomy; the reported DSCs then overstate anatomical accuracy. The Discussion explicitly acknowledges that external validation was not performed, but even within this single-center cohort the reliability of the reference is unmeasured. The assessment set has a related but distinct issue: it compares model output to a manual refinement of that same output, so it measures editing effort, not independent accuracy. A human-human DSC baseline would let the reader interpret all reported numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports the development and clinical integration of an nnU-Net-based automated segmentation method for hepatic anatomy—liver parenchyma, tumors, portal vein, hepatic vein, and biliary tree—from the hepatobiliary phase of gadoxetic acid-enhanced MRI. Manual segmentations from 90 patients were used to train (n=72) and test (n=18) the model, with an additional assessment cohort of 10 patients where model outputs were manually refined for clinical use. The authors report test-set Dice similarity coefficients of 0.97 for parenchyma, 0.80 for hepatic vein, 0.79 for biliary tree, 0.77 for tumors, and 0.74 for portal vein, together with a mean tumor detection rate of 76.6% and a median of one false positive per patient. In the assessment set, high DSC values are reported for parenchyma, portal vein, and hepatic vein after manual refinement, and the authors state that segmentation time decreased from several hours to about 15 minutes per patient. The paper concludes that the method enables accurate automated delineation and supports broader adoption of 3D planning in liver surgery.","tokens_in":9163,"tokens_out":3106,"duration_ms":34660,"significance":"If the reported performance holds under independent verification, the work has clear translational value: it addresses a real clinical bottleneck in liver surgery planning, uses a clinically relevant MRI phase, and includes a prospective assessment of workflow impact. The comparison with prior MRI-based liver segmentation studies (Zbinden et al., Ivashchenko et al., Oh et al.) is useful positioning. The strengths include a held-out test set, a consecutively collected single-center cohort, clinical integration with manual refinement, and reporting of tumor detection rate and false positives. However, the central quantitative claims rest on three points that are not currently established: the reliability of the manual ground truth, the contribution of the customized loss function, and the interpretation of the assessment-set DSC as evidence of accuracy. These are load-bearing because the weakest structures (portal vein, biliary tree, tumors) are exactly those where annotation variability and loss-function choices are most consequential.","major_comments":[{"comment":"The manuscript uses manual segmentations by two technical physicians, confirmed by a hepatobiliary surgeon, as ground truth for all five structures, but reports no inter-observer variability, no repeat-segmentation study, and no independent second contour for any structure. This is load-bearing because the structures with the lowest DSCs—portal vein (0.74±0.06), biliary tree (0.79±0.07), hepatic vein (0.80±0.04), and tumors (0.77±0.17)—are precisely the structures where boundary definition and lesion identification are most subjective. Without a human-human DSC baseline, the reader cannot separate model error from annotation noise, and the conclusion that the model provides 'accurate delineation' is not supported. Please provide an inter-observer variability measurement (e.g., DSC between independent manual segmentations) on a representative subset, or alternatively report the uncertainty this introduces into all quoted DSC values.","section":"Methods, Manual segmentations; Results, Quantitative evaluation"},{"comment":"The methodological claim that the combination of clDice and bootstrapped cross-entropy improves thin-structure delineation and topology preservation is not tested. The paper reports no ablation against a standard nnU-Net baseline (e.g., default cross-entropy plus Dice loss) on the same training and test splits. Since nnU-Net v1 with default settings is a strong baseline, the observed DSC values could be attributable to the architecture and data rather than to the customized loss. Moreover, the bootstrapped cross-entropy schedule (K growing from 15% to 50% over the final 100 epochs after a 400-epoch warm-up) introduces several free hyperparameters without sensitivity analysis. Please add an ablation study comparing (i) default nnU-Net, (ii) default nnU-Net plus clDice, and (iii) the full proposed loss combination, reporting DSC and topology metrics for at least the vascular and biliary structures.","section":"Methods, Automated segmentation model"},{"comment":"The assessment-set evaluation compares the automated segmentations to the manually refined versions of those same automated segmentations. A high DSC in this comparison primarily measures the amount of editing performed, not the anatomical accuracy of the model, and the manual refinement process is likely biased toward minimal changes rather than independent re-delineation. The statement that 'minor adjustments were required for clinical use' is therefore not equivalent to anatomical accuracy. Please either compare assessment-set outputs to an independent manual segmentation of the assessment scans, or reframe this analysis explicitly as a measure of editing effort, supported by quantitative edit-distance or time metrics rather than DSC.","section":"Results, Performance of the network in prospective use; Discussion"},{"comment":"The paper claims that 'the model detected three additional tumors initially missed by radiologists,' but this is an anecdotal observation from prospective clinical use without a systematic reference standard or an evaluation protocol for detection sensitivity in the assessment cohort. Given that the test-set tumor detection rate is only 76.6% with a median of one false positive per patient, this claim should be contextualized as a case observation, not a quantitative result. The external-validation limitation is explicitly acknowledged in the Discussion, and I agree it is important; however, given the conclusion advocating standard-of-care use, the absence of external or multi-center validation should also be reflected in the abstract and conclusion as a qualification of the generalizability claim.","section":"Discussion, Limitations; Results, Prospective tumor detection"}],"minor_comments":[{"comment":"The text reports a combined DSC of 0.84±0.06 for 'gallbladder (central biliary tree) and bile ducts (peripheral biliary tree),' while the overall biliary tree DSC is 0.79±0.07. Please clarify whether these are separate structures or pooled in the figure, and define how cholecystectomy cases are handled in the calculation of the combined value.","section":"Results, Quantitative evaluation; Figure 2"},{"comment":"The pixel-spacing row reads '-1.5 1' and the footnote explains that negative pixel spacing indicates partially overlapping slices, but this formatting is confusing. Please present the acquisition and reconstruction voxel sizes more clearly, and specify which spacing value is used by nnU-Net for resampling.","section":"Table 1"},{"comment":"The statistical analysis section states that DSC scores were compared using an unpaired Mann-Whitney test, but it does not state which groups were compared. Please specify the comparisons (e.g., central vs. peripheral vessels, training vs. test demographics) and correct the Python/SciPy version and the SPSS version if needed.","section":"Statistical analysis"},{"comment":"Reference 29 (Strahler, 1957) is listed in the references but does not appear to be cited in the main text. If it is used for the notion of vessel topology or branching order, please cite it explicitly where that concept is introduced.","section":"Introduction and Discussion"},{"comment":"The framework is referred to inconsistently as 'nnUNet' and 'nnU-Net'; please use a single consistent spelling, preferably 'nnU-Net' as in the reference to Isensee et al.","section":"Throughout"},{"comment":"The comparison with Oh et al. reports higher DSC values for the proposed method across all structures, but no statistical comparison or matching of annotation protocols is provided. Please state whether the comparison is descriptive only, and note any differences in tumor inclusion criteria, vessel definition, or evaluation methodology that could affect the comparison.","section":"Discussion, comparison with prior work"}],"recommendation":"major_revision","confidential_remarks":"This is a clinically relevant and clearly written manuscript, but the current evidence does not yet support the strength of the central conclusions. The two most important revisions are (1) an inter-observer variability analysis of the manual ground truth and (2) an ablation of the customized loss function against a default nnU-Net baseline. Without these, the reported DSC values and the claimed contribution of the loss design cannot be interpreted. The assessment-set analysis should also be repositioned as an editing-effort metric or validated against independent manual contours. I see no indication of circular evaluation (the test set is held out, and the loss schedule was not tuned on the test set), and the work fits the journal's scope. My recommendation is major revision rather than rejection because the dataset and clinical integration are valuable and the requested analyses are feasible within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you do liver MRI segmentation or surgical planning. This is a clinical integration study more than a method paper. The new pieces are a 90-patient hepatobiliary-phase MRI cohort with manual segmentations of five structures, a training setup that combines nnU-Net with clDice and bootstrapped cross-entropy, and a small prospective assessment of how much manual editing the model outputs need in practice. The time reduction from hours to roughly 15 minutes per patient is meaningful, and the prospective finding of three sub-centimeter tumors initially missed by radiologists shows the model can add real value.\n\nThe method part is thin: there is no ablation against a standard nnU-Net baseline, so we do not actually know what clDice and bootstrapped cross-entropy buy you. The test set is 18 patients and the assessment set is 10; that is small but acceptable for a clinical feasibility study.\n\nThe soft spot that matters most is label noise. Manual segmentations by two technical physicians are the reference for all five structures, and no inter-observer variability is reported. For thin vessels (portal vein DSC 0.74) and small tumors (0.77 with SD 0.17), a one-voxel boundary shift can change DSC substantially. Without a human-human DSC baseline, the reported numbers could overstate anatomical accuracy even if the model matches the manual labels closely. The stress-test note is right about this, and it is not a manufactured flaw: it is load-bearing for interpreting the numbers. The assessment-set DSCs (e.g., portal vein 0.98) compare model output to manual refinement of that same output, so they measure editing effort rather than independent accuracy. The paper should say that plainly.\n\nThe other issue is the conclusion. The authors acknowledge external validation is missing, yet the abstract and conclusion claim standard-of-care use for every patient. That overreaches. For a single center, the results support efficient assisted planning, not standard-of-care.\n\nCitation pattern is fine. Oh et al. is cited and discussed fairly, and the comparison is direct. The paper deserves a serious referee, though it needs revisions on the label-noise baseline and the conclusion. I would not desk-reject it.","headline":"Solid single-center clinical integration study for HBP-MRI liver segmentation, with real workflow gains, but the quantitative evaluation is thinner than the standard-of-care conclusion suggests.","tokens_in":9734,"tokens_out":1947,"would_cite":true,"duration_ms":21103,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single automated MRI segmentation model can produce clinically usable 3D liver models in about 15 minutes per patient.","keywords":["liver MRI","automated segmentation","patient-specific 3D models","surgical planning","liver surgery","nnU-Net","hepatobiliary phase","clDice"],"falsifier":"Segment the same 18 test scans with two or more independent expert annotators and compute their pairwise overlap scores for portal vein, hepatic vein, biliary tree, and tumors; if expert-to-expert agreement is no better than the model's scores on those structures, the model is operating at the limit of the reference standard rather than demonstrating true anatomical accuracy.","tokens_in":8653,"feed_emoji":"🩺","tokens_out":10531,"duration_ms":97136,"temperature":0.7,"pith_summary":"This paper sets out to show that one deep-learning segmentation network, trained on hepatobiliary-phase MRI scans, can produce the full set of anatomical outlines a liver surgeon needs before an operation: the liver itself, tumors, portal and hepatic veins, and the biliary tree. The authors trained an nnU-Net on 72 manually segmented patients with a topology-preserving loss aimed at keeping thin vessels connected, then tested it on 18 unseen patients. In prospective clinical use the automated segmentations needed only minor corrections for the main structures, reducing the 3D-modeling step from several hours to about 15 minutes per patient and identifying three sub-centimeter tumors that radiologists had initially missed. If these results generalize, 3D planning could become a routine component of liver-surgery workup rather than a specialized manual service.","feed_headline":"15 minutes: one MRI model maps liver anatomy for surgical planning","feed_subtitle":"Automated outlines needed only minor corrections in clinical use and caught three tumors radiologists initially missed","key_machinery":"The mechanism that carries the argument is an nnU-Net v1 segmentation network trained with a composite loss: clDice plus bootstrapped cross-entropy. clDice is a topology-preserving loss that skeletonizes the ground-truth and predicted vessel trees and penalizes disconnected or missing tubular structures; it is what keeps thin portal and hepatic vein branches and small bile ducts attached to the main tree. Bootstrapped cross-entropy selects only the hardest voxels for the loss, after a 400-epoch warm-up with ordinary cross-entropy, with the top-K fraction growing from 15% to 50% over the final 100 epochs. The input is the 20-minute hepatobiliary phase of a Gd-EOB-DTPA-enhanced 3T MRI, a sequence chosen because it shows both vascular anatomy and bile excretion. Manual segmentations made by two experienced technical physicians and confirmed by a hepatobiliary surgeon provide the reference standard for both training and evaluation.","core_discovery":"The central claim is that the hepatobiliary phase of gadoxetic acid-enhanced MRI carries enough contrast information for a self-configuring nnU-Net to delineate all five structures that matter for liver surgery planning, and that the resulting segmentations are clinically usable with only light manual touch-up. On the 18-patient test set, mean Dice similarity coefficients (0-to-1 overlap scores) were 0.97 for liver parenchyma, 0.80 for the hepatic vein, 0.79 for the biliary tree, 0.77 for tumors, and 0.74 for the portal vein; the average tumor detection rate was 76.6%, with a median of one false positive per patient. In the 10-patient assessment dataset collected after integration into clinical practice, the parenchyma reached 1.00, portal vein 0.98, hepatic vein 0.95, and tumor 0.80, with the largest manual corrections needed for tumors in patients with prior liver interventions, high tumor burden, or low scan quality. The workflow time for producing a 3D model fell from several hours to roughly 15 minutes per patient, and the network identified three sub-centimeter malignant lesions that radiologists had not initially reported.","pith_inferences":["A testable extension is external validation: running the trained network on multi-center MRI data with different scanners and protocols would show whether the reported vessel accuracy holds across imaging setups or is tied to one protocol.","Since inter-observer variability of the manual reference was not reported, computing pairwise expert-vs-expert Dice on the same test scans would reveal how much of the model's apparent error is actually reference noise, especially for thin vessels and small tumors.","The three incidentally detected sub-centimeter tumors suggest a second-reader role for the model, but establishing that requires a blinded prospective reader study comparing radiologists with and without the model's output.","A natural next step is adding diffusion-weighted imaging or arterial-phase input to raise tumor detection beyond the current 76.6%; the trade-off in false positives is directly testable with the same evaluation protocol."],"forward_implications":["If these results generalize, 3D liver models can be produced for every patient scheduled for liver surgery, not only those treated at centers that can afford hours of manual segmentation.","Preserving vessel-tree topology means the segmentations can serve as a map for image-guided procedures, where vessel bifurcations are used as registration landmarks.","Cutting the segmentation step from several hours to about 15 minutes makes 3D planning compatible with normal clinical scheduling on a per-patient basis.","Automated tumor outlining adds a safety check: in this cohort it surfaced three sub-centimeter lesions initially missed by the radiologist.","The network handles common anatomical variations such as an absent gallbladder after cholecystectomy and regenerated anatomy after prior resection, so it does not fail at the edges of a typical liver-surgery population."],"supporting_citations":[{"why":"Supplies the nnU-Net framework used to train the segmentation model, the central architecture of the method.","marker":"[22]"},{"why":"Defines clDice, the topology-preserving loss used to maintain the connectivity of thin vessels.","marker":"[25]"},{"why":"Introduces the bootstrapped cross-entropy strategy for focusing training on hard voxels in thin structures.","marker":"[23]"},{"why":"Provides the 3D Slicer platform used to create the manual reference segmentations.","marker":"[21]"},{"why":"Prior hepatobiliary-phase MRI segmentation method including tumors and bile ducts; its reported Dice scores serve as the main comparison baseline.","marker":"[18]"},{"why":"Prior non-contrast T1 MRI vessel segmentation study; its lower vessel Dice scores contextualize the improvement reported here.","marker":"[17]"},{"why":"Prior contrast-enhanced MRI vasculature segmentation method; provides another comparison point for portal and hepatic vein accuracy.","marker":"[26]"},{"why":"Prior tumor-detection method with diffusion-weighted MRI; discussed as a route to improve the model's tumor detection rate.","marker":"[27]"}],"fun_headline_variants":["15-minute MRI liver map for surgery, plus tumor finds","One MRI, 15 minutes: automated liver surgery roadmap","nnU-Net liver anatomy: 15-minute planning, extra tumor detection","Automated liver segmentation reduced to 15 minutes per case","MRI liver model catches three tumors in minutes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the manual outlines used for training and evaluation are correct for every structure; if those outlines are noisy, especially for thin vessels and small tumors, the reported overlap scores overstate how precisely the model matches true anatomy.","fun_headline_variants_meta":{"raw":{"variants":["15-minute MRI liver map for surgery, plus tumor finds","One MRI, 15 minutes: automated liver surgery roadmap","nnU-Net liver anatomy: 15-minute planning, extra tumor detection","Automated liver segmentation reduced to 15 minutes per case","MRI liver model catches three tumors in minutes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000315,"raw_usage":{"total_tokens":1911,"prompt_tokens":1193,"completion_tokens":718,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":809,"completion_tokens_details":{"reasoning_tokens":636}},"tokens_in":809,"tokens_out":718,"duration_ms":7915,"temperature":1.0,"reasoning_tokens":636,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:10:48.352769+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Segment the same 18 test scans with two or more independent expert annotators and compute their pairwise overlap scores for portal vein, hepatic vein, biliary tree, and tumors; if expert-to-expert agreement is no better than the model's scores on those structures, the model is operating at the limit of the reference standard rather than demonstrating true anatomical accuracy.","supporting_citations":[{"cited_title":"nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation","cited_arxiv_id":null,"evidence_quote":"Supplies the nnU-Net framework used to train the segmentation model, the central architecture of the method."},{"cited_title":"CLDICE - A novel topology-preserving loss function for tubular structure segmentation","cited_arxiv_id":null,"evidence_quote":"Defines clDice, the topology-preserving loss used to maintain the connectivity of thin vessels."},{"cited_title":"Deep interactive thin object selection","cited_arxiv_id":null,"evidence_quote":"Introduces the bootstrapped cross-entropy strategy for focusing training on hard voxels in thin structures."},{"cited_title":"Automated 3D liver segmentation from hepatobiliary phase MRI for enhanced preoperative planning","cited_arxiv_id":null,"evidence_quote":"Prior hepatobiliary-phase MRI segmentation method including tumors and bile ducts; its reported Dice scores serve as the main comparison baseline."},{"cited_title":"Convolutional neural network for automated segmentation of the liver and its vessels on non-contrast T1 vibe Dixon acquisitions","cited_arxiv_id":null,"evidence_quote":"Prior non-contrast T1 MRI vessel segmentation study; its lower vessel Dice scores contextualize the improvement reported here."},{"cited_title":"Optimization of hepatic vasculature segmentation from contrast-enhanced MRI, exploring two 3D Unet modifications and various loss functions","cited_arxiv_id":null,"evidence_quote":"Prior contrast-enhanced MRI vasculature segmentation method; provides another comparison point for portal and hepatic vein accuracy."},{"cited_title":"Liver segmentation and metastases detection in MR images using convolutional neural networks","cited_arxiv_id":null,"evidence_quote":"Prior tumor-detection method with diffusion-weighted MRI; discussed as a route to improve the model's tumor detection rate."}],"review_version":2}