{"id":"c7989a11-3abf-47d4-914b-54e6ac9a3cef","arxiv_id":"2606.31444","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Full 27-frame training best for low-contrast early phases in LA/LAA segmentation; physiologically selected subsets match later phases, with foreground normalization partially helping reduced sets.","lead":"This paper tests how choosing different numbers of time frames from dynamic heart CT scans affects computer segmentation of the left atrium and its appendage. A generalist might read it to see practical trade-offs between using all data versus smart subsets when training AI for time-varying medical images.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Single per-sequence annotation (propagated by registration) may inject systematic label noise that biases early-phase vs. filling-phase performance gaps","rationale":"The reader’s weakest_assumption pinpoints the identical vulnerability. Full-text access does not appear to add multi-annotator validation or registration-error metrics that would neutralize it, so the concern remains load-bearing and the UNVERDICTED status is appropriate.","tokens_in":1731,"tokens_out":322,"duration_ms":13970,"concrete_test":"Select 4–6 sequences; obtain independent manual segmentations on 3 low-contrast and 3 filling-phase frames per sequence; compute Dice between these and the registered single-annotation GT; if mean Dice < 0.85 (or Hausdorff > 3 mm) in the low-contrast subset, re-train the three nnUNet variants on the corrected labels and re-evaluate the reported performance deltas.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result—that full 27-frame training outperforms on low-contrast phases while the physiological subset matches later—rests on the same registered single annotation serving as GT for every frame and every training-set variant. If registration drift or contrast-dependent boundary ambiguity produces frame-specific label errors (especially in the early low-contrast regime where the full-set advantage is claimed), then the observed differences could be artifacts of how each training subset samples that noise rather than true effects of temporal diversity. The abstract acknowledges the “trade-off” but supplies no quantitative check on registration fidelity or inter-frame label consistency.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper evaluates temporal training-set designs for nnUNet segmentation of the left atrium (LA) and left atrial appendage (LAA) in dynamic contrast 4DCT. It compares a minimal two-frame set (standard practice), a physiologically selected subset, and the full 27-frame sequence, plus the effect of foreground-based normalization derived from the full dataset. The central empirical claim is that full-frame training yields the best performance in early low-contrast phases, the physiological subset matches from the filling phase onward, and normalization improves reduced datasets in low-contrast frames without fully closing the gap. The work highlights a trade-off between temporal diversity and label noise arising from single per-sequence annotations propagated by registration.","tokens_in":1871,"tokens_out":531,"duration_ms":16270,"significance":"If the ordering holds under proper controls, the results offer actionable guidance for efficient training of dynamic cardiac CT segmentations relevant to blood-stasis assessment in atrial fibrillation. The purely empirical, held-out comparison of training regimes is a strength; reproducible code or public data splits would further strengthen it.","major_comments":[{"comment":"Methods (annotation and registration subsection): the central performance ordering rests on a single manual annotation per registered sequence serving as ground truth for every temporal frame and every training-set variant. No quantitative validation of registration fidelity (e.g., landmark error, inter-frame Dice on propagated labels, or contrast-phase-specific boundary consistency) is reported. Systematic label noise that varies with contrast level could therefore artifactually inflate the reported advantage of the full 27-frame set in early phases.","section":"Methods"},{"comment":"Results (performance tables/figures): the abstract and summary state clear ordering claims, yet the provided text supplies neither patient counts, cross-validation scheme, nor statistical tests (paired t-tests or Wilcoxon with correction) on the reported metrics. Without these, it is impossible to judge whether the observed gaps exceed inter-patient variability or are driven by a few outlier cases.","section":"Results"}],"minor_comments":[{"comment":"Abstract: states performance ordering without any numerical values, patient numbers, or error bars; this should be supplemented with at least the key Dice or surface-distance figures.","section":"Abstract"},{"comment":"Notation: “foreground-based normalization” is introduced without an explicit equation or reference to the exact intensity statistics used; add a short methods paragraph or equation.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive review. Below we respond point-by-point to the major comments, indicating where revisions will be made.","responses":[{"response":"We agree that explicit quantitative registration validation is not reported. However, because the identical propagated labels serve as ground truth for every training-set variant, any contrast-dependent label noise affects all models equally; relative differences therefore reflect temporal diversity rather than differential noise. We will add a limitations paragraph discussing registration quality and its potential impact.","revision_made":"partial","referee_comment":"[Methods] Methods (annotation and registration subsection): the central performance ordering rests on a single manual annotation per registered sequence serving as ground truth for every temporal frame and every training-set variant. No quantitative validation of registration fidelity (e.g., landmark error, inter-frame Dice on propagated labels, or contrast-phase-specific boundary consistency) is reported. Systematic label noise that varies with contrast level could therefore artifactually inflate the reported advantage of the full 27-frame set in early phases."},{"response":"The revised manuscript will explicitly report the patient count, the cross-validation scheme, and the results of paired statistical tests (Wilcoxon signed-rank with Bonferroni correction) comparing the training regimes. These additions will allow readers to evaluate whether the reported ordering exceeds inter-patient variability.","revision_made":"yes","referee_comment":"[Results] Results (performance tables/figures): the abstract and summary state clear ordering claims, yet the provided text supplies neither patient counts, cross-validation scheme, nor statistical tests (paired t-tests or Wilcoxon with correction) on the reported metrics. Without these, it is impossible to judge whether the observed gaps exceed inter-patient variability or are driven by a few outlier cases."}],"tokens_in":1439,"tokens_out":382,"duration_ms":20549,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper finds that training nnUNet on all 27 frames of dynamic 4DCT gives the best results for left atrium and appendage segmentation in the early low-contrast phases. From the filling phase onward a physiologically selected subset performs comparably. Normalization parameters taken from the full dataset lift the smaller training sets in those early frames but do not close the gap entirely.\n\nThe concrete contribution is the direct empirical comparison of three temporal training designs on this exact task: the minimal two-frame set used in standard practice, a physiologically chosen subset, and the complete sequence. They also test foreground-based normalization. The setup addresses a real constraint in dynamic contrast CT, where only one annotation exists per registered sequence and contrast changes sharply over time.\n\nThe work is straightforward about the trade-off between temporal diversity for robustness and the risk of label noise from propagating a single annotation. That framing is useful for anyone building models on time-resolved cardiac CT.\n\nThe soft spot is the absence of any numbers or statistical details in the abstract, so the size and reliability of the reported ordering remain unclear. The stress-test point about registration drift or contrast-dependent boundary ambiguity is reasonable to raise; if the single ground-truth label is less accurate in low-contrast frames, the full-set advantage could partly reflect how each training variant samples that noise rather than a pure effect of frame count. The paper notes the trade-off but does not appear to include a quantitative check on registration fidelity or inter-frame label consistency.\n\nThis is aimed at researchers working on nnUNet or similar segmentation pipelines for dynamic cardiac CT. A reader facing the same annotation-efficiency problem would get practical guidance from the comparison.\n\nIt deserves peer review so the quantitative results, patient numbers, and any registration validation can be examined.","headline":"Full-frame training edges out subsets in early low-contrast phases for LA/LAA segmentation but a physiological subset matches later, though single-annotation label noise via registration could be inflating that gap.","tokens_in":2355,"tokens_out":438,"would_cite":false,"duration_ms":26752,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Training on full 27-frame sequences outperforms reduced sets for early low-contrast phases in LA and LAA segmentation from dynamic 4DCT.","keywords":["left atrium segmentation","left atrial appendage","dynamic 4DCT","temporal training strategies","nnUNet","contrast-enhanced CT","atrial fibrillation"],"falsifier":"Independent per-frame ground-truth annotations on a held-out test sequence, followed by retraining the three strategies and direct comparison of Dice or surface-distance metrics in the early low-contrast phases.","tokens_in":2658,"feed_emoji":"🫀","tokens_out":696,"duration_ms":21229,"temperature":0.7,"pith_summary":"The paper tests three temporal training-set designs for nnUNet segmentation of the left atrium and left atrial appendage in dynamic contrast-enhanced 4DCT: a minimal two-frame set, a physiologically selected subset, and the complete 27-frame sequence. Full-frame training delivers the strongest results in the initial low-contrast phases, while the physiologically selected subset reaches comparable accuracy once contrast filling begins. Foreground-based normalization taken from the full dataset boosts the reduced sets in difficult early frames but leaves a remaining performance gap. The work addresses the practical tension between needing temporal diversity for robustness and avoiding label noise from single annotations propagated across registered frames. These design choices directly affect the feasibility of time-resolved contrast analysis for applications such as blood-stasis assessment in atrial fibrillation.","feed_headline":"Full-frame training beats subsets for early low-contrast LA segmentation","feed_subtitle":"Physiologically selected frames match performance after filling begins, and full-set normalization helps reduced datasets in dynamic 4DCT.","key_machinery":"Temporal training-set design comparing minimal two-frame, physiologically selected, and full 27-frame datasets for nnUNet segmentation, together with foreground-based normalization derived from the complete sequence.","core_discovery":"The authors establish that training nnUNet models on the full 27-frame dynamic sequences yields the best segmentation performance in early low-contrast phases of the left atrium and left atrial appendage. A physiologically selected subset of frames achieves comparable performance from the filling phase onward. Applying normalization parameters derived from the full dataset improves performance of the reduced datasets in low-contrast frames but does not fully close the gap to the full-set model.","pith_inferences":["Reduced frame sets could lower annotation cost for clinical pipelines focused on post-filling phases without major accuracy loss.","The same temporal-design logic may apply to other dynamic contrast modalities where label propagation from one frame is common.","Explicit per-frame re-annotation experiments would quantify how much label noise currently limits the reduced-set approaches."],"forward_implications":["Full-frame training is required to achieve optimal robustness during early low-contrast phases.","Physiologically selected frames provide a practical trade-off for segmentation once contrast filling has started.","Normalization parameters from the full dataset partially compensate for smaller training sets but do not eliminate the advantage of temporal diversity.","Downstream time-resolved contrast analysis benefits when training data include the full range of temporal contrast states."],"fun_headline_variants":["Full 27-frame training improves early low-contrast LA segmentation","Physiologically selected frames match full after filling begins","Full dataset normalization helps reduced datasets in low-contrast frames","Temporal diversity affects nnUNet LA segmentation in dynamic 4DCT"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"A single annotation per registered sequence provides sufficient ground truth across all temporal frames without introducing label noise that systematically biases the comparison between training-set designs.","fun_headline_variants_meta":{"raw":{"variants":["Full 27-frame training improves early low-contrast LA segmentation","Physiologically selected frames match full after filling begins","Full dataset normalization helps reduced datasets in low-contrast frames","Temporal diversity affects nnUNet LA segmentation in dynamic 4DCT"]},"model":"grok-4.3","cost_usd":0.006785,"raw_usage":{"total_tokens":3168,"prompt_tokens":693,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":67849500,"prompt_tokens_details":{"text_tokens":693,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2409,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":693,"tokens_out":66,"duration_ms":18274,"temperature":1.0,"reasoning_tokens":2409,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T06:22:03.432968+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Independent per-frame ground-truth annotations on a held-out test sequence, followed by retraining the three strategies and direct comparison of Dice or surface-distance metrics in the early low-contrast phases.","supporting_citations":[],"review_version":1}