{"id":"bf26ccf4-4110-477c-8a8c-36717df95917","arxiv_id":"2411.15388","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A SynthSeg-based deep learning method segments the claustrum from ultra-high-resolution ex vivo MRI and generalizes to typical in vivo scans across contrasts and resolutions.","lead":"This paper trains an AI to outline the claustrum, a thin brain region, using ultra-high-resolution MRI scans of donated brains, then shows the same tool works on ordinary lower-resolution scans. The method generalizes across different MRI contrasts and scanners, which could make studies of this under-explored structure feasible on large existing datasets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In vivo accuracy claim rests on an unvalidated QC proxy; the paper never shows QC correlates with native-space Dice.","rationale":"The reader's weakest_assumption identified exactly this issue: the in vivo evaluation depends on the QC score and its assumptions about SynthMorph registration and similarity-to-labels as a proxy for accuracy. My stress-test confirms that this is the most load-bearing concern for the central claim. The paper's direct evidence for accuracy at ultra-high resolution (CV Dice 0.632) is moderate, but it is at least a direct overlap with manual labels; the 'accurate' claim for typical in vivo resolution rests entirely on the indirect QC metric. The paper explicitly discourages interpreting QC as accuracy, yet uses it for model selection and performance reporting. This is a measurement-validity gap, not a disagreement with consensus. The proposed concrete test would settle it by correlating QC with known native-space Dice on the existing CV data, requiring no new manual labeling. Since the reader's conditional verdict already accounts for this concern, my read does not change the verdict. I would maintain CONDITIONAL: the authors should either provide direct in vivo validation (e.g., manual labels on a small subset) or at minimum demonstrate that QC tracks native-space accuracy in the CV data where ground truth exists.","tokens_in":27861,"tokens_out":4734,"duration_ms":46266,"concrete_test":"Use the 18 high-resolution CV cases, for which native-space Dice against manual labels is already reported in Table 2. For each CV test case, compute the QC score exactly as in Sec. 3.4 (max Dice after SynthMorph nonlinear registration to MNI152 against the 18 manual labels). Then calculate the Spearman correlation between native-space Dice and QC across the 18 cases. If the correlation is weak (e.g., Spearman rho < 0.7) or if low-Dice cases do not consistently yield low QC, the QC-based epoch selection and IXI in vivo results would not support the accuracy claim; if the correlation is strong, the QC proxy is validated and the concern is resolved. This test reuses existing data and directly assesses the measurement tool underlying the in vivo conclusions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the method is 'accurate' on typical in vivo scans is supported, in the IXI experiments and epoch selection, almost entirely by the QC score defined in Sec. 3.4: the maximum Dice between the automatic segmentation (nonlinearly registered to MNI152 via SynthMorph) and each of the 18 manual labels in MNI space. The paper itself states that QC 'should not be interpreted as a direct measure of segmentation accuracy' and that even a perfect segmentation will not get QC = 1. Yet QC is used to select the final epoch (Sec. 3.4) and to report IXI performance (Sec. 5). The validity of this proxy depends on two untested assumptions: (a) SynthMorph accurately aligns new images and the 18 labels into a common space, and (b) similarity to the training labels in that space implies accurate segmentation in native space. Because the 18 labels are the same labels used for training and for building the atlas, QC may reward segmentations that are anatomically 'typical' rather than individually correct; a segmentation that misses the true claustrum but lands on the population-typical location could receive a high QC. No evidence is provided that QC correlates with true native-space Dice. Without such validation, the IXI statistics and the epoch selected by QC cannot substantiate the 'accurate' part of the claim for standard-resolution in vivo data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a SynthSeg-based 3D U-Net for claustrum segmentation, trained on manual labels from 18 ultra-high-resolution hemispheres (mostly ex vivo) and evaluated with 6-fold cross-validation (average Dice 0.632), a downsampling simulation, and in vivo experiments on the IXI, Miriad, and FSM datasets. The authors claim this is the first accurate, contrast- and resolution-agnostic method for ultra-high-resolution claustrum segmentation, and they release the method on GitHub and in FreeSurfer.","tokens_in":28034,"tokens_out":5025,"duration_ms":43678,"significance":"If the central claims held, this would be a practically valuable tool for claustrum research, especially because existing automatic methods are limited and the method is released with FreeSurfer integration and public in vivo evaluations. The paper also deserves credit for using a contrast-synthesis training strategy, for testing on multiple independent datasets, and for reporting test-retest and cross-modality consistency. However, the evidence for the 'accurate' part of the claim is incomplete: the CV Dice is below the reported inter-rater Dice, the in vivo accuracy argument depends on an unvalidated QC proxy, the resolution-robustness experiment stays within the training augmentation range, and the IXI volume inflation is unexplained. The study is therefore a solid engineering contribution but needs additional validation before the abstract-level claim is justified.","major_comments":[{"comment":"The in vivo accuracy claim rests on the QC score defined in Sec. 3.4, yet the paper itself states that QC 'should not be interpreted as a direct measure of segmentation accuracy' and that even a perfect segmentation will not reach QC = 1. Because the QC score is the maximum Dice against the same 18 manual labels used for training, after nonlinear registration to MNI152, it can reward segmentations that fall at the population-typical claustrum location rather than at the individual's true claustrum, and no experiment in the paper shows that QC correlates with native-space Dice. Since the final epoch is selected on a 20-subject IXI subset using this score and the IXI 'never critically failed' conclusion is based on QC plus visual inspection of extremes, the load-bearing 'accurate in vivo' part of the abstract claim is not yet supported. Please validate QC against native-space Dice on the 18 CV cases (or another labeled set) and report the relationship, or explicitly downgrade the in vivo claim to robustness/plausibility.","section":"Sec. 3.4 and Sec. 5 (IXI experiments and epoch selection)"},{"comment":"The resolution-robustness experiment downsamples the 18 high-resolution hemispheres to 0.4-3.5 mm and reports graceful Dice degradation, but the training procedure already simulates downsampling with resolutions sampled from U(0.35,3.5) isotropically and U(0.35,5) anisotropically. The tested range is therefore almost entirely inside the training augmentation envelope, and the simulated images inherit the high SNR and ex vivo contrast of the source scans rather than the noise and partial-volume properties of native ~1 mm in vivo acquisitions. This experiment supports robustness within the training distribution but does not by itself establish resolution-agnostic accuracy at typical in vivo resolutions; please either test on native-resolution in vivo images with manual labels or restrict the claim accordingly.","section":"Sec. 3.3 and Fig. 9"},{"comment":"The IXI automatic claustrum volumes average 1,793.92 ± 259.16 mm3 versus 1,253.05 ± 283.79 mm3 for the manual labels, a ~43% inflation that the paper leaves unexplained ('It is not clear why the IXI volumes are so much higher'). This unexplained systematic bias undermines the use of the method for quantitative in vivo volumetry and should be addressed (e.g., by validating against manual labels on native-resolution scans or by modeling partial-volume effects) or explicitly listed as a limitation of quantitative accuracy.","section":"Sec. 5 (IXI volumes)"},{"comment":"The primary ultra-high-resolution accuracy evidence is CV Dice 0.632 ± 0.061, which is substantially below the inter-rater Dice 0.805 ± 0.018 reported for the seven shared samples. The discussion acknowledges this gap but does not establish that 0.632 constitutes 'accurate' segmentation rather than moderate agreement; since the abstract's first claim is accuracy at ultra-high resolution, please provide a more direct argument (e.g., error analysis, comparison with a baseline on the same labels, or a stated acceptability threshold) or soften the claim.","section":"Sec. 5 and Table 2"}],"minor_comments":[{"comment":"The URL 'https://github.com/chiara-mauri/claustrum segmentation' contains a space; please provide the correct link.","section":"Abstract and Section 1"},{"comment":"The caption says 'ev vivo' and should read 'ex vivo'.","section":"Fig. 3 caption"},{"comment":"The text refers to 'Vichow-Robins spaces' and should read 'Virchow-Robin spaces'.","section":"Sec. 5"},{"comment":"The inline equations for Dice, IoU, TPR, FDR, and volumetric similarity are garbled by line wrapping; please typeset them as display equations.","section":"Appendix A"},{"comment":"The '?' entries for sample 15 are not explained; please add a footnote stating that postmortem interval and brain weight were unavailable.","section":"Table 1"},{"comment":"The notation FoV/FOV is used inconsistently (e.g., Sec. 3.2 'FoV' versus Fig. 4 'FOV'); please unify.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the QC proxy is well-founded and is the main reason for the recommendation. The paper is potentially salvageable with a native-space Dice validation on the CV folds or a clear limitation statement. I would not reject, but the abstract-level claim of 'accurate' segmentation needs to be either supported by such validation or proportionately softened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a useful, honest paper. The authors adapt SynthSeg to claustrum segmentation using 18 new ultra-high-resolution manual labels, add voxel-wise Gaussian noise to the synthetic intensity augmentation, and release code plus data. The evaluation is broad: 6-fold CV, simulated downsampling, test-retest on Miriad, cross-modality on FSM, and a 581-subject IXI pass. They report CV Dice 0.632 against inter-rater 0.805 without spin, and they explicitly discuss the gap. The comparison against Casamitjana et al. and Albishri et al. is fair and useful.\n\nThe main soft spot is the QC score. As the stress-test note says, QC is max Dice against the 18 training labels after SynthMorph to MNI, and the paper uses it for epoch selection and as the primary evidence that in vivo segmentation is accurate. The authors do state that QC 'should not be interpreted as a direct measure of segmentation accuracy,' so they are not hiding the limitation. But they still lean on it, and the concern is real: a segmentation that lands on the population-typical claustrum location could get a high QC even if it missed the individual's actual claustrum. No evidence is shown that QC correlates with native-space Dice. I would not call this fatal, because the paper also does visual inspection of the worst QC cases and reports test-retest Dice 0.781 and cross-modality Dice 0.696-0.809, which are direct. But the 'accurate on typical in vivo scans' claim is weaker than the abstract implies.\n\nTwo smaller issues. The 20 IXI subjects used for epoch selection are included in the reported IXI statistics; that should be fixed. The ~43% volume inflation on IXI is reported as unexplained; it may well be partial volume effects, but it makes one wonder whether the method is systematically over-segmenting on 1mm data. The resolution-robustness experiment simulates downsampling within the training augmentation range, so it is less of an external test than it looks.\n\nThis paper deserves peer review. It is a practical contribution with released code and data, and the authors are transparent about what they did and where the numbers come from. My recommendation: send it out, ask for a direct in vivo accuracy check if any manual labels exist, re-run IXI stats excluding the validation subset, and soften the 'accurate' claim to something like 'reliable and robust' until that check is done.","headline":"A practical, honestly-reported SynthSeg adaptation for claustrum segmentation with released code; the in vivo accuracy claim leans on an unvalidated QC proxy, but the authors say so themselves.","tokens_in":28679,"tokens_out":2464,"would_cite":true,"duration_ms":22067,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims to provide the first automatic claustrum segmentation method that is accurate at ultra-high resolution and robust to changes in contrast and resolution, trained on synthetic images and manual labels from…","keywords":["claustrum","segmentation","ultra-high-resolution MRI","synthetic images","contrast-agnostic","resolution-agnostic","convolutional neural network","in vivo MRI"],"falsifier":"Manually label a held-out set of in vivo T1-weighted scans at about 1 mm resolution (say 20-30 subjects) and run the released model; if the mean Dice against these manual labels is far below the reported quality-control estimate (~0.57) and below the reported cross-modal agreement (~0.7), then the in vivo accuracy claim fails.","tokens_in":27585,"feed_emoji":"🧠","tokens_out":7005,"duration_ms":61772,"temperature":0.7,"pith_summary":"The paper seeks to establish that a deep network trained on manual claustrum labels from 18 ultra-high-resolution scans (0.1-0.25 mm, mostly ex vivo) can segment the entire claustrum at 0.35 mm isotropic and, without retraining, also segment it in ordinary in vivo scans at roughly 1 mm resolution across multiple contrasts. The motivation is that the claustrum is a thin gray-matter band, barely visible at typical clinical resolutions, and existing automatic methods either cover only part of it, work within one dataset, or do not quantify accuracy. On its own high-resolution test cases the method reaches a cross-validated Dice score of 0.632, a mean surface distance of 0.458 mm, and a volumetric similarity of 0.867, while a separate inter-rater comparison gives 0.805, indicating the automatic result is below human agreement on a difficult structure. On in vivo data the method is reported to be stable across repeated scans (Dice 0.781) and across modalities (Dice 0.696-0.809 against T1-weighted segmentations). If these claims hold, researchers gain an off-the-shelf tool for studying a structure implicated in consciousness, attention, salience, and several neurological disorders.","feed_headline":"Claustrum segmentation now works across contrasts and resolutions","feed_subtitle":"A model trained on ultra-high-res ex vivo labels segments standard in vivo T1, T2, and qT1 scans without retraining.","key_machinery":"The load-bearing mechanism is label-conditioned synthetic image generation: a 3D U-Net is trained on pairs of heavily augmented label maps and intensity images synthesized from those labels with randomly sampled contrast and resolution, including anisotropic downsampling from 0.35 mm up to 5 mm and added voxel-wise Gaussian noise. Because the training intensities are synthetic, the network learns shape and context rather than scanner-specific intensity statistics, and the same weights can be applied to ex vivo and in vivo images of different contrasts. A contrast-insensitive affine registration into a standard space is used only to locate a cropping field of view around the claustrum; it plays no role in assigning labels.","core_discovery":"The central claim is that this is the first accurate automatic method for ultra-high-resolution claustrum segmentation that is robust to changes in contrast and resolution. The method uses a segmentation framework that requires only label maps for training: intensity images are synthesized on the fly with randomized contrast, bias field, noise, smoothing, and downsampling, so the network does not learn a fixed acquisition protocol. The authors manually labeled the claustrum in 18 ultra-high-resolution hemispheres, generated dense whole-field labels by adding surrounding structures, trained a 3D U-Net at 0.35 mm isotropic resolution, and then applied the same model to in vivo T1-weighted, T2-weighted, proton-density, and quantitative T1 scans at standard resolutions. They report that performance degrades gracefully when inputs are downsampled, remains stable in test-retest settings, and does not fail catastrophically across 581 subjects in an independent T1-weighted dataset.","pith_inferences":["If the registration-to-reference-label quality-control strategy is sound, the same approach could screen segmentations of other small, low-contrast structures whose manual labels are scarce, without requiring manual ground truth on every test image.","The resolution robustness down to about 1.4 mm suggests the method could be applied retrospectively to legacy datasets that lack ultra-high-resolution acquisitions, enabling large-scale claustrum morphometry.","Because the in vivo evaluation relies on similarity to training labels in a standard space rather than manual labels on the test scans, an independent test with manual in vivo ground truth would clarify how much of the reported generalization comes from the network itself versus the evaluation proxy."],"forward_implications":["A single trained model can segment the full dorsal and ventral claustrum in standard ~1 mm T1-weighted in vivo scans, including scans from different field strengths and manufacturers.","The same model transfers to T2-weighted, proton-density, and quantitative T1 images, so multimodal studies can use one segmentation pipeline without retraining.","Test-retest Dice of 0.781 supports use in longitudinal and clinical studies where scans are acquired weeks apart.","Released code and integration into a widely used neuroimaging software package would let other groups segment the claustrum and correct putamen overlabeling errors without building their own method."],"supporting_citations":[{"why":"Supplies the synthetic-image segmentation framework that makes contrast- and resolution-agnostic training possible.","marker":"Billot et al. (2023a)"},{"why":"Supplies the contrast-invariant registration method used to establish the field of view and to compute quality-control scores in MNI space.","marker":"Hoffmann et al. (2021)"},{"why":"Supplies the 3D U-Net architecture on which the segmentation network is built.","marker":"Ronneberger et al. (2015)"},{"why":"Provides a baseline deep-learning claustrum segmentation method; its pre-trained version scores 0.299 on the authors' in vivo labeled case versus 0.632 for the proposed method.","marker":"Albishri et al. (2022)"},{"why":"Provides a comparable manual claustrum labeling and volume estimates, including a direct volume comparison on case 13.","marker":"Coates and Zaretskaya (2024)"},{"why":"Supplies the Miriad test-retest dataset used to measure segmentation stability across repeated scans.","marker":"Malone et al. (2013)"},{"why":"Supplies the multimodal FSM dataset used to test generalization across T1-weighted, T2-weighted, proton-density, and quantitative T1 images.","marker":"Greve and Fischl (2024)"}],"fun_headline_variants":["Claustrum segmentation leaps across MRI contrasts and scales","First accurate ultra-high-res claustrum segmentation across contrasts","No retraining: claustrum model works on T1, T2, qT1, PD","Claustrum segmentation robust to any MRI acquisition","Single claustrum model handles ex vivo to in vivo, any contrast"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The in vivo performance claim rests on a quality-control score that assumes the nonlinear registration of a test brain into a standard space is accurate enough that looking like one of the 18 training labels after alignment means the segmentation is correct.","fun_headline_variants_meta":{"raw":{"variants":["Claustrum segmentation leaps across MRI contrasts and scales","First accurate ultra-high-res claustrum segmentation across contrasts","No retraining: claustrum model works on T1, T2, qT1, PD","Claustrum segmentation robust to any MRI acquisition","Single claustrum model handles ex vivo to in vivo, any contrast"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000581,"raw_usage":{"total_tokens":2802,"prompt_tokens":1077,"completion_tokens":1725,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":693,"completion_tokens_details":{"reasoning_tokens":1634}},"tokens_in":693,"tokens_out":1725,"duration_ms":12511,"temperature":1.0,"reasoning_tokens":1634,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:21:36.538606+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Manually label a held-out set of in vivo T1-weighted scans at about 1 mm resolution (say 20-30 subjects) and run the released model; if the mean Dice against these manual labels is far below the reported quality-control estimate (~0.57) and below the reported cross-modal agreement (~0.7), then the in vivo accuracy claim fails.","supporting_citations":[{"cited_title":"N., Iglesias, J","cited_arxiv_id":null,"evidence_quote":"Supplies the contrast-invariant registration method used to establish the field of view and to compute quality-control scores in MNI space."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 3D U-Net architecture on which the segmentation network is built."},{"cited_title":"and Zaretskaya, N","cited_arxiv_id":null,"evidence_quote":"Provides a comparable manual claustrum labeling and volume estimates, including a direct volume comparison on case 13."},{"cited_title":"B., Cash, D., Ridgway, G","cited_arxiv_id":null,"evidence_quote":"Supplies the Miriad test-retest dataset used to measure segmentation stability across repeated scans."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the multimodal FSM dataset used to test generalization across T1-weighted, T2-weighted, proton-density, and quantitative T1 images."}],"review_version":1}