{"id":"ce6b5cf1-ec53-412e-93ef-091564278b80","arxiv_id":"2607.05325","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":12,"one_line_summary":"A centre-guided Attention U-Net with physics-guided synthetic mimics achieves best local-comparison lesion-level F1 for CMB detection on VALDO and external AIBL SWI.","lead":"CenSynCMB combines a centre-map-guided detector with physics-based synthetic data to better detect cerebral microbleeds on MRI. It improves lesion-level F1 on the VALDO benchmark and external AIBL SWI cohort, though burden calibration remains limited.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The primary VALDO F1 improvement (p=0.020) rests on a single 5-fold CV partition of 72 subjects with no individual component reaching ablation significance; the result's sensitivity to partition choice is untested.","rationale":"The reader correctly identified the non-significant ablation results and the synthesis module's limitations as concerns, but focused primarily on the synthesis module's physics approximation. I think the more load-bearing concern is the robustness of the primary VALDO F1 result itself: with n=72, a single CV seed, overlapping bootstrap CIs, and no component reaching individual significance, the headline p=0.020 could be partition-sensitive. The AIBL result (n=370, p=0.0016) is more robust but is an averaged-over-folds evaluation with a substantial FP tradeoff. These concerns reinforce the reader's CONDITIONAL verdict rather than changing it — the paper is honest about its limitations, the evaluation protocol is fair and leakage-free, and the external validation on AIBL provides genuine additional evidence. But the evidence does not yet support a strong claim about the individual design components or the stability of the in-domain result. The paper remains a solid engineering contribution with transparent reporting, appropriately rated as conditional pending multi-seed validation and code release.","tokens_in":13528,"tokens_out":5073,"duration_ms":180962,"concrete_test":"Re-run the 5-fold CV on VALDO with at least 5 different random seeds for the partition. For each seed, recompute the paired Wilcoxon p-value for lesion-level F1 between CenSynCMB and the best local comparator. If the p-value exceeds 0.05 in more than 1 of 5 seeds, the headline VALDO F1 claim is partition-sensitive and should be reported as trend-level rather than significant. Separately, report per-fold (not averaged) AIBL metrics to show the spread across individual fold models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the F1 improvements on VALDO and AIBL are real and attributable to the proposed design. The AIBL result (n=370, p=0.0016) is reasonably robust, but the VALDO result (n=72, p=0.020) is based on a single random seed for the 5-fold CV partition (§IV-A: 'fixing the partitions with a single random seed'). With ~14 subjects per test fold, the per-subject metrics feeding the paired Wilcoxon test could vary substantially across partitions. The bootstrap CIs for CenSynCMB (74.3±8.8%) and the backbone alone (65.8±10.2%) overlap considerably, and no individual component reaches p<0.05 in ablation (centre-map p=0.0986, synthesis p>0.05 for all comparisons in Table IV). This means the overall improvement could reflect a combination of small, individually non-significant effects that happen to align on this particular partition. Without multi-seed CV or an independent held-out test set, the stability of the headline F1 gain on VALDO is uncertain. The AIBL external result partially mitigates this concern (larger n, significant p), but it comes with a precision tradeoff (FP/subject 2.4 vs 0.9 for DynUNet) and worse patient-level MAE (1.87 vs 0.93), so it supports candidate extraction but not calibrated burden estimation. Additionally, the AIBL metrics are reported as 'the mean over five fold checkpoints' (Table I caption), meaning all 5 fold models are applied to each AIBL subject and averaged — this is an ensemble-style evaluation that may overstate single-model deployment performance.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This paper proposes CenSynCMB, a centre-guided CMB detection framework combining an Attention U-Net backbone with auxiliary centre-map supervision, false-negative-driven reweighting, and fold-wise physics-guided synthesis of positive CMBs and labelled hard negatives (vessel-like and calcification-like mimics). The system is evaluated on VALDO Task 2 (5-fold CV, n=72) and externally on AIBL SWI (n=370) under a shared centroid-matching protocol with bootstrap CIs and paired Wilcoxon tests. The central claims are that CenSynCMB achieves the best local-comparison lesion-level F1 on VALDO (74.3%, p=0.020) and the highest local recall and F1 on AIBL (88.5%, p=0.0058; 65.0%, p=0.0016), supporting scalable CMB candidate extraction across GRE-to-SWI shifts. The evaluation protocol is clearly defined, comparison methods are appropriately separated into leakage-free reimplemented models and pretrained references, and synthesis priors are estimated only from training-fold subjects.","tokens_in":14485,"tokens_out":1222,"duration_ms":161061,"significance":"The paper addresses a clinically relevant problem (automated CMB detection for CSVD and ARIA-H monitoring) with a well-motivated design. Strengths include: (1) a clearly defined centroid-matching evaluation protocol with bootstrap CIs and paired statistical tests; (2) fold-wise synthesis that prevents test leakage, with priors estimated only from training subjects; (3) transparent separation of reimplemented (R) and pretrained (L) comparison methods; (4) honest reporting of ablation results including non-significant p-values; (5) external validation on a different sequence (GRE→SWI) and scanner. The physics-guided synthesis module is a reasonable approach to the hard-negative scarcity problem, and the magnitude-only static-dephasing limitation is openly acknowledged. The patient-level burden analysis (Table II) is a valuable addition that honestly delineates the boundary between candidate extraction and calibrated burden estimation.","major_comments":[{"comment":"§IV-A, Table I caption: The VALDO 5-fold CV uses a single random seed for partition assignment (n=72, ~14 subjects per test fold). The headline F1 gain (74.3% vs 65.8% backbone, p=0.020) has substantially overlapping bootstrap CIs (±8.8% vs ±10.2%), and no individual ablation component reaches p<0.05 (Table III: centre-map p=0.0986; Table IV: all synthesis comparisons p>0.05). Without multi-seed CV or an independent held-out test set, the stability of the VALDO F1 improvement across partition choices is untested. The AIBL external result (n=370, p=0.0016) partially mitigates this, but AIBL is evaluated as the mean over five fold checkpoints (Table I caption), which is an ensemble-style evaluation that may overstate single-model performance. The authors should either (a) run multi-seed CV on VALDO and report the mean and standard deviation of the F1 metric across seeds, or (b) clearly add","section":null}],"minor_comments":[{"comment":"§III-C, Eq. (2): The focal centre-map loss Lctr uses the notation (Cx + epsilon_c)^gamma_c, but it is unclear whether this exponent applies to the target Cx or to the full expression. Clarifying the grouping would help readers.","section":null},{"comment":"§III-D: The field-strength and echo-time ranges (1.5–3.0 T, 20–40 ms) are sampled when metadata is unavailable, but VALDO metadata includes scanner field strengths (1.5 T and 3 T). It would help to state how often metadata was actually unavailable and whether the sampled ranges match the VALDO acquisition distribution.","section":null},{"comment":"Table I: The asterisk marking convention (p<0.05 against best local comparator) is defined in the caption but the specific comparator for each asterisked metric is not named. Adding the comparator name or a footnote would improve clarity.","section":null},{"comment":"Table II: CenSynCMB's VALDO MAE (1.81) is bolded as best, but the FRST-style detector achieves 2.00 with a much higher Pearson r (0.84 vs 0.77). The caption should clarify whether bolding is per-metric or overall, and the text should note this tradeoff.","section":null},{"comment":"§V-B: The Grad-CAM visualisation (Fig. 4) is discussed qualitatively but no quantitative evidence links the centre-map head to improved localisation beyond the ablation in Table III. A brief statement acknowledging this as qualitative evidence would be appropriate.","section":null},{"comment":"Reference [23] is a medRxiv preprint (2026); if a peer-reviewed version is available, it should be updated.","section":null},{"comment":"§I-C: The contributions list combines multiple techniques (centre-map, focal regression, uncertainty weighting, FN-driven reweighting) into a single bullet point. Separating the detector design contributions from the synthesis contribution would improve readability.","section":null}],"recommendation":"major_revision","confidential_remarks":"The reader's stress-test concern about single-seed VALDO partition instability is well-founded and is the primary basis for the major revision recommendation. The AIBL ensemble-style evaluation is a secondary concern that should also be addressed. The paper is otherwise well-conducted with honest reporting of limitations; the synthesis module's lack of ablation significance is openly acknowledged and does not by itself warrant rejection, but the combination of small VALDO sample, single-seed partition, overlapping CIs, and no significant ablation components means the central VALDO claim needs additional support. The AIBL result is more robust but comes with the precision tradeoff and ensemble evaluation caveats noted above."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the careful reading and constructive statistical critique. The concern about single-seed partition stability on the small VALDO cohort is well-taken, and we will address it through both additional experiments and clearer reporting.","responses":[{"response":"The referee is correct that the current VALDO evaluation relies on a single partition seed, and that the overlapping bootstrap CIs and non-significant ablation p-values mean the VALDO improvement cannot be considered definitively stable across partition choices. We acknowledge this limitation honestly. We will run multi-seed 5-fold cross-validation on VALDO (at least 5 distinct partition seeds) and report the mean and standard deviation of F1 (and other lesion-level metrics) across seeds. This will directly test whether the F1 improvement over the backbone and over local comparators is stable across partition choices. We will also report the proportion of seeds in which CenSynCMB outperforms the best local comparator, providing an intuitive stability measure. If the multi-seed results show that the improvement is not robust, we will revise the headline claim accordingly and reframe the VALDO result as trend-level rather than confirmatory. We agree this is necessary before the community can rely on the reported F1 gain.","revision_made":"yes","referee_comment":"VALDO 5-fold CV uses a single random seed for partition assignment (n=72, ~14 subjects per test fold). Headline F1 gain has substantially overlapping bootstrap CIs, and no individual ablation component reaches p<0.05. Without multi-seed CV or an independent held-out test set, the stability of the VALDO F1 improvement across partition choices is untested."},{"response":"This is a fair and important observation. The current AIBL protocol averages predictions across five fold checkpoints, which is indeed an ensemble-style evaluation and does not reflect single-model deployment performance. We will address this in two ways. First, we will add a single-model AIBL evaluation: for each fold checkpoint, we will report AIBL metrics independently, and then report the mean and standard deviation across the five individual checkpoints. This will show the spread of single-model performance on AIBL. Second, we will clearly label the ensemble-style result as such in the table and text, so readers understand the distinction. We note that even under single-model evaluation, the AIBL cohort (n=370) is substantially larger than VALDO (n=72), so the external validation still provides meaningful evidence of cross-sequence transfer; however, the referee is right that the current presentation conflates single-model and ensemble performance, and we will correct this.","revision_made":"yes","referee_comment":"AIBL is evaluated as the mean over five fold checkpoints (Table I caption), which is an ensemble-style evaluation that may overstate single-model performance."},{"response":"We interpret option (b) as a request to clearly state the partition-stability limitation in the manuscript if multi-seed experiments are not feasible. We will in fact pursue option (a) as described above. Additionally, regardless of the multi-seed outcome, we will add an explicit limitation paragraph noting that VALDO is small (n=72), that the single-seed partition was the original design, and that ablation-level p-values did not reach conventional significance thresholds. We believe both actions together fully address the referee's concern.","revision_made":"yes","referee_comment":"The authors should either (a) run multi-seed CV on VALDO and report the mean and standard deviation of the F1 metric across seeds, or (b) clearly add [comment appears truncated]."}],"tokens_in":13136,"tokens_out":1175,"duration_ms":64044,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Short version: this is a competent, honestly-reported engineering paper on CMB detection. The main F1 gains on VALDO and AIBL are real enough to take seriously, but the ablations don't reach significance for any individual component, and the VALDO result rests on a single 5-fold partition of 72 subjects. The AIBL external result is the stronger evidence. I'd send it to a referee with a request for multi-seed CV and code release as conditions for acceptance.","headline":"Solid engineering for CMB detection; ablations are underpowered but the external AIBL result carries real weight.","tokens_in":14679,"tokens_out":158,"would_cite":true,"duration_ms":69635,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Centre maps and synthetic mimics cut microbleed detection errors","keywords":[],"falsifier":"If the synthetic mimics do not adequately represent the real false-positive distribution in external cohorts, the external recall gains on AIBL could be partially offset by the higher false-positive burden, undermining the practical utility of the detection improvement for burden estimation.","tokens_in":13723,"feed_emoji":"🧠","tokens_out":865,"duration_ms":105126,"temperature":0.7,"pith_summary":"CenSynCMB argues that two design choices fix the main failure modes of automated cerebral microbleed (CMB) detection: (1) an auxiliary centre-map head that trains the network to localise lesion centroids rather than voxel-level segmentation overlap, and (2) fold-wise physics-guided synthesis that inserts synthetic positive CMBs and labelled hard-negative mimics (vessel-like and calcification-like structures) into real MRI backgrounds. The centre-map head aligns the training signal with the centroid-matching criterion used in clinical evaluation, while the synthesis module supplies controlled negative evidence for the structures that most often generate false positives. Together, these mechanisms are shown to improve lesion-level F1 on the VALDO benchmark (74.3%, p=0.020) and on the external AIBL SWI cohort (65.0%, p=0.0016), where the model must transfer from T2*-GRE to susceptibility-weighted imaging. The external gains come with higher false-positive burden, and patient-level count calibration remains unreliable, so the authors frame the system as a high-throughput candidate-extraction layer rather than a finished burden-estimation tool.","feed_headline":"Centre maps and synthetic mimics cut microbleed detection errors","feed_subtitle":"A centroid-supervised detector with physics-guided hard-negative synthesis improves CMB detection across MRI sequences, but burden counting仍","key_machinery":"Centre-map supervision (Gaussian target around each lesion centroid), false-negative-driven crop reweighting (Tversky loss with α=0.1, β=0.9, plus per-lesion crop weights), and physics-guided synthesis (static-dephasing susceptibility model with dipole kernel convolution for positive CMBs; elongated anisotropic masks for vessel-like mimics; compact low-signal sources for calcification-like mimics).","core_discovery":"The paper's central claim is that coupling explicit centroid supervision with physics-guided synthesis of both positive lesions and labelled hard-negative mimics yields a better balance of recall and precision for CMB detection than segmentation-backbone or cascaded-detector approaches, across both in-domain and cross-sequence MRI settings. The centre-map head shifts the training objective toward the detection target (lesion centroids), and the synthesis module provides the scarce negative examples (vessel cross-sections, calcification-like foci) that real datasets lack. The combination improves lesion-level F1 on both VALDO and external AIBL, though the external recall gains are partially抵消","pith_inferences":[],"forward_implications":["CMB candidate extraction at scale becomes feasible for large unlabelled MRI cohorts, enabling semi-automated triage where a sensitive detector flags candidates and a human reader confirms or rejects them.","The centre-map supervision principle could transfer to other compact, sparse lesion types in neuroimaging (e.g., lacunes, enlarged perivascular spaces) where the evaluation target is a centroid rather than a segmentation mask.","The gap between lesion-level detection gains and patient-level burden calibration suggests that downstream cohort studies will need cohort-specific threshold tuning or uncertainty-aware review layers before using detector outputs as clinical counts.","The physics-guided synthesis approach could be extended to other susceptibility-sensitive markers or to phase/QSM-based rendering if the magnitude-only limitation is addressed."],"fun_headline_variants":["Centroid supervision and physics-guided synthesis improve microbleed detection","Physics-guided mimics improve CMB detection across MRI sequences","Centre-map head plus synthetic hard negatives lift microbleed F1 on two cohorts","Synthesizing lesions and mimics cuts microbleed false positives and negatives","Centroid-guided detector with fold-wise synthesis reaches 74% F1 for CMBs"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The physics-guided synthesis module uses a magnitude-only static-dephasing susceptibility approximation that cannot distinguish calcification from blood products by susceptibility sign, and the synthesis ablation did not reach statistical significance (p>0.05 for all comparisons), so the claim that synthetic mimics meaningfully reduce false positives rests on trend-level directional evidence only.","fun_headline_variants_meta":{"raw":{"variants":["Centroid supervision and physics-guided synthesis improve microbleed detection","Physics-guided mimics improve CMB detection across MRI sequences","Centre-map head plus synthetic hard negatives lift microbleed F1 on two cohorts","Synthesizing lesions and mimics cuts microbleed false positives and negatives","Centroid-guided detector with fold-wise synthesis reaches 74% F1 for CMBs","Physics-guided positive and hard-negative synthesis balances CMB precision-recall","Centre maps shift CMB detection toward lesion targets over voxel masks","Fold-safe synthesis of CMBs and mimics lifts external recall to 88.5%","Auxiliary centroid supervision narrows the CMB precision-recall gap on two datasets","Mimic-aware synthesis lets a single U-Net match cascaded CMB detectors"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1340,"prompt_tokens":545,"completion_tokens":795,"prompt_tokens_details":null},"tokens_in":545,"tokens_out":795,"duration_ms":42103,"temperature":1.0,"reasoning_tokens":614,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-07T17:05:00.303132+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the synthetic mimics do not adequately represent the real false-positive distribution in external cohorts, the external recall gains on AIBL could be partially offset by the higher false-positive burden, undermining the practical utility of the detection improvement for burden estimation.","supporting_citations":[],"review_version":1}