REVIEW 2 major objections 1 minor 59 references
Specificity- and Calibration-Aware Breast Ultrasound Segmentation via Entropy-Guided Boundary Supervision
T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Entropy scaling of boundary penalties reduces false positives on lesion-free breast ultrasound images while keeping lesion segmentation accuracy unchanged.
desk verdict The entropy-weighted boundary loss delivers a statistically backed drop in false positives on BUSI's 20 no-lesion images while holding Dice steady, but the tiny single-dataset no-lesion set leaves generalization unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Entropy-guided boundary supervision loss that multiplies the contour penalty term by per-pixel predictive entropy and the ground-truth boundary map.
What would settle it
Re-evaluate the same method on an independent multi-center breast ultrasound collection and test whether the false-positive count on no-lesion cases remains significantly lower than the two baselines.
Extended reading notes
Core claim
Rather than weighting every boundary pixel equally, the proposed loss scales contour penalties by per-pixel predictive entropy and the ground-truth boundary map, concentrating gradient emphasis on lesion margin locations where the network remains uncertain. This single change suppresses false-positive activations in no-lesion images while preserving segmentation quality on lesion-containing images.
Load-bearing premise
The no-lesion images and training setup in the BUSI dataset represent the false-positive failure modes that occur in real clinical breast ultrasound workflows.
Editorial extensions
If this is right
- Dice scores on lesion images stay statistically indistinguishable from the no-boundary baseline.
- False-positive activations on no-lesion images drop from 14–19 cases to 5 cases out of 20.
- Post-hoc spatial temperature scaling can reduce expected calibration error from 0.0201 to 0.0095 without changing the segmentation masks.
- The entropy weighting and boundary supervision operate as complementary training-level and inference-level refinements inside a U-Net.
Reading between the lines
- The same loss modification could be tested on other ultrasound modalities where speckle and shadowing produce similar false positives.
- If the reduction holds, clinical pipelines might rely less on separate false-positive filtering stages.
- The approach leaves open whether entropy weighting needs dataset-specific retuning when scanner hardware or acquisition protocols change.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes modifying the training loss for U-Net breast ultrasound segmentation by scaling boundary penalties with per-pixel predictive entropy and the ground-truth boundary map. This is claimed to reduce false-positive activations on no-lesion images while preserving Dice on lesion images. On the BUSI dataset, Dice is statistically equivalent (0.7624 vs. 0.7616, Wilcoxon p=0.27) across 97 lesion test images, while false-positive images drop to 5/20 from 14/20 and 19/20 (McNemar p=0.012, 0.0005) on 20 no-lesion images; a post-hoc spatial temperature scaling further lowers expected calibration error from 0.0201 to 0.0095.
Significance. If the entropy-guided weighting generalizes, the method provides a targeted way to improve specificity and calibration in BUS segmentation without Dice degradation, using appropriate statistical tests (Wilcoxon, McNemar, Wilson intervals) on held-out data. This addresses a practical clinical failure mode and could be complementary to existing U-Net pipelines.
major comments (2)
- [Evaluation / Results (no-lesion subset)] The specificity claim rests on false-positive reduction observed in only 20 no-lesion BUSI images. This small hold-out and single-dataset evaluation (with no cross-scanner, multi-center, or sensitivity analysis on entropy threshold/spatial scaling) is load-bearing for the central claim that the approach addresses real-world false-positive modes; the lesion Dice equivalence offers no safeguard if artifact statistics shift.
- [Methods] Exact loss formulation, data splits, ablation details, and hyper-parameter choices for the entropy weighting are not fully specified in the provided description, preventing verification that the reported gains are not due to post-hoc tuning or limited generalization testing.
minor comments (1)
- [Abstract / Calibration step] The post-hoc spatial temperature scaling is presented as complementary but its interaction with the training loss should be clarified to avoid implying it is part of the core contribution.
Simulated Author's Rebuttal
We thank the referee for the detailed and constructive report. We respond to each major comment below, indicating where revisions will be made to improve clarity and address concerns about evaluation scope and reproducibility.
read point-by-point responses
-
Referee: [Evaluation / Results (no-lesion subset)] The specificity claim rests on false-positive reduction observed in only 20 no-lesion BUSI images. This small hold-out and single-dataset evaluation (with no cross-scanner, multi-center, or sensitivity analysis on entropy threshold/spatial scaling) is load-bearing for the central claim that the approach addresses real-world false-positive modes; the lesion Dice equivalence offers no safeguard if artifact statistics shift.
Authors: We acknowledge the limited size of the no-lesion subset (n=20) in BUSI, which is the full available test split for this class in the standard benchmark. The observed reduction (14/20 and 19/20 to 5/20) is supported by McNemar tests (p=0.012, 0.0005) and non-overlapping Wilson intervals, indicating a statistically and practically meaningful effect within this dataset. In revision we will add an explicit limitations paragraph discussing single-dataset constraints and the value of future multi-center studies. We will also include a sensitivity analysis varying the entropy threshold and spatial scaling factor to assess robustness. Cross-scanner validation is noted as future work as it requires new data. revision: partial
-
Referee: [Methods] Exact loss formulation, data splits, ablation details, and hyper-parameter choices for the entropy weighting are not fully specified in the provided description, preventing verification that the reported gains are not due to post-hoc tuning or limited generalization testing.
Authors: We agree that additional methodological detail is needed for full reproducibility. The revised manuscript will include the exact loss formulation (the per-pixel entropy-weighted boundary term multiplied by the ground-truth boundary map), the precise train/validation/test splits on BUSI, complete ablation tables comparing all variants, and the specific hyper-parameter values used for entropy scaling and temperature. These additions will allow independent verification that gains are not attributable to post-hoc choices. revision: yes
Circularity Check
No circularity; empirical claims rest on held-out test evaluation
full rationale
The paper introduces an entropy-weighted boundary loss term and reports Dice preservation plus specificity gains via direct ablation on the BUSI held-out sets (97 lesion images, 20 no-lesion images). All quantitative results are computed from model outputs on unseen data using standard metrics and statistical tests; no equations reduce a claimed prediction to a fitted input by construction, no self-citations carry load-bearing uniqueness arguments, and no ansatz or renaming is smuggled in. The derivation chain is therefore self-contained against external benchmarks.
Assumptions & free parameters
assumptions (2)
- domain assumption U-Net architecture is an appropriate base model for the task
- domain assumption BUSI dataset splits and no-lesion images are representative of clinical false-positive scenarios
Cite this review
Pith. "Pith review of Specificity- and Calibration-Aware Breast Ultrasound Segmentation via Entropy-Guided Boundary Supervision." pith.science (2026). https://pith.science/paper/J43LWMGK
@misc{pith2026260622308,
author = {Pith},
title = {Pith review of: Specificity- and Calibration-Aware Breast Ultrasound Segmentation via Entropy-Guided Boundary Supervision},
year = {2026},
howpublished = {\url{https://pith.science/paper/J43LWMGK}},
note = {Machine review of arXiv:2606.22308}
}
read the original abstract
Lesion segmentation in breast ultrasound involves two related challenges. In images with lesions, speckle noise, low tissue contrast, and posterior acoustic shadowing cause boundary leakage and incomplete contour delineation. In images without lesions, those same artifacts generate false-positive activations in regions resembling solid lesion tissue. This study addresses both failure modes through a single modification to the training objective. Rather than weighting every boundary pixel equally, the proposed loss scales contour penalties by per-pixel predictive entropy and the ground-truth boundary map, concentrating gradient emphasis on lesion margin locations where the network remains uncertain. The loss was evaluated on the BUSI dataset through a controlled ablation against two baselines: a model without boundary supervision and a model with uniformly weighted boundary binary cross-entropy. Across 97 lesion-containing test images, mean Dice scores were statistically indistinguishable between the proposed method and the no-boundary baseline (0.7624 versus 0.7616, paired Wilcoxon p = 0.27), confirming that lesion segmentation quality is preserved. The primary effect appears in specificity. False-positive activations on 20 no-lesion test images fell from 14 of 20 and 19 of 20 for the two baselines to 5 of 20 with the proposed approach (McNemar p = 0.012 and 0.0005). Non-overlapping Wilson 95% confidence intervals confirm the difference is both statistically significant and practically substantial. A post-hoc spatial temperature scaling step further reduced expected calibration error from 0.0201 to 0.0095 without altering segmentation masks. Entropy-guided boundary supervision and spatial calibration thus function as complementary training-level and inference-level refinements that improve specificity and probability reliability within a U-Net framework.
Figures
Reference graph
Works this paper leans on
-
[1]
Abdellaoui, S
C. Abdellaoui, S. Belkacem, and N. Messaoudi, ”Deep learning ar- chitectures for medical image segmentation: an organized analysis of CNN-based models and uses,” Bulletin of Electrical Engineering and Informatics, vol. 15, no. 1, pp. 424-437, 2026
2026
-
[2]
Thomas, M
C. Thomas, M. Byra, R. Marti, M. H. Yap, and R. Zwiggelaar, ”BUS- Set: A benchmark for quantitative evaluation of breast ultrasound segmentation networks with public datasets,” Medical Physics, vol. 50, no. 5, pp. 3223-3243, 2023
2023
-
[3]
Al-Dhabyani, M
W. Al-Dhabyani, M. Gomaa, H. Khaled, and A. Fahmy, ”Dataset of breast ultrasound images,” Data in Brief, vol. 28, p. 104863, 2020
2020
-
[4]
X. Xiao, J. Zhang, Y . Shao, et al., ”Deep learning-based medical ultrasound image and video segmentation methods: overview, frontiers, and challenges,” Sensors, vol. 25, no. 8, p. 2361, 2025
2025
-
[5]
Ronneberger, P
O. Ronneberger, P. Fischer, and T. Brox, ”U-Net: Convolutional net- works for biomedical image segmentation,” in MICCAI, 2015, pp. 234- 241
2015
-
[6]
Isensee, P
F. Isensee, P. F. Jaeger, S. A. A. Kohl, J. Petersen, and K. H. Maier-Hein, ”nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation,” Nature Methods, vol. 18, no. 2, pp. 203-211, 2021
2021
-
[7]
Sulaiman et al., ”Attention based U-Net model for breast cancer segmentation using BUSI dataset,” Scientific Reports, vol
A. Sulaiman et al., ”Attention based U-Net model for breast cancer segmentation using BUSI dataset,” Scientific Reports, vol. 14, p. 22422, 2024
2024
-
[8]
Yang et al., ”Multilevel perception boundary-guided network for breast lesion segmentation in ultrasound images,” Medical Physics, 2025
X. Yang et al., ”Multilevel perception boundary-guided network for breast lesion segmentation in ultrasound images,” Medical Physics, 2025
2025
Show all 59 references
-
[9]
Kervadec, J
H. Kervadec, J. Bouchtiba, C. Desrosiers, E. Granger, J. Dolz, and I. Ben Ayed, ”Boundary loss for highly unbalanced segmentation,” Medical Image Analysis, vol. 67, p. 101851, 2021
2021
-
[10]
F. Sun, Z. Luo, and S. Li, ”Boundary difference over union loss for medical image segmentation,” in MICCAI, 2023, pp. 292-301
2023
-
[11]
Xu et al., ”Boundary guidance network for medical image segmen- tation,” Scientific Reports, vol
R. Xu et al., ”Boundary guidance network for medical image segmen- tation,” Scientific Reports, vol. 14, p. 17345, 2024
2024
-
[12]
Kendall and Y
A. Kendall and Y . Gal, ”What uncertainties do we need in Bayesian deep learning for computer vision?” in NeurIPS, 2017, pp. 5574-5584
2017
-
[13]
Scalco et al., ”Uncertainty quantification in multi-class segmentation: comparison between Bayesian and non-Bayesian approaches in a clinical perspective,” Medical Physics, vol
E. Scalco et al., ”Uncertainty quantification in multi-class segmentation: comparison between Bayesian and non-Bayesian approaches in a clinical perspective,” Medical Physics, vol. 51, no. 8, pp. 5460-5474, 2024
2024
-
[14]
C. Guo, G. Pleiss, Y . Sun, and K. Q. Weinberger, ”On calibration of modern neural networks,” in ICML, 2017, pp. 1321-1330
2017
-
[15]
Z. Ding, X. Han, P. Liu, and M. Niethammer, ”Local temperature scaling for probability calibration,” in ICCV , 2021, pp. 6889-6899
2021
-
[16]
Vallez, G
N. Vallez, G. Bueno, O. Deniz, M. A. Rienda, and C. Pastor, ”BUS- UCLM: Breast ultrasound lesion segmentation dataset,” Scientific Data, vol. 12, no. 242, 2025
2025
-
[17]
Z. Zhou, M. M. Rahman Siddiquee, N. Tajbakhsh, and J. Liang, ”UNet++: A nested U-Net architecture for medical image segmentation,” in DLMIA/MICCAI Workshops, LNCS 11045, 2018, pp. 3-11
2018
-
[18]
Oktay et al., ”Attention U-Net: Learning where to look for the pancreas,” arXiv:1804.03999, 2018
O. Oktay et al., ”Attention U-Net: Learning where to look for the pancreas,” arXiv:1804.03999, 2018
2018 arXiv
-
[19]
Chen et al., ”TransUNet: Transformers make strong encoders for medical image segmentation,” arXiv:2102.04306, 2021
J. Chen et al., ”TransUNet: Transformers make strong encoders for medical image segmentation,” arXiv:2102.04306, 2021
2021 arXiv
-
[20]
Cao et al., ”Swin-UNet: Unet-like pure transformer for medical image segmentation,” in ECCV Workshops, LNCS 13803, 2023, pp
H. Cao et al., ”Swin-UNet: Unet-like pure transformer for medical image segmentation,” in ECCV Workshops, LNCS 13803, 2023, pp. 205-218
2023
-
[21]
Abboud, H
Z. Abboud, H. Lombaert, and S. Kadoury, ”Sparse Bayesian networks: efficient uncertainty quantification in medical image analysis,” in MIC- CAI, 2024
2024
-
[22]
Xie et al., ”Entropy-guided contrastive learning for semi-supervised medical image segmentation,” IET Image Processing, 2024
X. Xie et al., ”Entropy-guided contrastive learning for semi-supervised medical image segmentation,” IET Image Processing, 2024
2024
-
[23]
M. H. Yap, G. Pons, J. Marti, S. Ganau, M. Sentis, R. Zwiggelaar, A. K. Davison, and R. Marti, ”Automated breast ultrasound lesions detection using convolutional neural networks,” IEEE J. Biomed. Health Inform., vol. 22, no. 4, pp. 1218-1226, 2018
2018
-
[24]
Gomez-Flores, M
W. Gomez-Flores, M. J. Gregorio-Calas, and W. Coelho de Albuquerque Pereira, ”BUS-BRA: A breast ultrasound dataset for assessing computer- aided diagnosis systems,” Medical Physics, vol. 51, no. 4, pp. 3110-3123, 2024
2024
-
[25]
J. Ma, J. Chen, M. Ng, R. Huang, Y . Li, C. Li, X. Yang, and A. L. Martel, ”Loss odyssey in medical image segmentation,” Medical Image Analysis, vol. 71, p. 102035, 2021
2021
-
[26]
Karimi and S
D. Karimi and S. E. Salcudean, ”Reducing the Hausdorff distance in medical image segmentation with convolutional neural networks,” IEEE Trans. Med. Imaging, vol. 39, no. 2, pp. 499-513, 2020
2020
-
[27]
Jungo, F
A. Jungo, F. Balsiger, and M. Reyes, ”Analyzing the quality and challenges of uncertainty estimations for brain tumor segmentation,” Frontiers in Neuroscience, vol. 14, p. 282, 2020
2020
-
[28]
Mehrtash, W
A. Mehrtash, W. M. Wells, C. M. Tempany, P. Abolmaesumi, and T. Kapur, ”Confidence calibration and predictive uncertainty estimation for deep medical image segmentation,” IEEE Trans. Med. Imaging, vol. 39, no. 12, pp. 3868-3878, 2020
2020
-
[29]
K. He, X. Zhang, S. Ren, and J. Sun, ”Deep residual learning for image recognition,” in CVPR, 2016, pp. 770-778
2016
-
[30]
Loshchilov and F
I. Loshchilov and F. Hutter, ”Decoupled weight decay regularization,” in ICLR, 2019
2019
-
[31]
Gal and Z
Y . Gal and Z. Ghahramani, ”Dropout as a Bayesian approximation: Representing model uncertainty in deep learning,” in ICML, 2016, pp. 1050-1059
2016
-
[32]
Salehi, A
A. Salehi, A. Erdogmus, and A. Gholipour, ”Tversky loss function for image segmentation using 3D fully convolutional deep networks,” in MLMI/MICCAI, LNCS 10541, 2017, pp. 379-387
2017
-
[33]
A. L. Simpson et al., ”A large annotated medical image dataset for the development and evaluation of segmentation algorithms,” arXiv:1902.09063, 2019
1902 arXiv
-
[34]
G. W. Brier, ”Verification of forecasts expressed in terms of probability,” Monthly Weather Review, vol. 78, no. 1, pp. 1-3, 1950
1950
-
[35]
Rohlfing, ”Image similarity and tissue overlaps as surrogates for image registration accuracy: widely used but unreliable,” IEEE Trans
T. Rohlfing, ”Image similarity and tissue overlaps as surrogates for image registration accuracy: widely used but unreliable,” IEEE Trans. Med. Imaging, vol. 31, no. 2, pp. 153-163, 2012
2012
-
[36]
Al-Dhabyani, M
W. Al-Dhabyani, M. Gomaa, H. Khaled, and A. Fahmy, ”Deep learning approaches for data augmentation and classification of breast masses using ultrasound images,” International Journal of Advanced Computer Science and Applications, vol. 10, no. 5, pp. 618–627, 2019, doi: 10.1456...
2019 doi
-
[37]
A. T. Stavros, C. Thickman, C. L. Rapp, M. A. Dennis, S. H. Parker, and G. A. Sisney, ”Solid breast nodules: use of sonography to distinguish between benign and malignant lesions,” Radiology, vol. 196, no. 1, pp. 123-134, 1995
1995
-
[38]
C. M. Rumack, S. R. Wilson, J. W. Charboneau, and D. Levine, Diagnostic Ultrasound, 4th ed. Philadelphia, PA: Elsevier Mosby, 2011
2011
-
[39]
Wilcoxon, ”Individual comparisons by ranking methods,” Biometrics Bulletin, vol
F. Wilcoxon, ”Individual comparisons by ranking methods,” Biometrics Bulletin, vol. 1, no. 6, pp. 80-83, 1945
1945
-
[40]
McNemar, ”Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol
Q. McNemar, ”Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol. 12, no. 2, pp. 153-157, 1947
1947
-
[41]
E. B. Wilson, ”Probable inference, the law of succession, and statistical inference,” Journal of the American Statistical Association, vol. 22, no. 158, pp. 209-212, 1927
1927
-
[42]
Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed
J. Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates, 1988
1988
-
[43]
Shrivastava, A
A. Shrivastava, A. Gupta, and R. Girshick, ”Training region-based object detectors with online hard example mining,” in CVPR, 2016, pp. 761- 769
2016
-
[44]
T. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, ”Focal loss for dense object detection,” in ICCV , 2017, pp. 2980-2988
2017
-
[45]
E. A. Krupinski, ”Current perspectives in medical image perception,” Attention, Perception, and Psychophysics, vol. 72, no. 5, pp. 1205-1217, 2010
2010
-
[46]
Tajbakhsh et al., ”Embracing imperfect datasets: A review of deep learning solutions for medical image segmentation,” Medical Image Analysis, vol
N. Tajbakhsh et al., ”Embracing imperfect datasets: A review of deep learning solutions for medical image segmentation,” Medical Image Analysis, vol. 63, p. 101693, 2020
2020
-
[47]
Bouthillier, C
P. Bouthillier, C. Laurent, and P. Vincent, ”Unreproducible research is reproducible,” in ICML, 2019, pp. 725-734
2019
-
[48]
W. A. Berg et al., ”Diagnostic accuracy of mammography, clinical examination, US, and MR imaging in preoperative assessment of breast cancer,” Radiology, vol. 233, no. 3, pp. 830-849, 2004
2004
-
[49]
Dodge and N
J. Dodge and N. Gane, ”Show your work: Improved reporting of experimental results,” in EMNLP, 2019, pp. 2185-2194
2019
-
[50]
Fornasa, ”Ultrasound-based breast lesion segmentation for surgical planning,” in Medical Imaging: Image Processing, SPIE, 2019
M. Fornasa, ”Ultrasound-based breast lesion segmentation for surgical planning,” in Medical Imaging: Image Processing, SPIE, 2019
2019
-
[51]
V . Buda, M. Maki, and M. A. Mazurowski, ”A systematic study of the class imbalance problem in convolutional neural networks,” Neural Networks, vol. 106, pp. 249-259, 2018
2018
-
[52]
P. Luc, C. Couprie, S. Chintala, and J. Verbeek, ”Semantic segmentation using adversarial networks,” in NIPS Workshop on Adversarial Training, 2016
2016
-
[53]
S. Meng, H. Zhang, C. Li, and Y . Zheng, ”Acoustic shadow detection in ultrasound images using deep learning,” in IEEE International Ultra- sonics Symposium, 2020
2020
-
[54]
T. F. Cootes, C. J. Taylor, D. H. Cooper, and J. Graham, ”Active shape models – their training and application,” Computer Vision and Image Understanding, vol. 61, no. 1, pp. 38-59, 1995
1995
-
[55]
Avendi, A
S. Avendi, A. Kheradvar, and H. Jafarkhani, ”A combined deep-learning and deformable-model approach to fully automatic segmentation of the left ventricle in cardiac MRI,” Medical Image Analysis, vol. 30, pp. 108-119, 2016
2016
-
[56]
S. Liu, D. Johns, and A. Davison, ”End-to-end multi-task learning with attention,” in CVPR, 2019, pp. 1871-1880
2019
-
[57]
Ovadia et al., ”Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift,” in NeurIPS, 2019, pp
S. Ovadia et al., ”Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift,” in NeurIPS, 2019, pp. 13991- 14002
2019
-
[58]
Ranschaert, S
R. Ranschaert, S. Morozov, and P. Algra, Artificial Intelligence in Medical Imaging. Cham, Switzerland: Springer, 2019
2019
-
[59]
Food and Drug Administration, ”Artificial intelligence and machine learning in software as a medical device,” FDA Discussion Paper, 2021
U.S. Food and Drug Administration, ”Artificial intelligence and machine learning in software as a medical device,” FDA Discussion Paper, 2021. [Online]. Available: https://www.fda.gov/media/145022/download
2021
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.