Pith. sign in

REVIEW 3 major objections 5 minor 59 references

Hybrid Attention Network for Accurate Breast Tumor Segmentation in Ultrasound Images

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A hybrid attention network that fuses transformer-style and spatial attention at the bottleneck reports 97.28 Dice and 94.75 Jaccard on BUSI, beating ten comparison models.

desk verdict Plausible architecture, but the headline numbers cannot be trusted: reported Dice/Jaccard pairs and dataset splits are mathematically impossible. read the letter →

arxiv 2506.16592 v1 pith:V6D6MNTP submitted 2025-06-19 eess.IV cs.AIcs.CV

classification eess.IVcs.AIcs.CV
keywords breastultrasoundsegmentationhybridattentionnetworkDenseNet121transformerself-attentionglobalspatialfeatureenhancementBCEandJaccardlossBUSIUDIATdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Breast ultrasound images are noisy and have fuzzy lesion boundaries, making automated tumor delineation unreliable. This paper proposes a hybrid attention network that combines a dense convolutional encoder, transformer-style attention at the bottleneck, and spatial-enhancement blocks on skip connections so that the decoder simultaneously preserves fine boundaries and global context. The authors report that the method outperforms ten recent segmentation models on BUSI and five on UDIAT, reaching 94.75 Jaccard and 97.28 Dice on BUSI and 86.71 Jaccard and 92.38 Dice on UDIAT. If correct, the result points toward a segmentation tool that works on low-contrast, speckle-corrupted ultrasound images without manual adjustment, and its modular design offers a reusable recipe for attention-based medical segmentation.

What carries the argument

The load-bearing object is the Transformer Attention Module, a self-aware attention block placed at the encoder bottleneck: it combines Global Spatial Attention, which computes a softmax-normalized dot product between spatially split feature halves to form a $(h\times w)\times(h\times w)$ spatial correlation map, with Transformer Self-Attention, which adds learnable position encoding before scaled dot-product attention so that relative spatial positions are encoded. The Spatial Feature Enhancement Block runs global max-pooling and average-pooling in parallel, concatenates them, and gates the result with a sigmoid-weighted $1\times 1$ convolution before adding it back to the input. These modules carry the argument because the ablation attributes the largest accuracy gains to their addition, and the hybrid loss is what ties the whole network to region-level overlap metrics.

What would settle it

Take the exact train/test split stated in the paper, 700/80 for BUSI and 133/33 for UDIAT, run the proposed model and each baseline on the original images and masks, and recompute Dice, Jaccard, sensitivity, and accuracy per image; if the Dice and IoU values for any method do not satisfy $Dice = 2\cdot IoU/(1+IoU)$, or if the stated image counts exceed the actual dataset sizes, then the reported comparison cannot be reproduced.

Watch

Extended reading notes

Core claim

The central claim is that a single architecture can resolve the two competing demands of ultrasound segmentation, local detail and long-range context, by fusing three attention mechanisms rather than choosing one. At the bottleneck, the Transformer Attention Module concatenates scaled dot-product self-attention with learnable position encoding, global spatial attention, and the original feature map, so channel correlations and spatial-position correlations are learned together. The Spatial Feature Enhancement Block at each skip connection re-weights pooled global features before they are fused with encoder details, and a BCE plus Jaccard loss drives pixel-level and region-level accuracy simultaneously. The paper's evidence is an ablation on BUSI in which each added component raises the reported Jaccard/Dice from a 90.44/94.37 baseline to 94.75/97.28, and full comparisons in which the proposed model records the highest Dice, Jaccard, and specificity among the listed methods on both datasets.

Load-bearing premise

The central claim depends on the reported accuracy numbers being real and mutually consistent, and several published Dice/IoU pairs and dataset split counts fail basic arithmetic checks, so if those numbers are misreported, the claimed superiority over prior models collapses.

Editorial extensions

If this is right

  • If the reported BUSI results are correct, a network with about 15.4 million parameters can exceed 99 percent pixel accuracy and reach 99.84 specificity on ultrasound tumor segmentation, a practical range for clinical decision support.
  • On the UDIAT cross-dataset test, the method generalizes to a different acquisition system and resolution, suggesting the attention design is not overfit to one scanner.
  • The component-wise ablation implies that skip-connection spatial enhancement and bottleneck attention contribute independently and additively, so each module is a separable improvement that other U-Net-style architectures could adopt.
  • The authors' Friedman and Nemenyi analysis implies the performance gap over UNet++, Attention U-Net, and BGRD-TransUNet on BUSI is larger than random variation would produce.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to fix the reported arithmetic, since Dice should equal $2\times IoU/(1+IoU)$, and the stated split counts, then rerun the same comparison; if the corrected numbers still favor the proposed model, the architectural claim survives, and if not, the margin may shrink.
  • The same hybrid attention recipe could be transferred to other low-contrast ultrasound targets such as thyroid nodules or liver lesions, where speckle and fuzzy boundaries create the same failure mode.
  • Because the authors deliberately used no data augmentation, their reported margins are a lower bound on what the architecture could achieve; adding standard augmentation would likely raise Dice further, but it would also make the comparison to prior augmentation-trained baselines less direct.
  • A deployment study measuring inference latency and radiologist reading time would test the paper's implicit promise of real-time clinical use, which the authors assert from parameter count rather than from measured runtime.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript proposes a hybrid attention network for breast tumor segmentation in ultrasound images, combining a pre-trained DenseNet121 encoder, transformer-inspired attention modules (GSA, SDPA, PE), a Spatial Feature Enhancement Block (SFEB) at skip connections, and a hybrid BCE+Jaccard loss. The method is evaluated on the BUSI and UDIAT datasets, and the authors claim state-of-the-art performance across Dice, Jaccard, sensitivity, accuracy, and specificity. The paper also includes ablation studies and statistical significance tests (Friedman and Nemenyi). The central claim is that the proposed method consistently outperforms all compared approaches on both datasets.

Significance. The architectural combination is an incremental engineering contribution built from existing modules (DenseNet, attention, transformers), and the paper would be of moderate interest if the empirical results were reliable. However, the reported quantitative evidence is internally inconsistent: the Dice and Jaccard values in the comparison tables violate the exact mathematical relationship defined in the paper, and the reported dataset splits do not match the actual dataset sizes or the paper's own text. Because the superiority claim rests entirely on these tables, the current manuscript does not provide a trustworthy basis for the claimed contribution. The paper also provides no code, trained models, or per-image prediction scores, making the results non-reproducible and the statistical claims unverifiable.

major comments (3)
  1. [Tables 3 and 4; Eqs. (19) and (23)] The reported Dice and Jaccard values are mathematically inconsistent with the definitions given in the paper. Equations (19) and (23) imply Dice = 2*IoU/(1+IoU) for any non-empty prediction, and hence Dice >= IoU. In Table 3, U-Net reports (Jaccard 76.54, Dice 83.13), but the identity gives Dice ≈ 86.72, and DDRA-Net reports (89.23, 75.32), where Dice < IoU, which is impossible. In Table 4, the proposed method reports (86.71, 92.38) while BGRD-TransUNet reports (86.61, 92.47): the proposed method has a higher Jaccard but a lower Dice, which also violates the monotonic relationship. These inconsistencies directly undermine the central claim of superiority over competing methods.
  2. [Table 1; Section 4.1; Section 4.2] The dataset splits in Table 1 contradict the text and the actual dataset sizes. Section 4.1 states that only benign and malignant BUSI images were retained, but the BUSI dataset contains 647 benign+malignant images, whereas Table 1 lists 700 training and 80 test images (totaling 780, the full dataset including normal images). For UDIAT, the dataset is described as containing 163 images, but Table 1 lists 133 training and 33 test images (totaling 166). Additionally, Table 1 shows no validation split, while Section 4.2 states that 20% of the training data was reserved for validation. These inconsistencies mean the experimental setup is not reproducible and the reported results cannot be attributed to the claimed data protocol.
  3. [Sections 5.1.1 and 5.2.1] The Friedman and Nemenyi tests are described as being performed on Dice scores for each image in the test set, but only aggregate Dice values are reported in Tables 3 and 4. No per-image scores, standard deviations, or number of repeated runs are provided anywhere in the manuscript. Without per-image prediction scores or access to the raw outputs, the claimed p-values (p < 0.001 and p < 0.005) cannot be verified, and the procedure appears to apply a rank-based statistical test to a single aggregate number per method, which is not a valid use of the Friedman test.
minor comments (5)
  1. [Throughout] The abbreviation for the Spatial Feature Enhancement Block is inconsistently written as SFEB in most of the text but SEFB in Table 2 and in the ablation study description (Section 4.4); please standardize.
  2. [Section 3.1.2, Eq. (5)] The definition of the GSA attention map in Eq. (5) is incomplete: the reshaping of F_cc into F1_cc and F2_cc, the normalization, and the role of the scaling factor are not clearly specified, making the formula difficult to reproduce.
  3. [Section 2.3] The sentence describing Swin-Net reads 'with feature refinement and enhancement module and hierarchical multi-scale feature fusion module module'; the duplicated 'module' is a typo, and the sentence should be rephrased for clarity.
  4. [Section 3; Figure 1] Figure 1 labels feature maps as '64x64 x channels', '128x128 x channels', etc., but does not specify the actual channel numbers; the caption and text should clarify the tensor dimensions at each stage.
  5. [Section 5.4; Table 2] The paper claims 'low computational complexity' and practicality for real-time clinical deployment, but only parameter counts are reported; no FLOPs, inference time, or memory usage is provided to support the efficiency claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's segmentation claims are empirical evaluations against public benchmarks, not derived from or defined by their inputs.

full rationale

The paper contains no derivation chain that could be circular. Its central claim—that the proposed Hybrid Attention Network outperforms comparison methods on BUSI and UDIAT—is an empirical claim supported by Tables 2–4 and statistical tests, not by an analytic derivation from a premise that already contains the conclusion. The architecture components (DenseNet121, GSA, SDPA, SFEB, BCE+Jaccard loss) are introduced with references to prior work; none of these references is by the same authors, and none is invoked as a uniqueness theorem. The few self-citations (Refs. 8–14) appear in the introduction and related work as background examples of deep learning or image quality assessment and are not load-bearing for the segmentation result. No parameter is fitted to a subset and then renamed a prediction; the model is evaluated on public datasets with separate train/test splits. The internal inconsistencies noted by the skeptic—Dice/IoU pairs in Tables 3–4 violating Dice = 2·IoU/(1+IoU), and split counts exceeding dataset sizes—are serious correctness and reproducibility concerns that undermine the empirical conclusion, but they are not instances of circularity: the conclusion is not equivalent to the evidence by construction. Section 5.4 acknowledges limitations (moderate dataset size, no augmentation), which further supports a non-circular but limited claim. Therefore, no circular step is present.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The model's learned weights are not counted as free parameters because the paper makes no analytical derivation; the listed entries are hand-chosen hyperparameters that influence training. The key assumptions are the reliability of the annotations and the transferability of pre-trained features. No new physical entities are introduced.

free parameters (4)
  • Loss combination weights = 1.0 (BCE) : 1.0 (Jaccard)
    Fixed at equal weights in Eq. 18; no ablation or search over the weighting is reported, so the choice is a hand-set hyperparameter.
  • Learning rate = 0.001
    Adam optimizer initial learning rate (Section 4.2), chosen without a reported tuning procedure.
  • Batch size = 10
    Selected in Section 4.2 without reported justification.
  • Attention channel reduction factor = c' = c/2
    Introduced in Section 3.1.2 for GSA to halve the channel dimension before computing the spatial attention map; no justification beyond simplification.
assumptions (4)
  • standard math Scaled dot-product attention and softmax produce meaningful feature aggregations.
    Invoked in Eq. 4 (TSA) and Eq. 5 (GSA), Section 3.1.
  • domain assumption The BUSI and UDIAT ground-truth annotations accurately represent tumor boundaries.
    Used as labels for training and evaluation in Section 4.1.
  • domain assumption Pre-trained ImageNet weights of DenseNet121 transfer to breast ultrasound images.
    The encoder initialization relies on this, Section 3 and ref 44.
  • standard math The selected evaluation metrics (Dice, IoU, sensitivity, etc.) reliably quantify segmentation quality.
    Metrics defined in Section 4.3 and used as the basis for all claims.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hybrid Attention Network for Accurate Breast Tumor Segmentation in Ultrasound Images." pith.science (2026). https://pith.science/paper/V6D6MNTP

@misc{pith2026250616592,
  author       = {Pith},
  title        = {Pith review of: Hybrid Attention Network for Accurate Breast Tumor Segmentation in Ultrasound Images},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V6D6MNTP}},
  note         = {Machine review of arXiv:2506.16592}
}
read the original abstract

Breast ultrasound imaging is a valuable tool for early breast cancer detection, but automated tumor segmentation is challenging due to inherent noise, variations in scale of lesions, and fuzzy boundaries. To address these challenges, we propose a novel hybrid attention-based network for lesion segmentation. Our proposed architecture integrates a pre-trained DenseNet121 in the encoder part for robust feature extraction with a multi-branch attention-enhanced decoder tailored for breast ultrasound images. The bottleneck incorporates Global Spatial Attention (GSA), Position Encoding (PE), and Scaled Dot-Product Attention (SDPA) to learn global context, spatial relationships, and relative positional features. The Spatial Feature Enhancement Block (SFEB) is embedded at skip connections to refine and enhance spatial features, enabling the network to focus more effectively on tumor regions. A hybrid loss function combining Binary Cross-Entropy (BCE) and Jaccard Index loss optimizes both pixel-level accuracy and region-level overlap metrics, enhancing robustness to class imbalance and irregular tumor shapes. Experiments on public datasets demonstrate that our method outperforms existing approaches, highlighting its potential to assist radiologists in early and accurate breast cancer diagnosis.

Figures

Figures reproduced from arXiv: 2506.16592 by the authors.

Figure 1
Figure 1. The details of the proposed method. The proposed method consists of a pre-trained encoder and a specific decoder, a spatial features enhancement block (SFEB), and a Transformer attention module. 5/15 [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. The details of the self-aware attention module. The top block shows the transformer self-attention, and the bottom shows the global self-attention block. 3.1.2 Global Spatial Attention (GSA) The TAM uses the GSA component to selectively aggregate global context with learned attributes and encode larger information. Incorporating contextual positioning information into local features enhances intra-class compactness … view at source ↗
Figure 3
Figure 3. Spatial Features Enhancement Block tensor, ultimately improving the model’s segmentation performance. I = R H×W×C (6) In Equation 6, the symbol I represents input, where H denotes height, W represents width, and C signifies its depth, resulting in dimensions H ×W ×C. These dimensions define the size and depth of the input tensor. I1 = ReLU(µ(f 3×3 (I)), (7) Equation 7 represents the output I1, which is acquired by c… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Qualitative comparison of segmentation results on the BUSI dataset. From left to right: original image, ground truth, U-Net45, UNet++58, Attention U-Net57, and the proposed method. Green represents true positives, red indicates false positives, and blue denotes false n…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

59 extracted references · 56 canonical work pages

  1. [1]

    & Shu, P

    Zhang, S., Jin, Z., Bao, L. & Shu, P. The global burden of breast cancer in women from 1990 to 2030: assessment and projection based on the global burden of disease study 2019. Front. Oncol. 14, 1364397 (2024)

  2. [2]

    & Jing, J

    Zheng, D., He, X. & Jing, J. Overview of artificial intelligence in breast cancer medical imaging. J. clinical medicine 12, 419 (2023)

  3. [3]

    & Vedantham, S

    Karellas, A. & Vedantham, S. Breast cancer imaging: a perspective for the next decade. Med. physics 35, 4878–4897 (2008)

  4. [4]

    & Harman, J

    Benson, S., Blue, J., Judd, K. & Harman, J. Ultrasound is now better than mammography for the detection of invasive breast cancer. The Am. journal surgery 188, 381–385 (2004)

  5. [5]

    Role of breast ultrasound for the detection and differentiation of breast lesions

    Madjar, H. Role of breast ultrasound for the detection and differentiation of breast lesions. Breast Care 5, 109–114 (2010)

  6. [6]

    Gonzaga, M. A. How accurate is ultrasound in evaluating palpable breast masses? Pan Afr. Med. J. 7 (2010)

  7. [7]

    Liu, L. et al. Automated breast tumor detection and segmentation with a novel computational framework of whole ultrasound images. Med. & biological engineering & computing 56, 183–199 (2018)

  8. [8]

    Ahmed, N., Asif, H. M. S. & Khalid, H. Piqi: perceptual image quality index based on ensemble of gaussian process regression. Multimed. Tools Appl. 80, 15677–15700 (2021)

Show all 59 references
  1. [9]

    & Ahmed, N

    Khalid, H., Ali, M. & Ahmed, N. Gaussian process-based feature-enriched blind image quality assessment. J. Vis. Commun. Image Represent. 77, 103092 (2021)

  2. [10]

    & Asif, H

    Ahmed, N. & Asif, H. M. S. Perceptual quality assessment of digital images using deep features. Comput. Informatics 39, 385–409 (2020)

  3. [11]

    Ahmed, N., Shahzad Asif, H., Bhatti, A. R. & Khan, A. Deep ensembling for perceptual image quality assessment. Soft Comput. 26, 7601–7622 (2022)

  4. [12]

    Aslam, M. A. et al. Vrl-iqa: Visual representation learning for image quality assessment. IEEE Access 12, 2458–2473 (2023)

  5. [13]

    Aslam, M. A. et al. Qualitynet: A multi-stream fusion framework with spatial and channel attention for blind image quality assessment. Sci. Reports 14, 26039 (2024)

  6. [14]

    Aslam, M. A. et al. Tqp: An efficient video quality assessment framework for adaptive bitrate video streaming. IEEE Access (2024)

  7. [15]

    & Chen, D

    Song, K., Feng, J. & Chen, D. A survey on deep learning in medical ultrasound imaging. Front. Phys. 12, 1398393 (2024)

  8. [16]

    Huang, Z., Wang, L. & Xu, L. Dra-net: Medical image segmentation based on adaptive feature extraction and region-level information fusion. Sci. Reports 14, 9714 (2024)

  9. [17]

    & Bendechache, M

    Anari, S., Sadeghi, S., Sheikhi, G., Ranjbarzadeh, R. & Bendechache, M. Explainable attention based breast tumor segmentation using a combination of unet, resnet, densenet, and efficientnet models. Sci. Reports 15, 1027 (2025)

  10. [18]

    & Han, X.-h

    Fang, W. & Han, X.-h. Spatial and channel attention modulated network for medical image segmentation. In Proceedings of the Asian conference on computer vision (2020). 13/15

  11. [19]

    & Okatani, T

    Murase, R., Suganuma, M. & Okatani, T. How can cnns use image position for segmentation? arXiv preprint arXiv:2005.03463 (2020)

  12. [20]

    Shen, X. et al. Dilated transformer: residual axial attention for breast ultrasound image segmentation. Quant. Imaging Medicine Surg. 12, 4512 (2022)

  13. [21]

    He, J. et al. Sab-net: Self-attention backward network for gastric tumor segmentation in ct images. Comput. Biol. Medicine 169, 107866 (2024)

  14. [22]

    Zhang, H. et al. Acl-dunet: A tumor segmentation method based on multiple attention and densely connected breast ultrasound images. PloS one 19, e0307916 (2024)

  15. [23]

    Byra, M. et al. Breast mass segmentation in ultrasound with selective kernel u-net convolutional neural network. Biomed. Signal Process. Control. 61, 102027 (2020)

  16. [24]

    Zhang, S. et al. Fully automatic tumor segmentation of breast ultrasound images with deep learning. J. Appl. Clin. Med. Phys. 24, e13863 (2023)

  17. [25]

    Michael, E., Ma, H., Li, H., Kulwa, F. & Li, J. Breast cancer segmentation methods: current status and future potentials. BioMed research international 2021, 9962109 (2021)

  18. [26]

    & Quan, R

    Xu, Y ., Liu, F., Xu, W. & Quan, R. Overview of graph theoretical approaches in medical image segmentation. In International Conference on Computational & Experimental Engineering and Sciences , 819–835 (Springer, 2024)

  19. [27]

    & Huang, B

    Li, L., Niu, Y ., Tian, F. & Huang, B. An efficient deep learning strategy for accurate and automated detection of breast tumors in ultrasound image datasets. Front. Oncol. 14, 1461542 (2025)

  20. [28]

    Pan, P. et al. Tumor segmentation in automated whole breast ultrasound using bidirectional lstm neural network and attention mechanism. Ultrasonics 110, 106271 (2021)

  21. [29]

    & Khan, N

    Abraham, N. & Khan, N. M. A novel focal tversky loss function with improved attention u-net for lesion segmentation. In 2019 IEEE 16th international symposium on biomedical imaging (ISBI 2019) , 683–687 (IEEE, 2019)

  22. [30]

    & Yap, M

    Chen, G., Li, L., Dai, Y ., Zhang, J. & Yap, M. H. Aau-net: an adaptive attention u-net for breast lesions segmentation in ultrasound images. IEEE Transactions on Med. Imaging 42, 1289–1300 (2022)

  23. [31]

    Chen, G. et al. Esknet: An enhanced adaptive selection kernel convolution for ultrasound breast tumors segmentation. Expert. Syst. with Appl. 246, 123265 (2024)

  24. [32]

    & Zhang, Q

    Xiao, H., Li, L., Liu, Q., Zhu, X. & Zhang, Q. Transformers in medical image segmentation: A review. Biomed. Signal Process. Control. 84, 104791 (2023)

  25. [33]

    Liu, Q. et al. Optimizing vision transformers for medical image segmentation. In ICASSP 2023-2023 IEEE international conference on acoustics, speech and signal processing (ICASSP), 1–5 (IEEE, 2023)

  26. [34]

    & Hei, X

    Zhang, J., Li, F., Zhang, X., Wang, H. & Hei, X. Automatic medical image segmentation with vision transformer. Appl. Sci. 14, 2741 (2024)

  27. [35]

    & Xie, M

    He, Q., Yang, Q. & Xie, M. Hctnet: A hybrid cnn-transformer network for breast ultrasound image segmentation. Comput. Biol. Medicine 155, 106629 (2023)

  28. [36]

    Ma, Z. et al. Atfe-net: axial transformer and feature enhancement-based cnn for ultrasound breast mass segmentation. Comput. Biol. Medicine 153, 106533 (2023)

  29. [37]

    Lin, A. et al. Ds-transunet: Dual swin transformer u-net for medical image segmentation. IEEE Transactions on Instrumentation Meas. 71, 1–15 (2022)

  30. [38]

    Zhu, C. et al. Swin-net: a swin-transformer-based network combing with multi-scale features for segmentation of breast tumor ultrasound images. Diagnostics 14, 269 (2024)

  31. [39]

    Zhao, Z. et al. Swinhr: Hemodynamic-powered hierarchical vision transformer for breast tumor segmentation. Comput. biology medicine 169, 107939 (2024)

  32. [40]

    Cao, W. et al. Neighbornet: Learning intra-and inter-image pixel neighbor representation for breast lesion segmentation.IEEE J. Biomed. Heal. Informatics (2024)

  33. [41]

    Zhang, H. et al. Hau-net: Hybrid cnn-transformer for breast ultrasound image segmentation. Biomed. Signal Process. Control. 87, 105427 (2024)

  34. [42]

    Wu, R., Lu, X., Yao, Z. & Ma, Y . Mfmsnet: A multi-frequency and multi-scale interactive cnn-transformer hybrid network for breast ultrasound image segmentation. Comput. Biol. Medicine 177, 108616 (2024)

  35. [43]

    & Tairi, H

    Tagnamas, J., Ramadan, H., Yahyaouy, A. & Tairi, H. Multi-task approach based on combined cnn-transformer for efficient segmentation and classification of breast tumors in ultrasound images. Vis. Comput. for Ind. Biomed. Art 7, 2 (2024)

  36. [44]

    & Weinberger, K

    Huang, G., Liu, Z., Van Der Maaten, L. & Weinberger, K. Q. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition , 4700–4708 (2017)

  37. [45]

    & Brox, T

    Ronneberger, O., Fischer, P. & Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18 , 234...

  38. [46]

    & Kong, A

    Chen, B., Liu, Y ., Zhang, Z., Lu, G. & Kong, A. W. K. Transattunet: Multi-level attention-guided u-net with transformer for medical image segmentation. IEEE Transactions on Emerg. Top. Comput. Intell. (2023)

  39. [47]

    A survey of loss functions for semantic segmentation

    Jadon, S. A survey of loss functions for semantic segmentation. In 2020 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), 1–7 (IEEE, 2020)

  40. [48]

    & Wiering, M

    van Beers, F., Lindström, A., Okafor, E. & Wiering, M. A. Deep neural networks with intersection over union loss for binary image segmentation. In ICPRAM, 438–445 (2019)

  41. [49]

    & Fahmy, A

    Al-Dhabyani, W., Gomaa, M., Khaled, H. & Fahmy, A. Dataset of breast ultrasound images. Data brief 28, 104863 (2020)

  42. [50]

    Yap, M. H. et al. Breast ultrasound region of interest detection and lesion localisation. Artif. intelligence medicine 107, 101880 (2020)

  43. [51]

    Kingma, D. P. & Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). 14/15

  44. [52]

    Hu, K. et al. Boundary-guided and region-aware network with global scale-adaptive for accurate segmentation of breast tumors in ultrasound images. IEEE J. Biomed. Heal. Informatics 27, 4421–4432 (2023)

  45. [53]

    & Ning, C

    Tang, R. & Ning, C. Mlfeu-net: A multi-scale low-level feature enhancement unet for breast lesions segmentation in ultrasound images. Biomed. Signal Process. Control. 100, 106931 (2025)

  46. [54]

    Cao, H. et al. Swin-unet: Unet-like pure transformer for medical image segmentation. In Computer Vision–ECCV 2022 Workshops: Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part III, 205–218 (Springer, 2023)

  47. [55]

    Qu, X. et al. Eh-former: Regional easy-hard-aware transformer for breast lesion segmentation in ultrasound images. Inf. Fusion 109, 102430 (2024)

  48. [56]

    Ji, Z. et al. Bgrd-transunet: A novel transunet-based model for ultrasound breast lesion segmentation. IEEE Access (2024)

  49. [57]

    Oktay, O. et al. Attention u-net: Learning where to look for the pancreas. arXiv preprint arXiv:1804.03999 (2018)

  50. [58]

    M., Tajbakhsh, N

    Zhou, Z., Rahman Siddiquee, M. M., Tajbakhsh, N. & Liang, J. Unet++: A nested u-net architecture for medical image segmentation. In Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support: 4th International Workshop, DLMIA 2018, and 8th In...

  51. [59]

    Sun, J. et al. Ddra-net: Dual-channel deep residual attention upernet for breast lesions segmentation in ultrasound images. IEEE Access (2024). 15/15

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.