Pith. sign in

REVIEW 4 major objections 4 minor 62 references

TAFM-Net: A Novel Approach to Skin Lesion Segmentation Using Transformer Attention and Focal Modulation

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read TAFM-Net combines transformer self-attention and focal modulation in a U-Net to reach state-of-the-art skin lesion segmentation scores.

desk verdict A plausible architecture-level contribution whose reported superiority is undercut by uncontrolled comparisons and internally inconsistent numbers; fix the evaluation before believing the margins. read the letter →

arxiv 2411.17556 v1 pith:UT5HCJVE submitted 2024-11-26 eess.IV cs.CV

classification eess.IVcs.CV
keywords skinlesionsegmentationdermoscopicimageanalysistransformerself-attentionfocalmodulationU-NetISICbenchmarkboundarylossmedical
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that adding a transformer self-attention module at the encoder-decoder bottleneck and focal modulation blocks inside the skip connections can push skin lesion segmentation accuracy well beyond previously published results. On the ISIC 2016, 2017, and 2018 benchmarks, the proposed TAFM-Net reports Jaccard scores of 93.64%, 86.88%, and 92.88%, respectively, surpassing earlier attention- and transformer-based networks by large margins. The same design is also smaller and faster than the comparison networks, with 20.6 million parameters and an 18.25 ms inference time per image. If these numbers hold under controlled comparison, the network would be a practical candidate for clinical computer-assisted diagnosis.

What carries the argument

The load-bearing component is the self-aware attention module placed at the encoder-decoder bottleneck. It concatenates three streams: transformer self-attention output, global spatial attention output, and the original encoder feature map, so that both channel-level and position-level dependencies are preserved. Around this, focal modulation blocks in the skip connections compute a global feature map using depthwise separable convolution and use it to modulate the local convolution output, effectively injecting global context into local feature extraction. The third piece of the machinery is the dynamic fused loss, which combines binary cross-entropy, Jaccard loss, and boundary loss, with the boundary term's weight increasing as training progresses.

What would settle it

Run TAFM-Net and all publicly available comparison methods on the same ISIC 2016, 2017, and 2018 train/test splits with identical preprocessing and evaluation protocols, and check whether TAFM-Net still leads on the majority of metrics.

Watch

Extended reading notes

Core claim

The central claim is that TAFM-Net, built from an EfficientNetV2B1 encoder, a self-aware attention module at the bottleneck, focal modulation in every skip connection, a densely connected decoder, and a dynamically weighted fused loss, consistently outperforms existing state-of-the-art methods on all three ISIC datasets and on PH2 after training on ISIC 2016. The reported improvements in Jaccard score are 6.2%–15.1% on ISIC 2016, 3.8%–17.2% on ISIC 2017, and 9.3%–19.7% on ISIC 2018 over the compared methods. The authors attribute the gains to the combination of global contextual reasoning from transformer attention, fine-grained feature emphasis from focal modulation, and the boundary-aware dynamic loss, and they present ablation experiments and Grad-CAM visualizations to support these attributions.

Load-bearing premise

The claim of state-of-the-art performance rests on comparing TAFM-Net's own numbers with scores copied from other papers that likely used different training sets, preprocessing, and evaluation rules, and this comparability is not established.

Editorial extensions

If this is right

  • Reported Jaccard scores of 93.64%, 86.88%, and 92.88% on ISIC 2016, 2017, and 2018 would make TAFM-Net the new reference point for skin lesion segmentation on these benchmarks.
  • With 20.6 million parameters and 18.25 ms inference per 256x256 image, the network is small and fast enough for clinical decision-support deployment.
  • The ablation results indicate that the transformer at the bottleneck and focal modulation in skip connections, rather than the backbone alone, drive most of the accuracy gain.
  • The dynamically weighted boundary loss improves segmentation in low-contrast and hair-occluded images, which are the hard cases that matter in practice.
  • Cross-dataset training from ISIC 2016 to PH2 shows the method generalizes to a new distribution, supporting its use beyond the training benchmark.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reported superiority margins may compress in a controlled re-run because many competitor scores were copied from their original papers, which likely used different training sets, preprocessing, and evaluation protocols; a head-to-head benchmark with identical splits and preprocessing is the natural next test.
  • Since focal modulation is a generic block, the same encoder-decoder recipe could transfer to other medical segmentation problems, such as retinal vessel or organ segmentation, without architectural changes.
  • The linear decay of the loss weight is a simple schedule; adaptive weighting based on validation boundary performance might yield further gains, especially when training epochs are limited.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript proposes TAFM-Net, a U-Net-style architecture for skin lesion segmentation. The encoder is EfficientNetV2B1; a transformer self-attention block and a global spatial attention block are inserted at the bottleneck; focal modulation blocks are placed in the skip connections; and the decoder uses seven upsampling blocks with dense connections. Training uses a dynamically weighted fusion of BCE, Jaccard, focal Tversky, Dice, and boundary losses. The paper reports Jaccard scores of 93.64% (ISIC2016), 86.88% (ISIC2017), 92.88% (ISIC2018), and 95.60% (PH2 cross-dataset), with 20.6M parameters and 18.25 ms inference time, and claims consistent state-of-the-art superiority.

Significance. The architecture is a reasonable design study: combining transformer attention, focal modulation, and a boundary-aware dynamic loss is directionally interesting, and the ablation over loss functions and network components is a strength. If the numerical results and comparisons are verified, a 20.6M-parameter model with the reported accuracy could be practically useful in clinical workflows. The lightweight claims in Table 7 are also attractive. However, the central claim of state-of-the-art performance currently rests on cross-paper comparisons and internally inconsistent numbers, so the significance cannot be fully assessed until those issues are resolved.

major comments (4)
  1. [Section 4.5, Table 5] The TAFM-Net row in Table 5 for ISIC2018 reports accuracy 99.08, sensitivity 96.90, specificity 98.19, Jaccard 93.08, and Dice 96.85, while Table 4 for the same training/testing setting and the abstract report accuracy 97.87, sensitivity 96.39, specificity 97.92, Jaccard 92.88, and Dice 96.53. These cannot both describe the same experiment. The discrepancy must be resolved and a single protocol stated for the numbers used in the comparison table.
  2. [Section 4.5.1 and Section 4.5.2] The claimed improvement ranges are not supported by the tables. For ISIC2018 the text reports Jaccard improvements of 9.3%–19.7%, but Table 5 gives differences ranging from 8.53 percentage points (vs. ARU-GD) to 13.20 (vs. CPFNet). For ISIC2017 the text claims 3.8%–17.2%, while Table 5 gives 3.18 (vs. Hyper-Fusion Net) to 11.19 (vs. U-Net). For ISIC2016 the text claims 6.2%–15.1%, while Table 5 gives 5.47 (vs. Hyper-Fusion Net) to 12.26 (vs. U-Net). Likewise, the PH2 claim of 9.1%–13.8% in Section 4.5.2 exceeds the 8.00–11.61 range in Table 6. The text and tables must be reconciled.
  3. [Section 4.5] The comparison with state-of-the-art methods is not a controlled experiment. The paper states that scores for comparison methods were taken from the original articles, meaning they were produced under different training/validation/test splits, preprocessing, post-processing, and evaluation protocols. The conclusion in Section 4.5.1 that TAFM-Net 'consistently outperformed existing methods by a considerable margin' is therefore not established by Table 5. The authors should re-run at least all publicly available methods under an identical protocol with the same evaluation code, or, failing that, explicitly present the comparison as indicative and temper the superiority claims.
  4. [Table 1 and Section 4.3] The early-stopping protocol is underspecified. Table 1 lists no validation set for ISIC2016 and ISIC2017, yet Section 4.3 describes early stopping monitored from epoch 10 with patience 9. The monitored data split is not identified. If the test set was used for early stopping, the reported scores are optimistic; if a validation split was carved from the training set, it should be described. Additionally, Tables 2–4 report 3-fold cross-validation means, but Table 1 gives fixed train/test partitions; please clarify how the folds relate to the fixed test set.
minor comments (4)
  1. [Table 5 caption] The caption names the first dataset as ISIC 2018 twice; it should read ISIC 2016, ISIC 2017, and ISIC 2018.
  2. [Section 3.5.4] With γ=1, Eq. (14) is identical to Eq. (13), so describing the used loss as 'focal Tversky' is misleading; either select γ>1 or call it the Tversky loss in the reported experiments.
  3. [Section 3.5.6 and Section 4.3] The dynamic loss schedule is not fully reproducible because the total number of training epochs is not specified, and the batch size and the exact point at which α reaches its lower bound are not given.
  4. [General] No code or trained model weights are provided. Given the paper's stated ambition to serve as a baseline, releasing an implementation would substantially improve reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: TAFM-Net's reported benchmark gains are an empirical result, not a derivation equivalent to its inputs.

full rationale

The paper's central claim is an empirical performance comparison. The architecture combines externally published components (EfficientNetV2B1, transformer self-attention [43], focal modulation [50,55], boundary loss [52], and focal Tversky loss [51]) with a hand-specified dynamic fusion schedule; none of these are defined in terms of the target segmentation scores, and no uniqueness theorem or author-specific prior result is invoked to force the design. The many self-citations in the introduction and related work are contextual and are not used to justify the reported improvements. The loss-function ablation in Section 4.4.1 selects L4 from the ISIC2016 cross-validation, and the same L4 configuration is later reported in Table 5; this is a model-selection or possible test-set-reuse limitation, not a constructional equivalence, because the reported Jaccard values are measurements produced by training the network, not algebraic consequences of the loss definition. The cross-paper comparison in Section 4.5, in which comparison scores are taken from original articles, raises a comparability threat, and Table 5's TAFM-Net ISIC2018 row (J=93.08, D=96.85, A=99.08) conflicts with Table 4 and the abstract (J=92.88, D=96.53, A=97.87); both are correctness and reproducibility concerns, not circularity. No step in the paper reduces a prediction to its input by definition, so the appropriate circularity score is 0.

Assumptions & free parameters 10 free parameters · 5 assumptions · 0 invented entities

The central claim rests on a series of hand-chosen hyperparameters, particularly the dynamic loss schedule and the focal Tversky weights. None of these are fitted with a principled procedure, and gamma=1 removes the focal effect. The validity of the comparison also depends on the untested assumption that published scores from other papers are directly comparable.

free parameters (10)
  • dynamic loss weight alpha initial value = 1
    Hand-chosen; used to blend region losses with boundary loss in L4-L6.
  • alpha decay step = 0.005
    Hand-chosen; decrements alpha each epoch from 1 to 0.01.
  • focal Tversky alpha = 0.3
    Sets false-negative weight in Tversky index; taken from [51], not tuned on the target data.
  • focal Tversky beta = 0.7
    Sets false-positive weight; taken from [51].
  • focal Tversky gamma = 1
    Gamma=1 makes the focal Tversky loss equivalent to the standard Tversky loss; the 'focal' effect claimed in section 3.5.4 is therefore absent.
  • learning rate = 0.001
    Standard Adam learning rate; chosen without reported tuning (section 4.3).
  • dropout rate = 0.5
    Applied in decoder blocks (section 3.4); not justified by experiments.
  • input resolution = 256x256
    All images reshaped to 256x256 (section 4.3); a resolution choice that affects all reported metrics.
  • early stopping patience = 9 epochs with monitoring from epoch 10
    Terminates training when the monitored metric does not improve for 9 epochs (section 4.3).
  • binarization threshold = maximizes F1 score
    The predicted map is thresholded to maximize F1 (section 3.4); this is a data-dependent choice that influences final metrics.
assumptions (5)
  • domain assumption The ISIC and PH2 datasets provide accurate ground truth for skin lesion segmentation.
    Section 4.1 relies on these public datasets as ground truth; label noise or misannotation would directly affect all metrics.
  • standard math The boundary loss approximation from Kervadec et al. is valid for this task.
    Section 3.5.5 uses Equation (18) as the boundary loss, citing [52]; no proof is provided in this paper.
  • domain assumption EfficientNetV2B1 pretrained on ImageNet provides a suitable feature extractor.
    Section 3.1 selects this backbone for performance; this assumes transfer learning works for dermoscopic images.
  • standard math Focal modulation blocks behave as described in the original Focal Modulation Networks paper.
    Section 3.3 adopts the FM mechanism from [55] without modification; correctness is assumed from the cited work.
  • domain assumption The reported comparison scores from other papers were obtained with protocols comparable to this work.
    Section 4.5 uses published scores for most SOTA methods, assuming cross-paper comparability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TAFM-Net: A Novel Approach to Skin Lesion Segmentation Using Transformer Attention and Focal Modulation." pith.science (2026). https://pith.science/paper/UT5HCJVE

@misc{pith2026241117556,
  author       = {Pith},
  title        = {Pith review of: TAFM-Net: A Novel Approach to Skin Lesion Segmentation Using Transformer Attention and Focal Modulation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UT5HCJVE}},
  note         = {Machine review of arXiv:2411.17556}
}
read the original abstract

Incorporating modern computer vision techniques into clinical protocols shows promise in improving skin lesion segmentation. The U-Net architecture has been a key model in this area, iteratively improved to address challenges arising from the heterogeneity of dermatologic images due to varying clinical settings, lighting, patient attributes, and hair density. To further improve skin lesion segmentation, we developed TAFM-Net, an innovative model leveraging self-adaptive transformer attention (TA) coupled with focal modulation (FM). Our model integrates an EfficientNetV2B1 encoder, which employs TA to enhance spatial and channel-related saliency, while a densely connected decoder integrates FM within skip connections, enhancing feature emphasis, segmentation performance, and interpretability crucial for medical image analysis. A novel dynamic loss function amalgamates region and boundary information, guiding effective model training. Our model achieves competitive performance, with Jaccard coefficients of 93.64\%, 86.88\% and 92.88\% in the ISIC2016, ISIC2017 and ISIC2018 datasets, respectively, demonstrating its potential in real-world scenarios.

Figures

Figures reproduced from arXiv: 2411.17556 by the authors.

Figure 1
Figure 1. Examples of challenging skin-lesion dermoscopy images: (a) variation in appearance, (b) presence [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Design of the proposed transformer attention focal modulation network (TAFM-Net). [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Design of the self-aware attention module. Top: The transformer self-attention (TSA) block. [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Design of the focal modulation block. where WQ, WK, WV are three embedding matrices for different linear projections. The scaled dot product of Q and K with Softmax normalisation gives ETSA ∈ R c×c , which represents the similarity between channels in Q and others. By …
Figure 5
Figure 5. Figure 5: Design of the decoder block. the following, P and G denote the prediction of the model and the segmentation of the ground truth, respectively, and p c i is the prediction of the model that the pixel i belongs to the class c, while g c i is the corresponding ground trut…
Figure 6
Figure 6. Figure 6: Visual evaluation of TAFM-Net for the different loss functions on three challenging example [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: Heatmaps showing the impact of the different network components of TAFM-Net. The three [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: Visual performance comparison of TAFM-Net with other SOTA methods on four example cases [PITH_FULL_IMAGE:figures/full_fig_p023_8.png]
Figure 9
Figure 9. Figure 9: Visual performance comparison of TAFM-Net with other SOTA methods on four example cases [PITH_FULL_IMAGE:figures/full_fig_p024_9.png]
Figure 10
Figure 10. Figure 10: Visual performance comparison of TAFM-Net with other SOTA methods on four example cases [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]
Figure 11
Figure 11. Figure 11: Visual performance of TAFM-Net on eight examples cases from the PH2 dataset with training [PITH_FULL_IMAGE:figures/full_fig_p025_11.png]
Figure 12
Figure 12. Figure 12: Performance trend of TAFM-Net and comparison methods on the ISIC 2016 dataset in terms of [PITH_FULL_IMAGE:figures/full_fig_p027_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

62 extracted references · 52 canonical work pages

  1. [1]

    T. M. Khan, S. S. Naqvi, E. Meijering, Esdmr-net: A lightweight network with expand-squeeze and dual multiscale residual connections for medical image segmentation, Engineering Applications of Artificial Intelligence 133 (2024) 107995

  2. [2]

    T. M. Khan, S. S. Naqvi, E. Meijering, Leveraging image complexity in macro-level neural network design for medical image segmentation, Scientific Reports 12 (1) (2022) 22286

  3. [3]

    T. M. Khan, A. Robles-Kelly, S. S. Naqvi, T-net: A resource-constrained tiny convolutional neural network for medical image segmentation, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2022, pp. 644–653

  4. [4]

    T. M. Khan, M. Arsalan, A. Robles-Kelly, E. Meijering, Mkis-net: a light-weight multi-kernel network for medical image segmentation, in: International Conference on Digital Image Computing: Tech- niques and Applications (DICTA), 10.1109/DICTA56598.2022.10034573, 2022, pp. 1–8

  5. [5]

    S. S. Naqvi, Z. A. Langah, H. A. Khan, M. I. Khan, T. Bashir, M. I. Razzak, T. M. Khan, Glan: Gan as- sisted lightweight attention network for biomedical imaging based diagnostics, Cognitive Computation 15 (3) (2023) 932–942. 28

  6. [6]

    Iqbal, T

    S. Iqbal, T. M. Khan, S. S. Naqvi, A. Naveed, M. Usman, H. A. Khan, I. Razzak, Ldmres-net: A lightweight neural network for efficient medical image segmentation on iot and edge devices, IEEE journal of biomedical and health informatics (2023)

  7. [7]

    Qayyum, I

    A. Qayyum, I. Razzak, M. Mazher, T. Khan, W. Ding, S. Niederer, Two-stage self-supervised con- trastive learning aided transformer for real-time medical image segmentation, IEEE Journal of Biomed- ical and Health Informatics (2023)

  8. [8]

    Javed, T

    S. Javed, T. M. Khan, A. Qayyum, A. Sowmya, I. Razzak, Advancing medical image segmentation with mini-net: A lightweight solution tailored for efficient segmentation of medical images, arXiv preprint arXiv:2405.17520 (2024)

Show all 62 references
  1. [9]

    Iqbal, T

    S. Iqbal, T. M. Khan, K. Naveed, S. S. Naqvi, S. J. Nawaz, Recent trends and advances in fundus image analysis: A review, Compt. in Biology and Medicine (2022) 106277

  2. [10]

    T. A. Soomro, M. A. Khan, J. Gao, T. M. Khan, M. Paul, N. Mir, Automatic retinal vessel extraction al- gorithm, in: 2016 International Conference on Digital Image Computing: Techniques and Applications (DICTA), IEEE, 2016, pp. 1–8

  3. [11]

    M. A. Khan, T. M. Khan, T. A. Soomro, N. Mir, J. Gao, Boosting sensitivity of a retinal vessel seg- mentation algorithm, Pattern Analysis and Applications 22 (2019) 583–599

  4. [12]

    T. M. Khan, F. Abdullah, S. S. Naqvi, M. Arsalan, M. A. Khan, Shallow vessel segmentation network for automatic retinal vessel segmentation, in: 2020 International Joint Conference on Neural Networks (IJCNN), IEEE, 2020, pp. 1–7

  5. [13]

    Arsalan, T

    M. Arsalan, T. M. Khan, S. S. Naqvi, M. Nawaz, I. Razzak, Prompt deep light-weight vessel segmenta- tion network (plvs-net), IEEE/ACM Transactions on Computational Biology and Bioinformatics 20 (2) (2022) 1363–1371

  6. [14]

    T. M. Khan, S. S. Naqvi, A. Robles-Kelly, E. Meijering, Neural network compression by joint sparsity promotion and redundancy reduction, in: International Conference on Neural Information Processing, Springer International Publishing Cham, 2022, pp. 612–623

  7. [15]

    T. M. Khan, S. S. Naqvi, A. Robles-Kelly, I. Razzak, Retinal vessel segmentation via a multi-resolution contextual network and adversarial learning, Neural Networks 165 (2023) 310–320

  8. [16]

    T. M. Khan, S. S. Naqvi, M. Arsalan, M. A. Khan, H. A. Khan, A. Haider, Exploiting residual edge information in deep fully convolutional neural networks for retinal vessel segmentation, in: 2020 In- ternational Joint Conference on Neural Networks (IJCNN), IEEE, 2020, pp. 1–8. 29

  9. [17]

    T. M. Khan, A. Robles-Kelly, S. S. Naqvi, A semantically flexible feature fusion network for retinal vessel segmentation, in: International Conference on Neural Information Processing, Springer, Cham, 2020, pp. 159–167

  10. [18]

    T. M. Khan, A. Robles-Kelly, S. S. Naqvi, A. Muhammad, Residual multiscale full convolutional network (rm-fcn) for high resolution semantic segmentation of retinal vasculature, in: Structural, Syn- tactic, and Statistical Pattern Recognition: Joint IAPR International Workshops...

  11. [19]

    T. M. Khan, A. Robles-Kelly, S. S. Naqvi, Rc-net: A convolutional neural network for retinal vessel segmentation, in: 2021 Digital Image Computing: Techniques and Applications (DICTA), IEEE, 2021, pp. 01–07

  12. [20]

    Naveed, S

    A. Naveed, S. S. Naqvi, T. M. Khan, I. Razzak, Pca: progressive class-wise attention for skin lesions diagnosis, Engineering Applications of Artificial Intelligence 127 (2024) 107417

  13. [21]

    Naveed, S

    A. Naveed, S. S. Naqvi, S. Iqbal, I. Razzak, H. A. Khan, T. M. Khan, Ra-net: Region-aware attention network for skin lesion segmentation, Cognitive Computation (2024) 1–18

  14. [22]

    Iqbal, M

    S. Iqbal, M. Zeeshan, M. Mehmood, T. M. Khan, I. Razzak, Tesl-net: A transformer-enhanced cnn for accurate skin lesion segmentation, arXiv preprint arXiv:2408.09687 (2024)

  15. [23]

    Naveed, S

    A. Naveed, S. S. Naqvi, T. M. Khan, S. Iqbal, M. Y . Wani, H. A. Khan, Ad-net: Attention-based dilated convolutional residual network with guided decoder for robust skin lesion segmentation, Neural Computing and Applications (2024) 1–23

  16. [24]

    Ronneberger, P

    O. Ronneberger, P. Fischer, T. Brox, U-Net: Convolutional networks for biomedical image segmenta- tion, in: Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2015, pp. 234– 241

  17. [25]

    Ghafoorian, N

    M. Ghafoorian, N. Karssemeijer, T. Heskes, I. W. M. van Uder, F. E. de Leeuw, E. Marchiori, B. van Ginneken, B. Platel, Non-uniform patch sampling with deep convolutional neural networks for white matter hyperintensity segmentation, in: IEEE International Symposium on Biomedic...

  18. [26]

    L. Yu, H. Chen, Q. Dou, J. Qin, P.-A. Heng, Automated melanoma recognition in dermoscopy images via very deep residual networks, IEEE Transactions on Medical Imaging 36 (4) (2017) 994–1004

  19. [27]

    Basak, R

    H. Basak, R. Kundu, R. Sarkar, MFSNet: A multi focus segmentation network for skin lesion segmen- tation, Pattern Recognition 128 (2022) 108673. 30

  20. [28]

    K. Wang, X. Zhang, X. Zhang, Y . Lu, S. Huang, D. Yang, Eanet: Iterative edge attention network for medical image segmentation, Pattern Recognition 127 (2022) 108636

  21. [29]

    Oktay, J

    O. Oktay, J. Schlemper, L. L. Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y . Hammerla, B. Kainz, B. Glocker, D. Rueckert, Attention U-Net: Learning where to look for the pancreas, arXiv:1804.03999 (2018)

  22. [30]

    Zhang, Y

    J. Zhang, Y . Xie, Y . Xia, C. Shen, Attention residual learning for skin lesion classification, IEEE Transactions on Medical Imaging 38 (9) (2019) 2092–2103

  23. [31]

    S. Woo, J. Park, J.-Y . Lee, I. S. Kweon, CBAM: Convolutional block attention module, in: European Conference on Computer Vision (ECCV), 2018, pp. 3–19

  24. [32]

    Farooq, Z

    H. Farooq, Z. Zafar, A. Saadat, T. M. Khan, S. Iqbal, I. Razzak, Lssf-net: Lightweight segmentation with self-awareness, spatial attention, and focal modulation, arXiv preprint arXiv:2409.01572 (2024)

  25. [33]

    Iqbal, T

    S. Iqbal, T. M. Khan, S. S. Naqvi, A. Naveed, E. Meijering, Tbconvl-net: A hybrid deep learning architecture for robust medical image segmentation, Pattern Recognition 158 (2025) 111028

  26. [34]

    T. M. Khan, S. Iqbal, S. S. Naqvi, I. Razzak, E. Meijering, Lmbf-net: A lightweight multipath bidirec- tional focal attention network for multifeatures segmentation, in: 2024 IEEE International Conference on Image Processing (ICIP), IEEE, 2024, pp. 2807–2813

  27. [35]

    Jiang, J

    X. Jiang, J. Jiang, B. Wang, J. Yu, J. Wang, SEACU-Net: Attentive ConvLSTM U-Net with squeeze- and-excitation layer for skin lesion segmentation, Computer Methods and Programs in Biomedicine 225 (2022) 107076

  28. [36]

    R. Azad, M. Asadi-Aghbolaghi, M. Fathy, S. Escalera, Bi-directional ConvLSTM U-Net with densely connected convolutions, in: IEEE/CVF International Conference on Computer Vision Workshops (IC- CVW), 2019, pp. 406–415

  29. [37]

    H. Song, W. Wang, S. Zhao, J. Shen, K.-M. Lam, Pyramid dilated deeper ConvLSTM for video salient object detection, in: European Conference on Computer Vision (ECCV), 2018, pp. 715–731

  30. [38]

    Z. Zhou, M. M. Rahman Siddiquee, N. Tajbakhsh, J. Liang, UNet++: A nested U-Net architecture for medical image segmentation, in: Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, 2018, pp. 3–11

  31. [39]

    D. Maji, P. Sigedar, M. Singh, Attention Res-UNet with guided decoder for semantic segmentation of brain tumors, Biomedical Signal Processing and Control 71 (2022) 103077. 31

  32. [40]

    E. K. Aghdam, R. Azad, M. Zarvani, D. Merhof, Attention Swin U-Net: Cross-contextual attention mechanism for skin lesion segmentation, arXiv:2210.16898 (2022)

  33. [41]

    H. Cao, Y . Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, M. Wang, Swin-Unet: Unet-like pure trans- former for medical image segmentation, in: European Conference on Computer Vision Workshops (ECCVW), 2023, pp. 205–218

  34. [42]

    M. Fiaz, M. Noman, H. Cholakkal, R. M. Anwer, J. Hanna, F. S. Khan, Guided-attention and gated- aggregation network for medical image segmentation, Pattern Recognition 156 (2024) 110812

  35. [43]

    B. Chen, Y . Liu, Z. Zhang, G. Lu, A. W. K. Kong, TransAttUnet: Multi-level attention-guided U- Net with Transformer for medical image segmentation, IEEE Transactions on Emerging Topics in Computational Intelligence 8 (1) (2024) 55–68

  36. [44]

    Huang, S

    Z. Huang, S. Cheng, L. Wang, Medical image segmentation based on dynamic positioning and region- aware attention, Pattern Recognition 151 (2024) 110375

  37. [45]

    Y . Dong, L. Wang, Y . Li, TC-Net: Dual coding network of Transformer and CNN for skin lesion segmentation, PLoS One 17 (11) (2022) e0277578

  38. [46]

    K. Feng, L. Ren, G. Wang, H. Wang, Y . Li, SLT-Net: A codec network for skin lesion segmentation, Computers in Biology and Medicine 148 (2022) 105942

  39. [47]

    F. Yuan, Z. Zhang, Z. Fang, An effective cnn and transformer complementary network for medical image segmentation, Pattern Recognition 136 (2023) 109228

  40. [48]

    X. Guo, X. Lin, X. Yang, L. Yu, K.-T. Cheng, Z. Yan, Uctnet: Uncertainty-guided cnn-transformer hybrid networks for medical image segmentation, Pattern Recognition 152 (2024) 110491

  41. [49]

    M. Tan, Q. V . Le, EfficientNetV2: Smaller models and faster training, arXiv:2104.00298 (2021)

  42. [50]

    Naderi, M

    M. Naderi, M. Givkashi, F. Piri, N. Karimi, S. Samavi, Focal-UNet: UNet-like focal modulation for medical image segmentation, arXiv:2212.09263 (2022)

  43. [51]

    Abraham, N

    N. Abraham, N. M. Khan, A novel focal Tversky loss function with improved attention U-Net for lesion segmentation, arXiv:1810.07842 (2018)

  44. [52]

    Kervadec, J

    H. Kervadec, J. Bouchtiba, C. Desrosiers, E. Granger, J. Dolz, I. Ben Ayed, Boundary loss for highly unbalanced segmentation, Medical Image Analysis 67 (2021) 101851

  45. [53]

    Mirikharaji, K

    Z. Mirikharaji, K. Abhishek, A. Bissoto, C. Barata, S. Avila, E. Valle, M. E. Celebi, G. Hamarneh, A survey on deep learning for skin lesion segmentation, Medical Image Analysis 88 (2023) 102863. 32

  46. [54]

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, Grad-CAM: Visual ex- planations from deep networks via gradient-based localization, in: IEEE International Conference on Computer Vision (ICCV), 2017, pp. 618–626

  47. [55]

    J. Yang, C. Li, X. Dai, J. Gao, Focal modulation networks, Advances in Neural Information Processing Systems 35 (2022) 4203–4217

  48. [56]

    Z. Wang, J. Lyu, X. Tang, autoSMIM: Automatic superpixel-based masked image modeling for skin lesion segmentation, IEEE Transactions on Medical Imaging 42 (12) (2023) 3501–3511

  49. [57]

    S. Feng, H. Zhao, F. Shi, X. Cheng, M. Wang, Y . Ma, D. Xiang, W. Zhu, X. Chen, CPFNet: Con- text pyramid fusion network for medical image segmentation, IEEE Transactions on Medical Imaging 39 (10) (2020) 3008–3018

  50. [58]

    B. Lei, Z. Xia, F. Jiang, X. Jiang, Z. Ge, Y . Xu, J. Qin, S. Chen, T. Wang, S. Wang, Skin lesion segmentation via generative adversarial networks with dual discriminators, Medical Image Analysis 64 (2020) 101716

  51. [59]

    L. Bi, M. Fulham, J. Kim, Hyper-fusion network for semi-automatic segmentation of skin lesions, Medical Image Analysis 76 (2022) 102334

  52. [60]

    K. Hu, J. Lu, D. Lee, D. Xiong, Z. Chen, AS-Net: Attention synergy network for skin lesion segmen- tation, Expert Systems with Applications 201 (2022) 117112

  53. [61]

    L. Bi, J. Kim, E. Ahn, A. Kumar, D. Feng, M. Fulham, Step-wise integration of deep class-specific learning for dermoscopic image segmentation, Pattern Recognition 85 (2019) 78–89

  54. [62]

    W. Cao, G. Yuan, Q. Liu, C. Peng, J. Xie, X. Yang, X. Ni, J. Zheng, ICL-Net: Global and local inter-pixel correlations learning network for skin lesion segmentation, IEEE Journal of Biomedical and Health Informatics 27 (1) (2023) 145–156. 33

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.