REVIEW 4 major objections 5 minor 41 references
J-CaPA : Joint Channel and Pyramid Attention Improves Medical Image Segmentation
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Jointly applying Channel Attention and Pyramid Attention in a transformer-based U-Net improves multi-organ CT segmentation beyond either mechanism alone, reaching 82.29% Mean Dice and 19.74 mm HD95 on the Synapse dataset.
desk verdict Sensible incremental idea and a coherent ablation, but the headline gains rest on a baseline that looks copied from the TransUNet paper, so the main claim isn't established as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing component is the J-CaPA module, a dual-attention block inserted into the encoder of a transformer-based U-Net. The Channel Attention branch reshapes the input feature map $X \in \mathbb{R}^{B \times C \times H \times W}$, derives query and key projections, applies softmax over a max-normalized energy matrix, and reweights the value representation. The Pyramid Attention branch computes query, key, and value projections at three spatial scales ($s=1$, $0.5$, $0.25$), applies dot-product attention at each scale, and upsamples the results back to full resolution. The two branches are combined by element-wise addition and refined with 3x3 convolutions, with each branch controlled by a learnable parameter $\gamma_{\mathrm{CA}}$ or $\gamma_{\mathrm{PA}}$ that starts at zero and is updated during training. CutMix augmentation, applied to 33% of images per batch with patch areas from 20% to 60%, compensates for the added model capacity. The ablation study isolates this module as the source of the claimed improvement: the joint version outperforms either branch alone and the unmodified baseline.
What would settle it
Re-run the comparison methods on the same Synapse train/test split with the same training protocol (150 epochs, batch size 8, same preprocessing and augmentation) and check whether the reported 82.29% Mean Dice and 19.74 mm HD95 hold; if a baseline matches or exceeds these numbers, or if Channel or Pyramid Attention alone with CutMix matches the joint model, the central claim would be overturned.
Extended reading notes
Core claim
The central claim is that jointly applying Channel Attention and Pyramid Attention in a transformer-based U-Net improves medical image segmentation more than applying either mechanism alone or using neither. The proposed J-CaPA module processes a feature map through both attention paths in parallel, fuses them by element-wise summation, and refines the result with convolutional layers; two learnable scaling parameters, initialized to zero, let the model decide how much each attention path contributes. On the Synapse dataset the full model reports 82.29% Mean Dice and 19.74 mm HD95, and the ablation study provides the key evidence: Channel Attention alone gives 79.70 Mean Dice, Pyramid Attention alone gives 78.14, and the joint model gives 82.29, with the largest per-organ gains in the gallbladder, kidneys, and pancreas.
Load-bearing premise
The claim that J-CaPA outperforms existing methods rests on comparing the paper's own runs with numbers taken from previously published papers, under the assumption that those methods were trained under equivalent conditions; the paper does not re-run them and does not specify the compared configurations.
Editorial extensions
If this is right
- A transformer-based U-Net with the joint attention module reaches 82.29% Mean Dice and 19.74 mm HD95 on Synapse, a 6.9% Dice gain and a 39.9% HD95 reduction over the same model without the enhancements.
- The ablation results imply that the two attention mechanisms are complementary: Channel Attention alone scores 79.70 Dice and Pyramid Attention alone scores 78.14, while the joint model scores 82.29.
- The full recipe includes CutMix augmentation, which alone raises the baseline from 76.90 to 79.80 Dice; the best result requires CutMix combined with joint attention.
- The model also reports higher mean Dice and lower HD95 than published comparison methods, including a SAM-based method, on the same dataset and split.
Reading between the lines
- Because the comparison numbers for other methods are taken from published papers and not re-run under the paper's own protocol, the state-of-the-art comparison is only as strong as the assumption that those runs used identical training conditions; a reproduction study that runs all methods in one setting would settle it.
- The two gamma parameters start at zero and are learned independently, so the model begins as the baseline and gradually turns on each attention branch; tracking how these gammas evolve might reveal why the joint model beats either branch alone, but the paper does not report those curves.
- The CutMix settings (33% of images per batch, 20–60% patch area) are likely not optimized; varying the mixing fraction or using anatomy-aware mixing masks could further improve generalization on this and other datasets.
- If the joint-attention effect is general, similar gains should appear on other multi-organ or cardiac/brain segmentation benchmarks and with other transformer backbones; the paper only tests Synapse with one backbone, so this transfer remains untested.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes J-CaPA, a TransUNet-based architecture that adds Channel Attention and Pyramid Attention modules and uses CutMix augmentation, evaluated on the Synapse multi-organ CT segmentation dataset. The authors report a Mean Dice of 82.29% and HD95 of 19.74 mm, corresponding to claimed improvements of 6.9% and 39.9% over a baseline without the enhancements. The ablation study in Table 2 is intended to show that jointly applying Channel and Pyramid Attention outperforms either attention module used alone, and the paper concludes that this joint application is the key contribution.
Significance. If the empirical results hold, the paper offers a simple, modular enhancement to TransUNet and provides an instructive ablation of two attention mechanisms plus CutMix. The direction of the ablation results is coherent: joint attention scores higher than either attention alone, and CutMix contributes positively. The visual comparisons in Figure 2 are also suggestive. However, the contribution is purely empirical, and the paper does not release code or trained models, does not report variance or statistical significance, and relies on numbers from other papers for the state-of-the-art comparison. The unresolved status of the baseline in Table 2 is load-bearing: the headline improvements and the ablation ranking all trace back to that baseline row.
major comments (4)
- [Section 4, Table 2] The baseline row in Table 2 reports 76.90 Mean Dice and 32.87 mm HD95, which are exactly the values of the TransUNet row in Table 1. Section 3.6 states only that 'all parameters... were kept consistent with original specifications' and does not explicitly say whether this baseline was re-trained in the same codebase, preprocessing, augmentation, optimizer, and seed settings as the J-CaPA runs. Because the headline 6.9%/39.9% improvements and the entire ablation ranking are computed against this baseline, the central claim is not established unless the baseline was produced in the same experimental pipeline. The paper must state this explicitly; if the baseline numbers were taken from the literature, the authors need to re-run the baseline in their own setting and report the resulting scores.
- [Section 5, Table 2] All numbers in Table 2 are single-run values with no variance, confidence intervals, or statistical tests. The differences between Joint Attention (80.30), Channel Attention (79.70), Pyramid Attention (78.14), and CutMix (79.80) are small relative to typical run-to-run variability in medical segmentation. The claim that joint attention is optimal would be substantially strengthened by reporting mean and standard deviation over at least three seeds and by a paired test over the 12 test volumes, such as a Wilcoxon signed-rank test on per-case Dice scores.
- [Section 3.6, Table 1] The state-of-the-art comparison appears to use numbers published in previous papers rather than results obtained under a common experimental protocol. The manuscript does not report the image resolution, optimizer hyperparameters, learning-rate schedule, augmentation, or hardware used for the comparison methods. Since these choices materially affect Dice and HD95, the abstract's claim that the method 'outperforming existing state of the art methods' is not supported unless the comparators are re-run under the same protocol, or the authors provide a detailed table of each method's training setup and clearly frame the comparison as cross-paper.
- [Section 3.2, Section 3.3] The architecture description lacks the precision needed for reproducibility. In Section 3.2.1, the Channel Attention module's 'max-value subtraction technique' is mentioned but never defined with an equation. In Section 3.2.2, the Pyramid Attention module does not specify how the three scales are generated (e.g., pooling versus strided convolution), how the attention outputs are upsampled back to the original resolution, or what the spatial dimensions of the subsequent 3x3 convolution and bilinear upsampling are. Since no code is provided, these details are essential and should be given as explicit equations and layer specifications.
minor comments (5)
- [Abstract and Section 4] The phrase 'a 6.9% improvement in Mean Dice score' is ambiguous: 76.90 to 82.29 is 5.39 percentage points in absolute terms and about 7.0% relative. Please state explicitly which convention is used.
- [Figure 1] Figure 1 is difficult to read: the J-CaPA block is repeated without an inset diagram of its internal structure, and the path from multi-scale features to the attention modules is not clearly drawn. A zoomed-in module diagram would help.
- [Section 3.4] The CutMix hyperparameters (applied to 33% of images, segment area 20–60%) are stated without justification or sensitivity analysis; a sentence explaining the choice or citing prior usage on Synapse would improve confidence in the augmentation setup.
- [Table 1] There is a typographical issue in the SAMed row ('72.1788.72') and inconsistent spacing in the title and Table 1 entries; please proofread the final PDF.
- [References] Reference [12] is incomplete ('A Vaswani' without coauthors) and several reference entries lack full author lists or DOIs; please standardize all entries.
Circularity Check
No circularity found; the paper's claims rest on measured experimental outcomes, not on definitions or self-citations.
full rationale
The manuscript contains no mathematical derivation whose output is defined in terms of its input. The attention modules are standard components cited from prior external work, the CutMix augmentation is an input-side regularizer, and the reported Mean Dice and HD95 values are empirical measurements taken after training. The baseline row in Table 2 is an external comparison point; even if those numbers were reused from the TransUNet publication rather than re-run in the same pipeline, that would be a reproducibility/fairness limitation and not a circularity, because the full model's scores are not constructed from the baseline values. There are no self-citations, no invoked uniqueness theorem, and no fitted parameter renamed as a prediction. The central claim that jointly applying Channel and Pyramid Attention improves segmentation is therefore supported by independent experimental content rather than by definitional equivalence.
Assumptions & free parameters
free parameters (5)
- CutMix application probability =
0.33
- CutMix segment area fraction =
0.20-0.60
- Transformer depth =
12
- Channel attention blend factor gamma_CA =
learned (initialized 0)
- Pyramid attention blend factor gamma_PA =
learned (initialized 0)
assumptions (3)
- domain assumption The preprocessed Synapse dataset and its train/test split match those used by TransUNet [14].
- domain assumption The comparison numbers for other methods, taken from their papers, were produced under conditions equivalent to the authors' training setup.
- domain assumption The described network architecture, including the attention modules, is correctly implemented as claimed.
Cite this review
Pith. "Pith review of J-CaPA : Joint Channel and Pyramid Attention Improves Medical Image Segmentation." pith.science (2026). https://pith.science/paper/6ONPCT66
@misc{pith2026241116568,
author = {Pith},
title = {Pith review of: J-CaPA : Joint Channel and Pyramid Attention Improves Medical Image Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/6ONPCT66}},
note = {Machine review of arXiv:2411.16568}
}
read the original abstract
Medical image segmentation is crucial for diagnosis and treatment planning. Traditional CNN-based models, like U-Net, have shown promising results but struggle to capture long-range dependencies and global context. To address these limitations, we propose a transformer-based architecture that jointly applies Channel Attention and Pyramid Attention mechanisms to improve multi-scale feature extraction and enhance segmentation performance for medical images. Increasing model complexity requires more training data, and we further improve model generalization with CutMix data augmentation. Our approach is evaluated on the Synapse multi-organ segmentation dataset, achieving a 6.9% improvement in Mean Dice score and a 39.9% improvement in Hausdorff Distance (HD95) over an implementation without our enhancements. Our proposed model demonstrates improved segmentation accuracy for complex anatomical structures, outperforming existing state-of-the-art methods.
Reference graph
Works this paper leans on
-
[1]
INTRODUCTION Medical image segmentation is a fundamental task in clinical applications, providing the precise identification of anatom- ical structures critical for various diagnoses and treatments. However, it remains challenging due to the varying size, shape, and appearance of different organs and pathologies. While convolutional neural networks (CNNs)...
-
[2]
RELA TED WORK 2.1. CNN-Based Methods for Medical Image Segmenta- tion CNNs, including FCNs [5] and U-Net variants [1], have shown strong segmentation performance. U-Net++ [6] nar- rows the semantic gap using dense skip connections, while Attention U-Net [7] employs attention gates for feature se- lection. Models like Res-UNet [3] and R2U-Net [4] in- trodu...
-
[3]
METHODOLOGY 3.1. Model Architecture Overview The overall architecture of our model is a Transformer based structure, with an encoder-decoder design as shown in Fig- ure 1. The encoder employs Transformer blocks to capture global context, while the decoder reconstructs detailed seg- mentation maps. The encoder features a modified ResNetV2 backbone that ext...
-
[4]
Results are provided in Table 1
RESULTS We conducted an evaluation to compare the performance of our model to multiple state-of-the-art methods when segment- ing organs in abdominal CT scans. Results are provided in Table 1. We follow prior work and report results using the metrics Mean Dice score and Hausdorff Distance (HD95), facilitating a direct comparison with baseline models. Our ...
-
[5]
ABLA TION STUDY We conducted an ablation study to assess the impact of Chan- nel Attention, Pyramid Attention, and CutMix data augmen- tation on our model’s performance. The results, shown in Table 2, compare a baseline implementation without our en- hancements to the model with each enhancement added. We find that while each of Channel Attention and Pyra...
-
[6]
CONCLUSION In this paper, we propose a Transformer-based U-Net model for medical image segmentation, integrating Joint Channel and Pyramid Attention mechanisms (J-CaPA) to improve fea- ture representation. Our approach demonstrates notable im- provements in segmentation accuracy and generalization, par- ticularly for challenging organs in abdominal CT scans
-
[7]
No further ethical approval was required
COMPLIANCE WITH ETHICAL STANDARDS The Synapse multi-organ segmentation dataset used in this study is publicly available and contains de-identified, anonymized data. No further ethical approval was required
-
[8]
ACKNOWLEDGEMENT We are grateful to Prof. Yuyin Zhou for her guidance and sug- gestions throughout the duration of this project, and Vanshika Vats for her valuable feedback. No financial conflicts exist
Show all 41 references
-
[9]
Despite these advances, CNNs still struggle with model- ing long-range dependencies
and KiU-Net [2] targeted specific medical challenges. Despite these advances, CNNs still struggle with model- ing long-range dependencies. Attempts to integrate self- attention mechanisms [10, 11] improved feature selection but were insufficient in capturing global context. Th...
2024 arXiv
-
[10]
U-Net: Convo- lutional Networks for Biomedical Image Segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convo- lutional Networks for Biomedical Image Segmentation,” in Medical image computing and computer-assisted intervention– MICCAI 2015, 2015, pp. 234–241
2015
-
[11]
KiU-Net: Accurate Segmentation of Biomedical Im- ages using Over-complete Representations,
J. M. J. Valanarasu, V . A. Sindagi, I. Hacihaliloglu, and V . M. Patel, “KiU-Net: Accurate Segmentation of Biomedical Im- ages using Over-complete Representations,” in Med. Image Comput. Comput.-Assist. Intervent., 2020, pp. 363–373
2020
-
[12]
Weighted Res-UNet for High-Quality Retina Vessel Segmentation,
Xiao Xiao, Shen Lian, Zhiming Luo, and Shaozi Li, “Weighted Res-UNet for High-Quality Retina Vessel Segmentation,” in 2018 9th international conference on information technology in medicine and education (ITME) . IEEE, 2018, pp. 327–331
2018
-
[13]
Recurrent Residual Convolutional Neural Network based on U-Net (R2U-Net) for Medical Image Segmentation,
M. Z. Alom, M. Hasan, C. Yakopcic, T. M. Taha, and V . K. Asari, “Recurrent Residual Convolutional Neural Network based on U-Net (R2U-Net) for Medical Image Segmentation,” arXiv preprint arXiv:1802.06955, 2018
2018 arXiv
-
[14]
Fully Convolutional Networks for Semantic Segmentation,
J. Long, E. Shelhamer, and T. Darrell, “Fully Convolutional Networks for Semantic Segmentation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2015, pp. 3431–3440
2015
-
[15]
UNet++: A Nested U-Net Architecture for Medical Image Segmentation,
Z. Zhou, M. M. Siddiquee, N. Tajbakhsh, and J. Liang, “UNet++: A Nested U-Net Architecture for Medical Image Segmentation,” in Deep Learning Med. Image Anal. Multi- modal Clin. Decision Support , 2018, pp. 3–11
2018
-
[16]
Attention U-Net: Learning Where to Look for the Pan- creas,
O. Oktay, J. Schlemper, L. Le Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y . Hammerla, B. Kainz, et al., “Attention U-Net: Learning Where to Look for the Pan- creas,” arXiv preprint arXiv:1804.03999, 2018
2018 arXiv
-
[17]
DoubleU-Net: A Deep Convolutional Neural Net- work for Medical Image Segmentation,
D. Jha, M. A. Riegler, D. Johansen, P. Halvorsen, and H. D. Johansen, “DoubleU-Net: A Deep Convolutional Neural Net- work for Medical Image Segmentation,” in Proc. IEEE Int. Symp. Comput.-Based Med. Syst. (CBMS) , 2020, pp. 558–564
2020
-
[18]
PraNet: Parallel Reverse Attention Network for Polyp Segmentation,
D.-P. Fan, G.-P. Ji, T. Zhou, G. Chen, H. Fu, J. Shen, and L. Shao, “PraNet: Parallel Reverse Attention Network for Polyp Segmentation,” in Med. Image Comput. Comput.-Assist. Intervent., 2020, pp. 263–273
2020
-
[19]
Non-local Neural Networks,
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He, “Non-local Neural Networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7794–7803
2018
-
[20]
Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images,
J. Schlemper, O. Oktay, M. Schaap, M. Heinrich, B. Kainz, B. Glocker, and D. Rueckert, “Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images,” Med. Image Anal., vol. 53, pp. 197–207, 2019
2019
-
[21]
Attention Is All You Need,
A Vaswani, “Attention Is All You Need,” Advances in Neural Information Processing Systems, 2017
2017
-
[22]
An Image is Worth 16x16 Words: Trans- formers for Image Recognition at Scale,
Alexey Dosovitskiy, “An Image is Worth 16x16 Words: Trans- formers for Image Recognition at Scale,” arXiv preprint arXiv:2010.11929, 2020
2010 arXiv
-
[23]
TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation,
J. Chen, Y . Lu, Q. Yu, X. Luo, E. Adeli, Y . Wang, L. Lu, A. L. Yuille, and Y . Zhou, “TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation,” arXiv preprint arXiv:2102.04306, 2021
2021 arXiv
-
[24]
Swin-Unet: Unet-like Pure Transformer for Med- ical Image Segmentation,
H. Cao, Y . Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang, “Swin-Unet: Unet-like Pure Transformer for Med- ical Image Segmentation,” in Eur . Conf. Comput. Vis., 2022, pp. 205–218
2022
-
[25]
DS- TransUNet: Dual Swin Transformer U-Net for Medical Image Segmentation,
A. Lin, B. Chen, J. Xu, Z. Zhang, G. Lu, and D. Zhang, “DS- TransUNet: Dual Swin Transformer U-Net for Medical Image Segmentation,” IEEE Trans. Instrum. Meas., vol. 71, pp. 1–15, 2022
2022
-
[26]
AA-TransUNet: Attention Aug- mented TransUNet For Nowcasting Tasks,
Y . Yang and S. Mehrkanoon, “AA-TransUNet: Attention Aug- mented TransUNet For Nowcasting Tasks,” in Int. Joint Conf. Neural Netw. (IJCNN), 2022, pp. 01–08
2022
-
[27]
DA-TransUNet: Integrating Spatial and Channel Dual Attention with Transformer U-Net for Medical Image Segmentation,
G. Sun, Y . Pan, W. Kong, Z. Xu, J. Ma, T. Racharak, L.-M. Nguyen, and J. Xin, “DA-TransUNet: Integrating Spatial and Channel Dual Attention with Transformer U-Net for Medical Image Segmentation,” Front. Bioeng. Biotechnol., vol. 12, pp. 1398237, 2024
2024
-
[28]
Segment Anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, et al., “Segment Anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 4015– 4026
2023
-
[29]
SegGPT: Towards Segmenting Everything in Context,
X. Wang, X. Zhang, Y . Cao, W. Wang, C. Shen, and T. Huang, “SegGPT: Towards Segmenting Everything in Context,” in Proc. IEEE/CVF Int. Conf. Comput. Vis., 2023, pp. 1130–1140
2023
-
[30]
STU-Net: Scalable and Transferable Medical Image Segmentation Models Empow- ered by Large-Scale Supervised Pre-training,
Z. Huang, H. Wang, Z. Deng, J. Ye, Y . Su, H. Sun, J. He, Y . Gu, L. Gu, S. Zhang, et al., “STU-Net: Scalable and Transferable Medical Image Segmentation Models Empow- ered by Large-Scale Supervised Pre-training,” arXiv preprint arXiv:2304.06716, 2023
2023 arXiv
-
[31]
Segment Anything in Medical Images,
Jun Ma, Yuting He, Feifei Li, Lin Han, Chenyu You, and Bo Wang, “Segment Anything in Medical Images,” Nature Communications, vol. 15, no. 1, pp. 654, 2024
2024
-
[32]
Segment Any- thing Model for Medical Image Segmentation: Current ap- plications and future directions,
Yichi Zhang, Zhenrong Shen, and Rushi Jiao, “Segment Any- thing Model for Medical Image Segmentation: Current ap- plications and future directions,” Computers in Biology and Medicine, p. 108238, 2024
2024
-
[33]
How to efficiently adapt large segmentation model (SAM) to medical images,
Xinrong Hu, Xiaowei Xu, and Yiyu Shi, “How to efficiently adapt large segmentation model (SAM) to medical images,” arXiv preprint arXiv:2306.13731, 2023
2023 arXiv
-
[34]
Customized Segment Any- thing Model for Medical Image Segmentation,
Kaidong Zhang and Dong Liu, “Customized Segment Any- thing Model for Medical Image Segmentation,” arXiv preprint arXiv:2304.13785, 2023
2023 arXiv
-
[35]
LoRA: Low-Rank Adaptation of Large Language Models,
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” arXiv preprint arXiv:2106.09685 , 2021
2021 arXiv
-
[36]
TransNorm: Transformer Provides a Strong Spatial Normal- ization Mechanism for a Deep Segmentation Model,
R. Azad, M. T. Al-Antary, M. Heidari, and D. Merhof, “TransNorm: Transformer Provides a Strong Spatial Normal- ization Mechanism for a Deep Segmentation Model,” IEEE Access, vol. 10, pp. 108205–108215, 2022
2022
-
[37]
IB-TransUNet: Combin- ing Information Bottleneck and Transformer for Medical Im- age Segmentation,
G. Li, D. Jin, Q. Yu, and M. Qi, “IB-TransUNet: Combin- ing Information Bottleneck and Transformer for Medical Im- age Segmentation,” J. King Saud Univ. Comput. Inf. Sci. , vol. 35, no. 3, pp. 249–258, 2023
2023
-
[38]
Dual Attention Network for Scene Segmentation,
J. Fu, J. Liu, H. Tian, Y . Li, Y . Bao, Z. Fang, and H. Lu, “Dual Attention Network for Scene Segmentation,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2019, pp. 3146–3154
2019
-
[39]
TA-Net: Triple Attention Network for Medical Image Segmentation,
Yang Li, Jun Yang, Jiajia Ni, Ahmed Elazab, and Jianhuang Wu, “TA-Net: Triple Attention Network for Medical Image Segmentation,” Computers in Biology and Medicine , vol. 137, pp. 104836, 2021
2021
-
[40]
A Multi- Class COVID-19 Segmentation Network with Pyramid Atten- tion and Edge Loss in CT Images,
F. Yu, Y . Zhu, X. Qin, Y . Xin, D. Yang, and T. Xu, “A Multi- Class COVID-19 Segmentation Network with Pyramid Atten- tion and Edge Loss in CT Images,” IET Image Process. , vol. 15, no. 11, pp. 2604–2613, 2021
2021
-
[41]
Cut- Mix: Regularization Strategy to Train Strong Classifiers With Localizable Features,
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y . Yoo, “Cut- Mix: Regularization Strategy to Train Strong Classifiers With Localizable Features,” in Proc. IEEE/CVF Int. Conf. Comput. Vis., 2019, pp. 6023–6032
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.