REVIEW 4 major objections 6 minor 37 references
A4-Unet: Deformable Multi-Scale Attention Network for Brain Tumor Segmentation
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper proposes A4-Unet, a U-shaped convolutional network that reports 94.4% Dice on BraTS 2020 and claims new state-of-the-art results on three brain tumor segmentation benchmarks.
desk verdict Competent assembly of known attention modules for BraTS, but the reported mIoU is arithmetically impossible given the paper's own DSC, so the SOTA claims do not survive contact with the equations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the A4-Unet architecture, a U-Net with four interacting attention and multi-scale components. DLKA uses deformable convolution together with a large depth-wise dilated kernel to move sampling points onto irregular tumor boundaries. SSPP replaces the dilated convolutions of an ASPP-style bottleneck with Swin Transformer blocks of different window sizes, then applies cross-contextual attention to reweight the multi-scale features. CAM computes channel weights from Discrete Cosine Transform coefficients under an orthogonality constraint and computes spatial weights by convolutional element-wise multiplication of channel max and average. Attention gates in the skip connections multiply encoder features by gating coefficients derived from coarser decoder features. The paper's argument is that these mechanisms jointly supply shape adaptivity, long-range context, channel selectivity, and background suppression, and that the whole is what produces the reported scores.
What would settle it
Running the released code on the official BraTS 2020 test set through the challenge evaluation portal would settle the claim: the reported 94.47% Dice must reproduce. A second concrete check is metric consistency, since the standard foreground/background relationship would put IoU near 89.5% for a 94.47% Dice, not 99.68%; a per-class IoU table or the exact mIoU definition used would resolve the discrepancy.
Extended reading notes
Core claim
The central discovery the paper is trying to establish is that four previously separate mechanisms work synergistically inside a U-Net: DLKA adapts convolutional sampling to irregular tumor shapes, SSPP provides multi-scale global context through shifted-window transformers, CAM weights channels via DCT-based orthogonal attention and suppresses irrelevant spatial regions, and attention gates filter the skip-connection features. The paper argues that this combination, rather than a larger model or full 3D processing, is what pushes BraTS segmentation accuracy past published baselines. Its evidence is Table IV, where A4-Unet reaches 94.61% Dice on BraTS 2019, 94.47% on BraTS 2020, and 92.84% on BraTS 2021, alongside 84.18% Dice on a proprietary two-modality clinical dataset.
Load-bearing premise
The central claim depends on the assumption that the model was evaluated on the official held-out BraTS test labels with the standard definitions of Dice, mIoU, and HD95, so that Table IV is directly comparable to the published baselines.
Editorial extensions
If this is right
- If the reported results are correct, A4-Unet is the top published method on BraTS 2020 with 94.47% Dice, above nnU-Net's 91.18%.
- The architecture would show that CNN encoders with deformable large kernels can match or beat transformer encoders on brain tumor segmentation at comparatively low complexity, training on a single 24 GB GPU in about 30 hours for BraTS 2020.
- The 2D-slice design would make the method easy to reproduce and adapt to other MRI segmentation tasks without requiring 3D convolutions.
- The attention-gated skip connections and DCT-based channel attention would be directly reusable components for other medical segmentation networks.
Reading between the lines
- The paper does not report class-wise Dice for enhancing tumor, tumor core, and whole tumor, so a natural extension is to publish those three values; the overall Dice may hide large differences across BraTS sub-regions.
- Because the model is 2D and trained slice-by-slice, the results suggest neighboring-slice context can be compensated by large-kernel attention; one could test this by adding a third dimension and measuring whether Dice improves further.
- The DCT-based channel weighting may transfer to other small-lesion segmentation tasks where global average pooling washes out rare foreground classes; the paper does not test this transfer itself.
- The authors themselves note in Section IV-G that clinical applicability is limited by data diversity and annotation scarcity, and the proprietary dataset score of 84.18% Dice is well below the BraTS numbers, so the headline benchmark gains should be read as benchmark achievements rather than clinical readiness.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes A4-Unet, an encoder-decoder for brain tumor segmentation that combines deformable large-kernel attention (DLKA) in the encoder, a Swin spatial pyramid pooling bottleneck with cross-contextual attention (SSPP), a combined attention module (CAM) using DCT-based orthogonal channel attention, and attention gates in skip connections. The authors report DSC, mIoU, and HD95 on BraTS 2019, 2020, 2021, and a proprietary dataset, and claim multiple new state-of-the-art benchmarks, notably 94.47% DSC and 99.68% mIoU on BraTS 2020. The major problem is that the reported evaluation numbers are internally inconsistent with the manuscript's own metric definitions in Eqs. (12)-(13), and the evaluation protocol is too underspecified to support the performance and SOTA claims.
Significance. If the reported results were valid, the paper would demonstrate that a U-Net assembled from recently published attention and transformer components improves brain tumor segmentation over strong baselines; that would be a useful engineering contribution. The paper's strengths are its multi-benchmark scope, the inclusion of an ablation study, and the intent to release code. However, the architecture is largely an integration of existing modules (DLKA, SSPP, OrthoNet-based channel attention, and attention gates), and the quantitative gains are the main basis for the claimed significance. Because the central table is arithmetically impossible under the paper's own metric definitions and the evaluation lacks basic provenance information, the current manuscript does not establish a new state of the art.
major comments (4)
- [§IV-B, §IV-E (Table IV)] The BraTS 2020 pair (DSC 94.47%, mIoU 99.68%) is arithmetically impossible under the definitions in Eqs. (12) and (13). For any class, IoU = TP/(TP+FP+FN) and DSC = 2TP/(2TP+FP+FN), hence IoU = DSC/(2−DSC) ≤ DSC. If 94.47% is the mean DSC over the three foreground tumor classes, then even a perfect background class gives mIoU ≤ (1 + 3×0.9447)/4 ≈ 95.85%, far below 99.68%. If 94.47% is instead a whole-tumor Dice, the corresponding whole-tumor IoU is about 89.5% and the four-class mIoU is lower still. Since this table carries the abstract's claim of 'multiple new state-of-the-art benchmarks,' the central empirical claim is unsupported by the manuscript's own evidence.
- [§IV-A, §IV-E] The evaluation protocol is underspecified. The text states that the results 'represent the average of five independent runs and were subjected to cross-validation,' but it does not report standard deviations or per-fold results, and it never states whether the 'Testing set' entries in Table II are the BraTS challenge validation sets, the hidden challenge test sets, or locally held-out partitions. BraTS test labels are not generally released to users on request; if the authors had access to them, the source and terms of that access must be documented. Without this information, the numbers in Tables III-V cannot be independently checked, and the comparison with literature values in Table V may mix different evaluation protocols.
- [§IV-F, Table V] The text says A4-Unet outperforms nnU-Net on BraTS 2020, but the same table reports HD95 of 8.57 mm for A4-Unet versus 8.49 mm for nnU-Net, so A4-Unet is worse on that metric. Official BraTS ranking uses multiple metrics and a challenge champion is not necessarily the best on every metric; the claim of superiority should therefore be limited to the metrics in which A4-Unet actually improves, and ideally supported by matched evaluation settings.
- [§IV-D, Table III] The ablation study reports only point estimates of DSC, yet the text claims the results are averages of five runs. The componentwise improvements (e.g., 1.3% for DLKA, 2.0% for SSPP, and 1.9% for CAM) are presented without variance or significance testing, so it is unclear whether these differences exceed run-to-run noise for brain tumor segmentation. This weakens the specific causal claims made for each module.
minor comments (6)
- [§III-D, Eq. (11)] The same symbol σ is used to denote both the sigmoid and ReLU functions, and the sigmoid expression contains an undefined factor a; use distinct symbols and write the sigmoid as 1/(1+e^{-x}).
- [§III-B, Eq. (1)] The parentheses in 'Conv 1×1(Conv DC (Conv DW (F ))' are unbalanced.
- [§III-B, Eq. (4)] The expression 'K/d' should be written with an integer division or floor operation, and the kernel-size variables KDW and KDC should be defined consistently with the dilations used later.
- [§IV-E] The phrase '95% might not be the optimal hyperparameter' is confusing because HD95 is an evaluation metric, not a hyperparameter; the explanation for cross-dataset HD95 differences should be revised.
- [Abstract and footnote] The GitHub URL in the full text contains a stray space ('WendyW AAAAANG'); the link should be corrected and verified before publication.
- [Table V] Several comparison rows have missing metric values, and the source and evaluation protocol for each cited number should be provided so that the comparison is reproducible and fair.
Circularity Check
No circularity: A4-Unet is an empirical architecture combination evaluated on external BraTS benchmarks; no derivation in the paper reduces to its own inputs.
full rationale
The paper does not claim to derive a theoretical result. Its methodology defines standard segmentation metrics (DSC in Eq. 12, mIoU in Eq. 13, HD in Eqs. 14-16), then reports empirical performance of a network assembled from previously published components (DLKA from Azad et al., SSPP from Azad et al., OCA from Salman et al., and additive attention gates). Each borrowed component is attributed to an external source rather than justified by a self-citation chain, so there is no load-bearing self-citation and no uniqueness theorem imported from the authors' own prior work. The ablation study compares the full model to a ResUnet baseline and to component combinations; this is standard empirical attribution and does not fit a parameter to a target result and then rename that fit a prediction. The central SOTA claim is carried by Table IV numbers. Those numbers are internally suspicious (mIoU of 99.68% is arithmetically incompatible with DSC of 94.47% under the paper's own Eqs. 12-13), and the paper does not document access to held-out BraTS test labels, but these are evaluation-validity and correctness concerns, not circularity: they do not show that any claimed result is equivalent to an input by definition. Accordingly, per the hard rule that non-findings are expected when warranted, the circularity score is 0.
Assumptions & free parameters
free parameters (5)
- learning rate =
1e-5
- batch size =
16
- training epochs =
30
- DLKA kernel size K and dilation d =
not reported
- SSPP Swin window sizes =
not reported
assumptions (4)
- ad hoc to paper BraTS held-out test labels are accessible to the authors
- domain assumption The mIoU metric is computed with the standard class-average definition
- standard math DCT orthogonality preserves channel information used for attention
- domain assumption 2D slicing preserves the 3D tumor structure adequately for segmentation
Cite this review
Pith. "Pith review of A4-Unet: Deformable Multi-Scale Attention Network for Brain Tumor Segmentation." pith.science (2026). https://pith.science/paper/WYOTBFJ7
@misc{pith2026241206088,
author = {Pith},
title = {Pith review of: A4-Unet: Deformable Multi-Scale Attention Network for Brain Tumor Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/WYOTBFJ7}},
note = {Machine review of arXiv:2412.06088}
}
read the original abstract
Brain tumor segmentation models have aided diagnosis in recent years. However, they face MRI complexity and variability challenges, including irregular shapes and unclear boundaries, leading to noise, misclassification, and incomplete segmentation, thereby limiting accuracy. To address these issues, we adhere to an outstanding Convolutional Neural Networks (CNNs) design paradigm and propose a novel network named A4-Unet. In A4-Unet, Deformable Large Kernel Attention (DLKA) is incorporated in the encoder, allowing for improved capture of multi-scale tumors. Swin Spatial Pyramid Pooling (SSPP) with cross-channel attention is employed in a bottleneck further to study long-distance dependencies within images and channel relationships. To enhance accuracy, a Combined Attention Module (CAM) with Discrete Cosine Transform (DCT) orthogonality for channel weighting and convolutional element-wise multiplication is introduced for spatial weighting in the decoder. Attention gates (AG) are added in the skip connection to highlight the foreground while suppressing irrelevant background information. The proposed network is evaluated on three authoritative MRI brain tumor benchmarks and a proprietary dataset, and it achieves a 94.4% Dice score on the BraTS 2020 dataset, thereby establishing multiple new state-of-the-art benchmarks. The code is available here: https://github.com/WendyWAAAAANG/A4-Unet.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Segnext: Rethinking convolutional attention design for semantic segmentation,
M.-H. Guo, C.-Z. Lu, Q. Hou, Z. Liu, M.-M. Cheng, and S.-M. Hu, “Segnext: Rethinking convolutional attention design for semantic segmentation,” Advances in Neural Information Processing Systems , vol. 35, pp. 1140–1156, 2022
2022
-
[2]
Densely connected convolutional networks,
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE confer- ence on computer vision and pattern recognition , 2017, pp. 4700–4708
2017
-
[3]
nnu- net: a self-configuring method for deep learning-based biomedical image segmentation,
F. Isensee, P. Jaeger, S. Kohl, J. Petersen, and K. Maier-Hein, “nnu- net: a self-configuring method for deep learning-based biomedical image segmentation,” Nature Methods, vol. 18, pp. 1–9, 02 2021
work page 2021
-
[4]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020
arXiv 2010
-
[5]
Segformer: Simple and efficient design for semantic segmentation with transformers,
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” Advances in neural information processing systems , vol. 34, pp. 12 077–12 090, 2021
2021
-
[6]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10 012–10 022
2021
-
[7]
Transattunet: Multi-level attention-guided u-net with transformer for medical image segmentation,
B. Chen, Y . Liu, Z. Zhang, G. Lu, and A. W. K. Kong, “Transattunet: Multi-level attention-guided u-net with transformer for medical image segmentation,” IEEE Transactions on Emerging Topics in Computational Intelligence, 2023
2023
-
[8]
Bottleneck transformers for visual recognition,
A. Srinivas, T.-Y . Lin, N. Parmar, J. Shlens, P. Abbeel, and A. Vaswani, “Bottleneck transformers for visual recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 16 519–16 529
work page 2021
Show all 37 references
-
[9]
Squeeze-and-excitation networks,
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141
2018
-
[10]
Fcanet: Frequency channel attention networks,
Z. Qin, P. Zhang, F. Wu, and X. Li, “Fcanet: Frequency channel attention networks,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 783–792
2021
-
[11]
Cbam: Convolutional block attention module,
S. Woo, J. Park, J.-Y . Lee, and I. S. Kweon, “Cbam: Convolutional block attention module,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 3–19
2018
-
[12]
A real-time algorithm for signal analysis with the help of the wavelet transform,
M. Holschneider, R. Kronland-Martinet, J. Morlet, and P. Tchamitchian, “A real-time algorithm for signal analysis with the help of the wavelet transform,” in Wavelets: Time-Frequency Methods and Phase Space Pro- ceedings of the International Conference, Marseille, France, Dece...
1987
-
[14]
Object detection with discriminatively trained part-based models,
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan, “Object detection with discriminatively trained part-based models,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 32, no. 9, pp. 1627–1645, 2010
2010
-
[15]
Spatial transformer networks,
M. Jaderberg, K. Simonyan, A. Zisserman et al. , “Spatial transformer networks,” Advances in neural information processing systems , vol. 28, 2015
2015
-
[16]
Deformable convolutional networks,
J. Dai, H. Qi, Y . Xiong, Y . Li, G. Zhang, H. Hu, and Y . Wei, “Deformable convolutional networks,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 764–773
2017
-
[17]
Multi-scale context aggregation by dilated convolutions,
F. Yu and V . Koltun, “Multi-scale context aggregation by dilated convolutions,” arXiv preprint arXiv:1511.07122 , 2015
2015 arXiv
-
[18]
Spatial pyramid pooling in deep convolutional networks for visual recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 37, no. 9, pp. 1904– 1916, 2015
1904
-
[19]
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 4, pp. 834–848, 2017
2017
-
[20]
Crossvit: Cross-attention multi- scale vision transformer for image classification,
C.-F. R. Chen, Q. Fan, and R. Panda, “Crossvit: Cross-attention multi- scale vision transformer for image classification,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 357– 366
2021
-
[21]
Multiscale vision transformers,
H. Fan, B. Xiong, K. Mangalam, Y . Li, Z. Yan, J. Malik, and C. Feichten- hofer, “Multiscale vision transformers,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 6824–6835
2021
-
[22]
Beyond self-attention: Deformable large kernel attention for medical image segmentation,
R. Azad, L. Niggemeier, M. H ¨uttemann, A. Kazerouni, E. K. Agh- dam, Y . Velichko, U. Bagci, and D. Merhof, “Beyond self-attention: Deformable large kernel attention for medical image segmentation,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer ...
2024
-
[23]
Visual attention network,
M.-H. Guo, C.-Z. Lu, Z.-N. Liu, M.-M. Cheng, and S.-M. Hu, “Visual attention network,” Computational Visual Media, vol. 9, no. 4, pp. 733– 752, 2023
2023
-
[24]
Location sensitive deep convolutional neural networks for segmentation of white matter hyperintensities,
M. Ghafoorian, N. Karssemeijer, T. Heskes, I. W. van Uden, C. I. Sanchez, G. Litjens, F.-E. de Leeuw, B. van Ginneken, E. Marchiori, and B. Platel, “Location sensitive deep convolutional neural networks for segmentation of white matter hyperintensities,” Scientific Reports , v...
2017
-
[25]
Segmentation of glioma tumors in brain using deep convolutional neural network,
S. Hussain, S. M. Anwar, and M. Majid, “Segmentation of glioma tumors in brain using deep convolutional neural network,” Neurocom- puting, vol. 282, pp. 248–261, 2018
2018
-
[26]
Rethinking atrous convolution for semantic image segmentation,
L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking atrous convolution for semantic image segmentation,” arXiv preprint arXiv:1706.05587, 2017
2017 arXiv
-
[27]
Transdeeplab: Convolution-free transformer- based deeplab v3+ for medical image segmentation,
R. Azad, M. Heidari, M. Shariatnia, E. K. Aghdam, S. Karimijafarbigloo, E. Adeli, and D. Merhof, “Transdeeplab: Convolution-free transformer- based deeplab v3+ for medical image segmentation,” in International Workshop on PRedictive Intelligence In MEdicine . Springer, 2022, p...
2022
-
[28]
Orthonets: Orthogonal channel attention networks,
H. Salman, C. Parks, M. Swan, and J. Gauch, “Orthonets: Orthogonal channel attention networks,” arXiv preprint arXiv:2311.03071 , 2023
2023 arXiv
-
[29]
Disan: Direc- tional self-attention network for rnn/cnn-free language understanding,
T. Shen, T. Zhou, G. Long, J. Jiang, S. Pan, and C. Zhang, “Disan: Direc- tional self-attention network for rnn/cnn-free language understanding,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
-
[30]
Transunet: Transformers make strong encoders for medical image segmentation,
J. Chen, Y . Lu, Q. Yu, X. Luo, E. Adeli, Y . Wang, L. Lu, A. L. Yuille, and Y . Zhou, “Transunet: Transformers make strong encoders for medical image segmentation,” arXiv preprint arXiv:2102.04306 , 2021
2021 arXiv
-
[31]
Swin-unet: Unet-like pure transformer for medical image segmenta- tion,
H. Cao, Y . Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang, “Swin-unet: Unet-like pure transformer for medical image segmenta- tion,” in European conference on computer vision . Springer, 2022, pp. 205–218
2022
-
[32]
Resunet+: A new convolutional and attention block-based approach for brain tumor segmentation,
S. Metlek and H. C ¸ etıner, “Resunet+: A new convolutional and attention block-based approach for brain tumor segmentation,” IEEE Access , vol. 11, pp. 69 884–69 902, 2023
2023
-
[33]
Two-stage cascaded u- net: 1st place solution to brats challenge 2019 segmentation task,
Z. Jiang, C. Ding, M. Liu, and D. Tao, “Two-stage cascaded u- net: 1st place solution to brats challenge 2019 segmentation task,” in Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries, A. Crimi and S. Bakas, Eds. Cham: Springer International Publishin...
2019
-
[34]
Unetr: Transformers for 3d medical image segmentation,
A. Hatamizadeh, Y . Tang, V . Nath, D. Yang, A. Myronenko, B. Land- man, H. R. Roth, and D. Xu, “Unetr: Transformers for 3d medical image segmentation,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision , 2022, pp. 574–584
2022
-
[35]
Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images,
A. Hatamizadeh, V . Nath, Y . Tang, D. Yang, H. R. Roth, and D. Xu, “Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images,” in International MICCAI Brainlesion Workshop. Springer, 2021, pp. 272–284
2021
-
[36]
Redundancy reduction in semantic segmentation of 3d brain tumor mris,
M. M. Rahman Siddiquee and A. Myronenko, “Redundancy reduction in semantic segmentation of 3d brain tumor mris,” inInternational MICCAI Brainlesion Workshop. Springer, 2021, pp. 163–172
2021
-
[37]
Optimized u- net for brain tumor segmentation,
M. Futrega, A. Milesi, M. Marcinkiewicz, and P. Ribalta, “Optimized u- net for brain tumor segmentation,” in International MICCAI brainlesion workshop. Springer, 2021, pp. 15–29
2021
-
[38]
Coupling nnu-nets with expert knowledge for accurate brain tumor segmentation from mri,
K. Kotowski, S. Adamski, B. Machura, L. Zarudzki, and J. Nalepa, “Coupling nnu-nets with expert knowledge for accurate brain tumor segmentation from mri,” in Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries , A. Crimi and S. Bakas, Eds. Cham: Spring...
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.