REVIEW 5 major objections 6 minor 55 references
A Cascaded Dilated Convolution Approach for Mpox Lesion Classification
T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A cascaded dilated-attention module claims top accuracy on three Mpox lesion datasets while cutting parameters by over a third.
desk verdict A modest, honestly assembled architecture with one clean dataset result; the MSLD claim is undermined by augmentation-before-split leakage, and the MCSI 'SOTA' gap is not statistically significant. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Cascaded Atrous Group Attention (CAGA) module, which stacks two mechanisms. The Cascaded Atrous Attention (CAA) part applies dilated convolutions (rates 1, 2, 3) to each attention head, runs self-attention per dilated map, and adds each map's output to the next to create multi-scale context. The Cascaded Group Attention (CGA) part splits the input into heads and gives each head the previous head's output added to it, so information flows across heads and head redundancy is reduced. This module replaces the standard self-attention in the EfficientViT-L1 backbone, turning it into EfficientViT-CAGA; the cascade and the atrous rates are what carry the multi-scale argument.
What would settle it
Split the original 228 MSLD images into train, validation, and test sets first, then apply data augmentation only to the training partition, retrain EfficientViT-CAGA, and measure test accuracy; if it drops substantially from 0.9969, the cross-dataset SOTA claim loses the MSLD pillar.
Extended reading notes
Core claim
The central claim is that CAGA, a module that nests Cascaded Atrous Attention (CAA) inside Cascaded Group Attention (CGA), lets an EfficientViT-L1 model beat all compared CNN and vision-transformer baselines on three Mpox lesion datasets. CAA runs dilated convolutions with rates 1, 2, 3 on each attention head, computes self-attention on each dilated feature map, then cascades the attention outputs across scales, while CGA adds each processed head's output to the next head to cut redundancy. Integrated with the EfficientViT-L1 backbone, the model reaches 98% accuracy on the MCSI dataset, outperforms baselines on MSID and MSLD, and uses 36.8M parameters and 4.86G FLOPs, about 37.5% fewer parameters than the backbone. The paper also reports balanced per-class accuracy on MSID and near-perfect Mpox recall on MSLD, and it offers an ablation study showing that the cascade and the CGA wrapper each contribute to the gain.
Load-bearing premise
The reported state-of-the-art accuracy assumes each benchmark's split is clean, and in particular that the augmented MSLD images used for training do not reappear in the test set; if they do, the 99.69% accuracy on that dataset would measure memorization, not generalization to new patients.
Editorial extensions
If this is right
- If the accuracy numbers hold, the model could run on edge devices for Mpox screening, since it uses only 4.86G FLOPs and 36.8M parameters, less than most compared architectures.
- The balanced class-wise accuracy on MSID suggests the module counteracts the class-specific overfitting that the compared CNNs and ViTs show, which would be valuable in real-world triage where all lesion types must be recognized.
- The module is defined independently of the task head, so it could be dropped into other vision-transformer backbones for skin-lesion classification or other fine-grained medical image tasks.
- The paper's ablation attributes part of the gain to the cascading and the CGA wrapper, implying that the hierarchy itself, not just enlarged receptive fields, is responsible for the improvement.
Reading between the lines
- The MSLD evaluation splits the data after augmentation, so the test set can contain near-duplicates of training images; re-running the comparison with augmentation applied only after splitting would show whether the reported 99.69% accuracy reflects real generalization or memorization of augmented variants.
- The dilation rates are fixed per dataset, so the module's success may depend on those specific rates; testing other rate schedules or making the rates learnable could separate the effect of the cascade structure from the effect of the chosen receptive fields.
- Because the parameter counts given for the same baseline models differ slightly across tables (e.g., ResNet-101 has 43.5M parameters in the MSID table but 47.4M in the MCSI table), the exact 37.5% reduction depends on reproducible parameter counting, which an independent re-count could verify.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a Cascaded Atrous Group Attention (CAGA) module that combines a cascaded atrous attention (CAA) mechanism with cascaded group attention (CGA), and integrates it with the EfficientViT-L1 backbone for Mpox lesion classification. The authors report state-of-the-art accuracy on three datasets (MCSI, MSID, MSLD) with reduced parameters and FLOPs relative to several baseline ViT and CNN models, and provide an ablation study and Grad-CAM visualizations. The central claim is that EfficientViT-CAGA achieves higher accuracy than existing methods across all evaluated datasets while being computationally efficient.
Significance. If the claims are validated, the CAGA module would be a lightweight, accurate alternative for Mpox classification, which is clinically useful given the need for fast and reliable screening. The paper's strengths include evaluation on three public datasets, comparison with a reasonable set of modern baselines, an ablation study, and an explicit efficiency analysis. However, the supporting evidence for the central SOTA claim is weakened by methodological concerns in the MSLD evaluation, a lack of statistical significance testing on the MCSI result, and inconsistencies in reported parameter counts. The architecture itself is a plausible incremental improvement over existing attention mechanisms, but the empirical validation currently does not meet the bar for a definitive SOTA claim.
major comments (5)
- [IV.C.2] Section IV.C.2 states that for MSLD the dataset was first expanded by augmentation and then divided into training, validation, and testing sets using a 70:20:10 split. Augmentations of the same original image (rotation, translation, reflection, shear, etc.) can therefore appear in both training and test partitions, creating near-duplicate leakage. The near-perfect MSLD accuracy (0.9969, 1.0 for the Mpox class) and the claimed consistent outperformance on MSLD are not reliable evidence of generalization under this protocol. A split-before-augmentation protocol is required before MSLD can support the central claim.
- [Table II] On the MCSI dataset, EfficientViT-CAGA achieves 0.98 ± 0.0229 accuracy, while the EfficientViT-L1 backbone achieves 0.9725 ± 0.018. The difference of 0.0075 is well within one standard deviation of either result, so the reported improvement is not statistically significant. The manuscript should provide a paired significance test (e.g., a paired t-test or Wilcoxon test across the 10 folds) or confidence intervals before claiming improved accuracy over the backbone.
- [Tables I, II, and IV] The parameter counts for the same architectures are inconsistent across tables: EfficientViT-L1 is reported as 58.9M in Tables I and II but 42M in Table IV, and EfficientViT-CAGA is reported as 37.8M in Table I but 36.8M in Tables II and IV. These inconsistencies (also affecting the stated 37.5% and 35.8% parameter reductions) make the efficiency claims difficult to verify. The authors should state the exact parameter counting method and ensure all tables use consistent model configurations.
- [Tables I and IV] Results on MSID (Table I) and MSLD (Table IV) are reported as point estimates with no standard deviations or confidence intervals, even though MCSI uses 10-fold cross-validation. Without error bars, the claimed 'consistent outperformance' over DeiT3-Medium (0.9545 vs 0.9481 on MSID; 0.9969 vs 0.9938 on MSLD) cannot be assessed for statistical significance. The authors should report variance across multiple splits or folds for these datasets as well.
- [IV.B] The manuscript states that dilation rates are 'adjusted for each dataset' and that focal loss uses 'class-specific weights (alpha),' implying per-dataset tuning of these hyperparameters on the same benchmark test sets. Reporting tuned results as SOTA without external validation or a nested cross-validation procedure risks selection bias. The authors should either fix hyperparameters a priori, provide sensitivity analyses, or validate on a held-out external dataset.
minor comments (6)
- [V] The conclusion says 'the approach highlighted state-of-the-art results,' which is awkward phrasing; consider rewording to 'achieved state-of-the-art results.'
- [IV.B] The implementation details state that the learning rate is 'on the order of 10^-5,' which is imprecise; the exact value should be reported for reproducibility.
- [III.A] Equation (2) uses the index i in the left-hand side but the cascading in d is not fully defined; clarify the indexing of heads and dilations to avoid ambiguity.
- [Abstract] The phrase 'reducing model parameters by 37.5% compared to the original EfficientViT-L1' is only accurate for the MCSI configuration (58.9M vs 36.8M); the percentages differ for MSID and MSLD, so state the specific comparison.
- [IV.C.1] The statement that 'runtime data augmentation techniques were applied to the training set' and caused 'a marginal decrease in performance' is vague; specify which augmentation techniques were used and how this was applied consistently across models.
- [I] The code availability line says 'The code will be available here' but no link is provided; please include a URL or state a clear availability plan.
Circularity Check
MSLD split-after-augmentation makes the reported 99.69% accuracy a near-duplicate recognition result; otherwise the paper is empirical and not circular.
-
fitted input called prediction
[Section IV.C.2 (MSLD) and Table IV]
"These augmentations expanded the dataset to over 1,428 images for the Mpox class and 1,764 images for the Others class. The augmented dataset was then divided into training, validation, and testing sets using a 70:20:10 split ratio... Notably, EfficientViT-CAGA is the only model to achieve 100% accuracy on the Mpox class."
Because the split occurs after augmentation, augmented copies of the same original images appear in both the training and test partitions. The test predictions are therefore made on near-duplicates of training images, so the 0.9969 accuracy and 1.0 Mpox-class recall on MSLD measure memorization of augmented variants rather than generalization to new lesions. The paper uses MSLD as one of the three datasets supporting the abstract's claim that the model 'consistently outperforms existing approaches,' so this particular prediction is forced by the evaluation construction rather than by independent generalization. Comparing all baselines under the same leaky protocol does not neutralize the problem, because different models memorize augmented duplicates to different degrees.
full rationale
No analytic derivation is offered; the CAGA module is an empirical combination of dilated convolutions (after ASPP) and cascaded group attention (after EfficientViT's CGA). I found no self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation. The only circular step is the MSLD evaluation: splitting after augmentation places augmented copies of the same originals on both sides of the train/test boundary, so the reported MSLD accuracy is not independent evidence for the SOTA claim. The MCSI and MSID evaluations, while possibly affected by per-dataset hyperparameter adjustment, are not circular under the strict standard used here. Overall, one prediction (MSLD) reduces by construction, giving partial circularity; the central architecture claim still retains independent empirical content on the other two datasets.
Assumptions & free parameters
free parameters (3)
- Dilation rates d in CAA =
d ∈ {1,2,3}, adjusted per dataset
- Focal loss alpha (class-specific weights) =
not reported
- Number of heads, head dimension, and dqkv =
3 heads, embedding dim 16, dqkv=8
assumptions (4)
- domain assumption ImageNet pretraining transfers to medical skin lesion images
- domain assumption The three public datasets (MCSI, MSID, MSLD) are representative of real-world Mpox presentations
- standard math Standard scaled dot-product attention (Eq. 1) is accepted
- domain assumption MSLD augmentation before split creates independent test samples
Cite this review
Pith. "Pith review of A Cascaded Dilated Convolution Approach for Mpox Lesion Classification." pith.science (2026). https://pith.science/paper/VFVTGLPN
@misc{pith2026241210106,
author = {Pith},
title = {Pith review of: A Cascaded Dilated Convolution Approach for Mpox Lesion Classification},
year = {2026},
howpublished = {\url{https://pith.science/paper/VFVTGLPN}},
note = {Machine review of arXiv:2412.10106}
}
read the original abstract
The global outbreak of the Mpox virus, classified as a Public Health Emergency of International Concern (PHEIC) by the World Health Organization, presents significant diagnostic challenges due to its visual similarity to other skin lesion diseases. Traditional diagnostic methods for Mpox, which rely on clinical symptoms and laboratory tests, are slow and labor intensive. Deep learning-based approaches for skin lesion classification offer a promising alternative. However, developing a model that balances efficiency with accuracy is crucial to ensure reliable and timely diagnosis without compromising performance. This study introduces the Cascaded Atrous Group Attention (CAGA) framework to address these challenges, combining the Cascaded Atrous Attention module and the Cascaded Group Attention mechanism. The Cascaded Atrous Attention module utilizes dilated convolutions and cascades the outputs to enhance multi-scale representation. This is integrated into the Cascaded Group Attention mechanism, which reduces redundancy in Multi-Head Self-Attention. By integrating the Cascaded Atrous Group Attention module with EfficientViT-L1 as the backbone architecture, this approach achieves state-of-the-art performance, reaching an accuracy of 98% on the Mpox Close Skin Image (MCSI) dataset while reducing model parameters by 37.5% compared to the original EfficientViT-L1. The model's robustness is demonstrated through extensive validation on two additional benchmark datasets, where it consistently outperforms existing approaches.
Figures
Reference graph
Works this paper leans on
-
[1]
Multi-country monkeypox outbreak in non- endemic countries, May 2022
World Health Organization. Multi-country monkeypox outbreak in non- endemic countries, May 2022
work page 2022
-
[2]
Catharina E van Ewijk, Fuminari Miura, Gini van Rijckevorsel, Henry JC de Vries, Matthijs RA Welkers, Oda E van den Berg, Ingrid HM Friesema, Patrick R van den Berg, Thomas Dalhuisen, Jacco Wallinga, Diederik Brandwagt, Brigitte AGL van Cleef, Harry Vennema, Bettie V oordouw, Marion Koopmans, Annemiek A van der Eijk, Corien M Swaan, Margreet JM te Wierik,...
work page 2022
-
[3]
Overview of mpox outbreak in greece in 2022–2023: Is it over? Viruses, 15(6), 2023
Kassiani Mellou, Kyriaki Tryfinopoulou, Styliani Pappa, Kassiani Gkolfinopoulou, Sofia Papanikou, Georgia Papadopoulou, Evangelia Vas- sou, Evangelia-Georgia Kostaki, Kalliopi Papadima, Elissavet Mouratidou, Maria Tsintziloni, Nikolaos Siafakas, Zoi Florou, Antigoni Katsoulidou, Spyros Sapounas, George Sourvinos, Spyridon Pournaras, Efthymia Petinaki, Mar...
work page 2022
-
[4]
Vítor Borges, Maria P. Duque, Joana V . Martins, Paula Vasconcelos, Rita Ferreira, Duarte Sobral, Ana Pelerito, Isabel L. de Carvalho, Marta S. Núncio, Maria J. Borrego, Cornelia Roemer, Richard A. Neher, Megan O’Driscoll, Raquel Rocha, Sofia Lopo, Ricardo Neves, Patricia Palminha, Liliana Coelho, Andreia Nunes, Joana Isidro, and João Paulo Gomes. Viral g...
work page 2022
-
[5]
Catarina Krug, Arnaud Tarantola, Emilie Chazelle, Erica Fougère, Annie Velter, Anne Guinard, Yvan Souares, Anna Mercier, Céline François, Katia Hamdad, Laetitia Tan-Lhernould, Anita Balestier, Hana Lahbib, Nicolas Etien, Pascale Bernillon, Virginie De Lauzun, Julien Durand, Myriam Fayad, Investigation Team, Henriette De Valk, François Beck, Didier Che, Br...
work page 2022
-
[6]
Edouard Mathieu, Fiona Spooner, Saloni Dattani, Hannah Ritchie, and Max Roser. Mpox. Our World in Data , 2022. https://ourworldindata.org/mpox
work page 2022
-
[7]
Shania J. R. D. Silva, Alain Kohl, Luis Pena, and Keith Pardee. Clinical and laboratory diagnosis of monkeypox (mpox): Current status and future directions. iScience, 26(6):106759, 2023
work page 2023
-
[8]
Barboza, Bijaya Kumar Padhi, and Ranjit Sah
Shriyansh Srivastava, Sachin Kumar, Shagun Jain, Aroop Mohanty, Neeraj Thapa, Prabhat Poudel, Krishna Bhusal, Zahraa Haleem Al-qaim, Joshuan J. Barboza, Bijaya Kumar Padhi, and Ranjit Sah. The global monkeypox (mpox) outbreak: A comprehensive review. Vaccines, 11(6), 2023
work page 2023
Show all 55 references
-
[9]
Strahan, Lloyd C
Surbhi Prasad, Claudia Galvan Casas, Andrew G. Strahan, Lloyd C. Fuller, Kathryn Peebles, Alessandra Carugno, Kenneth S. Leslie, Jen- nifer L. Harp, Tudor Pumnea, David E. McMahon, Michael Rosenbach, Jonathan E. Lubov, George Chen, Lynn P. Fox, Allison McMillen, Hyun W. Lim, A...
2022
-
[10]
Dermatologist-level classification of skin cancer with deep neural networks
Andre Esteva, Brett Kuprel, Roberto A Novoa, Justin Ko, Susan M Swetter, Helen M Blau, and Sebastian Thrun. Dermatologist-level classification of skin cancer with deep neural networks. nature, 542(7639):115–118, 2017
2017
-
[11]
Man against machine: diagnostic performance of a deep learning convolutional neural network for dermoscopic melanoma recognition in comparison to 58 dermatologists
Holger A Haenssle, Christine Fink, Roland Schneiderbauer, Ferdinand Toberer, Timo Buhl, Andreas Blum, Aadi Kalloo, A Ben Hadj Hassen, Luc Thomas, Alexander Enk, et al. Man against machine: diagnostic performance of a deep learning convolutional neural network for dermoscopic m...
2018
-
[12]
Systematic review of machine learning for diagnosis and prognosis in dermatology
Kenneth Thomsen, Lars Iversen, Therese Louise Titlestad, and Ole Winther. Systematic review of machine learning for diagnosis and prognosis in dermatology. Journal of Dermatological Treatment , 31(5):496–510, 2020
2020
-
[13]
Skin lesion classification from dermoscopic images using deep learning techniques
Adria Romero Lopez, Xavier Giro-i Nieto, Jack Burdick, and Oge Marques. Skin lesion classification from dermoscopic images using deep learning techniques. In 2017 13th IASTED International Conference on Biomedical Engineering (BioMed) , pages 49–54, 2017
2017
-
[14]
Chadaga, S
K. Chadaga, S. Prabhu, N. Sampathila, S. Nireshwalya, S. S. Katta, R. S. Tan, and U. R. Acharya. Application of artificial intelligence techniques for monkeypox: A systematic review. Diagnostics (Basel, Switzerland) , 13(5):824, 2023
2023
-
[15]
S. Asif, M. Zhao, Y . Li, et al. Ai-based approaches for the diagnosis of mpox: Challenges and future prospects. Arch Computational Methods in Engineering, 31:3585–3617, August 2024. Received: 28 October 2023; Accepted: 04 February 2024; Published: 26 March 2024
2024
-
[16]
Jaradat, R.E
A.S. Jaradat, R.E. Al Mamlook, N. Almakayeel, N. Alharbe, A.S. Almuflih, A. Nasayreh, H. Gharaibeh, M. Gharaibeh, A. Gharaibeh, and H. Bzizi. Automated monkeypox skin lesion detection using deep learning and transfer learning techniques. International Journal of Environmental ...
2023
-
[17]
A transfer learning and explainable solution to detect mpox from smartphones images
Mattia Giovanni Campana, Marco Colussi, Franca Delmastro, Sergio Mascetti, and Elena Pagani. A transfer learning and explainable solution to detect mpox from smartphones images. Pervasive and Mobile Computing, 98:101874, 2024
2024
-
[18]
Mehedi Hassan, Anupam Kumar Bairagi, and Sheikh Mohammed Shariful Islam
Avi Deb Raha, Mrityunjoy Gain, Rameswar Debnath, Apurba Adhikary, Yu Qiao, Md. Mehedi Hassan, Anupam Kumar Bairagi, and Sheikh Mohammed Shariful Islam. Attention to monkeypox: An interpretable monkeypox detection technique using attention mechanism. IEEE Access, 12:51942–51965, 2024
2024
-
[19]
Thieme, Y
A.H. Thieme, Y . Zheng, G. Machiraju, et al. A deep-learning algorithm to classify skin lesions from mpox virus infection. Nature Medicine, 29:738–747, March 2023. Received: 05 August 2022; Accepted: 19 January 2023; Published: 02 March 2023
2023
-
[20]
Efficientvit: Multi-scale linear attention for high-resolution dense prediction, 2024
Han Cai, Junyan Li, Muyan Hu, Chuang Gan, and Song Han. Efficientvit: Multi-scale linear attention for high-resolution dense prediction, 2024
2024
-
[21]
Monkeynet: A robust deep convolutional neural network for monkeypox disease detection and classification
Diponkor Bala, Md Shamim Hossain, Mohammad Alamgir Hossain, Md Ibrahim Abdullah, Md Mizanur Rahman, Balachandran Manavalan, Naijie Gu, Mohammad S Islam, and Zhangjin Huang. Monkeynet: A robust deep convolutional neural network for monkeypox disease detection and classification...
2023
-
[22]
Tazuddin Ahmed, Joydip Paul, Tasnim Jahan, S
Shams Nafisa Ali, Md. Tazuddin Ahmed, Joydip Paul, Tasnim Jahan, S. M. Sakeef Sani, Nawshaba Noor, and Taufiq Hasan. Monkeypox skin lesion detection using deep learning models: A preliminary feasibility study. arXiv preprint arXiv:2207.03342 , 2022
2022 arXiv
-
[23]
Tazuddin Ahmed, Tasnim Jahan, Joydip Paul, S
Shams Nafisa Ali, Md. Tazuddin Ahmed, Tasnim Jahan, Joydip Paul, S. M. Sakeef Sani, Nawshaba Noor, Anzirun Nahar Asma, and Taufiq Hasan. A web-based mpox skin lesion detection system using state-of- the-art deep learning models considering racial diversity. arXiv preprint arXi...
2023 arXiv
-
[24]
Going deeper with convolutions, 2014
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions, 2014
2014
-
[25]
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In F. Pereira, C.J. Burges, L. Bottou, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems , volume 25. Curran Associates, Inc., 2012
2012
-
[26]
Aggregated residual transformations for deep neural networks, 2017
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks, 2017
2017
-
[27]
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition, 2015
2015
-
[28]
Xception: Deep learning with depthwise separable convolutions, 2017
François Chollet. Xception: Deep learning with depthwise separable convolutions, 2017
2017
-
[29]
Mingxing Tan and Quoc V . Le. Efficientnet: Rethinking model scaling for convolutional neural networks, 2020
2020
-
[30]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023
2023
-
[31]
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weis- senborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[32]
Do vision transformers see like convolutional neural networks?, 2022
Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. Do vision transformers see like convolutional neural networks?, 2022
2022
-
[33]
How do vision transformers work?, 2022
Namuk Park and Songkuk Kim. How do vision transformers work?, 2022
2022
-
[34]
When vision transformers outperform resnets without pre-training or strong data augmentations, 2022
Xiangning Chen, Cho-Jui Hsieh, and Boqing Gong. When vision transformers outperform resnets without pre-training or strong data augmentations, 2022
2022
-
[35]
Cmt: Convolutional neural networks meet vision transformers, 2022
Jianyuan Guo, Kai Han, Han Wu, Yehui Tang, Xinghao Chen, Yunhe Wang, and Chang Xu. Cmt: Convolutional neural networks meet vision transformers, 2022
2022
-
[36]
A convnet for the 2020s, 2022
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s, 2022
2022
-
[37]
Levit: a vision transformer in convnet’s clothing for faster inference, 2021
Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze. Levit: a vision transformer in convnet’s clothing for faster inference, 2021
2021
-
[38]
Swin transformer: Hierarchical vision transformer using shifted windows, 2021
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows, 2021
2021
-
[39]
Le, and Mingxing Tan
Zihang Dai, Hanxiao Liu, Quoc V . Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes, 2021
2021
-
[40]
Shannon Wongvibulsin and Adewole S. Adamson. Deep learning for mpox: Advances, challenges, and opportunities. Med, 4(5):283–284, 2023
2023
-
[41]
Metaheuristics optimization-based ensemble of deep neural networks for mpox disease detection
Sohaib Asif, Ming Zhao, Fengxiao Tang, Yusen Zhu, and Baokang Zhao. Metaheuristics optimization-based ensemble of deep neural networks for mpox disease detection. Neural Networks, 167:342–359, 2023
2023
-
[42]
Mobilenetv2: Inverted residuals and linear bottlenecks, 2019
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks, 2019
2019
-
[43]
Batch normalization: Accelerating deep network training by reducing internal covariate shift, 2015
Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift, 2015
2015
-
[44]
Efficientvit: Memory efficient vision transformer with cascaded group attention, 2023
Xinyu Liu, Houwen Peng, Ningxin Zheng, Yuqing Yang, Han Hu, and Yixuan Yuan. Efficientvit: Memory efficient vision transformer with cascaded group attention, 2023
2023
-
[45]
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence , 40(4):834–...
2018
-
[46]
Rep- resentation degeneration problem in training natural language generation models, 2019
Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. Rep- resentation degeneration problem in training natural language generation models, 2019
2019
-
[47]
Going deeper with image transformers, 2021
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. Going deeper with image transformers, 2021
2021
-
[48]
Deepvit: Towards deeper vision transformer, 2021
Daquan Zhou, Bingyi Kang, Xiaojie Jin, Linjie Yang, Xiaochen Lian, Zihang Jiang, Qibin Hou, and Jiashi Feng. Deepvit: Towards deeper vision transformer, 2021
2021
-
[49]
Refiner: Refining self- attention for vision transformers, 2021
Daquan Zhou, Yujun Shi, Bingyi Kang, Weihao Yu, Zihang Jiang, Yuan Li, Xiaojie Jin, Qibin Hou, and Jiashi Feng. Refiner: Refining self- attention for vision transformers, 2021
2021
-
[50]
The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy, 2022
Tianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah, and Zhangyang Wang. The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy, 2022
2022
-
[51]
Le, and Hartwig Adam
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V . Le, and Hartwig Adam. Searching for mobilenetv3, 2019
2019
-
[52]
Deit iii: Revenge of the vit, 2022
Hugo Touvron, Matthieu Cord, and Hervé Jégou. Deit iii: Revenge of the vit, 2022
2022
-
[53]
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fan...
2019
-
[54]
Pytorch image models
Ross Wightman. Pytorch image models. https://github.com/rwightman/ pytorch-image-models, 2019
2019
-
[55]
Focal loss for dense object detection, 2018
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection, 2018
2018
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.