REVIEW 3 major objections 5 minor 41 references
SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual Recognition
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A lightweight CNN with 3.8M parameters reaches 77.7% top-1 accuracy on ImageNet-1k and, as a detection backbone, lifts RT-DETR to 50.0% AP on COCO.
desk verdict A capable lightweight backbone whose reported gains are real but not yet isolated from resolution and training-recipe differences. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a two-module block. The spatial SMixer decomposes the input into fine-grained and global feature streams, then applies convolutions with different receptive fields (a 3x3 convolution and a dilated 3x3 convolution with dilation 2, equivalent to a 5x5 effective field) to capture multi-order spatial information. The wave-based CMixer reformulates each channel as a complex-valued wave, expressed via Euler's formula, and applies a superposition mechanism that amplifies channels aligned with the maximally activated channel while suppressing trivial ones; a complex weight modulates the waves, enabling dynamic context-aware channel aggregation. Together these modules replace the standard MLP-based mixer in modern CNN blocks, allowing a lower expansion ratio and better parameter efficiency.
What would settle it
Train SpaRTAN-T and the leading baselines (e.g., ConvNeXt-XT, MogaNet-XT, Swin-1G) under identical conditions with 224x224 input and the same 300-epoch recipe, then check whether SpaRTAN-T still matches or beats their published accuracy at equal or lower parameter counts. A failure to maintain the advantage would indicate that the reported gains stem from training details rather than the proposed modules.
Extended reading notes
Core claim
The paper claims that a carefully designed convolutional network can extract the middle-order spatial interactions that both CNNs and transformers tend to undervalue, and that doing so yields a superior accuracy-efficiency trade-off. The discovery is demonstrated by SpaRTAN, which uses stacked 3x3 convolutions with dilation to cover multiple receptive fields and a channel-aggregation module modeled on wave superposition. On ImageNet-1k, SpaRTAN-T reaches 77.7% at 256x256 input with 3.8M parameters, and SpaRTAN-XT reaches 74.4% with only 2.2M parameters. As the backbone for RT-DETR, SpaRTAN-T achieves 50.0% AP on COCO, exceeding the ResNet-34 variant by 1.2% while using 10M fewer parameters. The authors attribute these gains to the combination of the spatial SMixer and wave-based CMixer, which together improve parameter utilization and reduce channel-wise redundancy.
Load-bearing premise
The paper's headline results are only meaningful if the comparison against published baselines is fair; the main risk is that SpaRTAN's 256x256 input, 300-epoch training schedule, and strong augmentation recipe, rather than the architecture itself, are what drive the accuracy and parameter-efficiency advantages.
Editorial extensions
If this is right
- If the reported results are correct, a small-kernel convolutional architecture can match or beat attention-based and large-kernel transformers in both classification and detection at a fraction of the parameters and FLOPs.
- The wave-based channel aggregation suggests that modeling channels as waves with learnable complex weights is a viable alternative to high-expansion-ratio MLPs, reducing redundancy without sacrificing accuracy.
- The architecture's efficiency at low parameter counts could make it applicable to mobile and embedded settings where both accuracy and latency matter.
- The claimed mid-order interaction capture implies that such networks may be more robust to occlusion and background clutter, as suggested by the Grad-CAM visualizations.
Reading between the lines
- The paper's implicit claim that multi-order interactions are the cause of the accuracy gains is not directly evidenced; the gains could also stem from the specific training recipe, data augmentation, or the hybrid convolution strategy.
- A direct comparison with baselines trained under identical settings (same resolution, same epochs, same augmentation) would be a stronger test of whether the architecture itself, rather than the training setup, is responsible for the improvements.
- The wave-based CMixer's reliance on the maximally activated channel as a reference point could be sensitive to outlier channels; testing on adversarial or distorted inputs might reveal failure modes.
- If the middle-order interaction hypothesis is correct, SpaRTAN's design could be combined with other lightweight modules or formally analyzed to quantify interaction order, but the paper does not provide such a mechanism.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SpaRTAN, a lightweight CNN architecture composed of two modules: a spatial SMixer that combines small and large effective receptive fields via stacked convolutions with different dilations, and a wave-based CMixer that treats channels as real/imaginary wave components and aggregates them through a superposition mechanism guided by a maximally activated channel. The authors report 77.7% ImageNet-1k top-1 accuracy with 3.8M parameters and about 1.0 GFLOPs, and 50.0% COCO AP when used as an RT-DETR backbone with 21.5M parameters, claiming superior parameter efficiency over existing lightweight models. The paper includes architectural ablations, Grad-CAM visualizations, and a public code release.
Significance. If the accuracy and efficiency claims survive controlled comparison, SpaRTAN would be a genuinely parameter-efficient convolutional design, especially because it avoids large kernels and high MLP expansion ratios. The paper's strengths include a clear engineering motivation, internal ablation studies that support the module-level choices (Table IV through Table VIII), a public code release, and evaluation on two standard benchmarks. However, the central quantitative claims currently rest on comparisons against baselines trained at different input resolutions and with different training recipes; this fairness gap is the main barrier to accepting the architecture-level conclusions.
major comments (3)
- [Section IV-A.2, Table II, Abstract] The headline result of 77.7% top-1 accuracy with 3.8M parameters and "approximately 1.0 GFLOPs" corresponds to SpaRTAN-T evaluated at 256x256 input (1.08 GFLOPs), whereas most baselines in Table II are reported at 224x224 input. At 224x224 the same model achieves 77.1% with 0.83 GFLOPs. The abstract's presentation is therefore misleading without specifying the input resolution. More importantly, the comparison is not controlled for training recipe: SpaRTAN is trained for 300 epochs with RandAugment, Mixup, CutMix, Random Erasing, and Stochastic Depth, while many cited baselines (e.g., ResNet-18, PVT-T, ConvNeXt-XT) use shorter or different schedules. Since the margins over MogaNet-XT and ConvNeXt-XT are 0.5 and 0.2 points respectively, these unstated differences could account for the advantage. Please either retrain key baselines under the same recipe, or substantially qualify the comparison claims.
- [Section IV-B, Table III] The COCO comparison compares SpaRTAN backbones pretrained on ImageNet for 300 epochs with a strong augmentation recipe against RT-DETR ResNet-18/34 baselines that use standard (much shorter) ImageNet pretraining as published in [38]. The detector training is matched, but the backbone pretraining is not. The reported 1.2 AP improvement over ResNet-34 and 2.1 AP over ResNet-18 are therefore confounded by pretraining schedule and augmentation. Please retrain the ResNet backbones with the same 300-epoch recipe, or present a controlled comparison where only the architecture differs. Without this, the claim that the backbone architecture is responsible for the detection gains is not established.
- [Section III-C, Equations (7)-(9)] The wave-based CMixer is not specified at the level needed for reproduction. The text explains the conceptual wave interpretation, but the actual forward computation of W(·) is left underspecified: the definition of Fmax as "the maximum value of C" is ambiguous (presumably a channel-wise max across the channel dimension?), the split of channels into sine and cosine halves is not formalized, the role of the complex weight Wc in the actual tensor operations is not written out, and the promised "linear approximation of sinusoidal waves using a point convolution with a non-linear activation function" is not described concretely. The paper should provide a precise layer-by-layer definition of the wave-based aggregation module, or a pseudocode block, to make the contribution self-contained despite the public code.
minor comments (5)
- [Abstract] The text contains a typographical spacing error in "77. 7%" which should be "77.7%".
- [Table III] The column headers "APval APval 50 APval 75 APval S APval M APval L" are unclear due to missing subscripts and spacing; please reformat them as AP, AP50, AP75, APS, APM, APL.
- [Section IV-A.1] The augmentation "Random Resized Crop" should be typeset consistently with the standard name "RandomResizedCrop" to avoid ambiguity.
- [Table V] The table reports kernel sizes 3 and 5, but the text refers to replacing a 5x5 dilation-2 convolution with stacked 3x3 dilation-2 convolutions; please clarify that the "5" entry denotes the single 5x5 kernel and the "3" entry denotes the stacked 3x3 replacement.
- [Section III-B, Equation (6)] The feature maps FH and FL are introduced in the text as outputs of the high- and low-frequency branches, but they are not formally defined before appearing in Equation (6); please add explicit definitions.
Circularity Check
No significant circularity: SpaRTAN is an empirical architecture paper whose claims are validated by held-out ImageNet/COCO benchmarks, not by construction or self-citation.
full rationale
This is an empirical architecture paper. The central claims — 77.7% top-1 on ImageNet-1k and 50.0% AP on COCO — are measured results on held-out benchmarks, not quantities derived from fitted constants or from assumptions that already contain the conclusions. No parameter is fitted to a subset of benchmark data and then reported as a prediction of a closely related quantity. The wave-based CMixer formulation in Eqs. (8)-(10) is an architectural design inspired by external prior work (Wave-MLP [21]), and the paper makes no uniqueness claim and imports no theorem from the authors' own prior work. The authors do not cite themselves in any load-bearing way; all cited methods (MogaNet, HorNet, Wave-MLP, RT-DETR, etc.) are independent external works. The internal ablations in Tables IV-VIII support component efficacy but do not generate or rename the headline results. The potential concern that Table II mixes 224x224 and 256x256 comparisons, and that the 300-epoch training recipe differs from some published baselines, is a question of experimental fairness and benchmark control, not circularity under the stated criteria. Because the derivation chain is self-contained and benchmark-driven, the appropriate score is 0.
Assumptions & free parameters
free parameters (6)
- Stage channels and block counts =
32/64/96/192 channels and 3/3/10/2 blocks for XT; 32/64/128/256 channels and 3/3/12/2 blocks for T
- Expand ratio r for channel mixing =
4 for stages 1-2, 2 for stages 3-4
- SMixer kernel configuration =
3x3 dilation 1 branch and stacked 3x3 dilation 2 branch
- Convolution type schedule =
full convolution in stages 1-2, depthwise in stages 3-4
- Activation and normalization pairings =
SiLU in patch embedding, GELU in blocks; BatchNorm after convolutions, LayerNorm before SMixer and CMixer
- Training recipe =
300 epochs, AdamW, batch 2048, base LR 2.5e-3, weight decay 0.03, warmup 20 epochs, RandAugment, Mixup, CutMix, Random…
assumptions (5)
- standard math Euler's formula allows a real feature channel to be written as a complex wave and split into sin and cos components with half the channels each.
- standard math Multiplication in the frequency domain is equivalent to global circular convolution in the spatial domain, so complex weight modulation captures both short- and long-range channel interactions.
- domain assumption The game-theoretic 'middle-order interaction' account of CNN limitations is a valid description of visual concept learning.
- ad hoc to paper Approximating sinusoidal waves with a pointwise convolution plus nonlinearity preserves the benefits of wave modulation while stabilizing training.
- domain assumption ImageNet top-1 accuracy and COCO AP are sufficient proxies for the general quality of a visual backbone.
invented entities (1)
-
Maximally activated channel wave (F_max) as a reference for channel superposition
Cite this review
Pith. "Pith review of SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual Recognition." pith.science (2026). https://pith.science/paper/2R5SVBEI
@misc{pith2026250710999,
author = {Pith},
title = {Pith review of: SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/2R5SVBEI}},
note = {Machine review of arXiv:2507.10999}
}
read the original abstract
The resurgence of convolutional neural networks (CNNs) in visual recognition tasks, exemplified by ConvNeXt, has demonstrated their capability to rival transformer-based architectures through advanced training methodologies and ViT-inspired design principles. However, both CNNs and transformers exhibit a simplicity bias, favoring straightforward features over complex structural representations. Furthermore, modern CNNs often integrate MLP-like blocks akin to those in transformers, but these blocks suffer from significant information redundancies, necessitating high expansion ratios to sustain competitive performance. To address these limitations, we propose SpaRTAN, a lightweight architectural design that enhances spatial and channel-wise information processing. SpaRTAN employs kernels with varying receptive fields, controlled by kernel size and dilation factor, to capture discriminative multi-order spatial features effectively. A wave-based channel aggregation module further modulates and reinforces pixel interactions, mitigating channel-wise redundancies. Combining the two modules, the proposed network can efficiently gather and dynamically contextualize discriminative features. Experimental results in ImageNet and COCO demonstrate that SpaRTAN achieves remarkable parameter efficiency while maintaining competitive performance. In particular, on the ImageNet-1k benchmark, SpaRTAN achieves 77. 7% accuracy with only 3.8M parameters and approximately 1.0 GFLOPs, demonstrating its ability to deliver strong performance through an efficient design. On the COCO benchmark, it achieves 50.0% AP, surpassing the previous benchmark by 1.2% with only 21.5M parameters. The code is publicly available at [https://github.com/henry-pay/SpaRTAN].
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[38]
Detrs beat yolos on real-time object detection,
Y . Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y . Liu, and J. Chen, “Detrs beat yolos on real-time object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 16 965–16 974
work page 2024
-
[1]
Imagenet classification with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural informa- tion processing systems , vol. 25, 2012
2012
-
[2]
Very deep convolutional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in 3rd International Conference on Learning Representations (ICLR 2015) . Computational and Biological Learning Society, 2015, pp. 1–14
work page 2015
-
[3]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
-
[4]
Efficientnet: Rethinking model scaling for con- volutional neural networks,
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for con- volutional neural networks,” in International conference on machine learning. PMLR, 2019, pp. 6105–6114
2019
-
[5]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” ICLR, 2021
2021
-
[6]
Imagenet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
-
[7]
Microsoft coco: Common objects in context,
T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 . Springer, 2014, pp. 740–755
2014
Show all 41 references
-
[8]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V . Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garn...
2017
-
[9]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10 012–10 022
2021
-
[10]
Training data-efficient image transformers & distillation through attention,
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. J ´egou, “Training data-efficient image transformers & distillation through attention,” in International conference on machine learning . PMLR, 2021, pp. 10 347–10 357
2021
-
[11]
Rethinking attention with performers,
K. M. Choromanski, V . Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller, “Rethinking attention with performers,” in International Conference on Learning Representations , 2021
2021
-
[12]
Video swin transformer,
Z. Liu, J. Ning, Y . Cao, Y . Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 3202–3211
2022
-
[13]
A convnet for the 2020s,
Z. Liu, H. Mao, C.-Y . Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2022, pp. 11 976–11 986
2022
-
[14]
Scaling up your kernels to 31x31: Revisiting large kernel design in cnns,
X. Ding, X. Zhang, J. Han, and G. Ding, “Scaling up your kernels to 31x31: Revisiting large kernel design in cnns,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 11 963–11 975
2022
-
[15]
More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity,
S. Liu, T. Chen, X. Chen, X. Chen, Q. Xiao, B. Wu, T. K ¨arkk¨ainen, M. Pechenizkiy, D. C. Mocanu, and Z. Wang, “More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity,” in The Eleventh International Conference on Learning Representations , 2023
2023
-
[16]
Conv2former: A simple transformer-style convnet for visual recognition,
Q. Hou, C.-Z. Lu, M.-M. Cheng, and J. Feng, “Conv2former: A simple transformer-style convnet for visual recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
-
[17]
Hornet: Efficient high-order spatial interactions with recursive gated convolu- tions,
Y . Rao, W. Zhao, Y . Tang, J. Zhou, S. N. Lim, and J. Lu, “Hornet: Efficient high-order spatial interactions with recursive gated convolu- tions,” Advances in Neural Information Processing Systems , vol. 35, pp. 10 353–10 366, 2022
2022
-
[18]
Moganet: Multi-order gated aggregation network,
S. Li, Z. Wang, Z. Liu, C. Tan, H. Lin, D. Wu, Z. Chen, J. Zheng, and S. Z. Li, “Moganet: Multi-order gated aggregation network,” in The Twelfth International Conference on Learning Representations , 2023
2023
-
[19]
Mlp-mixer: An all-mlp architecture for vision,
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Un- terthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreitet al., “Mlp-mixer: An all-mlp architecture for vision,” Advances in neural information processing systems, vol. 34, pp. 24 261–24 272, 2021
2021
-
[20]
Resmlp: Feed- forward networks for image classification with data-efficient training,
H. Touvron, P. Bojanowski, M. Caron, M. Cord, A. El-Nouby, E. Grave, G. Izacard, A. Joulin, G. Synnaeve, J. Verbeek et al. , “Resmlp: Feed- forward networks for image classification with data-efficient training,” IEEE transactions on pattern analysis and machine intelligence, ...
2022
-
[21]
An image patch is a wave: Phase-aware vision mlp,
Y . Tang, K. Han, J. Guo, C. Xu, Y . Li, C. Xu, and Y . Wang, “An image patch is a wave: Phase-aware vision mlp,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 935–10 944
2022
-
[22]
A game-theoretic taxonomy of visual concepts in dnns,
X. Cheng, C. Chu, Y . Zheng, J. Ren, and Q. Zhang, “A game-theoretic taxonomy of visual concepts in dnns,” arXiv preprint arXiv:2106.10938, 2021
2021 arXiv
-
[23]
Discovering and explaining the representation bottleneck of dnns,
H. Deng, Q. Ren, H. Zhang, and Q. Zhang, “Discovering and explaining the representation bottleneck of dnns,” in International Conference on Learning Representations, 2022
2022
-
[24]
Towards a unified game- theoretic view of adversarial perturbations and robustness,
J. Ren, D. Zhang, Y . Wang, L. Chen, Z. Zhou, Y . Chen, X. Cheng, X. Wang, M. Zhou, J. Shi, and Q. Zhang, “Towards a unified game- theoretic view of adversarial perturbations and robustness,” in Advances in Neural Information Processing Systems, A. Beygelzimer, Y . Dauphin, P....
2021
-
[25]
An impartial take to the cnn vs transformer robustness contest,
F. Pinto, P. H. Torr, and P. K. Dokania, “An impartial take to the cnn vs transformer robustness contest,” in European Conference on Computer Vision. Springer, 2022, pp. 466–480
2022
-
[26]
Efficientformer: Vision transformers at mobilenet speed,
Y . Li, G. Yuan, Y . Wen, J. Hu, G. Evangelidis, S. Tulyakov, Y . Wang, and J. Ren, “Efficientformer: Vision transformers at mobilenet speed,” Ad- vances in Neural Information Processing Systems , vol. 35, pp. 12 934– 12 949, 2022
2022
-
[27]
Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer,
S. Mehta and M. Rastegari, “Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer,” in International Conference on Learning Representations, 2022
2022
-
[28]
Mobilenetv2: Inverted residuals and linear bottlenecks,
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 4510–4520
2018
-
[29]
Shufflevitnet: Mobile-friendly vision transformer with less-memory,
X. Zhao and J. Lu, “Shufflevitnet: Mobile-friendly vision transformer with less-memory,” in 2024 International Joint Conference on Neural Networks (IJCNN). IEEE, 2024, pp. 1–7
2024
-
[30]
Visual attention network,
M.-H. Guo, C.-Z. Lu, Z.-N. Liu, M.-M. Cheng, and S.-M. Hu, “Visual attention network,” Computational Visual Media, vol. 9, no. 4, pp. 733– 752, 2023
2023
-
[31]
Squeeze-and-excitation networks,
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141
2018
-
[32]
Global filter networks for image classification,
Y . Rao, W. Zhao, Z. Zhu, J. Lu, and J. Zhou, “Global filter networks for image classification,” Advances in neural information processing systems, vol. 34, pp. 980–993, 2021
2021
-
[33]
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 568–578
2021
-
[34]
Early convolutions help transformers see better,
T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Doll ´ar, and R. Girshick, “Early convolutions help transformers see better,” Advances in neural information processing systems , vol. 34, pp. 30 392–30 400, 2021
2021
-
[35]
Rethinking vision transformers for mobilenet size and speed,
Y . Li, J. Hu, Y . Wen, G. Evangelidis, K. Salahi, Y . Wang, S. Tulyakov, and J. Ren, “Rethinking vision transformers for mobilenet size and speed,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 16 889–16 900
2023
-
[36]
Mobileone: An improved one millisecond mobile backbone,
P. K. A. Vasu, J. Gabriel, J. Zhu, O. Tuzel, and A. Ranjan, “Mobileone: An improved one millisecond mobile backbone,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 7907–7917
2023
-
[37]
Deformable {detr}: Deformable transformers for end-to-end object detection,
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable {detr}: Deformable transformers for end-to-end object detection,” in International Conference on Learning Representations , 2021
2021
-
[39]
DAB-DETR: Dynamic anchor boxes are better queries for DETR,
S. Liu, F. Li, H. Zhang, X. Yang, X. Qi, H. Su, J. Zhu, and L. Zhang, “DAB-DETR: Dynamic anchor boxes are better queries for DETR,” in International Conference on Learning Representations , 2022
2022
-
[40]
Dn-detr: Accelerate detr training by introducing query denoising,
F. Li, H. Zhang, S. Liu, J. Guo, L. M. Ni, and L. Zhang, “Dn-detr: Accelerate detr training by introducing query denoising,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 13 619–13 627
2022
-
[41]
Grad-cam: Visual explanations from deep networks via gradient-based localization,
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 618–626
2017
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.