Aggregated Residual Transformations for Deep Neural Networks

Kaiming He; Piotr Doll\'ar; Ross Girshick; Saining Xie; Zhuowen Tu

arxiv: 1611.05431 · v2 · pith:FFR3VKHOnew · submitted 2016-11-16 · 💻 cs.CV

Aggregated Residual Transformations for Deep Neural Networks

Saining Xie , Ross Girshick , Piotr Doll\'ar , Zhuowen Tu , Kaiming He This is my paper

classification 💻 cs.CV

keywords cardinalityclassificationtransformationsarchitectureincreasingmodelsnetworkresnext

0 comments

read the original abstract

We present a simple, highly modularized network architecture for image classification. Our network is constructed by repeating a building block that aggregates a set of transformations with the same topology. Our simple design results in a homogeneous, multi-branch architecture that has only a few hyper-parameters to set. This strategy exposes a new dimension, which we call "cardinality" (the size of the set of transformations), as an essential factor in addition to the dimensions of depth and width. On the ImageNet-1K dataset, we empirically show that even under the restricted condition of maintaining complexity, increasing cardinality is able to improve classification accuracy. Moreover, increasing cardinality is more effective than going deeper or wider when we increase the capacity. Our models, named ResNeXt, are the foundations of our entry to the ILSVRC 2016 classification task in which we secured 2nd place. We further investigate ResNeXt on an ImageNet-5K set and the COCO detection set, also showing better results than its ResNet counterpart. The code and models are publicly available online.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

NASTaR: NovaSAR Automated Ship Target Recognition Dataset
cs.CV 2025-12 accept novelty 7.0

NASTaR is a new dataset of 3415 AIS-labeled ship patches from NovaSAR S-band SAR imagery with 23 classes, inshore/offshore splits, and wake annotations, validated via benchmark deep learning models.
Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025
cs.CV 2023-05 unverdicted novelty 2.0

The paper summarizes results from the SurgToolLoc and SurgVU challenges held at MICCAI conferences from 2022 to 2025.