Pith. sign in

REVIEW 2 cited by

ConTNet: Why not use convolution and transformer at the same time?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.13497 v3 pith:FERXPWGX submitted 2021-04-27 cs.CV

classification cs.CV
keywords contnettasksaugmentationsbackboneconvnetsdatadatasetresnet
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although convolutional networks (ConvNets) have enjoyed great success in computer vision (CV), it suffers from capturing global information crucial to dense prediction tasks such as object detection and segmentation. In this work, we innovatively propose ConTNet (ConvolutionTransformer Network), combining transformer with ConvNet architectures to provide large receptive fields. Unlike the recently-proposed transformer-based models (e.g., ViT, DeiT) that are sensitive to hyper-parameters and extremely dependent on a pile of data augmentations when trained from scratch on a midsize dataset (e.g., ImageNet1k), ConTNet can be optimized like normal ConvNets (e.g., ResNet) and preserve an outstanding robustness. It is also worth pointing that, given identical strong data augmentations, the performance improvement of ConTNet is more remarkable than that of ResNet. We present its superiority and effectiveness on image classification and downstream tasks. For example, our ConTNet achieves 81.8% top-1 accuracy on ImageNet which is the same as DeiT-B with less than 40% computational complexity. ConTNet-M also outperforms ResNet50 as the backbone of both Faster-RCNN (by 2.6%) and Mask-RCNN (by 3.2%) on COCO2017 dataset. We hope that ConTNet could serve as a useful backbone for CV tasks and bring new ideas for model design

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ProtoOcc: Accurate, Efficient 3D Occupancy Prediction Using Dual Branch Encoder-Prototype Query Decoder

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ProtoOcc achieves 39.56% mIoU single-frame and 45.02% mIoU multi-frame on Occ3D-nuScenes using a dual-branch encoder and a prototype query decoder that skips iterative decoding.

  2. Optimizing Local-Global Dependencies for Accurate 3D Human Pose Estimation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A hybrid local-global network using skeleton selective refine attention reports state-of-the-art 3D pose accuracy on Human3.6M and MPI-INF-3DHP.

Pith tools